Multi-objective optimization scheduling method and system for tug
By combining multidimensional real-number encoding and Q-learning algorithm with the global-local collaborative optimization model of the Jaya algorithm, the problems of local optima and multi-objective collaborative optimization in tugboat scheduling are solved, achieving efficient, economical and environmentally friendly multi-objective optimization results in tugboat scheduling.
Patent Information
- Application Number
- CN202511026456.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-18
AI Technical Summary
Existing metaheuristic algorithms are prone to getting stuck in local optima in tugboat scheduling, making it difficult to effectively explore and develop the solution space in large-scale task scenarios. Furthermore, they are difficult to coordinate and optimize economic efficiency, timeliness, and environmental benefits, and thus cannot meet the multi-objective requirements of modern ports.
A multidimensional real-number encoding method is used to represent the tugboat allocation scheme. By combining the Q-learning algorithm and the Jaya algorithm, a global-local collaborative optimization solution model is constructed. Through the dynamic decision-making capability of the Q-learning algorithm and the elite solution-oriented characteristic of the Jaya algorithm, the tugboat scheduling decision framework is optimized and the optimal scheduling scheme is output.
It achieves multi-objective collaborative optimization of tugboat scheduling in large-scale task scenarios, reduces operating costs, improves scheduling efficiency, and meets the needs of modern ports for economy, timeliness and environmental benefits.
Smart Images

Figure CN120975446A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of port scheduling, and particularly relates to a tug multi-objective optimization scheduling method and system. BACKGROUND
[0002] In the modern port operation system, as the core guarantee resource for ensuring the safety of ships to berth and unberth, the scheduling efficiency of the tug directly affects the overall operation efficiency of the port. The traditional scheduling mode relying on manual experience faces significant challenges: on the one hand, it is difficult to optimize the tug operation cost, operation time and service reliability in a complex port operation environment; on the other hand, it is even more difficult to effectively respond to the increasingly urgent demand for green and low-carbon development of the port, and to achieve precise control and reduction of carbon emissions. Therefore, breaking through the limitations of traditional manual scheduling and introducing and applying intelligent optimization algorithms have become a key breakthrough and inevitable choice to improve the scheduling efficiency of the tug, reduce the operation cost and promote the construction of green ports.
[0003] At present, the intelligent optimization algorithms used in the scheduling of the tug mainly include meta-heuristic algorithms such as particle swarm optimization (PSO), ant colony optimization (ACO), artificial bee colony optimization (ABC), genetic algorithm (GA), grey wolf optimization (GWO) and Jaya algorithm. Although these intelligent optimization algorithms improve the scheduling efficiency of the tug and reduce the operation cost, they focus on single objective optimization and are difficult to meet the demand for the collaborative optimization of economy, timeliness and environmental benefits in modern ports. In addition, in the large-scale task scenario of the tug scheduling, the meta-heuristic algorithm is prone to fall into local optimum, and there are still deficiencies in the balance mechanism of solution space exploration and development capability, and there are still obvious limitations in dynamic constraint processing and multi-objective collaborative optimization. SUMMARY
[0004] In view of the problems in the prior art that the meta-heuristic algorithm is prone to fall into local optimum, there are still deficiencies in the balance mechanism of solution space exploration and development capability, and there are still obvious limitations in dynamic constraint processing and multi-objective collaborative optimization, the present application provides a tug multi-objective optimization scheduling method and system.
[0005] To solve the above technical problems, the present application provides the following technical solutions: A tug multi-objective optimization scheduling method, comprising the following steps: S1, obtaining port tug base data, active tug data and all berthing and unberthing ship data within a period; S2, establishing and embedding a mathematical model of tug multi-objective collaborative optimization scheduling based on the obtained port tug base data, active tug data and all berthing and unberthing ship data within a period, and embedding the core business constraints of port operation; S3, using multi-dimensional real number coding to represent the tug allocation scheme; setting the state space and discrete action space based on the tug allocation scheme; S4, adopting a Q-learning algorithm to construct a tugboat scheduling multi-objective optimization decision framework, and updating a state space and a discrete action space of a current tugboat allocation scheme based on the tugboat scheduling multi-objective optimization decision framework; S5, coupling an elite solution guiding characteristic of a Jaya algorithm with a dynamic decision-making capability of Q-learning to construct a global-local collaborative optimization solving model; S6, solving a mathematical model of tugboat multi-objective collaborative optimization scheduling based on the global-local collaborative optimization solving model, and outputting an optimal scheduling scheme.
[0006] Further, the port tugboat base data includes: the number, location, and capacity of the tugboat base; the active tugboat data includes: the number, horsepower, average speed, towage operation cost, non-towage operation cost, and the tugboat base to which the tugboat belongs; and the berthing and unberthing ship data includes: the ship arrival time, starting position, target position, required number of tugboats, and tugboat horsepower requirement.
[0007] Further, the step S1 includes: S11, classifying the active tugboats in the port based on the horsepower of the tugboat base and the tugboat, to obtain classified tugboats; S12, classifying the ships based on the ship length, the required number and type of tugboats under the light load and heavy load of the ship, to obtain classified ships; S13, establishing a matching rule table between the classified tugboats and the classified ships.
[0008] Further, the mathematical model of tugboat multi-objective collaborative optimization scheduling in the step S2 includes a multi-objective function and a constraint condition; The multi-objective function is: (1) (2) (3) (4) wherein, is the total operation cost of the tugboat, is the total operation time of the tugboat, is the operation cost of the tugboat serving a single ship, is the operation time of the tugboat serving a single ship, is the number of the tugboat and , is the number of the task and , is the number of the tugboat base and , is the number of the working task of the tugboat , Cp is the unit cost of operation per nautical mile for the tugboat in non-towing condition, Ct is the unit cost of operation per nautical mile for the tugboat in towing condition, Cp is the unit cost of operation per nautical mile for the tugboat in non-towing condition, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, , D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, D is the distance from the base to the mission starting point, The above established multi-objective function is normalized into a single objective function by using linear weighting method:
[0009] where Cmin is the minimum total operation cost of the tugboat, Cmin is the minimum total operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, Cmin and Cmin are the optimal values of the single objective optimization of the model for the minimum operation cost and minimum operation time of the tugboat, (10) (11) (12) (13) (14) (15) (16) wherein, is the tugboat number and , is the task number and , is the tugboat base number and , is the number of tasks the tugboat works and , is the number of tugboats needed for the task , is the horsepower of the tugboat , is the horsepower of the tugboat needed for the task , is the distance from the base to the start of the task , is the start time of the task by the tugboat , is the end time of the task by the tugboat , is the time of departure from the base for the first task by the tugboat , is the average speed of the tugboat when not towing, is the decision variable indicating the assignment of the tugboat to the task , is the decision variable indicating the return of the tugboat to the base after completing the nth task, i.e., the task ,
[0010] Further, the state space in the step S3 comprises:
[0011] wherein S is the state, is the current value of the normalized objective function, to normalize the objective function, to a fine-tuning stage of the optimal solution, to a medium-term optimization stage, to an initial optimization stage. The discrete action space includes: action keeping the tugboat allocation of the current candidate solution unchanged, action randomly selecting a tugboat allocation of the task to replace the allocation of the corresponding task in the global optimal solution, action re-randomly generating a tugboat combination that satisfies the quantity and type constraints for the selected task.
[0012] Further, the step S3 of representing the tugboat allocation scheme by using multi-dimensional real number coding includes: using multi-dimensional real number coding, the tugboat allocation scheme is represented by The ships are numbered in chronological order, and the tugboats are allocated in priority according to the principle of first come, first served.
[0013] Further, the step S4 includes: Step S41, based on the set state space, the relative quality difference between the current solution and the global optimal solution is calculated for state evaluation, and then the corresponding discrete action space is selected by ε-greedy selection; Step S42, based on the Q-learning algorithm, the Q value associated with the state space and the discrete action space is calculated, and the action with the highest Q value is selected; after each iteration, the Q value is updated, and the updated Q value formula is as follows:
[0014] wherein Q (S, A) is the Q value of action A taken in state S, is the learning rate, is the discount factor, and r is the actual reward received by action A, is the highest expected reward of all possible operations in state S; Step S43, based on the two objectives of minimizing the total tugboat operation cost and minimizing the total tugboat operation time, a reward function is set to calculate the feedback reward value, and the reward function is as follows:
[0015] wherein is the reward function, is the difference between the optimal solution of minimizing the total tugboat operation cost and the current solution, is the difference between the optimal solution of minimizing the total tugboat operation time and the current solution, is a reference value to eliminate dimensional differences, and are the corresponding weight factors of the two objectives, Take a value between [0, 1].
[0016] Further, the global-local collaborative optimization solution model comprises a global search layer and a local optimization layer; step S5 comprises: Step S51, the global search layer adopts Jaya algorithm to iteratively update the tug scheduling population, performs coarse-grained exploration in the solution space, and guides the population to move to the potential area by comparing the difference vector of the current tug scheduling optimal solution and the worst solution; wherein the solution update formula of Jaya algorithm is as follows:
[0017] Wherein , is a random number in , is the updated tug allocation scheme, is the current tug allocation scheme, , are the current optimal and worst tug allocation schemes respectively; Step S52, the local optimization layer embeds Q learning algorithm, and realizes fine-grained development through designing state-action-reward mapping mechanism; Step S53, the knowledge sharing mechanism establishes dynamic association between the global optimal solution and the Q value, directly refers to the allocation mode of the global optimal solution when the Q learning algorithm executes the action, and realizes cross-level migration of experience knowledge.
[0018] Further, the step S6 comprises the following sub-steps: Step S61, the maximum number of tugs required by a single ship task is the lateral length of the individual, the number of ship tasks at the present stage is the longitudinal length of the individual, for each to-be-generated individual, under the premise of meeting the specification constraints, a corresponding number of tug units are randomly selected from the available tug set to ensure that each initial individual constitutes a feasible solution; Step S62, input the obtained port tug base data, active tug data and all berthing and unberthing ship task data within the period; Step S63, a certain number and horsepower of tugs are matched for a single ship task, and according to the task sorting and the constraint conditions of the model, the set of currently available tugs is obtained, and finally the initial candidate solution set is randomly generated according to the currently available tug set; Step S64, the global-local collaborative optimization solution model is used to iteratively update through Jaya algorithm, quickly locate the high-quality solution area, and then the Q learning algorithm is used to locally optimize each candidate solution, and the output optimal solution is the optimal tug scheduling scheme of the ship task.
[0019] On the other hand, the present application provides a tug multi-objective optimization scheduling system, comprising: The data acquisition module is used for acquiring port tug base data, active tug data and all berthing and unberthing ship data in a period; The mathematical model construction module is used for establishing and embedding a mathematical model of tug multi-objective collaborative optimization scheduling based on the acquired port tug base data, active tug data and all berthing and unberthing ship data in a period and the core business constraints of port operation; The encoding module is used for representing a tug allocation scheme by using multi-dimensional real number encoding; and setting a state space and a discrete action space based on the tug allocation scheme; The multi-objective optimization decision framework construction module is used for constructing a tug scheduling multi-objective optimization decision framework by using a Q learning algorithm, and updating the state space and the discrete action space of the current tug allocation scheme based on the tug scheduling multi-objective optimization decision framework. The solving model construction module is used for coupling the elite solution-oriented characteristics of the Jaya algorithm and the dynamic decision-making ability of Q learning, and constructing a global-local collaborative optimization solving model; The scheduling scheme output module is used for solving the mathematical model of tug multi-objective collaborative optimization scheduling based on the global-local collaborative optimization solving model, and outputting an optimal scheduling scheme.
[0020] Compared with the prior art, the present application has the following beneficial effects: The tug multi-objective optimization scheduling method fully considers the problem that the traditional tug scheduling is focused on single target optimization and the algorithm is easy to fall into local optimum in a large-scale task scenario, constructs a mathematical model of minimum total operation cost and total operation time of tugs, and obtains an optimal scheduling scheme of port tugs by using the Jaya algorithm fused with Q learning. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0022] Figure 1 The tug multi-objective optimization scheduling flowchart of the present application.
[0023] Figure 2 The Q learning framework designed by the present application.
[0024] Figure 3 The Jaya algorithm fused with Q learning is used to solve the tug scheduling flowchart of the present application.
[0025] Figure 4This is a comparison chart of the optimal value convergence of the Jaya-QL algorithm of this invention with other algorithms.
[0026] Figure 5 This is a Gantt chart of the optimal scheduling scheme of this invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0028] Example 1 The implementation of the present invention will now be further described with reference to the accompanying drawings, such as... Figure 1 As shown, this invention provides a tugboat multi-objective optimization scheduling method that integrates Q-learning and the Jaya algorithm. The method includes the following steps: S1. Obtain port tugboat base data, active tugboat data, and data on all berthing and departing vessels within the period. The tugboat base data includes the number, location, and capacity of tugboat bases. The active tugboat data includes the number of tugboats, horsepower, average speed, towing operation cost, non-towing operation cost, and the tugboat base to which they belong. The vessel data includes the vessel's arrival time, starting position, target position, required number of tugboats, and tugboat horsepower requirements.
[0029] S2. Based on the acquired port tugboat base data, active tugboat data, and all vessel data within the period, establish a mathematical model for multi-objective collaborative optimization scheduling of tugboats with the dual objective functions of minimizing the total tugboat operation cost and minimizing the total tugboat operation time, and embedding core business constraints of port operations. The core business constraints are used to ensure the feasibility of the scheduling scheme and the multi-objective collaborative optimization capability.
[0030] S3. Multidimensional real number encoding is used to describe the tugboat allocation problem. Initial individuals and population are set, and the state space and discrete action space are set based on the tugboat allocation scheme. S4. A multi-objective optimization decision-making framework for tugboat scheduling is constructed using the Q-learning algorithm. The state space is divided into three stages based on the comprehensive quality difference ratio between the current scheduling solution and the global optimal solution in terms of job cost and job time. The action space selects tugboat scheduling actions through an ε-greedy strategy. A Q-table is established to map the state and action value matrix, and a reward function based on a dual-objective scheduling strategy evaluation mechanism is designed.
[0031] S5, by deeply coupling the elite solution-oriented characteristics of the Jaya algorithm and the dynamic decision-making ability of Q-learning, a global-local collaborative tug scheduling optimization mechanism is constructed.
[0032] S6, the Jaya algorithm with fusion Q-learning is used to solve the constructed tug scheduling model to obtain the optimal solution of the objective function, and the optimal solution is taken as the optimal scheme of the port tug scheduling.
[0033] In the example of the application, S1 comprises the following sub-steps: S11, based on the port tug base and the horsepower of the tug, the port active tug is classified, as shown in Table 1: Table 1. Classified tug table
[0034] S12, based on the ship length, the number of tugs required under the empty and heavy loads of the ship and the required tug type, classification is carried out, as shown in Table 2: Table 2. Ship classification standard table
[0035] S13, based on the classified tug and the classified ship, the matching rule table between the ship and the active tug is established according to the example, as shown in Table 3: Table 3. Matching rule table
[0036] In the example of the application, S2 comprises the following sub-steps: S21, according to the starting position, the end position and the task start time of the ship in the port ship schedule table, a target function is established, in order to improve the port operation efficiency, the target function is the minimum total tug operation cost and the minimum total tug operation time. The established target function is as follows: (1) (2) (3) (4) Wherein, is the total operation cost of the tug, is the total operation time of the tug, is the operation cost of a single ship served by the tug, is the operation time of a single ship served by the tug, is the tug number and , is the task number and , is the tug base number and , is the number of tasks for the tugboat, , is the unit nautical mile operation cost of the tugboat in non-towing condition, is the unit nautical mile operation cost of the tugboat in towing condition, is the distance from the base to the starting point of the task, is the distance from the starting point of the task to the base, is the starting time of the task performed by the tugboat, is the ending time of the task performed by the tugboat, is the average speed of the tugboat in non-towing condition, is the average speed of the tugboat in towing condition, is the starting time of the task performed by the tugboat, is the ending time of the task performed by the tugboat, is the average speed of the tugboat in non-towing condition, is the average speed of the tugboat in towing condition, is the decision variable indicating the assignment of the tugboat to the task, is the decision variable indicating the completion of the task by the tugboat, , is the decision variable indicating the completion of the task by the tugboat, is the decision variable indicating the completion of the task by the tugboat,
[0037] Equation (1) is to minimize the operation cost of all tugboats. Equation (2) is to minimize the operation time of all tugboats. Equation (3) is the calculation method of the operation cost of the tugboat, which is the sum of the non-towing operation cost and the towing operation cost. Equation (4) is the calculation method of the operation time of the tugboat, which is the total time from the base to the starting position of the task, from the starting position of the task to the ending position of the task, and from the ending position of the task to the base of the tugboat.
[0038] S22, the above established multi-objective function is reduced to a single objective function by using a linear weighting method, but since the two objectives have a large difference in unit quantity, normalization is required. The mathematical model of the entire linear weighting measurement method is as follows:
[0039] wherein is the minimum total operation cost of the tugboat, is the minimum total operation time of the tugboat, and is the minimum operation cost of the tugboat and the minimum operation time of the tugboat the optimal value of the single-objective optimization, and are the respective weight factors of the two objectives, taking values between [0, 1].
[0040] S23, the objective function needs to satisfy the core business constraint conditions, and the specific constraint conditions are as shown in the following formula: (5) (6) (7) (8) (9) (10) (11) (12) (13) (14) (15) (16) wherein, is the tugboat number and , is the task number and , is the tugboat base number and , is the number of tugboat working tasks and , is the number of tugboats required for the task , is the horsepower of the tugboat , is the horsepower of the tugboat required for the task , is the distance from the base of the tugboat to the starting point of the task , is the starting time of the tugboat performing the task , is the end time of the tugboat performing the task , is the time when the tugboat starts from the base to perform the first task, is the time when the tugboat Average speed when not towed Decision variables refer to tugboats Assigned to task , Decision variable refers to tugboat Complete the first Sub-task (i.e., task) Return to base.
[0041] Equation (5) states that the number of tugboats assigned to each task must meet the requirements of that task. Equation (6) states that the number of tugboats assigned to each task must meet the horsepower requirements of that task. Equation (7) states that each tugboat can be assigned to a maximum of one task at a time to avoid the phenomenon of tugboats being assigned to multiple tasks simultaneously. Equation (8) states that the tugboat assigned to the first task should arrive at the starting point of the task before the task begins. Equations (9) and (10) restrict the task order of the tugboats, with each tugboat being assigned priority according to the planned task order. Equation (11) ensures that the tugboat meets the requirement of returning to its tugboat base before starting the next task. Equation (12) states that the number of tugboats docked at the base should meet the number of tugboats required for the task. Equation (13) indicates the location of the tugboat at the tugboat base before performing the first task. Equation (14) states that if there are no tasks in the current stage, the tugboat should return to the tugboat base. Equations (15) and (16) define the range of values for the variables.
[0042] In an example of the present invention, S3 includes: Using multidimensional real number encoding, for The vessels are numbered chronologically, and tugboats are assigned priority based on a first-come, first-served principle. The dimension of the solution vector depends on the total number of vessel tasks and the maximum number of tugboats required for a given task. Table 4 provides examples of specific encodings. Currently, there are six vessels waiting for tugboat service. A tugboat number of 0 indicates that no tugboat is needed. For example, tugboats numbered 7 and 9 provide service to vessel number 1, requiring only two tugboats. Tugboats numbered 4, 11, 12, and 13 provide service to vessel number 2, requiring a total of four tugboats.
[0043] Table 4. Individual Coding
[0044] For multi-objective optimization scheduling of tugboat operations, considering both cost and time, Q-learning's state space design employs a three-stage partitioning. The optimization state is dynamically determined by calculating the ratio of the combined cost and time difference between the current solution and the global optimal solution. The definitions of the three states are as follows:
[0045] Where S represents the state. The current value of the normalized objective function. to normalize the objective function.
[0046] According to the state definition formula, the ratio less than 0.1 is in the fine solution stage of the optimal solution, and the stability of the excellent solution tends to retain the current solution; the ratio between 0.1 and 0.5 is in the medium-term optimization stage, and the current solution needs to be significantly improved; the ratio is more than 0.5 and is in the initial optimization stage that needs large-scale adjustment, and the current solution needs more aggressive adjustment; the state evaluation benchmark is updated automatically with the algorithm iteration.
[0047] The discrete action space includes: action Keep the tugboat allocation of the current candidate solution unchanged, action Randomly select the tugboat allocation of the task to replace the allocation of the corresponding task in the global optimal solution, action Randomly generate a tugboat combination that meets the quantity and type constraints for the selected task.
[0048] In the examples of the present application, S4 includes the following sub-steps: S41, the action space of Q-learning designs a discrete action space including three actions, each action corresponding to a different solution adjustment strategy. Based on the above designed state space, the relative quality difference between the current solution and the global optimal solution is calculated for state evaluation, and then the adjustment action is selected through ε-greedy. Action Keep the tugboat allocation of the current candidate solution unchanged, action Randomly select the tugboat allocation of the task to replace the allocation of the corresponding task in the global optimal solution, action Randomly generate a tugboat combination that meets the quantity and type constraints for the selected task. The design of the action space and the division of the state space form a synergistic mechanism, and their relevance is shown in Table 5 below: Table 5. Association design of action space and state space
[0049] S42, based on the above state space and action space design, Q-learning calculates the Q value associated with each state and action combination, so as to select the action with the highest Q value. During iteration, the Q table used for local guided search selection is shown in Table 6 below; Table 6. Q table
[0050] If the action Q value executed in the Q table is higher, the individual will have a high probability of selecting the corresponding action. After each iteration of the algorithm, the Q value is modified, and the updated Q value formula is as follows:
[0051] wherein Make the Q value of action A taken in state S. It's the learning rate. It is the discount factor, and r is the actual reward received for action A. It is the highest expected reward for all possible operations in state S.
[0052] S43. Based on the two objectives of minimizing the total tugboat operation cost and minimizing the total tugboat operation time, design a reward function to calculate the feedback reward value, which is used to evaluate actions and optimize strategies. The reward function is as follows:
[0053] in For the reward function, To minimize the difference between the optimal solution and the current solution for the total operating cost of the tugboat, To minimize the difference between the optimal solution and the current solution for the total operating cost of the tugboat, A benchmark value to eliminate dimensional differences. and These are the corresponding weighting factors for the two objectives. Take a value between [0, 1].
[0054] Based on the above, a multi-objective optimization scheduling decision framework for tugboats is constructed using the Q-learning algorithm. After environment initialization, an initial scheduling scheme for tugboats is generated, and its total operation cost and total operation time are calculated. The state is divided based on the comprehensive difference ratio between the current solution and the optimal solution in terms of both cost and time objectives. Action selection is based on an ε-greedy strategy, and state transitions are determined by the action execution results. The executed actions are output to the environment, which returns the designed dual-objective reward function value and the new state. The Q-table is updated in real time based on feedback. The implementation process is as follows: Figure 2 As shown, the scheduling scheme can be continuously improved through the dynamic coordination of three-stage state division and three types of optimization actions.
[0055] In this embodiment of the invention, step S5 includes the following sub-steps: S51. The global search layer uses the Jaya algorithm for iterative updates of the tugboat scheduling population. The population size is set to 50. Leveraging its fast convergence characteristic, it performs coarse-grained exploration in the solution space. By comparing the difference vector between the current optimal and worst tugboat scheduling solutions, the population is guided to move towards potential regions. The solution update formula for the Jaya algorithm is as follows:
[0056] in , yes Random numbers within, For the updated tugboat allocation scheme, For the current tugboat allocation plan, , These are the current optimal and worst-case tugboat allocation schemes, respectively.
[0057] S52, the local optimization layer embeds a Q-learning adaptive adjustment stage, setting the exploration rate to 0.8, the learning rate to 0.25, and the discount factor to 0.95. Fine-grained development is achieved through a state-action-reward mapping mechanism. The state space is divided based on the difference between candidate solutions and the global optimum; the action set includes three operations: retention, copying a fragment of the optimal solution, and random perturbation; and the reward function integrates the improvement levels of both cost and time objectives.
[0058] S53. The knowledge sharing mechanism establishes a dynamic association between the global optimal solution and the Q table. When Q learns to execute actions, it can directly refer to the allocation pattern of the global optimal solution, realizing cross-level transfer of experiential knowledge.
[0059] In this embodiment of the invention, S6 includes the following sub-steps: S61. The maximum number of tugboats required for a single vessel task is the transverse length of the individual, while the current number of vessel tasks is the longitudinal length of the individual. For each individual to be generated, under the premise of satisfying the specification constraints, a corresponding number of tugboat units are randomly selected from the available tugboat set to ensure that each initial individual constitutes a feasible solution.
[0060] S62. Input the port tugboat base data, active tugboat data, and all vessel mission data obtained in step 1. In this example, the vessel timetable for entering and leaving the main port area of Lianyungang Port from 12:00 on October 14, 2024 to 12:00 on October 16, 2024 is selected as the case. During this period, 34 vessels require tugboat assistance for berthing, unberthing, and shifting operations.
[0061] S63. The task involves matching a certain number and horsepower of tugboats. The tugboat configuration for each vessel is set according to Table 3. Based on the task order and model constraints, a set of currently available tugboats is obtained. Finally, an initial candidate solution set is randomly generated based on the current set of available tugboats.
[0062] S64. Solve using the Jaya-QL algorithm with fused Q-learning. Set the number of iterations to 500. Iterate and update using the Jaya algorithm to quickly locate high-quality solution regions. Then, Q-learning optimizes the local strategy for each candidate solution. The entire flowchart is shown below. Figure 3 As shown in the figure, the results obtained by the Jaya-QL algorithm are compared with those of the Jaya, Artificial Bee Colony (ABC), and Quantum Particle Swarm Optimization (QPSO) algorithms. The convergence comparison graph is shown in the figure. Figure 4The specific target value comparison is shown in Table 7 below. In terms of operation cost optimization, compared with ABC, QPSO and standard Jaya algorithm, Jaya-QL respectively realizes a cost reduction of 24.36%, 17.92% and 22.84%. At the same time, in terms of the total operating time of the tugboat, the algorithm achieves an optimization effect of 1.58%, 0.29% and 0.93% respectively.
[0063] Table 7. Comparison of algorithms for solving different values
[0064] By comparison, the optimal solution output by the Jaya-QL algorithm is the optimal tugboat scheduling scheme for this case, and the specific tugboat allocation Gantt chart is shown in Figure 5 .
[0065] It should be noted that, according to the needs of implementation, each step / component described in the present application can be split into more steps / components, or two or more steps / components or part of the operation of the steps / components can be combined into a new step / component to achieve the purpose of the present application.
[0066] Although the present application uses terms such as artificial potential field, gravitational node, steel chain node and rigid rod more frequently, it does not exclude the possibility of using other terms. The use of these terms is only to facilitate the description and explanation of the essence of the present application; any kind of additional limitation is contrary to the spirit of the present application.
[0067] Example 2 The present embodiment provides a tugboat multi-objective optimization scheduling system, comprising: a data acquisition module for acquiring port tugboat base data, active tugboat data and all berthing and unberthing ship data within a period; a mathematical model construction module for establishing and embedding a mathematical model of tugboat multi-objective collaborative optimization scheduling based on the acquired port tugboat base data, active tugboat data and all berthing and unberthing ship data within a period, and embedding the core business constraints of port operation; an encoding module for representing a tugboat allocation scheme using multi-dimensional real number encoding; setting a state space and a discrete action space based on the tugboat allocation scheme; a multi-objective optimization decision framework construction module for constructing a tugboat scheduling multi-objective optimization decision framework using a Q-learning algorithm, and updating the state space and discrete action space of the current tugboat allocation scheme based on the tugboat scheduling multi-objective optimization decision framework; a solution model construction module for coupling the elite solution-oriented characteristics of the Jaya algorithm and the dynamic decision-making ability of Q-learning to construct a global-local collaborative optimization solution model; The dispatching scheme output module is configured to solve a mathematical model of the tugboat multi-objective collaborative optimization dispatching based on the global-local collaborative optimization solving model, and output an optimal dispatching scheme.
[0068] It should be understood that parts not elaborated in the specification are all prior art.
[0069] It should be understood that the above description of the preferred embodiments is more detailed and is not considered as a limitation to the patent protection scope of the present application. It is not necessary or possible to enumerate all the embodiments. Those skilled in the art can make substitutions or modifications without departing from the scope of the present application, which falls within the protection scope of the present application. The patent protection scope of the present application should be subject to the appended claims.
Claims
1. A multi-objective optimization scheduling method for tugboats, characterized in that, Includes the following steps: S1. Obtain port tugboat base data, active tugboat data, and data on all berthing and departure vessels within the period; S2. A mathematical model for multi-objective collaborative optimization scheduling of tugboats, based on the acquired port tugboat base data, active tugboat data, and data of all berthing and departure vessels within the period, and embedded with the core business constraints of port operations. S3. The tugboat allocation scheme is represented by multi-dimensional real number encoding; the state space and discrete action space are set based on the tugboat allocation scheme. S4. A multi-objective optimization decision-making framework for tugboat scheduling is constructed using the Q-learning algorithm. Based on the multi-objective optimization decision-making framework for tugboat scheduling, the state space and discrete action space of the current tugboat allocation scheme are locally optimized and updated. S5. By coupling the elite solution-oriented characteristics of the Jaya algorithm with the dynamic decision-making ability of Q-learning, a global-local collaborative optimization solution model is constructed. S6. Solve the mathematical model of multi-objective cooperative optimization scheduling of tugboats based on the global-local cooperative optimization solution model, and output the optimal scheduling scheme.
2. The tugboat multi-objective optimization scheduling method according to claim 1, characterized in that, The port tugboat base data includes: the number, location, and capacity of tugboat bases; the active tugboat data includes: the number of tugboats, horsepower, average speed, towing operation cost, non-towing operation cost, and the tugboat base to which they belong; the berthing and departure vessel data includes: vessel arrival time, starting position, target position, required number of tugboats, and tugboat horsepower requirements.
3. The tugboat multi-objective optimization scheduling method according to claim 2, characterized in that, Step S1 includes: S11. Classify the active tugboats in the port based on the tugboat base and the horsepower of the tugboats to obtain the classified tugboats; S12. Classify the ships based on their length, the number of tugboats required under both empty and heavy load conditions, and the required tugboat types to obtain the classified ships. S13. Establish a matching rule table between ships and active tugboats based on the classified tugboats and classified vessels.
4. The tugboat multi-objective optimization scheduling method according to claim 3, characterized in that, The mathematical model for the multi-objective cooperative optimization scheduling of tugboats in step S2 includes a multi-objective function and constraints. The multi-objective function is: (1) (2) (3) (4) in, The total operating cost of the tugboat. Total tugboat operating time, The cost of operating a single tugboat. The operating time for a single tugboat service vessel. Number the tugboat and , Number the task and , Number the tugboat base and , For the number of tugboat work tasks and , This represents the unit nautical mile operating cost for tugboats in non-towing conditions. This refers to the unit nautical mile operating cost under tugboat towing conditions. For tugboats From the base To the starting point of the mission distance, For tugboats From the starting point of the mission Arrive at the base distance, For tugboats Execute the task The start time, For tugboats Execute the task End time, For tugboats Average speed when not being towed For tugboats Average speed during towing Decision variables refer to tugboats Assigned to task , Decision variable refers to tugboat Complete the first Sub-task is the same as task Return to base; The multi-objective function established above is normalized into a single-objective function using the linear weighting method: in To minimize the total operating cost of the tugboat, Minimum total tugboat operating time, and It is the model's minimum operating cost for the tugboat. and minimum operation time The optimal value of single-objective optimization. and These are the corresponding weighting factors for the two objectives. Take values between [0, 1]; The constraints include: (5) (6) (7) (8) (9) (10) (11) (12) (13) (14) (15) (16) in, Number the tugboat and , Number the task and , Number the tugboat base and , For the number of tugboat work tasks and , For the task Number of tugboats required For tugboats horsepower, Task Required tug horsepower For tugboats From the base To the starting point of the mission distance, For tugboats Execute the task The start time, For tugboats Execute the task End time, For tugboats Execute the first mission from the base Departure time For tugboats Average speed when not being towed Decision variables refer to tugboats Assigned to task , Decision variable refers to tugboat Complete the first Sub-task is the same as task Return to base.
5. The tugboat multi-objective optimization scheduling method according to claim 4, characterized in that, The state space in step S3 includes: Where S represents the state. The current value of the normalized objective function. To find the optimal value of the normalized objective function, For the refinement stage of the optimal solution, For the mid-term optimization phase, This is the initial optimization phase. Discrete action space includes: actions Keep the tugboat allocation of the current candidate solution unchanged, and proceed with the action. The randomly selected tugboat assignment is replaced with the assignment of the corresponding task from the globally optimal solution, and the action... For the selected task, randomly generate a combination of tugboats that meets the quantity and type constraints.
6. The tugboat multi-objective optimization scheduling method according to claim 5, characterized in that, The multi-dimensional real number encoding method used in step S3 to represent the tugboat allocation scheme includes: Using multidimensional real number encoding, for The vessels are numbered in chronological order, and tugboats with earlier numbers are assigned priority based on the principle of first-come, first-served.
7. The tugboat multi-objective optimization scheduling method according to claim 6, characterized in that, Step S4 includes: Step S41: Based on the set state space, calculate the relative quality difference between the current solution and the global optimal solution to evaluate the state, and then select the corresponding discrete action space through ε-greedy. Step S42: Calculate the Q-values associated with the state space and discrete action space based on the Q-learning algorithm, and select the action with the highest Q-value; after each iteration, update the Q-values, and the formula for the updated Q-values is as follows: in Let Q be the value of action A taken in state S. It's the learning rate. It is a discount factor. r It is the actual reward received for action A. It is the highest expected reward for all possible actions in state S; Step S43: Based on the two objectives of minimizing the total tugboat operation cost and minimizing the total tugboat operation time, a reward function is set to calculate the feedback reward value. The reward function is as follows: in For the reward function, To minimize the difference between the optimal solution and the current solution for the total operating cost of the tugboat, To minimize the difference between the optimal solution and the current solution for the total operating cost of the tugboat, A benchmark value to eliminate dimensional differences, and These are the corresponding weighting factors for the two objectives. Take a value between [0, 1].
8. The tugboat multi-objective optimization scheduling method according to claim 7, characterized in that, The global-local collaborative optimization solution model includes a global search layer and a local optimization layer; step S5 includes: Step S51: The global search layer uses the Jaya algorithm for iterative updates of the tugboat scheduling population. It performs coarse-grained exploration in the solution space, guiding the population towards potential regions by comparing the difference vectors between the current optimal and worst tugboat scheduling solutions. The solution update formula for the Jaya algorithm is as follows: in , yes Random numbers within, For the updated tugboat allocation scheme, For the current tugboat allocation plan, , These are the current optimal and worst-case tugboat allocation schemes, respectively; Step S52: Embed the Q-learning algorithm in the local optimization layer to achieve fine-grained development by designing a state-action-reward mapping mechanism; Step S53: The knowledge sharing mechanism establishes a dynamic relationship between the global optimal solution and the Q value. When the Q-learning algorithm performs an action, it directly refers to the allocation pattern of the global optimal solution to realize cross-level transfer of experiential knowledge.
9. A tugboat multi-objective optimization scheduling method according to claim 8, characterized in that, Step S6 includes the following sub-steps: Step S61: The maximum number of tugboats required for a single ship mission is the transverse length of the individual. At the current stage, the number of ship missions is the longitudinal length of the individual. For each individual to be generated, under the premise of meeting the specification constraints, a corresponding number of tugboat units are randomly selected from the available tugboat set to ensure that each initial individual constitutes a feasible solution. Step S62: Input the obtained port tugboat base data, active tugboat data, and all berthing and departure vessel task data within the period; Step S63: Match a certain number and horsepower of tugboats for a single ship task, and obtain the set of currently available tugboats according to the task order and model constraints. Finally, randomly generate an initial candidate solution set based on the set of currently available tugboats. Step S64: Using the global-local co-optimization solution model, the Jaya algorithm is used for iterative updates to quickly locate high-quality solution regions. Subsequently, the Q-learning algorithm performs local policy optimization on each candidate solution, and the output of the optimal solution is the optimal tugboat scheduling scheme for the ship mission.
10. A tugboat multi-objective optimization scheduling system, characterized in that, include: Data acquisition module: It is used to acquire port tugboat base data, active tugboat data, and data on all berthing and departing vessels within the period; Mathematical model building module: It is used to establish and embed a mathematical model for multi-objective collaborative optimization scheduling of tugboats based on the acquired port tugboat base data, active tugboat data and all berthing and departure data within the period, and to embed the core business constraints of port operations. Encoding module: It is used to represent the tugboat allocation scheme using multidimensional real number encoding; and to set the state space and discrete action space based on the tugboat allocation scheme; Multi-objective optimization decision framework construction module: It is used to construct a multi-objective optimization decision framework for tugboat scheduling using the Q-learning algorithm, and locally optimize and update the state space and discrete action space of the current tugboat allocation scheme based on the multi-objective optimization decision framework for tugboat scheduling. The solution model building module is used to couple the elite solution-oriented characteristics of the Jaya algorithm with the dynamic decision-making ability of Q-learning to build a global-local co-optimization solution model. Scheduling scheme output module: It is used to solve the mathematical model of multi-objective cooperative optimization scheduling of tugboats based on the global-local cooperative optimization solution model, and output the optimal scheduling scheme; The tugboat multi-objective optimization scheduling system is used to execute the steps in the tugboat multi-objective optimization scheduling method according to any one of claims 1-9.