Container terminal multi-resource collaborative scheduling optimization method based on DCRLQ-GWO algorithm
The DCRLQ-GWO algorithm optimizes the multi-resource scheduling of container terminals, solving the problem of collaborative scheduling of multiple resources in container terminals. It reduces port pollution gas emissions and improves operational efficiency, and provides flexible scheduling strategies to meet different needs.
Patent Information
- Application Number
- CN202511006979.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies are insufficient to effectively solve the problem of multi-resource collaborative scheduling at container terminals, especially in reducing port pollutant emissions while ensuring operational efficiency. Furthermore, conventional algorithms are inadequate for solving complex multi-resource collaborative scheduling models.
The DCRLQ-GWO algorithm is adopted, combined with Circle chaotic mapping, reverse learning mechanism, quantum potential well search and Levy flight strategy, to optimize the integrated scheduling of waterway-berth-quay crane-container truck. The Grey Wolf algorithm is improved by reinforcement learning to enhance global and local search capabilities.
It significantly reduces the time ships spend in port, lowers pollutant emissions, improves port operational efficiency and service quality, and provides flexible scheduling strategies to balance the cost of pollutant emission treatment with the time ships spend in port.
Smart Images

Figure CN120952647A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of port logistics scheduling technology, and specifically relates to a multi-resource collaborative scheduling optimization method for container terminals based on the DCRLQ-GWO (DDPG-Circle-Reverse-Levy-Quantum-Grey Wolf Optimization) algorithm. Background Technology
[0002] With the deepening of global economic integration, the shipping industry, as a vital pillar of international trade, plays an irreplaceable role in global economic prosperity. Ports, as key nodes connecting sea and land transportation, directly impact the sustainable development of regional economies through their operational efficiency. In recent years, my country's port throughput has continued to grow, making it one of the world's largest port cargo handling nations. However, the problem of pollutant emissions generated during port operations has become increasingly serious, posing a threat to the surrounding environment and residents' health. To address this challenge, the Chinese government has introduced a series of green port construction policies, promoting the development of ports towards low-carbon, environmentally friendly, and intelligent directions. Against this backdrop, port scheduling, as a core aspect of operation and management, is particularly important for optimization research. How to further reduce port area pollutant emissions while ensuring operational efficiency and reducing the total time ships spend in port has become one of the core key issues in green port construction.
[0003] Currently, maritime transport handles approximately 90% of global international trade. With the development of trade and the increase in ship size, ports and the shipping industry face challenges such as ship traffic congestion, berth resource shortages, and increased pollutant emissions. Upgrading port infrastructure requires substantial funding and time; therefore, given certain hardware conditions, enhancing information-based and intelligent scheduling strategies is crucial for the rational allocation of resources such as berths, waterways, quay cranes, and storage yards, thereby improving loading and unloading efficiency and port competitiveness.
[0004] Existing research largely focuses on independent optimization of single resources, with insufficient attention paid to the systematic collaborative scheduling of multiple resources. Furthermore, collaborative scheduling models considering operations involving waterways, berths, quay cranes, and container trucks are complex and difficult to solve effectively using conventional algorithms. Meanwhile, port pollutants mainly originate from ships entering and leaving the port, ships berthing in port, and wharf operations, including particulate matter, sulfur oxides, and nitrogen oxides. In the context of green port construction, achieving energy conservation and emission reduction at the operational level can save costs. Excellent energy conservation and emission reduction strategies can be combined with technological equipment to further improve effectiveness. Therefore, in-depth research is urgently needed on the integrated scheduling problem of waterways and berths under complex navigation conditions that consider pollutant emissions. Summary of the Invention
[0005] To address the problems existing in the prior art, the present invention aims to provide a multi-resource collaborative scheduling optimization method for container terminals based on the DCRLQ-GWO algorithm. To efficiently solve the optimization model, this invention introduces a population initialization strategy based on Circle chaotic mapping, a reverse learning mechanism, a quantum potential well search mechanism, a Lévy flight strategy, and a reinforcement learning algorithm on the basis of the traditional Grey Wolf optimization algorithm, constructing the DCRLQ-GWO (DDPG-Circle-Reverse-Levy-Quantum-Grey Wolf Optimization) algorithm, thus enhancing the global and local search capabilities of the Grey Wolf algorithm. The integrated scheduling optimization of waterways, berths, quay cranes, and container trucks implemented in this invention can not only effectively reduce ship port time and pollutant emissions, but also significantly improve the overall port operating efficiency and service quality, which is of great significance for promoting the construction of green and smart ports in my country.
[0006] The objective of this invention is achieved through the following technical solution:
[0007] This invention provides a multi-resource collaborative scheduling optimization method for container terminals based on the DCRLQ-GWO algorithm, comprising the following steps:
[0008] Step 1, Estimation of pollutant gas emissions
[0009] Based on the activity-based calculation method, the emission of a certain pollutant gas is estimated by multiplying the emission factor of the pollutant gas by the energy consumption generated by the activity process. Specifically, this includes:
[0010] Pollutant emissions from ship navigation E 1,z Pollutant emissions from ships waiting 2,z Pollutant emissions from quay crane operations 3,z And pollutant emissions from truck operations E 4,z ;
[0011] Step 2, Model Building
[0012] S21, Make model assumptions, including that the expected arrival time of the ship is known, the channel depth is calculated according to the tide table, the port access channel is a two-way channel, and each ship is allowed to berth once;
[0013] S22, Establish the objective function and constraints. The specific objective function is as follows:
[0014] F=min(ω1·k1·F1+ω2·k2·F2)
[0015]
[0016] In the formula, F is the objective function, ω1 and ω2 are weight adjustment factors, k1 and k2 are quantity balance coefficients, F1 is the total time ships spend in port, v is the number of ships, and T is the weight adjustment factor. F,i It is the time when ship i departs from the port, T A,i F2 is the estimated arrival time of vessel i, F2 is the total cost of treating pollutant emissions, and p z E is the unit treatment cost of the z-th pollutant. 1,z It refers to the amount of polluting gases emitted by ships during navigation, E 2,z It is the emission of pollutants from ships while they are waiting, E 3,z It is the emission of pollutants from quay crane operations, E 4,z It is the emission of polluting gases from container truck operations;
[0017] The constraints include ship constraints, quay crane constraints, container truck constraints, and waterway constraints;
[0018] Step 3: Solve the model using the DCRLQ-GWO algorithm.
[0019] The DCRLQ-GWO algorithm is designed to solve the model. The algorithm includes: introducing a chaotic initialization strategy, a reverse learning mechanism, a quantum potential well search mechanism, a Lévy flight strategy, and a reinforcement learning algorithm on the basis of the traditional Grey Wolf algorithm.
[0020] Step 4, Obtaining the optimal scheduling scheme
[0021] The specific scheduling scheme is obtained based on the algorithm described in step 3. The scheduling scheme includes the berthing position of the ship, the number of quay cranes, the number of container trucks, the operation priority, and the speed of the ship entering and leaving the port.
[0022] The port multi-resource joint scheduling optimization method and system based on the reinforcement learning-improved gray wolf algorithm provided by this invention has the following significant advantages:
[0023] 1. Improve port operation efficiency and service level: By optimizing the scheduling problem of waterway-berth-quay crane-truck in continuous berths and two-way channels, the method can rationally allocate berth resources, waterway resources, and loading and unloading equipment resources at the terminal, reduce the time ships spend in port and waiting time at anchor, improve the loading and unloading efficiency and traffic efficiency of the terminal, thereby increasing the cargo throughput of the container terminal and enhancing the terminal's competitiveness and profitability. For example, in the case study, the scheduling scheme obtained by this method can effectively plan the berthing positions of ships, the number of quay cranes and trucks, and significantly optimize the time ships spend in port.
[0024] 2. Achieving Green and Sustainable Port Development: This approach considers the impact of environmental factors (polluting gas emissions) on terminal operations and incorporates them into the optimization objective function. While meeting the operational needs of the terminal, it aims to minimize pollutant emissions from ships entering and leaving the port and from terminal equipment, thereby improving air quality in and around the terminal and mitigating the harmful effects of the greenhouse effect and climate change. By accurately estimating and optimizing pollutant emissions at each stage, including ship navigation, waiting, quay crane operations, and container truck operations, the cost of pollutant treatment can be reduced. As demonstrated in the example, this method effectively controls the cost of pollutant emission treatment.
[0025] 3. Enhanced Algorithm Optimization Capabilities: The designed DCRLQ-GWO algorithm enhances the global and local search capabilities of the Grey Wolf algorithm by introducing a Circle chaotic mapping initialization strategy, a reverse learning mechanism, a quantum potential well search mechanism, a Lévy flight strategy, and a reinforcement learning algorithm. Compared with traditional algorithms, this algorithm has significant advantages in solution accuracy, convergence speed, and stability. As demonstrated in the experimental examples, in the first set of examples, the average optimal value of this algorithm is 35.65% better than the traditional GWO algorithm, and in the second set of examples, it is 37.53% better, enabling more efficient solutions to complex port multi-resource joint scheduling optimization models.
[0026] 4. Provides flexible scheduling strategies: Establishes a dual-objective optimization model for pollutant gas treatment costs and ship waiting time. By adjusting the weights of the objective function, the scheduling strategy can be flexibly switched, providing quantitative decision support for green scheduling. It can focus on reducing pollutant gas emission treatment costs or shortening ship waiting time, or achieve a balance between the two, meeting the port scheduling needs of different scenarios. Attached Figure Description
[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments:
[0028] Figure 1 A graph showing the change over time in the maximum draft of vessels that can safely navigate the channel under the influence of tides;
[0029] Figure 2 , Figure 3 The following are Gantt charts of the scheduling schemes for the first and second sets of examples. The orange area represents ship unloading operations, the blue area represents ship loading operations, the red numbers at the top represent the number of quay cranes, and the blue numbers at the bottom represent the number of container trucks.
[0030] Figure 4 , Figure 5 The graphs show the changes in resource usage over time in the first and second sets of examples, respectively.
[0031] Figure 6 , Figure 7 These are the iterative convergence curves for the first set of examples;
[0032] Figure 8 Statistical graphs of experimental results for different schemes (the objective function is represented on the left axis, and the port time and emission cost are represented on the right axis, where the emission cost is the result after multiplying by the quantity balance coefficient). Detailed Implementation
[0033] Example 1
[0034] This embodiment provides a multi-resource collaborative scheduling optimization method for container terminals based on the DCRLQ-GWO algorithm, including the following steps:
[0035] Step 1, Estimation of pollutant gas emissions
[0036] Based on the activity-based calculation method, the emission of a certain pollutant gas is estimated by multiplying the emission factor of the pollutant gas by the energy consumption generated by the activity process. Specifically, this includes:
[0037] Pollutant emissions from ship navigation E 1,z Pollutant emissions from ships waiting 2,z Pollutant emissions from quay crane operations 3,z And pollutant emissions from truck operations E 4,z .
[0038] (1) Pollutant emissions from ship navigation
[0039] The amount of pollutants produced by a ship during navigation depends on its fuel consumption. When a ship passes through a port channel, the rate of fuel consumption is mainly affected by its speed. The amount of pollutants emitted is equal to the fuel consumption multiplied by the emission factor corresponding to that pollutant. Therefore, the amount of pollutants emitted by a ship during navigation is E. 1,z The calculation formula is as follows.
[0040]
[0041] In the formula, E s1,i,z E s2,i,z E s3,i,z denoted as z, representing the emissions of the z-th pollutant from the main engine, auxiliary engine, and boiler of the i-th ship; v represents the number of ships.
[0042] E s1,i,z =K 1,i ×κ s,i ×T s,i ×A s1,i,z (2)
[0043] E s2,i,z =K 2,i ×κ s,i ×T s,i×A s2,i,z (3)
[0044] E s3,i,z =K s3,i ×T s,i ×A s3,i,z (4)
[0045] K 1,i K 2,i K s3,i These are the rated power (kW) of the main engine, auxiliary engine, and boiler of the ship during navigation, respectively; κ s,i T represents the load factor of ship i during navigation. s,i Let A be the sailing time (h) of vessel i; s1,i,z A s2,i,z A s3,i,z These are the emission factors (g / kWh) of the z-th pollutant from the main engine, auxiliary engine, and boiler during the i-th voyage of the ship.
[0046] When a ship is sailing, its main engine load factor LF 航 The calculation formula is as follows.
[0047]
[0048] In the formula, V s and V max,s These are the ship's actual speed and maximum design speed (kn), respectively.
[0049] (2) Pollutant emissions during ship waiting
[0050] Emissions from ships while in port are one of the main sources of pollutants affecting the port and its surrounding environment. The total emissions during a ship's berthing period can be calculated by combining the emission factors of pollutants with the unit emission cost. The pollutant emissions from ships while waiting at berths or anchorages are calculated using the following formula.
[0051]
[0052] E w2,i,z =K 2,i ×κ w,i ×A w2,i,z ×T w,i (7)
[0053] E w3,i,z =K w3,i ×A w3,i,z ×T w,i (8)
[0054] In the formula, E w2,i,z E w3,i,zThese are the emissions (g) of the zth pollutant from the ship's auxiliary engine and boiler, respectively; K w3,i This refers to the rated power (kW) of the boiler when the ship is waiting; κ w,i A represents the average load factor of the auxiliary engines while ship i is waiting; w2,i,z A w3,i,z The emission factors (g / kWh) of the z-th pollutant gas from the i-th auxiliary engine and boiler of the ship are respectively; T w,i It is the waiting time (h) for ship i.
[0055] (3) Pollutant emissions from quay crane operations
[0056] Quay cranes are crucial mechanical equipment for loading and unloading containers. During this process, the pollutants generated by shore power are a significant source of air pollution in ports. Quay crane operation time mainly consists of net working time and waiting time. Based on relevant data, this embodiment uses the following formula to estimate the pollutant emissions E generated by the quay crane. 3,z :
[0057]
[0058] In the formula, A c,z K represents the emission factor (kg / kWh) of the z-th pollutant during quay crane operations. c1 and K c0 T represents the energy consumption (kWh / h) of the quay crane during operation and waiting, respectively. c1,i and T c0,i These represent the total working time (h) and total waiting time (h) of all quay cranes serving vessel i, respectively.
[0059] (4) Pollutant emissions from truck operations
[0060] During ship loading and unloading operations, container trucks horizontally transport containers between the ship and the yard. This process emits a large amount of polluting gases. The operation of container trucks includes three phases: heavily loaded truck driving, unloaded truck driving, and truck waiting time. Based on relevant data, this embodiment uses the following formula to estimate the amount of polluting gas emissions E. 4,z :
[0061]
[0062] In the formula, A g0,z A g1,z A g2,z and represent the emission factors (kg / h) of the z-th pollutant when the truck is waiting, operating empty, and operating with a heavy load, respectively. T g0,i T g1,i T g2,iThese represent the total waiting, empty, and loaded times (h) for all container trucks serving vessel i.
[0063] Step 2, Model Building
[0064] S21, Make model assumptions, including:
[0065] 1. The estimated arrival time of the vessel is known.
[0066] 2. Considering the impact of tides on ships entering and leaving the port, the channel depth is calculated based on the tide table, and all ships are only allowed to enter and leave the port when the channel depth meets the requirements for safe navigation.
[0067] 3. The port's entrance channel is a two-way channel. Vessels can enter and leave the port simultaneously as long as they meet the safety distance requirement. Vessels traveling in the same direction must meet the safety distance based on their length, and vessels traveling in opposite directions must meet the safety distance based on their width.
[0068] 4. Each vessel is allowed to moor once, and the berth may not be moved after mooring.
[0069] 5. The berths are continuous berths, and vessels must maintain a safe distance based on their length when berthing.
[0070] 6. After the vessel begins operation, the working efficiency of a single quay crane remains unchanged, while the working efficiency decreases at a fixed decay rate when multiple quay cranes are operating simultaneously.
[0071] 7. The number of quay cranes configured to operate on board or off board must meet the requirements of vessel operation restrictions.
[0072] 8. Quay cranes in the same harbor basin can be moved to adjacent berths when not in use, but quay cranes in different harbor basins cannot be moved.
[0073] 9. The number of yard cranes in the container yard is sufficient, and the loading and unloading efficiency remains unchanged.
[0074] 10. After the vessel arrives at the port, it will enter the port as soon as possible and leave the port as soon as the operation is completed. If there is a conflict between the vessel's entry and exit times, the vessel with higher priority will enter and exit the port first.
[0075] The model variables are shown in Tables 1 to 3 below.
[0076] Table 1. Statistics of State Variables
[0077]
[0078] Table 2 Process Variables Table
[0079]
[0080] Table 3 Decision Variables Table
[0081]
[0082] S22, Establish the objective function and constraints. The specific objective function is as follows:
[0083] F=min(ω1·k1·F1+ω2·k2·F2) (1)
[0084]
[0085] In the formula, F is the objective function, ω1 and ω2 are weight adjustment factors, and k1 and k2 are quantity balance coefficients to ensure that the influence weight of each sub-objective function is equal. In this embodiment, k1 = 1, k2 = 100, F1 is the total time the ship spends in port, v is the number of ships, and T is the total time the ship spends in port. F,i It is the time when ship i departs from the port, T A,i F2 is the estimated arrival time of vessel i, F2 is the total cost of treating pollutant emissions, and p z E is the unit treatment cost of the z-th pollutant. 1,z It refers to the amount of polluting gases emitted by ships during navigation, E 2,z It is the emission of pollutants from ships while they are waiting, E 3,z It is the emission of pollutants from quay crane operations, E 4,z It is the emission of polluting gases from container truck operations.
[0086] The constraints include ship constraints, quay crane constraints, container truck constraints, and waterway constraints.
[0087] (1) Ship constraints
[0088]
[0089] V min,s ≤V E,i ≤V max,s (7)
[0090] V min,s ≤V F,i ≤V max,s (8)
[0091]
[0092] max(V en,i V ex,i )≤G D,j j = VO i (16)
[0093]
[0094] 0≤V B,i ≤G L,j j=V O,i(18)
[0095]
[0096] Equation (4) stipulates that the time for a ship to enter the channel upon entering the port must be after the ship has arrived at the port. Equations (5) and (6) respectively give the calculation formulas for the time when a ship enters and leaves the port. Equations (7) and (8) stipulate that the speed of a ship entering and leaving the port must be within the limit range. Equation (9) gives the calculation formula for the start time of operations. Equations (10) and (11) give the calculation formulas for the time consumed in loading and unloading operations. Equations (12) and (13) stipulate that the actual loading and unloading efficiency is equal to the smaller value between the quay crane and truck operation efficiency. Equation (14) gives the theoretical operation efficiency of a single truck serving a ship i in port basin j, which is the container volume transported by a single truck in a single trip divided by the turnaround time of a single trip (truck empty and loaded time, yard operation time, quay crane operation time). Equation (15) gives the calculation formula for the time when a ship enters the channel upon leaving the port. Equation (16) stipulates that the draft of a ship entering or leaving the port must be less than or equal to the safe water depth of the port basin. The safe distance between two ships when berthing is 0.3 times the length of the larger ship. Equation (17) stipulates that at any given time, the sum of the total length of all ships berthed in harbor basin j and the safe distance between the ships cannot exceed the total length of the berth. Equation (18) stipulates that the berthing position of ship i in harbor basin j must be within the length of the harbor basin berth. Equations (19) and (20) stipulate that each ship can only berth at one berth.
[0097] (2) Restraints of quay cranes
[0098]
[0099] Equation (21) stipulates that a quay crane can serve a maximum of one vessel at a time. Equation (22) stipulates that the operation time of the quay crane is within the vessel's loading and unloading operation time. Equation (23) stipulates that the number of quay cranes configured for vessel i should meet the upper and lower limits of the number of quay cranes under the work restriction. Equation (24) stipulates that the total number of quay cranes operating in the harbor basin j cannot exceed the total number of quay cranes in the harbor basin.
[0100] (3) Truck constraints
[0101]
[0102] Equation (25) stipulates that a container truck can serve a maximum of one vessel at the same time. Equation (26) stipulates that the time a container truck serves vessel i should be within the vessel's loading and unloading operation time. Equation (27) stipulates that the total number of container trucks operating at the terminal at the same time cannot exceed the total number of container trucks available at the terminal. Equations (28), (29), and (30) respectively give the calculation methods for the total empty, loaded, and waiting time of the container trucks serving vessel i.
[0103] (4) Channel constraints
[0104]
[0105]
[0106] Equations (31) and (32) stipulate that the draft of vessels entering and leaving the port must not exceed the safe depth of the channel. Equations (33) and (34) stipulate that no two vessels may overtake each other during the entry and exit phases. Equations (35) and (36) stipulate that vessels traveling in the same direction must maintain a safe distance. Assuming that the speed of the vessels in the channel is constant, the closest distance between the two vessels can be calculated by taking the minimum difference between the time the two vessels enter and leave the channel, multiplying it by the speed of the following vessel. The closest distance must meet the safe distance requirement of the channel, which is the sum of the lengths of the two vessels plus the length of the larger vessel. Equation (37) stipulates that vessels traveling in opposite directions must maintain a safe distance, which is twice the beam of the two vessels plus the beam of the larger vessel should be less than the total width of the channel.
[0107] Step 3: Solve the model using the DCRLQ-GWO algorithm.
[0108] The DCRLQ-GWO algorithm is designed to solve the model. The algorithm includes: introducing a chaotic initialization strategy, a reverse learning mechanism, a quantum potential well search mechanism, a Lévy flight strategy, and a reinforcement learning algorithm on the basis of the traditional Grey Wolf algorithm.
[0109] 1. Grey Wolf Algorithm
[0110] As a widely used heuristic optimization algorithm, the Grey Wolf algorithm mainly updates the population based on the following formula.
[0111] X i =X p -O1·d p ,i=1,2,3,p=α,β,δ (38)
[0112] O i =2(2-2·t / t) max )·o2-(2-2·t / t max ), i = 1, 2 (39)
[0113] d p =|2o3·X P -X|, p=α,β,δ (40)
[0114] d p =|O2·X P -X| (41)
[0115]
[0116] In the formula, t represents the t-th iteration, tmax Indicates the maximum number of iterations, X, X new Then X represents the position of the gray wolf in the current and next iterations, respectively. p Let d represent the positions of α, β, and δ wolves, respectively. p O1 and O2 represent the distance between the gray wolf and its prey. O1 and O2 are two adjustment variables that simulate the disturbance caused by uncertainty. O2 and O3 are two random number vectors with values between [0,1].
[0117] 2. Chaos Initialization
[0118] Chaotic mapping initialization of the population leverages the characteristics of chaotic systems to generate an initial population within the search space. Chaotic systems exhibit randomness, ergodicity, and sensitivity to initial conditions; utilizing these characteristics allows for a broader coverage of the search space. Compared to traditional random initialization, this method offers advantages such as a more uniform population distribution, greater diversity, and improved global search capabilities, preventing premature entrapment in local optima.
[0119] The Circle chaotic mapping is a special type of chaotic mapping. It is implemented using the following iterative formula:
[0120]
[0121] Where, x t+1 Let x be the population array in the (t+1)th iteration based on the chaotic initialization strategy. t Let be the population array for the t-th iteration based on the chaotic initialization strategy, α and b are control parameters, commonly 0.5 and 0.2, mod is the modulo function, and the initial population array x0 is a set of randomly generated random numbers, with the range of values based on the Circle chaotic mapping being (0,1).
[0122] 3. Reverse learning mechanism
[0123] In the Grey Wolf Algorithm, the reverse learning mechanism is mainly used in the search process. For a good solution generated in the current iteration, its reverse solution is calculated, and the fitness of the two solutions is compared. The better solution is selected to participate in subsequent iterations. In this way, the algorithm is guided to converge to the global optimum more quickly. This embodiment, based on the traditional reverse learning strategy and the Grey Wolf Algorithm update formula, designs the following population update formula based on the reverse learning mechanism.
[0124]
[0125] X i =X p -O i ·d p ,i=1,2,3,p=α,β,δ (45)
[0126] O i =2(2-2·t / t) max )·o2-(2-2·t / t max ), i = 1, 2, 3 (46)
[0127] d p =|2o3·X P -X|, p=α、β、δ (47)
[0128] X' new X1' is the new population of the gray wolf algorithm obtained iteratively based on the inverse learning mechanism, calculated using the inverse learning mechanism based on the α wolf position; X'2 is the gray wolf algorithm population calculated using the inverse learning mechanism based on the β wolf position; and X3' is the gray wolf algorithm population calculated using the inverse learning mechanism based on the δ wolf position. α Let X be the position of the α wolf. β For the position of the β wolf, X δ For the δ wolf position, O1-O3 are adjustment variables for the disturbance caused by the uncertainty of the simulated gray wolf algorithm, and d α d β d δ These represent the distances between the α wolf, β wolf, and δ wolf and their prey, respectively, and o1-o3 are all random numbers between 0 and 1.
[0129] In this way, a reverse learning strategy suitable for the Grey Wolf algorithm is established. At the same time, this method does not rely on the boundary information of the search space and can solve the scheduling optimization problem studied in this embodiment quite well.
[0130] 4. Levi's Flight Strategy
[0131] The Levy flight strategy significantly improves upon the shortcomings of the gray wolf algorithm. Its long stride helps wolves escape local optima and perform better global searches; its short stride facilitates refined searches within potentially optimal regions, leading to better local searches. Applying the Levy flight strategy, while preserving the original social hierarchy and hunting mechanism of the gray wolf optimization algorithm, significantly enhances its global search capability and convergence speed. Introducing Levy flight into the gray wolf optimization algorithm yields the following gray wolf population update formula:
[0132]
[0133] In the formula, X new The new population is obtained through the iteration of the gray wolf algorithm. X1, X2, and X3 are the gray wolf algorithm populations calculated based on the positions of α wolf, β wolf, and δ wolf, respectively, and θ is a random number between [0,2]. Let represent point-to-point multiplication, and levy(θ) represent a path following a Lévy distribution. As the weights for controlling the step size, based on the content of this study, we set...
[0134] 5. Quantum potential well search mechanism
[0135] By introducing a quantum potential well search mechanism to simulate particle behavior within a quantum potential well, the Grey Wolf algorithm's ability to escape local optima is significantly enhanced, allowing for a more effective balance between global and local searches. In quantum space, the particle's position and velocity cannot be simultaneously determined; therefore, a wave function is introduced to represent the particle's state, with each dimension P... i、j By establishing a one-dimensional delta potential well, the wave function of the particle in any dimension is obtained as follows:
[0136]
[0137] In the formula, X represents the particle position, P represents the center of the potential well, σ represents the characteristic length of the δ-potential well, t is the iteration number, and u i,j (t) represents a random number uniformly distributed in [0,1]. U(0,1) represents a random number uniformly distributed between 0 and 1. The "±" is determined by the random number. When the random number is greater than 0.5, "+" is used, and "-" is used in other cases.
[0138] After establishing the basic evolutionary model of particles in quantum space, it is also necessary to set a reasonable σ. i,j To ensure algorithm convergence, we introduce the average optimal position Z(t), which is the average of the optimal positions of all particles. Based on this, we can obtain the following equation for the evolution of particle positions in N dimensions:
[0139]
[0140] In the formula, ψ(t) is the expansion-contraction coefficient, and N is the total number of particles.
[0141] The feasible region of the solution using the gray wolf algorithm is considered as a quantum space, and each individual gray wolf is considered as a quantum particle. The positions X of α, β, and δ wolves are defined. α X β X δ Treating the individual as a potential well point and representing the individual state with a wave function, we can obtain the positional evolution equations for α, β, and δ wolves.
[0142]
[0143] In the formula, X i,j Let Z represent the value of the j-th dimension of the i-th wolf, where i = α, β, δ. i,j (t) represents the average optimal value of the i-th wolf in the j-th dimension, u i,jψ(t) represents a random number uniformly distributed in [0,1]. ψ(t) is the expansion-contraction coefficient, whose value decreases linearly from m to n with each iteration. Typically, m = 1 and n = 0.5.
[0144] 6. Strengthen the learning mechanism
[0145] In the traditional Grey Wolf Algorithm, increasing the weight of the α wolf leads the population to locally search towards the optimal wolf, while increasing the weights of the β and δ wolves leads to a global search across the solution space. While improvements to the chaotic initialization and quantum well search mechanisms enhance population diversity, the weights of the α, β, and δ wolves remain constant regardless of environmental changes. Over-reliance on the α wolf can lead to premature convergence, while increasing the weights of the β and δ wolves reduces local search capability. This static update, lacking interaction with the environment, struggles to balance exploration and development. The introduction of reinforcement learning endows GWO with "environmental perception-decision-feedback" capabilities. By interacting with the environment, the algorithm can selectively adjust the weights of the α, β, and δ wolves. This dynamic collaborative mechanism overcomes the limitation of fixed leader weights in traditional swarm intelligence algorithms, establishing a strong feedback link between the current population state and algorithm parameter adjustments. This enables the Grey Wolf population to continuously evolve, achieving a dynamic balance between global search and local optimization in a complex multimodal solution space.
[0146] This embodiment uses the improved Grey Wolf algorithm update process as the environment, realizes model training through interaction with the network environment, outputs the directed mutation scheme of the genetic algorithm, and uses the DDPG algorithm to optimize GWO. DDPG makes up for the shortcomings of traditional reinforcement learning algorithms that can only handle discrete and low-dimensional action spaces. It is an algorithm based on deterministic policy gradient and actor-critic architecture that can run in continuous action space.
[0147] The interaction process between DDPG and the environment can be represented by the following quadruple.<S,A,R,γ> Where S represents the state, A represents the action, R is the reward, and γ is the discount factor for future rewards, which is set to 1 in this embodiment, meaning that the fitness value is not discounted in consecutive iterations. The specific calculation methods for S, A, and R are given below.
[0148] (1) State space
[0149] The state space consists of a discrete real number and a continuous real number. Let t0 represent the number of iterations in which the optimal fitness of the population remains unchanged. n Let X represent the current iteration number. t ={x1,x2,x3,......x n} represents the population at the t-th iteration, x nLet S represent the nth individual. Considering both population diversity and average fitness, state S can be represented in the following form.
[0150]
[0151] S i =[s1,s2] (56)
[0152] In the formula, Var represents the standard deviation of the array. x represents the average fitness value of all individuals in the population. i Let be the fitness value of each individual in the population. s1 is used to measure the iterative update progress, thereby adjusting the weighting of global and local searches. s2 measures the diversity of the current population by calculating the overall population variance. i Let ξ represent the environment state in the i-th iteration, and ξ represent the total number of individuals in the population. The iterative fitness itself has a certain degree of randomness; therefore, in reinforcement learning, the current iteration update and population diversity are used instead of population fitness to measure the iterative state of the population.
[0153] (2) Action space
[0154] The action space consists of three consecutive real numbers, A1, A2, and A3, representing the weights of the three wolves α, β, and δ during population update, respectively. The sum of A1, A2, and A3 is 1. The population update formula of the Gray Wolf Algorithm can be expressed in the following form:
[0155] X′=A1·X1+A2·X2+A3·X3 (57)
[0156] A1 + A2 + A3 = 1 (58)
[0157] In the formula, X' represents the updated new population.
[0158] (3) Rewards and Returns
[0159] The reward R consists of a single continuous real number. This embodiment uses reinforcement learning to optimize the Gray Wolf algorithm so that it can adaptively adjust its update strategy based on environmental changes, balance the exploration-development phase, achieve a better objective function, and improve convergence speed. Therefore, the reward function is set as follows, taking into account both changes in population diversity and changes in optimal fitness.
[0160] R = (Var(c) t+1 )-Var(c t )) / Var(c t )+(min(c t )-min(c t+1 )) / min(c t (59)
[0161] In the formula, c t and c t+1 Let represent the sets of fitness of all individuals in the population at iterations t and t+1, respectively.
[0162] Step 4, Obtaining the optimal scheduling scheme
[0163] The specific scheduling scheme is obtained based on the algorithm described in step 3. The scheduling scheme includes the berthing position of the ship, the number of quay cranes, the number of container trucks, the operation priority, and the speed of the ship entering and leaving the port.
[0164] Example Experiment:
[0165] 1. Example Data
[0166] Based on experimental data from relevant literature, and combined with AIS and Clarkson data, two experimental datasets of different sizes were constructed. The port entrance channel is a two-way channel, and the safe draft of vessels for safe navigation at different times is as follows: Figure 1 As shown, in this embodiment, the time window for tidal level changes is set to 0.5 hours using an interpolation method; the width of the two-way channel studied is 200 meters, the ship speed range for entering the port is 5-15 nm / h, and the auxiliary engine load rate is approximately 20% when ships are waiting; this embodiment considers the emission of pollutants during berth scheduling, and the emission factors of the pollutants are shown in Table 4. The treatment costs for NOx, SOx, CO2, and PM are RMB 74.55 / kg, RMB 86.005 / kg, RMB 0.20232 / kg, and RMB 536.25 / kg, respectively.
[18] Because quay crane unloading directly transfers containers from the ship to the terminal, while quay crane loading needs to consider container sorting and space constraints, the loading efficiency of quay cranes is slightly slower than their unloading efficiency. A quay crane's individual loading efficiency is 32 TEU / h, and its individual unloading efficiency is 35 TEU / h. When multiple quay cranes operate simultaneously on the same ship, the efficiency decreases by 0.9%, and the operating power of the quay crane is 120 kWh. Yard crane operations have a fixed efficiency, with trucks spending an average of 2 minutes in the yard.
[0167] Table 4. Statistical Table of Pollutant Emission Factors
[0168]
[0169] The relevant data for the port basin and berths in the first and second sets of calculation examples are shown in Tables 5 and 6, respectively. The static data of the vessels are shown in Tables 7 and 8, and the relevant data for berth scheduling are shown in Tables 9 and 10, respectively.
[0170] Table 5 shows the relevant data for each berth in the first set of examples.
[0171]
[0172] Table 6 shows the relevant data for each berth in the second set of examples.
[0173]
[0174] Table 7. Static data of ships in the first set of examples.
[0175]
[0176] Table 8. Static data of ships in the second set of examples.
[0177]
[0178] Table 9. Ship berth scheduling data for the first set of examples.
[0179]
[0180] Table 10: Ship berth scheduling data for the second set of examples
[0181]
[0182] 2. Calculation Results
[0183] Based on the example data, the DCRLQ-GWO algorithm was applied to solve the problem. The maximum number of iterations was set to 300. If the optimal value remained unchanged for 50 consecutive iterations, the algorithm had converged early and the iteration stopped. Tables 11 and 13 show the optimal encoding schemes for the first and second sets of examples obtained through iteration, respectively. Based on the original encoding schemes in Tables 11 and 13, the decoding process was applied to obtain the specific scheduling schemes shown in Tables 12 and 14. The corresponding Gantt charts were then plotted, as shown below. Figure 2 and Figure 3 As shown. The specific scheduling plan includes the operational details of each vessel from arrival to departure, measured in minutes. Due to operational conflicts and tidal influences, some vessels may need to wait at anchor for a period of time after arrival before entering the port for operations. Furthermore, due to limited resources such as quay cranes, the number of vessels involved in loading and unloading operations may be adjusted during the operation. Figure 6 As shown in Ships 7 and 9, since Ship 9 has a higher operation priority, when the quay crane resources are limited, the operation of Ship 9 needs to be met first. After its operation is completed, the quay crane of Ship 9 is moved to Ship 7 for operation. In addition, since the quay crane movement takes time, the moved quay crane officially participates in the operation of Ship 6 30 minutes after Ship 5 finishes its operation.
[0184] Table 11 Optimal coding scheme for the first set of examples
[0185]
[0186] Table 12 Specific scheduling scheme for the first set of examples
[0187]
[0188] Table 13 shows the optimal coding scheme for the second set of examples.
[0189]
[0190] Table 14 Specific scheduling schemes for the second set of examples
[0191]
[0192] Along with the aforementioned scheduling scheme, the corresponding objective functions were also obtained. The overall objective function for the first set of examples is 40567.46, where the total ship port time F1 is 19732 and the total pollutant emission cost F2 is 208.3546. The overall objective function for the second set of examples is 157267.108, with F1 at 68034 and F2 at 892.3311. Furthermore, based on the scheduling plan, specific data on the usage of quay cranes and container trucks during the entire scheduling planning period can be obtained. Figure 4 Figures (a) and (b) respectively demonstrate the usage of quay cranes and container trucks in the first set of examples. Figure 5 Figures (a) and (b) illustrate the usage of quay cranes and container trucks in the second set of examples. As ships arrive at port gradually over time, the coordinated operation of quay cranes can interfere with each other and reduce operational efficiency. Furthermore, high-intensity quay crane operations generate a large amount of polluting gases. Therefore, the efficiency of quay cranes is not simply linearly related to overall operational efficiency, but the utilization rate of quay cranes can reflect overall operational efficiency to some extent. Typically, quay cranes are a scarce resource at the terminal, while container trucks are generally plentiful. The scheduling results in this embodiment also reflect this. Figure 4 As can be seen from (b), the maximum number of trucks operating simultaneously during the entire scheduling cycle is less than the number of trucks available at the terminal. Based on the model and solution results of this embodiment, the terminal can statistically analyze the number of trucks required for terminal operations over a longer period of time. With a certain number of spare trucks in reserve, the number of standby trucks at the terminal can be appropriately increased or decreased, thereby further optimizing operating costs while ensuring efficient terminal operations.
[0193] 3. Comparative Analysis of Different Algorithms
[0194] To further analyze the effectiveness of the DCRLQ-GWO algorithm proposed in this embodiment, seven algorithms, including CRLQ-GWO, LQ-GWO, traditional GWO, Particle Swarm Optimization (PSO), Genetic Algorithm (GA), Arithmetic Optimization Algorithm (AOA), and Whale Optimization Algorithm (WOA), were used as comparative algorithms in a comparative experiment. The parameters of the comparative algorithms were the same as those of the DCRLQ-GWO algorithm: a population size of 50, a maximum number of iterations of 300, and convergence was defined as no change in fitness after 50 consecutive iterations. The comparative experiment was conducted based on two sets of case data, with 10 experiments performed on each set. Tables 15 and 16 respectively summarize the optimal fitness values of the experimental results for the two sets of cases and plot the results as shown in the figure. Figure 6 and Figure 7 The iterative convergence curve.
[0195] Table 15 Optimal Fitness Statistics for the First Group of Examples
[0196]
[0197] Table 16: Statistical Table of Optimal Fitness for the Second Group of Examples
[0198]
[0199] As shown in Tables 15 and 16, the innovative DCRLQ-GWO algorithm in this embodiment exhibits significant advantages in all test scenarios. In the first set of low-complexity examples, its average optimal value is reduced to 39887.43, an improvement of 35.65% compared to the traditional GWO algorithm. It also maintains a stable advantage in the more complex second set of examples, with an average optimal value of 156094.46, an improvement of 37.53% compared to GWO. The convergence curves show that the DCRLQ-GWO algorithm has limited convergence speed in the early stages of iteration, as reinforcement learning requires a certain amount of time and the algorithm tends to explore the solution space. However, after approximately 30 iterations, the DCRLQ-GWO algorithm outperforms other algorithms in both sets of examples, indicating that the DCRLQ-GWO algorithm has good iterative efficiency. In terms of convergence efficiency, PSO, GA, and the traditional GWO algorithm have extremely high convergence speeds, achieving convergence in an average of 50 iterations. However, the convergence results show that this is mainly due to insufficient global search capability and the ability to escape local optima, leading to premature convergence due to local optima. In contrast, although the DCRLQ-GWO algorithm has a slightly slower convergence speed, its optimal value at the same number of iterations indicates that the DCRLQ-GWO algorithm has better iterative performance, better global and local search capabilities, and is less prone to getting trapped in local optima. Compared with the other seven algorithms, the DCRLQ-GWO algorithm in this embodiment can better solve the complex joint scheduling optimization model constructed in this embodiment. The above results fully demonstrate that the improved method proposed in this embodiment effectively alleviates the defect of the traditional GWO algorithm being prone to getting trapped in local optima, and enhances the algorithm's search capability and convergence accuracy.
[0200] 4. Algorithm Stability Analysis
[0201] To further analyze the iterative stability of the algorithm, the convergence results of 10 sets of experimental cases were statistically analyzed. Tables 17 and 18 respectively show the convergence values and convergence times of different algorithms in the two sets of experiments. According to the statistics, the DCRLQ-GWO algorithm exhibits good stability in both sets of cases. DCRLQ-GWO has the lowest average convergence value and minimum convergence value, while maintaining a low standard deviation of convergence values, indicating that the algorithm can solve problems stably and efficiently. Although the average number of iterations for DCRLQ-GWO is higher than other algorithms, combined with... Figure 6 and Figure 7 It can be seen that the algorithm has better continuous optimization ability when exploring the solution space and achieves better results with the same number of iterations. The slow convergence speed is mainly due to the algorithm's strong global and local search capabilities. This also shows that the introduced reinforcement learning algorithm effectively enhances the diversity and search ability of the Grey Wolf Optimization Algorithm, and exhibits higher efficiency and stability when solving complex optimization problems.
[0202] Table 17. Statistics of convergence values and convergence times for different algorithms in the first group of examples.
[0203]
[0204] Table 18. Statistics of convergence values and convergence times for different algorithms in the second group of examples.
[0205]
[0206] To verify the effectiveness of the proposed method of solving different scheduling schemes based on adjusting the weights of w1 and w2 in the objective function, this embodiment designs three different scheduling strategies. w1 and w2 correspond to the total ship time in port and the total cost of treating pollutant emissions, respectively. Scheme 1: w1 = 0.5, w2 = 1.5, this scheme focuses on reducing the cost of treating pollutant emissions. Scheme 2: w1 = 1, w2 = 1, this scheme gives equal importance to both ship time in port and the cost of treating pollutant emissions. Scheme 3: w1 = 1.5, w2 = 0.5, this scheme focuses on reducing the ship time in port. The algorithm for each strategy is iterated 10 times and the average value is taken. Figure 8 The statistical graphs (a) and (b) show the average results of the objective function and sub-objective functions F1 and F2 under the two schemes.
[0207] Experimental results show that adjusting the weights of w1 and w2 in the objective function can effectively guide the algorithm to optimize different sub-objectives. Strategy 1 significantly reduces the cost of pollutant emissions by increasing w2, but the time ships spend in port increases significantly; Strategy 3 minimizes the time ships spend in port by increasing w1, but the pollution cost increases significantly; Strategy 2, as a balancing scheme, achieved the optimal or near-optimal overall objective function value in both sets of examples, indicating that weight allocation can effectively guide the algorithm to achieve controllable adjustment between time cost and environmental benefits, verifying the effectiveness of the multi-objective scheduling strategy design based on weight adjustment.
[0208] in conclusion:
[0209] This application focuses on the integrated scheduling system of container terminal channels, berths, quay cranes, and container trucks, balancing operational efficiency with environmental protection costs. The main contents and innovative points are as follows:
[0210] I. A joint scheduling optimization model considering the distribution of multiple port basins, continuous berths, and two-way channels was constructed, incorporating tidal depth variations into the constraints to achieve accurate modeling of resource scheduling under complex environments. A full-process scheduling scheme acquisition process was proposed, breaking through the limitations of traditional segmented scheduling and promoting the overall coordination and optimized allocation of port resources.
[0211] II. A DCRLQ-GWO solution algorithm was designed. By using reinforcement learning to initialize the wolf pack weights of the Grey Wolf algorithm with Circle chaotic mapping, reverse learning, quantum potential well search, and Lévy flight strategy, the algorithm enhances population diversity, expands the search space, and improves global and local search capabilities. Experiments show that the fitness of this algorithm is 35.65% and 37.53% higher than that of the traditional GWO algorithm, respectively, and it has significant advantages in convergence speed, accuracy, and stability.
[0212] Third, establish a dual-objective optimization model for pollutant gas treatment costs and ship waiting time, incorporate emission issues into the scheduling objectives, and achieve flexible switching of scheduling strategies by adjusting the weights of the objective function, providing quantitative decision support for green scheduling.
[0213] Finally, it should be noted that the above is only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention (such as the application of various formulas, the order of steps, etc.) without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-resource collaborative scheduling optimization method for container terminals based on the DCRLQ-GWO algorithm, characterized in that, Includes the following steps: Step 1, Estimation of pollutant gas emissions Based on the activity-based calculation method, the emission of a certain pollutant gas is estimated by multiplying the emission factor of the pollutant gas by the energy consumption generated by the activity process. Specifically, this includes: Pollutant emissions from ship navigation E 1,z Pollutant emissions from ships waiting 2,z Pollutant emissions from quay crane operations 3,z And pollutant emissions from truck operations E 4,z ; Step 2, Model Building S21, Make model assumptions, including that the expected arrival time of the ship is known, the channel depth is calculated according to the tide table, the port access channel is a two-way channel, and each ship is allowed to berth once; S22, Establish the objective function and constraints. The specific objective function is as follows: F=min(ω1·k1·F1+ω2·k2·F2) In the formula, F is the objective function, ω1 and ω2 are weight adjustment factors, k1 and k2 are quantity balance coefficients, F1 is the total time ships spend in port, v is the number of ships, and T is the weight adjustment factor. F,i It is the time when ship i departs from the port, T A,i F2 is the estimated arrival time of vessel i, F2 is the total cost of treating pollutant emissions, and p z E is the unit treatment cost of the z-th pollutant. 1,z It refers to the amount of polluting gases emitted by ships during navigation, E 2,z It is the emission of pollutants from ships while they are waiting, E 3,z It is the emission of pollutants from quay crane operations, E 4,z It is the emission of polluting gases from container truck operations; The constraints include ship constraints, quay crane constraints, container truck constraints, and waterway constraints; Step 3: Solve the model using the DCRLQ-GWO algorithm. The DCRLQ-GWO algorithm is designed to solve the model. The algorithm includes: introducing a chaotic initialization strategy, a reverse learning mechanism, a quantum potential well search mechanism, a Lévy flight strategy, and a reinforcement learning algorithm on the basis of the traditional Grey Wolf algorithm. Step 4, Obtaining the optimal scheduling scheme The specific scheduling scheme is obtained based on the algorithm described in step 3. The scheduling scheme includes the berthing position of the ship, the number of quay cranes, the number of container trucks, the operation priority, and the speed of the ship entering and leaving the port.
2. The method according to claim 1, characterized in that, In step 3, the chaotic initialization strategy is implemented through the following iterative formula: In the formula, x t+1 Let x be the population array in the (t+1)th iteration based on the chaotic initialization strategy. t Let x0 be the population array for the t-th iteration based on the chaotic initialization strategy, where α and b are control parameters, mod is the modulo function, and the initial population array x0 is a set of randomly generated numbers.
3. The method according to claim 1, characterized in that, In step 3, the population update formula for the reverse learning mechanism is: X i '=X p +(O i ·d p )×o1,i=1,2,3,p=α,β,δ O i =2(2-2·t / t max )·o2-(2-2·t / t max ),i=1,2,3 d p =|2o3·X P -X|,p=α、β、δ In the formula, X' new X1' is the new population of the gray wolf algorithm obtained iteratively based on the inverse learning mechanism, calculated using the inverse learning mechanism based on the α wolf position; X'2 is the gray wolf algorithm population calculated using the inverse learning mechanism based on the β wolf position; and X3' is the gray wolf algorithm population calculated using the inverse learning mechanism based on the δ wolf position. α Let X be the position of the α wolf. β For the position of the β wolf, X δ Let δ represent the wolf position, and O1, O2, and O3 be adjustment variables for the disturbance caused by the uncertainty of the simulated gray wolf algorithm. α d β d δ o1, o2, and o3 represent the distances between the α wolf, β wolf, and δ wolf and their prey, respectively, and are all random numbers between [0,1].
4. The method according to claim 1, characterized in that, In step 3, a quantum potential well search mechanism is introduced to simulate the particle behavior in the quantum potential well. A wave function is introduced to represent the particle state, and each dimension of the coordinate P... i、j By establishing a one-dimensional delta potential well, the wave function of the particle in any dimension is obtained as follows: In the formula, X represents the particle position, P represents the center of the potential well, σ represents the characteristic length of the δ-potential well, t is the iteration number, and u i,j (t) represents a random number uniformly distributed in [0,1], U(0,1) represents a random number uniformly distributed between 0 and 1, and "±" is determined by the random number. When the random number is greater than 0.5, "+" is used, and "-" is used in other cases. After establishing the basic evolutionary model of particles in quantum space, a reasonable σ is set. i,j To ensure algorithm convergence, we introduce the average optimal position Z(t), which is the average of the optimal positions of all particles. This yields the following equation for the evolution of particle positions in N dimensions: In the formula, ψ(t) is the expansion-contraction coefficient, and N is the total number of particles; The feasible region of the solution using the gray wolf algorithm is considered as a quantum space, and each individual gray wolf is considered as a quantum particle. The positions X of α, β, and δ wolves are defined. α X β X δ Treating the individual as a potential well point and representing the individual state using wave functions, we obtain the positional evolution equations for α, β, and δ wolves: In the formula, X i,j Let Z represent the value of the j-th dimension of the i-th wolf, where i = α, β, δ. i,j (t) represents the average optimal value of the i-th wolf in the j-th dimension, u i,j ψ(t) represents a random number uniformly distributed in [0,1], and ψ(t) is the expansion-contraction coefficient, whose value decreases linearly from m to n with each iteration.
5. The method according to claim 1, characterized in that, In step 3, the Levy flight strategy is introduced into the gray wolf algorithm, resulting in the following gray wolf population update formula: In the formula, X new The new population is obtained through the iteration of the gray wolf algorithm. X1, X2, and X3 are the gray wolf algorithm populations calculated based on the positions of α wolf, β wolf, and δ wolf, respectively, and θ is a random number between [0,2]. Let represent point-to-point multiplication, and levy(θ) represent a path following a Lévy distribution. As a weight to control the step size.
6. The method according to claim 1, characterized in that, In step 3, the Grey Wolf algorithm is optimized using the DDPG algorithm from reinforcement learning algorithms, specifically including: The interaction process between DDPG and the environment is represented by the following quadruple.<S,A,R,γ> Where S represents the state, A represents the action, R is the reward, and γ is the discount factor for future rewards; State S can be represented in the following form: S i =[s1,s2] In the formula, Var represents the standard deviation of the array. x represents the average fitness value of all individuals in the population. i Let be the fitness value of each individual in the population. s1 is used to measure the iterative update progress, thereby adjusting the weighting of global and local searches. s2 measures the diversity of the current population by calculating the overall population variance. i Let ξ be the environmental state in the i-th iteration, and let ξ be the total number of individuals in the population. The action space consists of three consecutive real numbers, A1, A2, and A3, which represent the weights of the three wolves α, β, and δ during population update, respectively. The sum of A1, A2, and A3 is 1. The population update formula of the Gray Wolf Algorithm is expressed in the following form: X′=A1·X1+A2·X2+A3·X3 A1 + A2 + A3 = 1 In the formula, X' represents the updated population, and X1, X2, and X3 represent the gray wolf populations calculated based on the positions of α, β, and δ wolves, respectively. The reward R consists of a single continuous real number. The reward function takes into account both changes in population diversity and changes in optimal fitness, and is set as follows: R=(Var(c t+1 )-Var(c t )) / Var(c t )+(min(c t )-min(c t+1 )) / min(c t ) In the formula, c t and c t+1 Let represent the sets of fitness of all individuals in the population at iterations t and t+1, respectively.
Citation Information
Patent Citations
Grey wolf optimization method based on dimension learning strategy and Levy flight
CN113705761A
Lithium battery model parameter identification method based on chaotic quantum sparrow search algorithm
CN114970332A
Automatic container port multi-agent collaborative optimization operation method
CN115983746A
Optimal dynamic reverse learning strategy-based grey wolf algorithm hybrid optimization method
CN117035002A