Inland river port vehicle transfer and charging scheduling method based on unified information interaction

Through the inland port vehicle transshipment and charging scheduling method with unified information interaction, and utilizing reinforcement learning and Markov decision process models, precise matching of tasks and vehicles and charging scheduling optimization are achieved, solving the problem of information islands in inland ports and improving operational efficiency and battery life.

CN120764896APending Publication Date: 2025-10-10STATE GRID JIANGSU ELECTRIC POWER CO LTD CHANGZHOU BRANCH +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510809453.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Inland ports lack unified information interaction in vehicle transshipment and charging scheduling, resulting in unreasonable task allocation, lack of systematic and optimized charging scheduling, affecting operational efficiency and accelerating battery aging.

Method used

A vehicle transshipment and charging scheduling method for inland ports based on unified information interaction is adopted. Tasks and vehicles are matched through a reinforcement learning framework, a Markov decision process model is constructed, the optimal scheduling plan is generated, and a unified information interaction plan is designed to achieve coordinated optimization of tasks, vehicles and batteries.

Benefits of technology

It improves the accuracy and efficiency of vehicle transfer and charging scheduling, reduces charging costs, extends battery life, and improves port operating revenue and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764896A_ABST
    Figure CN120764896A_ABST
Patent Text Reader

Abstract

The invention relates to the field of internet big data and port energy scheduling, in particular to an inland port vehicle transfer and charging scheduling method based on unified information interaction, which comprises the following steps: acquiring tasks, vehicles and battery information of an inland port; task and vehicle matching is carried out through an integer programming-criticer solution framework based on reinforcement learning; constructing a Markov decision process model which aims at maximizing the income of completing all tasks, minimizing the charging cost and minimizing the battery capacity aging rate; solving the Markov decision process model through a deep reinforcement learning algorithm; information interaction requirements are included in the target task allocation scheme and the optimal transfer and charging scheduling scheme, and tasks of the inland river port are completed; and executing an information unified interaction scheme to realize information interaction among tasks, vehicles and batteries. According to the invention, the accuracy and efficiency of inland port vehicle transfer and charging scheduling can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of Internet big data and port energy scheduling, and in particular to a method for scheduling vehicle transshipment and charging in inland ports based on unified information interaction. Background Art

[0002] As important hubs for both water and land transportation, inland river ports play a crucial role in regional economic development. They serve not only as key nodes for cargo distribution and transit, but also as bridges connecting inland areas with coastal and international markets. With the continued development of the global economy and the continuous growth of trade, the cargo throughput of inland river ports is increasing, and the complexity and importance of their operational management are becoming increasingly prominent.

[0003] The transshipment tasks of inland river ports are diverse and complex. On the one hand, they handle the loading, unloading, storage, and transshipment of a wide variety of cargoes, each with its own distinct characteristics and demands for different transshipment equipment and operational procedures. On the other hand, ports must also rationally arrange vehicle routes and schedules based on cargo flow and transportation needs to ensure timely and accurate delivery of cargo to its destination.

[0004] With the rise of electric vehicles, inland ports are also gradually adopting electric transfer vehicles. These vehicles carry batteries and require recharging at charging stations. Therefore, inland port scheduling requires not only task allocation but also the planning and scheduling of vehicle transfer and charging. Currently, inland ports lack comprehensive consideration of task, vehicle, and battery information when allocating transfer tasks, resulting in irrational task allocation. Furthermore, existing charging scheduling methods lack systematicity and optimization, often charging vehicles only when their battery is low, without fully considering factors such as charging costs and battery capacity aging. This not only leads to excessively high charging costs but also accelerates battery aging, shortening battery life. Furthermore, most current port information systems suffer from information silos, preventing timely and accurate information sharing and transmission between various parties. The lack of effective information exchange mechanisms hinders the smooth flow of information between vehicles, batteries, and tasks, hindering the coordinated optimization of charging scheduling with task allocation and vehicle transfer, impacting the overall operational efficiency of ports.

[0005] Therefore, how to design an effective method to realize vehicle transshipment and charging scheduling in inland ports is a technical problem that needs to be solved urgently. Summary of the Invention

[0006] In view of the shortcomings of the above-mentioned existing technologies, the technical problem to be solved by the present invention is: how to provide an inland port vehicle transshipment and charging scheduling method based on unified information interaction, first obtaining the inland port task, vehicle and battery information, and matching the tasks and vehicles through a reinforcement learning framework to obtain the optimal allocation plan; then constructing a Markov decision process model, and generating the optimal scheduling plan through a deep reinforcement learning algorithm; at the same time, incorporating information interaction requirements, designing and implementing a unified interaction plan, thereby improving the accuracy and efficiency of inland port vehicle transshipment and charging scheduling.

[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0008] The inland port vehicle transshipment and charging scheduling method based on unified information interaction includes:

[0009] Obtain mission information, vehicle information, and battery information for inland ports;

[0010] Using a reinforcement learning-based integer programming-critic solution framework, we match tasks with vehicles by combining task, vehicle, and battery information to obtain the target task allocation solution. Based on the optimal task allocation solution, we establish a preliminary information interaction relationship between tasks, vehicles, and batteries.

[0011] Construct a Markov decision process model with the goal of maximizing the benefits of completing all tasks, minimizing charging costs, and minimizing battery capacity aging rate;

[0012] Based on the target task allocation plan, a deep reinforcement learning algorithm is used to solve the Markov decision process model that takes task information, vehicle information, and battery information as input to generate the optimal transport and charging scheduling plan;

[0013] Incorporate information exchange requirements into the target task allocation plan and the optimal transshipment and charging scheduling plan to ensure the smooth and secure flow of information between all parties during the scheduling decision-making process; dispatch vehicles to complete the tasks of the inland port based on the target task allocation plan and the optimal transshipment and charging scheduling plan;

[0014] Based on the preliminary information interaction relationship, a unified information interaction plan is designed to ensure the coordinated optimization of task allocation, vehicle charging scheduling and battery status; the unified information interaction plan is executed to realize information interaction between tasks, vehicles and batteries.

[0015] Preferably, the steps for matching tasks with vehicles by combining task information, vehicle information, and battery information through an integer programming-critic solution framework based on reinforcement learning to obtain a target task allocation solution are as follows:

[0016] S201: Define state s, action a and reward r;

[0017] 1) State s includes task information and vehicle information;

[0018] 2) Action a is the matching decision X of task and vehicle m,b,i ; X m,b,i represents whether task m is assigned to vehicle b at position i in the vehicle task list, X m,b,i = 1 represents assignment, X m,b,i = 0 represents no assignment;

[0019] 3) The formula of reward r is represented as:

[0020]

[0021] In the formula: r represents the reward of vehicle completing a task once; represents the income of vehicle completing a task once; represents the cost of power purchase of vehicle completing a task once; represents the battery capacity aging rate of vehicle completing a task once; w grid , w job , w deg represents the corresponding weight coefficient; p b represents the charging power; i represents the discharging times; δ cyc represents the battery attenuation rate;

[0022] S202: input the current state s into the policy improvement model, select the optimal action X m,b,i ; calculate the reward through the multi-objective model

[0023] S203: calculate the corresponding Q value estimate based on the optimal action X m,b,i through the Q function Q(m, b, i);

[0024] S204: update the Q function Q(m, b, i) based on the reward and the Q value estimate;

[0025] The formula is represented as:

[0026]

[0027] In the formula: represents the observed immediate reward when assigning task m at position i in the vehicle task list to vehicle b; α is the learning rate; represents the TD error, which is used to quantify the difference between the reward and the Q value estimate;

[0028] S205: repeat steps S202 to S204, and iteratively update the Q function Q(m, b, i);

[0029] S206: Calculate the loss function based on all the current rewards and Q value estimates, and optimize the parameters of the Q function Q(m, b, i) based on the loss function;

[0030] The formula of the loss function is:

[0031]

[0032] In the formula: Z is the number of samples in each small batch;

[0033] The formula for optimizing the parameters of the Q function Q(m, b, i) is:

[0034]

[0035] In the formula: θ represents the parameters of the Q function to be optimized; lr represents the learning rate of parameter update; represents the gradient of the loss function with respect to θ;

[0036] S207: Repeat steps S202 to S206, and iteratively train the Q function Q(m, b, i) until convergence;

[0037] S208: Input the current state of the vehicle into the trained Q function Q(m, b, i), and output the target task allocation scheme.

[0038] Preferably, the objective function of the policy improvement model is represented as:

[0039]

[0040] The constraint condition is represented as:

[0041]

[0042] Where: the objective function is to maximize the total reward in vehicle-task matching; the first constraint requires each task m to be assigned only once in all vehicles b, task positions i, and actions a, ensuring that each task occupies a unique position in the vehicle's task list; the second constraint condition forces each vehicle b to handle only one task at each task position i, preventing conflicts or overloading in the vehicle's task list; the third constraint condition ensures that each decision variable X m,b,i is a binary variable, ensuring the discrete selection of task allocation.

[0043] Preferably, the preliminary information interaction relationship includes: establishing the basic association relationship between tasks and vehicles and batteries, i.e., the matching relationship between task m and vehicle b and the binding relationship between vehicle b and battery.

[0044] Preferably, the Markov decision process model includes state variables, action space, and reward function:

[0045] 1) State variables:

[0046]

[0047] Where: s b represents the state of battery b; t represents the time step; Indicates the current itinerary; Indicates the next trip; c b Indicates the current SoC; τ b Indicates the remaining time of the current trip;

[0048] 2) Action space, including cargo transfer trips, empty load dispatch trips, trips to charging stations, and charging trips;

[0049] 2.1) Battery operation space during cargo transfer:

[0050]

[0051] Where: represents the action space of battery b in a cargo transfer trip; Indicates the speed of the next trip; Indicates the start time of the next trip; Indicates the trip after the next trip; Respectively represent the lower and upper limits of the speed; Indicate itinerary The set of next potential trips; Indicates cargo transfer trips, empty load dispatch trips, and trips to charging stations;

[0052] 2.2) Battery operation space during no-load dispatching trip:

[0053]

[0054] Where: represents the action space of battery b in an unloaded dispatch trip;

[0055] 2.3) The action space of the battery during the trip to the charging station:

[0056]

[0057] Where: represents the action space of battery b during a trip to the charging station; Indicates the charging rate of the charging trip; Indicates the lower and upper limits of the charging rate;

[0058] 2.4) Battery operation space during charging process:

[0059]

[0060] wherein: represents the action space of battery b in one charging trip;

[0061] 3) reward function:

[0062]

[0063] wherein: r represents the reward of the vehicle completing one task; represents the revenue of the vehicle completing one task; represents the cost of power purchase of the vehicle completing one task; represents the battery capacity aging rate of the vehicle completing one task; w grid , w job , w deg represents the corresponding weight coefficient; p b represents the charging power; i represents the discharging times; δ cyc represents the battery decay rate.

[0064] Preferably, the state transition process of the Markov decision process model is represented as:

[0065] 1) state transition of the goods transfer trip

[0066] The action of the goods transfer trip is represented as:

[0067]

[0068] The state transition of one goods transfer trip is represented as:

[0069]

[0070] wherein: and represent the energy consumed per unit time in the current trip and the next trip respectively; m 0 , v 0 represent the vehicle weight and speed of executing the current trip respectively; m 1 , v 1 represent the vehicle weight and speed of executing the next trip respectively; E represents the energy capacity of the battery; represents the distance of the next trip κ 1 ;

[0071] 2) state transition of the empty scheduling trip The action of the empty scheduling trip is represented as:

[0072] The state transition of the empty dispatch trip is represented as:

[0073] 3) The state transition of the de-charging station trip is represented as:

[0074] The state transition of the de-charging station trip:

[0075] 4) The state transition of the charging trip is represented as:

[0076]

[0077] The state transition of the charging trip is represented as:

[0078]

[0079] Preferably, the battery capacity aging rate is calculated by constructing a battery aging quantification model.

[0080] The calculation formula of the battery aging quantification model is:

[0081]

[0082] In the formula, δ deg represents the battery capacity aging rate under the non-design cycle condition; N seq represents the number of non-design discharges; δ cyc represents the cycle aging capacity attenuation rate of the i-th non-design discharge; Δδ represents the capacity attenuation error under the non-design cycle condition, which is a correction term for describing the influence of small-scale discharge depth (DoD) on battery capacity degradation;

[0083] wherein:

[0084]

[0085] In the formula, β is a constant parameter; q represents the order of the battery aging quantification model; DoD q represents the discharge depth (DoD) value of the q-th order; represents the discharge depth value of the i-th non-design discharge; represents the total discharge depth value of the entire non-design cycle condition; DoD tot represents the change of SOC from the end of charging to the start of the next charging, regardless of how many discharges are experienced in between; N seq represents the number of discharges.

[0086] Preferably, the deep reinforcement learning algorithm is an Actor-Critic algorithm.

[0087] The processing steps of the Actor-Critic algorithm include:

[0088] S401: Initialize the strategy (Actor) π of the Actor-Critic algorithm θ (a|s) and value function (Critic); where the policy refers to the mapping from state to corresponding action;

[0089] S402: By strategy π θ (a|s) selects an action a based on the vehicle’s current state s;

[0090] S403: Execute the current action a and perform state transition through the Markov decision process model to obtain the next state s ′ and the corresponding reward r;

[0091] S404: Estimate the value V of the current state s and the next state s′ through the value network w (s) and V w (s′); and then calculate the temporal difference error (TD Error);

[0092] The formula is:

[0093] δ=r+γV w (s′)-V w (s);

[0094] Where: δ represents the time difference error; γ represents the discount factor; w represents the value function parameter;

[0095] S405: Update the value function parameter w by gradient descent combined with the time difference error δ;

[0096] The formula is:

[0097]

[0098] Where: a w represents the learning rate;

[0099] S406: Update strategy π through policy gradient θ (a|sθ’s policy parameter θ;

[0100] The formula is:

[0101]

[0102] Where: a θ represents the learning rate; represents the policy gradient;

[0103] S407: Repeat steps S402 to S406 until the strategy converges to the optimal solution or reaches a preset number of training rounds;

[0104] S408: Input the current state of the vehicle into the trained strategy π θ In (a|s), output the optimal transshipment scheduling plan.

[0105] Preferably, the information interaction requirements include:

[0106] 1) Task requirements: task ID, cargo type, starting and ending ports, priority, and remaining time;

[0107] 2) Vehicle requirements: vehicle ID, current location, current battery level, and load status;

[0108] 3) Battery requirements: battery ID, capacity decay rate, charging rate, and health status.

[0109] Preferably, the unified information interaction scheme includes: collecting task information, vehicle information and battery information in real time through the central controller of the inland port; using all collected information as input, executing steps S2 to S5 to generate a target task allocation scheme and an optimal transshipment and charging scheduling scheme; issuing scheduling instructions to vehicles, charging stations and batteries based on the target task allocation scheme and the optimal transshipment and charging scheduling scheme; monitoring the status of vehicles, charging stations and batteries after executing the scheduling instructions.

[0110] Compared with the prior art, the inland port vehicle transshipment and charging scheduling method based on unified information interaction in the present invention has the following beneficial effects:

[0111] This invention utilizes a reinforcement learning-based integer programming-critic solution framework to match tasks with vehicles, combining task information, vehicle information, and battery information to generate a target task allocation solution. First, this framework comprehensively considers the complex task requirements of inland ports, the diverse performance of vehicles, and the varying battery states, achieving precise matching between tasks and vehicles. Compared to traditional simple allocation methods, this effectively improves the accuracy and rationality of task allocation, thereby enhancing the accuracy of inland port vehicle transshipment and charging scheduling. Second, integer programming enables quantitative allocation of vehicle resources, ensuring that every vehicle is fully utilized and preventing idle or overused vehicles. Simultaneously, the critic solution framework continuously learns and optimizes, adjusting allocation strategies to improve resource utilization and enhance the efficiency of inland port vehicle transshipment and charging scheduling. Finally, the reinforcement learning mechanism enables the entire framework to perceive changes in the inland port environment in real time and dynamically adjust task allocation solutions. This adaptive capability enables the port vehicle scheduling system to respond to various emergencies, maintain efficient and stable operation, and enhance the flexibility of inland port vehicle transshipment and charging scheduling.

[0112] Based on task allocation, the present invention constructs a Markov decision process model with the goal of maximizing the benefits of completing all tasks, minimizing the charging cost, and minimizing the battery capacity aging rate, and solves the model through a deep reinforcement learning algorithm to generate the optimal transshipment and charging scheduling plan. First, with the goal of maximizing the benefits of completing all tasks, it is possible to reasonably arrange the vehicle transshipment routes and times, improve cargo transportation efficiency, increase the cargo throughput of the port, and thus directly improve the operating income of the port; at the same time, by optimizing the scheduling plan, the idle time and empty mileage of the vehicle are reduced, and the cost of vehicle transshipment and charging scheduling at the inland port is reduced. Secondly, with the goal of minimizing the charging cost of completing all tasks, the model can formulate the optimal charging policy based on factors such as the vehicle's remaining power, battery charging characteristics, and fluctuations in grid electricity prices, thereby reasonably arranging the charging time period and charging costs. Then, with the goal of achieving the lowest battery capacity aging rate for all tasks, the model optimizes the vehicle's charging and discharging processes, avoiding adverse operations such as overcharging, over-discharging, and high-temperature charging. This reduces battery capacity decay and extends battery life. This not only lowers battery replacement costs but also reduces vehicle downtime due to battery failure, improving vehicle reliability and availability. Finally, deep reinforcement learning algorithms, capable of processing complex Markov decision process models, generate optimal transshipment and charging scheduling plans by learning and analyzing large amounts of historical data and real-time information, thereby improving the quality and efficiency of vehicle transshipment and charging scheduling at inland ports.

[0113] The present invention incorporates information interaction requirements into the target task allocation plan and the optimal transfer and charging scheduling plan to ensure the smoothness of information flow among all parties, enabling real-time and accurate information sharing between various links such as task allocation, vehicle scheduling and battery management, improving scheduling coordination, and avoiding scheduling confusion and conflicts caused by poor information flow.

[0114] The present invention designs a unified information interaction scheme based on the preliminary information interaction relationship, and executes the unified information interaction scheme to realize information interaction between tasks, vehicles and batteries. First, the unified information interaction scheme can organically combine task allocation, vehicle charging scheduling and battery status management to achieve coordinated optimization of the three. Through real-time information interaction, task allocation and charging scheduling schemes can be dynamically adjusted according to the battery status. At the same time, task allocation and charging scheduling will also affect the battery usage and maintenance strategy, thereby achieving the overall optimal effect. Secondly, through the information interaction between tasks, vehicles and batteries, each other's status and needs can be understood in a timely manner, the intermediate links of information transmission can be reduced, and operational efficiency can be improved; at the same time, through unified information interaction, the resource status of tasks, vehicles and batteries can be fully grasped, and the optimal allocation of resources can be achieved. Finally, the unified information interaction scheme can monitor the operating status of tasks, vehicles and batteries in real time, discover potential problems and risks in a timely manner, and take corresponding measures to make adjustments to ensure the stable operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] In order to make the purpose, technical solutions and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0116] Figure 1 This is a logical block diagram of the inland port vehicle transshipment and charging scheduling method based on unified information interaction.

[0117] Figure 2 Schematic diagram of four types of itineraries. DETAILED DESCRIPTION

[0118] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but only represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0119] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not require further definition or explanation in subsequent figures. In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer" indicate positions or relationships based on the positions or relationships shown in the figures, or the positions or relationships in which the inventive product is typically placed when in use. These terms are intended solely to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation, and are therefore not to be construed as limiting the present invention. Furthermore, the terms "first," "second," and "third," etc., are used solely to distinguish descriptions and are not to be construed as indicating or implying relative importance. Furthermore, terms such as "horizontal" and "vertical" do not imply that a component is absolutely horizontal or overhanging, but rather may be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but rather may be slightly tilted. In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0120] The following is a further detailed description through specific implementation methods:

[0121] Example 1:

[0122] This embodiment discloses a method for inland port vehicle transshipment and charging scheduling based on unified information interaction.

[0123] like Figure 1 As shown in FIG, the inland port vehicle transshipment and charging scheduling method based on unified information interaction includes:

[0124] Obtain mission information, vehicle information, and battery information for inland ports;

[0125] In this embodiment, task information includes task ID, cargo type, starting and ending port coordinates, task priority (based on deadline and revenue), estimated duration, and task status (e.g., pending or in progress). Vehicle information includes vehicle ID, current location (GPS coordinates), current battery level, maximum payload, and task queue status. Battery information includes battery ID, capacity decay rate, charge rate, historical depth of discharge records, and battery health score (SOH).

[0126] Using a reinforcement learning-based integer programming-critic solution framework, we match tasks with vehicles by combining task, vehicle, and battery information to obtain the target task allocation solution. Based on the optimal task allocation solution, we establish a preliminary information interaction relationship between tasks, vehicles, and batteries.

[0127] In this embodiment, the preliminary information interaction relationship includes: establishing a basic association relationship between the task and the vehicle and the battery, that is, a matching relationship between task m and vehicle b and a binding relationship between vehicle b and the battery.

[0128] Construct a Markov decision process model with the goal of maximizing the benefits of completing all tasks, minimizing charging costs, and minimizing battery capacity aging rate;

[0129] Based on the target task allocation plan, a deep reinforcement learning algorithm is used to solve the Markov decision process model that takes task information, vehicle information, and battery information as input to generate the optimal transport and charging scheduling plan;

[0130] Incorporate information exchange requirements into the target task allocation plan and the optimal transshipment and charging scheduling plan to ensure the smooth and secure flow of information between all parties during the scheduling decision-making process; dispatch vehicles to complete the tasks of the inland port based on the target task allocation plan and the optimal transshipment and charging scheduling plan;

[0131] Based on the preliminary information interaction relationship, a unified information interaction plan is designed to ensure the coordinated optimization of task allocation, vehicle charging scheduling and battery status; the unified information interaction plan is executed to realize information interaction between tasks, vehicles and batteries.

[0132] This invention utilizes a reinforcement learning-based integer programming-critic solution framework to match tasks with vehicles, combining task information, vehicle information, and battery information to generate a target task allocation solution. First, this framework comprehensively considers the complex task requirements of inland ports, the diverse performance of vehicles, and the varying battery states, achieving precise matching between tasks and vehicles. Compared to traditional simple allocation methods, this effectively improves the accuracy and rationality of task allocation, thereby enhancing the accuracy of inland port vehicle transshipment and charging scheduling. Second, integer programming enables quantitative allocation of vehicle resources, ensuring that every vehicle is fully utilized and preventing idle or overused vehicles. Simultaneously, the critic solution framework continuously learns and optimizes, adjusting allocation strategies to improve resource utilization and enhance the efficiency of inland port vehicle transshipment and charging scheduling. Finally, the reinforcement learning mechanism enables the entire framework to perceive changes in the inland port environment in real time and dynamically adjust task allocation solutions. This adaptive capability enables the port vehicle scheduling system to respond to various emergencies, maintain efficient and stable operation, and enhance the flexibility of inland port vehicle transshipment and charging scheduling.

[0133] Based on task allocation, the present invention constructs a Markov decision process model with the goal of maximizing the benefits of completing all tasks, minimizing the charging cost, and minimizing the battery capacity aging rate, and solves the model through a deep reinforcement learning algorithm to generate the optimal transshipment and charging scheduling plan. First, with the goal of maximizing the benefits of completing all tasks, it is possible to reasonably arrange the vehicle transshipment routes and times, improve cargo transportation efficiency, increase the cargo throughput of the port, and thus directly improve the operating income of the port; at the same time, by optimizing the scheduling plan, the idle time and empty mileage of the vehicle are reduced, and the cost of vehicle transshipment and charging scheduling at the inland port is reduced. Secondly, with the goal of minimizing the charging cost of completing all tasks, the model can formulate the optimal charging policy based on factors such as the vehicle's remaining power, battery charging characteristics, and fluctuations in grid electricity prices, thereby reasonably arranging the charging time period and charging costs. Then, with the goal of achieving the lowest battery capacity aging rate for all tasks, the model optimizes the vehicle's charging and discharging processes, avoiding adverse operations such as overcharging, over-discharging, and high-temperature charging. This reduces battery capacity decay and extends battery life. This not only lowers battery replacement costs but also reduces vehicle downtime due to battery failure, improving vehicle reliability and availability. Finally, deep reinforcement learning algorithms, capable of processing complex Markov decision process models, generate optimal transshipment and charging scheduling plans by learning and analyzing large amounts of historical data and real-time information, thereby improving the quality and efficiency of vehicle transshipment and charging scheduling at inland ports.

[0134] The present invention incorporates information interaction requirements into the target task allocation plan and the optimal transfer and charging scheduling plan to ensure the smoothness of information flow among all parties, enabling real-time and accurate information sharing between various links such as task allocation, vehicle scheduling and battery management, improving scheduling coordination, and avoiding scheduling confusion and conflicts caused by poor information flow.

[0135] The present invention designs a unified information interaction scheme based on the preliminary information interaction relationship, and executes the unified information interaction scheme to realize information interaction between tasks, vehicles and batteries. First, the unified information interaction scheme can organically combine task allocation, vehicle charging scheduling and battery status management to achieve coordinated optimization of the three. Through real-time information interaction, task allocation and charging scheduling schemes can be dynamically adjusted according to the battery status. At the same time, task allocation and charging scheduling will also affect the battery usage and maintenance strategy, thereby achieving the overall optimal effect. Secondly, through the information interaction between tasks, vehicles and batteries, each other's status and needs can be understood in a timely manner, the intermediate links of information transmission can be reduced, and operational efficiency can be improved; at the same time, through unified information interaction, the resource status of tasks, vehicles and batteries can be fully grasped, and the optimal allocation of resources can be achieved. Finally, the unified information interaction scheme can monitor the operating status of tasks, vehicles and batteries in real time, discover potential problems and risks in a timely manner, and take corresponding measures to make adjustments to ensure the stable operation of the system.

[0136] In order to better introduce the technical solution of the present invention, this embodiment is described through the following parts.

[0137] 1. Matching Tasks and Vehicles

[0138] In this embodiment, the integer programming-critic solution framework based on reinforcement learning is used to match tasks with vehicles by combining task information, vehicle information, and battery information. The processing steps for obtaining the target task allocation solution are as follows:

[0139] S201: Define state s, action a and reward r;

[0140] 1) The state s includes task information and vehicle information;

[0141] 2) Action a is the matching decision X between the task and the vehicle m,b,i ;X m,b,i Indicates whether task m is assigned to vehicle b at position i in the vehicle task list, X m,b,i =1 means allocation, X m,b,i =0 means no allocation;

[0142] 3) The formula for reward r is:

[0143]

[0144] Where: represents the immediate reward observed when assigning task m to vehicle b at position i in the vehicle's task list; Indicates the profit of the vehicle completing a task; Indicates the cost of electricity purchased by the vehicle to complete a mission; Indicates the battery capacity aging rate of the vehicle after completing a mission; w grid 、w job 、w deg Indicates the corresponding weight coefficient; p b represents the charging power; n represents the number of discharges; δ cyc Indicates the battery attenuation rate;

[0145] S202: Execute integer programming: Input the current state s into the policy improvement model and select the optimal action X m,b,i , and calculate the reward

[0146] S203: Based on the optimal action X through the Q function Q(m,b,i) m,b,i Calculate the corresponding Q value estimate;

[0147] S204: Executing the critic solution: Reward-based And Q value estimate update Q function Q(m,b,i);

[0148] The formula is:

[0149]

[0150] Where: represents the immediate reward observed when assigning task m to vehicle b at position i in the vehicle task list; α is the learning rate, which is used to control the magnitude of each update; represents the TD error, which is used to quantify the difference between the observed reward and the Q-value estimate;

[0151] S205: Repeat steps S202 to S204 to iteratively update the Q function Q(m,b,i);

[0152] S206: Calculate the loss function based on all current rewards and Q values, and reversely optimize the parameters of the Q function Q(m, b, i) based on the loss function;

[0153] The formula of the loss function is expressed as:

[0154]

[0155] Where: Z is the number of samples in each mini-batch;

[0156] The formula for optimizing the parameters of the Q function Q(m,b,i) is expressed as:

[0157]

[0158] Where: θ represents the Q function parameter to be optimized; lr represents the learning rate of parameter update; Represents the gradient of the loss function with respect to θ;

[0159] S207: Repeat steps S202 to S206 to iteratively train the Q function Q(m, b, i) until convergence;

[0160] S208: Input the current state of the vehicle into the trained Q function Q(m,b,i) and output the target task allocation plan.

[0161] Specifically, the objective function of the strategy improvement model is expressed as:

[0162]

[0163] The constraints are expressed as:

[0164]

[0165] The objective function aims to maximize the total reward in vehicle-task matching; the first constraint requires that each task m can only be assigned once among all vehicles b, task positions i, and actions a, ensuring that each task occupies a unique position in the vehicle's task list; the second constraint forces each vehicle b to handle only one task at each task position i, preventing conflicts or overloading in the vehicle's task list; the third constraint ensures that each decision variable X m,b,i It is a binary variable to ensure the discreteness of task assignment.

[0166] 2. Markov Decision Process Model

[0167] In this embodiment, the Markov decision process model includes state variables, action space, and reward function:

[0168] 1) State variables:

[0169]

[0170] Where: s b represents the state of battery b; t represents the time step; Indicates the current itinerary; Indicates the next trip; c b Indicates the current SoC; τ b Indicates the remaining time of the current trip;

[0171] 2) Action space, including cargo transfer trips, empty load dispatch trips, trips to charging stations, and charging trips;

[0172] Each battery operates in four types of trips, namely cargo transfer, no-load dispatch, de-charging and charging trips, e.g. Figure 2As shown (the red line indicates that the next trip is certain). The battery must be running in one of the trips. For the first three types of trips (freight transfer, empty dispatch, and trip to the charging station), the battery is equipped on the transfer vehicle and powers it, while the battery is charged at the station in the charging trip. Specifically, in the freight transfer trip, the battery powers the transfer vehicle that transports the goods. In the empty dispatch trip, the transfer vehicle travels from another node to the loading node of the freight transfer trip, as shown in Figure 2. Figure 2 When the battery is low, the transfer vehicle on the charging trip will drive from the end of the delivery trip to the nearest station for charging, as shown in the red line between "R" and "D". Figure 2 This is shown as the red line between "B" and "C" in the diagram. After the charging journey is complete, the battery leaves the transport vehicle and begins its charging journey.

[0173] 2.1) Each battery b has four types of actions in four types of trips. When battery i operates in a cargo transfer trip, that is, The battery needs to decide the speed and start time of the next trip and the trips after the next one. The action space of the battery in a cargo transfer trip is:

[0174]

[0175] Where: represents the action space of battery b in a cargo transfer trip; Indicates the speed of the next trip; Indicates the start time of the next trip; Indicates the trip after the next trip; Respectively represent the lower and upper limits of the speed; Indicate itinerary The set of next potential trips; Indicates cargo transfer trips, empty load dispatch trips, and trips to charging stations;

[0176] 2.2) In an empty transfer trip, that is The battery needs to decide the speed and start time of the next trip. Figure 2 As shown in Figure 2, the next trip of the no-load transfer trip is a certain transfer trip. The action space of the battery in the no-load dispatch trip is:

[0177]

[0178] Where: represents the action space of battery b in an unloaded dispatch trip;

[0179] 2.3) In one charging trip, that is The battery needs to determine the charging power and start time for the next charging trip and the next trip after that. The action space of the battery during the trip to the charging station is:

[0180]

[0181] Where: represents the action space of battery b during a trip to the charging station; Indicates the charging rate of the charging trip; Indicates the lower and upper limits of the charging rate;

[0182] 2.4) In one charging trip, that is The battery needs to decide the speed and start time of the next relocation trip and the trip after the next one. The action space of the battery operation during the charging trip is:

[0183]

[0184] Where: represents the action space of battery b during a charging trip;

[0185] 3) Reward function:

[0186] Each cell will be rewarded at the end of each time interval. The reward b for each cell consists of completing the transfer task R job The income is used to purchase electricity C grid The cost and battery capacity aging rate δ deg Different types of trips have different reward functions. Given the income from completing the transfer task (R job ), the cost of purchasing electricity (C grid ) and battery capacity aging rate (δ deg ) are significantly different in magnitude, the reward function for battery b is not simply the sum of these three factors. Instead, it is a weighted sum where each factor is multiplied by a specific weight that reflects its relative importance in the overall reward calculation.

[0187]

[0188] Where: r represents the reward for the vehicle to complete a task; Indicates the profit of the vehicle completing a task; Indicates the cost of electricity purchased by the vehicle to complete a mission; Indicates the battery capacity aging rate of the vehicle after completing a mission; w grid 、w job 、w deg Indicates the corresponding weight coefficient; p brepresents the charging power; n represents the number of discharges; δ cyc Indicates the battery degradation rate.

[0189] If the remaining time of the current trip is less than 1, the equation is equal to 1; if the remaining time is greater than or equal to 1, the equation is equal to 0;

[0190] If the current trip is a transfer trip, the equation is equal to 1; if the current trip is not a transfer trip, the equation is equal to 0;

[0191] If the current trip is a trip to charge, the equation is equal to 1; if the current trip is not a trip to charge, the equation is equal to 0;

[0192] If the current trip is a charging trip, the equation is equal to 1; if the current trip is not a charging trip, the equation is equal to 0;

[0193] If the current trip is a de-charging or charging trip, the equation is equal to 1; if the current trip is neither a de-charging nor a charging trip, the equation is equal to 0.

[0194] Different types of trips have different action spaces and state transition functions. The state transition process of the Markov decision process model is expressed as:

[0195] 1) Status transition of cargo transshipment trip

[0196] The cargo transfer trip determines the speed of the next trip The start time of the next trip and the itinerary after the next trip

[0197] The actions of the cargo transfer trip are expressed as:

[0198]

[0199] The state transition of a cargo transfer trip is expressed as:

[0200]

[0201] Current itinerary and next trip Will not change with the remaining time τ b ≥1. If the current trip is about to end, that is, the remaining time τ b , current itinerary and next trip will become and like Figure 2 As shown in Figure 2, a delivery trip can transition to a delivery, idle dispatch, and charging trip. Since subsequent trips consume energy, the SoC will decrease if the current trip is a delivery trip.

[0202] Where: and Represents the energy consumed per unit time for the current trip and the next trip respectively; m 0 ,v 0 Respectively indicate the execution of the current trip Vehicle weight and speed; m 1 ,v 1 Respectively indicate the execution of the next trip The vehicle weight and speed; E represents the energy capacity of the battery; Indicates the next trip κ 1 distance;

[0203] 2) State transition of no-load dispatching trip

[0204] Repositioning the trip will determine the speed of the next trip The start time of the next trip and the trip after the next trip

[0205] The action of the no-load dispatch trip is expressed as:

[0206]

[0207] The state transition of the no-load dispatch trip is expressed as:

[0208]

[0209] The state transition function between the current trip and the next trip is the same as that of the cargo transfer trip. Figure 2 As shown in Figure 2, an empty-load dispatch trip can only be converted into a cargo transfer trip, which consumes energy. The state transition function of the current SoC and the remaining time of the current trip is the same as that of the cargo transfer trip.

[0210] 3) State transition for the trip to the charging station

[0211] like Figure 2 As shown, the next trip of the charging trip is the charging trip, and the charging rate of the charging trip is Charging start time of the charging trip and the trip after the charging trip Make decisions.

[0212] The action of traveling to the charging station is represented as:

[0213]

[0214] State transitions for the trip to the charging station:

[0215]

[0216] The state transition function between the current trip and the next trip is the same as that of the cargo transfer trip. Figure 2 As shown, the charging trip can only be converted to the charging trip, and the energy increases during charging. Therefore, the battery SoC first decreases during the charging trip and then increases during the charging trip. The increase in SoC is determined by the charging rate. The charging time is determined by the battery capacity E. The remaining time for the next charging trip is determined by the amount of battery to be replenished. Charging rate and battery charging time Decide.

[0217] 4) State transition of charging trip

[0218] Since the next charging trip is a relocation trip, Figure 2 As shown, the charging trip speed to the next trip The start time of the next trip and the itinerary after the next trip Make decisions.

[0219] The action of the charging trip is expressed as:

[0220]

[0221] The state transition of the charging trip is expressed as:

[0222]

[0223] The state transition function between the trip and the next trip is the same as that of the delivery trip. Figure 2 As shown in , the charging trip can only be converted into a repositioning trip, which consumes energy. Therefore, the battery SoC first increases in the charging trip and then decreases in the repositioning trip. b = 1, the change of SoC is determined by the charging rate Charging time τ b , the energy consumed in the next trip and battery capacity E, while the remaining time of the next trip is determined by the total trip time and the time the vehicle has been traveling Decide.

[0224] 3. Battery Aging Quantification Model

[0225] In this embodiment, the battery capacity aging rate is calculated by constructing a battery aging quantification model;

[0226] The calculation formula of the battery aging quantitative model is:

[0227]

[0228] Where: δ deg Indicates the battery capacity aging rate under non-design cycle conditions; N seq Indicates the non-design discharge times of non-design cycle conditions; represents the cycle aging capacity attenuation rate of the nth non-design discharge; Δδ represents the capacity attenuation error of the non-design cycle condition, which is a correction term used to describe the effect of small-scale depth of discharge (DoD) on battery capacity degradation;

[0229] in:

[0230]

[0231] Where: is a constant parameter; q represents the order of the battery aging quantization model; DoD q Indicates the depth of discharge (DoD) value of the qth stage; Indicates the discharge depth value of the nth non-design discharge; Indicates the total depth of discharge value of the entire non-design cycle condition; DoD tot Indicates the change in SOC from the end of charging to the start of the next charge, regardless of how many times the discharge has occurred in the middle; N seq Indicates the total number of discharges.

[0232] Specifically, the correction term of the battery aging quantification model is updated as follows:

[0233]

[0234] Where: DoD j Indicates non-design cycle conditions The depth of discharge value of the jth discharge behavior (the process or method of releasing stored energy during use) under the condition; DoD tot Indicates the design cycle condition ( The battery is specified in a laboratory environment for a single cycle fixed discharge depth, such as the discharge depth value under the condition of each discharge from 1 to 0.2); i represents the nth coefficient;

[0235]

[0236] Where: c b,trepresents the SoC of battery b at time t; SoC represents the SoC of battery b at time t+1; Indicates the current itinerary; Indicates the next trip; τ b Indicates the remaining time of the current trip; Indicates the start time of the next trip; Indicates the speed of the next trip; E represents the energy capacity of the battery; p b represents the charging rate of battery b; It is a function whose input includes parameters such as current trip, next trip, speed, etc., and its output is the change in battery state of charge; represents the set of D (transfer trip), R (dispatch trip), and B (de-charging trip), and k belongs to one of these three trip types; C represents the charging trip. If k belongs to the charging trip, 1_piC=1; if k does not belong to the charging trip, 1_piC=0.

[0237] 4. Actor-Critic Algorithm

[0238] In this embodiment, the deep reinforcement learning algorithm is the Actor-Critic algorithm;

[0239] The processing steps of the Actor-Critic algorithm include:

[0240] S401: Initialize the strategy (Actor) π of the Actor-Critic algorithm θ (a|s) and value function (Critic); where the policy refers to the mapping from state to corresponding action;

[0241] S402: By strategy π θ (a|s) selects an action a based on the vehicle’s current state s;

[0242] S403: Execute the current action a and perform state transition through the Markov decision process model to obtain the next state s′ and the corresponding reward r;

[0243] S404: Estimate the value V of the current state s and the next state s′ through the value network w (s) and V w (s′); and then calculate the temporal difference error (TD Error);

[0244] The formula is:

[0245] δ=r+γV w (s′)-V w (s);

[0246] wherein: δ represents the time-difference error; γ represents the discount factor; w represents the value function parameter;

[0247] S405: updating the value function parameter w by the gradient descent method combined with the time-difference error δ;

[0248] The formula is represented as:

[0249]

[0250] wherein: a w represents the learning rate;

[0251] S406: updating the policy parameter θ of the policy π θ (a|s) by the policy gradient;

[0252] The formula is represented as:

[0253]

[0254] wherein: a θ represents the learning rate; represents the policy gradient;

[0255] S407: repeating steps S402 to S406 until the policy converges to an optimal solution or reaches a preset training round;

[0256] S408: inputting the current state of the vehicle into the trained policy π θ (a|s) to output the optimal transfer scheduling scheme.

[0257] V. Information Interaction Requirement

[0258] In this embodiment, the information interaction requirement refers to ensuring the real-time, integrity and security of data required in the task allocation, vehicle scheduling, battery charging and other links in the transfer scheduling process. The information interaction requirement includes: 1) task requirement: task ID, cargo type, start and end wharf, priority, remaining time; 2) vehicle requirement: vehicle ID, current position, current power, load status; 3) battery requirement: battery ID, capacity attenuation rate, charging rate, health status.

[0259] The process of incorporating the information interaction requirement is as follows: 1) data collection: the central controller collects the state data of tasks, vehicles and batteries in real time; 2) requirement mapping: converting the information interaction requirement into constraint conditions of the scheduling model (such as low-power vehicles being given priority for charging); 3) scheme generation: generating a scheduling scheme by a deep reinforcement learning algorithm to ensure the synchronous optimization of data flow and task flow; 4) dynamic adjustment: updating the scheduling scheme according to real-time feedback (such as task changes or battery abnormalities).

[0260] 6. Unified Information Interaction Solution

[0261] In this embodiment, the unified information interaction scheme includes: collecting task information, vehicle information, and battery information in real time through the central controller of the inland port; using all collected information as input, executing steps S2 to S5 to generate a target task allocation scheme and an optimal transshipment and charging scheduling scheme; issuing scheduling instructions to vehicles, charging stations, and batteries based on the target task allocation scheme and the optimal transshipment and charging scheduling scheme; and monitoring the status of vehicles, charging stations, and batteries after executing the scheduling instructions.

[0262] 7. Experimental Description

[0263] In order to better illustrate the advantages of the technical solution of the present invention, this embodiment discloses the following experiments.

[0264] 1. Model port (15 work tasks, 4 electric transfer vehicles)

[0265] For small-scale port applications, the operation of electric transfer vehicles is relatively simple, and their main task is to complete a small amount of material transfer tasks. At this time, the logistics scheduling problem is relatively small, and the frequency of battery charging and replacement is low. For this scenario, the multi-objective optimization framework of the present invention effectively reduces the battery decay rate through an accurate battery decay prediction model (such as CDECM), and reduces the charging cost and the time to complete the task by coordinating the optimization of energy scheduling and logistics scheduling. In this application scenario, by adopting a method based on deep reinforcement learning (DRL), the present invention can complete the task in a shorter time and can efficiently cope with the scheduling challenges of a small number of work tasks.

[0266] Table 1 shows that in this application scenario, the DRL-MOA (multi-objective optimization algorithm based on deep reinforcement learning) method shows significant advantages, with a higher hypervolume value and a relatively short running time:

[0267] Table 1

[0268]

[0269] This result shows that in a small-scale port environment, the use of a deep reinforcement learning optimization framework can effectively improve the overall system operating efficiency and battery life.

[0270] 2. Large-scale port (120 work tasks, 8 electric transfer vehicles)

[0271] In medium-sized port applications, the workload of electric transfer vehicles increases significantly, and the frequency of charging increases, making the coordination between logistics scheduling and energy scheduling more complex. Traditional optimization methods, faced with a large number of task scheduling and battery degradation, may struggle to ensure efficient and accurate completion of scheduling tasks. The present invention combines an efficient multi-objective optimization algorithm, particularly accurate battery degradation prediction based on the CDECM model, to effectively reduce battery degradation, optimize charging and recharging strategies, and reduce operating costs.

[0272] The results in Table 2 show that at this scale, the DRL-MOA optimization method not only improves the hypervolume value but also significantly reduces the computation time:

[0273] Table 2

[0274]

[0275] In medium-sized ports, the technical solution of the present invention can handle more complex scheduling tasks and obtain efficient optimization results in a relatively short time.

[0276] 3. Large-scale port (200 work tasks, 12 electric transfer vehicles)

[0277] For large-scale port applications, electric transfer vehicles face more complex scheduling tasks, with higher frequency of energy replenishment. Task completion time and battery decay rate become key factors. In this scenario, traditional methods struggle to meet real-time optimization requirements due to their large computational workload and slow convergence. This invention, through the adaptive capabilities of deep reinforcement learning, can rapidly learn and optimize the battery decay prediction model, while simultaneously improving the synergistic efficiency of energy and task scheduling, thereby effectively enhancing the overall operational efficiency of large-scale ports.

[0278] Table 3 shows that in this scenario, the DRL-MOA method demonstrates excellent performance, especially in terms of the balance between solution efficiency and optimization results, significantly surpassing traditional optimization methods:

[0279] Table 3

[0280]

[0281] In a large-scale port environment, the present invention can effectively perform task scheduling and, through precise battery management technology, optimize battery life, reduce operating costs, and improve the operating efficiency of the entire system.

[0282] 4. Comparison of different charging rules

[0283] In the energy scheduling of electric transshipment vehicles in ports, the choice of charging rules has a direct impact on energy costs, task duration, and battery degradation. As shown in Table 4, by comparing the effects of two different charging rules, partial charging and full charging, we can see that although the full charging rule slightly increases charging costs, it shows better results in terms of task duration and less battery degradation:

[0284] Table 4

[0285]

[0286] In practical applications, charging efficiency and battery life must be considered comprehensively when selecting charging rules to achieve optimal economic and environmental goals.

[0287] 5. Comparison of different routing rules

[0288] In the path planning of electric transport vehicles, different routing rules have a significant impact on charging costs, task completion time, and battery degradation rate. As shown in Table 5, by comparing the proposed routing rule with high-priority and low-priority routing rules, it can be seen that the proposed rule performs better in terms of charging costs and task duration, and has a lower battery degradation rate, demonstrating a strong comprehensive optimization capability:

[0289] Table 5

[0290]

[0291] In different application scenarios, the optimization strategy of the present invention can achieve coordinated optimization of energy and logistics scheduling through reasonable routing selection, charging strategy and task scheduling.

[0292] Example 2:

[0293] This embodiment discloses an inland river port vehicle transshipment and charging scheduling system based on unified information interaction, which is implemented based on the inland river port vehicle transshipment and charging scheduling method based on unified information interaction in Example 1.

[0294] The inland port vehicle transshipment and charging scheduling system based on unified information interaction includes:

[0295] Acquisition module, used to obtain mission information, vehicle information and battery information of inland ports;

[0296] In this embodiment, task information includes task ID, cargo type, starting and ending terminal coordinates, task priority (based on deadline and revenue), estimated duration, and task status (e.g., pending or in progress). Vehicle information includes vehicle ID, current location (GPS coordinates), current battery level, maximum load, and task queue status. Battery information includes battery ID, capacity decay rate, charge rate, historical depth of discharge records, and battery health score (SOH).

[0297] A preliminary information interaction building module is used to match tasks with vehicles using a reinforcement learning-based integer programming-critic solution framework, combining task information, vehicle information, and battery information to obtain a target task allocation solution. The module also establishes preliminary information interaction relationships between tasks, vehicles, and batteries based on the optimal task allocation solution.

[0298] In this embodiment, the preliminary information interaction relationship includes: establishing a basic association relationship between the task and the vehicle and the battery, that is, a matching relationship between task m and vehicle b and a binding relationship between vehicle b and the battery.

[0299] A model building module is used to build a Markov decision process model with the goal of maximizing the benefits of completing all tasks, minimizing the charging cost, and minimizing the battery capacity aging rate;

[0300] The calculation module is used to generate the optimal transport and charging scheduling plan based on the target task allocation plan by solving the Markov decision process model with task information, vehicle information and battery information as input through deep reinforcement learning algorithm;

[0301] An information interaction requirement association module is used to incorporate information interaction requirements into the target task allocation plan and the optimal transshipment and charging scheduling plan; based on the target task allocation plan and the optimal transshipment and charging scheduling plan, vehicles are dispatched to complete the tasks of the inland port;

[0302] The information unified interaction implementation module is used to design and execute the information unified interaction plan based on the preliminary information interaction relationship to realize the information interaction between tasks, vehicles and batteries.

[0303] The application matches tasks and vehicles based on the reinforcement learning-based integer programming-critic solution framework, combines task information, vehicle information and battery information to obtain a target task allocation scheme. Firstly, the framework can comprehensively consider the complex task demand of the inland port, the diverse performance of the vehicle and the different states of the battery, realize the accurate matching between the task and the vehicle, and effectively improve the accuracy and rationality of the task allocation compared with the traditional simple allocation method, thereby assisting in improving the accuracy of the inland port vehicle transfer and charging scheduling. Secondly, the integer programming can quantitatively allocate vehicle resources to ensure that each vehicle can be fully utilized, avoid vehicle idling or overuse, and the critic solution framework can adjust the allocation strategy through continuous learning and optimization to improve resource utilization efficiency and assist in improving the efficiency of the inland port vehicle transfer and charging scheduling. Finally, the reinforcement learning mechanism enables the entire framework to perceive the changes in the inland port environment in real time and dynamically adjust the task allocation scheme, so that the port vehicle scheduling system can adapt to various emergencies and maintain efficient and stable operation, thereby assisting in improving the flexibility of the inland port vehicle transfer and charging scheduling.

[0304] On the basis of task allocation, a Markov decision process model is constructed to maximize the revenue of completing all tasks, minimize the charging cost and minimize the battery capacity aging rate, and a deep reinforcement learning algorithm is used to solve the model to generate an optimal transfer and charging scheduling scheme. Firstly, the maximum revenue of completing all tasks can reasonably arrange the transfer route and time of the vehicle, improve the efficiency of cargo transportation and increase the cargo throughput of the port, thereby directly improving the operating revenue of the port; at the same time, by optimizing the scheduling scheme, the idle time and empty mileage of the vehicle are reduced, thereby reducing the cost of the inland port vehicle transfer and charging scheduling. Secondly, the minimum charging cost of completing all tasks enables the model to develop an optimal charging strategy according to the remaining battery capacity, battery charging characteristics and power grid price fluctuations, thereby reasonably arranging the charging time period and charging cost. Then, the minimum battery capacity aging rate of completing all tasks enables the model to optimize the charging and discharging process of the vehicle, avoid adverse operations such as excessive charging, excessive discharging and high-temperature charging, reduce the attenuation of battery capacity and prolong the service life of the battery, which not only reduces the battery replacement cost, but also reduces the vehicle downtime caused by battery failure, thereby improving the reliability and availability of the vehicle. Finally, the deep reinforcement learning algorithm can process complex Markov decision process models, learn and analyze a large amount of historical data and real-time information to generate an optimal transfer and charging scheduling scheme, thereby improving the quality and efficiency of the inland port vehicle transfer and charging scheduling.

[0305] The present invention incorporates information interaction requirements into the target task allocation plan and the optimal transfer and charging scheduling plan to ensure the smoothness of information flow among all parties, enabling real-time and accurate information sharing between various links such as task allocation, vehicle scheduling and battery management, improving scheduling coordination, and avoiding scheduling confusion and conflicts caused by poor information flow.

[0306] The present invention designs a unified information interaction scheme based on the preliminary information interaction relationship, and executes the unified information interaction scheme to realize information interaction between tasks, vehicles and batteries. First, the unified information interaction scheme can organically combine task allocation, vehicle charging scheduling and battery status management to achieve coordinated optimization of the three. Through real-time information interaction, task allocation and charging scheduling schemes can be dynamically adjusted according to the battery status. At the same time, task allocation and charging scheduling will also affect the battery usage and maintenance strategy, thereby achieving the overall optimal effect. Secondly, through the information interaction between tasks, vehicles and batteries, each other's status and needs can be understood in a timely manner, the intermediate links of information transmission can be reduced, and operational efficiency can be improved; at the same time, through unified information interaction, the resource status of tasks, vehicles and batteries can be fully grasped, and the optimal allocation of resources can be achieved. Finally, the unified information interaction scheme can monitor the operating status of tasks, vehicles and batteries in real time, discover potential problems and risks in a timely manner, and take corresponding measures to make adjustments to ensure the stable operation of the system.

[0307] In the preliminary information interaction building module of this embodiment, the integer programming-critic solution framework based on reinforcement learning is used to match tasks with vehicles by combining task information, vehicle information, and battery information. The processing steps for obtaining the target task allocation solution are as follows:

[0308] A201: Define state s, action a, and reward r;

[0309] 1) The state s includes task information and vehicle information;

[0310] 2) Action a is the matching decision X between the task and the vehicle m,b,i ;X m,b,i Indicates whether task m is assigned to vehicle b at position i in the vehicle task list, X m,b,i =1 means allocation, X m,b,i =0 means no allocation;

[0311] 3) The formula for reward r is:

[0312]

[0313] Where: r represents the reward for the vehicle to complete a task; Indicates the profit of the vehicle completing a task; Indicates the cost of electricity purchased by the vehicle to complete a mission; Indicates the battery capacity aging rate of the vehicle after completing a mission; w grid 、w job 、w deg Indicates the corresponding weight coefficient; p b represents the charging power; i represents the number of discharges; δ cyc Indicates the battery attenuation rate;

[0314] A202: Input the current state s into the policy improvement model and select the optimal action X m,b,i , and calculate the reward

[0315] A203: Based on the optimal action X through the Q function Q(m,b,i) m,b,i Calculate the corresponding Q value estimate;

[0316] A204: Reward-based And Q value estimate update Q function Q(m,b,i);

[0317] The formula is:

[0318]

[0319] Where: represents the immediate reward observed when assigning task m to vehicle b at position i in the vehicle task list; α is the learning rate; Represents the TD error, which is used to quantify the difference between the reward and the Q value estimate;

[0320] A205: Repeat steps A202 to A204 to iteratively update the Q function Q(m,b,i);

[0321] A206: Calculate the loss function based on all current rewards and Q values, and reversely optimize the parameters of the Q function Q(m,b,i) based on the loss function;

[0322] The formula of the loss function is expressed as:

[0323]

[0324] Where: Z is the number of samples in each training batch;

[0325] The formula for optimizing the parameters of the Q function Q(m,b,i) is expressed as:

[0326]

[0327] Where: θ represents the Q function parameter to be optimized; lr represents the learning rate of parameter update; Represents the gradient of the loss function with respect to θ;

[0328] A207: Repeat steps A202 to A206 to iteratively train the Q function Q(m, b, i) until convergence;

[0329] A208: Input the current state of the vehicle into the trained Q function Q(m,b,i) and output the target task allocation plan.

[0330] In step SA02, the objective function of the strategy improvement model is expressed as:

[0331]

[0332] The constraints are expressed as:

[0333]

[0334] The objective function is to maximize the total reward in vehicle-task matching; the first constraint requires that each task m can only be assigned once among all vehicles b, task positions i, and actions a, ensuring that each task occupies a unique position in the vehicle's task list; the second constraint forces each vehicle b to handle only one task at each task position i, preventing conflicts or overloading in the vehicle's task list; the third constraint ensures that each decision variable X m,b,i It is a binary variable to ensure the discreteness of task assignment.

[0335] In the model building module of this embodiment, the Markov decision process model includes state variables, action space and reward function:

[0336] 1) State variables:

[0337]

[0338] Where: s b represents the state of battery b (since the vehicle and the battery are associated and bound, the battery can also be described by b); t represents the time step; Indicates the current itinerary; Indicates the next trip; c b Indicates the current SoC; τ b Indicates the remaining time of the current trip;

[0339] 2) Action space, including cargo transfer trips, empty load dispatch trips, trips to charging stations, and charging trips;

[0340] 2.1) Battery operation space during cargo transfer:

[0341]

[0342] Where: represents the action space of battery b in a cargo transfer trip; Indicates the speed of the next trip; Indicates the start time of the next trip; Indicates the trip after the next trip; Respectively represent the lower and upper limits of the speed; Indicate itinerary The set of next potential trips; Indicates cargo transfer trips, empty load dispatch trips, and trips to charging stations;

[0343] 2.2) Battery operation space during no-load dispatching trip:

[0344]

[0345] Where: represents the action space of battery b in an unloaded dispatch trip;

[0346] 2.3) The action space of the battery during the trip to the charging station:

[0347]

[0348] Where: represents the action space of battery b during a trip to the charging station; Indicates the charging rate of the charging trip; Indicates the lower and upper limits of the charging rate;

[0349] 2.4) Battery operation space during charging process:

[0350]

[0351] Where: represents the action space of battery b during a charging trip;

[0352] 3) Reward function:

[0353]

[0354]

[0355] Where: r represents the reward for the vehicle to complete a task; Indicates the profit of the vehicle completing a task; Indicates the cost of electricity purchased by the vehicle to complete a mission; Indicates the battery capacity aging rate of the vehicle after completing a mission; w grid 、w job 、w deg Indicates the corresponding weight coefficient; pb represents the charging power; i represents the number of discharges; δ cyc Indicates the battery degradation rate.

[0356] Specifically, the state transition process of the Markov decision process model is expressed as:

[0357] 1) Status transition of cargo transshipment trip

[0358] The actions of the cargo transfer trip are expressed as:

[0359]

[0360] The state transition of a cargo transfer trip is expressed as:

[0361]

[0362] Where: and Represents the energy consumed per unit time for the current trip and the next trip respectively; m 0 ,v 0 Respectively indicate the execution of the current trip Vehicle weight and speed; m 1 ,v 1 Respectively indicate the execution of the next trip The vehicle weight and speed; E represents the energy capacity of the battery; Indicates the next trip κ 1 distance;

[0363] 2) State transition of no-load dispatching trip

[0364] The action of the no-load dispatch trip is expressed as:

[0365]

[0366] The state transition of the no-load dispatch trip is expressed as:

[0367]

[0368] The action of traveling to the charging station is represented as:

[0369]

[0370] State transitions for the trip to the charging station:

[0371]

[0372] The action of the charging trip is expressed as:

[0373]

[0374] The state transition of the charging trip is represented as:

[0375]

[0376] In the model construction module, a battery aging quantification model is constructed to calculate the battery capacity aging rate;

[0377] The calculation formula of the battery aging quantification model is:

[0378]

[0379] In the formula: δ deg represents the battery capacity aging rate under the non-design cycle condition; N seq represents the non-design discharge times under the non-design cycle condition; represents the cycle aging capacity decay rate of the nth non-design discharge; Δδ represents the capacity decay error under the non-design cycle condition, which is a correction term for describing the influence of small-scale discharge depth on battery capacity degradation;

[0380] wherein:

[0381]

[0382] In the formula: is a constant parameter; q represents the order of the battery aging quantification model; DoD q represents the discharge depth value of the qth order; represents the discharge depth value of the nth non-design discharge; represents the total discharge depth value of the entire non-design cycle condition; DoD tot represents the change of SOC from the end of charging to the start of the next charging; N seq represents the total number of discharges.

[0383] In the calculation module of the present embodiment, the deep reinforcement learning algorithm is an Actor-Critic algorithm;

[0384] The processing steps of the Actor-Critic algorithm include:

[0385] A401: Initialize the policy π θ of the Actor-Critic algorithm; wherein the policy refers to the mapping from the state to the corresponding action;

[0386] A402: Select an action a based on the current state s of the vehicle through the policy π θ (a|s);

[0387] A403: Execute the current action a through the Markov decision process model and perform state transition to obtain the next state s' and the corresponding reward r;

[0388] A404: Estimate the value V of the current state s and the next state s′ through the value network w (s) and V w (s′); and then calculate the time difference error;

[0389] The formula is:

[0390] δ=r+γV w (s′)-V w (s);

[0391] Where: δ represents the time difference error; γ represents the discount factor; w represents the value function parameter;

[0392] A405: Update the value function parameter w by gradient descent combined with the time difference error δ;

[0393] The formula is:

[0394]

[0395] Where: a w represents the learning rate;

[0396] A406: Update policy π via policy gradient θ The policy parameter θ of (a|s);

[0397] The formula is:

[0398]

[0399] Where: a θ represents the learning rate; represents the policy gradient;

[0400] A407: Repeat steps A402 to A406 until the strategy converges to the optimal solution or reaches the preset training rounds;

[0401] A408: Input the current state of the vehicle into the trained strategy π θ In (a|s), output the optimal transshipment scheduling plan.

[0402] In this embodiment, information exchange requirements refer to ensuring the real-time, integrity, and security of data required for task assignment, vehicle scheduling, battery charging, and other aspects of the transshipment and scheduling process. Information exchange requirements include: 1) Task requirements: task ID, cargo type, starting and ending terminals, priority, and remaining time; 2) Vehicle requirements: vehicle ID, current location, current battery level, and load status; 3) Battery requirements: battery ID, capacity decay rate, charging rate, and health status.

[0403] The process of incorporating information interaction requirements: 1) Data acquisition: The central controller collects task, vehicle, and battery status data in real time; 2) Demand mapping: Converts information interaction requirements into constraints for the scheduling model (such as prioritizing charging for vehicles with low battery power); 3) Plan generation: Generates scheduling plans through deep reinforcement learning algorithms to ensure synchronous optimization of data flow and task flow; 4) Dynamic adjustment: Updates the scheduling plan based on real-time feedback (such as task changes or battery anomalies).

[0404] In this embodiment, the unified information interaction scheme includes: collecting task information, vehicle information, and battery information in real time through the central controller of the inland port; using all collected information as input, executing steps S2 to S5 to generate a target task allocation scheme and an optimal transshipment and charging scheduling scheme; issuing scheduling instructions to vehicles, charging stations, and batteries based on the target task allocation scheme and the optimal transshipment and charging scheduling scheme; and monitoring the status of vehicles, charging stations, and batteries after executing the scheduling instructions.

[0405] Example 3:

[0406] This embodiment discloses a computer device.

[0407] The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of each of the aforementioned embodiments of the inland port vehicle transshipment and charging scheduling method based on unified information interaction. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in each of the aforementioned system embodiments.

[0408] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the power system dynamic reactive power compensation device.

[0409] The computer device may be a desktop computer, a notebook computer, a PDA, a cloud server, etc. The computer device may include, but is not limited to, a processor and a memory.

[0410] Example 4:

[0411] This embodiment discloses a computer-readable storage medium.

[0412] A computer-readable storage medium having a computer program stored thereon, which, when executed, performs the steps in each of the above-mentioned embodiments of the inland port vehicle transshipment and charging scheduling method based on unified information interaction.

[0413] The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of each of the above-mentioned method embodiments may be implemented. The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.

[0414] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the technical solutions. Those skilled in the art should understand that modifications or equivalent replacements of the technical solutions of the present invention that do not depart from the purpose and scope of the technical solutions of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A method for dispatching inland river port vehicles for transshipment and charging based on unified information interaction, characterized in that: include: Obtain mission information, vehicle information, and battery information at inland ports; Through the reinforcement learning-based integer programming-critic solution framework, the task and vehicle information are combined to match the task and vehicle, and the target task allocation solution is obtained; Establish preliminary information interaction between tasks, vehicles and batteries based on the optimal task allocation party; Construct a Markov decision process model with the goal of maximizing the benefits of completing all tasks, minimizing charging costs, and minimizing battery capacity aging rate; Based on the target task allocation plan, a deep reinforcement learning algorithm is used to solve the Markov decision process model that takes task information, vehicle information, and battery information as input to generate the optimal transport and charging scheduling plan; Incorporate information exchange requirements into target task allocation plans and optimal transport and charging scheduling plans; Dispatching vehicles to complete inland port tasks based on target task allocation schemes and optimal transshipment and charging scheduling schemes; Based on the preliminary information interaction relationship, a unified information interaction plan is designed and implemented to realize information interaction between tasks, vehicles and batteries.

2. The method for inland port vehicle transshipment and charging scheduling based on unified information interaction according to claim 1, characterized in that: Using the reinforcement learning-based integer programming-critic solution framework, we match tasks with vehicles by combining task information, vehicle information, and battery information. The following steps are used to obtain the target task allocation solution: S201: Define state s, action a and reward r; 1) The state s includes task information and vehicle information; 2) Action a is the matching decision X between the task and the vehicle m,b,i ;X m,b,i Indicates whether task m is assigned to vehicle b at position i in the vehicle task list, X m,b,i =1 means allocation, X m,b,i =0 means no allocation; 3) The formula for reward r is: Where: represents the immediate reward observed when assigning task m to vehicle b at position i in the vehicle's task list; Indicates the profit of the vehicle completing a task; Indicates the cost of electricity purchased by the vehicle to complete a mission; Indicates the battery capacity aging rate of the vehicle after completing a mission; w grid 、w job 、w deg Indicates the corresponding weight coefficient; p b represents the charging power; n represents the number of discharges; δ cyc Indicates the battery attenuation rate; S202: Input the current state s into the policy improvement model and select the optimal action X m,b,i , and calculate the reward S203: Based on the optimal action X through the Q function Q(m,b,i) m,b,i Calculate the corresponding Q value estimate; S204: Reward-based And Q value estimate update Q function Q(m,b,i); The formula is: Where: represents the immediate reward observed when assigning task m to vehicle b at position i in the vehicle task list; α is the learning rate; Represents the TD error, which is used to quantify the difference between the reward and the Q value estimate; S205: Repeat steps S202 to S204 to iteratively update the Q function Q(m,b,i); S206: Calculate the loss function based on all current rewards and Q values, and reversely optimize the parameters of the Q function Q(m, b, i) based on the loss function; The formula of the loss function is expressed as: Where: Z is the number of samples in each training batch; The formula for optimizing the parameters of the Q function Q(m,b,i) is expressed as: Where: θ represents the Q function parameter to be optimized; lr represents the learning rate of parameter update; Represents the gradient of the loss function with respect to θ; S207: Repeat steps S202 to S206 to iteratively train the Q function Q(m, b, i) until convergence; S208: Input the current state of the vehicle into the trained Q function Q(m,b,i) and output the target task allocation plan.

3. The method for inland river port vehicle transshipment and charging scheduling based on unified information interaction as claimed in claim 2, characterized in that: The objective function of the policy improvement model is expressed as: The constraints are expressed as: The objective function is to maximize the total reward in vehicle-task matching; the first constraint requires that each task m can only be assigned once among all vehicles b, task positions i, and actions a, ensuring that each task occupies a unique position in the vehicle's task list; the second constraint forces each vehicle b to handle only one task at each task position i, preventing conflicts or overloading in the vehicle's task list; the third constraint ensures that each decision variable X m,b,i It is a binary variable to ensure the discreteness of task assignment.

4. The method for inland river port vehicle transshipment and charging scheduling based on unified information interaction according to claim 1 is characterized in that: The preliminary information interaction relationship includes: establishing a basic association relationship between tasks, vehicles and batteries, that is, the matching relationship between task m and vehicle b and the binding relationship between vehicle b and the battery.

5. The method for inland river port vehicle transshipment and charging scheduling based on unified information interaction according to claim 1 is characterized in that: The Markov decision process model includes state variables, action space, and reward function: 1) State variables: Where: s b represents the state of battery b (since the vehicle and the battery are associated and bound, the battery can also be described by b); t represents the time step; Indicates the current itinerary; Indicates the next trip; c b Indicates the current SoC; τ b Indicates the remaining time of the current trip; 2) Action space, including cargo transfer trips, empty load dispatch trips, trips to charging stations, and charging trips; 2.1) Battery operation space during cargo transfer: Where: represents the action space of battery b in a cargo transfer trip; Indicates the speed of the next trip; Indicates the start time of the next trip; Indicates the trip after the next trip; Respectively represent the lower and upper limits of the speed; Indicate itinerary The set of next potential trips; Indicates cargo transfer trips, empty load dispatch trips, and trips to charging stations; 2.2) Battery operation space during no-load dispatching trip: Where: represents the action space of battery b in an unloaded dispatch trip; 2.3) The action space of the battery during the trip to the charging station: Where: represents the action space of battery b during a trip to the charging station; Indicates the charging rate of the charging trip; Indicates the lower and upper limits of the charging rate; 2.4) Battery operation space during charging process: Where: represents the action space of battery b during a charging trip; 3) Reward function: Where: r represents the reward for the vehicle to complete a task; Indicates the profit of the vehicle completing a task; Indicates the cost of electricity purchased by the vehicle to complete a mission; Indicates the battery capacity aging rate of the vehicle after completing a mission; w grid 、w job 、w deg Indicates the corresponding weight coefficient; p b represents the charging power; n represents the number of discharges; δ cyc Indicates the battery degradation rate.

6. The method for inland river port vehicle transshipment and charging scheduling based on unified information interaction according to claim 5 is characterized in that: The state transition process of the Markov decision process model is expressed as: 1) Status transition of cargo transshipment trip The actions of the cargo transfer trip are expressed as: The state transition of a cargo transfer trip is expressed as: Where: and Represents the energy consumed per unit time for the current trip and the next trip respectively; m 0 ,v 0 Respectively indicate the execution of the current trip Vehicle weight and speed; m 1 ,v 1 Respectively indicate the execution of the next trip The vehicle weight and speed; E represents the energy capacity of the battery; Indicates the next trip κ 1 distance; 2) State transition of no-load dispatching trip The action of the no-load dispatch trip is expressed as: The state transition of the no-load dispatch trip is expressed as: 3) State transition for the trip to the charging station The action of traveling to the charging station is represented as: State transitions for the trip to the charging station: 4) State transition of charging trip The action of the charging trip is expressed as: The state transition of the charging trip is expressed as:

7. The method for inland port vehicle transshipment and charging scheduling based on unified information interaction according to claim 1, characterized in that: Calculate the battery capacity aging rate by building a battery aging quantitative model; The calculation formula of the battery aging quantitative model is: Where: δ deg Indicates the battery capacity aging rate under non-design cycle conditions; N seq Indicates the non-design discharge times of non-design cycle conditions; It represents the cycle aging capacity attenuation rate of the nth non-design discharge; Δδ represents the capacity attenuation error of the non-design cycle condition, which is a correction term used to describe the effect of small-scale discharge depth on battery capacity degradation; in: Where: is a constant parameter; q represents the order of the battery aging quantization model; DoD q Indicates the discharge depth value of the qth order; Indicates the discharge depth value of the nth non-design discharge; Indicates the total depth of discharge value of the entire non-design cycle condition; DoD tot Indicates the SOC change from the end of charging to the start of the next charging; N seq Indicates the total number of discharges.

8. The method for inland river port vehicle transshipment and charging scheduling based on unified information interaction as claimed in claim 5, characterized in that: The deep reinforcement learning algorithm is the Actor-Critic algorithm; The processing steps of the Actor-Critic algorithm include: S401: Initialize the strategy π of the Actor-Critic algorithm θ (a|s) and value function; where policy refers to the mapping from state to corresponding action; S402: By strategy π θ (a|s) selects an action a based on the vehicle’s current state s; S403: Execute the current action a and perform state transition through the Markov decision process model to obtain the next state s′ and the corresponding reward r; S404: Estimate the value V of the current state s and the next state s′ through the value network w (s) and V w (s′); and then calculate the time difference error; The formula is: δ=r+γV w (s′)-V w (s); Where: δ represents the time difference error; γ represents the discount factor; w represents the value function parameter; S405: Update the value function parameter w by gradient descent combined with the time difference error δ; The formula is: Where: a w represents the learning rate; S406: Update strategy π through policy gradient θ The policy parameter θ of (a|s); The formula is: Where: a θ represents the learning rate; represents the policy gradient; S407: Repeat steps S402 to S406 until the strategy converges to the optimal solution or reaches a preset number of training rounds; S408: Input the current state of the vehicle into the trained strategy π θ In (a|s), output the optimal transshipment scheduling plan.

9. An inland river port vehicle transshipment and charging scheduling system based on unified information interaction, implemented based on the inland river port vehicle transshipment and charging scheduling method based on unified information interaction in claim 1; the system comprises: Acquisition module, used to obtain mission information, vehicle information and battery information of inland ports; A preliminary information interaction building block is used to match tasks with vehicles by combining task information, vehicle information, and battery information through a reinforcement learning-based integer programming-critic solution framework to obtain the target task allocation plan; Establish preliminary information interaction between tasks, vehicles and batteries based on the optimal task allocation party; A model building module is used to build a Markov decision process model with the goal of maximizing the benefits of completing all tasks, minimizing the charging cost, and minimizing the battery capacity aging rate; The calculation module is used to generate the optimal transport and charging scheduling plan based on the target task allocation plan by solving the Markov decision process model with task information, vehicle information and battery information as input through deep reinforcement learning algorithm; An information interaction requirement association module is used to incorporate information interaction requirements into the target task allocation plan and the optimal transport and charging scheduling plan; Dispatching vehicles to complete inland port tasks based on target task allocation schemes and optimal transshipment and charging scheduling schemes; The information unified interaction implementation module is used to design and execute the information unified interaction plan based on the preliminary information interaction relationship to realize the information interaction between tasks, vehicles and batteries.

10. The inland port vehicle transshipment and charging scheduling system based on unified information interaction as claimed in claim 9, characterized in that: In the preliminary information interaction building module, the reinforcement learning-based integer programming-critic solution framework combines task information, vehicle information, and battery information to match tasks with vehicles. The processing steps to obtain the target task allocation solution are as follows: A201: Define state s, action a, and reward r; 1) The state s includes task information and vehicle information; 2) Action a is the matching decision X between the task and the vehicle m,b,i ;X m,b,i Indicates whether task m is assigned to vehicle b at position i in the vehicle task list, X m,b,i =1 means allocation, X m,b,i =0 means no allocation; 3) The formula for reward r is: Where: r represents the reward for the vehicle to complete a task; Indicates the profit of the vehicle completing a task; Indicates the cost of electricity purchased by the vehicle to complete a mission; Indicates the battery capacity aging rate of the vehicle after completing a mission; w grid 、w job 、w deg Indicates the corresponding weight coefficient; p b represents the charging power; i represents the number of discharges; δ cyc Indicates the battery attenuation rate; A202: Input the current state s into the policy improvement model and select the optimal action X m,b,i , and calculate the reward A203: Based on the optimal action X through the Q function Q(m,b,i) m,b,i Calculate the corresponding Q value estimate; A204: Reward-based And Q value estimate update Q function Q(m,b,i); The formula is: Where: represents the immediate reward observed when assigning task m to vehicle b at position i in the vehicle task list; α is the learning rate; Represents the TD error, which is used to quantify the difference between the reward and the Q value estimate; A205: Repeat steps A202 to A204 to iteratively update the Q function Q(m, b, i); A206: Calculate the loss function based on all current rewards and Q values, and reversely optimize the parameters of the Q function Q(m,b,i) based on the loss function; The formula of the loss function is expressed as: Where: Z is the number of samples in each training batch; The formula for optimizing the parameters of the Q function Q(m,b,i) is expressed as: Where: θ represents the Q function parameter to be optimized; lr represents the learning rate of parameter update; Represents the gradient of the loss function with respect to θ; A207: Repeat steps A202 to A206 to iteratively train the Q function Q(m, b, i) until convergence; A208: Input the current state of the vehicle into the trained Q function Q(m,b,i) and output the target task allocation plan.

11. The inland port vehicle transshipment and charging scheduling system based on unified information interaction as claimed in claim 10, characterized in that: In step SA02, the objective function of the strategy improvement model is expressed as: The constraints are expressed as: The objective function is to maximize the total reward in vehicle-task matching; the first constraint requires that each task m can only be assigned once among all vehicles b, task positions i, and actions a, ensuring that each task occupies a unique position in the vehicle's task list; the second constraint forces each vehicle b to handle only one task at each task position i, preventing conflicts or overloading in the vehicle's task list; the third constraint ensures that each decision variable X m,b,i It is a binary variable to ensure the discreteness of task assignment.

12. The inland port vehicle transshipment and charging scheduling system based on unified information interaction according to claim 9, characterized in that: In the model building module, the Markov decision process model includes state variables, action space, and reward function: 1) State variables: Where: s b represents the state of battery b (since the vehicle and the battery are associated and bound, the battery can also be described by b); t represents the time step; Indicates the current itinerary; Indicates the next trip; c b Indicates the current SoC; τ b Indicates the remaining time of the current trip; 2) Action space, including cargo transfer trips, empty load dispatch trips, trips to charging stations, and charging trips; 2.1) Battery operation space during cargo transfer: Where: represents the action space of battery b in a cargo transfer trip; Indicates the speed of the next trip; Indicates the start time of the next trip; Indicates the trip after the next trip; Respectively represent the lower and upper limits of the speed; Indicate itinerary The set of next potential trips; Indicates cargo transfer trips, empty load dispatch trips, and trips to charging stations; 2.2) Battery operation space during no-load dispatching trip: Where: represents the action space of battery b in an unloaded dispatch trip; 2.3) The action space of the battery during the trip to the charging station: Where: represents the action space of battery b during a trip to the charging station; Indicates the charging rate of the charging trip; Indicates the lower and upper limits of the charging rate; 2.4) Battery operation space during charging process: Where: represents the action space of battery b during a charging trip; 3) Reward function: Where: r represents the reward for the vehicle to complete a task; Indicates the profit of the vehicle completing a task; Indicates the cost of electricity purchased by the vehicle to complete a mission; Indicates the battery capacity aging rate of the vehicle after completing a mission; w grid 、w job 、w deg Indicates the corresponding weight coefficient; p b represents the charging power; i represents the number of discharges; δ cyc Indicates the battery degradation rate.

13. The inland port vehicle transshipment and charging scheduling system based on unified information interaction as claimed in claim 12, characterized in that: The state transition process of the Markov decision process model is expressed as: 1) Status transition of cargo transshipment trip The actions of the cargo transfer trip are expressed as: The state transition of a cargo transfer trip is expressed as: Where: and Represents the energy consumed per unit time for the current trip and the next trip respectively; m 0 ,v 0 Respectively indicate the execution of the current trip Vehicle weight and speed; m 1 ,v 1 Respectively indicate the execution of the next trip The vehicle weight and speed; E represents the energy capacity of the battery; Indicates the next trip k 1 distance; 2) State transition of no-load dispatching trip The action of the no-load dispatch trip is expressed as: The state transition of the no-load dispatch trip is expressed as: 3) State transition of the trip to the charging station The action of the trip to the charging station is expressed as: State transitions for the trip to the charging station: 4) State transition of charging trip The action of charging trip is expressed as: The state transition of the charging trip is expressed as:

14. The inland port vehicle transshipment and charging scheduling system based on unified information interaction according to claim 9, characterized in that: In the model building module, the battery capacity aging rate is calculated by building a battery aging quantification model; The calculation formula of the battery aging quantitative model is: Where: δ deg Indicates the battery capacity aging rate under non-design cycle conditions; N seq Indicates the non-design discharge times of non-design cycle conditions; It represents the cycle aging capacity attenuation rate of the nth non-design discharge; Δδ represents the capacity attenuation error of the non-design cycle condition, which is a correction term used to describe the effect of small-scale discharge depth on battery capacity degradation; in: Where: is a constant parameter; q represents the order of the battery aging quantization model; DoD q Indicates the discharge depth value of the qth order; Indicates the discharge depth value of the nth non-design discharge; Indicates the total depth of discharge value of the entire non-design cycle condition; DoD tot Indicates the SOC change from the end of charging to the start of the next charging; N seq Indicates the total number of discharges.

15. The inland port vehicle transshipment and charging scheduling system based on unified information interaction according to claim 9, characterized in that: In the computing module, the deep reinforcement learning algorithm is the Actor-Critic algorithm; The processing steps of the Actor-Critic algorithm include: A401: Initialize the strategy π of the Actor-Critic algorithm θ (a|s) and value function; where policy refers to the mapping from state to corresponding action; A402: Through Strategy π θ (a|s) selects an action a based on the vehicle’s current state s; A403: Execute the current action a and perform state transition through the Markov decision process model to obtain the next state s′ and the corresponding reward r; A404: Estimate the value V of the current state s and the next state s′ through the value network w (s) and V w (s′); and then calculate the time difference error; The formula is: δ=r+γV w (s′)-V w (s); Where: δ represents the time difference error; γ represents the discount factor; w represents the value function parameter; A405: Update the value function parameter w by gradient descent combined with the time difference error δ; The formula is: Where: a w represents the learning rate; A406: Update policy π via policy gradient θ The policy parameter θ of (a|s); The formula is: Where: a θ represents the learning rate; represents the policy gradient; A407: Repeat steps A402 to A406 until the strategy converges to the optimal solution or reaches the preset training rounds; A408: Input the current state of the vehicle into the trained strategy π θ In (a|s), output the optimal transshipment scheduling plan.

16. A computer device, characterized in that: include: one or more processors; The processor is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the inland port vehicle transshipment and charging scheduling method based on unified information interaction as described in any one of claims 1 to 8 is implemented.

17. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed, it implements the inland port vehicle transshipment and charging scheduling method based on unified information interaction as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Unmanned mine card charging dispatching method and system based on cloud platform

    CN121352432A