Multimodal transport intelligent scheduling optimization method, device, equipment and storage medium

Through modeling and intelligent scheduling optimization of multimodal transport networks, multi-layer feedforward neural networks and improved deep deterministic strategy gradient algorithms are used to solve the problems of dynamic network changes and low resource utilization in multimodal transport scheduling, and efficient and flexible scheduling decisions and rapid response are achieved.

CN119167789BActive Publication Date: 2025-08-22SICHUAN CANGLAN HONGHAN SUPPLY CHAIN MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411440001.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-08-22
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

Traditional multimodal transport scheduling methods are difficult to effectively deal with large-scale and dynamically changing transportation networks, with low resource utilization and lack of accurate predictions of future freight demand, resulting in low scheduling efficiency and insufficient adaptability.

Method used

By modeling multiple transportation nodes and paths, building a network topology structure, using a multi-layer feedforward neural network to predict freight demand, combining the improved dual-delay depth deterministic strategy gradient algorithm to generate scheduling decisions, introducing dynamic time windows and adaptive adjustment mechanisms, and optimizing multimodal transport scheduling schemes.

Benefits of technology

It improves the operational efficiency of multimodal transport, enhances the accuracy of forecasting future freight demand, realizes efficient scheduling decision-making and resource balance in complex environments, and improves the system's response speed and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119167789B_ABST
    Figure CN119167789B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent scheduling technology, and discloses a multimodal intelligent scheduling optimization method, device, equipment and storage medium. The method comprises: modeling multiple transport nodes and transport paths to obtain a network topology structure; dividing the scheduling cycle into time periods based on the network topology structure to obtain a time window scheduling framework; inputting historical freight data into a multi-layer feedforward neural network for training to obtain a demand forecaster; dynamically adjusting the time window of the time window scheduling framework to obtain a dynamically adjusted time window; inputting the state space and action space into an improved double-delay deep deterministic policy gradient algorithm for training to obtain a policy generator; and generating multimodal scheduling decisions based on the demand forecaster, the dynamically adjusted time window and the policy generator to obtain a multimodal scheduling solution. The present invention realizes efficient and flexible scheduling decision generation and improves the operational efficiency of multimodal transport.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent scheduling, and in particular to a multimodal transport intelligent scheduling optimization method, device, equipment and storage medium. Background Art

[0002] As an efficient and environmentally friendly mode of logistics and transportation, traditional intermodal transport scheduling methods are no longer able to meet the current rapidly changing market demands. The complexity of the intermodal transport network, the diversity of transportation modes, and the uncertainty of cargo demand all pose significant challenges to scheduling optimization.

[0003] Existing intermodal transport scheduling methods often have several major problems: it is difficult to effectively handle large-scale, dynamically changing transportation networks, resulting in low scheduling efficiency; secondly, insufficient consideration is given to the coordination and balance between multiple modes of transport, resulting in low resource utilization; thirdly, the lack of accurate predictions of future freight demand makes scheduling decisions lack adaptability and foresight. Summary of the Invention

[0004] The present invention provides a multimodal transport intelligent scheduling optimization method, device, equipment and storage medium, which realizes efficient and flexible scheduling decision generation and improves the operational efficiency of multimodal transport.

[0005] In a first aspect, the present invention provides a multimodal transport intelligent scheduling optimization method, the multimodal transport intelligent scheduling optimization method comprising:

[0006] Model multiple transportation nodes and transportation paths to obtain the network topology;

[0007] Dividing the scheduling period into time periods based on the network topology to obtain a time window scheduling framework;

[0008] The historical freight data is fed into a multi-layer feedforward neural network for training to obtain a demand forecaster;

[0009] Dynamically adjusting the time window of the time window scheduling framework to obtain a dynamically adjusted time window;

[0010] The state space and action space are input into the improved double-delayed deep deterministic policy gradient algorithm for training to obtain the policy generator;

[0011] An intermodal transport scheduling decision is generated based on the demand forecaster, the dynamically adjusted time window, and the strategy generator to obtain an intermodal transport scheduling solution.

[0012] In a second aspect, the present invention provides a multimodal transport intelligent scheduling optimization device, the multimodal transport intelligent scheduling optimization device comprising:

[0013] A modeling module is used to model multiple transportation nodes and transportation paths to obtain a network topology structure;

[0014] A partitioning module, configured to partition the scheduling period into time periods based on the network topology to obtain a time window scheduling framework;

[0015] A training module is used to input historical freight data into a multi-layer feedforward neural network for training to obtain a demand forecaster;

[0016] An adjustment module, configured to dynamically adjust the time window of the time window scheduling framework to obtain a dynamically adjusted time window;

[0017] A processing module is used to input the state space and action space into the improved double-delay deep deterministic policy gradient algorithm for training to obtain a policy generator;

[0018] A decision module is used to generate a multimodal transport scheduling decision based on the demand forecaster, the dynamically adjusted time window and the strategy generator to obtain a multimodal transport scheduling solution.

[0019] The third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned multimodal transport intelligent scheduling optimization method.

[0020] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned multimodal transport intelligent scheduling optimization method.

[0021] In the technical solution provided by the present invention, a complete network topology is obtained by comprehensively modeling multiple transport nodes and transport paths. The scheduling cycle is divided into time periods based on the network topology, and a dynamic adjustment mechanism is introduced to enable the scheduling plan to better adapt to changes in actual transportation demand. A multi-layer feedforward neural network is used to train historical freight data, and an adaptive adjustment mechanism is introduced to improve the accuracy of predicting future freight demand. The policy generator obtained by training with an improved double-delay deep deterministic policy gradient algorithm is capable of generating high-quality scheduling decisions in complex state spaces. By introducing multimodal coordination factors and effect evaluation indicators, while optimizing the total transportation cost and time, the balance between different modes of transportation is also taken into account. Mechanisms such as dynamic time window adjustment and adaptive learning rate are adopted to enable the scheduling system to better adapt to environmental changes and uncertainties. Based on the demand forecaster, the dynamically adjusted time window and the policy generator, multimodal transport scheduling plans can be generated and updated in real time, improving the response speed of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 Schematic diagram of the steps of the multimodal transport intelligent scheduling optimization method in an embodiment of the present invention;

[0024] Figure 2 Schematic diagram of the structure of the multimodal transport intelligent scheduling optimization device in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] Embodiments of the present invention provide a method, apparatus, device and storage medium for intelligent scheduling optimization of multimodal transport. The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products or devices.

[0026] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 An embodiment of the multimodal transport intelligent scheduling optimization method in the embodiment of the present invention includes:

[0027] Step S1: Model multiple transportation nodes and transportation paths to obtain a network topology structure;

[0028] It is understandable that the execution subject of the present invention can be a multimodal transport intelligent scheduling optimization device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking the server as the execution subject as an example.

[0029] Specifically, attributes are defined for each transport node to form a comprehensive set of node attributes. This node attribute set encompasses multiple key elements of each node, such as location coordinates, node type, and handling capacity. Location coordinates determine the node's specific location in geographic space, while node type distinguishes the node's functional role within the intermodal transport network, such as whether it is a port, railway station, or distribution center. Handling capacity reflects the upper limit of the cargo volume that a node can handle within a given timeframe, such as loading and unloading speeds and storage capacity. Simultaneously, attributes are defined for each transport route to form a complete set of route attributes. The route attribute set includes the start and end points of the transport route—the two nodes connected by the route. This information helps determine the specific route that cargo takes from one node to another. The mode of transport describes the type of transportation used along the route. Transit time refers to the time required to transport cargo along the route, while transportation cost reflects the economic cost of transporting cargo along the route, including fuel costs, tolls, and other related expenses. The capacity limit defines the maximum cargo volume that the route can handle per unit time. Together, these attributes describe the characteristics of a transport route. Based on node and route attribute information, as well as actual geographic information and transportation network data, the connectivity between transport nodes is determined to form a complete multimodal transport network topology. Historical freight data is collected and organized to determine the inflow and outflow of cargo at each node and the actual transport volume of each transport route. The accumulation and analysis of historical data helps identify frequently used transport routes and nodes prone to bottlenecks within the network. An objective function is constructed based on total transport cost and total transport time to determine the network optimization goal. This objective function is expressed as: F = min(α total transport cost + β total transport time), where α and β are weighting coefficients for transport cost and transport time, respectively. By minimizing the objective function, under certain weightings, the optimal balance between overall transport cost and time is sought to maximize the overall network efficiency. The weighting coefficients α and β in the objective function are initialized to obtain initial weight values. The weighting values ​​are set based on the actual operational requirements and management objectives of the transport network. For example, in time-sensitive transport scenarios, total transport time is given a higher weight, while in cost-sensitive scenarios, total transport cost is given a higher weight. Based on the node and path attribute sets, various network constraints are constructed to obtain the initial network constraints. Constraints include upper limits on node processing capabilities, capacity limits on paths, and priority requirements for specific paths. This ensures that the solution obtained during the optimization process is feasible and meets the physical and management constraints of actual operations. The network optimization objectives, initial weights, and initial network constraints are combined to obtain a complete network topology.

[0030] Step S2: Divide the scheduling period into time periods based on the network topology to obtain a time window scheduling framework;

[0031] Specifically, based on the network topology, the entire scheduling cycle T is evenly divided into several discrete time periods {t1, t2, ..., tK}, where K is the preset number of time periods. Discretizing the continuous scheduling cycle into several time segments capable of executing specific scheduling operations facilitates subsequent scheduling decisions and optimization. The length of each time period is determined based on actual transportation demand and scheduling flexibility, ensuring that scheduling tasks can be effectively processed within a time period without being too long or too short. A state vector is defined for each transportation node Ni in each time period tk, resulting in a node state vector Si,k. The definition of the node state vector includes two key elements: the current cargo volume Qi,k and the remaining processing capacity Ri,k. The current cargo volume Qi,k represents the amount of cargo at node Ni in time period tk, reflecting the node's cargo demand or cargo flow during that time period. The remaining processing capacity Ri,k represents the node's processing margin within that time period, namely, the amount of cargo that the node can still process during the current time period. These two parameters together constitute the node state vector, which dynamically reflects the node's load and processing capacity in each time period. For each transport route Rj, a state vector is defined in each time period tk, resulting in a path state vector Pj,k. The path state vector primarily consists of the current transport volume Vj,k and the remaining capacity Uj,k. The current transport volume Vj,k represents the actual amount of cargo being transported on route Rj in time period tk, while the remaining capacity Uj,k represents the maximum cargo flow that the route can still handle during the current time period. These two metrics provide real-time insights into each route's usage and carrying capacity, enabling scheduling decisions to fully utilize existing transport resources while avoiding overload or resource waste. Decision variables xi,j,k are introduced to represent the cargo volume shipped from node Ni to node Nj in time period tk, resulting in a set of scheduling decision variables. This set of scheduling decision variables describes all possible scheduling operations within the entire scheduling system. By rationally configuring these variables, optimal scheduling is achieved for the entire transportation network. These decision variables are subject to various constraints, including node capacity constraints, path capacity constraints, and cargo balance constraints. Node capacity constraints stipulate that the processing capacity of each node within each time period cannot exceed its upper limit, while path capacity constraints limit the cargo volume of each transportation path within each time period to its maximum capacity. Cargo balance constraints ensure that the inflow and outflow of cargo between nodes and time periods are balanced, meaning that cargo cannot increase or decrease out of thin air. These constraints together constitute the time window scheduling constraint set, ensuring the feasibility of the scheduling scheme in all aspects. To meet the timeliness requirements of actual transportation, each batch of cargo is set with an earliest departure time, TEi, and a latest arrival time, TLi. This results in the time window constraint TEi ≤ tk ≤ TLi - tj, where tj is the transportation time from node i to node j.Time window constraints ensure that each shipment is delivered within the specified timeframe, preventing early or delayed arrival of shipments due to inappropriate scheduling, which could impact overall logistics efficiency. The optimization objective function is refined and formulated as F=min(∑i,j,k(cjxi,j,k) +γ∑imax(0, Qi,K - Qi,0)), where cj represents the unit cost of transporting from node Ni to node Nj, and γ represents the penalty coefficient for shipments not delivered on time. This objective function considers the trade-off between minimizing transportation costs and imposing penalties for shipments not delivered on time. By minimizing transportation costs, transportation costs are minimized while ensuring timeliness. The penalty for shipments not delivered within the last timeframe penalizes shipments not delivered within the last timeframe, encouraging the scheduling strategy to complete shipments within the specified timeframe. The discrete timeframes, node state vectors, path state vectors, a set of scheduling decision variables, a set of time window scheduling constraints, time window constraints, and the optimization objective function are combined to form a complete time window scheduling framework.

[0032] Step S3: Input historical freight data into a multi-layer feedforward neural network for training to obtain a demand forecaster;

[0033] Specifically, historical freight data is preprocessed to convert the raw data into standardized input features that can be effectively learned and understood by the neural network model. These standardized input features include historical freight volumes, seasonal factors, economic indicators, and weather conditions, comprehensively reflecting the multi-dimensional factors influencing freight demand. During preprocessing, the raw data is cleaned to remove outliers and missing values, and all numerical features are normalized to a mean of 0 and a variance of 1 to eliminate the impact of different feature scales on model training. A multi-layer feedforward neural network structure is designed to form the neural network model. This neural network model consists of an input layer, multiple hidden layers, and an output layer. The number of neurons in the input layer matches the dimensionality of the input features, with each neuron corresponding to a single input feature. The hidden layers use a multi-layer structure to capture the complex nonlinear relationships between the input features. The number of neurons in each hidden layer is set based on the complexity of the problem. To enhance the model's expressiveness, the hidden layers use the ReLU activation function. The nonlinear nature of the ReLU function enables the network to learn more complex patterns while also alleviating the vanishing gradient problem. The output layer uses a linear activation function. Because demand forecasting is a regression problem, linear activation functions can directly output predicted values, better reflecting the continuity and changing trends of freight demand. After the model structure is designed, the standardized input features are divided into a training set and a validation set in an 8:2 ratio. This effectively prevents overfitting during model training and facilitates evaluation of model performance on unseen data. The training set data is normalized, scaling all feature values ​​to the range [0, 1]. This smooths the gradients during model training, improving training efficiency and model stability. The normalized training data is input into the neural network model and trained using the backpropagation algorithm and the Adam optimizer. The backpropagation algorithm calculates the gradient of the loss function with respect to the model parameters to guide the update of the model parameters, thereby gradually reducing the prediction error. The Adam optimizer is an adaptive learning rate optimization algorithm that calculates first- and second-order moment estimates to adaptively adjust the learning rate and quickly converge to the global optimal solution. The mean squared error (MSE) loss function is used during training. This loss function measures the average squared difference between the model's predicted value and the true value, reflecting the accuracy of the model's predictions. To prevent overfitting during training, we applied dropout to the model parameters. During each training iteration, we randomly discarded some of the outputs of hidden layer neurons, with a dropout rate of 0.5. This means that 50% of the neuron outputs were temporarily ignored during each training session, preventing the model from becoming overly dependent on the weights of specific neurons. After training, we used k-fold cross-validation to fully evaluate the model, with a k value of 5.Cross-validation divides the dataset into k non-overlapping subsets, sequentially using one of these subsets as the validation set and the remaining k-1 subsets as the training set for training and validation. This effectively reduces random errors introduced by the data partitioning method and yields more stable model performance metrics. Performance metrics include prediction accuracy, mean absolute error, and root mean square error, which measure different aspects of the model's predictive performance. To address the dynamic changes in freight demand and fluctuations in model performance in real-world applications, an adaptive adjustment mechanism is designed to improve the model's predictive performance. This mechanism triggers model retraining when the model's prediction error exceeds a preset threshold, and updates the neural network weights using an adaptive learning rate. The adaptive learning rate is dynamically adjusted using the formula η(t) = η0 / (1+λt), where η0 is the initial learning rate, λ is the decay coefficient, t is the number of training steps, and η(t) is the adaptive learning rate. As the number of training steps t increases, the learning rate gradually decays, enabling the model to more stably converge to a local optimum in the later stages of training. The adaptive adjustment mechanism enables the demand forecaster to maintain high prediction performance in the face of different freight environment changes, and ultimately obtains a demand forecaster model with stable performance and the ability to adapt to actual demand fluctuations.

[0034] Step S4: dynamically adjust the time window of the time window scheduling framework to obtain a dynamically adjusted time window;

[0035] Specifically, a time window adjustment factor δi,k is defined. This factor quantifies the adjustment demand within each time period. The adjustment calculation formula is δi,k = min(ρ*(Di,k - Ci,k), δmax), where Di,k represents the forecasted demand within time period tk, Ci,k represents the current transportation or handling capacity, ρ is the adjustment coefficient, which controls the sensitivity of the adjustment, and δmax is the maximum adjustment, which limits the magnitude of a single adjustment to prevent excessive adjustments from causing scheduling disruptions. This formula considers the gap between actual demand and existing capacity and quantifies it into an actionable adjustment value, providing the basis for subsequent time window adjustments. After calculating the adjustment δi,k, the time window is updated based on this adjustment, resulting in the updated time window. The update formulas are TEi,k+1 = TEi,k - max(0, δi,k) and TLi,k+1 = TLi,k + max(0, δi,k), where TEi,k and TLi,k represent the earliest shipment time and latest arrival time, respectively, in time period tk. By adjusting the earliest shipment time and latest arrival time, the duration of each time window and scheduling flexibility are flexibly controlled to adapt to dynamically changing transportation demand and resource conditions. For example, when the forecast demand exceeds the current capacity, δi,k becomes positive, the earliest shipment time TEi,k decreases accordingly, and the latest arrival time TLi,k increases accordingly, freeing up more time to handle additional demand. When the forecast demand falls below the current capacity, δi,k becomes negative, shortening the time window to reduce resource waste. Constraints are set on the updated time windows to prevent them from over-adjusting beyond a reasonable range. The specific constraints are TEmin ≤ TEi,k ≤ TEmax and TLmin ≤ TLi,k ≤ TLmax, where TEmin and TEmax represent the minimum and maximum limits for the earliest shipment time, and TLmin and TLmax represent the minimum and maximum limits for the latest arrival time. These upper and lower constraints ensure that the dynamic adjustment of the time window occurs within a controllable range, preventing excessively long or short windows from decreasing scheduling efficiency or reducing service levels. At the same time, to achieve better coordination and resource balance between different transportation modes, a multimodal coordination factor μj is introduced. The calculation formula for the multimodal coordination factor is μj = (cj / cmax)*(tj / tmax), where cj is the transportation cost of path j, tj is the transportation time of path j, and cmax and tmax are the maximum cost and maximum time among all paths, respectively. By standardizing and combining the transportation cost and time factors, μj reflects the comprehensive balance between different transportation paths, effectively regulating the allocation and utilization of resources across different modes in scheduling decisions. Integrating the multimodal coordination factor μj into the optimization objective function yields a new optimization objective function F'.The new optimization objective function, expressed as F'=min(∑i,j,k(μjcjxi,j,k)+γ*∑imax(0,Qi,K-Qi,0)), adapts to the coordinated scheduling needs of different modes of transport in an intermodal transport network and improves the overall scheduling efficiency of the system. To evaluate the effectiveness of dynamically adjusted time windows, an evaluation metric, E, is designed, and the evaluation formula, E=w1*(actual cost / expected cost)+w2*(actual time / expected time), is given. w1 and w2 are weight coefficients, representing the importance of cost and time in the evaluation metric, respectively. The effectiveness of time window adjustment is quantified by calculating the ratios of actual cost to expected cost and actual time to expected time. When E>1, the actual cost and time exceed expectations and require further optimization; when E<1, the actual situation exceeds expectations. Based on the results of the evaluation formula, the time window adjustment strategy for the next cycle is adaptively adjusted to more accurately match changes in transportation demand and resource supply. When E > 1, it indicates that both actual costs and time exceed expectations. In this case, the time window adjustment for the next cycle should be increased to better adapt to changing demand and alleviate resource pressure. When E < 1, the current scheduling strategy is ideal, and adjustments should be appropriately reduced to maintain stable system operation. The adaptive adjustment strategy can flexibly adjust the time window based on actual operating conditions, improving the dynamic responsiveness of the scheduling strategy. Combining the constrained time window, the new optimization objective function, and the adaptive adjustment strategy yields a dynamically adjusted time window. This time window can dynamically adapt to changes in transportation demand within different time periods and achieve efficient resource allocation across different modes of transport. Through continuous optimization through the adaptive adjustment strategy, the overall scheduling level and operational efficiency of the intermodal transport system are improved.

[0036] Step S5: input the state space and action space into the improved double-delayed deep deterministic policy gradient algorithm for training to obtain a policy generator;

[0037] Specifically, a state space S and an action space A are defined. The state space S includes the current cargo volume of all transport nodes, the current cargo volume of all transport routes, and time window information, reflecting the current state of the entire intermodal transport network. The action space A represents the decision on how to allocate cargo volume to each route—that is, the decision to transport a specific amount of cargo from one node to another via a specific route within each time period. The definition of the action space determines the scheduling system's behavior under different states. A deep neural network is constructed to approximate the policy and value functions. In the improved DDPG algorithm, an actor network fθ(s) and two critic networks Qφ1(s,a) and Qφ2(s,a) are constructed. These networks are used to estimate the policy value and value of taking an action a given state s. The actor network fθ(s) takes the current state s as input and outputs a specific action a, namely, the policy π(s). The critic networks Qφ1 and Qφ2 are used to estimate the Q-value of a given state-action pair (s,a). Through this structural design, the network continuously adjusts the policy π(s) and value function Q(s,a) during training to learn the optimal scheduling policy. To improve the stability and convergence speed of network training, the target network fθ' and Qφ1' and Qφ2' are introduced. The parameters of the target network are updated via a soft update mechanism. The soft update mechanism uses the parameter update formulas θ'←τθ+(1-τ)θ' and φi'←τφi+(1-τ)φi', where τ is the soft update coefficient. The soft update mechanism avoids training instability caused by rapid changes in the target network parameters while ensuring smooth convergence during training. During action generation, Gaussian noise ε is introduced to explore more of the state-action space during the initial training phase. This noise is added to the policy generation to obtain the exploration policy μ(s). The specific calculation formula is μ(s)=π(s)+ε, where ε follows a normal distribution with mean 0 and standard deviation σ. By adding noise to the action generation, the policy generator explores as much of the state space as possible during the initial training phase, thus avoiding falling into local optima. An adaptive noise attenuation mechanism is designed to control the balance between exploration and exploitation during training. The noise standard deviation update formula is σt=σ0*exp(-λt), where σ0 is the initial noise standard deviation, λ is the decay rate, and t is the number of training steps. As training progresses, the noise standard deviation gradually decreases, and the policy generator gradually shifts from extensive exploration in the early stage to precise policy optimization in the later stage. In order to make full use of the existing experience data, an experience replay buffer D is defined. This buffer is used to store the state-action-reward-next-state quadruple accumulated during training. During training, a sampling mechanism is adopted to randomly extract mini-batch data from D for training, breaking the temporal correlation between data and improving the efficiency and stability of training.Randomly sampled mini-batches of data include multiple state-action pairs, their corresponding rewards, and next-state information. Batch sampling training allows for more comprehensive updates to network parameters, improving model generalization. In the critic network, the target Q-value is calculated as y=r+ω*min(Qφ1'(s',π(s')),Qφ2'(s',π(s'))), where r is the current reward, s' is the next state after executing action a, and ω represents the discount factor. The critic network's parameters φi are updated by minimizing the difference between the target Q-value and the current Q-value. The specific update formula is φi←φi-η▽φiL(φi), where L(φi) is the loss function. Simultaneously, the actor network's update formula is θ←θ+η▽aQφ1(s,a)|a=π(s)▽θπ(s). This updates the policy network's parameters θ by maximizing the Q-value, improving the quality of the actions generated by the policy network. To ensure adaptive learning rate adjustment during training, an adaptive learning rate adjustment mechanism is designed. The learning rate update formula is η=η0 / (1+β*t), where η0 is the initial learning rate, β is the decay coefficient, and t is the number of training steps. As the number of training steps increases, the learning rate gradually decreases, and the pace of network parameter updates becomes more gradual, making the training process more stable and avoiding oscillations. The neural network structure, soft update mechanism, exploration strategy, noise standard deviation update formula, experience replay sampling mechanism, network update formula, and learning rate update formula are organically combined to form a complete policy generator. Based on the input state space and action space, this policy generator can continuously adjust and optimize the multimodal transport scheduling strategy through a deep reinforcement learning training process, realizing intelligent scheduling decisions for multiple transportation modes in complex transportation networks.

[0038] Step S6: Based on the demand forecaster, the dynamically adjusted time window and the strategy generator, a multimodal transport scheduling decision is generated to obtain a multimodal transport scheduling solution.

[0039] Specifically, key system parameters are initialized, including the basic structure of the intermodal transport network, the initial time window settings, and the initial weights of the neural network used for policy generation. Network structure initialization involves determining the number of transport nodes in the intermodal transport system and the path connections between them. These connections reflect the connectivity and path selection between different transport modes. Time window initialization includes the initial earliest departure time (TE) and latest arrival time (TL) for each node. The time window provides the time boundary for subsequent scheduling decisions. Neural network weight initialization includes the initial weights of the actor network and the two critic networks. These weights influence the policy generator's exploration behavior during the initial training phase. The current system state is input into the demand forecaster to predict freight demand and capacity supply for the next cycle. The current state includes the current cargo volume at each node, the current transport volume on each transport route, and the current time window. This information reflects the real-time operating status of the system. Based on historical data and current state information, the demand forecaster uses a multi-layer feedforward neural network to predict freight demand and capacity supply for the next cycle. The prediction results include the inflow and outflow of goods for each node in the future time period, as well as the future capacity supply of each path. Based on the prediction results of the demand forecaster, the dynamically adjusted time window is updated to obtain a more accurate time window setting. The update process uses the time window adjustment formula proposed in claim 5, that is, the earliest shipment time TE and the latest arrival time TL are adaptively adjusted by adjusting the factors δi,k. When the predicted demand for a node in the future time period exceeds the current capacity, the earliest shipment time of the node will be reduced or the latest arrival time will be increased to provide a time window for more transportation needs, and vice versa. Through dynamic adjustment, the time window can respond more flexibly to fluctuations in transportation demand and improve the scheduling efficiency and service level of the entire system. After the time window update is completed, the current state st is input into the policy generator to generate the scheduling decision at under the current state. The policy generator learns and generates policies based on a modified double-delayed deep deterministic policy gradient (DDPG) algorithm. The specific calculation process is at = μ(st) = π(st) + ε, where π(st) is the scheduling policy generated by the actor network based on the current state st, and ε is the exploration noise, which follows a normal distribution with mean 0 and standard deviation σ. The introduction of exploration noise enhances the policy generator's exploration capability during training and prevents premature convergence to local optima. The generated scheduling decision at includes the transport volume allocation for each transport route in each time period, specifically the amount of cargo to be transported from one node to another during different time periods. Based on the scheduling decision, the transportation process of the entire intermodal transport system is simulated. The simulation process includes updating the cargo volume at each node and the transport volume of each route, and calculating the total transportation cost and total transportation time.During each time period, based on the scheduling decision, the cargo volume on each route is updated to the corresponding node, and the transportation cost and time overhead incurred by the scheduling solution are recorded. This simulation process helps verify the effectiveness of the current scheduling strategy and provides reference data for subsequent strategy optimization. After the transportation simulation is complete, the reward value rt for the scheduling strategy is calculated, and the expression of the reward function is derived. The reward function is designed as rt = -F'(st,at), where F' is the new optimization objective function. By minimizing total transportation cost and total transportation time and incorporating resource utilization and service level into considerations, the reward function effectively reflects the performance of the scheduling strategy. The introduction of the negative sign indicates that the system aims to minimize cost and time, maximizing the reward value to guide the strategy generator towards optimal system performance. After obtaining the reward value for the current decision, the new state st+1 after executing the current scheduling decision is observed, and the cargo volume at each node, the transportation volume of each route, and the time window information are updated to form a complete state transition sample (st,at,rt,st+1). State transition samples fully record the entire process from the current state to the new state through the execution of a scheduling decision, including the gains or losses resulting from the scheduling decision. State transition samples are stored in the experience replay buffer D. The size of the experience pool is a preset value M. When the experience pool is full, old samples are replaced by new ones to ensure sample diversity. This mechanism effectively avoids temporal correlation issues with samples and improves the learning effect of the policy generator. During training, N mini-batches of data are randomly sampled from the updated experience pool to obtain a training sample set. These N randomly sampled sample sets include several state transition samples (st, at, rt, st+1). These samples provide a rich set of state-action pairs and their corresponding feedback information, thus supporting the training process of the policy generator. Based on the training sample set, the network parameters of the policy generator are updated. The update process first calculates the target Q-value using the formula y=r+γ*min(Qφ1'(s',π(s')),Qφ2'(s',π(s'))), where r is the current reward, γ is the discount factor, Qφ1' and Qφ2' represent the target critic network, and π(s') represents the target actor network. The critic network parameters φi are updated by minimizing the difference between the target Q-value and the actual Q-value using the formula φi←φi-η▽φiL(φi), where L(φi) is the loss function and η is the learning rate. The actor network parameters are updated every d steps using the formula θ←θ+η▽aQφ1(s,a)|a=π(s)▽θπ(s). This update maximizes the Q-value to update the policy network parameters θ, thereby improving the quality of the actions generated by the policy network. The updated policy generator is then applied to the entire intermodal transport system to generate a specific scheduling solution.In the process of generating a scheduling plan, the cargo volume of each node and the transportation capacity of each route are taken into consideration, as well as the dynamic adjustment of the time window and the prediction of future demand, so that the scheduling plan can be more in line with the actual situation.

[0040] In an embodiment of the present invention, a complete network topology is obtained by comprehensively modeling multiple transport nodes and transport routes. The scheduling cycle is divided into time periods based on the network topology, and a dynamic adjustment mechanism is introduced to enable the scheduling plan to better adapt to changes in actual transportation demand. A multi-layer feedforward neural network is used to train historical freight data, and an adaptive adjustment mechanism is introduced to improve the accuracy of predicting future freight demand. The policy generator obtained by training with an improved double-delayed deep deterministic policy gradient algorithm is capable of generating high-quality scheduling decisions in complex state spaces. By introducing multimodal coordination factors and effect evaluation indicators, the balance between different modes of transportation is considered while optimizing total transportation cost and time. Mechanisms such as dynamic time window adjustment and adaptive learning rate are adopted to enable the scheduling system to better adapt to environmental changes and uncertainties. Based on the demand forecaster, the dynamically adjusted time window, and the policy generator, multimodal transport scheduling plans can be generated and updated in real time, improving the system's response speed.

[0041] In a specific embodiment, the process of executing step S1 may specifically include the following steps:

[0042] Define the attributes of each transport node to obtain a node attribute set, which includes location coordinates, node type and processing capacity indicators;

[0043] Define the attributes of each transport route to obtain a set of route attributes, which includes the starting point, end point, transport mode, transport time, transport cost and transport capacity upper limit;

[0044] Based on actual geographic information and transportation network data, the connection relationship between each node is determined to obtain a complete multimodal transport network topology. Historical freight data is collected and organized to obtain the cargo inflow and outflow of each node and the actual transportation volume of each transportation route.

[0045] Based on the total transportation cost and total transportation time, the objective function is constructed to obtain the network optimization goal. The objective function is F=min(α*total transportation cost+β*total transportation time);

[0046] Initialize the weight coefficients α and β in the objective function to obtain the initial weight values, and construct network constraints based on the node attribute set and the path attribute set to obtain the initial network constraints;

[0047] The network optimization objective, initial weight value and initial network constraint are combined to obtain the network topology.

[0048] Specifically, attributes are defined for each transport node to form a node attribute set. The construction of a node attribute set typically includes three core elements: location coordinates, node type, and processing capacity indicators. Location coordinates describe the specific location of each node in geographic space. Expressed in latitude and longitude coordinates, they accurately reflect the distribution of nodes in the actual geographic environment. Node types distinguish different functional roles of nodes, such as ports, railway hubs, and distribution centers. Different types of nodes have different functions and roles and participate in different modes of transportation. Processing capacity indicators represent the volume of cargo each node can handle per unit time, reflecting the node's loading, unloading, storage, and transshipment capabilities. Attributes are defined for each transport route to construct a route attribute set. The route attribute set encompasses multiple key data points in the transport process, including origin and destination, transport mode, transport time, transport cost, and capacity ceiling. Based on actual geographic information and transportation network data, the connectivity between nodes is determined to obtain a complete multimodal transport network topology. Multiple data sources, such as geographic information system (GIS) data, traffic management data, and historical transportation data, are integrated to accurately describe the physical connectivity between nodes. For each node, the number and type of connected paths directly affect the flexibility and efficiency of transportation. By integrating this information, a comprehensive network topology is formed to reflect the actual transportation network layout and function. Historical freight data is collected and organized to obtain the cargo inflow and outflow of each node, as well as the actual transportation volume of each transportation path. Historical data helps to identify frequently used paths and nodes, predict future transportation demand, and provide real operational data for network optimization. In order to optimize the entire transportation network, an objective function is constructed to measure the pros and cons of different options. The objective function should take into account the total transportation cost and total transportation time, and the goal is to minimize the weighted sum of these two factors. Define the objective function as:

[0049] ;

[0050] in, represents the total transportation cost, Indicates the total transport time, and is the corresponding weight coefficient. Total transportation cost The cost of transporting goods on all routes is calculated by multiplying the transport cost of each route by the transport volume. For example, for a route , if the transportation cost is , the amount of cargo transported on the route is , then the total cost of the path is Total shipping time The total time required to transport goods on all routes is the transport time multiplied by the transport volume to reflect the time burden of the route. The objective function optimizes the entire system by adjusting the balance between transport cost and time. In order to enhance the flexibility of the objective function, the weight coefficient and Perform initialization settings. For example, in some time-sensitive transportation tasks, setting Larger, and Small, thus giving priority to shortening transportation time; on the contrary, in cost-sensitive scenarios, increasing The weights are set to reduce total transportation costs. Reasonable weight settings balance the system's needs in different scenarios, improving the flexibility and accuracy of optimization. Based on the node attribute set and the path attribute set, corresponding network constraints are set. Initial network constraints typically include the node's processing capacity limit and the path's upper capacity limit. For example, if a node has a daily processing capacity of 5,000 tons, the total inflow and outflow of cargo from the node cannot exceed this upper limit. Similarly, the path's upper capacity limit cannot be exceeded, otherwise it will cause transportation delays or system crashes. Constraints ensure the feasibility of the optimization solution in actual operations and avoid unreasonable scheduling arrangements. Other constraints include the priority of specific paths, time window restrictions, etc. The constructed network optimization objective, initialized weight values, and initial network constraints are combined to obtain a complete network topology that reflects the nodes, paths, and their attributes in the multimodal transport system. The setting of objective functions and constraints provides information for the system's optimal scheduling.

[0051] In a specific embodiment, the process of executing step S2 may specifically include the following steps:

[0052] The entire scheduling period T is evenly divided into K discrete time periods {t1, t2, ..., tK}, where K is the number of preset time periods;

[0053] Define the state vector for each transport node Ni in each time period tk, and obtain the node state vector Si,k, which includes the current cargo volume Qi,k and the remaining processing capacity Ri,k;

[0054] Define the state vector for each transport route Rj in each time period tk to obtain the route state vector Pj,k, which includes the current transport volume Vj,k and the remaining transport capacity Uj,k;

[0055] Introduce decision variables xi,j,k to represent the amount of cargo shipped from node Ni to node Nj in time period tk, and obtain the set of scheduling decision variables;

[0056] Based on the node capacity constraints, path capacity constraints and cargo balance constraints, the constraints are constructed to obtain the time window scheduling constraint set.

[0057] For each batch of goods, the earliest shipping time TEi and the latest arrival time TLi are set, and the time window constraint TEi≤tk≤TLi-tj is obtained, where tj is the transportation time from node i to node j;

[0058] The objective function is refined into F=min(∑i,j,k(cj*xi,j,k)+γ*∑imax(0,Qi,K-Qi,0)), and the optimization objective function is obtained, where γ is the penalty coefficient for goods that are not transported on time;

[0059] The discrete time period, node state vector, path state vector, scheduling decision variable set, time window scheduling constraint set, time window constraints and optimization objective function are combined to obtain a time window scheduling framework.

[0060] Specifically, the entire scheduling cycle Divide evenly and get discrete time periods ,in The number of preset time periods. The continuous scheduling cycle is divided into several discrete time slices, so that each time period can make independent scheduling decisions. Discretization processing makes it easy to transform complex continuous scheduling problems into several easy-to-solve discrete optimization problems. For each transportation node In each time period Define the state vector Get the node state vector. Node state vector Including current cargo volume and remaining processing capacity Current cargo volume refers to the time period node The total amount of cargo on the node reflects the cargo reserve status of the node. It indicates the remaining cargo volume that the node can handle during the time period, that is, the cargo load that the node can still bear in transportation, loading and unloading operations. In the time period The current cargo volume of the node is 500 tons, and its maximum handling capacity is 1000 tons / hour. If there are no other influencing factors, the node The remaining processing capacity is 500 tons. Through the definition of state vector, the cargo status and processing capacity of each node are described in different time periods, providing accurate data support for scheduling decisions. Similarly, for each transportation path In each time period Define the state vector Get the path state vector. Path state vector Including current transport volume and remaining capacity Current transport volume refers to the time period path The actual amount of cargo transported on the route reflects the actual load of the route. It refers to the maximum amount of cargo that can still be transported by the route within the current time period. For example, if a railway transport route The upper limit of the transport capacity is 1000 tons / hour, and the current transport volume is 800 tons. By defining the path state vector, the usage of each transportation path is dynamically monitored to ensure that there is no overload or waste of resources during the transportation process. Indicates the time period Slave nodes Sent to node The amount of goods. The set of decision variables It includes the transportation decisions between all node pairs and is the core variable of the scheduling optimization problem. By optimizing these variables, the distribution and transportation volume of goods between each node in each time period are determined, and the optimal scheduling of the entire multimodal transport system is achieved. In order to make the scheduling plan meet various constraints in actual operation, constraints are constructed based on node capacity constraints, path capacity constraints and cargo balance constraints to obtain the time window scheduling constraint set. The node capacity constraint stipulates that the inflow and outflow of goods at each node in each time period cannot exceed its processing capacity, that is, , which represents the node The total amount of cargo inflow minus the total amount of cargo outflow cannot exceed its remaining processing capacity. The path capacity constraint requires that the transportation volume on each transportation path cannot exceed its capacity limit, that is, , which represents the node To Node The cargo volume cannot exceed the path The cargo balance constraint requires that in each time period, the total inflow and outflow of cargo at all nodes should be balanced, that is, , which represents the node In the time period The total outflow of goods should be equal to its current cargo volume. Constraints ensure that the scheduling plan is feasible in actual execution and complies with the physical and management limitations of the system. On the basis of satisfying the above constraints, the earliest shipping time is set for each batch of goods. and latest arrival time , get the time window constraint ,in For slave nodes To Node The time window constraint ensures that each batch of goods can complete the transportation task within the specified time range, avoiding the situation where goods arrive early or late due to unreasonable scheduling arrangements, thereby improving the logistics service level. The optimization objective function is refined to minimize transportation costs and time during the scheduling process. The optimization objective function is expressed as:

[0061] ;

[0062] in, Indicates the path The transportation cost from node To Node The unit transportation cost, is a decision variable, representing the time period Slave nodes Sent to node volume of cargo. is the penalty coefficient for goods that are not transported on time, reflecting the cost of delayed or unfinished tasks. and Represents nodes respectively The amount of goods in the last time period and the initial time period. The first part of the objective function reflects the minimization of total transportation costs, the second part This represents a penalty for failing to complete a transport task on time, forcing the system to complete cargo transport tasks as timely as possible. By solving the optimization objective function, the optimal cargo transport scheduling solution is obtained while satisfying all constraints. A comprehensive time window scheduling framework is formed by organically combining discrete time periods, node state vectors, path state vectors, a set of scheduling decision variables, a set of time window scheduling constraints, time window constraints, and the optimization objective function. This framework dynamically adjusts the scheduling solutions for each node and path based on the actual operation of the multimodal transport system, ensuring that cargo is transported within the optimal time and cost range.

[0063] In a specific embodiment, the process of executing step S3 may specifically include the following steps:

[0064] Preprocess historical freight data to obtain standardized input features, including historical freight volumes, seasonal factors, economic indicators, and weather conditions;

[0065] Design a multi-layer feedforward neural network structure to obtain a neural network model. The neural network model includes an input layer, multiple hidden layers, and an output layer. The hidden layer uses a ReLU activation function, and the output layer uses a linear activation function.

[0066] The standardized input features are divided into a training set and a validation set in a ratio of 8:2. The training set data is normalized to obtain the normalized training data. The normalization process scales all feature values ​​to the range of [0, 1].

[0067] The normalized training data is input into the neural network model and trained using the back propagation algorithm and Adam optimizer to obtain the initial model parameters. The mean square error is used as the loss function in the training process.

[0068] Apply dropout technology to the initial model parameters to obtain model parameters that prevent overfitting, and the dropout rate is set to 0.5;

[0069] The k-fold cross-validation method was used to evaluate the model and obtain the model performance indicators. The k value was set to 5. The performance indicators included prediction accuracy, mean absolute error, and root mean square error.

[0070] An adaptive adjustment mechanism is designed to obtain a demand forecaster. The adaptive adjustment mechanism includes: when the prediction error exceeds a preset threshold, model retraining is triggered, and the neural network weights are updated using an adaptive learning rate, where the adaptive learning rate η(t)=η0 / (1+λt), η0 is the initial learning rate, λ is the decay coefficient, t is the number of training steps, and η(t) is the adaptive learning rate.

[0071] Specifically, historical freight data is preprocessed to convert the raw data into standardized input features that can be understood and utilized by the neural network model. Input features include historical freight volumes, seasonal factors, economic indicators, and weather conditions. Historical freight volumes reflect the freight transportation situation at each node and route over different time periods. Seasonal factors capture fluctuations in transportation demand due to seasonal changes. Economic indicators include macroeconomic data such as GDP growth rate and industrial production index, which reflect the impact of macroeconomic conditions on freight demand. Weather conditions include information such as temperature, rainfall, and wind speed. During preprocessing, the data is cleaned and outliers are processed, and missing values ​​are filled or deleted. All numerical features are normalized to a mean of 0 and a variance of 1 to eliminate the impact of scale differences between features on model training. After normalization, the input features are better utilized by the neural network model, improving training efficiency and prediction accuracy. A multi-layer feedforward neural network structure is designed to construct the neural network model. The model consists of an input layer, multiple hidden layers, and an output layer. The number of neurons in the input layer matches the dimensionality of the standardized input features, with each neuron receiving a single feature value. The hidden layer is composed of several layers, each containing a certain number of neurons, which are used to capture the complex nonlinear relationship between input features. The activation function of the hidden layer uses the ReLU function, that is, ReLU , the activation function can effectively alleviate the gradient vanishing problem and improve the training speed and expression ability of the model. The output layer uses a linear activation function to output the predicted freight demand value. The form of the linear activation function is , suitable for regression problems, and can directly output predicted values ​​without additional nonlinear transformations. Through this network structure design, the model can learn complex mapping relationships under different feature combinations, providing accurate model support for demand forecasting. After completing the design of the neural network model, the standardized input features are divided into training sets and validation sets in a ratio of 8:2, that is, 80% of the data is used for model training and 20% of the data is used for model validation. The training set is used to learn model parameters, while the validation set is used to evaluate the model's generalization ability and performance on unseen data. In order to improve the training efficiency of the model, the training set data is normalized, that is, all feature values ​​are scaled to the [0,1] interval. The formula for normalization is:

[0072]

[0073] in, represents the original eigenvalue, and are the minimum and maximum values ​​of the feature, respectively. is the normalized eigenvalue. Normalization ensures that all input eigenvalues ​​are within the same numerical range, which helps the model converge faster and prevents eigenvalues ​​that are too large or too small from adversely affecting gradient calculations during model training. The normalized training data is input into the neural network model and trained using the backpropagation algorithm and the Adam optimizer. The backpropagation algorithm calculates the gradient of the loss function relative to the model parameters and gradually adjusts the model parameters to minimize the loss function. The mean square error is used as the loss function, and its formula is:

[0074] ;

[0075] in, is the sample size, is the true value, is the model's predicted value. The mean squared error measures the average squared difference between the model's predicted value and the true value. The Adam optimizer is an adaptive learning rate optimization algorithm that adjusts the learning rate based on the first and second moment estimates of the gradient at each parameter update. The Adam optimizer calculates the gradient , that is, the loss function for the current model parameters Partial derivatives of :

[0076] ;

[0077] in, Indicates that the loss function is at the current parameter The value at . The first-order moment (momentum) of the Adam optimizer on the gradient and the second moment (the square of the gradient) Perform exponentially weighted moving average, the formula is as follows:

[0078] ;

[0079]

[0080] in, and are the decay rates of the first-order and second-order moment estimates, which are generally set to 0.9 and 0.999. and Perform bias correction to obtain the corrected value:

[0081] ;

[0082] ;

[0083] Adam optimizer uses the corrected gradient momentum and the squared gradient after correction Update model parameters

[0084] ;

[0085] in, is the learning rate, is a very small constant (usually ) to prevent division by zero errors. The parameter update formula maintains stable updates under different gradient scales, avoiding training instability caused by excessively large or small learning rates. After model training is complete, dropout is applied to the initial model parameters to prevent overfitting. Dropout is a regularization method that randomly discards some neurons. In each training iteration, the outputs of some neurons are randomly set to zero with a certain probability, reducing the model's dependence on certain specific features and improving its generalization ability. Setting the dropout rate to 0.5 means that in each training iteration, 50% of the neurons are randomly blocked, leaving only the remaining half participating in training. To comprehensively evaluate model performance, k-fold cross-validation is used. k-fold cross-validation divides the dataset into k non-overlapping subsets, sequentially using one of the subsets as the validation set and the remaining k-1 subsets as the training set for training and validation. This is repeated k times to effectively reduce random errors introduced by data partitioning. Choosing k=5 means that the data is divided into five subsets, with one subset used as the validation set and the other four used as the training set. Through cross-validation, more stable model performance indicators are obtained, including prediction accuracy, mean absolute error, and root mean square error. The formula for mean absolute error is:

[0086] ;

[0087] in, is the sample size, is the true value, For the predicted value, MAE measures the average absolute difference between the predicted value and the true value. The formula for the root mean square error is:

[0088] ;

[0089] RMSE is the square root of the mean square error and can more intuitively reflect the average size of the prediction error. These two indicators are used to comprehensively evaluate the prediction performance of the model in different aspects. To improve the dynamic adaptability of the model, an adaptive adjustment mechanism is designed. When the model's prediction error exceeds a preset threshold, the model is retrained and the weights of the neural network are updated using an adaptive learning rate. The update formula for the adaptive learning rate is:

[0090] ;

[0091] in, is the initial learning rate, is the learning rate decay coefficient, is the number of training steps, is the adaptive learning rate. As the number of training steps increases As the learning rate increases, it gradually decays, allowing the model to more stably converge to the optimal solution in the later stages of training. Through an adaptive adjustment mechanism, when the model prediction error is large, the learning rate and weight update strategy are dynamically adjusted to continuously improve the model's prediction accuracy and stability. This results in a demand forecaster with stable performance that can adapt to changing environments.

[0092] In a specific embodiment, the process of executing step S4 may specifically include the following steps:

[0093] Define the time window adjustment factor δi,k and obtain the adjustment amount calculation formula: δi,k=min(ρ*(Di,k-Ci,k),δmax), where Di,k is the predicted demand, Ci,k is the current capacity, ρ is the adjustment coefficient, and δmax is the maximum adjustment amount.

[0094] The time window is updated based on the adjustment calculation formula to obtain the updated time window. The update formula is TEi,k+1=TEi,k-max(0,δi,k), TLi,k+1=TLi,k+max(0,δi,k), where TEi,k is the earliest shipment time and TLi,k is the latest arrival time.

[0095] Adjust the upper and lower limits of the updated time window to obtain the constrained time window. The constraints are TEmin≤TEi,k≤TEmax, TLmin≤TLi,k≤TLmax;

[0096] Introducing the multimodal coordination factor μj, the multimodal balance formula is obtained. The multimodal balance formula is μj=(cj / cmax)*(tj / tmax), where cj is the transportation cost of path j, tj is the transportation time of path j, cmax and tmax are the maximum cost and maximum time among all paths respectively;

[0097] Integrate the multi-mode coordination factor into the objective function to obtain a new optimization objective function F', which is F'=min(∑i,j,k(μj*cj*xi,j,k)+γ*∑imax(0,Qi,K-Qi,0));

[0098] Design effect evaluation index E and obtain the evaluation formula: E=w1*(actual cost / expected cost)+w2*(actual time / expected time), where w1 and w2 are weight coefficients;

[0099] Based on the calculation results of the evaluation formula, the time window adjustment strategy for the next cycle is adjusted to obtain an adaptive adjustment strategy. When E>1, the time window adjustment strength for the next cycle is increased; when E<1, the adjustment amplitude for the next cycle is reduced.

[0100] The constrained time window, the new optimization objective function and the adaptive adjustment strategy are combined to obtain the dynamically adjusted time window.

[0101] Specifically, define a time window adjustment factor , used to quantify the measure of time window adjustment at different nodes and time periods. The calculation formula of the time window adjustment factor is:

[0102] ;

[0103] in, Indicates the time period Internal Node The forecast demand represents how much cargo is expected to be handled or transported during that time period. Indicates the processing capacity of the current node, reflecting the maximum processing capacity of the node in this time period. The sensitivity of the control adjustment determines how quickly the adjustment changes when there is a difference between the demand and the processing capacity. A higher setting makes the response more sensitive to differences in demand and capacity. The maximum adjustment is used to limit the amplitude of each adjustment to prevent excessive adjustment from causing scheduling chaos. Through this formula, the time window adjustment required at each time period and each node is determined. Greater than current processing capacity When the adjustment amount A positive value indicates that the time window needs to be increased; conversely, a zero or negative value indicates that the time window does not need to be increased or may even be decreased. Based on the above adjustment amount calculation formula, the time window of each node is updated to obtain the updated time window. The update process is as follows:

[0104] ;

[0105] ;

[0106] in, and Represents nodes respectively In the time period The earliest shipping time and the latest arrival time. Update the time window, when adjusting the amount When it is a positive value, the earliest shipping time Decrease means that it is necessary to ship in advance to accommodate more transportation needs; the latest arrival time Increasing means allowing the goods to arrive in a longer time to disperse the transportation pressure. If the adjustment amount is zero or negative, the time window remains unchanged or shortened. After updating the time window, set upper and lower bounds on the updated time window to avoid excessive adjustment of the time window and unreasonable scheduling. The constraints are:

[0107] ;

[0108] ;

[0109] in, and are the minimum and maximum limits of the earliest shipping time, and The upper and lower limits ensure that the time window adjustment is carried out within a controllable range, preventing excessively long or short time windows from affecting scheduling efficiency and transportation service levels. In the time window adjustment strategy, in order to coordinate the relationship between multiple transportation modes, a multi-mode coordination factor is introduced. The multimodal coordination factor is used to balance the cost and time factors of different transportation modes. Its calculation formula is:

[0110] ;

[0111] in, Indicates the path The unit transportation cost, Indicates the path transportation time, and The maximum transportation cost and maximum transportation time among all paths are quantified into a comprehensive indicator through standardization. The larger the index is, the higher the comprehensive transportation cost and time of the route is; otherwise, it is lower. The introduction of multimodal coordination factors can better balance the resource allocation between different modes of transportation in scheduling decisions and avoid the overuse or neglect of a certain mode of transportation. Integrate it into the optimization objective function to obtain a new optimization objective function:

[0112] ;

[0113] in, Indicates time period Slave nodes Sent to node The decision variable of cargo volume is The penalty coefficient for goods that are not transported on time. and Represents nodes respectively The amount of cargo in the last time period and the initial time period. The first part of the optimization objective function It represents the minimization of the total transportation cost after the multimodal coordination factor is introduced, taking into account the cost and time factors of different paths. It is a penalty for failing to complete the transportation task in time, forcing the system to complete the transportation task as soon as possible while ensuring the lowest cost. In order to evaluate the effectiveness of the time window adjustment strategy, the effect evaluation index is designed. , and its calculation formula is

[0114] ;

[0115] in, and Respectively represent the weight coefficients of actual cost and actual time in the evaluation index. Actual cost and expected cost are the total cost of the current actual execution of the scheduling plan and the total cost predicted by the optimization model, respectively. Actual time and expected time are the total transportation time of the current actual Su line and the total transportation time predicted by the model. When the value of is greater than 1, it means that the actual execution cost and time are beyond expectations and the scheduling strategy needs to be adjusted; When it is less than 1, it means that the actual implementation is better than expected, and the adjustment range should be appropriately reduced. The calculation results of , the time window adjustment strategy of the next cycle is adaptively adjusted. When , it means that the current strategy cannot effectively meet the actual needs, and it is necessary to increase the time window adjustment strength of the next cycle, that is, to increase the time window adjustment factor The value of When the current scheduling strategy is better than expected, the adjustment effort can be appropriately reduced to maintain system stability and scheduling flexibility. The adaptive adjustment strategy can dynamically adjust the time window based on actual operating conditions, thereby improving the responsiveness and operational efficiency of the scheduling system. The constrained time window, the new optimization objective function, and the adaptive adjustment strategy are organically combined to form a complete dynamic time window scheduling framework. This framework can flexibly adjust the time windows of each node and path within different time periods and achieve efficient resource allocation across different transportation modes. The adaptive adjustment strategy continuously optimizes the adjustment amplitude and direction of the time window.

[0116] In a specific embodiment, the process of executing step S5 may specifically include the following steps:

[0117] Define the state space S and action space A to obtain the complete input space, where the state space contains the current cargo volume of all nodes, the current transportation volume of all routes, and the time window information, and the action space represents the transportation volume allocation decision for each route;

[0118] Construct the Actor network fθ(s) and two Critic networks Qφ1(s,a) and Qφ2(s,a) to obtain the neural network structure, where θ and φ are network parameters respectively;

[0119] Introducing the target network fθ' and Qφ1', Qφ2', we get the soft update mechanism, which uses the parameter update formula θ'←τθ+(1-τ)θ', φi'←τφi+(1-τ)φi', where τ is the soft update coefficient;

[0120] Gaussian noise ε is added to the action generation process to obtain the exploration strategy μ(s). The calculation formula of the exploration strategy is μ(s)=π(s)+ε, where ε follows a normal distribution with mean 0 and standard deviation σ.

[0121] Design an adaptive noise attenuation mechanism and obtain the noise standard deviation update formula σt=σ0*exp(-λt), where σ0 is the initial noise standard deviation, λ is the attenuation rate, and t is the number of training steps;

[0122] Define the experience replay buffer D and obtain the sampling mechanism, which randomly extracts mini-batch data from D for training;

[0123] Calculate the target Q value y=r+ω*min(Qφ1'(s',π(s')),Qφ2'(s',π(s'))) and get the network update formula. The update formula includes the Critic network update φi←φi-η▽φiL(φi) and the Actor network update θ←θ+η▽aQφ1(s,a)|a=π(s)▽θπ(s), where ω represents the discount factor and s' is the next state after executing action a.

[0124] Design an adaptive learning rate adjustment mechanism and obtain the learning rate update formula η=η0 / (1+β*t), where η0 is the initial learning rate, β is the decay coefficient, and t is the number of training steps;

[0125] The neural network structure, soft update mechanism, exploration strategy, noise standard deviation update formula, sampling mechanism, network update formula and learning rate update formula are combined to obtain a strategy generator.

[0126] Specifically, define the state space and action space , describing the current state of the multimodal transport system and possible decision operations. State space It contains the current cargo volume of all transportation nodes, the current transportation volume of all transportation routes, and time window information. The current cargo volume of each node represents the amount of cargo waiting to be processed or transported at the node, reflecting the node's load; the current transportation volume of each route represents the transportation tasks currently scheduled on the route, reflecting the utilization of the route; the time window information includes the earliest shipping time and the latest arrival time of each node, which is used to determine the timeliness constraints of each batch of cargo during the transportation process. These state variables together constitute the state space of the system at a certain moment. , which can fully describe the current operating status of the entire transportation network. Action space It represents the decision on the transportation volume allocation for each path, that is, how much cargo should be transported from one node to another node via a certain path in the current state. By defining the state space and action space, a complete input space is constructed. Constructing the neural network structure in deep reinforcement learning, namely the Actor network and two Critic networks and ,in and Represents the parameters of the network respectively. Actor network The role is to receive the current state As input, and output an action , that is, strategy , represents the optimal transport volume allocation decision critic network under the current state and It is used to evaluate a given state-action pair The value, that is, Q value, reflects the current state Next action The design of the dual critic network can effectively alleviate the overestimation bias problem caused by the single Q value network and improve the stability and convergence speed of the algorithm. In order to improve the stability and convergence effect of model training, the self-standard network is used for soft update mechanism. Introducing the target actor network and target critic network and , and their parameters are and The parameter update of the target network adopts the soft update mechanism, and its update formula is:

[0127]

[0128] in, is the soft update coefficient, which is usually small (such as 0.005) to ensure that the target network parameters change slowly. The soft update mechanism makes the parameters of the target network gradually approach the parameters of the actual network, thereby maintaining the stability of the target network during training and avoiding training instability caused by rapid parameter changes. In the process of strategy generation, in order to improve the exploration ability of the strategy, Gaussian noise is added when generating actions. , forming an exploration strategy The specific calculation formula is:

[0129] ;

[0130] in, The mean is 0 and the standard deviation is The normal distribution of By adding noise, the policy generator is guided to explore more state space in the early stages of training to avoid falling into local optimal solutions too early. In order to gradually reduce exploration during training, an adaptive noise attenuation mechanism is designed, and the update formula for the noise standard deviation is:

[0131] ;

[0132] in, is the initial noise standard deviation, λ is the attenuation rate, is the number of training steps. As the training progresses, the noise standard deviation Gradually decreases, the strategy generator gradually shifts from the initial extensive exploration to the later precise strategy optimization. In order to effectively utilize the experience data, define the experience replay buffer , used to store the state-action-reward-next-state quadruple during training During each training, mini-batch data is randomly extracted from the experience pool for training to break the temporal correlation between samples and improve the efficiency and stability of training. The sampling mechanism of experience replay is random sampling to ensure the diversity of training samples, avoid the model from being overly dependent on certain specific state-action pairs, and improve the generalization ability of the model. In each training, according to the current state and actions generated by the policy , and the rewards obtained after performing actions in the environment and the next state , calculate the target Q value , and its calculation formula is:

[0133] ;

[0134] in, is the discount factor, usually taken as It is used to measure the trade-off between current rewards and future rewards. Indicates the current state Next action After that, the sum of the immediate reward and future long-term reward can be obtained. By minimizing the target Q value and the current Q value and The mean square error between them is used to update the parameters of the Critic network. The update formula of the Critic network is:

[0135] ;

[0136] in, is the loss function, that is, the mean square error between the target Q value and the current Q value, is the learning rate. For the Actor network, the parameters of the policy network are updated by maximizing the Q value. The update formula is:

[0137] ;

[0138] By increasing the Q value of the output of the policy network in the critic network, the quality of the policy is improved. In order to ensure the rationality of the learning rate at different training stages, an adaptive learning rate adjustment mechanism is designed. The learning rate update formula is:

[0139] ;

[0140] in, is the initial learning rate, is the learning rate decay coefficient, is the number of training steps. As the number of training steps increases, the learning rate Gradually decrease, so that the model can converge to the local optimal solution more stably in the later stage of training. The neural network structure, soft update mechanism, exploration strategy, noise standard deviation update formula, sampling mechanism, network update formula and learning rate update formula are combined to obtain the strategy generator.

[0141] In a specific embodiment, the process of executing step S6 may specifically include the following steps:

[0142] Initialize the system parameters to obtain the initialization parameter set, which includes the network structure, time window, and neural network weights. The network structure includes the number of nodes and path connection relationships, the time window includes the initial earliest shipping time and the latest arrival time, and the neural network weights include the initial weight values ​​of the actor network and the critic network.

[0143] The current state is input into the demand forecaster to obtain the forecast values ​​of freight demand and capacity supply for the next cycle. The current state includes the current cargo volume of each node, the current transportation volume of each path, and the current time window information.

[0144] updating the dynamically adjusted time window based on the predicted value to obtain an updated time window, wherein the updating process uses the time window adjustment formula in claim 5;

[0145] The current state st is input into the policy generator to obtain the scheduling decision at. The scheduling decision includes the transportation volume allocation of each path in each time period. The specific calculation process is at=μ(st)=π(st)+ε, where π(st) is the output of the Actor network and ε is the exploration noise.

[0146] Based on the scheduling decision, the transportation process simulation is performed to obtain the actual transportation cost and time. The simulation process includes updating the cargo volume of each node and the transportation volume of each route, and calculating the total transportation cost and total transportation time.

[0147] Calculate the reward value rt and obtain the reward function, which is rt=-F'(st,at), where F' is the new optimization objective function;

[0148] Observe the new state st+1 after the decision is executed and obtain the state transition sample (st, at, rt, st+1). The new state includes the updated cargo volume of each node, the transportation volume of each path, and the time window information;

[0149] The state transition samples are stored in the experience replay buffer D to obtain an updated experience pool, the size of which is the preset value M;

[0150] Randomly sample N mini-batch data from the updated experience pool to obtain a training sample set, where N is the preset batch size;

[0151] The network parameters of the policy generator are updated based on the training sample set to obtain the multimodal transport scheduling plan. The updating process includes: calculating the target Q value y=r+γ*min(Qφ1'(s',π(s')),Qφ2'(s',π(s'))), updating the critic network parameters φi←φi-η▽φiL(φi), and updating the actor network parameters θ←θ+η▽aQφ1(s,a)|a=π(s)▽θπ(s) every d steps, where η is the adaptive learning rate.

[0152] Specifically, the system parameters are initialized to obtain an initialization parameter set. The initialization parameter set includes the basic structure of the multimodal transport network, the time window setting, and the initial weight value of the neural network. The network structure describes the number of nodes and path connection relationships in the transportation system. For example, a multimodal transport system contains three main nodes: ports, railway hubs, and distribution centers. These three nodes are connected by roads, railways, and waterways to form a complete transportation network. The connection relationship of each path defines the accessibility and transportation mode between nodes. The initialization of the network structure can provide a basic framework for the entire scheduling optimization. The time window initialization includes the initial earliest shipping time of each node ( ) and the latest arrival time ( ). The neural network weight initialization includes the initial weight values ​​of the Actor network and the Critic network. The reasonable initialization of these weight values ​​can speed up the convergence of the network and improve the learning effect of the strategy generator. After completing the initialization of the system parameters, the current system state is input into the demand forecaster to predict the freight demand and capacity supply in the next cycle. The current state includes the current cargo volume of each transport node, the current transport volume of each transport path, and the current time window information. This information can reflect the real-time operating status of the system and provide effective input data for the demand forecaster. Based on historical data and current state information, the demand forecaster predicts the freight demand and capacity supply value of the next cycle through the operation of a multi-layer feedforward neural network. Based on the prediction results of the demand forecaster, the dynamically adjusted time window is updated to obtain a more accurate time window setting. The update process uses the time window adjustment formula in claim 5, that is, by adjusting the factor Earliest shipping time and latest arrival time Make adaptive adjustments. The adjustment amount of the time window Based on forecast demand and current capacity The adjustment formula is:

[0153] ;

[0154] in, is the adjustment coefficient, which indicates the sensitivity to changes in demand. is the maximum adjustment amount, which is used to limit the amplitude of a single adjustment. Based on the calculated adjustment amount, the updated time window is:

[0155] ;

[0156] ;

[0157] When the demand is expected to increase, the shipping time will be appropriately advanced and the arrival time will be delayed, thereby alleviating the transportation pressure of the node and improving the flexibility of scheduling. Input strategy generator to generate scheduling decisions under the current state The policy generator learns and generates policies based on a deep deterministic policy gradient algorithm. The specific calculation process is as follows:

[0158] ;

[0159] in, For the Actor network based on the current state The generated scheduling policy output, To explore noise, the mean is 0 and the standard deviation is Normal distribution This strategy generation method can guide the strategy generator to explore more state space in the early stage of training, avoiding premature convergence to the local optimal solution. Including the transportation volume allocation of each path in each time period. Based on the generated scheduling decision , performs transportation process simulation in the transportation network and calculates the actual transportation cost and time. The transportation process simulation includes updating the cargo volume of each node and the transportation volume of each path, and calculating the total transportation cost and total transportation time. Based on the scheduling decision, the transportation volume on each path is allocated to the corresponding node, and the transportation cost and time overhead caused by the scheduling plan are recorded. After the transportation simulation is completed, the reward value of the scheduling strategy is calculated. , the reward function is in the form of:

[0160] ;

[0161] in, is the new optimization objective function, indicating that Execute scheduling decisions The total cost and time caused by the negative sign indicates that the system hopes to minimize the total cost and time, that is, the larger the reward value, the better the system performance. The design of the reward function is to guide the policy generator to generate a scheduling policy that can maximize the overall benefit of the system. After obtaining the current reward value, observe the new state after executing the current scheduling decision. , record the new status information, including the updated cargo volume of each node, the transportation volume of each path, and the time window information. , scheduling decision , reward value and the new state Composition state transition sample , and store it in the experience replay buffer , get the updated experience pool. The size of the experience pool is the preset value , used to store the experience data accumulated during the training process. Randomly sample from the updated experience pool The randomly sampled training sample set includes several state transition samples. These samples can provide rich state-action pairs and their corresponding reward information, thus supporting the training process of the policy generator. Based on the training sample set, the network parameters of the policy generator are updated, including the parameter updates of the Critic network and the Actor network. Calculate the target Q value , and its calculation formula is:

[0162] ;

[0163] in, is the discount factor, which indicates the degree to which future rewards are discounted. and They are target Critic network, is the target Actor network. Target Q value Indicates the current state Next action After that, the sum of the immediate reward and future long-term reward that can be obtained. By minimizing the target Q value and the current Q value and The mean square error between them is used to update the parameters of the Critic network. The update formula is:

[0164] ;

[0165] in, is the loss function, that is, the mean square error between the target Q value and the current Q value, is the learning rate. For the Actor network, the parameters of the policy network are updated by maximizing the Q value. The update formula is:

[0166] ;

[0167] That is, by increasing the Q value of the output of the policy network in the Critic network, the quality of the policy is improved. The parameters of the Actor network are updated once every step, so that the Actor network can continuously optimize its strategy generation ability. During the training process, in order to maintain the rationality of the learning rate at different training stages, an adaptive learning rate adjustment mechanism is designed. The learning rate update formula is:

[0168] ;

[0169] in, is the initial learning rate, is the learning rate decay coefficient, is the number of training steps. As the number of training steps increases, the learning rate gradually decreases, allowing the model to more stably converge to the optimal solution in the later stages of training. The updated results of the policy generator are applied to the entire intermodal transport system to generate a specific scheduling plan. This scheduling plan, taking into account multiple factors such as node cargo volume, route transportation volume, and time windows, can effectively respond to dynamic changes in transportation demand, providing an intelligent and highly responsive scheduling optimization solution for the intermodal transport system.

[0170] The above describes the multimodal transport intelligent scheduling optimization method according to the embodiment of the present invention. The following describes the multimodal transport intelligent scheduling optimization device according to the embodiment of the present invention. Figure 2 In one embodiment of the present invention, an intelligent multimodal transport scheduling optimization device includes:

[0171] A modeling module is used to model multiple transportation nodes and transportation paths to obtain a network topology structure;

[0172] The partitioning module is used to divide the scheduling period into time periods based on the network topology to obtain a time window scheduling framework;

[0173] A training module is used to input historical freight data into a multi-layer feedforward neural network for training to obtain a demand forecaster;

[0174] An adjustment module is used to dynamically adjust the time window of the time window scheduling framework to obtain a dynamically adjusted time window;

[0175] A processing module is used to input the state space and action space into the improved double-delay deep deterministic policy gradient algorithm for training to obtain a policy generator;

[0176] The decision module is used to generate intermodal transport scheduling decisions based on the demand forecaster, the dynamically adjusted time window and the strategy generator to obtain the intermodal transport scheduling plan.

[0177] Through the collaborative efforts of the aforementioned components and comprehensive modeling of multiple transport nodes and routes, a complete network topology is derived. Based on this network topology, the scheduling cycle is divided into time periods, and a dynamic adjustment mechanism is introduced to enable the scheduling plan to better adapt to changes in actual transportation demand. A multi-layer feedforward neural network is trained on historical freight data, and an adaptive adjustment mechanism is introduced to improve the accuracy of future freight demand forecasts. The policy generator, trained using an improved double-delayed deep deterministic policy gradient algorithm, is capable of generating high-quality scheduling decisions in complex state spaces. By introducing a multimodal coordination factor and performance evaluation metrics, the system optimizes total transportation cost and time while also considering the trade-offs between different modes of transport. Mechanisms such as dynamic time window adjustment and adaptive learning rate enable the scheduling system to better adapt to environmental changes and uncertainties. Based on the demand forecaster, dynamically adjusted time windows, and policy generator, multimodal transport scheduling plans can be generated and updated in real time, improving the system's responsiveness.

[0178] The present invention also provides a computer device, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the multimodal transport intelligent scheduling optimization method in the above-mentioned embodiments.

[0179] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the multimodal transport intelligent scheduling optimization method.

[0180] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0181] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0182] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal transport intelligent scheduling optimization method, characterized in that: The method comprises: Model multiple transportation nodes and transportation paths to obtain the network topology; Dividing the scheduling period into time periods based on the network topology to obtain a time window scheduling framework; The historical freight data is fed into a multi-layer feedforward neural network for training to obtain a demand forecaster; Dynamically adjusting the time window of the time window scheduling framework to obtain a dynamically adjusted time window; The state space and action space are input into the improved double-delayed deep deterministic policy gradient algorithm for training to obtain the policy generator; An intermodal transport scheduling decision is generated based on the demand forecaster, the dynamically adjusted time window, and the strategy generator to obtain an intermodal transport scheduling solution.

2. The multimodal transport intelligent scheduling optimization method according to claim 1 is characterized in that: The network topology structure is obtained by modeling multiple transport nodes and transport paths, including: Defining attributes for each transport node to obtain a node attribute set, wherein the node attribute set includes location coordinates, node type, and processing capacity indicators; Defining attributes for each transport route to obtain a set of route attributes, wherein the set of route attributes includes a starting point, an end point, a transport mode, a transport time, a transport cost, and a transport capacity upper limit; Based on actual geographic information and transportation network data, the connection relationship between each node is determined to obtain a complete multimodal transport network topology. Historical freight data is collected and organized to obtain the cargo inflow and outflow of each node and the actual transportation volume of each transportation route. An objective function is constructed based on the total transportation cost and the total transportation time to obtain a network optimization target, wherein the objective function is F=min(α*total transportation cost+β*total transportation time); Initialize the weight coefficients α and β in the objective function to obtain the initial weight values, and construct network constraints based on the node attribute set and the path attribute set to obtain the initial network constraints; The network optimization target, the initial weight value and the initial network constraint are combined to obtain a network topology structure.

3. A multimodal transport intelligent scheduling optimization device, characterized in that: For executing the multimodal transport intelligent scheduling optimization method according to claim 1 or 2, the multimodal transport intelligent scheduling optimization device comprises: A modeling module is used to model multiple transportation nodes and transportation paths to obtain a network topology structure; A partitioning module, configured to partition the scheduling period into time periods based on the network topology to obtain a time window scheduling framework; A training module is used to input historical freight data into a multi-layer feedforward neural network for training to obtain a demand forecaster; An adjustment module, configured to dynamically adjust the time window of the time window scheduling framework to obtain a dynamically adjusted time window; A processing module is used to input the state space and action space into the improved double-delay deep deterministic policy gradient algorithm for training to obtain a policy generator; A decision module is used to generate a multimodal transport scheduling decision based on the demand forecaster, the dynamically adjusted time window and the strategy generator to obtain a multimodal transport scheduling solution.

4. A computer device, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and is characterized in that when the processor executes the computer program, it implements the multimodal transport intelligent scheduling optimization method according to claim 1 or 2.

5. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, enables the processor to execute the multimodal transport intelligent scheduling optimization method according to claim 1 or 2.

Citation Information

Patent Citations

  • Multimodal transport path optimization method considering uncertain conditions

    CN111626477A

  • Railway guarantee transportation task planning method based on NT-A* algorithm

    CN117910741A