Scheduling method, equipment and storage medium for logistics platform

The logistics order data is converted into a state vector through a fully connected feedforward neural network, and the target value and expected value are calculated by combining the main network and the target network, which solves the problem of low accuracy of logistics order scheduling and realizes efficient logistics scheduling in complex environments.

CN120297688BActive Publication Date: 2025-09-30GUANGZHOU PINGYUN CRAFTSMAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510749730.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-30
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing logistics order scheduling system is difficult to accurately schedule based on complex and changeable internal and external conditions, resulting in low order scheduling accuracy.

Method used

A fully connected feedforward neural network is used to convert logistics order data into a state vector of preset dimensions. The target value is calculated by the main network and the expected value is calculated by the target network. The target decision is determined by combining the weighted fusion results to realize logistics scheduling.

Benefits of technology

It improves the accuracy and efficiency of logistics scheduling, and can make reasonable and efficient scheduling decisions in complex and changing environments, reduce costs, and ensure that goods are delivered on time and in good quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297688B_ABST
    Figure CN120297688B_ABST
Patent Text Reader

Abstract

The present application discloses a scheduling method, device and storage medium for a logistics platform. The present application relates to the field of data processing technology. The scheduling method for a logistics platform includes: forming a state vector of preset dimensions from logistics order data in a standard form; inputting the state vector into a main network, and calculating the target value of each decision to be selected. The main network is a fully connected feedforward neural network, comprising an input layer, at least two hidden layers and an output layer. The input layer dimension is consistent with the preset dimension, the hidden layer adopts a target activation function, and the output layer dimension is consistent with the number of decisions to be selected; inputting the state vector and the decision to be selected into a target network, and calculating the expected value of each decision to be selected. The target network has the same structure as the main network and the parameter update lags; determining the target decision based on the weighted fusion result of the target value and the expected value, so as to issue the logistics scheduling task based on the target decision. The present application can accurately perform logistics scheduling in combination with multi-source business data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a scheduling method, device and storage medium for a logistics platform. Background Art

[0002] Currently, logistics order scheduling primarily relies on data generated by logistics operators during actual operations to plan vehicle allocation, route arrangements, and other tasks. This data is therefore limited to the logistics operations themselves, failing to capture the diverse factors influencing order scheduling. This makes it difficult to schedule orders based on complex and changing internal and external conditions, resulting in low order scheduling accuracy. Summary of the Invention

[0003] The main purpose of this application is to provide a scheduling method, equipment and storage medium for a logistics platform, aiming to solve the technical problem that it is difficult to schedule orders according to complex and changeable internal and external conditions, which leads to low order scheduling accuracy.

[0004] To achieve the above objectives, the present application provides a scheduling method for a logistics platform, which includes:

[0005] The logistics order data in standard form is organized into a state vector of preset dimensions;

[0006] Input the state vector into the main network and calculate the target value of each candidate decision. The main network is a fully connected feedforward neural network, including an input layer, at least two hidden layers, and an output layer. The dimension of the input layer is consistent with the preset dimension. The hidden layer uses the target activation function. The dimension of the output layer is consistent with the number of candidate decisions.

[0007] Inputting the state vector and the candidate decision into a target network, and calculating the expected value of each candidate decision, wherein the target network has the same structure as the main network and has a parameter update lag;

[0008] A target decision is determined according to a weighted fusion result of the target value and the expected value, so as to issue a logistics scheduling task based on the target decision.

[0009] In one embodiment, the step of forming a state vector of a preset dimension from the standardized form-based logistics order data includes:

[0010] If the logistics order data meets the order splitting conditions, determine multiple sub-orders corresponding to the logistics order data;

[0011] A state vector of the preset dimension is generated based on each of the sub-orders.

[0012] In one embodiment, the step of inputting the state vector and the candidate decisions into the target network and calculating the expected value of each candidate decision includes:

[0013] Based on the target values ​​of the candidate decisions, sort and select at least one target candidate decision;

[0014] Based on the state vector and the target network, the expected value corresponding to each target candidate decision is calculated.

[0015] In one embodiment, before the step of inputting the state vector into the main network and calculating the target value of each candidate decision, the following steps are included:

[0016] Establish training data sets based on internal and external data;

[0017] Construct an initial network based on one input layer, at least two hidden layers, and one output layer;

[0018] Determining hyperparameters of the initial network based on the training data set to construct a first network and a second network;

[0019] The first network and the second network are trained according to the training data set to generate the main network and the target network.

[0020] In one embodiment, the step of establishing a training data set based on the endogenous data and the exogenous data includes:

[0021] Associating the endogenous data and the exogenous data into a data relationship set based on a data relationship graph;

[0022] Determine the quantitative indicators of factors affecting logistics scheduling;

[0023] The quantitative value of each data in the data relationship set is determined based on the quantitative index, and then the training data set is established.

[0024] In one embodiment, the initial network is a fully connected layer structure, and the step of determining the hyperparameters of the initial network based on the training data set to construct the first network and the second network includes:

[0025] Determine the initial values ​​of the hyperparameters in the experience replay pool, sampling batch, discount factor, soft update target network decision parameter, learning rate, action space dimension, and state space dimension;

[0026] The initial network is initialized based on the initial value to construct the first network and the second network.

[0027] In one embodiment, the step of training the first network and the second network according to the training data set to generate the main network and the target network includes:

[0028] Storing the training data set in an experience replay pool;

[0029] Obtaining sample data from the experience replay pool based on a sampling batch;

[0030] Based on each batch of the sample data input to the first network, the action target value is calculated; based on each batch of the sample data input to the second network, the action expected value is calculated;

[0031] Updating parameters of the first network according to the action target value and the action expected value;

[0032] Updating parameters of the second network based on the soft update parameters and the first network parameters;

[0033] If the training converges, the main network is generated based on the first network, and the target network is generated based on the second network.

[0034] In one embodiment, the steps of calculating the action target value based on each batch of sample data input into the first network, and calculating the action expected value based on each batch of sample data input into the second network, include:

[0035] Calculating the sample data through the first network to obtain an action for the next state and an action target value for the action;

[0036] Calculating the action value of the action based on the sample data by the second network;

[0037] The action expected value corresponding to the action value is calculated by combining the discount factor and the reward function.

[0038] In addition, to achieve the above-mentioned purpose, the present application also provides a scheduling device for a logistics platform, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the scheduling method for the logistics platform as described above.

[0039] In addition, to achieve the above-mentioned purpose, the present application also provides a storage medium, which is a computer-readable storage medium, and the computer-readable storage medium stores a program for implementing the scheduling method of the logistics platform. The program for implementing the scheduling method of the logistics platform is executed by the processor to implement the steps of the scheduling method of the logistics platform as described above.

[0040] The present application provides a scheduling method for a logistics platform. The present application first forms a state vector of preset dimensions from standardized form logistics order data; inputs the state vector into a main network to calculate the target value of each candidate decision; the main network is a fully connected feedforward neural network, comprising an input layer, at least two hidden layers, and an output layer, wherein the input layer dimension is consistent with the preset dimension, the hidden layer uses a target activation function, and the output layer dimension is consistent with the number of candidate decisions; inputs the state vector and the candidate decisions into a target network to calculate the expected value of each candidate decision; the target network has the same structure as the main network and has a parameter update lag; the target decision is determined based on the weighted fusion result of the target value and the expected value, and the logistics scheduling task is issued based on the target decision. This solves the technical problem of difficulty in scheduling orders according to complex and changeable internal and external conditions, which leads to low order scheduling accuracy, and achieves the technical effect of accurately executing logistics scheduling by combining multi-source business data. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0042] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0043] Figure 1 A flow chart of the first embodiment of the scheduling method for the logistics platform of this application;

[0044] Figure 2 A flowchart of the fourth embodiment of the scheduling method for the logistics platform of this application is provided;

[0045] Figure 3 A flow chart of the seventh embodiment of the scheduling method for the logistics platform of this application;

[0046] Figure 4 This is a schematic diagram of the hardware structure involved in the scheduling equipment embodiment of the logistics platform of this application.

[0047] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0048] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0049] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0050] The main solution of this application is: standard form logistics order data is formed into a state vector of preset dimensions; the state vector is input into the main network to calculate the target value of each decision to be selected, the main network is a fully connected feedforward neural network, comprising an input layer, at least two hidden layers and an output layer, the input layer dimension is consistent with the preset dimension, the hidden layer adopts a target activation function, and the output layer dimension is consistent with the number of the decisions to be selected; the state vector and the decisions to be selected are input into the target network to calculate the expected value of each decision to be selected, the target network has the same structure as the main network and the parameter update is lagged; the target decision is determined according to the weighted fusion result of the target value and the expected value, so as to issue the logistics scheduling task based on the target decision.

[0051] Currently, logistics order scheduling primarily relies on data generated by logistics operators during actual operations to plan vehicle allocation, route arrangements, and other tasks. Consequently, this data is limited to the logistics operations themselves, failing to capture the diverse factors influencing order scheduling. This makes it difficult to schedule orders based on complex and changing internal and external conditions, resulting in low order scheduling accuracy.

[0052] This application uses multi-source data fusion technology to model and integrate platform, carrier, driver, customer and external data to form platform data resources for intelligent agents to learn and use, making decisions more accurate and more in line with actual business scenarios.

[0053] This application addresses the poor robustness of expert algorithms and traditional Q-learning techniques. By designing separate action selection and value assessment networks, the logistics decision-making process is decoupled into two independent systems, allowing them to independently interfere with each other. Furthermore, the hidden layers of the deep network can factor in decision factors not explicitly enumerated, increasing the robustness of the agent's decisions and speeding up training convergence. This allows for more accurate decisions using less training data while consuming less computing power.

[0054] It should be noted that the execution entity of this embodiment can be a logistics platform, or a computing service device with data processing, network communication, and program execution capabilities, such as a tablet computer, personal computer, or mobile phone, or a logistics platform scheduling device capable of performing the above functions. This embodiment does not specifically limit this. The following uses the logistics platform as an example to illustrate this embodiment and the following embodiments.

[0055] Based on this, the present application embodiment 1 proposes a scheduling method for a logistics platform, please refer to Figure 1The scheduling method of the logistics platform includes steps S10 to S40:

[0056] Step S10: The logistics order data in the standard form is organized into a state vector of a preset dimension.

[0057] In this embodiment, standardized logistics order data refers to logistics order-related information recorded in a unified format, covering details such as the total number of ordered goods, cargo type, destination, and order urgency. The pre-defined state vector is a fixed-dimensional vector determined based on factors influencing logistics scheduling operations. Here, the pre-defined dimension is 30, which is used to comprehensively describe the state of the logistics scheduling scenario.

[0058] As an optional implementation, standardized logistics order data is first extracted from the logistics system database. This data is then combined with internal data sources (such as supplier inventory status, warehouse capacity availability, logistics vehicle information, and order management statistics) and external data sources (such as real-time ship positioning information, traffic navigation data, and weather conditions). This data is then integrated and quantified into a 30-dimensional state vector according to established feature engineering rules. For example, weather severity can be quantified as a value between 0 and 10, and road congestion can be converted into a congestion index, ultimately forming the state vector.

[0059] Step S20: input the state vector into the main network and calculate the target value of each candidate decision. The main network is a fully connected feedforward neural network, which includes an input layer, at least two hidden layers and an output layer. The dimension of the input layer is consistent with the preset dimension, the hidden layer adopts the target activation function, and the dimension of the output layer is consistent with the number of the candidate decisions.

[0060] In this embodiment, the main network is a neural network model used to calculate the target value of the candidate decision. A fully connected feedforward neural network means that the nodes in each layer of the network are connected to all nodes in the previous layer, and data propagates unidirectionally from the input layer to the output layer. The target activation function here refers to the ELU activation function, which is used to enhance the network's nonlinear expression capabilities. The candidate decision is a series of predefined executable actions for the logistics scheduling business, such as splitting an order into multiple waybills or assigning a waybill to a specific driver. A fully connected feedforward neural network (FCNN) is a neural network in which all nodes in all layers are fully connected. It consists of an input layer, a hidden layer, and an output layer. The nodes in each layer are connected to all nodes in the previous layer, but there are no connections between nodes within the same layer. Data propagates unidirectionally from the input layer to the output layer in the network, without feedback or loop connections.

[0061] As an optional implementation, the constructed 30-dimensional state vector is input into the input layer of the main network. After the input layer receives the vector, the data passes through two hidden layers in sequence (the first hidden layer has 128 nodes and the second hidden layer has 64 nodes). In the hidden layer, the data is nonlinearly transformed through the ELU activation function to enhance the network's ability to learn complex relationships. Finally, the data reaches the output layer. The output layer calculates the target values ​​corresponding to the number of decisions to be made (10 types) based on the input data. These target values ​​represent the expected value of taking each action in the current state.

[0062] Step S30: input the state vector and the candidate decision into a target network, and calculate the expected value of each candidate decision. The target network has the same structure as the main network and its parameters are updated with a lag.

[0063] In this embodiment, the target network is a neural network with the same structure as the main network, but its parameter updates are slower than those of the main network. This is used to provide a relatively stable evaluation reference and avoid possible overestimation biases caused by the main network.

[0064] As an optional implementation, the same 30-dimensional state vector and the candidate decisions calculated by the main network are input into the target network. The target network processes these inputs with the same structure as the main network (one input layer, two hidden layers, one output layer, the same number of nodes and activation function settings) and calculates the expected value of each candidate decision in the current state. Due to the lag in its parameter update, it can provide more stable evaluation results.

[0065] In this embodiment, by constructing two independent loop networks - the main network and the target network, the slow convergence, radical decision-making, and local optimality caused by the single network of traditional DQN technology are solved, and the action selection a and Q value evaluation are decoupled.

[0066] Step S40: determining a target decision according to a weighted fusion result of the target value and the expected value, so as to issue a logistics scheduling task based on the target decision.

[0067] In this embodiment, weighted fusion combines the target value calculated by the primary network and the expected value calculated by the target network using a weight coefficient to balance their contributions. Target decisioning is the final logistics scheduling action determined after comprehensively considering the target and expected values. Logistics scheduling tasks are then distributed to the driver software and the carrier system.

[0068] As an optional implementation, a weighted fusion formula (Qfinal = α × Qmain + (1 − α) × Qtarget) is used, where α is a weight coefficient ranging from 0.6 to 0.9 (for example, 0.7), Qmain is the target value output by the main network, and Qtarget is the expected value output by the target network. This formula calculates the final Q value of each candidate decision, selects the candidate with the highest Q value as the target decision, and then generates detailed logistics scheduling task instructions based on the target decision, such as vehicle scheduling instructions, including the vehicle's departure location, pickup location, and estimated arrival time. These instructions are then sent to the logistics execution system for execution.

[0069] For example, consider a logistics order for perishable goods, with a high degree of urgency and a remote destination. Order information, warehouse inventory (endogenous data), current road congestion, and weather information (exogenous data) are integrated to construct a 30-dimensional state vector. This state vector is then fed into the main network, which calculates target values ​​for pending decisions, such as "select a high-speed transport route" and "add refrigerated vehicles." The target network calculates expected values ​​based on the same state vector and the pending decisions. Finally, through weighted fusion, "select a high-speed transport route and add refrigerated vehicles" is determined as the target decision. The corresponding dispatch task is issued, arranging a suitable refrigerated vehicle to depart from the warehouse and travel along the highway to the destination. This embodiment, by leveraging the real-time performance of the main network and the stability of the target network, effectively improves the accuracy and reliability of logistics scheduling decisions, avoiding the potential overestimation bias of single-network decisions. In complex and changing logistics environments, more reasonable and efficient scheduling decisions can be made based on multi-source data, improving logistics operational efficiency, reducing costs, and ensuring the timely and safe delivery of goods to their destinations.

[0070] This application first forms the logistics order data in a standard form into a state vector of preset dimensions; inputs the state vector into the main network, calculates the target value of each candidate decision, the main network is a fully connected feedforward neural network, including an input layer, at least two hidden layers and an output layer, the input layer dimension is consistent with the preset dimension, the hidden layer uses a target activation function, and the output layer dimension is consistent with the number of the candidate decisions; inputs the state vector and the candidate decisions into the target network, calculates the expected value of each candidate decision, the target network has the same structure as the main network and the parameter update lags; determines the target decision based on the weighted fusion result of the target value and the expected value, and issues the logistics scheduling task based on the target decision. This solves the technical problem of difficulty in scheduling orders according to complex and changeable internal and external conditions, which leads to low order scheduling accuracy, and achieves the technical effect of accurately executing logistics scheduling by combining multi-source business data.

[0071] Based on any of the above embodiments, in the second embodiment of the present application, step S10 includes:

[0072] Step S11: If the logistics order data meets the order splitting conditions, determine multiple sub-orders corresponding to the logistics order data.

[0073] In this embodiment, the order splitting condition is a standard set according to logistics business rules and actual conditions for determining whether an order needs to be split into multiple sub-orders.

[0074] As an optional implementation, the logistics order data is determined to meet the splitting conditions when the total number of ordered goods in the order exceeds a preset threshold (e.g., 1,000 pieces), the volume or weight of the goods exceeds the carrying capacity of a single transport vehicle (e.g., exceeds the maximum load of a truck), or the order includes at least two types of goods requiring different transportation modes (e.g., goods requiring both sea and land transport). The order is then split into multiple sub-orders based on factors such as cargo type and destination. For example, an order containing goods destined for different destinations can be split into multiple sub-orders based on destination.

[0075] Step S12: generating a state vector of the preset dimension based on each of the sub-orders.

[0076] In this embodiment, the state vector of the preset dimension is a fixed-dimensional vector used to describe the state of the logistics scheduling scenario associated with each sub-order. As an optional implementation, for each sub-order, internal source data related to the sub-order (such as the inventory status of the corresponding warehouse and information related to the sub-order in the order management system) and external source data (such as weather data at the sub-order destination and real-time road conditions along the transportation route) are collected. This data is then integrated and quantized into a 30-dimensional state vector using the same feature engineering principles as for the overall order. For example, for a sub-order, the severity of the weather at its destination is quantified as a value between 0 and 10, and the storage location information of the sub-order in the warehouse is converted into a corresponding code, ultimately forming a 30-dimensional state vector for the sub-order.

[0077] For example, consider a large logistics order containing multiple types of goods, with a total weight exceeding the capacity of a single truck. Some of the goods require sea transport, while others require land transport. Based on the order splitting criteria, the order is split into two sub-orders: one for sea transport and the other for land transport. A 30-dimensional state vector is generated for each of these sub-orders. For the sea transport sub-order, data such as cargo information, warehouse inventory status, vessel transportation information, and weather and tidal conditions at the destination port are collected and quantized to generate a state vector. For the land transport sub-order, data such as cargo information, warehouse information, transport vehicle information, and road and weather conditions along the transport route are collected and quantized to generate a state vector.

[0078] Through reasonable order splitting operations and sub-order state vector generation, this embodiment can handle complex logistics orders more finely, make more accurate scheduling decisions based on the characteristics of different sub-orders, improve the utilization efficiency of logistics resources, ensure that all types of goods can be properly transported, and further enhance the overall efficiency of logistics scheduling.

[0079] Based on any of the above embodiments, in the third embodiment of the present application, the step of inputting the state vector and the candidate decision into the target network and calculating the expected value of each candidate decision includes:

[0080] Step S31: sort and select at least one target decision to be selected based on the target values ​​of the decision to be selected.

[0081] In this embodiment, the target value of each candidate decision is a numerical value calculated by the primary network for each candidate decision, representing the expected value of that decision in its current state. The target candidate decisions are selected from all candidate decisions, sorted by target value, for further evaluation.

[0082] As an optional implementation, first obtain the target values ​​of each candidate decision calculated by the primary network, and then sort all the candidate decisions from highest to lowest target value. Based on actual needs, at least one candidate decision with the highest ranking is selected as the target candidate decision. For example, the top three candidate decisions with the highest target values ​​are selected as the target candidate decisions.

[0083] Step S32: Calculate the expected value corresponding to each target candidate decision based on the state vector and the target network.

[0084] In this embodiment, the state vector is a vector containing various information related to logistics orders and used to describe the current logistics scheduling scenario. The target network is a neural network with the same structure as the main network but with delayed parameter updates, used to provide relatively stable evaluation. The expected value is the expected value calculated by the target network for each target candidate decision in its current state.

[0085] As an optional implementation, the determined target candidate decisions and state vectors are input into the target network. Based on its internal network structure and parameters, the target network calculates each target candidate decision under the current state vector and outputs the corresponding expected value. For example, the target network processes the state vector, combining its own weights and biases to calculate the expected value of each target candidate decision. These expected values ​​reflect the expected effect of each target candidate decision in the current logistics scheduling scenario.

[0086] For example, suppose the main network calculates the target values ​​for 10 candidate decisions. After sorting, it selects the top three with the highest target values ​​as the target decisions: "Select the fastest transport route," "Add additional transport vehicles," and "Prioritize this order." These three target decisions and the current 30-dimensional state vector are then input into the target network. Based on its own structure and parameters, the target network calculates each target decision and finds that the expected value of "Select the fastest transport route" is 80, indicating the likely effect of selecting this route in the current scenario. The expected value of "Add additional transport vehicles" is 70, and the expected value of "Prioritize this order" is 75.

[0087] This embodiment first filters the target value calculated by the main network and then uses the target network to calculate the expected value. This can reduce the amount of calculation while more specifically evaluating key candidate decisions, thereby improving the efficiency and accuracy of decision-making, helping to quickly make more reasonable decisions in complex logistics scheduling, optimize the allocation and utilization of logistics resources, and improve the overall logistics service quality.

[0088] Based on any of the above embodiments, in the fourth embodiment of the present application, refer to Figure 2 , before the step of inputting the state vector into the main network and calculating the target value of each candidate decision, including:

[0089] Step A10: Establish a training data set based on the internal source data and the external source data.

[0090] In this example, endogenous data refers to data related to logistics operations from within the logistics platform, including supplier information, warehouse management data, and logistics vehicle status. Exogenous data refers to data related to the logistics environment from outside the logistics platform, such as ship positioning information and weather data. The training dataset is a collection of samples used to train the neural network model.

[0091] As an optional implementation, historical logistics order data, including various order attribute information, is first extracted from the logistics platform's database. Simultaneously, corresponding endogenous data is collected, such as changes in supplier inventory status over time and warehouse operating load rates; and exogenous data, such as weather conditions and road congestion during order transportation, is collected. Data relationship modeling is then performed on this multi-source data to clarify the connections and influences between the data. Next, metrics are quantified, converting various non-numeric data into numerical indicators. Data cleaning is then performed to remove outliers and noise data, ultimately forming a training dataset containing a large number of samples, each of which contains information such as state, action, reward, and next state.

[0092] As an optional implementation, internal source data includes data from supplier management systems, warehouse management systems, logistics operations systems, customer management systems, and order management systems. External source data includes ship positioning data, route location data, map service data, overseas logistics data, and weather data. These internal and external source data are integrated into the logistics platform.

[0093] Step A20: constructing an initial network based on one input layer, at least two hidden layers, and one output layer.

[0094] In this embodiment, the initial network is the basic architecture of the neural network, the input layer is used to receive external data, the hidden layer is used to extract and convert features of the data, and the output layer is used to output prediction results.

[0095] As an optional implementation, a fully connected feedforward neural network consisting of one input layer, two hidden layers, and one output layer is constructed according to the preset network structure design. The dimension of the input layer is set to be consistent with the preset dimension, that is, 30 dimensions, corresponding to the dimension of the state vector. The first hidden layer is set to 128 nodes, the second hidden layer is set to 64 nodes, and the dimension of the output layer is set to be consistent with the number of candidate decisions, that is, 10 dimensions, corresponding to 10 candidate decisions. At the same time, a suitable activation function, such as the ELU activation function, is selected in the hidden layer to enhance the nonlinear expression ability of the network.

[0096] Step A30: Determine the hyperparameters of the initial network based on the training data set to construct a first network and a second network.

[0097] In this embodiment, hyperparameters are parameters that control the training process and performance of the neural network, such as learning rate, discount factor, etc. The first network and the second network are two neural networks with the same structure but possibly different parameters, corresponding to the subsequent main network and target network, respectively.

[0098] As an optional implementation, the training dataset is used to determine the hyperparameters of the initial network through experimentation and tuning. For example, the learning rate can be set to 5e-4, the discount factor to 0.99, the capacity of the experience replay pool to int(1e5), and the batch size to 64. Based on the determined hyperparameters, the structure of the initial network is replicated to create two independent networks, named the first network and the second network.

[0099] Step A40: Train the first network and the second network according to the training data set to generate the main network and the target network.

[0100] In this example, training refers to the process of continuously adjusting network parameters to enable the network to learn patterns and regularities from the training data set. The primary network is the neural network used for actual decision-making, while the target network is the neural network used to provide a stable evaluation reference.

[0101] As an optional implementation, first build an experience replay pool, store the samples in the training dataset in a specific format (such as Experience (state, action, reward, next_state, done, order_id)) into the experience replay pool, called transition, and store it in a fixed-capacity experience pool (Replay Buffer). <done>To mark completion, a Boolean value, distinguished by 0 / 1, is used to mark the end state of a sample, avoiding incorrect assessments and expectations of completed states. In this method, it is used to indicate whether all order goods have been dispatched. Next, data is randomly sampled in batches (typically 32 / 64 groups) from the experience replay pool to break temporal correlation and improve training stability. This data is then used to train the first and second networks. During training, the first network calculates the action choices for the next state, and the second network evaluates the value of these actions. The target Q-value is calculated by combining the reward and discount factor. The first network then predicts the Q-value for the current state. The mean squared error between the two is calculated, and the parameters of the first network are updated through backpropagation. Simultaneously, the parameters of the second network are periodically updated using a soft update mechanism, slowly aligning them with those of the first network. This training process continues until convergence, at which point the first and second networks become the primary and target networks, respectively, and can be used for practical logistics scheduling decisions.

[0102] In this embodiment, the loss function is calculated by the temporal difference error (TD error), and the target Q value is generated by the target network (TargetNetwork) using the formula:

[0103] .

[0104] Among them, r is the immediate reward, γ is the discount factor, Q target is the output value of the target network. The parameters of the target network are periodically synchronized from the training network (soft / hard synchronization) to reduce training fluctuations. Loss function (mean square error MSE):

[0105] .

[0106] Among them, Q θ is the main network, used to select actions; Q θ- is the target network, used to estimate the target Q value.

[0107] In this embodiment, action selection and value evaluation are decoupled.

[0108] Action selection: Use current network Q target (…; θ) selects the optimal action a*.

[0109] Value assessment: Using the target network Q target (……;θ - ) calculates the value of the action.

[0110] Significantly reduce the Q-value overestimation bias and improve strategy stability.

[0111] For example, in a practical application scenario on a logistics platform, a large amount of historical logistics order data, along with corresponding endogenous and exogenous data, was first collected and processed to create a training dataset. An initial network was then constructed, with a 30-dimensional input layer, two hidden layers of 128 nodes and 64 nodes, respectively, and a 10-dimensional output layer. Hyperparameters, such as a learning rate of 5e-4 and a discount factor of 0.99, were experimentally determined, and the first and second networks were constructed based on these parameters. These two networks were then trained using the training dataset. During training, data was continuously sampled from the experience replay pool for learning and parameter updates. After multiple iterations of training, the networks gradually converged, ultimately generating the master and target networks. When new logistics orders require scheduling, these two networks can be used to make decisions, improving the accuracy and efficiency of logistics scheduling.

[0112] Through a systematic training process, this embodiment enables the generated main network and target network to accurately learn the complex patterns and rules in logistics scheduling. The dual-network structure effectively avoids the overestimation bias that may occur in a single network, improves the stability and reliability of decision-making, provides strong support for the efficient operation of the logistics platform, reduces logistics costs, and improves customer satisfaction.

[0113] Based on any of the above embodiments, in the fifth embodiment of the present application, step A10 includes:

[0114] Step A11: associate the endogenous data and the exogenous data into a data relationship set based on a data relationship graph.

[0115] In this embodiment, the data relationship diagram is a graphical representation of the relationship between internal source data and external source data, used to guide data integration and association. The data relationship set is a data set composed of internal source data and external source data according to their internal relationship.

[0116] As an optional implementation, the logical relationships between internal data sources (such as supplier information, warehouse data, and logistics vehicle status) and external data sources (such as ship positioning and weather data) are first analyzed to determine association rules. For example, order data and warehouse inventory data can be linked through the goods information in the order, and transport vehicle data and road condition data can be linked through the transport route. Then, based on these association rules, the internal and external data sources are linked to form a complete data relationship set, ensuring data relevance and consistency.

[0117] This embodiment provides an example of a data relationship diagram. A logistics order includes at least the order number, pickup date, destination, delivery date, pickup address, and product information. The logistics order can then be split into several logistics waybills. Each logistics waybill includes the waybill number, waybill rating, driver information, product information, transportation method, and destination. The logistics waybill corresponds to the delivery region, including administrative level, road information, and weather information. The cargo carried on the logistics waybill includes the bulk-to-weight ratio, product name, actual weight, insurance information, and volume. The warehouse where the cargo is picked up / delivered includes the address, location, loading / unloading time, and platform. The container corresponding to the cargo includes the container number, container type, and company. Container consolidation methods include aircraft, ship, train, and train. Aircraft include the flight number and aircraft model. Ships include the IMO and MMSI. Trains include the train number and carriage model. Trains include the license plate number, operating license, vehicle model, and carrying capacity. The logistics waybill is then issued to the carrier, who dispatches the aircraft, ship, train, and train. The logistics waybill also assigns a driver, including the driver's name, vehicle type, professional qualifications, and driver's license number. The driver drives the truck.

[0118] Step A12: Determine the quantitative indicators of the factors affecting logistics scheduling.

[0119] In this embodiment, the factors affecting logistics scheduling refer to various factors that affect logistics scheduling decisions, such as the severity of weather conditions, road congestion, etc. Quantitative indicators are specific metrics that convert these factors into numerical values.

[0120] As an optional implementation, we conduct an in-depth analysis of logistics scheduling operations to identify 30 key factors influencing logistics scheduling, such as weather severity, road congestion, driver service star rating, order urgency, and platform order backlog. We develop quantitative rules for each influencing factor. For example, we classify weather severity into different levels, each corresponding to a specific numerical range; and we convert road congestion into a congestion index, which is calculated using traffic data.

[0121] For example, there are 30 factors: in-transit weather severity, road congestion, driver service star rating, order urgency, platform order backlog, on-time fulfillment rate, order gross profit margin, total number of ordered items, cargo type, mode of transport, whether multimodal transport is required, empty truck fuel consumption, fully loaded truck fuel consumption, truck length, driver acceptance status, available space in the truck compartment, vehicle load, whether the vehicle is returning, number of multi-stage transport segments, heavy / bulky cargo classification, cargo weight, cargo volume, full truckload / full container transport type, distance to the destination warehouse, availability of warehouse docks, warehouse operating load rate, carrier accident rate, carrier in-transit orders, carrier performance rating, and carrier capacity load rate. Table 1 below provides the numerical ranges for the quantitative indicators of each factor.

[0122] Table 1 Quantitative indicators of various factors

[0123]

[0124] Step A13: determining the quantitative value of each data in the data relationship set based on the quantitative index, and then establishing the training data set.

[0125] In this embodiment, the quantized value is a specific numerical value obtained by quantizing the data according to a quantitative index. A training dataset is a collection of quantized samples used to train a neural network model. As an optional implementation, each data point in the data relationship set is quantized according to a specific quantitative index. For example, weather data is converted to a corresponding numerical value based on a quantization rule for weather severity; order urgency is assigned a corresponding quantitative value based on a preset scoring criteria. The quantized data is organized into samples according to a specific format, with each sample containing information such as state, action, reward, and next state, ultimately forming the training dataset.

[0126] In this embodiment, based on business experience and actual conditions, and to facilitate intelligent decision-making, the following 10-element action set is abstracted for actual logistics scheduling operations: splitting an order into multiple waybills, assigning a waybill to a certain driver, assigning a waybill to a certain carrier, dispatching a return truck to a certain location to pick up goods, dispatching a return truck to a certain warehouse to deliver goods, dispatching a first-leg truck to a certain warehouse to deliver goods, loading a certain order of goods first when loading a vehicle, connecting a vehicle's goods to another vehicle, giving priority to the delivery of goods on a certain waybill, and transferring a carrier's waybill to another carrier.

[0127] This embodiment converts complex logistics scheduling influencing factors into computable numerical values ​​through a systematic data association and quantification process, providing a high-quality data foundation for the training of the neural network model, enabling the model to better learn the patterns and rules in logistics scheduling, improving the accuracy and efficiency of logistics scheduling decisions, optimizing the allocation and utilization of logistics resources, and reducing logistics costs.

[0128] Based on any of the above embodiments, in Embodiment 6 of the present application, step A30 includes:

[0129] Step A31, determine the initial values ​​of each hyperparameter in the experience replay pool, sampling batch, discount factor, soft update target network judgment parameter, learning rate, action space dimension, and state space dimension.

[0130] In this embodiment, the experience replay pool is a buffer for storing experience samples generated by the interaction between the agent and the environment. Its capacity affects the stability and efficiency of model training. The sampling batch is the number of samples extracted from the experience replay pool for training each time. The discount factor is used to calculate the present value of future rewards, balancing short-term and long-term rewards. The soft-update target network decision parameter controls the frequency and amplitude of target network parameter updates. The learning rate determines the step size of parameter updates. The action space dimension corresponds to the number of candidate decisions; the state space dimension corresponds to the dimension of the state vector. As an optional implementation, the initial values ​​of each hyperparameter are determined based on the characteristics of the logistics scheduling business and training requirements. For example, the experience replay pool capacity is set to int(1e5), the sampling batch size is set to 64, the discount factor is set to 0.99, the soft-update target network decision parameter is set to 1e-3, the learning rate is set to 5e-4, the action space dimension is set to 10 (corresponding to 10 candidate decisions), and the state space dimension is set to 30 (corresponding to a 30-dimensional state vector).

[0131] Step A32: Initialize the initial network based on the initial value to construct the first network and the second network.

[0132] In this embodiment, initialization refers to the process of setting and adjusting the parameters of the initial network according to determined hyperparameters. The first network and the second network are two neural networks with the same structure but possibly different parameters, corresponding to the subsequent main network and target network respectively. As an optional implementation, the initial network is initialized using the determined initial values ​​of the hyperparameters. First, the dimension of the input layer is set to 30 according to the state space dimension, and the dimension of the output layer is set to 10 according to the action space dimension. Then, the weights and biases of the network are initialized using a random initialization method to ensure that the network parameters have a suitable initial distribution. Based on the same initialization configuration, two independent network instances are created, namely the first network and the second network. The two networks have the same structure and initial parameters, but the parameters will be updated independently during the subsequent training process.

[0133] For example, in an actual training scenario for a logistics scheduling system, the initial values ​​of each hyperparameter are first determined based on business needs and experience. For example, the experience replay pool capacity is set to 100,000, the sampling batch size is set to 64, the discount factor is set to 0.99, the soft update parameter is set to 0.001, the learning rate is set to 0.0005, the action space dimension is set to 10, and the state space dimension is set to 30. These initial values ​​are then used to initialize the initial network. Two network instances are created as the first network and the second network, with a structure of a 30-dimensional input layer, a first hidden layer of 128 nodes, a second hidden layer of 64 nodes, and a 10-dimensional output layer. During subsequent training, the first network will serve as the main network for frequent parameter updates, while the second network will serve as the target network, slowly updating its parameters through a soft update mechanism to provide a stable reference for training.

[0134] This embodiment lays a good foundation for the subsequent training process by reasonably setting hyperparameters and initializing the network, enabling the model to learn effective decision-making strategies in logistics scheduling tasks, improving the accuracy and efficiency of logistics scheduling, and optimizing the allocation and utilization of logistics resources.

[0135] Based on any of the above embodiments, in the seventh embodiment of the present application, refer to Figure 3 , step A40, comprising:

[0136] Step A41: storing the training data set into an experience replay pool.

[0137] In this embodiment, the experience replay pool is a data structure used to store experience samples generated by the interaction between the agent and the environment. Its function is to break the correlation between samples and improve the stability and efficiency of training.

[0138] As an optional implementation, each sample in the training dataset is stored in the experience replay pool in the format of Experience(state, action, reward, next_state, done, order_id). The capacity of the experience replay pool is set to int(1e5). When the pool is full, the earliest sample that enters will be replaced.

[0139] In this embodiment, by introducing the experience replay mechanism, data correlation is broken, training variance is reduced, sample utilization is improved, data distribution is smoothed, local optimality is avoided, and stable updates of the target network are supported. Storage mechanism: The experience tuples (s t , a t , r t , s t+1 , <done>), called transition, is stored in a fixed-capacity experience pool (ReplayBuffer). <done>To mark completion, a Boolean value, distinguished by 0 / 1, is used to mark the end state of a sample, avoiding incorrect assessments and expectations of completed states. In this example, it is used to mark whether all order goods have been dispatched. Random Sampling: Uniformly and randomly extracts a batch of experience (typically 32 / 64 groups) from the experience pool to break the temporal correlation of the data and improve training stability.

[0140] Step A42: acquiring sample data from the experience replay pool based on the sampling batch.

[0141] In this example, a sampling batch refers to the number of samples randomly drawn from the experience replay pool for training. Alternatively, a batch of 64 samples can be randomly drawn from the experience replay pool. This random sampling approach reduces correlation between samples and enables the model to learn from a wider range of experiences.

[0142] Step A43: Calculate the action target value based on each batch of sample data input into the first network, and calculate the action expected value based on each batch of sample data input into the second network.

[0143] In this embodiment, the action target value is the future cumulative reward predicted by the first network for the current state and action, and the action expected value is the second network's evaluation of the possible actions in the next state. As an optional implementation, for each batch of sample data, the state is first input into the first network to calculate the Q value of each action in the current state, i.e., the action target value. The first network then calculates the action selection for the next state. The next state and the selected action are then input into the second network to calculate the action expected value.

[0144] Step A44: Update the parameters of the first network according to the action target value and the action expected value.

[0145] In this embodiment, updating the parameters of the first network is achieved by optimizing the loss function, with the aim of making the prediction of the first network closer to the target value. As an optional implementation, the target Q value is calculated by combining rewards and discount factors, with the formula Q_targets=rewards+(gamma*Q_targets_next*(1-dones)), where gamma is a discount factor with a value of 0.99. Then, the mean square error loss between the Q value predicted by the first network and the target Q value is calculated, i.e., loss=nn.MSELoss()(Q_expected, Q_targets), and the parameters of the first network are updated through the backpropagation algorithm.

[0146] Step A45: Update the parameters of the second network based on the soft update parameters and the first network parameters.

[0147] In this embodiment, soft update is a mechanism for slowly updating target network parameters, which can improve the stability of training.

[0148] As an optional implementation, the parameters of the second network are updated according to the soft update formula theta_target=tau*theta_main+(1-tau)*theta_target, where tau is the soft update coefficient, which takes a value of 1e-3, theta_target is the second network parameter, and theta_main is the first network parameter. In this way, the parameters of the second network slowly approach the parameters of the first network.

[0149] Furthermore, if the soft update target network determination parameter satisfies the soft update condition, step A45 is executed. If the soft update target network determination parameter is a sampling batch, the soft update condition is 20 times, 30 times, 40 times, or 50 times. This is not specifically limited in this embodiment.

[0150] Step A46: If the training converges, generate the main network based on the first network, and generate the target network based on the second network.

[0151] In this embodiment, training convergence means that the performance of the model tends to be stable after multiple iterations and the loss function no longer decreases significantly.

[0152] As an optional implementation, the training process from steps A42 to A45 is continuously executed, monitoring the loss function and model performance on the validation set. Training is considered complete when the loss function converges and the model performance stabilizes. At this point, the first network becomes the trained master network, used for actual logistics scheduling decisions, and the second network becomes the trained target network, providing a stable evaluation reference for decision-making.

[0153] As an optional implementation, convergence is considered achieved and training is terminated when the following conditions are met simultaneously: The agent is considered converged when the average reward for 100 consecutive episodes exceeds 450. The agent's order splitting error rate on the test set is less than 5%. The agent's order fulfillment rate in the simulation environment reaches over 95%.

[0154] For example, in an intelligent scheduling system of a logistics platform, a large amount of historical logistics data is first processed into a training data set and stored in an experience replay pool. After the training starts, 64 samples are randomly drawn from the experience replay pool each time, the action target value is calculated by the first network, and the action expected value is calculated by the second network. Then, the parameters of the first network are updated based on the difference between the two, and the parameters of the second network are updated in a soft update manner. After tens of thousands of iterative training, the loss function of the model gradually converges. At this time, the first network and the second network are determined as the main network and the target network respectively. When there is a new logistics order that needs to be scheduled, the main network can quickly give the optimal scheduling decision, such as selecting the best transportation route and the appropriate means of transportation. The target network provides a stable value assessment for the decision, thereby improving the accuracy and efficiency of the entire logistics scheduling system.

[0155] This embodiment effectively improves the stability of model training and the accuracy of decision-making through experience replay and dual network structure, enabling the logistics scheduling system to better cope with complex and changeable actual scenarios, optimize the allocation of logistics resources, reduce logistics costs, and improve customer satisfaction.

[0156] Based on any of the above embodiments, in Embodiment 8 of the present application, step A43 includes:

[0157] Step A431 : Calculate the sample data through the first network to obtain the action of the next state and the action target value of the action.

[0158] In this embodiment, the first network acts as the main network, used to predict the optimal action and its value based on the current state. The action target value is the first network's estimate of the expected cumulative reward of each action in the current state. The next state S is selected by the Main Network. t+1 The optimal action a*. Among them, .

[0159] As an optional implementation, the current state vector in the sample data is input into the first network, and forward propagation is used to calculate the Q values ​​of all possible actions. The action with the highest Q value is selected as the next state action, and this Q value becomes the corresponding target value for the action. For example, for a sample containing a 30-dimensional state vector, the output layer of the first network generates 10 Q values ​​(corresponding to 10 possible decisions), and the action corresponding to the highest value is selected as the predicted action.

[0160] Step A432: Calculate the action value of the action based on the sample data through the second network.

[0161] In this embodiment, the second network serves as the target network to provide a stable value assessment. The action value is the target network's estimate of the expected value of the action selected by the first network in the future state. As an optional implementation, the next state vector in the sample data and the action selected by the first network are input into the second network, and the Q value of the action under the second network, i.e., the action value, is calculated through forward propagation. Since the parameter update of the second network lags behind that of the first network, it can provide a more stable value assessment and reduce fluctuations during training. The Q value of the selected action a* is calculated by the ‌target network‌. Among them, When the state mark is terminated, done=1, the next state S t+1 The Q value of the terminal state is no longer involved in the current reward calculation (that is, the future reward weight is 0), avoiding invalid estimation of the Q value of the terminal state.

[0162] Step A433: Calculate the action expected value corresponding to the action value by combining the discount factor and the reward function.

[0163] In this embodiment, the discount factor is used to balance the importance of short-term rewards and long-term rewards. The reward function defines the immediate reward for each action based on the logistics scheduling business rules. The expected value of an action is an estimate of the value that takes into account the current reward and the expected future reward. When the state is marked as aborted, done = 1, and the next state S t+1 The Q value of the terminal state is no longer involved in the current reward calculation (that is, the future reward weight is 0), avoiding invalid estimation of the Q value of the terminal state.

[0164] As an optional implementation, the Bellman equation uses the formula action expected value = reward + discount factor × action value × (1-completion flag) to calculate. For example, if the current action receives a reward of 10, the discount factor is 0.99, the action value calculated by the second network is 80, and the task is not completed (completion flag is 0), then the action expected value is 10 + 0.99 × 80 × (1-0) = 89.2.

[0165] For example, in a logistics delivery scenario, an order needs to be shipped from warehouse A to customer B. The state vector of the current sample data contains information such as warehouse inventory, vehicle availability, and road conditions. The first network calculates and selects "dispatch vehicle C to perform the transport task" as the next state action, with a target value of 75. Based on the same state vector and this action, the second network calculates the action value to be 82. Assuming the immediate reward for this action is 5 (e.g., on-time dispatch), the discount factor is 0.99, and the task is not completed, the expected value of the action is 5 + 0.99 × 82 × 1 = 86.18. This process simulates the decision-making and evaluation mechanism of an intelligent agent in real-world scheduling. Through the dual-network structure and discount accumulation, it balances short-term rewards with long-term planning.

[0166] Furthermore, after the agent is trained and put into production, hyperparameter adjustments can be made based on actual operational performance to stabilize production environment decisions and prevent excessive external input factors from distorting system decisions. Specific strategies are as follows: Learning rate adjustment: Initial value: 5e-4. Optimization strategy: Use a learning rate scheduler (such as torch.optim.lr_scheduler.StepLR) with a 0.1x decay every 50 episodes. Discount factor adjustment: Initial value: 0.99. Optimization strategy: For long-tail logistics orders, the discount factor can be adjusted to 0.95 to balance immediate and future rewards. Experience replay pool size adjustment: Initial value: 1e5. Optimization strategy: For high-concurrency order scenarios, this can be increased to 5e5 to capture a wider range of state distributions. Target network update frequency: Initial value: Update every 4 steps. Optimization strategy: For logistics scenarios with high volatility, this can be adjusted to update every 8 steps to enhance stability. Exploration strategy optimization: Initial ε-greedy strategy: eps_start=1.0, eps_end=0.01, eps_decay=0.995. Optimization strategy: Introduce a noise network to replace ε-greedy to reduce randomness in exploration.

[0167] This embodiment uses dual networks to collaboratively calculate action target values ​​and expected values, effectively leveraging the real-time learning capabilities of the primary network and the stable evaluation capabilities of the target network. This improves decision-making accuracy and robustness, enabling the logistics scheduling system to make better scheduling decisions in complex environments, reducing transportation costs and increasing customer satisfaction. Using multi-source data fusion technology, platform, carrier, driver, customer, and external data are modeled and integrated to form a platform data resource for intelligent agents to learn from, making decisions more accurate and more aligned with real-world scenarios. This addresses the robustness issues of expert algorithms and traditional Q-learning techniques. By designing separate action selection and value assessment networks, the logistics decision-making process is decoupled into two independent systems, ensuring they do not interfere with each other. Furthermore, the hidden layers of the deep network can factor in decision factors not explicitly enumerated, increasing the robustness of the agent's decisions and speeding up training convergence. This allows for more accurate decisions using less training data while consuming less computing power.

[0168] The present application provides a scheduling device for a logistics platform, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the scheduling method for the logistics platform in the above-mentioned embodiment one.

[0169] Reference below Figure 4 , which shows a schematic diagram of the structure of a scheduling device suitable for implementing the logistics platform of the embodiments of the present application. The scheduling device of the logistics platform in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptops, tablet computers, etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The scheduling equipment of the logistics platform shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0170] like Figure 4 As shown, the logistics platform's scheduling equipment may include a processing device 1001 (e.g., a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the logistics platform's scheduling equipment. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems may be connected to I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication devices 1009 may allow the logistics platform's dispatching equipment to communicate wirelessly or wired with other devices to exchange data. While the diagram illustrates the logistics platform's dispatching equipment with various systems, it should be understood that implementation or presence of all illustrated systems is not required. More or fewer systems may alternatively be implemented or present.

[0171] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0172] The logistics platform scheduling device provided in this application, which utilizes the logistics platform scheduling method described in the aforementioned embodiment, can resolve the technical issue of difficulty in scheduling orders based on complex and changing internal and external conditions, which in turn leads to low order scheduling accuracy. Compared to the prior art, the beneficial effects of the logistics platform scheduling device provided in this application are the same as those of the logistics platform scheduling device provided in the aforementioned embodiment, and the other technical features of the logistics platform scheduling device are the same as those disclosed in the method of the previous embodiment, and are not further described here.

[0173] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0174] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0175] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer program) stored thereon, and the computer-readable program instructions are used to execute the scheduling method of the logistics platform in the above-mentioned embodiment.

[0176] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0177] The above-mentioned computer-readable storage medium may be included in the scheduling device of the logistics platform; or it may exist independently without being assembled into the scheduling device of the logistics platform.

[0178] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the scheduling device of the logistics platform, the scheduling device of the logistics platform: forms a state vector of a preset dimension from the logistics order data in a standard form;

[0179] Input the state vector into the main network and calculate the target value of each candidate decision. The main network is a fully connected feedforward neural network, including an input layer, at least two hidden layers, and an output layer. The dimension of the input layer is consistent with the preset dimension. The hidden layer uses the target activation function. The dimension of the output layer is consistent with the number of candidate decisions.

[0180] Inputting the state vector and the candidate decision into a target network, and calculating the expected value of each candidate decision, wherein the target network has the same structure as the main network and has a parameter update lag;

[0181] A target decision is determined according to a weighted fusion result of the target value and the expected value, so as to issue a logistics scheduling task based on the target decision.

[0182] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0183] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0184] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0185] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned logistics platform scheduling method. This computer-readable storage medium can address the technical issue of difficulty in scheduling orders based on complex and changing internal and external conditions, which in turn leads to low order scheduling accuracy. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the logistics platform scheduling method provided in the aforementioned embodiments and are not further elaborated here.

[0186] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the scheduling method for the logistics platform as described above.

[0187] The computer program product provided in this application can address the technical problem of difficulty in scheduling orders based on complex and ever-changing internal and external circumstances, which in turn leads to low order scheduling accuracy. Compared to the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as those of the logistics platform scheduling method provided in the aforementioned embodiments, and will not be further elaborated here.

[0188] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the present application.< / done> < / done> < / done>

Claims

1. A scheduling method for a logistics platform, characterized in that: The scheduling method of the logistics platform includes: Based on the data relationship diagram, internal source data and external source data are associated into a data relationship set; Determine the quantitative indicators of factors affecting logistics scheduling. The number of factors affecting logistics scheduling is the same as the dimension of the state vector, which is 30. The factors affecting logistics scheduling include the severity of weather conditions during transportation, the degree of road congestion, the star rating of driver service, the urgency of orders, the backlog of platform orders, the timely fulfillment rate of waybills, the gross profit margin of orders, the total number of order goods, the type of goods, the mode of transportation, whether multimodal transport is required, the fuel consumption of empty trucks, the fuel consumption of full trucks, the length of trucks, the status of driver acceptance of orders, the remaining space in the carriage, the vehicle load, whether the vehicle is returning, the number of multi-stage transportation sections, the distinction between heavy and bulky goods, the weight of goods, the volume of goods, the type of full truck / full container transportation, the distance to the target warehouse, whether there are vacant platforms at the warehouse, the warehouse operating load rate, the carrier's accident rate, the number of carriers' in-transit orders, the carrier's performance rating, and the carrier's capacity load rate. Determine the quantitative value of each data in the data relationship set based on the quantitative index, and then establish a training data set; Construct an initial network based on one input layer, at least two hidden layers, and one output layer; Determining hyperparameters of the initial network based on the training data set to construct a first network and a second network; Each sample in the training dataset is stored in the experience replay pool in the format of Experience(state, action, reward, next_state, done, order_id). Done is the completion mark, which is distinguished by 0 / 1 to mark whether all order goods have been dispatched. Obtaining sample data from the experience replay pool based on a sampling batch; Based on each batch of the sample data input to the first network, the action target value is calculated; based on each batch of the sample data input to the second network, the action expected value is calculated; Updating parameters of the first network according to the action target value and the action expected value; Updating parameters of the second network based on the soft update parameters and the first network parameters; If the training converges, generating a main network based on the first network, and generating a target network based on the second network; Extracting standardized logistics order data from the logistics system database, combining internal and external data, and integrating and quantifying it into a 30-dimensional state vector. The state vector is a fixed-dimensional vector determined by factors affecting logistics scheduling operations. Input the state vector into the main network and calculate the target value of each candidate decision. The main network is a fully connected feedforward neural network, including an input layer, at least two hidden layers, and an output layer. The input layer dimension is consistent with the preset dimension, the hidden layer uses the target activation function, and the output layer dimension is consistent with the number of candidate decisions. Inputting the state vector and the candidate decision into a target network, and calculating the expected value of each candidate decision, wherein the target network has the same structure as the main network and has a parameter update lag; The target decision is determined according to the weighted fusion result of the target value and the expected value, so as to issue the logistics scheduling task based on the target decision. The weighted fusion formula Qfinal=α×Qmain+(1-α)×Qtarget, α is the weight coefficient, Qmain is the target value output by the main network, and Qtarget is the expected value output by the target network. The final Q value of each candidate decision is calculated by the formula, and the candidate decision with the highest Q value is selected as the target decision.

2. The logistics platform scheduling method according to claim 1, characterized in that: If the logistics order data meets the order splitting conditions, determine multiple sub-orders corresponding to the logistics order data; A state vector of the preset dimension is generated based on each of the sub-orders.

3. The logistics platform scheduling method according to claim 1, characterized in that: The step of inputting the state vector and the candidate decisions into a target network and calculating the expected value of each candidate decision comprises: Based on the target values ​​of the candidate decisions, sort and select at least one target candidate decision; Based on the state vector and the target network, the expected value corresponding to each target candidate decision is calculated.

4. The method for scheduling a logistics platform according to claim 1, wherein: The initial network is a fully connected layer structure, and the step of determining the hyperparameters of the initial network based on the training data set to construct the first network and the second network includes: Determine the initial values ​​of the hyperparameters in the experience replay pool, sampling batch, discount factor, soft update target network decision parameter, learning rate, action space dimension, and state space dimension; The initial network is initialized based on the initial value to construct the first network and the second network.

5. The logistics platform scheduling method according to claim 1, characterized in that: The steps of inputting each batch of sample data into the first network to calculate the action target value, and inputting each batch of sample data into the second network to calculate the action expected value, include: Calculating the sample data through the first network to obtain an action for the next state and an action target value for the action; Calculating the action value of the action based on the sample data by the second network; The action expected value corresponding to the action value is calculated by combining the discount factor and the reward function.

6. A dispatching device for a logistics platform, characterized in that: The scheduling device of the logistics platform includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the scheduling method of the logistics platform as described in any one of claims 1 to 5.

7. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the scheduling method for the logistics platform as described in any one of claims 1 to 5 are implemented.