Logistics scene resource scheduling method and system, electronic equipment and storage medium

Through the column constraint generation algorithm and deep reinforcement learning model, the dynamic adaptability and robustness of logistics scenario resource scheduling in a dynamic environment is solved, and the system stability and operation efficiency are improved.

CN120373703APending Publication Date: 2025-07-25WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325275.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Logistics scenario resource scheduling has problems in insufficient adaptability to dynamic environments, coordinated optimization of multi-resources and robust guarantees in dynamic environments, resulting in inefficient resource utilization and operational conflicts, and delays in automated vehicle tasks and imbalance in resource allocation.

Method used

The column constraint generation algorithm is used to predict the job bit allocation results, and the vehicle scheduling results are predicted through the deep reinforcement learning model. The preset deep reinforcement learning model is built with the deep Q network and the probability strategy optimization algorithm, and the automated vehicle scheduling is dynamically adjusted to meet cost conditions.

Benefits of technology

The system stability and operation efficiency of logistics scenario resource scheduling in complex disturbance scenarios is improved, and task delays and resource allocation imbalances are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373703A_ABST
    Figure CN120373703A_ABST
Patent Text Reader

Abstract

The invention discloses a logistics scene resource scheduling method and system, electronic equipment and a storage medium. The method comprises the following steps: acquiring preset job information; predicting through a column constraint generation algorithm according to the preset job information to obtain an expected job position distribution result; according to the expected operation position distribution result and preset loading and unloading information, an expected carrier scheduling result is obtained through prediction of a preset deep reinforcement learning model; wherein the preset deep reinforcement learning model is constructed by a deep Q network algorithm and a probability strategy optimization algorithm; and when it is determined that the expected carrier scheduling result and the expected operation position distribution result meet a preset cost condition, determining a target scheduling distribution result according to the expected carrier scheduling result and the expected operation position distribution result. According to the embodiment of the invention, the system stability and operation efficiency of logistics scene resource scheduling in a complex disturbance scene can be effectively improved. The method can be widely applied to the technical field of logistics automation and resource scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of logistics automation and resource scheduling, and in particular, to a method, system, electronic device, and storage medium for resource scheduling in a logistics scenario. Background Art

[0002] With the rapid development of the global economy and the popularization of e-commerce, logistics automation scenarios (such as air cargo hubs, port terminals, intelligent manufacturing factories, etc.) are facing the dual challenges of improving operational efficiency and resource collaborative management. In related technologies, there are deficiencies in the dynamic environment adaptability, multi-resource collaborative optimization, and robustness guarantee in the process of logistics scenario resource scheduling. For example, the resource utilization is inefficient and there are operation conflicts, making it difficult to adapt to multi-source disturbances (such as flight delays, fluctuations in ship arrivals, changes in order priorities, etc.). Moreover, related automated vehicles (such as AGVs, straddle carriers, automated forklifts, etc.) under disturbances such as task delays, path blockages, and equipment failures exacerbate task execution delays and resource allocation imbalances.

[0003] In summary, the technical problems existing in the related technologies need to be improved. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a method, system, electronic device, and storage medium for resource scheduling in a logistics scenario, which can effectively improve the system stability and operation efficiency of logistics scenario resource scheduling in complex disturbance scenarios.

[0005] To achieve the above object, on the one hand, an embodiment of this application proposes a method for resource scheduling in a logistics scenario, and the method includes the following steps:

[0006] Obtain preset job information;

[0007] Predict the expected job position allocation result through a column constraint generation algorithm according to the preset job information;

[0008] Predict the expected vehicle scheduling result according to the expected job position allocation result and preset loading and unloading information through a preset deep reinforcement learning model; wherein, the preset deep reinforcement learning model is constructed by a deep Q-network algorithm and a probabilistic policy optimization algorithm;

[0009] When it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost condition, determine the target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result.

[0010] In some embodiments, the preset job information includes dynamic task unit information and job position information;

[0011] The predicted desired job position allocation result according to the preset job information through the column constraint generation algorithm includes:

[0012] Construct a preset mixed-integer programming model based on the dynamic task unit information and job position information, and calculate the initial job position allocation data through the preset mixed-integer programming model;

[0013] Identify a preset conflict scenario based on the initial job position allocation data;

[0014] Determine the key conflict constraints according to the conflict constraint conditions corresponding to the preset conflict scenario;

[0015] Generate candidate columns according to the key conflict constraints, and add the candidate columns to the preset mixed-integer programming model for optimization to obtain the target mixed-integer programming model;

[0016] Predict the desired job position allocation result through the target mixed-integer programming model.

[0017] In some embodiments, the determining the key conflict constraints according to the conflict constraint conditions corresponding to the preset conflict scenario includes:

[0018] Screen the conflict constraint conditions through a variable neighborhood search algorithm to obtain the key conflict constraints.

[0019] In some embodiments, before performing the prediction of the desired vehicle scheduling result through the preset deep reinforcement learning model according to the desired job position allocation result and the preset loading and unloading information, the method further includes:

[0020] Input the preset state data into the initial deep Q-network model to select a preset scheduling strategy;

[0021] Optimize the preset scheduling strategy through the probability policy optimization algorithm to update the initial deep Q-network model to obtain the preset deep reinforcement learning model.

[0022] In some embodiments, the optimizing the preset scheduling strategy through the probability policy optimization algorithm to update the initial deep Q-network model to obtain the preset deep reinforcement learning model includes:

[0023] Calculate a preset advantage value according to the preset scheduling strategy through a distributed robust optimization algorithm, and perform parameter optimization according to the preset advantage value to obtain an optimized strategy;

[0024] Generate corresponding policy actions according to the optimized strategy, and perform environment interaction according to the policy actions to calculate policy reward data;

[0025] When it is determined that the policy reward data is greater than the historical reward data, the network parameters of the initial deep Q-network model are updated according to the optimization policy to obtain the preset deep reinforcement learning model.

[0026] In some embodiments, calculating a preset advantage value according to the preset scheduling policy through a distributed robust optimization algorithm to perform parameter optimization according to the preset advantage value to obtain an optimization policy includes:

[0027] Construct an uncertainty set;

[0028] Perform a worst-case advantage estimation on the preset scheduling policy through the uncertainty set to obtain a robust advantage value, and perform policy parameter optimization according to the robust advantage value to obtain an optimization policy.

[0029] In some embodiments, when it is determined that the expected vehicle scheduling result and the expected job position allocation result satisfy a preset cost condition, determining a target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result includes:

[0030] Perform cost calculation according to the expected vehicle scheduling result and the expected job position allocation result to obtain preset cost data; wherein, the preset cost data includes delay cost data and operation cost data;

[0031] When it is determined that the preset cost data is less than a preset cost threshold, use the expected vehicle scheduling result and the expected job position allocation result as the target scheduling allocation result.

[0032] To achieve the above object, another aspect of the embodiments of the present application proposes a logistics scenario resource scheduling system, and the system includes:

[0033] A first module for obtaining preset job information;

[0034] A second module for predicting an expected job position allocation result according to the preset job information through a column constraint generation algorithm;

[0035] A third module for predicting an expected vehicle scheduling result according to the expected job position allocation result and preset loading and unloading information through a preset deep reinforcement learning model; wherein, the preset deep reinforcement learning model is constructed by a deep Q-network algorithm and a probabilistic policy optimization algorithm;

[0036] A fourth module for determining a target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result when it is determined that the expected vehicle scheduling result and the expected job position allocation result satisfy a preset cost condition.

[0037] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes:

[0038] At least one processor;

[0039] At least one memory for storing at least one program;

[0040] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0041] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0042] The embodiments of the present application at least include the following beneficial effects: The present application provides a method, system, electronic device and storage medium for resource scheduling in a logistics scenario. This solution obtains preset job information, and predicts the expected job position allocation result through a column constraint generation algorithm according to the preset job information. Then, according to the expected job position allocation result and the preset loading and unloading information, the expected vehicle scheduling result is predicted through a preset deep reinforcement learning model constructed by a deep Q-network algorithm and a probabilistic policy optimization algorithm. Finally, when it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost condition, the embodiment of the present invention determines the target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result, realizes the resource scheduling of the logistics scenario, and can effectively improve the system stability and operation efficiency of the logistics scenario resource scheduling in a complex disturbance scenario. Description of the Drawings

[0043] Figure 1 is a flowchart of the method for resource scheduling in a logistics scenario provided by an embodiment of the present invention;

[0044] Figure 2 is a schematic diagram of the algorithm architecture for predicting the expected job position allocation result through a column constraint generation algorithm according to the preset job information provided by an embodiment of the present invention;

[0045] Figure 3 is a schematic diagram of the algorithm architecture for constructing a preset deep reinforcement learning model provided by an embodiment of the present invention;

[0046] Figure 4 is a schematic diagram of the application scenario of the method for apron allocation and AGV scheduling in an air cargo hub under multiple disturbances provided by an embodiment of the present invention;

[0047] Figure 5 is a framework diagram of the overall process of the two-stage model provided by an embodiment of the present invention;

[0048] Figure 6 It is a schematic structural diagram of the logistics scenario resource scheduling system provided by an embodiment of the present invention;

[0049] Figure 7 It is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.

[0051] It can be understood that the terms "first", "second", etc. used in the present application can be used in this document to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".

[0052] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0054] Before elaborating on the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application will be described first. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0055] With the rapid development of the global economy and the popularization of e-commerce, logistics automation scenarios (such as air cargo hubs, port terminals, intelligent manufacturing factories, etc.) are facing the dual challenges of improving operational efficiency and resource collaborative management. In related technologies, there are deficiencies in the adaptability of the logistics scenario resource scheduling process to dynamic environments, multi-resource collaborative optimization, and robustness guarantee. For example, inefficient resource utilization and operation conflicts make it difficult to adapt to multi-source disturbances (such as flight delays, fluctuations in ship arrivals, changes in order priorities, etc.). Moreover, related automated vehicles (such as AGVs, straddle carriers, automated forklifts, etc.) under disturbances such as task delays, path blockages, and equipment failures exacerbate task execution delays and resource allocation imbalances.

[0056] Based on this, the embodiments of the present invention provide a logistics scenario resource scheduling method, which can effectively improve the system stability and operation efficiency of logistics scenario resource scheduling in complex disturbance scenarios. Correspondingly, Figure 1 is an optional flowchart of the logistics scenario resource scheduling method provided by the embodiments of the present invention, Figure 1 and the method in it may include but are not limited to steps S110 to S140.

[0057] Step S110: Obtain preset operation information.

[0058] Step S120: Predict the expected job position allocation result through a column constraint generation algorithm according to the preset operation information.

[0059] Step S130: Predict the expected vehicle scheduling result according to the expected job position allocation result and the preset loading and unloading information through a preset deep reinforcement learning model. Among them, the preset deep reinforcement learning model is constructed by a deep Q-network algorithm and a probabilistic policy optimization algorithm.

[0060] Step S140: When it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost condition, determine the target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result.

[0061] During the operation of this specific embodiment, the embodiment of the present invention first obtains preset operation information, and predicts the expected job position allocation result through a column constraint generation algorithm according to the preset operation information. Specifically, the preset operation information in the embodiment of the present invention refers to the operation data of the corresponding logistics scenario, such as flight information, time data, etc. in an airport freight hub. Correspondingly, the column constraint generation algorithm in the embodiment of the present invention refers to a dynamic optimization algorithm, that is, by dynamically adding variables (columns) to gradually construct the constraint conditions of a linear programming problem, so as to achieve the purpose of optimizing the problem. Among them, the embodiment of the present invention predicts the expected job position allocation result by gradually introducing scenarios that violate the constraints according to the preset operation information, thereby improving the solution efficiency. Next, the embodiment of the present invention predicts the expected vehicle scheduling result through a preset deep reinforcement learning model according to the expected job position allocation result and the preset loading and unloading information. Specifically, the preset deep reinforcement learning model in the embodiment of the present invention is constructed through a deep Q-network algorithm and a probabilistic policy optimization algorithm. Among them, the deep Q-network (DQN) algorithm in the embodiment of the present invention is a reinforcement learning algorithm that combines deep learning and Q-learning, and approximates the Q-value function through a deep neural network, so as to solve the reinforcement learning problem in a high-dimensional state space. At the same time, the probabilistic policy optimization (PPO) algorithm in the embodiment of the present invention is a class of policy-based reinforcement learning algorithms, where the policy is a probability distribution from states to actions. It is easy to understand that the path of the automated vehicle in the embodiment of the present invention is regarded as known information and no path conflict analysis is considered. Only different automated vehicles are considered to carry goods such as containers loaded and unloaded at different job positions, and the arrival time of logistics transportation (such as airplanes, cargo ships, etc.) is uncertain. Therefore, the automated vehicle scheduling needs to have the ability to dynamically adjust. Based on this, the embodiment of the present invention introduces deep reinforcement learning (DQN+PPO) to solve this problem. Correspondingly, the embodiment of the present invention uses a deep Q-network (DQN) to learn the scheduling decision of the automated vehicle, and learns the automated vehicle scheduling strategy in different state environments based on probabilistic policy optimization (PPO), and constructs a preset deep reinforcement information model to combine the expected job position allocation result and the preset loading and unloading information to predict the expected vehicle scheduling result. Among them, the preset loading and unloading information refers to the loading and unloading data of the goods, such as loading and unloading time, vehicle parameter data, etc. Finally, the embodiment of the present invention analyzes whether the predicted expected vehicle scheduling result and the expected job position allocation result meet the preset cost conditions. Among them, the predicted cost conditions are obtained by pre-customized settings. Correspondingly, when it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost conditions, the embodiment of the present invention determines the target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result, that is, takes the expected vehicle scheduling result and the expected job position allocation result as the target scheduling allocation result.

[0062] In some embodiments of the present invention, the preset operation information includes dynamic task unit information and operation position information. Correspondingly, in the embodiments of the present invention, the expected operation position allocation result is predicted through a column constraint generation algorithm according to the preset operation information, including but not limited to the following steps:

[0063] Construct a preset mixed-integer programming model according to the dynamic task unit information and the operation position information, so as to calculate the initial operation position allocation data through the preset mixed-integer programming model.

[0064] Identify a preset conflict scenario based on the initial operation position allocation data.

[0065] Determine the key conflict constraints according to the conflict constraint conditions corresponding to the preset conflict scenario.

[0066] Generate candidate columns according to the key conflict constraints, and add the candidate columns to the preset mixed-integer programming model for optimization to obtain the target mixed-integer programming model.

[0067] Predict the expected operation position allocation result through the target mixed-integer programming model.

[0068] In this specific embodiment, the dynamic task unit information refers to the dynamic plan data of the logistics task unit (such as an aircraft flight, a shipping vessel), such as the planned arrival time, the planned departure time, the dynamic task unit model, etc. In addition, the operation position information in the embodiments of the present invention refers to the operation position parameter data, such as the operation position size, the position information. For example, in the scenario of an airport freight hub, the operation position information includes the apron size. Correspondingly, as Figure 2 shown, in the embodiments of the present invention, a preset mixed-integer programming model is first constructed according to the dynamic task unit information and the operation position information, so as to calculate the initial operation position allocation data through the preset mixed-integer programming model. Specifically, in the embodiments of the present invention, according to the known dynamic task unit information (including the planned arrival time, the planned departure time, the dynamic task unit model) and the operation position information (including the operation position size, the position), an initial mixed-integer programming model is constructed to solve the initial operation position allocation plan and obtain the initial operation allocation data. Then, in the embodiments of the present invention, a preset conflict scenario is identified based on the initial operation position allocation data, and the key conflict constraints are determined according to the conflict constraints corresponding to the preset conflict scenario. Specifically, the preset conflict scenario in the embodiments of the present invention refers to a scenario that violates the constraints, such as the operation position conflict caused by the delay of the dynamic task unit. Among them, in the embodiments of the present invention, under the current operation position allocation plan, all violated constraint sets C = {c1, c2,..., c n}, a preset conflict scenario is obtained. Correspondingly, these scenarios in the embodiments of the present invention are generated by an uncertainty set (such as the fluctuation of the actual arrival time of dynamic task units). Then, the embodiments of the present invention screen out the key conflict constraints from the constraints (conflict constraint conditions) violated by these preset conflict scenarios, that is, the constraints whose impact on the current solution is greater than a preset threshold. For example, the embodiments of the present invention define multiple sets of domain structures N k {k = 1, 2, …, K}, for example, N1 represents selecting the top 5 constraints with the highest violation degree each time (sorted by delay cost), N2 represents randomly selecting 10% of the conflict constraints, and N3 represents grouping based on the constraint coupling relationship (such as constraints related to the same job position are grouped as one set). Then for each domain N k the constraint subset C k ∈ C to calculate its criticality score. Among them, the scoring rule in the embodiments of the present invention is shown in the following formula (1):

[0069] S = μ1 * W st + μ2 * S ci (1)

[0070] where, in the formula, W st represents the constraint violation degree value, S ci is the expected improvement value of the objective function (total cost) after adding this constraint, and μ1, μ2 are weight coefficients, which are set to 0.2 and 0.8 respectively. According to the scoring rule, in the neighborhood N k select the 3 constraints with the highest scores as candidate key constraints and give priority to adding them to the model.

[0071] Furthermore, the embodiments of the present invention generate candidate columns according to the key conflict constraints, add the candidate columns to the preset mixed integer programming model for optimization to obtain the target mixed integer programming model, and then predict the expected job position allocation result through the target mixed integer programming model. Specifically, the embodiments of the present invention generate new candidate columns from the screened key conflict constraints and add them to the optimization model to re-solve the MIP model. If the constraints screened in the current neighborhood can significantly improve the objective function, these constraints are retained and the current neighborhood structure is fixed. Correspondingly, the embodiments of the present invention re-identify the preset conflict scenario to determine whether there are new conflict constraints generated. If so, continue to execute the subsequent steps of generating candidate columns and optimizing the preset mixed integer programming model until no new conflict constraints are generated, construct the target mixed integer programming model, and output the optimal job position allocation plan (expected job position allocation result).

[0072] Exemplarily, in the problem of allocating job positions to task units, the parameters and variables of the task unit allocation to job position model in the embodiments of the present invention are shown in Table 1 below:

[0073] Table 1

[0074]

[0075]

[0076] Accordingly, the constructed model constraints include:

[0077] Constraining the non - overlap of time and space of two task units, as shown in the following formula (2):

[0078]

[0079] Constraining the sequence and position relationship between two dynamic task units α and β in the mixed - integer programming model, as shown in the following formula (3):

[0080] x αβ ·y αβ =1 (3)

[0081] Ensuring that at most one dynamic task unit can be parked at a job position at the same time, as shown in the following formula (4):

[0082]

[0083] Ensuring that for two dynamic task units at the same job position, the latter dynamic task unit can enter the job position only after the former arriving dynamic task unit has finished loading and left, as shown in the following formula (5):

[0084]

[0085] Constraining the size matching between the dynamic task unit and the job position, as shown in the following formula (6):

[0086]

[0087] In some embodiments of the present invention, determining the key conflict constraints according to the conflict constraints corresponding to the preset conflict scenarios includes, but is not limited to, the following steps:

[0088] Screening the conflict constraints through a variable - neighborhood search algorithm to obtain the key conflict constraints.

[0089] In this specific embodiment, after the embodiment of the present invention dynamically identifies a preset conflict scenario (such as a conflict scenario like flight delay) by using the column constraint generation (CCG) algorithm, it combines variable neighborhood search (VNS) to screen key constraints and generates a robust job position allocation scheme. Specifically, the variable neighborhood search algorithm in the embodiment of the present invention refers to an improved local search algorithm, which escapes from local optima by switching between different neighborhood structures and gradually improves the quality of the solution. Correspondingly, the embodiment of the present invention screens the conflict constraint conditions through the variable neighborhood search algorithm, and in the way of introducing the variable neighborhood search algorithm, screens out the most critical conflict constraints and reduces redundant calculations. Correspondingly, the variable neighborhood search algorithm quickly finds the constraint that has the greatest impact on the current solution through the combination of local search and global search.

[0090] In some embodiments of the present invention, before executing the expected vehicle scheduling result predicted by the preset deep reinforcement learning model according to the expected job position allocation result and the preset loading and unloading information, the logistics scenario resource scheduling method provided by the embodiment of the present invention further includes but is not limited to the following steps:

[0091] Input the preset state data into the initial deep Q-network model to select a preset scheduling strategy.

[0092] Optimize the preset scheduling strategy through the probability policy optimization algorithm to update the initial deep Q-network model and obtain the preset deep reinforcement learning model.

[0093] In this specific embodiment, the embodiment of the present invention first inputs preset state data into the initial deep Q-network model to select a preset scheduling strategy, and then optimizes the preset scheduling strategy through a probabilistic policy optimization algorithm to update the initial deep Q-network model, obtaining a preset deep reinforcement learning model. Specifically, the embodiment of the present invention constructs a deep reinforcement learning model (DQN+PPO) according to preset state data, such as dynamic job position allocation results (expected job position allocation results), vehicle capacity constraints, and task real-time status, to schedule vehicles (such as AGVs) in real time to complete container handling tasks. Among them, the reinforcement learning state space covers vehicle position, task progress, and job position occupancy rate, the action space includes automated vehicle allocation and priority adjustment, and the reward function integrates multiple objectives such as transportation efficiency and delay penalty. Among them, in the embodiment of the present invention, the path of the automated vehicle is regarded as known information and no path conflict analysis is considered. Only the handling of goods such as containers loaded and unloaded at different job positions by different automated vehicles is considered, and the arrival time of task units (such as airplanes) is uncertain. Therefore, the automated vehicle scheduling needs to have the ability of dynamic adjustment. Based on this, the embodiment of the present invention introduces deep reinforcement learning (DQN+PPO) to solve this problem. Correspondingly, the embodiment of the present invention learns the scheduling decision of the automated vehicle through a deep Q-network (DQN), and learns the automated vehicle scheduling strategy in different state environments based on probabilistic policy optimization (PPO), thereby constructing a preset deep reinforcement learning model.

[0094] Exemplarily, such as Figure 3As shown in the figure, taking the airport freight hub as an example, in the embodiments of the present invention, the allocation of the current dynamic operation positions of each task unit type, the loading and unloading of goods such as containers in the task unit type, and information such as the number and load-bearing capacity of automated vehicles are input as the initial state. The scheduling strategy learned through DQN in the early stage selects the scheduling scheme of the vehicle (such as AGV) as the action output according to the current state, and calculates the reward value through environmental feedback. In addition, for the design of the reward function, the embodiments of the present invention mainly consider minimizing the total transportation time (the shorter the time for the automated vehicle to complete all transportation tasks, the higher the reward) and delay penalty (if the automated vehicle fails to complete the container handling task on time, it will receive a negative reward). During the training process, DQN adopts an experience replay mechanism to avoid data correlation problems and improve training stability. For example, the embodiments of the present invention input the allocation of the current dynamic task units' operation positions, the container loading and unloading information in the dynamic task units, the number of automated vehicles, the load-bearing capacity of automated vehicles, etc. as the initial state, and learn the scheduling strategy of the automated vehicle through DQN, and select the scheduling scheme of the automated vehicle as the action output according to the multi-dimensional real-time information set state of the current scenario's dynamic environment (such as the real-time state of the vehicle, the dynamic information of the task unit, the occupation of the operation position and resources, and environmental disturbances, etc.). Among them, in the embodiments of the present invention, DQN estimates the long-term return of each action through the Q-value function and selects the action with the highest Q value. Then, the embodiments of the present invention calculate the reward value according to the environmental feedback, and the reward function is the combination of minimizing the total transportation time and delay penalty, as shown in the following formula (7):

[0095]

[0096] Wherein, is the actual start time of container i being carried by automated vehicle j, and P f is the delay penalty of dynamic task unit f.

[0097] In addition, in the embodiments of the present invention, the deep Q-network model adopts an experience replay mechanism to store historical states, actions, rewards, and next states to avoid data correlation problems and improve training stability.

[0098] In addition, for the model of scheduling the automated vehicle to carry the containers in the dynamic task units parked in each operation position, the parameters and variables are shown in Table 2 below:[[]]

[0099] Table 2

[0100]

[0101] Correspondingly, the model constraints constructed by the embodiments of the present invention at this stage are as follows:

[0102] First, in the embodiments of the present invention, an automated vehicle can only carry one container at a time, as shown in the following formula (8):

[0103]

[0104] Meanwhile, the time interval between the handling tasks of the same automated vehicle should be greater than or equal to the handling time plus the relaxation time (which can be understood as the time for the container to get on and off the automated vehicle), as shown in the following formula (9):

[0105]

[0106] In addition, for the handling constraints of the container: for any container, it can only be carried by one automated vehicle, as shown in the following formula (10):

[0107]

[0108] In addition, for the load-bearing constraints of the automated vehicle: for any automated vehicle, the containers carried need to be within the load-bearing range, as shown in the following formula (11):

[0109]

[0110] Meanwhile, the order of the containers carried by the same automated vehicle within the same dynamic task unit is emphasized, as shown in the following formula (12):

[0111] (t kj -t ij )·S ik >0 (12)

[0112] Correspondingly, in the embodiments of the present invention, within the same dynamic task unit, after the containers that need to be moved out within the dynamic task unit are carried, the containers that are to be carried into the task unit can start to be carried, as shown in the following formula (13):

[0113]

[0114] In some embodiments of the present invention, the preset scheduling strategy is optimized through a probability strategy optimization algorithm to update the initial deep Q-network model, and a preset deep reinforcement learning model is obtained, including but not limited to the following steps:

[0115] The preset advantage value is calculated through a distributed robust optimization algorithm according to the preset scheduling strategy, and parameter optimization is performed according to the preset advantage value to obtain an optimized strategy.

[0116] The corresponding policy actions are generated according to the optimized strategy, and environmental interaction is performed according to the policy actions to calculate the policy reward data.

[0117] When it is determined that the policy reward data is greater than the historical reward data, the network parameters of the initial deep Q-network model are updated according to the optimized policy to obtain a preset deep reinforcement learning model.

[0118] In this specific embodiment, the embodiment of the present invention first calculates a preset advantage value through a distributed robust optimization algorithm according to a preset scheduling policy, and performs parameter optimization according to the preset advantage value to obtain an optimized policy. Specifically, after the DQN training converges, the embodiment of the present invention further optimizes the automated vehicle scheduling policy through the PPO algorithm. Among them, PPO limits the update amplitude by clipping the probability ratio, and calculates the probability ratio r t (θ). Among them, the probability ratio in the embodiment of the present invention: r t (θ) = π θ1 (a t |s t ) / π θ0 (a t |s t ). Accordingly, the embodiment of the present invention limits the clipping probability in the interval [1 - ε, 1 + ε] to prevent the policy update from being too large and causing training oscillation. The embodiment of the present invention sets ε = 0.2. Among them, the loss function for policy optimization in the embodiment of the present invention is: In the formula represents the generalized advantage estimate value. At the same time, in order to cope with the uncertainty of the arrival time of dynamic task units, the embodiment of the present invention introduces a distributed robust optimization (DRO) framework, calculates the preset advantage value corresponding to the preset scheduling policy through the distributed robust optimization algorithm, and thus performs parameter optimization, that is, policy update, to obtain an optimized policy. Then, the embodiment of the present invention generates corresponding policy actions according to the optimized policy, and performs environment interaction according to the policy actions to calculate the policy reward function. Specifically, the embodiment of the present invention determines the corresponding policy actions according to the optimized policy obtained after optimization, and executes the corresponding policy actions, so as to calculate the policy reward data according to the data feedback by the environment interaction. Then, the embodiment of the present invention determines whether the policy reward data corresponding to the policy action is greater than the historical reward data. Among them, the historical reward data in the embodiment of the present invention refers to the policy reward data corresponding to the previous policy action, that is, it is determined whether the policy reward data corresponding to the current policy action is greater than the policy reward data corresponding to the previous policy action. Accordingly, when it is determined that the policy reward data corresponding to the current policy action is greater than the historical reward data, the embodiment of the present invention updates the network parameters of the initial deep Q-network model according to the policy optimization strategy to obtain a preset deep reinforcement model.

[0119] In some embodiments of the present invention, calculating a preset advantage value through a distributed robust optimization algorithm according to a preset scheduling policy, and performing parameter optimization according to the preset advantage value to obtain an optimized policy includes, but is not limited to, the following steps:

[0120] Construct an uncertainty set.

[0121] Perform a worst-case dominance estimation on a preset scheduling strategy through the uncertainty set to obtain a robust dominance value, and optimize the policy parameters according to the robust dominance value to obtain an optimized strategy.

[0122] In this specific embodiment, the embodiment of the present invention first constructs an uncertainty set, performs a worst-case dominance estimation on a preset scheduling strategy through the uncertainty set to obtain a robust dominance value, and then optimizes the policy parameters according to the robust dominance value to obtain an optimized strategy. Specifically, as Figure 3 shown, after the training of DQN converges, the embodiment of the present invention further optimizes the strategy through the PPO algorithm. The update process of the PPO algorithm adopts a dual optimization method and combines the distributed robust optimization (DRO) framework to ensure that the scheduling scheme still has strong adaptability and robustness in the case of uncertain arrival times of task unit types. Correspondingly, in each iteration of the embodiment of the present invention, the DRO algorithm constructs an uncertainty set, considers the worst-case dynamic task unit arrival time fluctuations, generates robust constraints according to the uncertainty set, and ensures the robustness of the scheduling scheme under uncertainty. At the same time, by real-time monitoring the changes in the arrival times of dynamic task units, the automated vehicle scheduling plan is dynamically adjusted to ensure that the scheduling scheme under multiple disturbances has high adaptability and robustness.

[0123] Exemplarily, what the embodiment of the present invention considers is that the arrival time of dynamic task units is uncertain due to various factors such as weather conditions or mechanical failures. It is assumed that the probability distribution of the arrival time is unknown, and the arrival time changes within the uncertainty set. The embodiment of the present invention considers an uncertainty set, which is defined as shown in the following formula (14):

[0124]

[0125] wherein, in the formula, EA f is the expected arrival time of the dynamic task unit, is the maximum allowable error of the dynamic task unit. Based on this, the model in this stage of the embodiment of the present invention will need to be adjusted. At the same time, since the results of the first stage and the second stage affect each other, for the problem of job position allocation and automated vehicle scheduling for the arrival time uncertainty of the entire cargo dynamic task unit, the embodiment of the present invention will adjust the above two-stage models according to the uncertainty and integrate them into a main problem model in a mathematical model. The variables that need to be redefined are as follows:

[0126] The actual arrival time of the dynamic task unit f;

[0127] The actual time required for the container to be loaded and unloaded in the dynamic task unit f;

[0128] The actual queuing time of the dynamic task unit f;

[0129] The time when the container i starts to be actually transported by the automated vehicle j;

[0130] The actual time interval during which the container in the dynamic task unit f can be transported by the automated vehicle;

[0131] O w : The minimum - maximum total actual completion time of work.

[0132] Correspondingly, while keeping the decision variables unchanged, with the goal of minimizing the total delay cost of the dynamic task unit and the airport operation cost as the overall objective of the model, the overall robust model is established as shown in the following formula (15):

[0133]

[0134] It is easy to understand that the core algorithm of the embodiment of the present invention is based on a two - stage distributed robust optimization framework, aiming to solve the problem of dynamic job position allocation and automated vehicle scheduling in an airport freight hub under uncertainty. The algorithm decomposes the problem into a master problem and a sub - problem, respectively processes the dynamic job position and automated vehicle scheduling, and realizes the global optimal solution through iterative optimization.

[0135] In some embodiments of the present invention, when it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost condition, the target scheduling allocation result is determined according to the expected vehicle scheduling result and the expected job position allocation result, including but not limited to the following steps:

[0136] Cost calculation is performed according to the expected vehicle scheduling result and the expected job position allocation result to obtain the preset cost data. Among them, the preset cost data includes delay cost data and operation cost data.

[0137] When it is determined that the preset cost data is less than the preset cost threshold, the expected vehicle scheduling result and the expected job position allocation result are used as the target scheduling allocation result.

[0138] In this specific embodiment, the embodiment of the present invention first calculates the cost based on the expected vehicle scheduling result and the expected job position allocation result to obtain preset cost data, and then determines whether the preset cost data is less than the preset cost threshold. When it is determined that the preset cost data is less than the preset cost threshold, the embodiment of the present invention uses the expected vehicle scheduling result and the expected job position allocation result as the target scheduling allocation result. Specifically, the preset cost data in the embodiment of the present invention includes delay cost data and operation cost data. Among them, the delay cost data refers to the total delay cost of the task unit, and the operation cost data refers to the operation cost of the operation scenario. Correspondingly, the embodiment of the present invention determines whether the currently predicted expected vehicle scheduling result and the expected job position allocation result meet the requirements by judging whether the total delay cost of the task unit and the operation cost of the operation scenario are less than the preset cost threshold. For example, when the embodiment of the present invention determines that the total delay cost of the current task unit and the operation cost of the operation scenario reach the minimum, the currently determined expected vehicle scheduling result and the expected job position allocation result are used as the target scheduling allocation result.

[0139] Next, in combination with a specific application example of logistics scenario resource scheduling, the solution of the embodiment of the present invention will be introduced and described in detail:

[0140] Exemplarily, taking an air cargo hub as an example, the cargo loading and unloading process includes two key steps (such as Figure 4 ): First is the dynamic job position allocation - the aircraft docks at the apron waiting for loading and unloading; second is the automated vehicle scheduling - the AGV (carries the container to the yard or transports it in the reverse direction. Such as Figure 5As shown in the figure, the embodiment of the present invention constructs a two-stage optimization framework based on the above process. In the first stage, based on the task unit type attributes (aircraft flight information), dynamic operation position information (gate size), and task planned time (planned time), the operation position allocation is optimized through a mixed integer programming model. The goal is to maximize the utilization rate of the operation position and minimize the queuing waiting time of the operation unit. During this process, the embodiment of the present invention uses the column constraint generation (CCG) algorithm to dynamically identify conflict scenarios such as flight delays, and combines variable neighborhood search (VNS) to screen key constraints to generate a robust operation position allocation scheme. In the second stage, a deep reinforcement learning model (DQN+PPO) is constructed based on the dynamic operation position allocation result, vehicle capacity constraint, and real-time task status to schedule AGVs to complete container handling tasks in real time. Among them, the reinforcement learning state space covers vehicle position, task progress, and operation position occupancy rate. The action space includes automated vehicle allocation and priority adjustment. The reward function integrates multiple objectives such as transportation efficiency and delay penalty. At the same time, distributed robust optimization (DRO) is introduced to address the uncertainty of the arrival time of the operation unit to ensure the strong adaptability of the scheduling scheme. Among them, in the embodiment of the present invention, the two stages are collaboratively optimized through two-way feedback. Its closed-loop feedback mechanism includes: dynamically adjusting the dynamic operation position allocation strategy based on the real-time status of the automated vehicle (such as working status, workload), and generating a robust constraint set through distributed robust optimization (DRO). In the forward link, the operation position layout directly affects the AGV transportation distance and efficiency; in the reverse link, the real-time status of the AGV (such as task backlog) is fed back to the first stage, triggering dynamic adjustment of the gate allocation, such as preferentially allocating proximal gates to reduce the global transportation cost. Correspondingly, the closed-loop mechanism in the embodiment of the present invention can be extended to port terminals and intelligent manufacturing scenarios. For example, in the port scenario, the dynamic operation position corresponds to the ship berth, and the automated vehicle is the straddle carrier. Through collaborative optimization, the ship waiting time is reduced. In the manufacturing scenario, the dynamic operation position is the production line workstation, and the automated vehicle is the material AGV. Collaborative optimization reduces the risk of production line stagnation. The two-stage framework of the embodiment of the present invention supports rapid cross-industry adaptation through general model design and modular parameter configuration, significantly improving the resource collaboration efficiency and system stability in multi-disturbance scenarios, and providing a highly robust and low-cost intelligent scheduling solution for the airport, port, and manufacturing fields. For the above two-stage model, the embodiment of the present invention extracts the sub-problem model and the main problem model of the model. The sub-problem is the automated vehicle scheduling plan for loading and unloading goods for dynamic task unit types with uncertain arrival times in the second stage of the model, and the main problem is the dynamic operation position allocation problem that needs to consider the output of the sub-problem.

[0141] It is easy to understand that through the innovative general model design and modular architecture in the embodiments of the present invention, efficient optimization scheduling of air cargo hubs and multi-industry logistics scenarios in complex dynamic environments is achieved, significantly improving the system operation efficiency and intelligent level. This resource scheduling method supports enterprises to quickly adapt to scenarios such as aviation, port terminals, and intelligent manufacturing by inputting industry-specific parameters (such as job position size, vehicle type, task unit attributes), greatly reducing the threshold of intelligent upgrading and realizing low-cost deployment. At the same time, the embodiments of the present invention combine a two-stage collaborative optimization framework and distributed robust optimization (DRO) technology. Firstly, in dynamic job position allocation, the column constraint generation (CCG) algorithm and variable neighborhood search (VNS) technology are used to accurately avoid risks of size mismatch and time overlap, improve the utilization rate of parking positions, and reduce job congestion caused by delays of dynamic task units. Secondly, through deep reinforcement learning (DQN+PPO), the path planning and task priority of automated vehicles are optimized in real time. An uncertainty set is constructed based on DRO to cope with disturbances such as flight delays and fluctuations in ship arrivals, reducing the task delay rate. At the same time, the closed-loop feedback mechanism realizes the two-way coordination of dynamic job position layout and vehicle scheduling - optimizing the distance of parking positions directly improves the AGV transportation efficiency, while the vehicle state (such as working state) triggers the dynamic adjustment of job positions in reverse, forming a global synergistic effect and improving the overall collaboration efficiency. And it is proved by mathematical induction that the column constraint generation (CCG) algorithm of the main problem converges to the optimal solution when the number of iterations N≤50, and the volatility of the Q-value function of the sub-problem is lower than 5% after the training cycle T = 10 4 The embodiments of the present invention can reduce the empty driving rate of AGVs in air cargo, reduce the demurrage cost of ships in port scenarios, and reduce the stagnation risk in the production line of manufacturing scenarios. The embodiments of the present invention use a highly robust and low-cost intelligent scheduling solution to break through the adaptability bottleneck of traditional methods in dynamic environments and reach the leading industry standards.

[0142] Please refer to Figure 6 , the embodiments of the present application also provide a logistics scenario resource scheduling system, which can implement the above-mentioned logistics scenario resource scheduling method. The system includes:

[0143] The first module 210 is used to obtain preset job information.

[0144] The second module 220 is used to predict the expected job position allocation result according to the preset job information through the column constraint generation algorithm.

[0145] The third module 230 is used to predict the expected vehicle scheduling result according to the expected job position allocation result and the preset loading and unloading information through a preset deep reinforcement learning model. Among them, the preset deep reinforcement learning model is constructed by the deep Q-network algorithm and the probabilistic policy optimization algorithm.

[0146] The fourth module 240 is configured to determine a target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result when it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost condition.

[0147] It can be understood that the content in the above method embodiments is applicable to this system embodiment. The functions specifically implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0148] This application embodiment also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned logistics scenario resource scheduling method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0149] It can be understood that the content in the above method embodiments is applicable to this device embodiment. The functions specifically implemented in this device embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0150] Please refer to Figure 7 , Figure 7 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0151] A processor 310, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in this application embodiment;

[0152] A memory 320, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 320 can store an operating system and other application programs. When implementing the technical solutions provided in this specification embodiment through software or firmware, the relevant program codes are stored in the memory 320 and are called by the processor 310 to execute the logistics scenario resource scheduling method of this application embodiment;

[0153] An input / output interface 330, which is used to implement information input and output;

[0154] A communication interface 340 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0155] A bus 350 for transmitting information between various components of the device (such as a processor 310, a memory 320, an input / output interface 330, and a communication interface 340);

[0156] Among them, the processor 310, the memory 320, the input / output interface 330, and the communication interface 340 achieve communication connections with each other inside the device through the bus 350.

[0157] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned logistics scenario resource scheduling method.

[0158] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0159] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0160] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0161] Those skilled in the art can understand that the technical solutions shown in the figure do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than shown in the figure, or combine certain steps, or different steps.

[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0163] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0164] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0165] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0166] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0167] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0168] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0169] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0170] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A method for scheduling logistics scenario resources, characterized in that The method includes the following steps: Obtain preset job information; Predict an expected job position allocation result through a column constraint generation algorithm according to the preset job information; Predict an expected vehicle scheduling result through a preset deep reinforcement learning model according to the expected job position allocation result and preset loading and unloading information; wherein, the preset deep reinforcement learning model is constructed by a deep Q-network algorithm and a probabilistic policy optimization algorithm; When it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost condition, determine a target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result.

2. The method according to claim 1, characterized in that, The preset job information includes dynamic task unit information and job position information; The predicting the expected job position allocation result through a column constraint generation algorithm according to the preset job information includes: Construct a preset mixed integer programming model according to the dynamic task unit information and the job position information, and calculate initial job position allocation data through the preset mixed integer programming model; Identify a preset conflict scenario according to the initial job position allocation data; Determine key conflict constraints according to the conflict constraint conditions corresponding to the preset conflict scenario; Generate candidate columns according to the key conflict constraints, and add the candidate columns to the preset mixed integer programming model for optimization to obtain a target mixed integer programming model; Predict the expected job position allocation result through the target mixed integer programming model.

3. The method according to claim 2, wherein The determining the key conflict constraints according to the conflict constraint conditions corresponding to the preset conflict scenario includes: Screen the conflict constraint conditions through a variable neighborhood search algorithm to obtain the key conflict constraints.

4. The method according to claim 1, wherein Before performing the predicting the expected vehicle scheduling result through a preset deep reinforcement learning model according to the expected job position allocation result and preset loading and unloading information, the method further includes: Input preset state data into an initial deep Q-network model to select a preset scheduling strategy; Perform policy optimization on the preset scheduling strategy through the probabilistic policy optimization algorithm to update the initial deep Q-network model to obtain the preset deep reinforcement learning model.

5. The method according to claim 4, wherein The performing policy optimization on the preset scheduling strategy through the probabilistic policy optimization algorithm to update the initial deep Q-network model to obtain the preset deep reinforcement learning model includes: Calculate a preset advantage value through a distributed robust optimization algorithm according to the preset scheduling strategy, and perform parameter optimization according to the preset advantage value to obtain an optimized strategy; Generate corresponding policy actions according to the optimized strategy, and perform environment interaction according to the policy actions to calculate policy reward data; When it is determined that the policy reward data is greater than the historical reward data, perform network parameter update on the initial deep Q-network model according to the optimized strategy to obtain the preset deep reinforcement learning model.

6. The method according to claim 5, characterized in that The calculating a preset advantage value through a distributed robust optimization algorithm according to the preset scheduling strategy, and performing parameter optimization according to the preset advantage value to obtain an optimized strategy includes: Construct an uncertainty set; Perform a worst-case dominance estimation on the preset scheduling strategy through the uncertainty set to obtain a robust dominance value, and optimize the policy parameters according to the robust dominance value to obtain an optimized strategy.

7. The method according to claim 1, characterized in that When it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost condition, determining the target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result includes: Perform cost calculation according to the expected vehicle scheduling result and the expected job position allocation result to obtain preset cost data; wherein, the preset cost data includes delay cost data and operation cost data; When it is determined that the preset cost data is less than the preset cost threshold, use the expected vehicle scheduling result and the expected job position allocation result as the target scheduling allocation result.

8. A logistics scenario resource scheduling system, characterized in that, The system includes: A first module for obtaining preset job information; A second module for predicting an expected job position allocation result according to the preset job information through a column constraint generation algorithm; A third module for predicting an expected vehicle scheduling result according to the expected job position allocation result and preset loading and unloading information through a preset deep reinforcement learning model; wherein, the preset deep reinforcement learning model is constructed by a deep Q-network algorithm and a probabilistic policy optimization algorithm; A fourth module for determining the target scheduling allocation result according to the expected vehicle scheduling result and the expected job position allocation result when it is determined that the expected vehicle scheduling result and the expected job position allocation result meet the preset cost condition.

9. An electronic device, characterized in that, Includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program implements the method according to any one of claims 1 to 7 when executed by a processor.