Motorcade multi-target dynamic scheduling method and system based on edge calculation
By building a four-dimensional dynamic programming model and dynamically adjusting target weights on the edge computing platform, the limitations of the existing logistics scheduling system in multi-objective collaborative optimization and resource utilization efficiency are solved, and efficient fleet multi-objective dynamic scheduling is achieved.
Patent Information
- Application Number
- CN202510535867.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing logistics scheduling systems have limitations in multi-objective collaborative optimization, real-time decision-making capabilities and resource utilization efficiency, and are unable to effectively deal with task splitting timeliness, return-haul empty utilization, load balancing and special scenario constraints.
The fleet multi-objective dynamic scheduling method is adopted based on edge computing, and data is collected in real time through on-board sensors, vehicle state vectors are constructed by edge devices, and technologies such as natural language processing and deep reinforcement learning are used to build a four-dimensional dynamic programming model, dynamically adjust the target weights, and achieve multi-objective collaborative optimization.
It significantly improves the overall efficiency of the system, ensures the balanced satisfaction of multi-target needs in complex scenarios, improves vehicle resource utilization and scheduling timeliness, and realizes efficient allocation of transportation resources.
Smart Images

Figure CN120069722A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation and logistics scheduling, and specifically to a multi-objective dynamic scheduling method and system for vehicle fleets based on edge computing. Background Art
[0002] With the rapid development of Internet of Things and artificial intelligence technologies, the logistics and transportation industries are undergoing a profound intelligent transformation. However, the limitations of traditional scheduling models in aspects such as multi-objective collaborative optimization, real-time decision-making ability, and resource utilization efficiency have restricted the improvement of the overall efficiency of the industries.
[0003] Existing logistics scheduling systems mostly focus on a single objective (such as transportation cost or timeliness), lacking comprehensive consideration of the timeliness of task splitting, the utilization rate of empty vehicles on the return journey, load balance, and special scenario constraints. For example, the traditional radiation-based distribution model only diverges paths with the distribution center as the core, resulting in an empty vehicle rate of up to 40% and insufficient utilization of return resources, making it difficult to achieve multi-objective collaborative optimization.
[0004] Cloud-based centralized scheduling relies on data upload and instruction issuance, with decision-making delays generally exceeding 10 seconds, and is unable to handle abnormal scenarios such as sudden road conditions and vehicle failures; traditional scheduling systems lack the ability to continuously optimize driven by data, and model parameter updates rely on manual experience, making them unable to adapt to dynamic environmental changes; special scenarios such as dangerous goods transportation and emergency tasks pose higher requirements on scheduling systems, but existing models lack targeted optimization.
[0005] In view of the above problems, the present invention proposes a multi-objective dynamic scheduling method and system for vehicle fleets based on edge computing to solve the above-mentioned problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a multi-objective dynamic scheduling method and system for vehicle fleets based on edge computing to solve the problems raised in the prior art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions: A multi-objective dynamic scheduling method for vehicle fleets based on edge computing, comprising the following steps: S1. On-vehicle sensors collect data on vehicle position, load, remaining mileage, and tank cleanliness in real time, and edge devices synchronously receive cloud task requirements and road condition information to construct a vehicle state vector; S2. Based on the cloud task requirement information included in the vehicle state vector constructed in S1, use natural language processing technology to parse task instructions, extract cargo type, transportation timeliness, and destination characteristics, and classify tasks into three categories: emergency transportation, regular transportation, and empty vehicle tasks on the return journey; S3. Based on the above data, construct a four-dimensional dynamic programming model including task splitting, utilization of empty return trips, load balancing, and tank cleaning costs; based on the vehicle state vector constructed in S1 and the task categories divided in S2, deploy a four-dimensional state transition equation including task splitting T, utilization of empty return trips R, load balancing L, and tank cleaning costs C in the edge computing chip, and convert the multi-objective into a comprehensive cost function with dynamic weights; S4. According to the regional empty vehicle rate reflected by the road condition information in S1, combined with the task urgency and cargo type characteristics extracted in S2, the fuzzy neural network dynamically adjusts the four-dimensional objective weights in the comprehensive cost function of S3. Combining with the deep reinforcement learning DRL algorithm, construct a policy library based on the four-dimensional state transition equation of S3, and update the optimal scheduling plan regularly; S5. When abnormal load sensors or road congestion feedback in the vehicle data collected in S1 occur, trigger a rule-based emergency mechanism. Based on the four-dimensional state transition equation of S3 and the policy library of S4, reallocate tasks through the backtracking search algorithm and update the four-dimensional model parameters; S6. After the task collection is completed, continuously record the task completion time, energy consumption and other data in S1, online evaluate the achievement degree of the four-dimensional objectives in the comprehensive cost function of S3, and use the Bayesian optimization algorithm to update the four-dimensional model parameters of S3 and the four-dimensional objective weights of S4 to form a closed-loop iterative system of data collection - S3 / S4 optimization - execution feedback.
[0008] The above-mentioned S1 further includes the following steps: S1.1: On-vehicle sensors collect vehicle position, load, remaining mileage, and tank cleanliness data; S1.2: The edge device synchronously receives the cloud task requirements and real-time road condition information through the 5G network; S1.3: Use the spatio-temporal alignment algorithm to eliminate data noise and construct a vehicle state vector; specifically: first, generate derivative features such as vehicle speed and load change through mathematical modeling of the vehicle position, load, remaining mileage, tank cleanliness data, cloud tasks, and road condition data collected by on-vehicle sensors in steps S1.1 and S1.2, and establish the correlation features between cargo type and cleanliness, and road condition and energy consumption; then perform standardization processing on the multi-dimensional data, and use the dimensionality reduction technology to construct a structured vector space covering vehicle state, task characteristics, and environmental constraints for the subsequent four-dimensional dynamic programming model, and use it as input data to support the optimization calculation of vehicle scheduling decisions.
[0009] The above-mentioned step S2 further includes the following steps: S2.1: Parse the cloud task requirement information contained in the vehicle status vector using natural language processing techniques; specifically: use a pre-trained BERT model to parse the task instruction text, identify the cargo type, transportation time limit, and destination coordinates, then encode the cargo type as a unique ID, store the transportation time limit as a value in minutes, and convert the destination coordinates to floating-point latitude and longitude data; S2.2: Generate derivative features related to the task based on the parsed task instructions. The derivative features include the distance between the task origin and destination, time window constraints, and the volume-weight ratio of the cargo; specifically: calculate the distance using the Euclidean distance formula, determine the penalty coefficients for early arrival and late arrival according to the required arrival time of the task, and obtain the volume-weight ratio by dividing the cargo volume by the weight; the time window constraints are specifically: set the required arrival time of the task as t req , the early arrival time as t early , and the late arrival time as t late , then the early arrival penalty coefficient is , and the late arrival penalty coefficient is ; S2.3: Use the K-means clustering algorithm to divide the derivative features such as transportation time limit, distance between origin and destination, and volume-weight ratio into different categories; the classification criteria are: transportation with a time limit of less than 4 hours and a distance of no more than 100 km is emergency transportation; transportation with a time limit of 4 - 24 hours and a distance of 100 - 500 km is regular transportation; a return empty vehicle task is one where the path overlap between the task destination and the current vehicle position is not less than 60%.
[0010] The step S3 further includes the following steps: S3.1: Establish a four-dimensional dynamic programming model for the state containing four-dimensional objectives to quantify the current state of the vehicle and task constraints; specifically: first define a four-dimensional state vector S = [s T , s R , s L , s C , and the specific meanings of each dimension are as follows: s T represents the possibility of task splitting, calculated by the ratio of the remaining load capacity of the vehicle to the cargo volume of the task. The calculation formula is as follows: ; s R represents the return empty vehicle rate, statistically calculated based on the regional empty vehicle distribution between the current vehicle position and the task destination. The calculation formula is as follows: ; s L represents the load balance degree, quantified by the sum of the squares of the load differences between adjacent vehicles. The calculation formula is as follows: ; s C Represents the tank cleaning demand, which is obtained by matching the tank cleanliness sensor data with the cargo type pollution level. When the pollution level ≥ the threshold, tank cleaning is triggered; S3.2: Define executable scheduling actions; specifically: define a four-dimensional action vector A = [a T , a R , a L ,a C ], the operation rules of each dimension are as follows: a T Represents the task split ratio, which is determined according to the remaining load and cargo integrity constraints, and the split ratio range is {0, 70%}; a R Represents the return task matching, and prioritizes the return tasks whose path overlap with the current task is ≥ 60%. The path overlap calculation formula is as follows: ; a L Represents the load transfer amount, which is calculated based on the load difference between adjacent vehicles and the road weight limit constraint. The single transfer amount does not exceed the preset standard; a C Represents the tank cleaning decision. When the cargo type is switched or the tank contamination level is ≥80%, the tank cleaning operation is forcibly triggered; S3.3: Convert the four-dimensional target into a unified quantitative indicator; specifically: construct a comprehensive cost function with dynamic weights, the formula is as follows: ; Among them, C T represents the task splitting cost, which is positively correlated with the cargo integrity loss rate. The formula is as follows: ; C R Represents the cost of empty driving, which is calculated by the return empty driving mileage × unit fuel consumption. The unit fuel consumption is determined based on the vehicle type and road conditions. C L represents the load imbalance cost, which is quantified by the sum of the squares of the load differences between adjacent vehicles, and the formula is as follows: ; C C Represents the tank cleaning cost, which is positively correlated with the tank contamination level and cleaning time. The formula is as follows: ; S3.4: Describe the changing rules of the vehicle state after executing the action, and provide a mathematical basis for dynamic programming; specifically: establish a four-dimensional state transfer matrix P(S{t+1} | St, At), and the transfer probability of each dimension is determined by historical data training; Among them, P(s{T,t+1} | s{T,t}, a{T,t}) is the probability of change of the remaining load after task splitting, which is based on the vehicle load distribution fitting; P(s{R,t+1} | s{R,t}, a{R,t}) is the probability of change of the empty vehicle rate after the return task matching, which is based on the historical data statistics of regional traffic flow; P(s{L,t+1} | s{L,t}, a{L,t}) is the probability of change of balance after load transfer, which is modeled through load distribution experimental data; P(s{C,t+1} | s{C,t}, a{C,t}) is the probability of change of cleanliness after tank washing decision, which is based on the historical data of tank cleanliness sensor.
[0011] The step S4 further includes the following contents: S4.1: Based on the real-time features of the above-mentioned road condition information, the regional empty vehicle rate calculated, the task urgency mapped by task classification, the cargo type converted by the parsed cargo type ID, the road grade extracted by the road condition, and the remaining load ratio calculated by the vehicle state vector, a three-tier fuzzy neural network is constructed: the input layer contains 12 feature nodes, the hidden layer adopts a two-layer ReLU activation structure with 32 neurons each, and the output layer generates a four-dimensional target weight vector ω=[ω T ,ω R , ω L , ω C ], where ω T The weight of the task splitting target; ω R is the weight of the empty return vehicle utilization target; ω L is the weight of the load balancing target; ω C is the weight of the tank washing cost target; the four-dimensional target weight ω i satisfy Normalization constraint; The network implements weight constraints through preset rules: mandatory allocation of ω for urgent transportation tasks T ≥0.4, empty return missionω R ≥0.3, dangerous goods transportation taskω C ≥0.2, ensuring the priority of multi-objective collaborative optimization in different scenarios; S4.2: Construct a deep reinforcement learning policy library. Specifically, based on the vehicle state vector, task classification results, and S1 environment parameters, construct a three-dimensional state space, and design a discrete action space containing four-dimensional actions, including task splitting ratio, return trip task matching, load transfer volume, and tank cleaning decision; adopt a reward mechanism coupled with the S3 comprehensive cost function, define a reward function, and the formula is as follows: ; where λ i is the reward weight for the four-dimensional target, which is determined after normalizing the four-dimensional target weight ω i described in S4.1, C i is the task splitting cost C T calculated by S3, the empty return trip cost C R the load imbalance cost C L and the tank cleaning cost C C ; where λ T is the reward weight for the task splitting target, λ R is the reward weight for the empty return trip utilization target, λ L is the reward weight for the load balancing target, λ C is the reward weight for the tank cleaning cost target; this policy library is online trained through the PPO algorithm in an environment containing millions of state-action combinations formed by the vehicle state vector in S1, the task features in S2, the four-dimensional state space and action space in S3 to generate an optimal scheduling policy sequence; Adopt an online learning mechanism. Based on the four-dimensional state transition equation, use the PPO algorithm to train the policy network to generate an optimal action sequence; at the same time, use an experience replay buffer to store millions of real-time state-action-reward triple data from the vehicle data collected in real time in step S1, the task instructions parsed in S2, and the cost and state transition results generated in S3, and improve the sample utilization rate through the prioritized experience replay strategy.
[0012] The step S5 further includes the following steps: S5.1: Real-time monitor vehicle state and road condition data, identify abnormal events and trigger emergency responses; specifically, detect abnormal load sensor data through sliding window filtering and the 3σ criterion, and combine the vehicle speed continuously being lower than the threshold and the road congestion index to double-verify the road congestion state to ensure the accurate identification and triggering of abnormal events; S5.2: Based on the detected abnormal type, activate the corresponding emergency strategy; reorder the tasks by constructing a priority queue, and call the action sequence that has successfully coped with similar abnormalities in the S4 policy library in the past, and combine the four-dimensional state transition equation to pre-act the state changes of alternative strategies to provide a feasible solution for subsequent task allocation; S5.3: Traverse the four-dimensional state space of schedulable vehicles using a backtracking search algorithm, preferentially select high-reward actions in the policy library, optimize the search path through pruning conditions, and the task splitting follows the rules of adjacent allocation for emergency subtasks and delayed execution for ordinary subtasks; S5.4: According to the execution data after abnormal events, refit the load state transition probability of abnormal vehicles, and update the empty vehicle rate transition probability based on historical congestion data; adjust the four-dimensional target weights in step S4 through the online gradient descent algorithm, gradually increase the task splitting priority according to the preset standard, and reduce the load balancing weight to optimize the adaptability of the model to abnormal scenarios; S5.5: Adopt a quantization verification system, calculate the on-time rate of abnormal tasks based on the vehicle position data in step S1 and the transportation timeliness in step S2, and set the target value to not be lower than 95%; at the same time, calculate the load change rate within the adjacent preset time period through the load sensor data in step S1, and require that the fluctuation control target reaches within 5%; when the above indicators cannot be met for three consecutive abnormal event handling, the system will trigger the manual intervention threshold, automatically generate an abnormal report and push it to the dispatcher terminal; at the same time, the system stores the current abnormal handling strategy in the experience replay buffer in real time, updates the abnormal response action sequence in the policy library in step S4 and increases its priority coefficient, and at the same time refits the abnormal probability matrix in the four-dimensional state transition equation based on the current abnormal data, and integrates the verification data into the model parameter update process in steps S3 / S4 to form a closed-loop management mechanism covering detection - response - verification and optimization.
[0013] The said step S6 further includes the following steps: S6.1: Based on the task completion time and energy consumption data continuously recorded in step S1, construct a four-dimensional target achievement evaluation system, including four indicators: task splitting timeliness, empty vehicle utilization rate on the return journey, load balancing deviation rate, and tank washing energy consumption ratio; S6.2: Use the Bayesian optimization algorithm to synchronously update the probability parameters of the four-dimensional state transition equation and the four-dimensional target weights, and the optimization objective function is defined as the weighted sum of the four-dimensional target achievements; the formula of the optimization objective function is as follows: ; where, i is the index of the four-dimensional target. When i = 1, γ 1 is the priority coefficient of the task splitting target, and the achievement degree 1 is determined by the task splitting timeliness index; when i = 2, γ 2 is the priority coefficient of the empty vehicle utilization target on the return journey, and the achievement degree 2 is determined by the empty vehicle utilization rate index on the return journey; when i = 3, γ 3 is the priority coefficient of the load balancing target, and the achievement degree 3 is determined by the load balancing deviation rate index; when i = 4, γ4 is the priority coefficient of the washing tank cost target, and the achievement degree 4 is determined by the washing tank energy consumption ratio index, γ i is the priority coefficient of the four-dimensional target, and is determined by mapping through a preset rule from the four-dimensional target weight ω in step S4 i ; S6.3: Push the optimized parameters to the modules in steps S3 / S4 in real time to form a complete closed loop including data acquisition, target evaluation, parameter optimization, and policy update, ensuring that the system completes model iteration in each task cycle.
[0014] A multi-objective dynamic scheduling system for a fleet based on edge computing, including a data acquisition and preprocessing module, a task parsing and classification module, a multi-objective optimization model module, a policy generation and optimization module, an emergency management module, and a closed-loop iteration module; The data acquisition and preprocessing module is responsible for collecting vehicle status, cloud tasks, and road condition data in real time, eliminating noise through a spatio-temporal alignment algorithm and generating derivative features, and constructing a structured vector space to provide standardized input for subsequent models; The task parsing and classification module uses natural language processing technology to parse task instructions, extract features such as cargo type, timeliness, and destination coordinates, generate derivative indicators such as distance, time window constraints, and cargo volume-to-weight ratio, and divides tasks into three categories: urgent, regular, and empty return vehicles through a clustering algorithm; The multi-objective optimization model module is used to construct a four-dimensional state transition equation, define the action space and state change rules, and transform the multi-objective into a comprehensive cost function with dynamic weights; The policy generation and optimization module dynamically adjusts the four-dimensional target weights through a fuzzy neural network, constructs a policy library in combination with a deep reinforcement learning algorithm, and regularly updates the optimal scheduling plan; The emergency management module monitors vehicle anomalies and road condition mutations in real time, reallocates tasks based on the four-dimensional model and policy library after triggering the emergency mechanism, updates model parameters and verifies the effectiveness of the plan, and forms an abnormal handling closed loop; The closed-loop iteration module evaluates the achievement degree of the four-dimensional target based on task completion time and energy consumption data, and uses the Bayesian optimization algorithm to synchronously update the four-dimensional model parameters and the four-dimensional target weights to achieve continuous iteration of data acquisition - optimization - execution.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. Multi-objective collaborative optimization: The present invention constructs a four-dimensional dynamic programming model including task splitting, utilization of empty return vehicles, load balancing, and tank washing cost, and combines a fuzzy neural network to dynamically adjust the weights of each objective to achieve the quantification and unification of multiple objectives. This mechanism breaks through the limitation of traditional scheduling models that only focus on a single objective. Through the collaborative optimization of task splitting timeliness, utilization rate of empty return vehicles, load balancing deviation rate, and tank washing energy consumption ratio, it significantly improves the overall efficiency of the system and ensures the balanced satisfaction of multi-objective requirements in complex scenarios.
[0016] 2. Improve the utilization rate of vehicle resources: Based on the matching rule of the quantification index of the empty return vehicle rate and the path coincidence degree, combined with the online optimization of the deep reinforcement learning strategy library, it dynamically matches the return tasks and preferentially assigns actions with a high path coincidence degree. This mechanism effectively reduces the empty driving mileage of vehicles, improves the utilization rate of empty return vehicle resources, optimizes the vehicle operation path, avoids local overload or empty load phenomena, and realizes the efficient allocation of transportation resources.
[0017] 3. Enhance scheduling timeliness: Deploy a four-dimensional state transition equation and a strategy library on the edge computing chip, and use the 5G network to process vehicle status and road condition data in real time to achieve millisecond-level strategy updates and exception responses. This local decision-making architecture breaks through the latency bottleneck of cloud scheduling, ensures fast response capabilities in complex road conditions, improves the order on-time rate and task execution accuracy, and continuously improves the system's adaptability in dynamic environments through an online parameter optimization mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the method flow chart of the multi-objective dynamic scheduling method for a fleet based on edge computing of the present invention; Figure 2 is the four-dimensional dynamic modeling flow chart of the multi-objective dynamic scheduling method for a fleet based on edge computing of the present invention; Figure 3 is the strategy generation and optimization flow chart of the multi-objective dynamic scheduling method for a fleet based on edge computing of the present invention; Figure 4 is the emergency handling flow chart of the multi-objective dynamic scheduling method for a fleet based on edge computing of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] Embodiment: As Figures 1-4 shown, the present invention provides a technical solution A multi-objective dynamic scheduling method for a vehicle fleet based on edge computing, such as Figure 1 shown, includes the following steps: S1. The on-vehicle sensors collect data on vehicle position, load, remaining mileage, and tank cleanliness in real time. The edge device synchronously receives the cloud task requirements and road condition information, and constructs a vehicle state vector; S2. Based on the cloud task requirement information contained in the vehicle state vector constructed in S1, use natural language processing technology to parse the task instructions, extract the cargo type, transportation timeliness, and destination characteristics, and divide the tasks into three categories: emergency transportation, regular transportation, and empty vehicle return tasks; S3. Based on the above data, construct a four-dimensional dynamic programming model including task splitting, utilization of empty vehicles on the return journey, load balancing, and tank cleaning cost; based on the vehicle state vector constructed in S1 and the task categories divided in S2, deploy a four-dimensional state transition equation including task splitting T, utilization of empty vehicles on the return journey R, load balancing L, and tank cleaning cost C in the edge computing chip, and transform the multi-objective into a comprehensive cost function with dynamic weights; S4. According to the regional empty vehicle rate reflected by the road condition information in S1, combined with the task urgency and cargo type characteristics extracted in S2, the fuzzy neural network dynamically adjusts the four-dimensional target weights in the comprehensive cost function of S3, and combines with the deep reinforcement learning DRL algorithm to construct a policy library based on the four-dimensional state transition equation of S3, and regularly update the optimal scheduling plan; S5. When an abnormality occurs in the load sensor in the vehicle data collected in S1 or the road condition information feedback indicates road congestion, trigger a rule-based emergency mechanism, and based on the four-dimensional state transition equation of S3 and the policy library of S4, re-allocate tasks through the backtracking search algorithm and update the four-dimensional model parameters; S6. After the collection task is completed, continuously record the task completion time, energy consumption and other data in S1, online evaluate the achievement degree of the four-dimensional objectives in the comprehensive cost function of S3, and use the Bayesian optimization algorithm to update the four-dimensional model parameters of S3 and the four-dimensional target weights of S4 to form a closed-loop iterative system of data collection - S3 / S4 optimization - execution feedback.
[0021] The above S1 further includes the following steps: S1.1: The on-vehicle sensors collect data on vehicle position, load, remaining mileage, and tank cleanliness; S1.2: The edge device synchronously receives the cloud task requirements and real-time road condition information through the 5G network; S1.3: Eliminate data noise using the spatio-temporal alignment algorithm and construct a vehicle state vector. Specifically: First, generate derivative features such as vehicle speed and load changes through mathematical modeling of the vehicle position, load, remaining mileage, tank cleanliness data, cloud tasks, and road condition data collected by on-vehicle sensors in steps S1.1 and S1.2, and establish the correlation features between cargo type and cleanliness, and road conditions and energy consumption. Subsequently, perform standardization processing on the multi-dimensional data, and use dimensionality reduction technology to construct a structured vector space covering vehicle state, task features, and environmental constraints for the subsequent four-dimensional dynamic programming model, as input data to support the optimization calculation of vehicle scheduling decisions.
[0022] Step S2, as Figure 2 shown, further includes the following steps: S2.1: Parse the cloud task requirement information contained in the vehicle state vector using natural language processing technology. Specifically: Use the pre-trained BERT model to parse the task instruction text, identify the cargo type, transportation time limit, and destination coordinates, then encode the cargo type as a unique ID, store the transportation time limit as a minute-level value, and convert the destination coordinates into latitude and longitude floating-point data. S2.2: Generate derivative features related to the task according to the parsed task instructions. The derivative features are divided into the distance between the task origin and destination, time window constraint, and cargo volume-weight ratio. Specifically: Calculate the distance through the Euclidean distance formula, determine the penalty coefficients for early arrival and late arrival according to the required arrival time of the task, and obtain the volume-weight ratio by dividing the cargo volume by the weight. The time window constraint is specifically: Set the required arrival time of the task as t req , the early arrival time as t early , the late arrival time as t late , then the early arrival penalty coefficient is , the late arrival penalty coefficient is ; S2.3: Use the K - means clustering algorithm to divide the derivative features such as transportation time limit, distance between origin and destination, and volume-weight ratio into different categories. The classification criteria are: Emergency transportation if the transportation time limit is less than 4 hours and the distance does not exceed 100 km; Regular transportation if the time limit is between 4 - 24 hours and the distance is between 100 - 500 km; Return empty vehicle task if the path overlap degree between the task destination and the current vehicle position is not less than 60%.
[0023] Step S3 further includes the following steps: S3.1: Establish a four-dimensional dynamic programming model including the state of four-dimensional objectives to quantify the current vehicle state and task constraints. Specifically: First, define the four-dimensional state vector S = [s T , s R, s L , sC , and the specific meanings of each dimension are as follows: s T represents the task splitting possibility, which is calculated by the ratio of the remaining load capacity of the vehicle to the volume of the task goods. The calculation formula is as follows: ; s R represents the empty vehicle rate on the return journey, which is statistically calculated based on the empty vehicle distribution in the area between the current vehicle location and the task destination. The calculation formula is as follows: ; s L represents the load balance degree, which is quantified by the sum of the squares of the load differences between adjacent vehicles. The calculation formula is as follows: ; s C represents the tank cleaning requirement, which is obtained by matching the data of the tank cleanliness sensor with the pollution level of the goods type. When the pollution level ≥ the threshold, the tank cleaning is triggered; S3.2: Define executable scheduling actions; specifically: Define the four-dimensional action vector A = [a T , a R , a L ,a C , and the operation rules of each dimension are as follows: a T represents the task splitting ratio, which is determined according to the remaining load capacity and the cargo integrity constraint. The splitting ratio range is {0, 70%}; a R represents the matching of the return journey task. Preferentially match the return journey task with a path overlap degree ≥ 60% with the current task. The calculation formula for the path overlap degree is as follows: ; a L represents the load transfer amount, which is calculated based on the load difference between adjacent vehicles and the road weight limit constraint. The single transfer amount does not exceed the preset standard; a C represents the tank cleaning decision. When the goods type is switched or the tank pollution level ≥ 80%, the tank cleaning operation is forced to be triggered; S3.3: Convert the four-dimensional goal into a unified quantitative index; specifically: Construct a comprehensive cost function with dynamic weights. The formula is as follows: ; Among them, C T represents the task splitting cost, which is positively correlated with the cargo integrity loss rate. The formula is as follows: ; C Rrepresents the empty running cost, which is calculated by multiplying the empty running mileage of the return trip by the unit fuel consumption. The unit fuel consumption is determined based on the vehicle type and road conditions; C L represents the load imbalance cost, which is quantified by the sum of squares of the load differences between adjacent vehicles. The formula is as follows: ; C C represents the tank cleaning cost, which is positively correlated with the tank pollution level and the cleaning time. The formula is as follows: ; S3.4: Describe the change law of the vehicle state after performing the action, providing a mathematical basis for dynamic programming; specifically: establish a four-dimensional state transition matrix P(S{t+1} | St, At), and the transition probabilities of each dimension are determined through training with historical data; Among them, P(s{T,t+1} | s{T,t}, a{T,t}) is the change probability of the remaining load weight after task splitting, which is fitted based on the vehicle load distribution; P(s{R,t+1} | s{R,t}, a{R,t}) is the change probability of the empty vehicle rate after the return trip task matching, which is statistically based on the historical data of regional vehicle flow; P(s{L,t+1} | s{L,t}, a{L,t}) is the change probability of the balance degree after load transfer, which is modeled through load distribution experimental data; P(s{C,t+1} | s{C,t}, a{C,t}) is the change probability of the cleanliness after the tank cleaning decision, which is based on the historical data of the tank cleanliness sensor.
[0024] The step S4 further includes the following content: S4.1: Based on the real-time features of these dimensions such as the regional empty vehicle rate calculated from the above road condition information, the task urgency mapped by task classification, the cargo type converted from the parsed cargo type ID, the road grade extracted from the road conditions, and the remaining load weight ratio calculated from the vehicle state vector, construct a three-layer fuzzy neural network: the input layer contains 12 feature nodes, the hidden layer adopts a ReLU activation structure with two layers of 32 neurons each, and the output layer generates a four-dimensional target weight vector ω = [ω T , ω R , ω L , ω C through the Softmax function. Among them, ω T is the weight of the task splitting target; ω R is the weight of the empty vehicle utilization target for the return trip; ω L is the weight of the load balance target; ω C is the weight of the tank cleaning cost target; the four-dimensional target ω i satisfies the normalization constraint; this network realizes weight constraint through preset rules: forcefully assign ω for urgent transportation tasksT ≥ 0.4, empty vehicle return task ω R ≥ 0.3, dangerous goods transportation task ω C ≥ 0.2, to ensure the multi - objective collaborative optimization priority under different scenarios; S4.2: Construct a deep reinforcement learning policy library. Specifically: Based on the vehicle state vector, task classification results, and S1 environment parameters, construct a three - dimensional state space, and design a discrete action space containing four - dimensional actions, including task splitting ratio, return task matching, load transfer volume, and tank cleaning decision; Adopt a reward mechanism coupled with the S3 comprehensive cost function, define the reward function, and the formula is as follows: ; where λ i is the reward weight of the four - dimensional objective, which is determined by normalizing the four - dimensional objective weights ω i described in S4.1, C i is the task splitting cost C T calculated by S3, the empty vehicle running - back cost C R , the load imbalance cost C L and the tank cleaning cost C C ; among them, λ T is the reward weight of the task splitting objective, λ R is the reward weight of the empty vehicle utilization objective for return trips, λ L is the reward weight of the load balancing objective, λ C is the reward weight of the tank cleaning cost objective; This policy library is online - trained through the PPO algorithm in an environment containing millions of state - action combinations formed by the vehicle state vector in S1, the task features in S2, the four - dimensional state space and action space in S3, to generate an optimal scheduling policy sequence; Adopt an online learning mechanism. Based on the four - dimensional state transition equation, use the PPO algorithm to train the policy network to generate an optimal action sequence; At the same time, use an experience replay buffer to store millions of real - time state - action - reward triple data from the vehicle data collected in real - time in step S1, the task instructions parsed in S2, and the cost and state transition results generated in S3, and improve the sample utilization rate through the prioritized experience replay strategy.
[0025] The said step S5, as Figure 3 , Figure 4 shown, further includes the following steps: S5.1: Real - time monitor vehicle states and road condition data, identify abnormal events and trigger emergency responses; Specifically, detect abnormal load sensor data through sliding window filtering and the 3σ criterion, and combine the vehicle speed continuously being lower than the threshold and the road congestion index to double - verify the road congestion state, ensuring the accurate identification and triggering of abnormal events; S5.2: Based on the detected abnormal type, activate the corresponding emergency strategy; reorder the tasks by constructing a priority queue, and call the action sequence that successfully coped with similar abnormalities in the historical strategy library in step S4. Combine the four-dimensional state transition equation to preview the state changes of alternative strategies, and provide a feasible solution for subsequent task allocation; S5.3: Use the backtracking search algorithm to traverse the four-dimensional state space of schedulable vehicles, preferentially select high-reward actions in the strategy library, optimize the search path through pruning conditions, and the task splitting follows the rules of adjacent allocation of urgent subtasks and delayed execution of ordinary subtasks; S5.4: According to the execution data after the abnormal event, refit the load state transition probability of the abnormal vehicle, and update the empty vehicle rate transition probability based on historical congestion data; adjust the four-dimensional target weights in step S4 through the online gradient descent algorithm, gradually increase the task splitting priority according to the preset standard, and reduce the load balancing weight to optimize the adaptability of the model to abnormal scenarios; S5.5: Adopt a quantitative verification system, calculate the on-time rate of abnormal tasks based on the vehicle position data in step S1 and the transportation timeliness in step S2, and set the target value to be not less than 95%; at the same time, calculate the load change rate within adjacent preset time periods through the load sensor data in step S1, and require a fluctuation control target of within 5%; when the above indicators cannot be met for three consecutive abnormal event handling, the system will trigger the manual intervention threshold, automatically generate an abnormal report and push it to the dispatcher terminal; at the same time, the system stores the current abnormal handling strategy in the experience replay buffer in real time, updates the abnormal response action sequence in the strategy library in step S4 and increases its priority coefficient, and at the same time refits the abnormal probability matrix in the four-dimensional state transition equation based on the current abnormal data, and integrates the verification data into the model parameter update process of steps S3 / S4 to form a closed-loop management mechanism covering detection - response - verification and optimization.
[0026] The said step S6 further includes the following steps: S6.1: Based on the task completion time and energy consumption data continuously recorded in step S1, construct a four-dimensional target achievement evaluation system, including four indicators: task splitting timeliness, empty vehicle utilization rate on the return journey, load balancing deviation rate, and tank cleaning energy consumption ratio; S6.2: Use the Bayesian optimization algorithm to synchronously update the probability parameters of the four-dimensional state transition equation and the four-dimensional target weights. The optimization objective function is defined as the weighted sum of the four-dimensional target achievements; the formula of the optimization objective function is as follows: ; where i is the index of the four-dimensional target. When i = 1, γ 1 is the priority coefficient of the task splitting target, and the achievement degree 1 is determined by the task splitting timeliness index; when i = 2, γ 2is the priority coefficient and achievement degree for the goal of empty vehicle utilization on the return trip 2 is determined by the empty vehicle utilization rate index on the return trip; when i = 3, γ 3 is the priority coefficient and achievement degree for the goal of load balancing 3 is determined by the load balancing deviation rate index; when i = 4, γ 4 is the priority coefficient and achievement degree for the goal of flushing tank cost 4 is determined by the flushing tank energy consumption ratio index, γ i is the priority coefficient for the four-dimensional goal, which is determined by mapping the four-dimensional goal weights ω in step S4 i through a preset rule S6.3: Push the optimized parameters to the modules in step S3 / S4 in real time to form a complete closed loop including data collection, target evaluation, parameter optimization, and strategy update, ensuring that the system completes model iteration within each task cycle
[0027] A multi-objective dynamic scheduling system for a fleet based on edge computing includes a data collection and preprocessing module, a task parsing and classification module, a multi-objective optimization model module, a strategy generation and optimization module, an emergency management module, and a closed-loop iteration module The data collection and preprocessing module is responsible for collecting vehicle status, cloud tasks, and road condition data in real time, eliminating noise through a spatio-temporal alignment algorithm and generating derivative features, and constructing a structured vector space to provide standardized input for subsequent models The task parsing and classification module uses natural language processing technology to parse task instructions, extract features such as cargo type, time limit, and destination coordinates, generate derivative indicators such as distance, time window constraints, and cargo volume-weight ratio, and divides tasks into three categories: urgent, regular, and empty vehicle on the return trip through a clustering algorithm The multi-objective optimization model module is used to construct a four-dimensional state transition equation, define the action space and state change rules, and transform the multi-objective into a comprehensive cost function with dynamic weights The strategy generation and optimization module dynamically adjusts the four-dimensional goal weights through a fuzzy neural network, constructs a strategy library in combination with a deep reinforcement learning algorithm, and updates the optimal scheduling plan regularly The emergency management module monitors vehicle anomalies and sudden road condition changes in real time. After triggering the emergency mechanism, it reallocates tasks based on the four-dimensional model and the strategy library, updates model parameters, and verifies the effectiveness of the plan to form a closed loop for anomaly handling The closed-loop iteration module evaluates the achievement degree of the four-dimensional goal based on task completion time and energy consumption data, and uses the Bayesian optimization algorithm to synchronously update the four-dimensional model parameters and the four-dimensional goal weights to achieve continuous iteration of data collection - optimization - execution
[0028] Suppose a cold-chain logistics company has 3 refrigerated trucks with load capacities of 20 tons, 25 tons, and 30 tons respectively, and needs to complete three types of transportation tasks: emergency medicine transportation (5m³ / 2 hours / 100km, pollution level 85%); regular fresh food distribution (8m³ / 8 hours / 200km, pollution level 30%); return empty truck task (6m³ / 6 hours, path overlap 70%). The regional empty truck rate is 25%, and the unit fuel consumption on the highway is 0.2L / km; On-vehicle sensors continuously obtain vehicle location, load (15 tons / 20 tons / 25 tons), remaining mileage (300km / 250km / 400km), and tank cleanliness (85% / 30% / 50%).
[0029] The BERT model identifies emergency tasks as dangerous goods and generates time window constraints (early arrival penalty coefficient 0, late arrival penalty coefficient 1.25); the volume-weight ratio of regular tasks is 0.53m³ / ton; the path overlap of return tasks is 70%.
[0030] The possibility of task splitting s T车辆1 = 3 tons / m³ (remaining load capacity 15 tons / cargo volume 5m³); The return empty truck rate s R = 25%; The load balance degree is calculated according to the formula s L = 50.
[0031] Action selection: Assign vehicle 3 to the emergency task (remaining mileage 400km), without splitting the goods; Split 30% of the regular task (vehicle 2 transports 6m³, vehicle 1 supplements 2m³); Vehicle 1 matches the return task, with a path overlap of 70%. The fuzzy neural network outputs ω T =0.4 (emergency task), ω R =0.3 (return task), ω C =0.2 (dangerous goods).
[0032] If the load sensor of vehicle 2 is abnormal (fluctuation exceeds 5%), trigger the backtracking search algorithm, reassign tasks, and update the four-dimensional state transition probability.
[0033] The task on-time rate is 93% (the emergency task is delivered 15 minutes in advance); The empty truck rate drops from 40% to 22%, reducing the empty driving mileage by 30km; The energy consumption ratio of tank washing is 8% (achievement rate 90%); Bayesian optimization updates the four-dimensional model parameters, and the load balance deviation rate drops from 25% to 8%.
[0034] Comparison result: Compared with the traditional single-objective optimization, the present invention increases the vehicle utilization rate from 60% to 85% and improves the comprehensive transportation efficiency by 32%, verifying the synergistic optimization effect of four-dimensional dynamic programming and real-time weight adjustment.
[0035] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claimed invention.
Claims
1. A multi-objective dynamic scheduling method for a fleet based on edge computing, characterized by: The following steps are involved: S1. On-board sensors collect vehicle location, load, remaining mileage, and tank cleanliness data in real time. Edge devices simultaneously receive cloud-based task requirements and road condition information to build a vehicle state vector. S2, based on the cloud-based task demand information contained in the vehicle state vector constructed in S1, uses natural language processing technology to parse task instructions, extract cargo type, transportation timeliness, and destination characteristics, and divides tasks into three categories: emergency transportation, regular transportation, and return empty vehicle tasks; S3. Based on the above data, a four-dimensional dynamic programming model including task splitting, return empty vehicle utilization, load balancing, and tank washing cost is constructed; based on the vehicle state vector constructed by S1 and the task categories divided by S2, a four-dimensional state transfer equation including task splitting T, return empty vehicle utilization R, load balancing L, and tank washing cost C is deployed in the edge computing chip, and multiple objectives are converted into a comprehensive cost function with dynamic weights; S4, based on the regional empty vehicle rate reflected by the road condition information in S1, combined with the task urgency and cargo type characteristics extracted by S2, the fuzzy neural network dynamically adjusts the four-dimensional target weights in the comprehensive cost function of S3, combined with the deep reinforcement learning DRL algorithm, and builds a strategy library based on the four-dimensional state transfer equation of S3, and regularly updates the optimal scheduling plan; S5. When the load sensor is abnormal or the road condition information feedback shows road congestion in the vehicle data collected by S1, the rule-based emergency mechanism is triggered. Based on the four-dimensional state transfer equation of S3 and the strategy library of S4, the task is reallocated and the four-dimensional model parameters are updated through the backtracking search algorithm; S6 collects the task completion time, energy consumption and other data continuously recorded by S1 after the task is completed, evaluates the achievement of the four-dimensional goal in the comprehensive cost function of S3 online, and uses the Bayesian optimization algorithm to update the four-dimensional model parameters of S3 and the four-dimensional target weights of S4, forming a closed-loop iterative system of data collection-S3 / S4 optimization-execution feedback.
2. The multi-objective dynamic scheduling method for a fleet based on edge computing according to claim 1 is characterized in that: The S1 further comprises the following steps: S1.1: On-board sensors collect data on vehicle location, load, remaining mileage, and tank cleanliness; S1.2: Edge devices receive cloud task requirements and real-time traffic information synchronously through the 5G network; S1.3: Use the spatiotemporal alignment algorithm to eliminate data noise and construct the vehicle state vector. Specifically: First, the vehicle position, load, remaining mileage, tank cleanliness data, cloud tasks and road condition data collected by the on-board sensors in steps S1.1 and S1.2 are used to generate derivative features such as vehicle speed and load change through mathematical modeling, and establish the correlation characteristics between cargo type and cleanliness, road conditions and energy consumption; then standardize the multi-dimensional data, and use dimensionality reduction technology to construct a structured vector space covering vehicle status, task characteristics and environmental constraints, which is used for the subsequent four-dimensional dynamic programming model as input data to support the optimization calculation of vehicle scheduling decisions.
3. The multi-objective dynamic scheduling method for a fleet based on edge computing according to claim 2 is characterized in that: The step S2 further comprises the following steps: S2.1: Use natural language processing technology to parse the cloud-based task requirement information contained in the vehicle state vector; specifically: use the pre-trained BERT model to parse the task instruction text, identify the cargo type, transportation time and destination coordinates, then encode the cargo type into a unique ID, store the transportation time as a minute-level value, and convert the destination coordinates into longitude and latitude floating-point data; S2.2: Based on the parsed task instructions, generate derived features related to the task. The derived features are divided into the distance between the task start and destination, the time window constraint, and the cargo volume-to-weight ratio. Specifically, the distance is calculated using the Euclidean distance formula, the penalty coefficients for early and late arrivals are determined based on the time required to arrive at the task, and the volume-to-weight ratio is obtained by dividing the cargo volume by the weight. The time window constraint is specifically: the task required arrival time is set to t req , the early arrival time is t early , the late arrival time is t late , then the early arrival penalty coefficient is , the late arrival penalty coefficient is ; S2.3: Using the K-means clustering algorithm, the derived features such as transportation time, distance between the origin and the destination, and volume-to-weight ratio are divided into different categories; the classification criteria are: transportation time of less than 4 hours and distance not exceeding 100km is emergency transportation; transportation time of 4-24 hours and distance of 100-500km is regular transportation; and the task destination and the current vehicle location path overlap by no less than 60% are return empty tasks.
4. The multi-objective dynamic scheduling method for a fleet based on edge computing according to claim 3 is characterized in that: The step S3 further comprises the following steps: S3.1: Establish a four-dimensional dynamic programming model containing the state of the four-dimensional target to quantify the current state of the vehicle and the task constraints; Specifically: First, define the four-dimensional state vector S = [s T , s R ,s L , s C ], the specific meanings of each dimension are as follows: s T Represents the possibility of task splitting, which is calculated by the ratio of the vehicle's remaining load and the task cargo volume. The calculation formula is as follows: ; s R Represents the return empty vehicle rate, based on the empty vehicle distribution statistics of the current vehicle location and the task destination area. The calculation formula is as follows: ; s L Represents the load balance, which is quantified by the sum of the squares of the load differences between adjacent vehicles. The calculation formula is as follows: ; s C Represents the tank cleaning demand, which is obtained by matching the tank cleanliness sensor data with the cargo type pollution level. When the pollution level ≥ the threshold, tank cleaning is triggered; S3.2: Define executable scheduling actions; specifically: define a four-dimensional action vector A = [a T , a R , a L , a C ], the operation rules of each dimension are as follows: a T Represents the task split ratio, which is determined according to the remaining load and cargo integrity constraints, and the split ratio range is {0, 70%}; a R Represents the return task matching, and prioritizes the return tasks whose path overlap with the current task is ≥ 60%. The path overlap calculation formula is as follows: ; a L Represents the load transfer amount, which is calculated based on the load difference between adjacent vehicles and the road weight limit constraint. The single transfer amount does not exceed the preset standard; a C Represents the tank cleaning decision. When the cargo type is switched or the tank contamination level is ≥80%, the tank cleaning operation is forcibly triggered; S3.3: Convert the four-dimensional target into a unified quantitative indicator; specifically: construct a comprehensive cost function with dynamic weights, the formula is as follows: ; Among them, C T represents the task splitting cost, which is positively correlated with the cargo integrity loss rate. The formula is as follows: ; C R Represents the cost of empty driving, which is calculated by the return empty driving mileage × unit fuel consumption. The unit fuel consumption is determined based on the vehicle type and road conditions. C L represents the load imbalance cost, which is quantified by the sum of the squares of the load differences between adjacent vehicles, and the formula is as follows: ; C C Represents the tank cleaning cost, which is positively correlated with the tank contamination level and cleaning time. The formula is as follows: ; S3.4: Describe the changing rules of the vehicle state after executing the action, and provide a mathematical basis for dynamic programming; specifically: establish a four-dimensional state transfer matrix P(S{t+1} | St, At), and the transfer probability of each dimension is determined by historical data training; Among them, P(s{T,t+1} | s{T,t}, a{T,t}) is the probability of change of the remaining load after task splitting, which is based on the vehicle load distribution fitting; P(s{R,t+1} | s{R,t}, a{R,t}) is the probability of change of the empty vehicle rate after the return task matching, which is based on the historical data statistics of regional traffic flow; P(s{L,t+1} | s{L,t}, a{L,t}) is the probability of change of balance after load transfer, which is modeled through load distribution experimental data; P(s{C,t+1} | s{C,t}, a{C,t}) is the probability of change of cleanliness after tank washing decision, which is based on the historical data of tank cleanliness sensor.
5. The multi-objective dynamic scheduling method for a fleet based on edge computing according to claim 4 is characterized in that: The step S4 further includes the following contents: S4.1: Based on the real-time features of the above-mentioned road condition information, the regional empty vehicle rate calculated, the task urgency mapped by task classification, the cargo type converted by the parsed cargo type ID, the road grade extracted by the road condition, and the remaining load ratio calculated by the vehicle state vector, a three-tier fuzzy neural network is constructed: the input layer contains 12 feature nodes, the hidden layer adopts a two-layer ReLU activation structure with 32 neurons each, and the output layer generates a four-dimensional target weight vector ω=[ω T , ω R ,ω L , ω C ], where ω T The weight of the task splitting target; ω R is the weight of the empty return vehicle utilization target; ω L is the weight of the load balancing target; ω C is the weight of the tank washing cost target; the four-dimensional target weight ω i satisfy Normalization constraint; The network implements weight constraints through preset rules: mandatory allocation of ω for urgent transportation tasks T ≥0.4, empty return missionω R ≥0.3, dangerous goods transportation taskω C ≥0.2, ensuring the priority of multi-objective collaborative optimization in different scenarios; S4.2: Build a deep reinforcement learning strategy library. Specifically: build a three-dimensional state space based on the vehicle state vector, task classification results and S1 environmental parameters, and design a discrete action space containing four-dimensional actions, including task splitting ratio, return task matching, load transfer amount, and tank washing decision; adopt a reward mechanism coupled with the S3 comprehensive cost function to define the reward function. The formula is as follows: ; where λ i is the reward weight of the four-dimensional target, which is the four-dimensional target weight ω described in S4.1 i After normalization, C i The task split cost C calculated for S3 T , return trip empty cost C R , load imbalance cost C L and tank washing cost C C ; Among them, λ T is the reward weight of the task splitting objective, λ R is the reward weight for the return empty vehicle utilization target, λ L is the reward weight of the load balancing objective, λ C is the reward weight of the tank washing cost target; the strategy library is trained online through the PPO algorithm in an environment containing millions of state-action combinations formed by the vehicle state vector in S1, the task characteristics in S2, and the four-dimensional state space and action space in S3 to generate the optimal scheduling strategy sequence; An online learning mechanism is adopted, based on the four-dimensional state transfer equation, and the PPO algorithm is used to train the policy network to generate the optimal action sequence. At the same time, the experience replay buffer is used to store the vehicle data collected in real time in step S1, the task instructions parsed by S2, and the millions of real-time state-action-reward triples of cost and state transfer results generated by S3, and the sample utilization rate is improved by the priority experience replay strategy.
6. The multi-objective dynamic scheduling method for a fleet based on edge computing according to claim 5 is characterized in that: The step S5 further comprises the following steps: S5.1: Real-time monitoring of vehicle status and road condition data, identification of abnormal events and triggering of emergency response; specifically, the sliding window filter and 3σ criterion are used to detect abnormal load sensor data, and the vehicle speed is continuously below the threshold and the road congestion index is used to double-verify the road congestion status to ensure accurate identification and triggering of abnormal events; S5.2: Based on the detected anomaly type, activate the corresponding emergency strategy; reorder the tasks by building a priority queue, and call the action sequence of historical successful responses to similar anomalies in the strategy library in step S4, and combine the four-dimensional state transfer equation to preview the state changes of the alternative strategies to provide a feasible solution for subsequent task allocation; S5.3: Use the backtracking search algorithm to traverse the four-dimensional state space of dispatchable vehicles, give priority to high-reward actions in the policy library, optimize the search path through pruning conditions, and split the tasks according to the rules of neighboring allocation of urgent subtasks and delayed execution of ordinary subtasks; S5.4: Refit the load state transition probability of abnormal vehicles based on the execution data after the abnormal event, and update the empty vehicle rate transition probability based on the historical congestion data; adjust the four-dimensional target weight of step S4 through the online gradient descent algorithm, gradually increase the task splitting priority according to the preset standard, and reduce the load balancing weight to optimize the model's adaptability to abnormal scenarios; S5.5: A quantitative verification system is adopted to calculate the punctuality rate of abnormal tasks based on the vehicle position data of step S1 and the transportation time efficiency of step S2, and the target value is set to be no less than 95%; at the same time, the load change rate in adjacent preset time periods is calculated through the load sensor data of step S1, and the fluctuation control target within 5% is required to be achieved; when the handling of three consecutive abnormal events fails to meet the above indicators, the system will trigger the manual intervention threshold and automatically generate an abnormal report and push it to the dispatcher terminal; at the same time, the system stores the current abnormal handling strategy in the experience playback buffer in real time, updates the abnormal response action sequence in the strategy library of step S4 and increases its priority coefficient, and refits the abnormal probability matrix in the four-dimensional state transfer equation based on the abnormal data, and integrates the verification data into the step S3 / S4 model parameter update process to form a closed-loop management mechanism covering detection-response-verification and optimization.
7. The multi-objective dynamic scheduling method for a fleet based on edge computing according to claim 6 is characterized by: The step S6 further comprises the following steps: S6.1: Based on the task completion time and energy consumption data continuously recorded in step S1, a four-dimensional target achievement evaluation system is constructed, including four indicators: task splitting timeliness, return empty vehicle utilization rate, load balancing deviation rate, and tank washing energy consumption ratio; S6.2: The Bayesian optimization algorithm is used to synchronously update the probability parameters of the four-dimensional state transfer equation and the four-dimensional target weights. The optimization objective function is defined as the weighted sum of the four-dimensional target achievement degrees. The formula of the optimization objective function is as follows: ; Among them, i is the index of the four-dimensional goal. When i=1, γ1 is the priority coefficient of the task splitting goal, and the achievement degree 1 is determined by the task splitting timeliness index; when i=2, γ2 is the priority coefficient of the return empty car utilization goal, and the achievement degree 2 is determined by the return empty car utilization rate index; when i=3, γ3 is the priority coefficient of the load balancing goal, and the achievement degree 3 is determined by the load balancing deviation rate index; when i=4, γ4 is the priority coefficient of the tank washing cost goal, and the achievement degree 4 is determined by the tank washing energy consumption ratio index, γ i is the priority coefficient of the four-dimensional target, which is determined by the four-dimensional target weight ω in step S4. i Determined by preset rule mapping; S6.3: Push the optimized parameters to step S3 / S4 modules in real time, forming a complete closed loop including data collection, target evaluation, parameter optimization and strategy update, ensuring that the system completes model iteration within each task cycle.
8. A multi-objective dynamic scheduling system for a fleet based on edge computing, applied to the multi-objective dynamic scheduling method for a fleet based on edge computing according to claims 1-7, characterized in that: It includes data acquisition and preprocessing module, task analysis and classification module, multi-objective optimization model module, strategy generation and optimization module, emergency management module and closed-loop iteration module; The data acquisition and preprocessing module is responsible for real-time acquisition of vehicle status, cloud tasks and road condition data, eliminating noise and generating derivative features through spatiotemporal alignment algorithm, and constructing structured vector space to provide standardized input for subsequent models; The task analysis and classification module uses natural language processing technology to analyze task instructions, extract characteristics such as cargo type, timeliness and destination coordinates, generate derived indicators such as distance, time window constraints and cargo volume-to-weight ratio, and classify tasks into three categories: emergency, routine and return empty through clustering algorithms; The multi-objective optimization model module is used to construct a four-dimensional state transfer equation, define the action space and state change rules, and convert the multi-objectives into a comprehensive cost function with dynamic weights; The strategy generation and optimization module dynamically adjusts the four-dimensional target weights through a fuzzy neural network, builds a strategy library in combination with a deep reinforcement learning algorithm, and regularly updates the optimal scheduling solution; The emergency management module monitors vehicle anomalies and sudden changes in road conditions in real time, and after the emergency mechanism is triggered, it reallocates tasks based on the four-dimensional model and strategy library, updates model parameters and verifies the effectiveness of the solution, thus forming an abnormality handling closed loop; The closed-loop iteration module evaluates the achievement of the four-dimensional target based on the task completion time and energy consumption data, and uses the Bayesian optimization algorithm to synchronously update the four-dimensional model parameters and the four-dimensional target weights, thereby realizing continuous iteration of data collection-optimization-execution.
Citation Information
Patent Citations
Collaborative multi-objective algorithm-based multi-stage low-carbon logistics distribution network planning method
CN107833002A
Vehicle and goods matching method based on AHP-DBN
CN113379356A
Cited By
Unmanned forklift dynamic task scheduling method and system based on deep reinforcement learning of cold chain warehouse
CN120235559A
Chemical transportation industry chain management system and method
CN120338645A
Annular automatic seasoning distribution system and method
CN120793465A
Multi-agent collaborative decision-making system and method for intelligent manufacturing
CN121028855A
Hydrogen energy vehicle intelligent scheduling system based on edge calculation
CN121235329A