Intelligent Task Division Method for Collaborative Operation of Driverless Sanitation Vehicles

Through the combination of clustering algorithm and Actor-Critic reinforcement learning algorithm, urban sanitation vehicle tasks are intelligently divided, solving the problem that traditional methods are difficult to obtain global optimal solutions, and achieving efficient and energy-saving sanitation services.

CN118966684BActive Publication Date: 2025-06-20COWA TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411038769.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2025-06-20
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

When urban sanitation vehicles perform cleaning and garbage collection tasks, how to effectively plan routes and allocate tasks is a complex problem. Traditional methods are difficult to obtain global optimal solutions and are difficult to cope with complex and dynamic changing environments.

Method used

A task intelligent division method for collaborative operation of unmanned sanitation vehicles is proposed. The clustering algorithm is used to cluster the task sections, and combined with the Actor-Critic reinforcement learning algorithm and greedy algorithm, the subtask division of the task section is optimized, and factors such as the distance from the task section to the clustering center, the cost of the subtask path, the distance of the supply station and the warehouse are comprehensively considered.

Benefits of technology

The overall optimal division of urban sanitation vehicle tasks has been achieved, work efficiency has been improved, energy consumption has been reduced, and urban sanitation services have been optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118966684B_ABST
    Figure CN118966684B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent task partitioning method for collaborative operation of unmanned sanitation vehicles, including: clustering all task sections by using a clustering algorithm, where the number of clusters is a preset number of subtasks, to obtain initial cluster centers; initializing a policy function table and a value function table; selecting a corresponding subtask for each task section according to the current policy function table; calculating a reward value according to the selection result by using a reward function, and evaluating the expected return of the selection result; the reward function involves multiple factors; based on the update rule of Actor-Critic, updating the value function and the policy function of the current task section according to the reward value and the value function of the next task section; repeating the above steps to update the policy function table and the value function table until the optimal subtasks are assigned to each task section. The present invention can improve work efficiency, reduce energy consumption, and optimize urban sanitation services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of task division of unmanned sanitation vehicles, and particularly relates to an intelligent task division method for collaborative operation of unmanned sanitation vehicles. Background Art

[0002] With the booming development of autonomous driving technology, a large number of unmanned operation vehicles have emerged in the market, such as unmanned sweeping vehicles and unmanned sprinkler vehicles. These vehicles play an important role in improving efficiency, reducing costs, improving the working environment, and promoting environmental protection, bringing revolutionary changes to the fields of modern urban management and transportation. When urban sanitation vehicles perform tasks such as cleaning and garbage collection, how to effectively plan routes and allocate tasks is an important and complex issue. Traditional methods usually rely on manual experience or simple rules, such as the nearest neighbor rule, k-means algorithm, etc. These methods often cannot obtain the global optimal solution and are difficult to cope with complex and dynamically changing environments. Summary of the Invention

[0003] To solve the above technical problems, the present invention proposes an intelligent task division method for collaborative operation of unmanned sanitation vehicles.

[0004] To achieve the above object, the technical solution of the present invention is as follows:

[0005] In a first aspect, the present invention discloses an intelligent task division method for collaborative operation of unmanned sanitation vehicles, including:

[0006] Step S1: Using a clustering algorithm to cluster all task sections, where the number of clusters is the preset number of subtasks, to obtain the initial cluster centers;

[0007] Step S2: Initializing the policy function table and the value function table;

[0008] The policy function table stores the probability of each task section selecting each subtask;

[0009] The value function table stores the value function of each task section, and the value function is the expected return that can be obtained when selecting the corresponding subtask;

[0010] Step S3: Selecting a corresponding subtask for each task section according to the current policy function table;

[0011] Step S4: According to the selection result, using the reward function to calculate the reward value and evaluate the expected return of the selection result;

[0012] The reward function involves one or more factors such as the distance from the task section to the corresponding cluster center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse;

[0013] Step S5: Based on the Actor-Critic update rule, update the value function and policy function of the current task segment according to the reward value and the value function of the next task segment;

[0014] Step S6: Repeat Steps S3 - S5 to update the policy function table and the value function table until the optimal subtasks are assigned to each task segment.

[0015] Based on the above technical solution, the following improvements can be made:

[0016] As a preferred solution, the greedy algorithm is used in the reward function to calculate the optimal path cost of the subtask, including the following steps:

[0017] Step A: Obtain all the task segments corresponding to the subtask and the connection relationship between the task segments, and regard each task segment as a node;

[0018] Step B: Initialize a boolean list to mark whether the node has been visited;

[0019] Initialize the path length to 0, and set the starting node to 0;

[0020] Step C: Starting from the current node, find the next unvisited node to minimize the distance to reach this node, and mark the current node as visited;

[0021] Traverse all unvisited nodes, calculate the distance from the current node to each unvisited node, and update the minimum distance and the next node until the shortest path is found through iteration;

[0022] Step D: Update the path length by adding the minimum distance to the path length.

[0023] As a preferred solution, the Euclidean distance is used to calculate the distance from the task segment to the corresponding cluster center.

[0024] As a preferred solution, the calculation formula of the reward function R is as follows:

[0025]

[0026] where x i is the coordinate or feature vector of the i-th task segment;

[0027] c j is the coordinate of the j-th cluster center;

[0028] P is the optimal path length of the subtask calculated by the greedy algorithm;

[0029] n is the number of task segments of the current subtask;

[0030] is the shortest distance from the i-th task section to the supply station;

[0031] is the distance from the i-th task section to the warehouse.

[0032] As a preferred solution, in step S5, the value function is updated through the following formula;

[0033] V(s)←V(s)+α*(r+γ*V(s′)-V(s))

[0034] where V(s) is the value function of the current task section;

[0035] V(s′) is the value function of the next task section;

[0036] r is the reward value;

[0037] α is the learning rate;

[0038] γ is the discount factor;

[0039] Update the policy function table through the following formula;

[0040] π(a|s)←π(a|s)+α*(r+γ*V(s′)-V(s))

[0041] where π(a|s) is the probability of selecting the a-th subtask under the s-th task section.

[0042] As a preferred solution, the softmax function is used to process the probabilities in the policy function table.

[0043] In a second aspect, the present invention also discloses a task intelligent partitioning device for collaborative operation of unmanned sanitation vehicles, including:

[0044] a clustering module, configured to cluster all task sections by using a clustering algorithm, where the number of clusters is the preset number of subtasks, to obtain an initial clustering center;

[0045] an initialization module, configured to initialize the policy function table and the value function table;

[0046] The policy function table stores the probabilities of each task section selecting each subtask;

[0047] The value function table stores the value function of each task section, and the value function is the expected return that can be obtained when selecting the corresponding subtask;

[0048] a selection module, configured to select a corresponding subtask for each task section according to the current policy function table;

[0049] A reward value calculation module, which is used to calculate the reward value according to the selection result by using a reward function and evaluate the expected return of the selection result;

[0050] The reward function involves one or more factors among the distance from the task section to the corresponding cluster center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse;

[0051] An update module, which is used to update the value function and policy function of the current task segment based on the update rule of Actor-Critic according to the reward value and the value function of the next task segment;

[0052] A repetition module, which is used to repeatedly execute the methods in the selection module, the reward value calculation module, and the update module to update the policy function table and the value function table until the optimal subtask is assigned to each task segment.

[0053] As a preferred solution, the optimal path cost calculation unit calculates the optimal path cost of the subtask in the reward function by using the greedy algorithm, including:

[0054] An acquisition unit, which is used to acquire all the task sections corresponding to the subtask and the connection relationship between the task sections, and regard each task section as a node;

[0055] A greedy algorithm initialization unit, which is used to initialize a boolean list for marking whether the node has been visited;

[0056] Initialize the path length to 0 and set the starting node to 0;

[0057] An iterative search unit, which is used to start from the current node, search for the next unvisited node to make the distance to this node the shortest, and mark the current node as visited;

[0058] Traverse all unvisited nodes, calculate the distance from the current node to each unvisited node, and update the minimum distance and the next node until the shortest path is iteratively found;

[0059] A path update unit, which is used to update the path length and add the minimum distance to the path length.

[0060] As a preferred solution, the Euclidean distance is used to calculate the distance from the task section to the corresponding cluster center.

[0061] As a preferred solution, the calculation formula of the reward function R is as follows:

[0062]

[0063] Among them, x i is the coordinate or feature vector of the i-th task section;

[0064] c j is the coordinate of the j-th clustering center;

[0065] P is the optimal path length of the subtask calculated by the greedy algorithm;

[0066] n is the number of task sections of the current subtask;

[0067] is the shortest distance from the i-th task section to the supply station;

[0068] is the distance from the i-th task section to the warehouse.

[0069] As a preferred solution, in step S5, the value function is updated through the following formula;

[0070] V(s)←V(s)+α*(r + γ*V(s′)-V(s))

[0071] where V(s) is the value function of the current task section;

[0072] V(s′) is the value function of the next task section;

[0073] r is the reward value;

[0074] α is the learning rate;

[0075] γ is the discount factor;

[0076] The policy function table is updated through the following formula;

[0077] π(a|s)←π(a|s)+α*(r + γ*V(s′)-V(s))

[0078] where π(a|s) is the probability of selecting the a-th subtask under the s-th task section.

[0079] As a preferred solution, the softmax function is used to process the probabilities in the policy function table.

[0080] In a third aspect, the present invention also discloses a storage medium storing one or more computer-readable programs, and the one or more programs include instructions adapted to be loaded and executed by a memory to perform any one of the above-mentioned task intelligent partitioning methods for collaborative operation of unmanned sanitation vehicles.

[0081] The present invention discloses a task intelligent partitioning method for collaborative operation of unmanned sanitation vehicles, which has the following beneficial effects:

[0082] The present invention divides the task section into multiple subtasks and uses the Actor-Critic reinforcement learning algorithm to optimize the division method of each subtask, thereby realizing the optimization of the division of the entire task section. At the same time, the present invention comprehensively considers factors such as the distance from the task section to the corresponding clustering center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse.

[0083] By using the Actor-Critic reinforcement learning model and combining with greedy algorithms, clustering algorithms, etc., the present invention effectively divides the tasks of urban sanitation vehicles, can solve the problems of route planning and task allocation when urban sanitation vehicles perform tasks such as cleaning and garbage collection, improve work efficiency, reduce energy consumption, and optimize urban sanitation services. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0085] Figure 1 It is a flowchart of the task intelligent division method provided by the embodiment of the present invention.

[0086] Figure 2 It is a framework diagram of the Actor-Critic reinforcement learning model provided by the embodiment of the present invention.

[0087] Figure 3 (a), (b), and (c) are respectively schematic diagrams of the urban environment models of Scenario 1, Scenario 2, and Scenario 3 provided by the embodiments of the present invention.

[0088] Figure 4 (a), (b), and (c) are respectively diagrams showing the proportion of the total non-working path and the number of vehicles when 10* complete the operation in three real test scenarios provided by the embodiments of the present invention.

[0089] Figure 5 (a), (b), and (c) are respectively diagrams showing the proportion of the non-working path per 1 km traveled by each unmanned vehicle in three real test scenarios provided by the embodiments of the present invention.

[0090] Figure 6 (a), (b), and (c) are respectively schematic diagrams of the total cost (Cost) in three real test scenarios provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0091] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0092] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0093] The expression "including" an element is an "open-ended" expression, which only means that there are corresponding components or steps, and should not be construed as excluding additional components or steps.

[0094] In order to achieve the purpose of the present invention, in some embodiments of the task intelligent partitioning method for unmanned sanitation vehicle collaborative operation, as Figure 1 shown, the task intelligent partitioning method includes:

[0095] Step S1: Use a clustering algorithm to cluster all task sections, where the number of clusters is the preset number of subtasks, to obtain the initial clustering centers;

[0096] Step S2: Initialize the policy function table and the value function table;

[0097] The policy function table stores the probability of each task section selecting each subtask;

[0098] The value function table stores the value function of each task section, and the value function is the expected return that can be obtained when selecting the corresponding subtask;

[0099] Step S3: Select a corresponding subtask for each task section according to the current policy function table;

[0100] Step S4: According to the selection result, use the reward function to calculate the reward value and evaluate the expected return of the selection result;

[0101] The reward function involves one or more factors such as the distance from the task section to the corresponding clustering center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse;

[0102] Step S5: Based on the update rule of Actor-Critic, update the value function and the policy function of the current task section according to the reward value and the value function of the next task section;

[0103] Step S6: Repeat steps S3 - S5 to update the policy function table and the value function table until the optimal subtask is assigned to each task section.

[0104] The present invention will be described in detail below.

[0105] In step S1, the K-means clustering algorithm is used to cluster all task sections, and the number of clusters is the preset number of subtasks, obtaining the initial cluster centers.

[0106] K-means is a commonly used clustering algorithm for dividing samples in a dataset into a preset number of different groups or clusters. Its basic principle is to find the center points of the clusters through iterative optimization, minimizing the sum of the distances from the sample points to the center points of their respective clusters.

[0107] In recent years, with the development of artificial intelligence and machine learning technologies, reinforcement learning, as a method that can learn optimal strategies through interaction with the environment, has been widely applied to various optimization problems. In particular, the Actor-Critic model, as a reinforcement learning algorithm that combines the two methods of value iteration and policy iteration, has shown excellent performance in many problems.

[0108] However, how to apply reinforcement learning, especially the Actor-Critic model, to the task division problem of urban sanitation vehicles remains a challenging issue.

[0109] The present invention uses reinforcement learning to update the task division strategy according to the reward values of each task section (i.e., state), thereby finding an optimal task division scheme.

[0110] The present invention uses the Actor-Critic reinforcement learning model to select corresponding subtasks (i.e., actions) and update the value function table and policy function table. The Actor-Critic model combines the advantages of the two methods of value iteration and policy iteration, and can effectively balance exploration and exploitation, thereby improving the effect of task division.

[0111] As Figure 2 shown, the Actor-Critic model is a reinforcement learning algorithm that combines the two methods of value iteration (Critic) and policy iteration (Actor). In the Actor-Critic model, "Actor" and "Critic" are two main components, and they each have different responsibilities: the responsibility of the Actor is to select actions according to the current policy. The policy is usually expressed as the probability of performing each possible action in a given state. In each step, the Actor will select an action according to the current policy. The responsibility of the Critic is to evaluate the actions selected by the Actor. It calculates the expected return of each action and provides feedback to the Actor. This feedback is used to update the Actor's policy, making actions with higher expected returns more likely to be selected in the future.

[0112] The main steps of the Actor-Critic model are mainly reflected in steps S2 - S6.

[0113] Step S2: Initialize the policy function table and the value function table;

[0114] The policy function table stores the probabilities of selecting each subtask for each task section.

[0115] The value function table stores the value functions of each task section. The value function is the expected return that can be obtained when selecting the corresponding subtask.

[0116] Step S3: Select a corresponding subtask for each task section according to the current policy function table. This step is completed by the Actor.

[0117] Step S4: According to the selection result, calculate the reward value using the reward function and evaluate the expected return of this selection result. This step is completed by the Critic;

[0118] The reward function involves one or more of the following factors: the distance from the task section to the corresponding cluster center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse.

[0119] Step S5: According to the feedback of the Critic, based on the update rules of Actor-Critic, update the value function and the policy function of the current task section according to the reward value and the value function of the next task section.

[0120] Step S6: Repeat steps S3 - S5 to update the policy function table and the value function table until the optimal subtask is assigned to each task section.

[0121] Specifically, it is hoped to increase the probability of selecting actions with higher expected returns and decrease the probability of selecting actions with lower expected returns. At the same time, the present invention also updates the value function to more accurately estimate the expected return of each state.

[0122] In the Actor-Critic model, the present invention maintains two tables: the value function table and the policy function table. The value function table is used to store the value function of each state, representing the expected return that can be obtained in that state. The policy function table is used to store the probabilities of performing each possible action in each state.

[0123] Regarding the policy function table. The policy function table is a two-dimensional array, where each row corresponds to a state and each column corresponds to an action. Each element of the policy function table represents the probability of performing a given action in a given state. For example, assuming there are 3 states and 2 actions, then the policy function table may be as follows:

[0124] [[0.7,0.3],

[0125] [0.4,0.6],

[0126] [0.5,0.5]]

[0127] This means that the probability of performing Action 1 in State 1 is 0.7, and the probability of performing Action 2 is 0.3; the probability of performing Action 1 in State 2 is 0.4, and the probability of performing Action 2 is 0.6; the probabilities of performing Action 1 and Action 2 in State 3 are both 0.5.

[0128] Regarding the value function table. The value function table is a one-dimensional array, where each element corresponds to a state. Each element of the value function table represents the expected return that can be obtained in that state.

[0129] For example, assuming there are 3 states, then the value function table might be as follows:

[0130] [0.5,0.6,0.7]

[0131] This means that the expected return in State 1 is 0.5, the expected return in State 2 is 0.6, and the expected return in State 3 is 0.7.

[0132] In reinforcement learning, the goal is to find the optimal policy function table and value function table through learning, such that starting from any state, by following the actions selected in the policy function table, the maximum cumulative return can be obtained.

[0133] The above states represent task segments, and the above actions represent the selection of subtasks.

[0134] The following is a detailed introduction to the update process of the policy function table and the value function table:

[0135] In each step, the value function table and the policy function table are updated according to the update rules of Actor-Critic. In the Actor-Critic model, the feedback of the Critic is used to update the policy function table and the value function table. Specifically, the present invention uses the following update rules:

[0136] In the update process of the value function table, first calculate the Temporal Difference Error (TD Error). The TD Error is the difference between the actual reward and the value function of the current state, plus the discounted value of the value function of the next state. Then, this error is multiplied by the learning rate and added to the value function of the current state to update the value function.

[0137] Specifically, in step S5, the value function is updated through the following formula;

[0138] V(s) ← V(s) + α * (r + γ * V(s′) - V(s))

[0139] Among them, V(s) is the value function of the current task section;

[0140] V(s′) is the value function of the next task section;

[0141] r is the reward value;

[0142] α is the learning rate;

[0143] γ is the discount factor.

[0144] During the update process of the policy function table, the present invention adopts the same temporal difference error as the update of the value function table. Specifically, this error is multiplied by the learning rate and then added to the probability of executing the current action in the current state.

[0145] Specifically, in step S5, the policy function table is updated through the following formula;

[0146] π(a|s) ← π(a|s) + α * (r + γ * V(s′) - V(s))

[0147] Among them, π(a|s) is the probability of selecting the a-th sub-task in the s-th task section.

[0148] Through the above process, the policy function and the value function can be gradually improved, so that the reinforcement learning algorithm can better solve the problem.

[0149] The present invention adopts the combination of the softmax function and the random selection function for action selection.

[0150] The present invention uses the softmax function to process the probabilities in the policy function table. The softmax function can convert a set of numerical values into a probability distribution, so that these probabilities satisfy the properties of the probability distribution (that is, the sum of all probabilities is 1, and each probability is between 0 and 1).

[0151] The formula of the softmax function is:

[0152]

[0153] Among them, x is the input vector, w j and w k are the weight vectors, and K is the total number of categories.

[0154] Furthermore, the present invention uses a random selection algorithm (e.g., the random.choice function in the numpy library) to randomly select an action from the probability distribution processed by the softmax function. This random selection algorithm can randomly select an element according to the given probability distribution, thus ensuring that the probability of each action being selected is consistent with its probability in the policy function table. This method provides an effective way to balance the needs of exploration and exploitation, thereby achieving optimization in a complex decision-making environment.

[0155] This method combining the softmax function and the random selection function can not only ensure the randomness of the policy but also ensure the superiority of the policy, that is, the action with a larger value function has a higher probability of being selected, but the action with a smaller value function still has the possibility of being selected. This can balance exploration and exploitation to a certain extent, thereby improving the effect of reinforcement learning.

[0156] In reinforcement learning, the design of the reward function is very crucial. The reward function of the present invention takes into account multiple factors, including: the distance from the task section to the corresponding cluster center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse. This design of the reward function considering multiple factors makes our task division scheme more in line with the actual needs.

[0157] The goal of the greedy algorithm (greedy_shortest_path) is to calculate the optimal path cost of a given subtask. Using the greedy algorithm, each time it selects the task section with the optimal path cost among the unvisited subtask sections. The input of the function is a matrix corresponding to a subtask, and the output is the optimal path cost of the subtask.

[0158] Specifically, it includes the following steps:

[0159] Step A: Obtain all the task sections corresponding to the subtask and the connection relationships between the task sections, and regard each task section as a node;

[0160] Step B: Initialize a boolean list to mark whether the node has been visited;

[0161] Initialize the path length to 0 and set the starting node to 0;

[0162] Step C: Starting from the current node, find the next unvisited node with the shortest distance to reach it, and mark the current node as visited;

[0163] Traverse all unvisited nodes, calculate the distance from the current node to each unvisited node, and update the minimum distance and the next node until the shortest path is found through iteration;

[0164] Step D: Update the path length by adding the minimum distance to the path length.

[0165] The core idea of the above algorithm is to select the node with the shortest distance from the current node to the unvisited nodes as the next node to be visited each time until all nodes have been visited, thus obtaining the shortest path length of the entire subtask.

[0166] Furthermore, the Euclidean distance is used to calculate the distance from the task section to the corresponding cluster center.

[0167] The main steps for obtaining the Euclidean distance are as follows:

[0168] First, obtain two coordinates from the input parameters.

[0169] Each coordinate is a tuple containing two elements (x and y coordinates).

[0170] Then, calculate the differences between the two coordinates on the x-axis and y-axis.

[0171] Next, use the Euclidean distance formula to calculate the distance between the two coordinates.

[0172] The Euclidean distance formula is:

[0173]

[0174] where (x1, y1) and (x2, y2) are the two coordinates.

[0175] Finally, return the calculated distance.

[0176] The following is a detailed introduction to the reward calculation process:

[0177] In the present invention, the calculation of the reward covers multiple key factors, including the distance from the task section to the cluster center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse.

[0178] First, regarding the distance from the task section to the cluster center. This distance is calculated by the Euclidean distance formula, and the goal is to minimize this distance. Therefore, the obtained reward is the negative distance.

[0179] Second, the optimal path cost of the subtask. The optimal path cost of the subtask is obtained from the task sections of the same subtask, and a greedy algorithm is used to calculate the optimal path cost, with the goal of minimizing this length. Therefore, the obtained reward is the negative length.

[0180] Third, the shortest distance from the task section to the supply station, with the goal of minimizing this length. Therefore, the obtained reward is the negative distance.

[0181] Fourth, the distance from the task section to the warehouse. The goal is to make this distance as small as possible. Therefore, the obtained reward is the negative distance.

[0182] It should be emphasized that these rewards are all negative values because the goal of the present invention is to minimize these distances and lengths. In the framework of reinforcement learning, the strategy sought is to maximize the cumulative reward. Therefore, minimizing these distances and lengths is essentially equivalent to maximizing the cumulative reward.

[0183] Furthermore, the calculation formula of the reward function R is as follows:

[0184]

[0185] where x i is the coordinate or feature vector of the i-th task section;

[0186] c j is the coordinate of the j-th cluster center;

[0187] P is the optimal path length of the subtask calculated by the greedy algorithm;

[0188] n is the number of task sections of the current subtask;

[0189] is the shortest distance from the i-th task section to the supply station;

[0190] is the distance from the i-th task section to the warehouse.

[0191] dis(a, b) is the Euclidean distance, expressed as

[0192] The present invention has a wide range of application fields, including but not limited to path planning of urban driverless vehicles and logistics distribution problems. For each task section, select the id of a subtask. The basis for selecting the subtask is the comprehensive consideration of multiple factors, including: the distance from the task section to the corresponding cluster center, the cost-optimal path length corresponding to the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse. The present invention can find an optimal task division scheme while ensuring task efficiency, taking into account the distances from the task section to the supply station and the warehouse. This scheme can not only improve the working efficiency of urban sanitation vehicles, reduce energy consumption, but also optimize urban sanitation services and contribute to the cleanliness and environmental protection of the city.

[0193] The present invention can be implemented by means of technologies such as computer software and hardware. The specific operation steps can be programmed and implemented in an intelligent vehicle control system. In addition, through cloud services and other means, the optimization and management of job task division can be achieved. This diversified implementation method provides a flexible and efficient solution to meet different application requirements and environmental conditions.

[0194] The present invention can also adopt an evaluation method to evaluate the obtained task allocation scheme. The evaluation method is to assign each subtask to an urban unmanned sanitation vehicle and compare the indicators after path planning. This evaluation method can comprehensively reflect the effect of the task allocation scheme and provide a strong reference for further optimization.

[0195] The sample set adopted in the embodiment of the present invention includes a total of 3 job tasks in real scenarios, such as Figure 3 shown, including road network, warehouse location, location of supply stations, etc.

[0196] The specific experimental process is as follows:

[0197] I. Performance indicators

[0198] The embodiment of the present invention uses: the proportion of the total non-working path; the number of vehicles and the proportion of the non-working path per 1 km traveled by each unmanned vehicle; the total system cost: the optimal number of unmanned vehicles required to complete 1 km of work tasks and the total cost required for the corresponding optimal path.

[0199]

[0200] Among them:

[0201] F is the fixed travel cost of the unmanned sanitation vehicle, and the value range is 0 - 1000;

[0202] Q is the number of unmanned vehicles participating in the work task, that is, the number of subtasks;

[0203] β is the adjustment coefficient, generally taking a value of 10000;

[0204] w is the energy consumption required per unit working distance, generally taking a value of 1;

[0205] r is the energy consumption ratio of the transfer mode relative to the working mode, and the value range is 0 - 1;

[0206] Path w is the working path;

[0207] Path r is the transfer path.

[0208] II. Comparative experiment

[0209] The following methods are adopted in this embodiment for comparative evaluation:

[0210] K-means_ant: This method uses the k-means algorithm as the task partitioning part and the ant colony algorithm to find the optimal path.

[0211] K-means_greedy: This method uses the K-means algorithm for task allocation and the greedy algorithm to find the optimal path.

[0212] Naive_greedy: This method is a greedy algorithm without task allocation.

[0213] Actor_Critic: The method of the present invention.

[0214] Figure 4 is the proportion of the total non-working path when completing the operation 10* in three real test scenarios in the embodiment of the present invention, and the schematic diagram of the number of vehicles.

[0215] Figure 5 is the proportion of the non-working path per 1 km traveled by each unmanned vehicle in three real test scenarios in the embodiment of the present invention. The scatter plot in the figure shows the proportion of the non-working path of a single unmanned vehicle corresponding to different algorithms. It can be seen that the Actor-Critic model also has a good performance in the proportion of the non-working path of a single vehicle.

[0216] Figure 6 The total cost (Cost) in refers to the total cost of the optimal number of unmanned vehicles (Q) required to complete a 1-kilometer work task and its corresponding optimal path.

[0217] From Figures 4 - 6 it can be found that the present invention performs best in the following four indicators:

[0218] Indicator 1: The proportion of the total non-working path, that is, the ratio of the non-working path to the total path.

[0219] Indicator 2: The number of vehicles, that is, the number of unmanned vehicles participating in the work task.

[0220] Indicator 3: The proportion of the non-working path per 1 km traveled by each unmanned vehicle, that is, the proportion of the non-working path in the 1-kilometer distance traveled by each unmanned vehicle.

[0221] Indicator 4: The optimal number of unmanned vehicles required to complete a 1-KM work task and the total cost required for the corresponding optimal path, that is, the total cost of the optimal number of unmanned vehicles required to complete a 1-kilometer work task and its corresponding optimal path.

[0222] The method proposed by the present invention performs excellently in various indicators, proving its effectiveness and superiority in practical applications.

[0223] In addition, an embodiment of the present invention also discloses a task intelligent partitioning device for collaborative operation of unmanned sanitation vehicles, including:

[0224] A clustering module, configured to cluster all task sections by using a clustering algorithm, where the number of clusters is a preset number of subtasks, to obtain an initial clustering center;

[0225] An initialization module, configured to initialize a policy function table and a value function table;

[0226] The policy function table stores the probability of each task section selecting each subtask;

[0227] The value function table stores the value function of each task section, and the value function is the expected return that can be obtained when selecting the corresponding subtask;

[0228] A selection module, configured to select a corresponding subtask for each task section according to the current policy function table;

[0229] A reward value calculation module, configured to calculate a reward value by using a reward function according to the selection result, and evaluate the expected return of the selection result;

[0230] The reward function involves one or more factors among the distance from the task section to the corresponding clustering center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse;

[0231] An update module, configured to update the value function and the policy function of the current task section based on the update rule of Actor-Critic according to the reward value and the value function of the next task section;

[0232] A repetition module, configured to repeatedly execute the methods in the selection module, the reward value calculation module, and the update module to update the policy function table and the value function table until an optimal subtask is assigned to each task section.

[0233] Furthermore, an optimal path cost calculation unit calculates the optimal path cost of the subtask in the reward function by using a greedy algorithm, including:

[0234] An acquisition unit, configured to acquire all task sections corresponding to the subtask and the connection relationship between the task sections, and regard each task section as a node;

[0235] A greedy algorithm initialization unit, configured to initialize a boolean list for marking whether a node has been visited;

[0236] Initialize the path length to 0, and set the starting node to 0;

[0237] Iterative search unit, which is used to start from the current node, find the next unvisited node to minimize the distance to reach this node, and mark the current node as visited;

[0238] Traverse all unvisited nodes, calculate the distance from the current node to each unvisited node, and update the minimum distance and the next node until the shortest path is iteratively found;

[0239] Path update unit, which is used to update the path length by adding the minimum distance to the path length.

[0240] Furthermore, the Euclidean distance is used to calculate the distance from the task section to the corresponding cluster center.

[0241] Furthermore, the calculation formula of the reward function R is as follows:

[0242]

[0243] where x i is the coordinate or feature vector of the i-th task section;

[0244] c j is the coordinate of the j-th cluster center;

[0245] P is the optimal path length of the subtask calculated by the greedy algorithm;

[0246] n is the number of task sections of the current subtask;

[0247] is the shortest distance from the i-th task section to the supply station;

[0248] is the distance from the i-th task section to the warehouse.

[0249] Furthermore, in step S5, the value function is updated by the following formula;

[0250] V(s)←V(s)+α*(r+γ*V(s′)-V(s))

[0251] where V(s) is the value function of the current task section;

[0252] V(s′) is the value function of the next task section;

[0253] r is the reward value;

[0254] α is the learning rate;

[0255] γ is the discount factor;

[0256] The policy function table is updated by the following formula;

[0257] π(a|s) ← π(a|s) + α * (r + γ * V(s′) - V(s))

[0258] Among them, π(a|s) is the probability of selecting the a-th subtask under the s-th task section.

[0259] Furthermore, the softmax function is used to process the probabilities in the policy function table.

[0260] In this embodiment, the specific content of the task intelligent partitioning device for collaborative operation of unmanned sanitation vehicles is similar to the content of the task intelligent partitioning method for collaborative operation of unmanned sanitation vehicles disclosed in the above embodiment, and will not be elaborated here.

[0261] The embodiment of the present invention also discloses a storage medium, which stores one or more computer-readable programs. The one or more programs include instructions, and the instructions are adapted to be loaded and executed by the memory to perform the task intelligent partitioning method for collaborative operation of unmanned sanitation vehicles disclosed in any of the above embodiments.

[0262] The present invention discloses a task intelligent partitioning method for collaborative operation of unmanned sanitation vehicles, which has the following beneficial effects:

[0263] The present invention divides the task section into multiple subtasks and uses the Actor-Critic reinforcement learning algorithm to optimize the partitioning method of each subtask, thereby realizing the optimization of the partitioning of the entire task section. At the same time, the present invention comprehensively considers factors such as the distance from the task section to the corresponding clustering center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse.

[0264] By using the Actor-Critic reinforcement learning model and combining with greedy algorithms, clustering algorithms, etc., the present invention effectively partitions the tasks of urban sanitation vehicles, can solve the problems of route planning and task allocation when urban sanitation vehicles perform tasks such as cleaning and garbage collection, improve work efficiency, reduce energy consumption, and optimize urban sanitation services.

[0265] It should be understood that the various technologies described here can be implemented in combination with hardware or software, or a combination of them. Thus, the method and device of the present invention, or certain aspects or parts of the method and device of the present invention, can take the form of program code (i.e., instructions) embedded in a tangible medium, such as a floppy disk, a CD-ROM, a hard disk drive, or any other machine-readable storage medium, where when the program is loaded into a machine such as a computer and executed by the machine, the machine becomes a device for practicing the present invention.

[0266] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. An intelligent task division method for cooperative operation of unmanned sanitation vehicles, characterized in that: include: Step S1: cluster all task sections using a clustering algorithm, the number of clusters being the preset number of subtasks, and obtaining the initial cluster center; Step S2: Initialize the strategy function table and the value function table; The strategy function table stores the probability of selecting each subtask for each task section; The value function table stores the value function of each task section, and the value function is the expected return that can be obtained when the corresponding subtask is selected; Step S3: Select a corresponding subtask for each task section according to the current strategy function table; Step S4: Calculate the reward value using the reward function according to the selection result, and evaluate the expected return of the selection result; The reward function involves one or more factors of the distance from the task section to the corresponding cluster center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse; Step S5: Based on the Actor-Critic update rule, the value function and strategy function of the current task segment are updated according to the reward value and the value function of the next task segment; Step S6: Repeat steps S3 to S5 to update the strategy function table and the value function table until the optimal subtask is assigned to each task section; The greedy algorithm is used in the reward function to calculate the optimal path cost of the subtask, including the following steps: Step A: Obtain all task sections corresponding to the subtasks and the connection relationship between the task sections, and regard each task section as a node; Step B: Initialize a Boolean list to mark whether a node has been visited; Initialize the path length to 0 and set the starting node to 0; Step C: Starting from the current node, find the next unvisited node so that the distance to the node is the shortest, and mark the current node as visited; Traverse all unvisited nodes, calculate the distance from the current node to each unvisited node, and update the minimum distance and the next node until the shortest path is found iteratively; Step D: Update the path length by adding the minimum distance to the path length.

2. The task intelligent division method according to claim 1, characterized in that: The Euclidean distance is used to calculate the distance from the task section to the corresponding cluster center.

3. The task intelligent division method according to claim 1, characterized in that: The calculation formula of the reward function R is as follows: Among them, x i is the coordinate or feature vector of the i-th task section; c j is the coordinate of the jth cluster center; P is the optimal path length of the subtask calculated by the greedy algorithm; n is the number of task sections of the current subtask; is the shortest distance from the i-th mission section to the supply station; is the distance from the i-th task section to the warehouse.

4. The task intelligent division method according to claim 1, characterized in that: In step S5, the value function is updated by the following formula; V(s)←V(s)+α*(r+γ*V(s′)-V(s)) Among them, V(s) is the value function of the current task section; V(s′) is the value function of the next task segment; r is the reward value; α is the learning rate; γ is the discount factor; Update the strategy function table by the following formula; π(a|s)←π(a|s)+α*(r+γ*V(s′)-V(s)) Among them, π(a|s) is the probability of selecting the ath subtask under the sth task segment.

5. The task intelligent division method according to claim 4, characterized in that: Use the softmax function to process the probabilities in the policy function table.

6. Intelligent task division equipment for unmanned sanitation vehicle collaborative operation, characterized in that: include: The clustering module is used to cluster all task sections using a clustering algorithm. The number of clusters is the preset number of subtasks to obtain the initial cluster center. Initialization module, used to initialize the strategy function table and value function table; The strategy function table stores the probability of selecting each subtask for each task section; The value function table stores the value function of each task section, and the value function is the expected return that can be obtained when the corresponding subtask is selected; A selection module is used to select a corresponding subtask for each task section according to the current strategy function table; A reward value calculation module is used to calculate the reward value based on the selection result using the reward function and evaluate the expected return of the selection result; The reward function involves one or more factors of the distance from the task section to the corresponding cluster center, the optimal path cost of the subtask, the shortest distance from the task section to the supply station, and the distance from the task section to the warehouse; The update module is used to update the value function and strategy function of the current task segment based on the Actor-Critic update rule according to the reward value and the value function of the next task segment; A repetition module, used to repeatedly execute the methods in the selection module, the reward value calculation module, and the update module, and update the strategy function table and the value function table until the optimal subtask is assigned to each task section; The optimal path cost calculation unit uses a greedy algorithm to calculate the optimal path cost of the subtask in the reward function, including: An acquisition unit, used to acquire all task sections corresponding to the subtasks and the connection relationship between the task sections, and regard each task section as a node; Greedy algorithm initialization unit, used to initialize the Boolean list to mark whether the node has been visited; Initialize the path length to 0 and set the starting node to 0; Iterative search unit, used to start from the current node, find the next unvisited node so that the distance to the node is the shortest, and mark the current node as visited; Traverse all unvisited nodes, calculate the distance from the current node to each unvisited node, and update the minimum distance and the next node until the shortest path is found iteratively; A path updating unit is used to update the path length by adding the minimum distance to the path length.

7. The task intelligent division device according to claim 6, characterized in that: The Euclidean distance is used to calculate the distance from the task section to the corresponding cluster center.

8. The task intelligent division device according to claim 6, characterized in that: The calculation formula of the reward function R is as follows: Among them, x i is the coordinate or feature vector of the i-th task section; c j is the coordinate of the jth cluster center; P is the optimal path length of the subtask calculated by the greedy algorithm; n is the number of task sections of the current subtask; is the shortest distance from the i-th mission section to the supply station; is the distance from the i-th task section to the warehouse.

9. The task intelligent division device according to claim 6, characterized in that: The update module updates the value function by the following formula; V(s)←V(s)+α*(r+γ*V(s′)-V(s)) Among them, V(s) is the value function of the current task section; V(s′) is the value function of the next task segment; r is the reward value; α is the learning rate; γ is the discount factor; Update the strategy function table by the following formula; π(a|s)←π(a|s)+α*(r+γ*V(s′)-V(s)) Among them, π(a|s) is the probability of selecting the ath subtask under the sth task segment.

10. The task intelligent division device according to claim 9, characterized in that: Use the softmax function to process the probabilities in the policy function table.

11. A storage medium, characterized in that The storage medium stores one or more computer-readable programs, and the one or more programs include instructions, which are suitable for being loaded by the memory and executing the method for intelligent task division of unmanned sanitation vehicle collaborative operation as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Smart city street management method, Internet of Things system, device and storage medium

    CN115481987A

  • Mechanical arm path planning method and system based on depth deterministic strategy gradient

    CN116494247A