Crowd sensing task scheduling method and system based on multi-space modeling and fairness reinforcement learning

By adopting the task scheduling method of multi-space modeling and fairness reinforcement learning in the multi-space group intelligence perception system, the coupling relationship and collaborative scheduling requirements in multi-space task allocation are solved, and the task completion rate and fair utilization of resources are achieved.

CN120069491AActive Publication Date: 2025-05-30YANTAI UNIV

Patent Information

Application Number
CN202510561102.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing multi-space group intelligence-aware task allocation method is difficult to effectively handle the coupling relationship of multi-space tasks and the coordinated scheduling requirements between tasks, resulting in uneven resource allocation, low task completion rate and reduced overall system efficiency.

Method used

The task scheduling method based on multi-space modeling and fairness reinforcement learning is adopted, and efficient task allocation and fairness guarantee are achieved through multi-space area division, task map construction and critical path recognition, agent matching strategies and fairness-aware multi-objective scheduling algorithm.

Benefits of technology

It improves the task completion rate, scheduling efficiency and fairness of agent participation in the group intelligence perception system, can effectively identify task hotspots and spatial density differences, and optimize resource utilization and task allocation paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069491A_ABST
    Figure CN120069491A_ABST
Patent Text Reader

Abstract

The invention relates to the crossing field of the crowd sensing technology and the artificial intelligence technology, in particular to a crowd sensing task scheduling method and system based on multi-space modeling and fairness reinforcement learning, and the method comprises the steps: constructing a multi-layer space task structure of ground and air cooperation according to an obtained task; establishing a topological correlation model between tasks according to a space region division result, and determining a priority sequence of task scheduling through a critical path extraction mechanism; constructing a differentiated agent matching strategy, and performing reasonable configuration of a critical path task based on a priority sequence; performing dynamic allocation on non-critical path tasks by utilizing a fairness perception multi-target scheduling algorithm; according to the method, the search space is reduced through a space division and path screening mechanism, the task completion rate and the overall scheduling performance of the system are improved, and the method is suitable for efficient resource allocation requirements in a multi-space collaborative awareness task scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross - field of crowd - sensing and artificial intelligence optimization technology, and particularly to a crowd - sensing task scheduling method and system based on multi - space modeling and fairness reinforcement learning. Background Art

[0002] With the rapid development of information technology and the wide application of intelligent devices, crowd - sensing has become an important means for large - scale data collection and intelligent analysis, and is widely used in fields such as environmental monitoring, urban management, public safety, and traffic sensing. Traditional crowd - sensing task allocation is mostly based on a two - dimensional plane model, which is difficult to meet the actual needs of strong task heterogeneity and large dynamic changes in resources in the three - dimensional urban space. Especially in the multi - space coupling environment facing the ground, air, and underground, there are a variety of tasks with different priorities, and there are also significant differences in the capabilities of sensing devices, network bandwidth, and energy constraints, which pose higher requirements for task allocation.

[0003] Most of the existing task allocation methods focus on a single - space level or static environment settings, ignoring the coupling relationship of multi - space tasks and the collaborative scheduling requirements between tasks, which easily leads to problems such as uneven resource allocation, low task completion rate, and decreased overall system efficiency. In addition, in the multi - space sensing environment, the priority scheduling and fairness guarantee of tasks also face huge challenges. Traditional methods often fail to balance the task urgency and the resource load of agents, resulting in the over - use of high - load agents and the waste of resources of low - load agents.

[0004] On the other hand, although some studies have tried to introduce multi - objective optimization algorithms, reinforcement learning, and graph - structure modeling to improve task allocation efficiency, in the face of the scenario of multi - space and multi - level dynamic task flows, there is still a lack of systematic modeling and optimization of space - level division, task - group collaborative scheduling, critical - path identification, and long - term fairness constraints.

[0005] Therefore, there is an urgent need for a fairness - guaranteed multi - space crowd - sensing task allocation method and system that can integrate space clustering, reinforcement learning, and multi - objective optimization mechanisms, which can not only achieve the efficiency of task allocation but also ensure the fairness of task execution and the rationality of resource utilization to adapt to the increasingly complex crowd - sensing application scenarios. Summary of the Invention

[0006] To solve the problems existing in the task allocation process of existing multi-space crowdsensing, such as complex task space structure, frequent dynamic changes of resources, strong task heterogeneity, low task scheduling efficiency, and poor fairness of agent allocation, the present invention provides a crowdsensing task scheduling method and system based on multi-space modeling and fairness reinforcement learning, which can efficiently achieve task partitioning, critical path identification, and fairness-aware scheduling in a multi-layer space environment, and improve the task completion rate, scheduling efficiency, and fairness of agent participation of the crowdsensing system.

[0007] In the first aspect, a crowdsensing task scheduling method based on multi-space modeling and fairness reinforcement learning provided by the present invention adopts the following technical solutions: A crowdsensing task scheduling method based on multi-space modeling and fairness reinforcement learning includes: In the multi-space region partitioning stage, a multi-layer space task structure of ground-air collaboration is constructed according to the obtained tasks; In the task graph construction and critical path identification stage, a topological association model between tasks is established according to the space region partitioning result, and the priority sequence of task scheduling is clarified through a critical path extraction mechanism; In the task allocation stage, a differentiated agent matching strategy is constructed, and the critical path tasks are reasonably configured based on the priority sequence; In the reinforcement learning stage, a fairness-aware multi-objective scheduling algorithm is used to dynamically allocate non-critical path tasks; A dynamic fairness evaluation mechanism is used to evaluate and optimize the task allocation result.

[0008] Further, the construction of the multi-layer space task structure of ground-air collaboration according to the obtained tasks includes performing density clustering processing on the ground task area, using the DBSCAN algorithm to identify task hotspots, and dividing the ground sensing tasks into several spatially autonomous task clusters; subsequently, for the air task area, spatial stratification is performed according to the height information of the tasks, the airspace is divided into multiple levels of low altitude, medium altitude, and high altitude, and a quadtree structure is introduced in each layer to recursively divide the planar area to achieve hierarchical management and refined expression of air tasks.

[0009] Further, the construction of the multi-layer space task structure of ground-air collaboration according to the obtained tasks also includes introducing a Voronoi diagram construction mechanism, generating a space control area based on the spatial centroids of the divided task groups, and automatically mapping the unassigned tasks to the corresponding space task units according to the nearest centroid principle; for collaborative tasks that involve both ground and air execution requirements, a hybrid Voronoi diagram is constructed to achieve cross-layer coupling mapping, ensuring the consistency and integrity of the task space organization.

[0010] Furthermore, a topological association model between tasks is established according to the spatial region division result, and the priority sequence of task scheduling is clarified through a critical path extraction mechanism, including constructing a task area graph based on the spatially divided task groups. The task area graph is a weighted undirected graph, where nodes correspond to task groups, and the weight of the edge is jointly determined by the time urgency difference and spatial distance between task groups; calculating an urgency score for each task, where the urgency score is defined based on the time window of the task; and constructing a group urgency index for each task group, and comprehensively reflecting the scheduling priority of the task group by weighted aggregation of the urgency of tasks within the group.

[0011] Furthermore, establishing the topological association model between tasks according to the spatial region division result and clarifying the priority sequence of task scheduling through the critical path extraction mechanism further includes dynamically generating the association weight of the edge in the graph by combining the earliest start time difference and cross-space movement cost between task groups to characterize the dependency relationship and cost information between tasks; where the nodes of the Voronoi diagram are used as the vertices of the graph to construct a task group network, where each seed node corresponds to a task group, and when determining the edge weight value, all critical paths in the task group network are found by dynamically adjusting the weight value and combining the NetworkX framework.

[0012] Furthermore, constructing a differentiated agent matching strategy and reasonably configuring the critical path tasks based on the priority sequence includes using a greedy matching algorithm and a fairness-aware multi-objective reinforcement learning algorithm FA-PMOAQL as complementary scheduling mechanisms. For the tasks in the critical path, based on the pre-constructed task scheduling priority sequence, a greedy strategy is used to first match the most suitable agent node for it, and the optimization objectives during the matching process are to minimize the spatial distance between the task and the agent and maximize the device capacity matching degree.

[0013] Furthermore, using the fairness-aware multi-objective scheduling algorithm to dynamically allocate non-critical path tasks includes continuously iteratively updating the task-agent allocation strategy by constructing a state-action mapping, combining pheromone concentration and value function estimation to ensure the dynamic convergence of the multi-objective optimization objective. Among them, task allocation is modeled as a Markov decision process, and an ant colony optimization mechanism is introduced to enhance local search and global exploration capabilities. For Q learning, task allocation is modeled as a state-action process, and the Q table is updated using the Bellman equation.

[0014] Furthermore, using the fairness-aware multi-objective scheduling algorithm to dynamically allocate non-critical path tasks also includes dynamically adjusting the pheromone concentration for ant colony optimization and using the Pareto pheromone update criterion: only when the path belongs to the Pareto optimal solution set, its pheromone concentration is increased, and at the same time pheromone evaporation occurs; finally, the allocation probability of the task is calculated according to the pheromone concentration and Q value.

[0015] Further, the use of the dynamic fairness evaluation mechanism for task assignment result evaluation and optimization includes, after each round of assignment is completed, calculating the task load of the agents, calculating the Gini coefficient and the load variance based on the task load, and using them as fairness indicators to evaluate the scheduling fairness; and embedding the fairness indicators as part of the reward function into the reinforcement learning model to achieve task assignment fairness while maintaining task completion efficiency.

[0016] In a second aspect, a crowdsensing task scheduling system based on multi-space modeling and fairness reinforcement learning includes: A multi-space region division module, configured to, in the multi-space region division stage, construct a multi-layer space task structure for ground-air cooperation according to the obtained tasks; A critical path identification module, configured to, in the task graph construction and critical path identification stage, establish a topological association model between tasks according to the space region division result, and clarify the priority sequence of task scheduling through a critical path extraction mechanism; A task assignment module, configured to, in the task assignment stage, construct a differentiated agent matching strategy and reasonably allocate critical path tasks based on the priority sequence; A reinforcement learning module, configured to, in the reinforcement learning stage, use a fairness-aware multi-objective scheduling algorithm to dynamically allocate non-critical path tasks; An evaluation module, configured to use a dynamic fairness evaluation mechanism to evaluate and optimize the task assignment result.

[0017] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the above-mentioned crowdsensing task scheduling method based on multi-space modeling and fairness reinforcement learning.

[0018] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor to perform the fairness-guaranteed multi-space crowdsensing task assignment dynamic cooperation optimization method.

[0019] In summary, the present invention has the following beneficial technical effects: 1. The present invention constructs a space task region division mechanism for ground-air cooperation through the MCQVP algorithm, and uses DBSCAN clustering and quad-tree structure to finely model the multi-space task environment, which can effectively identify task hotspots and spatial density differences, realize spatial autonomy and regional separation of task grouping, and provide basic support for subsequent task graph modeling and cooperative scheduling.

[0020] 2. The GUMSP critical path recognition algorithm proposed by the present invention comprehensively considers task priorities and time urgency, and can efficiently identify multiple critical scheduling paths in the constructed task area graph, ensuring that important tasks are preferentially completed under resource-constrained conditions, and significantly improving the overall task completion rate and scheduling timeliness.

[0021] 3. The fairness-aware multi-objective scheduling algorithm FA-PMOAQL that combines reinforcement learning and ant colony optimization is used in the present invention to construct a multi-objective model centered on device matching degree, distance optimization, and fairness threshold guarantee, realizing the dynamic allocation of tasks and agents. This method effectively improves the resource utilization rate of agents, optimizes the task execution path, and at the same time improves the stability and scalability of the scheduling system.

[0022] 4. The present invention embeds fairness indicators such as the Gini coefficient and agent load variance in the task scheduling process, constructs a fairness reward function and embeds it into the learning process, so that task allocation takes into account the fair participation degree among multiple agents while ensuring efficiency, effectively preventing the situation of agent overload or idleness, and improving the overall sustainable operation ability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic diagram of the overall system framework of the fairness-guaranteed multi-space crowd intelligence perception task allocation method based on multi-objective collaborative optimization in the embodiment of the present invention; Figure 2 It is a schematic diagram of task area space division and multi-layer space modeling in the embodiment of the present invention; Figure 3 It is a schematic diagram of the process of constructing a task area graph and identifying critical paths in the embodiment of the present invention; Figure 4 It is a schematic diagram of multi-space task distribution in the embodiment of the present invention; Figure 5 It is a graph of the comparative experimental results of different algorithms under the task completion rate index as the number of agents changes in the embodiment of the present invention; Figure 6 It is a graph of the comparative experimental results of different algorithms under the task completion rate index as the number of tasks changes in the embodiment of the present invention; Figure 7 It is a graph of the comparative experimental results of different algorithms under the running time index in the embodiment of the present invention; Figure 8 It is a graph of the comparative experimental results of different algorithms under the resource utilization rate index in the embodiment of the present invention; Figure 9 It is a graph of the comparative experimental results of different algorithms under the fairness index in the embodiment of the present invention; Figure 10 It is a graph of the comparative experimental results of different algorithms under the cross-layer task completion rate index in the embodiment of the present invention; Figure 11 This is a comparison experiment result graph of different algorithms in the average task latency index in the embodiments of the present invention. Specific implementation manners

[0024] The present invention will be further described in detail below with reference to the accompanying drawings.

[0025] Embodiment 1 Referring to Figure 1 , a crowdsensing task scheduling method based on multi-space modeling and fairness reinforcement learning in this embodiment includes: Multi-space region division stage, constructing a multi-layer space task structure for ground-air cooperation according to the obtained tasks; Task graph construction and critical path identification stage, establishing a topological association model between tasks according to the space region division result, and clarifying the priority sequence of task scheduling through a critical path extraction mechanism; Task allocation stage, constructing a differential agent matching strategy, and reasonably allocating critical path tasks based on the priority sequence; Reinforcement learning stage, using a fairness-aware multi-objective scheduling algorithm to dynamically allocate non-critical path tasks; Using a dynamic fairness evaluation mechanism to evaluate and optimize the task allocation result.

[0026] As Figure 1 shown, the present invention designs a crowdsensing task scheduling framework based on multi-space modeling and fairness reinforcement learning. The system introduces three core links of space modeling, path identification and intelligent scheduling in the scheduling framework, corresponding to the space organization, priority ranking and dynamic resource allocation of tasks respectively. In the space organization stage, the platform identifies ground hot spots through the DBSCAN clustering algorithm, and constructs a multi-space collaborative sensing environment in combination with the three-layer structure division of air tasks; in the path identification stage, the platform constructs a task area graph according to the earliest start time and priority of tasks, and identifies critical path tasks to give priority to emergency scheduling.

[0027] In the intelligent scheduling link, the platform allocates critical path tasks to the agent node with the strongest device ability and the optimal distance, while non-critical tasks are allocated using a fairness-aware multi-objective optimization model FA-PMOAQL that combines Q-Learning and ant colony optimization algorithms. In this optimization process, the system takes the minimum task allocation distance, the maximum device matching degree and the minimum agent load variance as the objective function, and comprehensively considers factors such as task location, device requirements, and agent resource status to construct a task allocation candidate set. The device matching degree of its task-agent is through Calculations are performed, and the platform jointly constructs the allocation probability based on the Q-value and pheromone concentration between tasks and agents, and dynamically generates the task allocation result.

[0028] During the scheduling execution process, to ensure the overall fairness of the system, the platform introduces statistical indicators such as the Gini coefficient and load variance to construct a fairness reward function, and embeds it into the reinforcement learning model to achieve a dynamic balance between task efficiency and agent participation fairness. Finally, the platform outputs a set of multi-objective optimal solutions as the task-agent allocation strategy, and adjusts the subsequent scheduling strategy according to the historical task status and feedback results to maintain the stability, efficiency, and fairness of the system in a multi-space high-dynamic perception scenario.

[0029] The specific steps are as follows: Embodiment 1 of the present invention provides a crowdsensing task scheduling method based on multi-space modeling and fairness reinforcement learning, including a multi-space region division stage, a task graph construction and critical path identification stage, a task allocation stage, a reinforcement learning and optimization algorithm implementation, and a fairness evaluation mechanism.

[0030] Multi-space region division stage: In the multi-space region division stage, a multi-layer space task structure for ground-air cooperation is constructed through the MCQVP algorithm to achieve the spatial orderly organization and efficient management of crowdsensing tasks. Specifically, first, density clustering processing is performed on the ground task area, and the DBSCAN algorithm is used to identify the task hotspots, and the ground sensing tasks are divided into several spatially autonomous task clusters; subsequently, for the air task area, according to the height information of the tasks, spatial stratification is performed, and the airspace is divided into multiple layers of low altitude, medium altitude, and high altitude, and a quadtree structure is introduced in each layer to recursively divide the planar area to achieve hierarchical management and refined expression of air tasks. To further improve the mapping accuracy between tasks and spatial units, the system introduces a Voronoi diagram construction mechanism, generates a spatial control area based on the spatial centroids of the divided task groups, and automatically maps the unassigned tasks to the corresponding spatial task units according to the nearest centroid principle; for collaborative tasks that involve both ground and air execution requirements, a hybrid Voronoi diagram is constructed to achieve cross-layer coupling mapping to ensure the consistency and integrity of the task space organization.

[0031] Step 1: For the ground area: Let the historical task data set be , where each historical task contains three-dimensional spatial coordinates , but its is 0. The distance between task points is measured by the Euclidean distance, defined as: , and then the neighborhood of any point is . If the point Within the neighborhood of (minimum sample number), then is defined as a core point. The generation of clusters starts from the core point, and all points within its ε-neighborhood (including boundary points) are assigned to the same cluster, ultimately forming a dense hot spot area . The task set of the hot spot area is defined as: , where represents the total number of clusters identified by DBSCAN, and each cluster represents a ground hot spot area.

[0032] Step 2: For the airspace: Stratify based on the height information of the tasks. Let the three-dimensional coordinates of each air task be , where represents the height of the task. Stratify the airspace according to the height, into the low-altitude layer ( ), the mid-altitude layer ( ), and the high-altitude layer ( ). Within each layer, further divide the two-dimensional planar area using the quadtree structure. Let the bounding box of area A within a certain layer be . When the number of tasks within the area and the current division depth , divide B evenly into four sub-areas . For each sub-area, recursively apply the same division rule until is satisfied.

[0033] Step 3: Current task assignment and Voronoi diagram construction: Let the current batch of task set be , where each task . For the tasks within the divided area, for each task , if it satisfies the boundary condition of a certain area , that is , then directly assign it to that area. For the tasks that cannot be assigned to the predefined area, use the Voronoi diagram method based on the centroid of the task group. First, for each determined task group , calculate the centroid . Then construct the Voronoi diagram , divide the entire space into several areas, and each area is defined as: . For each unassigned task , let , and assign to the task group corresponding to .

[0034] Step 4: For collaborative tasks that involve both ground and air operations (such as "drone-vehicle" relay delivery): Let the collaborative task set be , a cross-layer task processing method is adopted to construct a Hybrid Voronoi Diagram (HVD). It is required that for any , if there are collaborative constraints, it is forced to be satisfied: .

[0035] Task graph construction and critical path identification phase: The GUMSP algorithm is used in the task graph construction and critical path identification phase to establish a topological association model between tasks based on the spatial partitioning results, and the priority sequence of task scheduling is clarified through the critical path extraction mechanism, so as to ensure the urgency response ability and global collaboration efficiency of the overall system scheduling. Specifically, first, according to the spatial task groups formed by the aforementioned partitioning, a task area graph is constructed. The task area graph is a weighted undirected graph, where each node in the graph corresponds to a task group, and the weight of the edge is jointly determined by the time urgency difference and spatial distance between task groups. Further, the system calculates the urgency score (US) for each task, and this score is defined based on the time window of the task; and a Group Urgency (GU) index is constructed for each task group, and the urgency of tasks within the group is weighted and aggregated to comprehensively reflect the scheduling priority of the task group. At the same time, combining the earliest start time difference between task groups and the cross-space movement cost, the associated weight of the edge in the graph is dynamically generated to characterize the dependency relationship and cost information between tasks.

[0036] Step 5: Calculate the urgency score (Urgency Score, US) of each task. The calculation formula is: , where represents the end time of the task, represents the start time of the task.

[0037] Step 6: Calculate the Group Urgency (GU) of the task group, which is obtained by weighted aggregating the urgency scores of all tasks within the task group. The calculation formula is: , where is the weight based on the number of tasks in the task group, ensuring that the influence weight of the case with a large number of tasks within the task group on the urgency is relatively high.

[0038] Step 7: Calculate the weight value in the task group network. The edge weight is defined based on the time difference and urgency between task groups. The calculation formula is: , where represents the time difference between the earliest start tasks of two task groups, is the cross-layer movement cost (such as the energy consumption of the drone's lifting and lowering). And is the source task group and the target task group Average urgency level: .

[0039] Step 8: Find Start Regions and End Regions. Start Regions: The set of task groups with the highest urgency level, which is defined as: . End Regions: The set of task groups with the lowest urgency level, which is defined as: .

[0040] Step 9: Construct a task group network with the nodes of the Venn diagram as the vertices of the graph, where each seed node corresponds to a task group. When determining the edge weights, dynamically adjust the weights through the above formula and, in combination with the NetworkX framework, find all the critical paths in the task group network.

[0041] Task allocation phase: In the task allocation phase, for the identified critical path tasks and ordinary tasks, construct a differentiated agent matching strategy, comprehensively consider multiple factors such as geographical distance, device capacity matching degree, and scheduling fairness, etc., to achieve the efficient and reasonable allocation of task resources. This phase distinguishes between critical path tasks and non-critical path tasks and adopts two complementary scheduling mechanisms: the greedy matching algorithm and the fairness-aware multi-objective reinforcement learning algorithm FA-PMOAQL. Specifically, for the tasks in the critical path, the system, based on the pre-constructed task scheduling priority sequence, adopts a greedy strategy to preferentially match the most suitable agent node for them. During the matching process, the optimization objectives are to minimize the spatial distance between the task and the agent and maximize the device capacity matching degree. For non-critical path tasks, the fairness-aware multi-objective reinforcement learning scheduling algorithm FA-PMOAQL proposed by the present invention performs dynamic matching. This algorithm takes minimizing the total distance cost of task allocation, maximizing the device matching degree, and meeting the preset fairness threshold as the joint optimization objectives, constructs a task-agent allocation probability model, and iteratively generates the optimal task allocation strategy according to the reinforcement learning framework.

[0042] Step 10: The platform calculates the following metrics for the task groups on the critical path to select the optimal agent: Geographical distance calculation: , where is the location information of the agent and the task group. Device matching degree calculation: , where K is the number of sensors possessed by the agent, p is the position of the sensors required by the task in the agent sensor list (1 means the most suitable for this agent), α is a positive number used to penalize tasks that do not conform to the sensor list, HeightMatch is the matching degree between the aerial agent height capability and the task height, and β is the height matching weight coefficient.

[0043] Step 11: Perform priority task allocation, and preferentially allocate the agent that best meets the device requirements for the task to minimize the distance between the agent to the task . The priority of task allocation is determined by the optimization function: , where represents the distance between agent and task , is the device matching score of agent for task type .

[0044] Step 12: For non-critical path tasks, use the fairness-aware multi-objective reinforcement learning scheduling algorithm FA-PMOAQL for dynamic scheduling. This algorithm jointly optimizes three objectives: minimizing the task distance cost ; maximizing the device matching degree ; ensuring that the fairness threshold is met, where represents the total travel distance of the agent, represents the device matching score between the agent and the task, represents the fairness index, is the fairness threshold.

[0045] Reinforcement learning and optimization algorithm implementation: The reinforcement learning and optimization algorithm implementation stage mainly focuses on the dynamic allocation problem of non-critical path tasks, and proposes a fairness-aware multi-objective scheduling algorithm that combines ant colony optimization and Q-learning mechanism - FA-PMOAQL. This algorithm continuously iteratively updates the task-agent allocation strategy by constructing a state-action mapping, jointly estimating the pheromone concentration and value function, to ensure the dynamic convergence of the multi-objective optimization goal. Specifically, the system models the task allocation as a Markov decision process, where the state s represents the current task-agent allocation state, the action a represents allocating task t to agent a, and the reward function measures the task relevance (matching degree), distance cost, and fairness performance respectively. At the same time, the system introduces an ant colony optimization mechanism to enhance the local search and global exploration capabilities. In each round of allocation, based on the historical interaction results between task t and agent a, the pheromone concentration is dynamically adjusted. In addition, the system continuously maintains the Pareto front solution set, performs non-dominated sorting on the results of each round of task allocation, eliminates inferior solutions, and only retains the solutions that perform optimally in the three objectives (distance, matching degree, fairness) as the target objects for subsequent pheromone enhancement.

[0046] Step 13: For Q-learning, task allocation is modeled as a state-action process, where the state s is the current task allocation state, which can be represented as the set of tasks S = {S1, S2,..., Sn} that have been assigned to the agent, and the action a is to assign task t to agent a, that is , and the reward R is to measure the quality of the task allocation scheme, combining relevance, distance, and fairness, and its definition is as follows: . Among them, R1: the normalized value of the total relevance , R2: the normalized value of the total distance and take the negative , R3: fairness reward . At the same time, use the Bellman equation to update the Q-table: , where s is the current state, a is the selected action, a′ is the possible next action, and R i is the i-th component of the reward vector, is the learning rate, and γ is the discount factor.

[0047] Step 14: For ant colony optimization, dynamically adjust the pheromone concentration , and its update formula is: , where, is the pheromone evaporation rate, which controls the rate of pheromone decay over time, is the increment of the newly added pheromone, which is used to strengthen the selection probability of high-quality solutions. To ensure that the pheromone only strengthens the Pareto optimal solutions, use the Pareto pheromone update criterion: only when the path (t, a) belongs to the Pareto optimal solution set, increase its pheromone concentration. The increment is calculated as: , where Q is the pheromone enhancement factor, and f(D, S) is the objective function, which measures the quality of the current solution. At the same time, perform pheromone evaporation: .

[0048] Step 15: Calculate the allocation probability of the task , which is jointly determined by the pheromone concentration and the Q value, and the formula is as follows: , where, is the pheromone concentration, which reflects the historical performance of task t assigned to agent a, is the value estimate of agent a for task t, which reflects the contribution of the agent's relevance and fairness, and α and β control the importance of the pheromone and the Q value. After each ant completes the allocation, update the Pareto front according to its reward vector R as follows: , where, when maintaining the Pareto front, judge the dominance relationship between the two reward vectors R1 and R2 , , if R1 dominates R2, then R2 can be removed from the Pareto front. If R1 is dominated by R2, then R1 cannot be added to the Pareto front.

[0049] Fairness evaluation mechanism: To ensure the balanced utilization of agents' resources and the long-term sustainability of task allocation, the system designs and introduces a dynamic fairness evaluation mechanism. This mechanism combines two statistical indicators, the Gini coefficient and the load variance, to comprehensively measure the fairness performance of task scheduling from both individual and overall dimensions, and embeds it into the reinforcement learning reward model to guide the agent to continuously converge its selected strategy towards fair allocation. This fairness evaluation mechanism not only participates in the reinforcement learning feedback in real time during the model training stage, but also can be used as a posterior evaluation indicator to conduct multiple rounds of evaluation and model improvement on the task allocation results, so as to achieve the collaborative optimization of the scheduling system to maintain long-term stability and agent participation fairness while pursuing high efficiency.

[0050] Step 16: After each round of allocation is completed, the platform calculates the task load of the agents and evaluates the scheduling fairness based on the following indicators: where, is the average load of the agent, , n is the number of agents, is the agent 's task load, is arranged in non-decreasing order, is the load variance of the ground agents, is the load variance of the aerial agents, is the Gini coefficient of the ground agents, is the Gini coefficient of the aerial agents, and λ and μ are dynamically adjusted according to the task distribution (e.g., decreased when aerial tasks are intensive).

[0051] Step 17: Embed the fairness index as part of the reward function into the reinforcement learning model to promote the system to achieve fair task allocation while maintaining the task completion efficiency. The calculation formula of the fairness reward function is: . When Var + G → 0 (i.e., both the load variance and the Gini coefficient are close to 0), R3 → 1. At this time, the system load is balanced, the resource allocation is fair, and the reward function gives the highest value. When Var + G → ∞, R3 → 0. At this time, the system load difference is huge, the resource allocation is extremely unfair, and the reward function tends to the lowest value.

[0052] Such as Figure 2As shown in the figure, the present invention first divides the sensing area into a ground task area and an air task area. The DBSCAN algorithm is used for clustering in the ground task area to identify task hot spots. The air tasks are divided into three levels: low altitude, medium altitude, and high altitude according to the altitude, and a quadtree structure is used for fine-grained spatial division. Each spatial area is finally constructed into a multi-layer spatial area structure diagram for subsequent scheduling modeling.

[0053] As Figure 3 shown, the platform constructs a task area graph for all spatial task groups. The nodes in the graph are spatial task groups, and the edges represent the sequential dependencies between tasks.

[0054] As Figure 4 shown, the platform maps the tasks to multi-spatial distribution scenarios, including ground hot spots, airspace monitoring zones, and collaborative coverage areas, etc. For the tasks on the critical path, the platform adopts a greedy strategy for allocation based on the geographical distance and device similarity between the task and the agent.

[0055] Embodiment 2 The following details the fair-guaranteed multi-spatial crowd-sourced sensing task scheduling problem through an example. Suppose a smart city platform receives sensing tasks from 10 different regions in urban monitoring applications, including 6 ground tasks and 4 air tasks. The platform first performs spatial clustering on the ground tasks according to the task locations, identifies 3 ground hot spots, and divides the air tasks into two levels: low altitude and medium altitude according to the altitude. The platform broadcasts the tasks within the candidate range of agents with a radius of 1 kilometer (for the ground) and 500 meters (for the air) around each task point. The candidate agents include unmanned vehicles, ground camera devices, and air drone nodes, etc.

[0056] After receiving the task information, each agent calculates the device matching degree and spatial distance with the task respectively according to its own device capabilities, historical task completion records, and task locations, and generates a task response willingness. After the platform aggregates the candidate agent information, it generates a list of 10 candidate agents based on the matching degree calculation between the task and the agent, and takes it as the input and sends it into the fairness-aware multi-objective optimization scheduling model FA-PMOAQL.

[0057] During the optimization process of this model, with the goals of minimizing the total distance of task allocation, maximizing the device similarity, and ensuring to meet the fairness threshold, it iterates under the strategy framework dominated by Q-learning combined with the ant colony search mechanism. After multiple rounds of strategy convergence, the system outputs an optimal allocation result including 10 tasks and 10 agents, where each task is matched with the optimal agent one by one, and it is superior to the traditional static allocation strategy in multiple key indicators such as task completion rate, fairness, and resource utilization rate.

[0058] Experimental verification: To verify the effectiveness of the method in this embodiment, the present invention conducts experiments on multiple real trajectory datasets, and the results are as Figures 5 to 10 shown. Among them, the above real datasets are from existing public data and obtained the consent of the information owners.

[0059] Among them, MSFTA represents the framework algorithm proposed by the present invention, and the remaining algorithms are comparative algorithms, namely the improved greedy algorithm (IGA), ant colony algorithm (ACO), FGT algorithm, GT+RL algorithm, etc.

[0060] (1) Task completion rate, As Figure 5 shown, with the increase in the number of agents, the task completion rates of all algorithms have increased significantly. Among them, the MSFTA algorithm is always superior to other algorithms, and the increase rate of its task completion rate reaches 12.8%. This is mainly due to the increase in the number of agents expanding the selection space of task allocation, thus effectively reducing the situation of task failure caused by insufficient resources. As Figure 6 shown, with the increase in the task volume, the task completion rates of all algorithms have decreased, indicating that the increase in the number of tasks has increased the processing difficulty. However, the MSFTA algorithm still maintains the best performance under high-load conditions. Its superiority is attributed to the multi-objective optimization and the critical path priority allocation strategy, which effectively matches tasks with agents and minimizes the decrease in the completion rate caused by task backlog.

[0061] (2) Running time, As Figure 7 shown, the running times of each algorithm show a linear increase with the increase in the number of agents, but there are significant differences in the growth rates. It is worth noting that the MSFTA algorithm always maintains a relatively low level in terms of running time, further verifying the high efficiency and applicability of the model.

[0062] (3) Resource utilization rate, As Figure 8 shown, the resource utilization rates of each algorithm decrease with the increase in the number of agents, but the MSFTA algorithm performs the best. This is because this algorithm fully considers the device correlation during the task allocation process and preferentially allocates tasks to the most suitable agents, thus effectively reducing the resource idle rate.

[0063] (4) Fairness index, As Figure 9As shown, the fairness index of each algorithm slightly decreases as the number of agents increases, but the MSFTA algorithm can still maintain a relatively high fairness level. The reason is that high-priority tasks may be concentratedly assigned to a few efficient agents, resulting in uneven load; while the MSFTA algorithm effectively alleviates this problem by introducing load variance and Gini coefficient into the reward function, avoiding excessive sacrifice of fairness.

[0064] (5) Cross-layer task completion rate, As Figure 10 shown, the cross-layer task completion rates of all algorithms show an upward trend as the number of agents increases. Because a smaller number of agents may lead to insufficient task assignment coverage and limited cross-layer cooperation efficiency; while increasing the number of agents can improve the flexibility of task scheduling and resource utilization through dynamic load balancing, thus increasing the completion rate.

[0065] (6) Average task delay, As Figure 11 shown, the average task delays of almost all algorithms show a downward trend as the time window increases. This is because the expansion of the time window provides a longer global optimization scheduling period for the algorithm, enabling the task assignment strategy to coordinate resources more fully, reduce task waiting conflicts, and improve the overall scheduling efficiency.

[0066] The present invention conducts a large number of experiments on two real datasets to verify the performance of the proposed fairness-guaranteed multi-space crowdsensing task assignment method based on multi-objective collaborative optimization. The results show that the fairness-guaranteed multi-space crowdsensing task assignment method based on multi-objective collaborative optimization is superior to other baseline methods.

[0067] Embodiment 3 The present embodiment provides a crowdsensing task scheduling system based on multi-space modeling and fairness reinforcement learning. The system structure is as Figure 1 shown, and it includes the following functional modules: A computer-readable storage medium, which stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the fairness-guaranteed multi-space crowdsensing task assignment dynamic cooperation optimization method.

[0068] A terminal device includes a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the fairness-guaranteed multi-space crowdsensing task assignment dynamic cooperation optimization method.

[0069] The above are all preferred embodiments of the present invention. The protection scope of the present invention is not limited thereto. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A crowd-sensing task scheduling method based on multi-space modeling and fairness reinforcement learning, characterized in that: include: In the multi-space region division phase, a multi-layer space task structure with ground and air coordination is constructed based on the acquired tasks; In the task graph construction and critical path identification phase, a topological association model between tasks is established based on the spatial area division results, and the priority sequence of task scheduling is clarified through the critical path extraction mechanism; In the task allocation phase, we build differentiated agent matching strategies and rationally configure key path tasks based on priority sequences; In the reinforcement learning stage, the fairness-aware multi-objective scheduling algorithm is used to dynamically allocate non-critical path tasks; A dynamic fairness evaluation mechanism is used to evaluate and optimize task allocation results.

2. According to claim 1, a method for group intelligence sensing task scheduling based on multi-space modeling and fairness reinforcement learning is characterized in that: The method constructs a multi-layer spatial task structure for ground and air collaboration based on the acquired tasks, including density clustering of ground task areas, identifying task hotspots using the DBSCAN algorithm, and dividing ground perception tasks into several spatially autonomous task clusters; then, for the air task area, spatial stratification is performed according to the altitude information of the tasks, the airspace is divided into multiple levels of low altitude, medium altitude and high altitude, and a quadtree structure is introduced in each layer to recursively divide the plane area, thereby realizing hierarchical management and refined expression of air tasks.

3. According to claim 2, a method for group intelligence sensing task scheduling based on multi-space modeling and fairness reinforcement learning is characterized in that: The method of constructing a multi-layer spatial task structure for ground and air collaboration based on the acquired tasks also includes introducing a Voronoi diagram construction mechanism, generating a spatial control area based on the spatial centroid of the divided task group, and automatically mapping unattributed tasks to corresponding spatial task units based on the nearest centroid principle; for collaborative tasks involving both ground and air execution requirements, cross-layer coupling mapping is achieved by constructing a hybrid Voronoi diagram to ensure the consistency and integrity of the task space organization.

4. According to claim 3, a method for group intelligence sensing task scheduling based on multi-space modeling and fairness reinforcement learning is characterized in that: The method establishes a topological association model between tasks based on the spatial area division results, and clarifies the priority sequence of task scheduling through a critical path extraction mechanism, including constructing a task area graph based on the spatial task groups formed by the division, wherein the task area graph adopts a weighted undirected graph, wherein nodes correspond to task groups, and the weights of edges are determined by the time urgency difference and spatial distance between task groups; an urgency score is calculated for each task, and the urgency score is defined based on the time window of the task; and a group urgency index is constructed for each task group, and the urgency of tasks within the weighted aggregation group is comprehensively reflected to reflect the scheduling priority of the task group.

5. According to claim 4, a method for group intelligence sensing task scheduling based on multi-space modeling and fairness reinforcement learning is characterized in that: The method establishes a topological association model between tasks based on the spatial area division results, and clarifies the priority sequence of task scheduling through a critical path extraction mechanism. It also includes dynamically generating the association weights of the edges in the graph in combination with the earliest starting time difference between task groups and the cross-space movement cost, which is used to characterize the dependency relationship and cost information between tasks. Among them, a task group network is constructed with the nodes of the Voronoi diagram as the vertices of the graph, in which each seed node corresponds to a task group. When determining the edge weights, all critical paths in the task group network are found by dynamically adjusting the weights and combining with the NetworkX framework.

6. According to claim 5, a method for group intelligence sensing task scheduling based on multi-space modeling and fairness reinforcement learning is characterized in that: The differentiated agent matching strategy is constructed to reasonably configure the critical path tasks based on the priority sequence, including the use of a greedy matching algorithm and a fairness-aware multi-objective reinforcement learning algorithm FA-PMOAQL as complementary scheduling mechanisms. For tasks in the critical path, based on the pre-constructed task scheduling priority sequence, a greedy strategy is used to preferentially match the most suitable agent node for them. During the matching process, the optimization goals are to minimize the spatial distance between the task and the agent and to maximize the matching degree of equipment capabilities.

7. The method for group intelligence sensing task scheduling based on multi-space modeling and fairness reinforcement learning according to claim 6 is characterized in that: The method uses a fairness-aware multi-objective scheduling algorithm to dynamically allocate non-critical path tasks, including constructing a state-action mapping, combining pheromone concentration and value function estimation, and continuously iteratively updating the task-agent allocation strategy to ensure the dynamic convergence of multi-objective optimization goals. Task allocation is modeled as a Markov decision process, and an ant colony optimization mechanism is introduced to enhance local search and global exploration capabilities. For Q learning, task allocation is modeled as a state-action process, and the Bellman equation is used to update the Q table.

8. The method for group intelligence sensing task scheduling based on multi-space modeling and fairness reinforcement learning according to claim 7 is characterized in that: The method uses a fairness-aware multi-objective scheduling algorithm to dynamically allocate non-critical path tasks, and also includes dynamically adjusting the pheromone concentration for ant colony optimization and using the Pareto pheromone update criterion: only when the path belongs to the Pareto optimal solution set, its pheromone concentration is increased and pheromone volatilization is performed at the same time; finally, the task allocation probability is calculated based on the pheromone concentration and the Q value.

9. The method for group intelligence sensing task scheduling based on multi-space modeling and fairness reinforcement learning according to claim 8 is characterized in that: The dynamic fairness evaluation mechanism is used to evaluate and optimize the task allocation results, including calculating the task load of the agent after each round of allocation is completed, calculating the Gini coefficient and load variance based on the task load and using them as fairness indicators to evaluate scheduling fairness; and embedding the fairness indicator into the reinforcement learning model as part of the reward function to achieve task allocation fairness while maintaining task completion efficiency.

10. A crowd-sensing task scheduling system based on multi-space modeling and fairness reinforcement learning, characterized in that: include: The multi-space region division module is configured to, in the multi-space region division phase, construct a multi-layer space task structure for ground and air coordination according to the acquired tasks; The critical path identification module is configured to,at the task graph construction and critical path identification stage,build a topological association model between tasks according to the spatial area division results,and clarify the priority sequence of task scheduling through the critical path extraction mechanism; The task allocation module is configured to,build differentiated agent matching strategies in the task allocation phase,and reasonably allocate the critical path tasks based on the priority sequence; The reinforcement learning module is configured to,in the reinforcement learning phase, dynamically allocate non-critical path tasks,using the fairness-aware multi-objective scheduling algorithm; The evaluation module is configured to evaluate and optimize the task allocation results using a dynamic fairness evaluation mechanism.

Citation Information

Patent Citations

  • Moving target traversal access sequence planning method based on improved pointer network

    CN116090688A

  • Three-party multi-objective-considered high-dimensional optimization task allocation method and system

    CN118536678A

  • Cluster system intelligent operation and maintenance system based on reinforcement learning

    CN119090492A

  • Multi-agent task planning method for high-order thinking development under mixed game

    CN119312876A

  • Self-adaptive task scheduling execution unit management method and system

    CN119376903A

Cited By

  • Marketing scene-oriented AI agent task scheduling execution method

    CN120803672A

  • Streaming computing task scheduling method based on space-time perception and multi-agent reinforcement learning

    CN121479359A

  • Low-altitude resource intelligent scheduling method and system based on deep learning

    CN121616061A

  • Mobile crowd sensing fair unit selection method based on multi-agent reinforcement learning

    CN121723855A

  • Crowd sensing task allocation method and system based on cue word and deduction decoupling

    CN122242772A