Two-stage excitation method based on multi-agent deep reinforcement learning
Through the two-stage incentive method of multi-agent deep reinforcement learning, the problems of flexibility and resource waste in mobile crowd sensing task allocation are solved, efficient task coverage and resource allocation are achieved, and the task completion rate and platform utility are improved.
Patent Information
- Application Number
- CN202510901283.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
Smart Images

Figure CN120806036A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of agent deep reinforcement learning, in particular to a two-stage incentive method based on multi-agent deep reinforcement learning. BACKGROUND
[0002] The prior art has many deficiencies in mobile crowd sensing task allocation and incentive mechanism. The traditional single task single allocation mode has poor flexibility, and the enumeration method has high computational complexity in large-scale scenarios. In resource scarce scenarios, the task "minimum feasibility coverage" is difficult to guarantee, and there is a lack of balance mechanism between budget limit and task coverage integrity. In addition, the sensing and computing energy consumption are not separated, the task complexity evaluation lacks multi-dimensional feature quantization, the dynamic pricing strategy is not intelligent enough, it is difficult to adapt to the changes of participant state, and the evaluation index is single, which does not comprehensively consider the participant and platform utility and task completion rate, resulting in resource waste and low task execution efficiency. The specific problems are as follows:
[0003] 1. In the multi-to-multi resource allocation mode, the dynamic adjustment efficiency is low, the different sensor energy consumption characteristics are not distinguished, the participant pricing strategy does not consider the income and energy consumption, and the task complexity evaluation lacks multi-dimensional feature quantization.
[0004] 2. The traditional single task single allocation mode has poor flexibility, the enumeration method has high time complexity in large-scale scenarios, the task "minimum feasibility coverage" is difficult to guarantee in resource scarce scenarios, there is a lack of balance mechanism between budget limit and task coverage integrity, and there is a lack of multi-dimensional feature such as sensing resource quantity, energy consumption quantity, and task complexity evaluation.
[0005] 3. The dynamic pricing strategy is not intelligent enough, lacks real-time adjustment mechanism based on deep reinforcement learning, the task allocation mode is limited to traditional single task single allocation, has poor flexibility, and lacks modeling and adaptation ability to task geographical distribution, participant trajectory and dynamic state change, and the evaluation index is single, which does not comprehensively consider the multi-dimensional benefits such as participant utility, platform utility and task completion rate. SUMMARY
[0006] The purpose of the present application is to provide a two-stage incentive method based on multi-agent deep reinforcement learning, through simulation experiments in specific areas, to simulate real perception task geographic distribution and trajectory planning with taxis as participants, to adopt deep Q learning, historical participation rate optimization and other comparison mechanisms, to combine multi-dimensional indicators such as average participant utility, platform utility, task completion rate, to verify the adaptability of the scheme in a dynamic environment, to combine different sensor characteristics to quantify the cost, to enable participants to accurately assess their own energy consumption, to maximize platform utility under budget constraints, and to ensure task "minimum feasibility coverage" and budget not overspending, to observe the actions of other agents during training to optimize collaboration, to rely only on local state during execution, to adapt to multi-participant competitive scenarios, and to solve the problems in the prior art.
[0007] To achieve the above purpose, the present application provides the following technical solutions:
[0008] A two-stage incentive method based on multi-agent deep reinforcement learning, comprising:
[0009] a dynamic pricing optimization stage and a task allocation optimization stage;
[0010] The dynamic pricing optimization stage is to evaluate the complexity of the task set through the analytic hierarchy process, to provide a reference basis for the pricing of participants, to combine task complexity and historical bidding data, and to use multi-agent deep deterministic policy gradient technology to train the optimal unit pricing strategy;
[0011] The task allocation optimization stage is to select key participants, confirm the core requirements of the task, and then optimize the task allocation according to the satisfaction of the participants.
[0012] Preferably, it further comprises:
[0013] Each task is assigned to multiple mobile participants, and each mobile participant participates in multiple perception tasks simultaneously, wherein the perception tasks include task broadcasting, mobile participant bidding, task allocation, task execution and reward reward;
[0014] In the perception area, the perception platform manages the task set T = {T1, T2, …, T m}, wherein each perception task has a required total amount of perception resources, and each task T j ∈T can be divided into multiple perception rounds R = {R1, R2, …, R k} for execution;
[0015] In each perception round R t ∈R, each participant needs to collect the corresponding amount of perception resources q j when performing the task T j ;
[0016] The sensing platform broadcasts to the participant set P = {P1, P2, ..., P n}Publish task information, each participant P i ∈P autonomously selects a task subset based on its own resource capabilities and preferences Participate and upload bidding information in is the total amount of selected sensing resources, is the total energy consumption, is the number of tasks, Quote for resource units;
[0017] The perception platform optimizes task allocation based on the collected bidding information and reasonably allocates tasks to a group of participants. After the task allocation is completed, the participants actively perform the perception tasks according to the task requirements and upload the collected perception resource volume. and receive corresponding remuneration and rewards;
[0018] For unfinished tasks These tasks will be included in the next round of R t+1 , continue to assign and execute until all tasks are completed;
[0019] Among them, T j Represented as a perception task; R t Represented as perception round; P i represented as participants; Represented as the perception wheel R t When participant P i The selected set of tasks; Represented as the perception wheel R t When participant P i Uploaded bidding information.
[0020] Preferably, the dynamic quotation optimization stage includes:
[0021] The perception energy consumption is calculated as follows: When performing multiple perception tasks, mobile participants use different types of sensors, including radar, GPS, and cameras;
[0022] Among them, GPS records data at a rate of bits per second, and the camera is used to obtain image data, which generates image data at millions of bits per second;
[0023] If δ i For participant P i The perception rate, then for the perception resource size q′ j , participant P i Complete the perception task T j Time required The formula is as follows:
[0024]
[0025] Participant P i Complete the perception task T j The formula for perceived energy consumption is as follows:
[0026]
[0027] Among them, p s Indicates unit energy consumption; Represented as participant P i For the perception task T j Perception energy consumption and computation energy consumption; q j Represented as performing perception task T j the amount of perceived resources required;
[0028] The calculation of energy consumption is: Participant P i When processing sensory data, computing resources are consumed;
[0029] If χ represents the number of CPU cycles required to process bit-perceptual data, then participant P i Complete the perception task T j Data calculation time The formula is as follows:
[0030]
[0031] Among them, f i Indicates the computing resources of the device, that is, the CPU frequency;
[0032] Processing perception task T j The formula for the raw data energy consumption is as follows:
[0033]
[0034] Where κ represents the effective switching capacitance; Represented as participant P i For the perception task T j The perceived energy consumption and computational energy consumption.
[0035] Preferably, the dynamic quotation optimization stage also includes:
[0036] Mobile Participant P i In the perception wheel R t By selecting and executing a set of tasks The formula for the utility obtained is as follows:
[0037]
[0038] in, Represents participant P i Perception Task T j The actual benefits obtained; Represents participant P i In the perception task T j The perceived resources provided in; c represents the unit cost of energy consumption; Represents participant P i Perform perception task T j Actual energy consumption;
[0039] At the same time, the benefit of the perception platform is the perceived quality contribution provided by the participants in completing the perception task. According to the heterogeneous characteristics of the participants' devices, different participants contribute differently to the perceived quality.
[0040] The participants’ perceived quality contributions on different tasks were evaluated using As a participant P i Task T j perceived contribution;
[0041] Among them, the weight ω ij Reflects the participant P i For Task T j The contribution of the unit-perceived resources;
[0042] The total revenue of the platform is composed of the cumulative perceived quality contributions achieved in all tasks, and the total cost is the sum of the incentive rewards paid by the platform to all participants. The utility of the perceived platform is the difference between the total revenue and the total cost. The difference formula is as follows:
[0043]
[0044] Participant P i The goal is to select a set of tasks Develop an appropriate unit quotation To maximize its net income, player P i The formula for the optimal bidding problem is as follows:
[0045]
[0046] Among them, θ ij Represents the perception task T j Is it assigned to participant P? i , if assigned, then θ ij =1, otherwise θ ij =0;
[0047] Participant P i Selected Task Set If there is already an assigned task, then all tasks in the task set selected by the participant have been assigned, i.e. θ i = 1.
[0048] The constraint is that the total energy consumption of the participant to complete the task set does not exceed the energy budget of its device
[0049] where θ ij is the indicator value of whether the platform selects the participant P i to execute the sensing task T j . i
[0050] Preferably, the dynamic pricing optimization stage further comprises:
[0051] An analytic hierarchy process is used to determine the influence of multi-dimensional features on the complexity of the task combination TS i , and the formula is as follows:
[0052]
[0053] wherein, is the total amount of the largest sensing resources uploaded by all participants; is the maximum energy consumption amount of all participants executing tasks; is the maximum number of tasks selected by all participants; w1, w2, and w3 are weights for evaluating the relative importance of different parameters, and the sum of w1, w2, and w3 is 1.
[0054] In the analytic hierarchy process, X1, X2, and X3 are used to represent the sensing resource amount, energy consumption amount, and task number of the task set selected by the participant, respectively;
[0055] The importance between different parameters is measured by constructing a pair-wise comparison matrix A = (a ij ) 3×3 , wherein each element a ij represents the importance of X i relative to X j , and the value is assigned to an appropriate integer between 1 and 9 by the proposed scale a ij , and varies according to different situations;
[0056] When a ij > 1, X i is more important than X j ; when a ij < 1, X i is less important than X j ; when a ij = 1, X i and X j are equally important;
[0057] Pairwise comparison matrix is A=(a ij ) 3×3 , a 13 =2 indicates that the perceived resource amount X1 is more important than the task number X2;
[0058] After normalizing each column of A, the normalized matrix is obtained, and the formula is as follows:
[0059]
[0060] Weight vector W=(w1,w2,w3) T The average value of each row element of the normalized matrix is calculated;
[0061] Preferably, the dynamic pricing optimization stage further comprises:
[0062] An optimal unit price of the participant is formulated using MADDPG to maximize the participant's utility, wherein the MADDPG is used for dynamic pricing strategy training in the first stage, and through a centralized training and distributed execution framework, each participant observes the state and action of other intelligent agents during training to optimize the strategy;
[0063] In the perception round R t , each intelligent agent will output an action according to the data in the observation space
[0064] After completing the task according to the action, an expected reward will be obtained, and the environment in the observation is adjusted, and the strategy is adjusted;
[0065] Wherein, the state observation space , the action space and the reward function of each participant are defined as follows;
[0066] The observation space is: according to the observation results of the previous L rounds, the environmental changes of each participant are confirmed, and in each perception round R t , the observation result of the participant P i is as follows:
[0067]
[0068] Wherein, represents the selected task combination complexity of the participant P t in the perception round R i ; represents the selected task combination complexity of the participant P t-1 in the perception round R i ; represents the selected task combination complexity of the participant Pt-1 Participant P i Unit quotation of Represents the perception wheel R t-1 Participant P i The utility of Represents the perception wheel R t-1 whether participants are selected by the perceived platform;
[0069] The action space is: on the perception wheel R t In the i Observe the platform's decision in the first L rounds of perception and give the unit price of the current sensing wheel based on the output of the Actor network. The formula is as follows:
[0070]
[0071] Among them, π i Represents an Actor network; Represents the gradient optimization parameters of the Actor network;
[0072] The reward function is: according to the action Participant P i In the perception wheel R t The reward obtained is expressed using the utility function, as follows:
[0073]
[0074] Among them, θ ij Represents the perception task T j Is it assigned to participant P? i ;
[0075] For the perception wheel R t After selecting the task composition set, participants get the optimal bid based on MADDPG decision
[0076] Then upload the bidding record containing the latest quotation information
[0077] Finally, the perception platform allocates tasks based on the bidding information.
[0078] Preferably, the task allocation optimization stage further includes:
[0079] According to the simplified scenario of one task and multiple participants, the goal of the perception platform is to assign tasks to participants to maximize the platform utility while meeting the condition of not exceeding the given budget limit. The optimization problem formula of the simplified scenario of perception platform efficiency is as follows:
[0080]
[0081] Among them, θ′ i Indicates whether the sensing platform assigns a task to a participant; Represents the perception wheel R t Participant P i The contribution provided to the perception platform by performing a certain task; Represents the perception wheel R t Participant P i Reward for completing a task; B t Represents the perception wheel R t The platform budget to complete a task;
[0082] Based on the goal of maximizing platform utility, the task allocation algorithm is used to calculate the process. The calculation steps are as follows:
[0083] S1: Conduct a detailed analysis of the participants’ task coverage capabilities and include participants who meet the task requirements into the candidate set;
[0084] If there is only one participant in the candidate set, it is identified as the key participant. In the absence of other options and if the budget allows, the key participant will be given priority, and the budget and task coverage status will be adjusted accordingly;
[0085] S2: For unassigned tasks, select candidates who can provide necessary coverage from the unselected participants and use the satisfaction index SA i , SA i The ratio of perceived contribution to bid price is as follows:
[0086]
[0087] S3: Rank the satisfaction of the unselected participants to ensure the optimal allocation within the budget limit, and select the top-ranked participants in turn until the budget is exhausted;
[0088] According to the calculation steps of the task allocation algorithm, the updated θ i To complete the task allocation strategy, participants perform the assigned tasks and submit the collected perception data. Then, the MCS platform allocates resources based on the actual resource contribution of the participants. and the best unit quotation Payment Rewards The total reward for the selected participants cannot be greater than the platform budget.
[0089] Preferably, the task allocation optimization stage further includes:
[0090] The goal of the perception platform is to assign tasks to appropriate participants to optimize the platform utility. The task assignment problem is transformed into the problem of maximizing the platform utility as follows:
[0091]
[0092] wherein B t is the platform budget at the perception round R t ;
[0093] The first constraint condition is to confirm that the total expenditure of the incentive of the perception platform at the perception round R t does not exceed the budget B t , and the second constraint condition is to confirm that each task T j is at least assigned to one participant to complete.
[0094] Preferably, it further comprises:
[0095] Based on the dynamic pricing optimization stage and the task allocation optimization stage, it further comprises a simulation setting, a comparison mechanism, an evaluation index, and a comparison and analysis;
[0096] The simulation setting is to select a specific area for simulation experiment, randomly determine 10 to 30 task points in the specific area, each task point is located on the road, and is used to simulate the geographical distribution of real perception tasks. The selection of the task points is based on the road network, and each task has corresponding perception resource requirements;
[0097] In the specific area, 10 to 30 taxis with longer trajectories are selected as participants, the participants plan routes and complete tasks according to their interest in the tasks, each participant is an agent, and learns how to adjust the price according to the task complexity, historical pricing data, and platform feedback through a strategy network and a value network;
[0098] The strategy network and the value network each consist of an input layer, three hidden layers, and an output layer, wherein the number of hidden layer neurons is 512, 256, and 32 respectively. The strategy network and the value network use the ReLU function as the activation of all hidden layers, and the value network uses the Tanh function as the activation of the output layer to generate a strategy;
[0099] The comparison mechanism is that in the comparative experiment, a dynamic pricing strategy generation mechanism based on deep Q learning, a dynamic pricing optimization mechanism of historical participation rate and income, and a random pricing and task allocation mechanism are adopted, and the random pricing and task allocation mechanism is used as the comparison mechanism.
[0100] Preferably, it further comprises:
[0101] The evaluation index is introduced, including the average participant utility, the average platform utility, and the task completion rate;
[0102] The comparison and analysis is that according to the first analysis chart, when the number of tasks is 10 and the platform budget is 25, the average participant utility changes under different numbers of participants;
[0103] According to the second analysis chart, when the number of tasks is 10 and the platform budget is 25, the change of the average platform utility under different numbers of participants is known;
[0104] According to the third analysis chart, when the number of tasks is 10 and the platform budget is 25, the change of the task completion rate under different numbers of participants is known;
[0105] Compared with the prior art, the beneficial effects of the present application are as follows:
[0106] 1. The two-stage incentive method based on multi-agent deep reinforcement learning provided by the present application quantifies the task complexity by the analytic hierarchy process (AHP), converts multi-dimensional characteristics such as the amount of sensing resources, the amount of energy consumption and the number of tasks into a unified index, and provides a scientific basis for participant bidding. Combined with the multi-agent deep deterministic policy gradient (MADDPG) technology, the participant acts as an agent to learn the optimal bidding strategy through "centralized training + distributed execution". During training, the actions of other agents can be observed to optimize collaboration, and during execution, only the local state is relied on to adapt to the multi-participant competitive scenario. This mechanism allows the participant to consider both task revenue and device energy cost when bidding, avoiding resource waste, and at the same time, the platform can improve resource allocation efficiency based on the three-stage task allocation algorithm.
[0107] 2. The two-stage incentive method based on multi-agent deep reinforcement learning provided by the present application designs a three-stage task allocation algorithm (TPTAA), first determines the key participants to ensure the core needs of the task, then allocates the remaining tasks based on the satisfaction index of "sensing contribution / bid price", and finally consumes the budget according to the satisfaction order. The "many-to-many" allocation mode breaks through the traditional single-task single-allocation limitation, and the uncompleted tasks automatically enter the next round of allocation, dynamically adapting to the changes in the state of the participants. Under the budget constraint, the platform utility is maximized, and through the two-way balance of budget and task coverage, the "minimum feasible coverage" of the task and the non-exceeding of the budget are ensured.
[0108] 3. The two-stage incentive method based on multi-agent deep reinforcement learning provided by the present application simulates the geographical distribution and trajectory planning of real sensing tasks by taking taxis as participants through simulation experiments in specific areas, adopts deep Q learning, historical participation rate optimization and other comparison mechanisms, combines multi-dimensional indexes such as average participant utility, platform utility and task completion rate, and verifies the adaptability of the scheme in dynamic environment. Experiments show that compared with traditional mechanisms such as random bidding, the task completion rate is significantly improved, the budget can be fully utilized, and the individual and global benefits are considered, providing efficient task allocation and incentive strategies for urban sensing scenarios, and through the combination of MADDPG and AHP, the system can dynamically adapt to the differences in participant devices and changes in task types, improving the generalization ability of the system. BRIEF DESCRIPTION OF DRAWINGS
[0109] Figure 1 A pair comparison matrix and a normalization matrix of the present application are shown in the following table:
[0110] Figure 2 A MADDPG framework of the present application is shown in the following table:
[0111] Figure 3 A first analysis diagram of the present application is shown in the following table:
[0112] Figure 4 A second analysis diagram of the present application is shown in the following table:
[0113] Figure 5 A third analysis diagram of the present application is shown in the following table: DETAILED DESCRIPTION
[0114] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0115] In order to solve the problems in the prior art, such as low dynamic adjustment efficiency in the many-to-many resource allocation mode, resource waste caused by not distinguishing different sensor energy consumption characteristics, not taking into account both income and energy consumption in the participant bidding strategy, and lack of multi-dimensional feature quantization in task complexity evaluation, please refer to Figures 1-5 The technical solutions of the present application are provided as follows:
[0116] A two-stage incentive method based on multi-agent deep reinforcement learning includes:
[0117] a dynamic bidding optimization stage and a task allocation optimization stage;
[0118] The dynamic bidding optimization stage is to evaluate the complexity of the task set by the analytic hierarchy process, provide a reference for the participant's bidding, combine the task complexity and historical bidding data, and train the optimal unit bidding strategy by using the multi-agent deep deterministic policy gradient technology.
[0119] The task allocation optimization stage is to select key participants, confirm the core requirements of the task, and then optimize the task allocation according to the satisfaction of the participants.
[0120] Specifically, the analytic hierarchy process (AHP) is used to convert multi-dimensional characteristics such as perceived resource demand, energy consumption, and task quantity of the task into a unified complexity index, solving the problem of different dimensional characteristics that cannot be directly input into the model. With the help of multi-agent deep deterministic policy gradient (MADDPG) technology, each participant is regarded as an agent, and the optimal bidding strategy is learned through the "centralized training + distributed execution" mode. During training, the actions of other agents are observed to optimize collaboration, and during execution, only local state is relied on to adapt to the multi-participant competitive scenario. For the NP-hard problem of maximizing platform utility, a three-phase task allocation algorithm (TPTAA) is designed: first, determine the key participants to ensure the core requirements of the task (such as the participant who can only cover a certain task), then allocate the remaining tasks based on the satisfaction index of "perceived contribution / bid price", and finally, consume the remaining budget according to the satisfaction order. This framework efficiently approximates the optimal solution through heuristic strategies, significantly improving task completion rate and platform utility rate compared to random allocation or single-objective optimization. During task allocation, both the perceived contribution of participants and the unit bid are considered, ensuring task execution quality (such as prioritizing participants with high perceived contribution) and controlling platform costs.
[0121] Also included are:
[0122] Each task is assigned to multiple mobile participants, and each mobile participant participates in multiple perception tasks simultaneously, wherein the perception tasks include task broadcasting, mobile participant bidding, task allocation, task execution, and reward reward;
[0123] In the perception area, the perception platform manages the task set T = {T1, T2, …, T m}, wherein each perception task has a required total amount of perception resources, and each task T j ∈T can be divided into multiple perception rounds R = {R1, R2, …, R k} for execution;
[0124] In each perception round R t ∈R, each participant needs to collect the corresponding amount of perception resources q j when executing the task T j ;
[0125] The perception platform broadcasts task information to the participant set P = {P1, P2, …, P n}, and each participant P i ∈P selects a subset of tasks according to its own resource capacity and preference, and uploads bidding information , wherein is the total amount of selected perception resources, is the total energy consumption, is the number of tasks, and Quoting for resource units;
[0126] The perception platform optimizes task allocation according to the collected bid information, reasonably allocates tasks to a group of participants, and after the task allocation is completed, the participants actively perform the perception task according to the task requirements and upload the collected perception resource quantity And get the corresponding reward;
[0127] For unfinished tasks These tasks will be re-included in the next round of R t+1 , continue to allocate and execute until all tasks are completed;
[0128] Where T j represents a perception task; R t represents a perception round; P i represents a participant; represents the task set selected by participant P t at perception round R i ; represents the bid information uploaded by participant P t at perception round R i .
[0129] Specifically, each task is assigned to multiple participants, and each participant can participate in multiple tasks at the same time, forming a "many-to-many" resource allocation mode. Participants autonomously select a task subset and upload bid information (including total perception resource quantity, energy consumption, task quantity, and unit price). The perception platform can dynamically adjust the allocation based on the bid data, reasonably disperse the tasks to multiple participants, automatically include the unfinished tasks in the next round of allocation in each round of perception, and ensure the final completion of the tasks through multiple rounds of iteration. For example, a task that is not completed in this round due to high bidding price or insufficient resources can be reallocated based on new bid information in the next round, avoiding task abandonment. This mechanism has strong robustness to sudden tasks or dynamically changing participant states (such as device offline or power depletion). In multiple rounds of perception, participants can optimize their bidding strategies based on the task allocation results and reward feedback from previous rounds, and the platform can accumulate historical bid data to optimize the allocation model. The standardized process design of task broadcasting, bidding, allocation, execution, and reward ensures clear and predictable interaction between participants and the platform. Each task can be divided into multiple perception rounds for execution, and the platform can dynamically split tasks based on total resource requirements and round budgets to achieve global resource planning.
[0130] The dynamic pricing optimization phase includes:
[0131] The calculation of perception energy consumption is as follows: when executing multiple perception tasks, mobile participants use different types of sensors, including radar, GPS, and camera;
[0132] Among them, GPS records data at a rate of bits per second, and the camera is used to obtain image data, which generates image data at millions of bits per second;
[0133] If δ i For participant P i The perception rate, then for the perception resource size q′ j , participant P i Complete the perception task T j Time required The formula is as follows:
[0134]
[0135] Participant P i Complete the perception task T j The formula for perceived energy consumption is as follows:
[0136]
[0137] Among them, p s Indicates unit energy consumption; Represented as participant P i For the perception task T j Perception energy consumption and computation energy consumption; q j Represented as performing perception task T j the amount of perceived resources required;
[0138] The calculation of energy consumption is: Participant P i When processing sensory data, computing resources are consumed;
[0139] If χ represents the number of CPU cycles required to process bit-perceptual data, then participant P i Complete the perception task T j Data calculation time The formula is as follows:
[0140]
[0141] Among them, f i Indicates the computing resources of the device, that is, the CPU frequency;
[0142] Processing perception task T j The formula for the raw data energy consumption is as follows:
[0143]
[0144] Where κ represents the effective switching capacitance; Represented as participant P i For the perception task T j The perceived energy consumption and computational energy consumption.
[0145] Specifically, for the characteristic differences of different types of sensors (radar, GPS, camera), energy consumption models are established respectively. By distinguishing the sensor types, the platform can allocate appropriate participants according to the task requirements (such as positioning or image recognition), avoiding the waste of resources caused by "high energy consumption sensors processing simple tasks". The platform can also predict the rationality of the offers of different participants based on the energy consumption model, optimize the task allocation decision, and the mathematical modeling of the energy consumption formula provides accurate state input for multi-agent deep reinforcement learning (such as MADDPG).
[0146] The dynamic offer optimization stage also includes:
[0147] Mobile participants P i In the perception round R t , by selecting and executing the task set The utility obtained is as follows:
[0148]
[0149] Among them, represents the actual income obtained by the participant P i in the perception task T j ; represents the perception resources provided by the participant P i in the perception task T j ; c represents the unit cost of energy consumption; represents the actual energy consumption of the participant P i in executing the perception task T j ;
[0150] At the same time, the benefit of the perception platform is the contribution of the perception quality provided by the participants in completing the perception task, wherein, according to the heterogeneous characteristics of the participant devices, different participants have different contributions to the perception quality;
[0151] The evaluation of the perception quality contribution of the participants in different tasks is: as the perception contribution of the participant P i to the task T j ;
[0152] Among them, the weight ω ij reflects the contribution of the participant P i to the unit perception resource of the task T j ;
[0153] The total benefit of the platform is composed of the cumulative perception quality contribution realized in all tasks, and the total cost is the sum of the incentive rewards paid by the platform to all participants. The utility of the perception platform is the difference between the total benefit and the total cost, and the difference formula is as follows:
[0154]
[0155] Participant P i The goal is to select a set of tasks Develop an appropriate unit price quote i t maximizes its net income, player P i The formula for the optimal bidding problem is as follows:
[0156]
[0157] Among them, θ ij Represents the perception task T j Is it assigned to participant P? i , if assigned, then θ ij =1, otherwise θ ij =0;
[0158] Participant P i Selected Task Set If there is an assigned task in the task set, then all tasks in the task set selected by the participant have been assigned, that is, θ i =1;
[0159] The constraint is that the total energy consumption of the participant to complete the task set does not exceed the energy budget of his device
[0160] Among them, θ ij Indicates whether the platform selects participant P i The index value of the perception task T j Whether P i implement.
[0161] Specifically, the participant utility formula directly links the actual income with the energy cost, driving the participants to consider both the task income and the device resource consumption when bidding. The platform utility formula takes the perceived quality contribution as the core of the income, and quantifies the global benefit in combination with the incentive cost. Among them, the weight reflects the device heterogeneity (such as high-end cameras contribute higher quality than GPS), so that the platform can preferentially select "high perceived quality / low incentive cost" participants to maximize the perceived value under the budget constraint. The constraint condition of the participant's optimal bidding problem forces the total energy consumption to be less than the device budget, avoiding the device failure or offline caused by the participant's excessive participation in the task. If there is an already allocated task in the task set selected by the participant, it is considered that all tasks have been allocated. This rule simplifies the platform decision logic through "task set overall allocation". The participant's objective function takes the unit bid as a continuous variable optimization, which can realize dynamic strategy convergence combined with reinforcement learning algorithms such as MADDPG. The weight in the platform utility quantifies the contribution difference of different devices to the same task (such as higher data quality collected by high-end sensors), so that the platform can objectively evaluate the value of the participants when allocating.
[0162] The dynamic bidding optimization stage also includes:
[0163] The analytic hierarchy process is used to determine the influence of multi-dimensional features on the task combination complexity TS i , and the formula is as follows:
[0164]
[0165] Among them, is the total amount of the largest perception resources uploaded by all participants; is the maximum energy consumption amount of all participants executing tasks; is the maximum number of tasks selected by all participants; w1, w2, and w3 are weights evaluating the relative importance of different parameters, and the sum of w1, w2, and w3 is 1;
[0166] In the analytic hierarchy process, X1, X2, and X3 are used to represent the perception resource amount, energy consumption amount, and task number of the task set selected by the participants, respectively;
[0167] The importance between different parameters is measured by constructing a pair-wise comparison matrix A=(a ij ) 3×3 , where each element a ij represents the importance of X i relative to X j , and the value is assigned an appropriate integer between 1 and 9 by the proposed scale a ij , and varies according to different situations;
[0168] When a ij >1, X i is more important than Xj More important; when a ij <1, X i Not as important as X j Important; when a ij =1, X i and X j are equally important;
[0169] The pair-wise comparison matrix is A=(a ij ) 3×3 , a 13 =2 indicates that the amount of perception resource X1 is more important than the number of tasks X2;
[0170] After normalizing each column of A, the normalized matrix is obtained, and the formula is as follows:
[0171]
[0172] The weight vector W=(w1, w2, w3) T is calculated by taking the average of each row element of the normalized matrix;
[0173] Specifically, the total amount of perception resource, the amount of energy consumption, and the number of tasks are included in the unified evaluation framework, the influence of each parameter on the task complexity is quantified through the weight vector, AHP allows domain experts to dynamically adjust the parameter weight according to the task characteristics by constructing a pair-wise comparison matrix, avoids the one-sidedness of single index dominating the decision, quantifies the relative importance between parameters through the 1-9 scale method, this numerical expression makes the expert experience into a calculable model input, makes the model more focus on the energy consumption factor when evaluating the task complexity, improves the pertinence of the decision, the task complexity index output by AHP can be used as the state input of MADDPG algorithm, which provides more rich environment information for the agent, this cooperative mechanism makes the multi-agent system more efficiently cope with complex environmental changes, and realizes the global utility maximization.
[0174] The dynamic pricing optimization stage also includes:
[0175] MADDPG is used to formulate the optimal unit price of the participant to maximize the participant utility, wherein the MADDPG is used for dynamic pricing strategy training in the first stage, through the centralized training and distributed execution framework, each participant is allowed to observe the state and action of other agents during training to optimize the strategy;
[0176] In the perception round R t , each agent will output an action according to the data in the observation space
[0177] After completing the task according to the action, the expected reward will be obtained, and the strategy will be adjusted according to the observation of the changing environment.
[0178] where the state observation space of each participant action space and reward function are defined as follows.
[0179] The observation space is defined as: according to the observation results of the previous L rounds, the environmental changes of each participant are confirmed, and in each perception round R t , the observation result of participant P i is as follows:
[0180]
[0181] where represents the selected task combination complexity of participant P t in the perception round R i ; represents the selected task combination complexity of participant P t-1 in the perception round R i ; represents the unit bid of participant P t-1 in the perception round R i ; represents the utility of participant P t-1 in the perception round R i ; represents whether the participant is selected in the perception round R t-1 ;
[0182] The action space is: in the perception round R t , each participant P i observes the decision of the platform in the previous L rounds of perception, and gives the unit bid of the current perception round according to the output of the Actor network, and the formula is as follows:
[0183]
[0184] where π i represents the Actor network; represents the gradient optimization parameter of the Actor network;
[0185] The reward function is: according to the action obtained by participant P i in the perception round R t , the utility function is used to represent, and the formula is as follows:
[0186]
[0187] where θ ij represents whether the perception task T j is assigned to participant Pi ;
[0188] For the perception wheel R t , the participants select a task set after the task set is selected, and the optimal bid is obtained based on the MADDPG decision
[0189] Upload the bid record containing the latest bid information again
[0190] Finally, the perception platform allocates tasks according to the bid information.
[0191] Specifically, the action space is defined as a continuous unit bid, which is more suitable for the actual scene than a discrete bid. The deterministic policy gradient of MADDPG can directly output the optimal continuous bid value, avoiding the policy shock caused by discretization. The reward function is directly related to the participant's utility, driving the agent to maximize the net income. The actor network outputs a deterministic bid strategy, and the critic network evaluates the joint action value (combined with all agent actions). The two cooperate to realize the closed loop of "policy generation-value evaluation-policy optimization".
[0192] In order to solve the problems of lack of flexibility in traditional single task single allocation mode, high time complexity of enumeration method in large-scale scene, difficulty in guaranteeing "minimum feasibility coverage" of tasks in resource scarce scene, lack of balance mechanism between budget limit and task coverage integrity, lack of task complexity evaluation of multi-dimensional features such as perception resource quantity and energy consumption, please refer to Figures 1-5 The embodiment provides the following technical solutions:
[0193] The task allocation optimization stage also includes:
[0194] According to a simplified scenario of one task and multiple participants, the perception platform aims to allocate tasks to participants to maximize platform utility while meeting the condition of not exceeding the given budget limit. The optimization problem formula of the perception platform efficiency simplified scenario is as follows:
[0195]
[0196] Wherein, θ′ i represents whether the perception platform allocates a certain task to a participant; represents the perception wheel R t , the contribution provided by the participant P i to the perception platform for executing a certain task; represents the perception wheel R t , the reward of the participant P i for executing a certain task; B t represents the platform budget for completing a certain task in the perception wheel R t ; and
[0197] Based on the goal of maximizing platform utility, the task allocation algorithm is used to calculate the process. The calculation steps are as follows:
[0198] S1: Conduct a detailed analysis of the participants’ task coverage capabilities and include participants who meet the task requirements into the candidate set;
[0199] If there is only one participant in the candidate set, it is identified as the key participant. In the absence of other options and if the budget allows, the key participant will be given priority, and the budget and task coverage status will be adjusted accordingly;
[0200] S2: For unassigned tasks, select candidates who can provide necessary coverage from the unselected participants and use the satisfaction index SA i , SA i The ratio of perceived contribution to bid price is as follows:
[0201]
[0202] S3: Rank the satisfaction of the unselected participants to ensure the optimal allocation within the budget limit, and select the top-ranked participants in turn until the budget is exhausted;
[0203] According to the calculation steps of the task allocation algorithm, the updated θ i To complete the task allocation strategy, participants perform the assigned tasks and submit the collected perception data. Then, the MCS platform allocates resources based on the actual resource contribution of the participants. and the best unit quotation Payment Rewards The total reward for the selected participants cannot be greater than the platform budget.
[0204] Specifically, step S1 assigns the task to the participant and adjusts the budget by analyzing the participant's task coverage capability. If there is only one participant (key participant) in the alternative set of a task, the task is assigned preferentially and the budget is adjusted. This mechanism is particularly important in resource scarce scenarios (such as remote area tasks), ensuring task "minimum viable coverage". After assigning the key participant, the budget and task coverage state are updated synchronously to avoid the risk of repeated allocation or exceeding the budget. Step S2 uses the satisfaction index to preferentially select participants with "high perceived quality / low bid", ensuring that the platform obtains higher perceived value with limited budget. For unassigned tasks, candidates are selected from unselected participants and the SA index is calculated to avoid "suboptimal allocation". Step S3 sorts the unselected participants in descending order of SA, and selects them in turn until the budget is exhausted. The constraint that the total reward does not exceed the platform budget and the requirement that each task is assigned at least once form a two-way balance. By reducing the problem to a knapsack problem, the three-stage heuristic algorithm completes the allocation in polynomial time, which is more suitable for large-scale scenarios than the enumeration method (exponential time). Unfinished tasks automatically enter the next round of allocation. Combined with real-time updating of participant bidding information (such as new participants and bid adjustments), the algorithm can dynamically optimize the allocation strategy.
[0205] The task allocation optimization phase also includes:
[0206] The goal of the perception platform is to assign tasks to suitable participants to optimize platform utility. The task allocation problem is transformed into a formula for maximizing platform utility as follows:
[0207]
[0208] where B t is the platform budget at the perception round R t .
[0209] The first constraint condition confirms that the total incentive expenditure of the perception platform at the perception round R t does not exceed the budget B t . The second constraint condition confirms that each task T j is assigned to at least one participant for completion.
[0210] Specifically, through the many-to-many mode of "task allocation to multiple participants, and participants participating in multiple tasks", the limitations of traditional single task single allocation are broken, and flexible scheduling of sensing resources is realized. The mechanism of automatically entering the next round of allocation for uncompleted tasks can dynamically adapt to changes in participant state (such as device offline, power consumption) or price fluctuations, avoiding task grounding, and the participant utility function directly links the income and energy cost, prompting it to consider task income and device resource limitations when pricing; The platform utility takes the quality of perception as the core, quantifies the global benefit combined with the incentive cost, and through AHP, multi-dimensional characteristics such as sensing resource quantity, energy consumption, and task quantity are included in task complexity evaluation. At the same time, the budget constraint and the condition of at least one task allocation in the platform utility optimization problem are introduced, which not only avoids over-spending on incentive spending, but also ensures the completeness of task coverage.
[0211] In order to solve the problems in the prior art that the dynamic pricing strategy is not intelligent enough, lacks a real-time adjustment mechanism based on deep reinforcement learning, the task allocation mode is limited to traditional single task single allocation, the flexibility is poor, the modeling and adaptation ability to task geographical distribution, participant trajectory and dynamic state change is insufficient, and the evaluation index is single and does not comprehensively consider the multi-dimensional benefits such as participant utility, platform utility and task completion rate, please refer to Figures 1-5 The embodiment provides the following technical solutions:
[0212] Also includes:
[0213] Based on the dynamic pricing optimization phase and the task allocation optimization phase, it also includes simulation settings, comparison mechanism, evaluation index, and comparison and analysis;
[0214] The simulation settings are: selecting a specific area for simulation experiment, randomly determining 10 to 30 task points in the specific area, each task point being located on the road to simulate the geographical distribution of real perception tasks, the selection of task points being based on road network, and each task having corresponding sensing resource requirements;
[0215] In the specific area, 10 to 30 taxis with long trajectories are selected as participants, and the participants plan routes and complete tasks according to their interest in the tasks, each participant being an agent that learns how to adjust the price according to the task complexity, historical pricing data and platform feedback through the strategy network and value network;
[0216] The strategy network and the value network each consist of an input layer, three hidden layers and an output layer, wherein the number of hidden layer neurons is 512, 256 and 32 respectively, the strategy network and the value network use the ReLU function as the activation of all hidden layers, and the value network uses the Tanh function as the activation of the output layer to generate the strategy. The simulation parameters are as shown in the figure:
[0217]
[0218] The comparison mechanism is that in the comparative experiment, the dynamic pricing strategy generation mechanism of deep Q learning, the dynamic pricing optimization mechanism of historical participation rate and income, and the random pricing and task allocation mechanism are used, and the random pricing and task allocation mechanism is used as the comparison mechanism.
[0219] Specifically, a specific area (latitude 39.9020°-39.9170°, longitude 116.3900°-116.4200°) of the Beijing taxi trajectory dataset is selected, the task points are randomly distributed based on the road network, the geographical constraints of the tasks in the city perception scene are simulated, 10-30 taxis with longer trajectories are selected as participants, the moving trajectories naturally cover multiple task points, and the route can be planned according to the interest (based on the path and task point overlap degree). This design makes the participant behavior closer to the real user (such as the driver's selective participation in the on-the-way task), improves the reference value of the simulation to the actual MCS system, uses an “input layer-3 hidden layer (512 / 256 / 32 neurons)-output layer” structure, a ReLU activation function to alleviate gradient disappearance, and a Tanh output layer to limit the price within a reasonable range. The total number of sets is 1000, and the time step is 50 to ensure that the model converges fully; a soft update factor of 0.01 and a small learning rate (strategy 0.001, value 0.002) stabilize the training process and avoid the collapse of the strategy due to parameter oscillation. The comparison mechanism uses the same simulation environment (number of tasks, budget, and participant size), only the core strategy is different, ensuring that the experimental results are due to algorithm design rather than scene bias. The average participant utility and platform utility balance individual and global interests, the task completion rate directly measures system efficiency, and the TSI-MADRL has a completion rate of 0.85 when the number of tasks is 30, which is significantly improved compared to the Random mechanism (0.6). The average remaining budget of the platform is used to evaluate the resource consumption efficiency. The remaining budget of TSI-MADRL is 12% lower than that of RADP-EWAB, indicating that it can make full use of the budget to improve the utility and avoid idle funds.
[0220] Also includes:
[0221] The evaluation index is introduced, including the average participant utility, the average platform utility, and the task completion rate.
[0222] Comparison and analysis: According to the first analysis chart, when the number of tasks is 10 and the platform budget is 25, the average participant utility changes under different participant quantities.
[0223] According to the second analysis chart, when the number of tasks is 10 and the platform budget is 25, the average platform utility changes under different participant quantities.
[0224] According to the third analysis chart, when the number of tasks is 10 and the platform budget is 25, the task completion rate changes under different numbers of participants.
[0225] Specifically, through the dynamic pricing mechanism and the MADDPG algorithm, the participants can autonomously adjust the unit price according to the energy consumption cost and the task income to maximize the utility. This mechanism avoids the decline of willingness to participate caused by "price solidification", especially when the number of tasks is fixed, the participants can adapt to task competition through policy iteration to maintain the utility level. The average platform utility takes the perception quality contribution as the core, and optimizes the allocation combined with the budget constraint to ensure the maximization of global benefits under limited resources. Comparative analysis shows that when the number of participants increases, the platform utility also increases synchronously, indicating that the resource pool of multiple participants can enable the platform to select the combination of "high perception quality / low price" to improve resource utilization. The task completion rate index reflects the robustness of the multi-round allocation mechanism: unfinished tasks automatically enter the next round of allocation, avoiding the shelving of tasks caused by participant state fluctuations (such as offline, energy consumption exceeding the limit), and ensuring that the total incentive expenditure does not exceed the budget through constraint conditions, while dynamically adjusting the task allocation strategy through multi-round allocation. This mechanism avoids the problems of "budget overruns" or "task omissions" in traditional solutions, especially in scenarios such as city monitoring and disaster rescue that are sensitive to budget, which can ensure task integrity under limited resources. The stability of the index under different numbers of participants and tasks in the comparative analysis shows that the scheme can dynamically adapt to environmental changes through the combination of MADDPG algorithm and AHP. This adaptability enables the scheme to remain efficient in heterogeneous environments (such as differences in participant device capabilities and varying task types), improving the system's generalization ability. The average participant utility, platform utility, and other indicators provide clear guidance for system optimization: participants can adjust the pricing strategy based on historical utility feedback, and the platform can optimize the allocation model based on the task completion rate and the remaining budget.
[0226] It should be noted that, in this text, relational terms such as first and second are used merely to distinguish one entity or action from another, without necessarily requiring or implying any actual such relationship or order between or among the entities or actions. Also, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0227] Although embodiments of the present application have been shown and described, it will be understood by those having ordinary skill in the art that various changes, modifications, alternatives, and variations can be made thereto without departing from the principles and spirit of the application.
Claims
1. A two-stage motivation method based on multi-agent deep reinforcement learning, characterized by: include: Dynamic quotation optimization stage and task allocation optimization stage; The dynamic bidding optimization phase uses the analytic hierarchy process to evaluate the complexity of the task set, providing a reference for participants' bids. It combines task complexity with historical bidding data and uses multi-agent deep deterministic policy gradient technology to train the optimal unit bidding strategy. The task allocation optimization stage is to select key participants, confirm the core requirements of the task, and then optimize the task allocation based on the participants' satisfaction.
2. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 1, characterized in that: Also includes: Each task is assigned to multiple mobile participants, and each mobile participant participates in multiple sensing tasks simultaneously. The sensing tasks include task broadcasting, mobile participant bidding, task allocation, task execution, and reward. In the sensing area, the sensing platform performs the task set T = {T1, T2, ..., T m } is managed, where each sensing task has the required total sensing resource amount, and each task T j ∈T can be divided into multiple perception rounds R = {R1, R2, ..., R k }Execute in; In each round of perception R t ∈R, each participant performs task T j The corresponding sensing resource quantity q needs to be collected j ; The sensing platform broadcasts to the participant set P = {P1, P2, ..., P n }Publish task information, each participant P i ∈P autonomously selects a task subset based on its own resource capabilities and preferences Participate and upload bidding information in is the total amount of selected sensing resources, is the total energy consumption, is the number of tasks, Quote for resource units; The perception platform optimizes task allocation based on the collected bidding information and reasonably allocates tasks to a group of participants. After the task allocation is completed, the participants actively perform the perception tasks according to the task requirements and upload the collected perception resource volume. and receive corresponding remuneration and rewards; For unfinished tasks These tasks will be included in the next round of R t+1 , continue to assign and execute until all tasks are completed; Among them, T j Represented as a perception task; R t Represented as perception round; P i represented as participants; Represented as the perception wheel R t When participant P i The selected set of tasks; Represented as the perception wheel R t When participant P i Uploaded bidding information.
3. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 2, characterized in that: The dynamic quotation optimization stage includes: The perception energy consumption is calculated as follows: When performing multiple perception tasks, mobile participants use different types of sensors, including radar, GPS, and cameras; Among them, GPS records data at a rate of bits per second, and the camera is used to obtain image data, which generates image data at millions of bits per second; If δ i For participant P i The perception rate, then for the perception resource size q′ j , participant P i Complete the perception task T j Time required The formula is as follows: Participant P i Complete the perception task T j The formula for perceived energy consumption is as follows: Among them, p s Indicates unit energy consumption; Represented as participant P i For the perception task T j Perception energy consumption and computation energy consumption; q j Represented as performing perception task T j the amount of perceived resources required; The calculation of energy consumption is: Participant P i When processing sensory data, computing resources are consumed; If χ represents the number of CPU cycles required to process bit-aware data, then participant P i Complete the perception task T j Data calculation time The formula is as follows: Among them, f i Indicates the computing resources of the device, that is, the CPU frequency; Processing perception task T j The formula for the raw data energy consumption is as follows: Where κ represents the effective switching capacitance; Represented as participant P i For the perception task T j The perceived energy consumption and computational energy consumption.
4. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 3, characterized in that: The dynamic quote optimization phase also includes: Mobile Participant P i In the perception wheel R t By selecting and executing a set of tasks The formula for the utility obtained is as follows: in, Represents participant P i Perception Task T j The actual benefits obtained; Represents participant P i In the perception task T j The perceived resources provided in; c represents the unit cost of energy consumption; Represents participant P i Perform perception task T j Actual energy consumption; At the same time, the benefit of the perception platform is the perceived quality contribution provided by the participants in completing the perception task. According to the heterogeneous characteristics of the participants' devices, different participants contribute differently to the perceived quality. The participants’ perceived quality contributions on different tasks were evaluated using As a participant P i Task T j perceived contribution; Among them, the weight ω ij Reflects the participant P i For Task T j The contribution of units’ perceived resources; The total revenue of the platform is composed of the cumulative perceived quality contributions achieved in all tasks, and the total cost is the sum of the incentive rewards paid by the platform to all participants. The utility of the perceived platform is the difference between the total revenue and the total cost. The difference formula is as follows: Participant P i The goal is to select a set of tasks Develop an appropriate unit quotation To maximize its net income, player P i The formula for the optimal bidding problem is as follows: Among them, θ ij Represents the perception task T j Is it assigned to participant P? i , if assigned, then θ ij =1, otherwise θ ij =0; Participant P i Selected Task Set If there is an assigned task in the task set, then all tasks in the task set selected by the participant have been assigned, that is, θ i =1; The constraint is that the total energy consumption of the participant to complete the task set does not exceed the energy budget of his device Among them, θ ij Indicates whether the platform selects participant P i The index value of the perception task T j Whether P i implement.
5. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 4, characterized in that: The dynamic quote optimization phase also includes: The hierarchical analysis method is used to determine the complexity of the multidimensional feature task combination TS i The influence of is as follows: in, is the maximum total amount of perceived resources uploaded by all participants; is the maximum energy consumption of all participants in performing the task; is the maximum number of tasks selected by all participants; w1, w2, w3 are weights for evaluating the relative importance of different parameters, and the sum of w1, w2, w3 is 1; In the AHP method, X1, X2, and X3 are used to represent the perceived resource amount, energy consumption, and number of tasks of the task set selected by the participants, respectively; By constructing a pairwise comparison matrix A = (a ij ) 3×3 To measure the importance between different parameters, where each element a ij Represents X i Relative to X j The importance of the value is given by the proposed scale a ij Assign a suitable integer between 1 and 9, and change it according to different situations; when a ij >1, X i X j More important; when a ij When X<1, i Not as good as X j important; when a ij =1, X i and X j equally important; The pairwise comparison matrix is, a 13 =2 means that the amount of perceived resources X1 is more important than the number of tasks X2; After normalizing each column of A, we get A=(a ij ) 3×3 Normalized matrix, the formula is as follows: Weight vector W = (w1, w2, w3) T Computed by taking the mean of each row element of the normalized matrix.
6. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 5, characterized in that: The dynamic quote optimization phase also includes: Use MADDPG to formulate the optimal unit bids of participants to maximize their utility, including using MADDPG to formulate the optimal unit bids of participants to maximize their utility. MADDPG is used for the first stage of dynamic bidding strategy training. Through centralized training and distributed execution framework, each participant is allowed to observe the states and actions of other agents during training to optimize the strategy. In the perception wheel R t In the observation space, each intelligent entity will Data to output an action After completing the task according to the action, you will get the expected reward Then observe the changing environment and adjust the strategy; Among them, the state observation space of each participant Action Space and the reward function The definition is as follows: The observation space is: Based on the observation results of the first L rounds, the environmental changes of each participant are confirmed. t , participant P i The formula for the observation result is as follows: in, Represents the perception wheel R t Participant P i The complexity of the selected task combination; Represents the perception wheel R t-1 Participant P i The complexity of the selected task combination; Represents the perception wheel R t-1 Participant P i Unit quotation of Represents the perception wheel R t-1 Participant P i The utility of Represents the perception wheel R t-1 whether participants are selected by the perceived platform; The action space is: on the perception wheel R t In the i Observe the platform's decision in the first L rounds of perception and give the unit quote for the current sensing wheel based on the output of the Actor network. The formula is as follows: Among them, π i Represents an Actor network; Represents the gradient optimization parameters of the Actor network; The reward function is: according to the action Participant P i In the perception wheel R t The reward obtained is expressed using the utility function, as follows: Among them, θ ij Represents the perception task T j Is it assigned to participant P? i ; For the perception wheel R t After selecting the task set, participants get the best bid based on MADDPG decision Then upload the bidding record containing the latest quotation information Finally, the perception platform allocates tasks based on the bidding information.
7. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 6, characterized in that: The task allocation optimization phase also includes: According to the simplified scenario of one task and multiple participants, the goal of the perception platform is to assign tasks to participants to maximize the platform utility while meeting the condition of not exceeding the given budget limit. The optimization problem formula of the simplified scenario of perception platform efficiency is as follows: Among them, θ′ i Indicates whether the sensing platform assigns a task to a participant; Represents the perception wheel R t Participant P i The contribution provided to the perception platform by performing a certain task; Represents the perception wheel R t Participant P i Reward for completing a task; B t Represents the perception wheel R t The platform budget to complete a task; Based on the goal of maximizing platform utility, the task allocation algorithm is used to calculate the process. The calculation steps are as follows: S1: Conduct a detailed analysis of the participants’ task coverage capabilities and include participants who meet the task requirements into the candidate set; If there is only one participant in the candidate set, it is identified as the key participant. In the absence of other options and if the budget allows, the key participant will be given priority, and the budget and task coverage status will be adjusted accordingly; S2: For unassigned tasks, select candidates who can provide necessary coverage from the unselected participants and use the satisfaction index SA i , SA i The ratio of perceived contribution to bid price is as follows: S3: Rank the satisfaction of the unselected participants to ensure the optimal allocation within the budget limit, and select the top-ranked participants in turn until the budget is exhausted; According to the calculation steps of the task allocation algorithm, the updated θ i To complete the task allocation strategy, participants perform the assigned tasks and submit the collected perception data. Then, the MCS platform allocates resources based on the actual resource contribution of the participants. and the best unit quotation Payment Rewards The total reward for the selected participants cannot be greater than the platform budget.
8. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 7, characterized in that: The task allocation optimization phase also includes: The goal of the perception platform is to assign tasks to appropriate participants to optimize the platform utility. The task assignment problem is transformed into the problem of maximizing the platform utility as follows: Among them, B t For the perception wheel R t Platform budget at the time; The first constraint is to confirm that the perception platform is on the perception wheel R t The total incentive expenditure does not exceed budget B t , The second constraint is to confirm that each task T j Assigned to at least one participant to complete.
9. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 8, characterized in that: Also includes: Based on the dynamic quotation optimization stage and the task allocation optimization stage, it also includes simulation settings, comparison mechanisms, evaluation indicators, and comparison and analysis; The simulation setup involves selecting a specific area for the simulation experiment and randomly determining 10 to 30 task points within the area. Each task point is located on the road to simulate the geographical distribution of real-world perception tasks. The task points are selected based on the road network, and each task has corresponding perception resource requirements. In a specific area, 10 to 30 taxis with long trajectories were selected as participants. Participants planned routes and completed tasks based on their interest in the task. Each participant, acting as an agent, learned through a policy network and a value network how to adjust the quote based on task complexity, historical quote data, and platform feedback. The policy network and value network each consist of an input layer, three hidden layers, and an output layer. The number of neurons in the hidden layers is 512, 256, and 32, respectively. The policy network and value network use the ReLU function as the activation of all hidden layers, while the value network uses the Tanh function as the activation of the output layer to generate policies. The comparison mechanism is: in the comparative experiment, the dynamic quotation strategy generation mechanism of deep Q learning, the dynamic quotation optimization mechanism of historical participation rate and income, and the random quotation and task allocation mechanism are adopted, and the random quotation and task allocation mechanism is used as the comparison mechanism.
10. A two-stage motivation method based on multi-agent deep reinforcement learning according to claim 9, characterized in that: Also includes: Introducing evaluation indicators, including average participant utility, average platform utility, and task completion rate; The comparison and analysis are as follows: According to the first analysis chart, when the number of tasks is 10 and the platform budget is 25, the average participant utility changes under different numbers of participants; According to the second analysis chart, when the number of tasks is 10 and the platform budget is 25, the average platform utility changes under different numbers of participants; According to the third analysis chart, when the number of tasks is 10 and the platform budget is 25, the task completion rate changes under different numbers of participants.
Citation Information
Cited By
Computing power center multi-agent energy consumption collaboration method and system
CN121635654A