Online Sparse Crowdsensing Method for Collaborative Execution of Multiple Tasks
Through the online sparse group intelligence perception method based on hierarchical multi-agent reinforcement learning, a variety of relevant sparse group intelligence perception tasks are coordinated to solve the problem of insufficient utilization of data information in the prior art, and lower cost and higher quality data inference are achieved.
Patent Information
- Application Number
- CN202311259394.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively utilize the correlation between multiple data, resulting in insufficient utilization of data information in multi-task sparse group intelligence perception scenarios, high cost and low data inference quality.
The online sparse group intelligence perception method based on hierarchical multi-agent reinforcement learning is adopted. Through three parts, data collection, data inference and training data update, a variety of relevant sparse group intelligence perception tasks are performed in a coordinated manner. This method uses real sparse data to guide task execution, updates training data in real time, and uses a multi-agent reinforcement learning model with a hierarchical architecture to allocate budgets and coordinated task scheduling, capturing the spatiotemporal correlation between multiple task data.
By collaborating on performing multiple sparse group intelligence perception tasks, the perceived cost is reduced and the quality of data inference is improved, and more efficient data utilization and more accurate data inference are achieved.
Smart Images

Figure CN117236447B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to model online update, data inference, budget allocation and task collaborative scheduling for multi-task sparse crowd sensing, and particularly relates to a data collection method based on hierarchical multi-agent reinforcement learning for realizing online sparse crowd sensing for collaborative execution of multiple tasks. Background Art
[0002] Sparse crowd sensing needs to recruit participants to collect data from a small number of sensing regions and infer data in the remaining un-sensed regions, which can reduce the cost of data collection. Prior work in the prior art generally only considered the independent collection and inference of single-type data. However, in real-world scenarios, there may be multiple types of data that can complement each other in spatio-temporal information. Therefore, the existing solutions have the problem of insufficient utilization of data information. Summary of the Invention
[0003] Aiming at the defects and deficiencies existing in the prior art, the present invention fully considers the correlation between multiple types of data and collaboratively executes multiple relevant sparse crowd sensing tasks, thereby further reducing the sensing cost and improving the quality of data inference. In addition, in order to be closer to the real application scenario of sparse crowd sensing, the present invention uses real sparse data collected in real time to guide the execution of tasks in the case of only a small amount of sparse historical data. For this purpose, an online sparse crowd sensing method based on hierarchical multi-agent reinforcement learning is proposed. This method consists of three parts: data collection, data inference, and training data update. First, by mining the spatio-temporal variation law and correlation of real data, the present invention can update the training data using multi-task data collected in real time to keep the model continuously updated. Then, the present invention proposes a multi-agent reinforcement learning model with a hierarchical architecture, attempts to allocate appropriate budgets for each sensing task, and executes collaborative collection of multiple task data. Finally, the present invention uses a data inference network to capture the spatio-temporal correlation between multiple task data and infer data in the uncollected regions, thereby further reducing the sensing cost and improving the quality of data inference.
[0004] This method can collaboratively execute multiple relevant sparse crowd sensing tasks, thereby further reducing the sensing cost and improving the quality of data inference.
[0005] The technical solution specifically adopted by the present invention to solve its technical problems is as follows:
[0006] An online sparse crowd sensing method for collaborative execution of multiple tasks, comprising the following steps:
[0007] Step S1: Cold start by randomly collecting real data for several cycles;
[0008] Step S2: Based on the latest collected real perception data, calculate the data change trend to update the training data;
[0009] Step S3: Use the latest training data to update the data inference and data collection model parameters;
[0010] Step S4: The budget allocation agent allocates the total budget to each perception task;
[0011] Step S5: Multiple agents cooperate to collect data for different perception tasks;
[0012] Step S6: Determine whether all tasks have exhausted the budget. If yes, go to Step S7; otherwise, continue with Step S5;
[0013] Step S7: Utilize the correlations of multiple task data to jointly infer the data in other uncollected regions;
[0014] Step S8: Enter the loop of the next cycle and return to Step S2.
[0015] Furthermore, the specific operation process of Step S1 is as follows:
[0016] Step S101: Divide the complete target perception area using the grid method into m sub-regions;
[0017] Step S102: Allocate the budget evenly to each perception task. Each data collection agent is responsible for collecting the data of one task, and the agents are randomly scheduled to collect the perception data of different sub-regions;
[0018] Step S103: Return the perceived data to the scheduling center until the cold start cycle ends.
[0019] Furthermore, the operation process of calculating the data change trend to update the training data based on the latest collected real perception data in Step S2 is as follows:
[0020] Step S201: Sequentially take out each task from all perception tasks and calculate the mean of the training data in the previous cycle, where TD k represents the training data of the k-th task data:
[0021] avg(TD k [:,-1])
[0022] Step S202: Calculate the mean of the data actually collected in the current cycle, where SparseGTD k stores the data actually collected for the k-th task:
[0023] avg(SparseGTD k [:,-1])
[0024] Step S203: Update the data of the sub-region that has not been collected in the training data. When the data of this sub-region i has not been collected by other tasks, update the training data of this sub-region with the following formula:
[0025]
[0026] Step S204: When the data of sub-region i is collected by other tasks, combine the change trend of other relevant data and update the training data of this sub-region with the following formula. Where λ t and represent the gradual change and similarity weights respectively, and
[0027]
[0028] Step S205: Repeat Steps S201 to S204 for all perception tasks until the training data is updated completely.
[0029] Furthermore, the operation process of updating the data inference and data collection model parameters using the latest training data in Step S3 is as follows:
[0030] Step S301: Train the data inference network using the training set updated in Step S2;
[0031] Step S302: Train the budget allocation agent in the data collection agent using the training set updated in Step S2. The parameter update formula is the reinforcement learning DQN algorithm update formula, as shown below; where is the reward obtained by the budget allocation agent for each exploration, and are the current state and the next state respectively, budget k is the budget obtained by task k, and γ is the reinforcement learning decay factor:
[0032]
[0033] Step S303: Train the region selection agent in the data collection agent. The parameter update formula is the multi-agent reinforcement learning QMIX algorithm update formula, as shown below; the number of region selection agents is equal to the number of perception tasks, and each region selection agent is responsible for collecting the data of one perception task:
[0034]
[0035] where is the reward obtained by all agents, o k,t is the local observation state of each agent, and They are the current global state and the next global state respectively.
[0036] Furthermore, the operation process of the budget allocation agent in step S4 for allocating the total budget to each sensing task is as follows:
[0037] Step S401: Obtain the current budget allocation status and allocate the budget to each sensing task in turn;
[0038] Step S402: According to the current budget allocation status, use the Q-value estimation network of the budget allocation agent to evaluate all budget allocation situations of the k-th task being allocated currently:
[0039]
[0040] Step S403: According to the output Q-value, take out the index of the maximum value, which is the budget allocated to task k this time;
[0041] Step S404: Update the budget allocation status, and repeat steps S401 to S403 until the budget allocation of all tasks is completed.
[0042] Furthermore, the operation process of multiple agents collaborating to collect data of different sensing tasks in step S5 is as follows:
[0043] Step S501: All region selection agents obtain their current local observation states and evaluate the maximum Q-values that can be obtained by collecting each sub-region respectively:
[0044] Q k = evalAgent C (o k,t )
[0045] Step S502: All region selection agents select the sub-region that can obtain the maximum Q-value. If the budget is sufficient, collect the sensing data of this sub-region, and the agents with exhausted budgets do not perform any operations.
[0046] Furthermore, step S6 is specifically as follows:
[0047] Step S601: After all region selection agents have collaboratively collected sensing data once, deduct the budget spent on this collection respectively;
[0048] Step S602: Judge whether all region selection agents have exhausted their budgets. If there are still agents with remaining budgets, return to step S5.
[0049] Furthermore, the operation process of using the correlations of multiple task data to jointly infer the data of other uncollected regions in step S7 is as follows:
[0050] Step S701: Input each type of perception task data into the graph convolutional neural network respectively to extract spatial features; where A represents the adjacency matrix. To make A obtain a self-connection structure, after adding an identity matrix, we get and further normalize it to get where, is the degree matrix; W 0 and W 1 represent the weight matrices of the first layer and the second layer respectively, and σ(·) represents the activation function:
[0051]
[0052] Step S702: The features extracted by each graph neural network are transposed and concatenated into X GRU , and then input into the gated recurrent unit GRU in sequence:
[0053]
[0054] r t =σ(W r [h t-1 ,X GRU [t,:]]+b r )
[0055] z t =σ(W z [h t-1 ,X GRU [t,:]]+b z )
[0056] c t =tanh(W c [r t *h t-1 ,X GRU [t,:]]+b c )
[0057] h t =(1 - z t )*h t-1 +z t *c t
[0058] Step S703: Through one layer of self-attention mechanism, capture the dependency relationships and shared information between different tasks and different moments:
[0059] Q, K, V = W q / k / v h
[0060]
[0061] Step S704: Output the data inference values of all perception tasks at the current moment through the last fully-connected layer.
[0062] Further, step S8 is specifically as follows: The collection budgets of all areas in the current region are exhausted, and the data of the remaining uncollected areas are inferred; After the tasks of the current cycle are completed, return to step S2 to enter the tasks of the next cycle.
[0063] Compared with the prior art, the present invention and its preferred solutions are mainly used to collaboratively execute multiple sparse crowd-sourced sensing tasks. In order to cope with complex data change trends and real data sparsity problems, this method uses the correlation of multiple types of data to update the training data to keep the model continuously updated. Through a multi-agent reinforcement learning model with a hierarchical architecture, the budget allocation and collaborative task collection of multiple tasks are completed. Capture the spatio-temporal correlation of the sparse data actually collected, and infer the data of other uncollected areas. This method fully considers the correlation between multiple types of data, collaboratively executes multiple relevant sparse crowd-sourced sensing tasks, thereby further reducing the sensing cost and improving the data inference quality. Brief Description of the Drawings
[0064] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0065] Figure 1 It is the overall flowchart of the embodiment of the present invention. Specific Embodiments
[0066] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0067] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0068] This embodiment provides an online sparse crowd-sourced sensing method for collaboratively executing multiple tasks, as Figure 1 shown, including the following steps:
[0069] Step S1: Cold start by randomly collecting real data for a small number of cycles;
[0070] Step S2: Based on the latest real sensing data collected, calculate the data change trend and update the training data;
[0071] Step S3: Update the data inference and data collection model parameters using the latest training data;
[0072] Step S4: The budget allocation agent allocates the total budget to each sensing task;
[0073] Step S5: Multiple agents cooperate to collect data for different perception tasks;
[0074] Step S6: Determine whether all tasks have exhausted the budget. If yes, go to S7; otherwise, continue with S5;
[0075] Step S7: Utilize the correlations of various task data to jointly infer the data of other uncollected areas;
[0076] Step S8: Enter the loop of the next cycle and return to S2.
[0077] In this embodiment, the operation process of cold starting by randomly collecting real data for a small number of cycles is as follows:
[0078] Step S101: Divide the complete target perception area using the grid method into m sub-areas;
[0079] Step S102: Randomly dispatch agents to different sub-areas to complete two tasks: 1. Sense the pedestrian flow in the area; 2. Execute tasks to serve the people in the area;
[0080] Step S103: Return the sensed pedestrian flow data to the dispatch center, and pre-fill the data of the remaining areas with 0.
[0081] In this embodiment, the operation process of calculating the data change trend and updating the training data based on the latest collected real perception data is as follows:
[0082] Step S201: Take out each type of task from all perception tasks in turn, and calculate the mean of the training data in the previous cycle, where TD k represents the training data of the k-th type of task data:
[0083] avg(TD k [:,-1])
[0084] Step S202: Calculate the mean of the real data collected in the current cycle, where the data stored in SparseGTD k is the real data collected for the k-th type of task:
[0085] avg(SparseGTD k [:,-1])
[0086] Step S203: Update the data of the uncollected sub-areas in the training data. When the data of sub-area i is not collected by other tasks, update the training data of this sub-area using the following formula:
[0087]
[0088] Step S204: When other tasks collect data for sub-region i, combine with the changing trends of other relevant data and update the training data for this sub-region using the following formula. Here, λ t and represent the gradual change weight and the similarity weight respectively, and
[0089]
[0090] Step S205: Repeat S201 to S204 for all perception tasks until the training data is updated completely.
[0091] In this embodiment, the operation process of updating the data inference and data collection model parameters using the latest training data is as follows:
[0092] Step S301: Train the data inference network using the training set updated in Step S2;
[0093] Step S302: Train the budget allocation agent in the data collection agent using the training set updated in Step S2. The parameter update formula is the reinforcement learning DQN algorithm update formula, as shown below. Here, is the reward obtained by the budget allocation agent for each exploration, and are the current state and the next state respectively, budget k is the budget obtained for task k, and γ is the reinforcement learning decay factor:
[0094]
[0095] Step S303: Train the region selection agent in the data collection agent. The parameter update formula is the multi-agent reinforcement learning QMIX algorithm update formula, as shown below. The number of region selection agents is equal to the number of perception tasks, and each region selection agent is responsible for collecting data for one perception task. Here, is the reward obtained by all agents, o k,t is the local observation state of each agent, and are the current global state and the next global state respectively:
[0096]
[0097] In this embodiment, the operation process of the budget allocation agent distributing the total budget to each perception task is as follows:
[0098] Step S401: Obtain the current budget allocation state and allocate budgets to each perception task in sequence;
[0099] Step S402: According to the current budget allocation status, use the Q-value estimation network of the budget allocation agent to evaluate all budget allocation situations of the k-th task being allocated currently:
[0100]
[0101] Step S403: According to the output Q-value, extract the index of the maximum value, which is the budget allocated for task k this time;
[0102] Step S404: Update the budget allocation status, and repeat S401 to S403 until the budget allocation of all tasks is completed;
[0103] In this embodiment, the operation process of multiple agents collaborating to collect different perception task data is as follows:
[0104] Step S501: All region selection agents obtain their current local observation states, and respectively evaluate the maximum Q-value that can be obtained by collecting each sub-region:
[0105] Q k = evalAgent C (o k,t )
[0106] Step S502: All region selection agents select the sub-region that can obtain the maximum Q-value. If the budget is sufficient, collect the perception data of this sub-region, and the agents with exhausted budgets do not perform any operations;
[0107] In this embodiment, the operation process of determining whether all tasks have exhausted the budget and going to S7 if the judgment is yes, otherwise continuing the operation process of S5 is as follows:
[0108] Step S601: After all region selection agents have collaboratively collected perception data once, deduct the budget spent on this collection respectively;
[0109] Step S602: Determine whether all region selection agents have exhausted the budget. If there are still agents with remaining budgets, return to Step S5;
[0110] In this embodiment, the operation process of jointly inferring the data of other uncollected regions by using the correlation of multiple task data is as follows:
[0111] Step S701: Input each type of perception task data into the graph convolutional neural network to extract spatial features. Among them, A represents the adjacency matrix. In order to make A obtain a self-connection structure, add an identity matrix to get Further normalization can obtain is the degree matrix. W 0 and W 1Respectively represent the weight matrices of the first layer and the second layer, and σ(·) represents the activation function:
[0112]
[0113] Step S702: The features extracted by each graph neural network will be transposed and concatenated into X GRU , and are sequentially input into the gated recurrent unit GRU:
[0114]
[0115] r t =σ(W r [h t-1 ,X GRU [t,:]]+b r )
[0116] z t =σ(W z [h t-1 ,X GRU [t,:]]+b z )
[0117] c t =tanh(W c [r t *h t-1 ,X GRU [t,:]]+b c )
[0118] h t =(1 - z t )*h t-1 +z t *c t
[0119] Step S703: Through one layer of self-attention mechanism, capture the dependencies and shared information between different tasks at different times:
[0120] Q, K, V = W q / k / v h
[0121]
[0122] Step S704: Through the last fully-connected layer, output the data inference values of all perception tasks at the current moment.
[0123] In this embodiment, the operation process of entering the next cycle and returning to S2 is as follows: The collection budgets of all current regions are exhausted, and the data of the remaining uncollected regions are inferred. After the tasks of the current cycle are completed, return to S2 to enter the tasks of the next cycle.
[0124] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0125] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0126] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realize the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0128] As described above, it is only the preferred embodiments of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
[0129] This patent is not limited to the above-mentioned best implementation modes. Anyone inspired by this patent can derive various other forms of online sparse crowd intelligence perception methods for synergistically performing multiple tasks. All equivalent changes and modifications made within the scope of the patent application of this invention shall fall within the scope covered by this patent.
Claims
1. An online sparse crowd sensing method for collaborative execution of multiple tasks, characterized in that, it includes the following steps: Step S1: Cold start by randomly collecting real data for several cycles; Step S2: Based on the latest collected real sensing data, calculate the data change trend to update the training data; Step S3: Use the latest training data to update the data inference and data collection model parameters; Step S4: The budget allocation agent allocates the total budget to each sensing task; Step S5: Multiple agents collaborate to collect data for different sensing tasks; Step S6: Determine whether all tasks have exhausted the budget. If yes, go to Step S7; otherwise, continue with Step S5; Step S7: Utilize the correlation of multiple task data to jointly infer the data in other uncollected regions; Step S8: Enter the loop of the next cycle and return to Step S2; The operation process of calculating the data change trend to update the training data based on the latest collected real sensing data in Step S2 is: Step S201: Take out each type of task from all perception tasks in sequence, and calculate the mean of the training data in the previous cycle, where TD k represents the training data of the k-th type of task data: avg(TD k [:,-1]) Step S202: Calculate the mean value of the data actually collected in the current period, where the data actually collected for the k-th task is stored in SparseGTD k : avg(SparseGTD k [:,-1]) Step S203: Update the data of the uncollected sub-region in the training data. When other tasks have not collected the data of sub-region i, use the following formula to update the training data of this sub-region: Step S204: When other tasks collect data for sub-region i, the training data for this sub-region is updated using the following formula in combination with the change trends of other relevant data; where λ t and represent the gradual change and similarity weights respectively, and Step S205: Repeat Steps S201 to S204 for all sensing tasks until the training data is updated; The operation process of using the latest training data to update the data inference and data collection model parameters in Step S3 is: Step S301: Use the training set updated in Step S2 to train the data inference network; Step S302: Train the budget allocation agent in the data collection agent using the training set updated in Step S2. The parameter update formula is the update formula of the reinforcement learning DQN algorithm as follows; where is the reward obtained by the budget allocation agent during each exploration, and are the current state and the next state respectively, and budget k is the budget obtained for task k, and γ is the reinforcement learning decay factor: Step S303: Train the region selection agent in the data collection agent. The parameter update formula is the multi-agent reinforcement learning QMIX algorithm update formula, as follows; The number of region selection agents is equal to the number of sensing tasks, and each region selection agent is responsible for collecting the data of one sensing task: where is the reward obtained by all agents, o k,t is the local observation state of each agent, and are the current global state and the next global state respectively; The operation process of utilizing the correlation of multiple task data to jointly infer the data in other uncollected regions in Step S7 is: Step S701: Input the data of each sensing task into the graph convolutional neural network respectively to extract spatial features; where A represents the adjacency matrix. To endow A with a self - connection structure, after adding an identity matrix, we get After further normalization, we obtain where is the degree matrix; W 0 and W 1 represent the weight matrices of the first layer and the second layer respectively, and σ(·) represents the activation function: Step S702: The features extracted by each graph neural network are transposed and concatenated into X GRU , and are sequentially input into the gated recurrent unit GRU: r t = σ(W r [h t-1 , X GRU [t, :]] + b r ) z t = σ(W z [h t-1 , X GRU [t, :]] + b z ) c t = tanh(W c [r t *h t-1 , X GRU [t, :]] + b c ) h t = (1 - z t ) * h t-1 + z t * c t Step S703: Pass through a self-attention mechanism to capture the dependence and shared information between different tasks at different times: Q, K, V = W q / k / v h Step S704: Output the data inference values of all sensing tasks at the current moment through the last fully connected layer.
2. The online sparse crowd sensing method for collaborative execution of multiple tasks according to claim 1, characterized in that: The specific operation process of Step S1 is: Step S101: Divide the complete target sensing area using the grid method into m sub-regions; Step S102: Allocate the budget evenly to each sensing task. Each data collection agent is responsible for collecting the data of one task, and the random scheduling agent collects the sensing data of different sub-regions; Step S103: Return the sensed data to the scheduling center until the cold start cycle ends.
3. The online sparse crowd sensing method for collaborative execution of multiple tasks according to claim 1, characterized in that: The operation process of the budget allocation agent allocating the total budget to each sensing task in Step S4 is: Step S401: Obtain the current budget allocation status and allocate budgets to each sensing task in sequence. Step S402: According to the current budget allocation status, use the Q-value estimation network of the budget allocation agent to evaluate all budget allocation situations of the k-th task being allocated currently. Step S403: According to the output Q-value, extract the index of the maximum value, which is the budget allocated to task k this time. Step S404: Update the budget allocation status, and repeat Step S401 to Step S403 until the budget allocation of all tasks is completed.
4. The online sparse crowd sensing method for collaborative execution of multiple tasks according to claim 3, characterized in that: The operation process of multiple agents collaboratively collecting data for different sensing tasks in Step S5 is as follows: Step S501: All region selection agents obtain their current local observation states and respectively evaluate the maximum Q-value that can be obtained by collecting each sub-region. Q k = evalAgent C (o k,t ) Step S502: All region selection agents select the sub-region that can obtain the maximum Q-value. If the budget is sufficient, collect the sensing data of this sub-region, and the agents with exhausted budgets do not perform any operations.
5. The online sparse crowd sensing method for collaborative execution of multiple tasks according to claim 4, characterized in that: Step S6 is specifically as follows: Step S601: After all region selection agents collaboratively collect sensing data once, deduct the budgets spent on this collection respectively. Step S602: Determine whether all region selection agents have exhausted their budgets. If there are still agents with remaining budgets, return to Step S5.
6. The online sparse crowd sensing method for collaborative execution of multiple tasks according to claim 1, characterized in that: Step S8 is specifically as follows: The collection budgets of all regions in the current area are exhausted, and the data of the remaining uncollected regions is inferred; the task execution of the current cycle is completed, and return to Step S2 to enter the next cycle task.
Citation Information
Patent Citations
Sparse crowd sensing online user recruitment method based on reinforcement learning
CN114912029A
Data collection method based on matrix completion and reinforcement learning in crowd-sourcing network
CN115510753A