Credibility-oriented MCS excitation method based on LSTM and PPO

By combining the incentive methods of LSTM and PPO, the problem of dynamic change in the existing MCS incentive mechanism is solved, and the problem of unconsidered time series long-term dependence in the maximization of participants' utility rewards and the improvement of data quality are achieved.

CN120216900APending Publication Date: 2025-06-27HENAN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510137339.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-20
Filing Date
2025-02-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing mobile group intelligence perceived MCS incentive mechanism has two problems in optimizing the long-term utility of participants: one is that fixed credibility is used to evaluate utility rewards in different time periods, ignoring the dynamic changes in credibility; the other is that it only relies on historical data at a single moment, and does not consider the long-term dependence of the time series.

Method used

The incentive method based on LSTM and PPO is adopted, and LSTM is used to capture long-term dependencies, improve prediction accuracy and decision stability, and derive the optimal perceived time strategy for each participant based on PPO to maximize its utility reward.

Benefits of technology

By dynamically updating participants’ credibility and optimizing perceived time strategies, the utility rewards are maximized, and the motivation and data quality of participants are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216900A_ABST
    Figure CN120216900A_ABST
Patent Text Reader

Abstract

The invention provides a credibility-oriented MCS excitation method based on LSTM and PPO, which can model a perception decision process of a participant as a non-cooperative game, and adopts a Markov decision process MDP to describe behaviors of the participant. Under the condition that global information is not known, an incentive model LSTM-PPO combining a long short-term memory network LSTM and a near-end strategy optimization algorithm PPO is utilized to formulate a most reasonable and effective perception duration strategy for each participant so as to maximize utility rewards. And after the task is completed, the credibility of the participant is dynamically updated by evaluating the quality of the uploaded data, so that the utility reward of the participant in the next stage is adjusted. On a real data set, a large number of simulation experiments are carried out on CIM-LP and other existing incentive mechanisms. Results show that the CIM-LP mechanism enables the average utility of the participants to be improved by 19.3% and the task completion rate to be improved by 12.8%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mobile crowd sensing, and in particular to a credibility-oriented MCS incentive method based on LSTM and PPO. Background Art

[0002] Mobile crowd sensing (MCS) is an emerging sensing paradigm that uses widely distributed intelligent mobile devices, such as sensing devices, to perform large-scale sensing of an open, dynamic, and complex physical environment. This sensing mode guides and feeds back the sensing group by intelligently analyzing a large amount of collected data, continuously emerging group intelligence, so as to assist in comprehensive decision-making. Compared with traditional methods, MCS systems exhibit advantages such as multi-source heterogeneous sensing data, wide and uniform coverage, strong scalability, and versatility. These characteristics have enabled MCS to be widely applied in fields such as smart city construction, intelligent transportation management, environmental monitoring, noise detection, facility construction, and public safety.

[0003] In the field of mobile crowd sensing (MCS), the design of the incentive mechanism has a significant impact on improving the participation enthusiasm of participants and the quality of the uploaded sensing data. However, there are two problems with the existing incentive mechanisms in optimizing the long-term utility of participants: one is that a fixed credibility is used to evaluate the utility rewards at different time periods, ignoring the dynamic changes of credibility; the other is that only the historical data at a single moment is relied on, without considering the long-term dependence of the time series. These problems will result in the utility rewards of participants not matching their actual contributions, thereby affecting their enthusiasm and data quality. Summary of the Invention

[0004] The purpose of the present invention is to provide a credibility-oriented MCS incentive method based on LSTM and PPO, which can use LSTM to capture long-term dependence relationships to improve prediction accuracy and decision-making stability, so as to derive an optimal sensing duration strategy for each participant based on PPO to maximize their utility rewards.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is: a credibility-oriented MCS incentive method based on LSTM and PPO, including the following steps: Step 1, the requester publishes a task on the sensing platform, and the sensing platform distributes the task to potential participants in the social network; Step 2, the distributed participants will actively choose whether to accept the sensing task. If accepted, they will upload their willingness to participate, and the sensing platform will select an optimal group of participants; Step 3, the participants collect sensing data and upload it. The sensing decision-making process of the participants is modeled as a non-cooperative game, and the Markov decision process (MDP) is used to describe their behavior; Step 4: Without knowing the global information, use the incentive model LSTM-PPO that combines the long short-term memory network (LSTM) and the proximal policy optimization algorithm (PPO) to formulate the most reasonable and effective perception duration strategy for each participant to maximize the utility reward. Step 5: The perception platform distributes rewards according to the effort level of the participants. The perception platform integrates the perception data and sends it to the requester. Step 6: After the task is completed, the credibility of the participants is dynamically updated by evaluating the quality of the uploaded data, thereby adjusting their utility rewards in the next stage.

[0006] Preferably, the utility of the participant includes monetary utility and social utility.

[0007] Preferably, the perception decision-making process focuses on the participant selecting the optimal perception duration strategy through deep reinforcement learning based on the state information provided by the perception platform to maximize their utility reward.

[0008] Preferably, it further includes a policy update process that optimizes the policy network and the value network through the PPO algorithm to ensure that the decisions of the participants are continuously improved with the accumulation of experience, thereby enhancing the overall quality of the perception task and the enthusiasm of the participants.

[0009] Preferably, the non-cooperative game maximizes its utility by optimizing the perception duration decision of each participant.

[0010] Preferably, the Markov decision process is represented as a quadruple including state, action, state transition probability matrix, and reward.

[0011] Preferably, the Markov decision process first observes the current state of the environment, selects the optimal strategy based on this state to execute an action; then, analyzes the impact of the action on the environment, obtains the next state based on this impact and the reward given by the environment after the action is implemented; finally, updates the strategy according to the next state and the reward.

[0012] Preferably, evaluating the credibility of the participants includes evaluating the credibility of the participants through perception data quality detection, motivating the participants to gradually contribute higher-quality perception data, thereby improving the incentive efficiency, and simulating the participant credibility model as a beta distribution.

[0013] Preferably, updating the credibility of the participants includes comparing the perception data with the quality benchmark value of the task, and then updating the credibility of the participants; the quality benchmark value of the task is calculated through a truth value estimation algorithm.

[0014] Preferably, the LSTM-PPO model includes four sub-networks: a feature network, a long short-term memory network, a policy network, and a value network.

[0015] The beneficial effects of the present invention are as follows:

[0016] 1. This solution proposes a participant credibility update mechanism based on quality detection, which uses the truth discovery algorithm to evaluate the true value of sensed data, and uses this as the basis for judging the quality of sensed data, and then evaluates and updates the credibility of participants.

[0017] 2. This solution designs an incentive strategy model combining LSTM and PPO. This model can derive the optimal sensing strategy for each participant in a dynamically changing sensing environment and under incomplete information, so as to maximize the utility.

[0018] 3. This solution proposes a credibility-oriented incentive mechanism CIM-LP based on LSTM and PPO, aiming to provide reasonable and optimal long-term utility rewards according to the participant's resource status and credibility level. We have conducted a series of simulation experiments using real datasets, which prove the convergence and effectiveness of the mechanism algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0020] Figure 1 It is a model diagram of the mobile crowd sensing system of the present invention.

[0021] Figure 2 It is an architecture diagram of the LSTM-PPO model of the present invention.

[0022] Figure 3 It is a specific implementation diagram of the CIM-LP mechanism of the present invention.

[0023] Figure 4 It is a comparison diagram of utility convergence of the present invention.

[0024] Figure 5 It is a diagram showing the influence of the number of participants of the present invention on the average participant utility, average monetary utility, average social utility, and task completion rate.

[0025] Figure 6 It is a diagram showing the influence of the average social intimacy of the present invention on the average participant utility, average monetary utility, average social utility, and task completion rate.

[0026] Figure 7 It is an analysis diagram of the influence and trend of credibility in the CIM-LP mechanism of the present invention.

[0027] Figure 8 This is a comparison chart of the energy consumption of different mechanisms of the present invention over a period of time. Detailed implementation manners

[0028] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] The present invention discloses a credibility-oriented MCS incentive method based on LSTM and PPO, simply referred to as CIM-LP. The embodiments include the following steps: Step 1: The requester publishes a task on the sensing platform, and the sensing platform distributes the task to potential participants in the social network; Step 2: The distributed participants will actively choose whether to accept the sensing task. If accepted, they will upload their willingness to participate, and the sensing platform will select an optimal group of participants; Step 3: The participants collect sensing data and upload it. The sensing decision-making process of the participants is modeled as a non-cooperative game, and the Markov decision process MDP is used to describe their behaviors; Step 4: Without knowing the global information, an incentive model LSTM-PPO that combines the long short-term memory network LSTM and the proximal policy optimization algorithm PPO is used to formulate the most reasonable and effective sensing duration strategy for each participant to maximize the utility reward; Step 5: The sensing platform distributes rewards according to the efforts of the participants, and the sensing platform integrates the sensing data and sends it to the requester; Step 6: After the task is completed, the credibility of the participants is dynamically updated by evaluating the quality of the uploaded data, so as to adjust their utility rewards in the next stage.

[0030] System model

[0031] The MCS system includes two parts: a sensing platform and a user platform. The requester on the user platform submits service requirements to the sensing platform. The sensing platform starts the recruitment process according to its resource status and task requirements, and collects appropriate sensing data from the participants on the user platform. When the participants upload sensing data, the sensing platform will provide appropriate rewards according to the efforts of the participants. The MCS system model involved in this embodiment is as Figure 1 shown, and the specific process is as follows:

[0032] 1. The requester publishes a task on the sensing platform.

[0033] 2. The sensing platform distributes the task to potential participants in the social network.

[0034] 3. The distributed participants will actively choose whether to accept the sensing task. If accepted, they will upload their willingness to participate.

[0035] 4. The sensing platform will select a group of optimal participants and provide some support information to help the participants better complete the tasks.

[0036] 5. The participants collect sensing data and upload it.

[0037] 6. The sensing platform distributes rewards according to the effort level of the participants.

[0038] 7. The sensing platform integrates the sensing data and sends it to the requester.

[0039] For each time period t ∈ [1, 2,..., T], the sensing platform publishes a sensing task j with a fixed reward budget of . The participants share the rewards of the task. To ensure data quality, each task j will have a quality detection threshold to check whether the quality of the participants' sensing data is qualified, so as to update the credibility of the participants. During the entire sensing process T, if the platform has a set of participants participating in sensing, denoted as A = {a1,..., a i ,..., a n}. At each time period t, each participant a i has a credibility value, which reflects the credibility of the participant submitting high-quality data and affects the participant's own utility gain, denoted as Before each participant a i participates in the sensing task j, according to the rewards at each t time period, it will independently select a sensing strategy that maximizes its own utility. In this embodiment, the sensing strategy is simplified to the duration of the participant collecting sensing data, denoted as In addition, this embodiment defines a sensing duration set i for each participant a The set represents the sensing durations of other participants except participant a i .

[0040] In this embodiment, the relationship of the social network will be represented by an undirected graph, and the social network among the participants is represented as a matrix G = [g ik n×n , where n represents the number of participants, and g ik ∈ (0, 1) represents the social closeness between participant a i and participant a k . The calculation formula of social intimacy can be realized through the Jaccard similarity, as shown in formula (1):

[0041] Among them, Set i and Set k respectively represent participant a​i and a k The set of connected participants, |Set i ∩Set k | is the size of the intersection of the two sets, |Set i ∪Set k | is the size of the union of the two sets. The average value of all elements in matrix G is defined as μ, which is used to measure the social intimacy of participants in the entire social network. The larger the value of μ, the closer the connection between participants in the network.

[0042] Before the detailed description, we show the main symbols involved in this embodiment through Table Ⅰ. Table Ⅰ Important Symbols

[0043] Utility model

[0044] In this embodiment, the utility of a participant consists of monetary utility and social utility and.

[0045] Monetary utility

[0046] When participant a i accepts the sensing task j at time period t, he will determine a sensing duration to complete the sensing task. During the entire sensing process, the participant will inevitably consume a certain amount of costs, such as resources like energy, computing, and storage. In this embodiment, the cost function i of participant a is defined as a strongly convex quadratic function, as shown in formula (2):

[0047] where a i ≥0, b i >0, c i >0 are predefined parameters related to the participant and the sensing task, respectively reflecting the cost growth rate brought by the increase in sensing duration, the cost linearly related to the sensing duration, and the basic cost.

[0048] When participant a i spends the duration to complete the sensing task and upload it to the sensing platform, he will obtain the reward from the platform. The monetary reward is estimated based on the credibility of the participant and the proportion of the sensing duration. Then, the reward obtained by participant a i for completing the sensing task j at time period t is as shown in formula (3):

[0049] It should be noted that the perception task needs to be completed collaboratively by multiple participants. Therefore, the reward should be shared by all participants who complete the same task and distributed according to each person's perception duration and data quality. In particular, the data quality is usually unknown in advance. In order to estimate the data quality to the greatest extent before the task starts, this embodiment considers the credibility for quality detection to measure the data quality status, and the credibility is updated in real time according to the perception quality of each participant in each round to provide a reference for the next round.

[0050] Therefore, the monetary utility of a participant consists of the monetary reward paid to the participant and the cost consumed by the participant itself, as shown in formula (4):

[0051] Social utility

[0052] Participant a i 's social utility depends on the degree of connection tightness among participants in the social network, as shown in formula (5):

[0053] where g ik represents the connection tightness between participant a i and a k , and and represent their perception durations.

[0054] Participant utility

[0055] In summary, the utility obtained by participant a i for completing the perception task j in time period t is as shown in formula (6):

[0056] Problem formulation

[0057] This embodiment models the problem of maximizing participant utility as a non - cooperative game. The goal of this game is to maximize its utility by optimizing the perception duration decision of each participant. Specifically, within time period t, the perception duration decision - making process of a participant is defined as a non - cooperative game process, and this game can be represented by a triple as where n is the total number of participants, is the set of perception durations of all participants within this time period, is the set of utilities of all participants.

[0058] Under this game framework, given the set of perception durations of other participants in time period t, i participant a i chooses a strategy to maximize its utility, as shown in Equation (7).

[0059] Assume

[0060] To ensure the effectiveness of the incentive mechanism, the following assumptions are proposed in this embodiment:

[0061] 1. All participants are rational. They will actively participate in the sensing tasks to obtain higher utility rewards and ensure stability during the sensing process without withdrawing randomly.

[0062] 2. The quality of the data uploaded by participants is proportional to their credibility. To obtain higher benefits, participants will compete to improve their own credibility.

[0063] 3. Participants performing the same sensing tasks are in the same social network, and they will exchange sensing information with each other to benefit from each other.

[0064] CIM-LP mechanism

[0065] To solve the problems in the existing incentive mechanisms, this embodiment proposes a CIM-LP mechanism, which mainly includes the following four parts. First, the definition and update algorithm of participants' credibility are introduced to illustrate how to update credibility by dynamically evaluating data quality. Second, the Markov decision process is described in detail so that the CIM-LP mechanism can learn and optimize the optimal strategy. Finally, the LSTM-PPO incentive model and the specific process of the CIM-LP mechanism are elaborated to improve the enthusiasm of participants and data quality.

[0066] Credibility design

[0067] Sensing tasks rely on sensing data provided by a large number of participants. However, due to factors such as device performance, sensing effort, and behavioral habits, the data quality submitted by different participants may vary, which in turn affects the results of sensing tasks and the rewards obtained. Existing studies usually rely on users' participation and task completion when evaluating credibility, while ignoring credibility evaluation based on quality detection. This limitation leads to low efficiency of the incentive mechanism. Therefore, this embodiment proposes a credibility update algorithm based on quality detection to improve the efficiency and fairness of the incentive mechanism.

[0068] Credibility definition

[0069] In this embodiment, the credibility reflects the trustworthiness of participants in submitting high-quality data and affects the benefits of participants. By evaluating the credibility of participants through perceived data quality detection, participants can be motivated to gradually contribute higher-quality perceived data, thereby improving the incentive efficiency. We simulate the participant credibility model as a beta distribution. For participant a i The credibility q i of a i represents the probability that a i provides high-quality data, that is, q i ~Beta(α i ). The credibility value q i is represented by the expected value of the beta distribution, as shown in Equation (8):

[0070] where α i represents the number of times a participant submits high-quality data, and β i represents the number of times a participant submits low-quality data. In particular, since the perception platform does not initially know the credibility of the initial participants, α i and β i are uniformly initialized to 1, and the beta distribution becomes an independent and identically distributed. According to Equation (8), the credibility of the initial participants is 0.5.

[0071] Credibility update

[0072] To accurately update the credibility of participants, we need to compare the perceived data with the true value of the task. However, the true value of the perception task is often unknown, and we need to rely on the perceived data submitted by participants to predict a true value. Therefore, we propose a true value estimation algorithm to predict the true value of the perception task and use it as the quality benchmark value. On this basis, we compare the perceived data submitted by participants with this quality benchmark value, and then update the credibility of participants.

[0073] For perception task j, the perceived data set submitted by the participant is D = {d1,..., d i ,..., d n}. Since the true value is unknown, the perception platform uses the perceived data of the participant to predict the true value of the perception task We define the predicted true value as a data that can minimize the weighted distance from all data, as shown in Equation (9):

[0074] where represents the distance function between any perceived data and the true value, and ω ijRepresents the weight of each data.

[0075] Weight ω ij Represents the importance degree of the sensed data. When the distance between two data is smaller, the weight of the data should be larger, and it is updated in real time through formulas (10) and (11):

[0076] Among them, ω ij ∈(0, 1) and Since the weights of each sensed data are not known at the beginning, the weight values of all sensed data are initialized to 1 / |D|, and the true value is predicted and the weights are updated through continuous iteration. When the change in the weight value of each data between two adjacent iterations is less than a certain threshold, the iteration process stops and the true value of the task is obtained

[0077] According to the true value Participant a i The sensed data d i Submitted by has the quality as shown in formula (12):

[0078] Among them, ι i = 1 indicates that the sensed data d i Submitted by a i Is consistent with the true value of the task ι i →0 indicates that the sensed data d i Submitted by a i Has a large deviation.

[0079] Whether the data quality inspection is qualified is measured by the current data quality ι i Submitted by participant a i High-quality data is closer to the true value of the task than low-quality data, which results in the data quality value being closer to the value 1. We represent the quality inspection result as ζ i , as shown in formula (13):

[0080] Among them, Represents the quality threshold of the sensing task. When the data quality ι i Submitted by the participant is greater than or equal to the threshold, it is considered that the data quality submitted by the participant is qualified; otherwise, the data quality is unqualified.

[0081] Next, update the participant's credibility according to the quality inspection result. Assume that participant a i Has performed Sum times of quality inspections. Among all the quality inspections, the number of times passing the quality inspection is represented as Therefore, participant ai The credibility follows a beta distribution Credibility The expected value is shown in Equation (14):

[0082] It should be noted that the updated credibility will be used as the evaluation criterion for the next round of optimized utility rewards to continuously motivate participants to submit high-quality perceptual data.

[0083] The pseudocode for credibility update is shown in Algorithm 1.

[0084] Markov decision process

[0085] In a non-cooperative game, participants need to adjust their perception duration strategies according to the changing dynamic environment. To effectively simulate this dynamic optimization process, this embodiment adopts a Markov decision process, which can be represented as a quadruple consisting of state, action, state transition probability matrix, and reward, such as MDP = {S, A, P, U}. In deep reinforcement learning, an agent is a decision maker. It first observes the current state of the environment, selects the optimal strategy based on this state to perform an action. Then, it analyzes the impact of the action on the environment, and obtains the next state based on this impact and the reward given by the environment after the action is implemented. Finally, it updates the strategy according to the next state and the reward. In this embodiment, the participant is the agent in the Markov decision process. He adopts an appropriate perception duration according to past experience, then calculates the obtained utility, and finally updates the strategy according to the utility.

[0086] State space:

[0087] At time period t, participant a i selects the optimal perception duration according to the state space s i t The state space is jointly composed of time series data and multiple discrete data, and is expressed as

[0088] Time series data: It represents the set of historical perception durations of the previous L time points of other participants at the current moment. These data are processed by an LSTM network, which can capture the time-dependent relationships therein to facilitate improving the decision-making efficiency.

[0089] Discrete data: It represents the data related to the decision. Among them, represents the credibility value of the participant, represents the remaining power of the device, Indicates the task quality detection threshold, Indicates the task reward.

[0090] It should be noted that when selecting the perception duration, all data in the state space should be comprehensively considered.

[0091] Action space:

[0092] In this embodiment, participant a i The action space at the t-th time period is defined as its perception duration For collecting perception data.

[0093] State transition matrix:

[0094] Perception task reward Participant credibility Remaining power States such as etc. change continuously between different time periods. Therefore, the state space transition follows the probability distribution defined as P. Assuming that the state space and action space at the current time period are s and x respectively, and the state space at the next time period is s′, then the state space conversion from s to s′ is Obviously, p(s′|s,x) ∈ [0,1].

[0095] Reward function:

[0096] The reward function represents the return given by the environment to the agent after the participant takes an action It is actually the utility obtained by the participant after performing the perception task. Therefore, the reward function of the Markov decision process is the utility of the participant, as shown in formula (15):

[0097] LSTM-PPO model architecture

[0098] The LSTM-PPO model consists of a feature network Long short-term memory network λ i , a policy network (with parameters θ i ) and a value network (with parameters ω i ) and consists of four sub-networks, as Figure 2 shown.

[0099] Feature network Aims to extract effective features from the state space of the Markov decision process. Its input includes time series data and discrete data The output is the feature vector extracted through the fully connected layer and

[0100] Long Short-Term Memory Network λ i Specifically designed to process sequential data and capture temporal dependencies in the data. Its input is time series data that has undergone feature extraction Its output is the hidden state h at the final time step t This hidden state is then used to predict the current state feature representation according to the feature network

[0101] Policy Network Also known as the actor, it aims to derive an approximately optimal perception duration to maximize the long-term cumulative reward. Its input is the comprehensive feature representation The output is the probability distribution over the action space The participant selects an action according to the policy probability distribution, and the action with a higher probability is more likely to be selected.

[0102] Value Network Also known as the critic, it aims to accurately estimate the expected cumulative utility reward that can be obtained by following the current policy starting from a specific state. In addition, the value network is also responsible for calculating the advantage function, which evaluates the additional value of performing an action in a given state compared to the average action in the policy, thereby guiding the optimization direction of the policy network The input of the value network is the comprehensive feature representation Its output is a scalar value

[0103] Specific implementation of the CIM-LP mechanism

[0104] The CIM-LP mechanism consists of a perception decision-making process and a policy update process. The perception decision-making process focuses on the participant selecting the optimal perception duration policy through deep reinforcement learning based on the state information provided by the perception platform to maximize its utility reward. The policy update process optimizes the policy network and the value network through the PPO algorithm to ensure that the participant's decision-making continuously improves with the accumulation of experience, thereby enhancing the overall quality of the perception task and the participant's enthusiasm. The specific implementation process of the CIM-LP mechanism is as Figure 3 shown.

[0105] Perception Decision-Making Process

[0106] The perception platform provides the participant with the current state information, including time series data and discrete data The participant processes the time series data through the long short-term memory network to capture temporal dependencies and extract the hidden state h t and then predicts the current state through the feature network Meanwhile, the discrete data undergoes feature extraction through the feature network and combined with the predicted current state in the fusion layer to form a comprehensive feature representation Based on the comprehensive features, the participants generate the probability distribution of the perception duration policy through the policy network (Actor) of deep reinforcement learning, and select the optimal policy according to the probability distribution In addition, the value network (Critic) evaluates the current state and outputs the expected cumulative reward for guiding policy update

[0107] After the perception task is completed, the perception data uploaded by the participants will undergo quality inspection, and the quality results will be used to update the credibility of the participants and affect their utility rewards in subsequent tasks. The relevant information of the task is stored as experience in buffer D i for policy update in the PPO algorithm to ensure that the participants continuously optimize their decisions in subsequent perception tasks and maximize the long-term utility reward

[0108] Policy update process

[0109] To further optimize the participants' policies, the CIM-LP mechanism uses the PPO algorithm for policy update. By sampling the stored experience data, calculating the gradients of the policy network and the value network, and performing parameter updates. Specifically, CIM-LP extracts a small batch of experience data of size B from buffer D i for calculating the gradient in the Actor network and the gradient in the Critic network The parameters of the two networks are θ i and ω i . The parameter update steps are as follows

[0110] First, we calculate the cumulative utility reward for each participant as shown in Equation (16):

[0111] where γ ∈ [0, 1] is the discount factor that defines the range of the value network. If γ = 0 it means it only cares about the current utility and not its long-term utility, while γ = 1 means it cares about the overall cumulative utility from time period t to T

[0112] Second, for the policy network, to measure the relative performance between the new policy and the old policy, the importance sampling ratio is used to adjust the magnitude of the policy network update to ensure the stability of the policy update. As shown in Equation (17):

[0113] Among them, b is the mini-batch index. The numerator and denominator respectively represent the probabilities of the participant taking an action for the state under the new and old policies.

[0114] To avoid instability caused by too large a policy update amplitude, a clipping technique is used to limit the range of variation to be within the interval [1 - ε, 1 + ε]. The clipped objective function J clip (θ i ) is shown in Equation (18):

[0115] where ε is a hyperparameter with a default value of 0.1. denotes the advantage function, which is used to measure how much better the current policy is than the old policy, and is defined as

[0116] Based on the objective function of the policy network, we can calculate the gradient of the policy network as shown in Equation (19):

[0117] where denotes the gradient of the logarithm probability of taking an action under the new policy given the state .

[0118] For the value network, the optimization objective is to minimize the error between the value prediction and the actual cumulative utility reward, and its loss function L(ω i ) takes the form of mean squared error, as shown in Equation (20):

[0119] Based on the loss function, we can calculate the gradient of the value network as shown in Equation (21):

[0120] where denotes the partial derivative of the value network with respect to the parameter ω i .

[0121] Finally, through the calculated gradients of the policy network and the value network, their respective parameters are updated. The policy network is updated using mini-batch stochastic gradient ascent, while the value network is updated using mini-batch stochastic gradient descent, as shown in Equations (22) and (23):

[0122] where \(l_1\) i,1 and \(l_2\) i,2 are the learning rates of the policy network and the value network respectively.

[0123] Through the above steps, the CIM-LP mechanism can ensure that in each round of policy update, it not only optimizes the decision-making ability of the participants but also guarantees the stability of the policy update, gradually improving the long-term utility reward.

[0124] The pseudo-code of the CIM-LP mechanism is shown in Algorithm 2.

[0125] Simulation and Result Analysis

[0126] To evaluate the performance of the CIM-LP mechanism, in this embodiment, the CIM-LP mechanism and three other existing incentive mechanisms are compared on a real dataset, and the simulation settings, comparison mechanisms, evaluation metrics, and performance comparison and analysis are introduced in detail.

[0127] Simulation Settings

[0128] This simulation is implemented based on Python 3.9 and uses the PyTorch 2.1 framework to build and train the mechanism model. For the dataset and parameter settings, the Gowalla dataset models the social relationships among participants in the real world. Gowalla is a location-based social network service provider that allows users to check in at specific locations and share location information, and forms a social network through the natural interactions of users. The dataset includes the geographical location check-in information of users, the activities of users at different times and locations, and a social network containing 196,591 user nodes and 950,327 edges. As shown in formula (1), we use the Jaccard similarity to calculate the social intimacy of each pair of users in the Gowalla dataset to generate a social intimacy matrix

[24] . Based on this matrix, a small social network subset with a specific average social intimacy \(\mu\) can be extracted to construct a specific social intimacy matrix \(G\). Compared with the social network constructed based on the normal distribution in the literature, this method can more realistically reflect the social relationships in the real world. In addition, the specific parameters used in the experimental simulation are given in Table II. Table II Simulation Parameters Parameter Value Total number of participants n [5,25] Average social intimacy μ [0.1,0.9] Total number of time periods T 50 Number of batches M 4 Minimum batch size B 125 Discount factor γ 0.99 Update range ε 0.1 Learning rate 0.0003

[0129] Comparison Mechanisms

[0130] In the comparative experiment, this embodiment will use PPO-DSIM, RLPM, and GSIM-SPD as the comparison mechanisms.

[0131] PPO-DSIM: This mechanism mainly uses PPO in deep reinforcement learning to solve the Nash equilibrium problem in the Stackelberg game, aiming to maximize the utility reward and thus encourage participants to actively participate in sensing tasks. We introduce PPO-DSIM as a comparison mechanism to demonstrate the effectiveness of the mechanism we proposed by combining with the LSTM network.

[0132] RLPM: This mechanism is specifically designed to solve the problem of utility maximization in games, mainly dealing with decision-making problems with discrete action spaces. We introduce RLPM as a comparison mechanism to demonstrate the superiority of the CIM-LP mechanism in dealing with continuous action spaces.

[0133] GSIM-SPD: This mechanism is based on dynamic programming and aims to solve the Nash equilibrium problem in the Stackelberg game. We introduce GSIM-SPD as a benchmark mechanism to show the advantages of CIM-LP based on deep reinforcement learning in dealing with game problems.

[0134] Evaluation Metrics

[0135] This embodiment will introduce evaluation metrics such as average participant utility, average monetary utility, average social utility, and task completion rate to evaluate the performance of the scheme:

[0136] 1. Average Participant Utility: It represents the average utility of all participants over all time periods, as shown in Equation (24):

[0137] where n represents the total number of participants, T represents the total number of time periods, represents the utility of participant a i at time period t.

[0138] 2. Average Monetary Utility: It represents the average monetary utility of all participants over all time periods, as shown in Equation (25):

[0139] where, represents the monetary utility of participant a i at different time periods.

[0140] 3. Average Social Utility: It represents the average social utility of all participants over all time periods, as shown in Equation (26):

[0141] where, represents the social utility of participant a i at different time periods.

[0142] 4. Task Completion Rate: To study the performance of different mechanisms in the perception task completion, the average task completion rate of the entire perception process T is shown in formula (27):

[0143] where the function is defined as This function is used to represent whether the total perception time in period t meets the task requirement threshold ψ. When the perception time is greater than or equal to the task threshold, the task is considered completed; otherwise, the task cannot be completed.

[0144] Comparison and Analysis

[0145] Figure 4 It shows the convergence of the average utility of participants obtained by various comparison mechanisms when n = 10 and μ = 0.9. According to the experimental results, the utility obtained by CIM-LP converges to about 2.1 after 350 rounds of training and stabilizes thereafter. Compared with PPO-DSIM and RLPM, these two benchmark mechanisms gradually converge to about 1.65 and 1.4 respectively when the training reaches about 400 rounds and 600 rounds. This indicates that the CIM-LP mechanism enables participants to obtain higher utility and faster convergence speed. This is because on the basis of deep reinforcement learning, a long short-term memory network is introduced. This improvement enhances the processing ability of the neural network for sequential data, especially data with time dependence, thereby improving the perception decision-making accuracy. Compared with the RLPM mechanism, the experimental results demonstrate the advantage of the CIM-LP mechanism in dealing with continuous action spaces. In addition, the utility obtained by the incentive mechanism based on deep reinforcement learning is significantly higher than that of the heuristic GSIM-SPD. This is because GSIM-SPD only considers the optimal solution in the current situation and ignores the global situation, lacking a clear strategy and goal orientation, thus resulting in low efficiency.

[0146] Figure 5 (a), Figure 5 (b), Figure 5 (c) and Figure 5 (d) respectively represent the effects of the change in the number of participants on the average participant utility, average monetary utility, average social utility, and task completion rate when the social intimacy μ = 0.9. From Figure 5(a) It can be seen that regardless of the change in the number of participants, the CIM-LP mechanism can always enable participants to obtain a relatively high average utility. For CIM-LP, when n = 10, its average utility is 16.43%, 70%, and 193.1% higher than that of PPO-DSIM, RLPM, and GSIM-SPD, respectively. This is due to the fact that participants can dynamically adjust the perception duration according to factors such as credibility, resource status, and task rewards. In addition, the introduction of the LSTM network enables a large amount of key time-series data to be incorporated into the decision-making process, thereby improving the accuracy of the perception duration. We also note that when n varies between 5 and 15, the average utility of participants will decrease. However, when n varies between 15 and 25, the average utility of participants will increase. For example, for CIM-LP, when n = 15, the average utility u reaches 0.49, which is 1.73 lower than when n = 5. When n = 25, the average utility u is 1.40, which is 0.91 higher than when n = 15. From Figure 5 (b) and Figure 5 (c) It can be seen that when the number of participants increases, the average monetary utility of participants will decrease, while the average social utility will increase. For example, in Figure 5 (b), for CIM-LP, when n = 15, the average monetary utility is 0.31, which is 1.85 lower than when n = 5. In Figure 5 (c), for CIM-LP, when n = 15, the average social utility is 0.18, which is 0.12 higher than when n = 5. This is because although the total amount of rewards provided by the perception platform remains unchanged as the number of participants increases, resulting in a decrease in the average monetary utility per participant. However, more participants mean a wider social network, which can bring a higher average social utility to each participant.

[0147] From Figure 5 (d) It can be seen that for the three deep reinforcement learning-based mechanisms (CIM-LP, PPO-DSIM, and RLPM), the task completion rate first increases and then decreases as the number of participants increases. Taking CIM-LP as an example, when n = 15, the task completion rate reaches 94.12%, which is 11.02% higher than when n = 5 and 12.21% higher than when n = 25. This is because when the number of participants increases from 5 to 15, the number of people participating in data collection increases, which helps the smooth completion of the task. However, when n continues to increase, social utility becomes dominant, and participants will tend to obtain more utility rewards by increasing the perception duration. But due to the limited energy of mobile devices, blindly increasing the perception duration will cause the device to run out of power before completing the task, thereby reducing the task completion rate. Therefore, when the perception platform publishes tasks, it should carefully consider the required number of participants.

[0148] Figure 6 (a), Figure 6 (b),Figure 6 (c) and Figure 6 (d) respectively show the effects of the change in average social intimacy on the average participant utility, average monetary utility, average social utility, and task completion rate when the number of participants n = 10. As can be seen from Figure 6 (a), CIM-LP obtains a near-optimal perception strategy and achieves the best utility reward. For example, when μ = 0.9, the average participant utility is 0.877, which is 0.117, 0.357, and 0.464 higher than those of PPO-DSIM, RLPM, and GSIM-SPD respectively. Additionally, as can be seen from Figure 6 (b), CIM-LP obtains more average monetary utility than other mechanisms.

[0149] As can be seen from Figure 6 (c), for all mechanisms, the average social utility increases with the increase in μ. For example, for GSIM-SPD, when μ = 0.9, the average social utility is 0.42, while when μ = 0.5, it is 0.23, and the former is 0.19 higher than the latter. Furthermore, combining Figure 6 (c) and Figure 6 (b) shows that if participants want to obtain higher social utility, they only need a simple perception decision mechanism such as GSIM-SPD. However, considering limited energy resources and different task pricing, a reasonable perception strategy mechanism is needed during the long-term perception process. The CIM-LP mechanism will guide participants to make appropriate perception strategies, achieve a balance between monetary utility and social utility, and achieve higher utility in the long term.

[0150] As can be seen from Figure 6 (c) and Figure 6(d) It can be seen that there are two reasons for the increase in the average social utility with the growth of μ. For RLPM, the average social utility obtained by this mechanism is due to the increase in the perceived duration, which can be verified by the fact that the task completion rate decreases as μ increases. This is because RLPM lacks foresight and blindly guides participants to increase their perceived duration to obtain more social utility at present, resulting in excessive energy consumption and affecting the completion of subsequent tasks. In contrast, for CIM-LP, PPO-DSIM, and GSIM-SPD, the increase in the average social utility is mainly caused by the improvement of social closeness, which can be confirmed by the fact that the task completion rate remains almost unchanged. For CIM-LP, when μ = 0.9, the average monetary utility brought by the perceived decision has reached 92% of the average participant utility. Therefore, the influence of social intimacy μ on the perceived decision is almost negligible, further confirming that the increase in the average social utility is mainly caused by the improvement of social closeness. In addition, when μ = 0.9, the task completion rate of CIM-LP reaches 98.35%, which is 4.56%, 48.96%, and 53.56% higher than that of PPO-DSIM (93.79%), RLPM (49.39%), and GSIM-SPD (44.79%) respectively. This advantage lies in that CIM-LP aims to maximize the overall utility in all perceived tasks by guiding participants to reasonably reserve device energy to obtain greater utility in the future. In addition, by integrating long- and short-term memory networks, CIM-LP optimizes the prediction of the perceived decisions of other participants, enabling it to intelligently adjust strategy execution, effectively manage energy consumption, and ensure the efficient completion of a large number of tasks.

[0151] For the CIM-LP mechanism, participants with different credibility levels obtain different utility rewards at a certain moment, and these differences will ultimately affect their cumulative utility rewards. Figure 7 (a) shows the influence of different credibility levels on the cumulative utility rewards under the same group when n = 5 and μ = 0.9. The experimental results show that the higher the initial credibility of the participants, the faster the growth of the cumulative utility rewards. For example, when the experimental round is 10, the cumulative utility reward with credibility q = 0.8 is 49.5, which is 34.3 higher than the cumulative utility reward with q = 0.2. It can be seen that participants with high credibility are more likely to obtain rewards under the CIM-LP mechanism, thus motivating them to continue to provide high-quality data in future tasks. At the same time, Figure 7(b) Further shows the average credibility change trends of different initial credibility groups in the experimental rounds, where each group consists of 5 participants. The experimental results show that the CIM-LP mechanism can effectively guide different groups to improve their data quality through a dynamic reward mechanism. Especially for the group with high initial credibility, their average credibility remains at a high level in the early stage of the experiment and tends to be stable during the task; while the credibility of the group with low initial credibility gradually increases during the task rounds. For example, for the average credibility q = 0.2, when the experimental round is 10, the average credibility is 0.578, which is 0.378 higher than the average credibility at the first experimental round. This dynamic change shows that the CIM-LP mechanism can improve the overall credibility level and reduce the volatility of data quality.

[0152] Figure 8 Shows the energy consumption performance of four different mechanisms over the entire time period. By comparing the remaining energy of the devices in each time slot, the advantages and disadvantages of each mechanism in energy management are revealed. The experimental parameters are the same as Figure 6 (a), with an average social intimacy of 0.9. For each mechanism, the initial energy of each device is 50. The CIM-LP mechanism shows the best performance, with the slowest energy consumption rate and still having remaining energy available after 50 time slots. This highlights its efficiency and long-term stability in managing energy consumption. By prioritizing high-return tasks and minimizing the energy consumption of low-return tasks, CIM-LP demonstrates its ability to optimize the sensing strategy. In contrast, the GSIM-SPD mechanism shows rapid early energy consumption, resulting in the tool device running out of energy after approximately 30 time slots and being unable to complete subsequent tasks, reflecting its inefficiency. The PPO-DSIM and RLPM mechanisms achieve a balance between short-term and long-term efficiency through reinforcement learning. However, their energy management is still not as good as that of CIM-LP.

[0153] It should be noted that the parts not described in detail in the above embodiments are all prior arts.

[0154] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention should be covered by the protection scope of the present invention.

Claims

1. A credibility-oriented MCS excitation method based on LSTM and PPO, characterized by: The following steps are involved: Step 1: The requester publishes the task on the perception platform, and the perception platform distributes the task to potential participants in the social network; Step 2: The assigned participants will actively choose whether to accept the perception task. If they accept, they will upload their willingness to participate, and the perception platform will select a group of optimal participants. Step 3: Participants collect and upload perception data, model the perception decision-making process of participants as a non-cooperative game, and use Markov decision process MDP to describe their behavior; Step 4: Without knowing the global information, the incentive model LSTM-PPO, which combines the long short-term memory network LSTM with the proximal policy optimization algorithm PPO, is used to formulate the most reasonable and effective perception duration strategy for each participant to maximize the utility reward; Step 5: The perception platform distributes rewards based on the participants’ efforts. The perception platform integrates the perception data and sends it to the requester. Step 6: After the task is completed, the credibility of the participant is dynamically updated by evaluating the quality of the uploaded data, thereby adjusting its utility reward for the next stage.

2. According to claim 1, a credibility-oriented MCS excitation method based on LSTM and PPO is characterized in that: The utility of the participants includes monetary utility and social utility.

3. The MCS excitation method based on LSTM and PPO for credibility according to claim 1, characterized in that: The perception decision-making process focuses on the participants selecting the optimal perception duration strategy through deep reinforcement learning based on the state information provided by the perception platform to maximize their utility rewards.

4. The MCS excitation method based on LSTM and PPO for credibility according to claim 3 is characterized by: It also includes a strategy updating process, which optimizes the strategy network and value network through the PPO algorithm to ensure that the participants' decisions are continuously improved as experience accumulates, thereby improving the overall quality of the perception task and the enthusiasm of the participants.

5. The MCS excitation method based on LSTM and PPO for credibility according to claim 1, characterized in that: The non-cooperative game maximizes the utility of each participant by optimizing the perceived duration decision.

6. The MCS excitation method based on LSTM and PPO for credibility according to claim 1, characterized in that: The Markov decision process is represented as a quadruple consisting of state, action, state transition probability matrix and reward.

7. The MCS excitation method based on LSTM and PPO for credibility according to claim 1, characterized in that: The Markov decision process first observes the current state of the environment and selects the optimal strategy to perform an action based on the state; then, the impact of the action on the environment is analyzed, and the next state is obtained based on the impact and the reward given by the environment after the action is implemented; finally, the strategy is updated based on the next state and the reward.

8. The credibility-oriented MCS excitation method based on LSTM and PPO according to claim 1, characterized in that: The evaluation of the credibility of the participants includes evaluating the credibility of the participants through perceptual data quality detection, motivating the participants to gradually contribute higher quality perceptual data, thereby improving the incentive efficiency, and simulating the participant credibility model as a Beta distribution.

9. The MCS excitation method based on LSTM and PPO for credibility according to claim 1, characterized in that: The updating of the credibility of the participant includes comparing the perception data with the quality benchmark value of the task, thereby updating the credibility of the participant; the quality benchmark value of the task is calculated by a true value estimation algorithm.

10. The credibility-oriented MCS excitation method based on LSTM and PPO according to claim 1, characterized in that: The LSTM-PPO model includes four sub-networks: feature network, long short-term memory network, strategy network and value network.

Citation Information

Cited By

  • Method for evaluating damage of air pollution to human skin health based on PPO-LSTM fusion algorithm

    CN121545756A