An information scene perception multi-task federated learning incentive method and system based on game theory and reinforcement learning
By employing game theory and multi-agent reinforcement learning methods in the federated learning system, and optimizing client pricing and server purchasing strategies, the problems of insufficient incentives and information asymmetry for IoT devices in multi-task federated learning are solved, thereby improving the system's adaptability and overall utility.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI UNIV
- Filing Date
- 2026-03-27
- Publication Date
- 2026-07-03
Smart Images

Figure CN122334545A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of federated learning technology, and in particular to an information scene perception multi-task federated learning incentive method and system based on game theory and reinforcement learning. Background Technology
[0002] With the rapid development of IoT technology, a large number of IoT devices with data collection capabilities have been deeply integrated into various fields. Meanwhile, artificial intelligence model training relies on large-scale data support, and the vast amounts of data generated by IoT devices provide a rich data source for optimizing AI models. However, this data is typically scattered across various IoT devices and may contain a large amount of sensitive information. Today, under increasingly stringent data security regulations, traditional centralized AI model training—which involves transmitting raw data from IoT devices to a central server for training—faces serious privacy risks.
[0003] Federated learning, as a distributed collaborative training paradigm with privacy-preserving advantages, offers a novel and feasible approach to training artificial intelligence models using massive amounts of data from local IoT devices. Despite its demonstrated superior privacy protection capabilities, federated learning still faces numerous challenges.
[0004] First, the initial challenge is insufficient client-side incentives. Specifically, the performance of the global AI model obtained through federated learning training is highly dependent on client participation. However, in federated learning, clients need to use their local resources to train and transmit their local AI models, resulting in significant resource consumption. Without satisfactory economic returns, clients' motivation to participate in federated learning training will decrease significantly, potentially further impacting the performance of the final global AI model due to reduced client-side data volume. It is worth emphasizing that most general-purpose IoT devices are typically equipped with multiple sensors and deployed across various domains, meaning a single device may hold multiple types of data. This fact provides the possibility for a single IoT device to participate in multiple different federated learning training tasks simultaneously. In this context, designing effective incentive mechanisms for multi-task federated learning in the IoT to balance the benefits and costs for clients and servers requires further research and exploration. Some studies have already attempted to address the incentive problem in federated learning. For example, patent application CN117010533A discloses a pricing and incentive method and apparatus for federated learning computing tasks. This method evaluates contributions by quantifying the number of client data samples and training accuracy, and designs incentive mechanisms to optimize resource allocation to improve the efficiency of federated learning training. However, this method can only achieve incentive optimization for single-task federated learning and contribution evaluation based on data volume, but cannot be adapted to scenarios where multiple tasks are participated in concurrently by IoT devices or to distinguish the heterogeneous value of multimodal data.
[0005] Another challenge is the information asymmetry between servers and clients. In existing federated learning incentive mechanisms, servers typically play the role of value assessment and resource allocation, developing reward distribution schemes to attract and integrate the value of dispersed clients. However, due to information asymmetry between clients and servers, servers tend to overestimate the resource costs of clients, making it difficult to recruit clients at near-optimal costs. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and provide an information scene perception multi-task federated learning incentive method and system based on game theory and reinforcement learning, which significantly improves the adaptability and overall effectiveness of the federated learning system in different information environments.
[0007] The objective of this invention can be achieved through the following technical solutions: A multi-task federated learning incentive method for information scene perception based on game theory and reinforcement learning includes the following steps: The utility of clients and servers participating in multi-task federated learning is obtained based on a pre-defined multi-task federated learning system model. The type of information scenario is determined based on the information sharing status between the client and the server. The types of information scenarios include information disclosure scenarios and information non-disclosure scenarios. In the context of information disclosure, a game theory-based incentive mechanism is adopted to model the utility optimization problem between the client and the server as a resource pricing and purchasing game. The Nash equilibrium solution is theoretically derived to determine the client pricing strategy and the server purchasing strategy. In scenarios where information is not publicly available, an incentive mechanism based on multi-agent reinforcement learning is adopted to model the utility optimization problem between clients and servers as a Markov decision process. The multi-agent reinforcement learning method is then used to iteratively explore client pricing strategies and server purchasing strategies to determine the client pricing strategies and server purchasing strategies. Formal federated learning training is performed based on the client pricing strategy and server purchase strategy until the predetermined training rounds are reached.
[0008] Furthermore, the preset multi-task federated learning system model includes multiple servers and multiple clients. Each server manages an independent federated learning task, and each client holds multiple types of data and participates in multiple federated learning tasks simultaneously.
[0009] Furthermore, the utility of clients participating in multi-task federated learning is as follows: In the formula, For the total utility of the j-th client, For the j-th client in the th... The utility gained on each server Total number of servers For the j-th client to the j-th Pricing for each server. For the j-th client to the j-th Local data volume for each server task Execute the first for the j-th client The local training cost of a single server task Execute the first for the j-th client The communication cost of each server task Indicates the first Does server purchase the j-th client as a task participant? ,in, Indicates the first The server purchases the j-th client as a task participant; otherwise, it does not purchase it. The utility of servers participating in multi-task federated learning is as follows: In the formula, For the first The utility of a server The utility gain coefficient. For the j-th client, For client collections.
[0010] Furthermore, in the context of information disclosure, a game theory-based incentive mechanism is adopted to model the utility optimization problem between the client and the server as a resource pricing and purchasing game. The Nash equilibrium solution is theoretically derived, and the specific steps for determining the client's pricing strategy and the server's purchasing strategy include: For a given server task, each client sets its pricing to the minimum price; The server evaluates the marginal utility gain that purchasing each client can bring based on the pricing and data volume of all clients, and selects the client that can bring the greatest marginal utility gain for purchase, which serves as the server purchasing strategy in the information disclosure scenario. Given that the client that brings the greatest marginal utility gain is purchased by the server, we increase its price and ensure that its marginal utility gain to the server is greater than that of other clients until a Nash equilibrium is reached, thus obtaining the client pricing strategy in the information disclosure scenario.
[0011] Furthermore, the formula for calculating the marginal utility gain is as follows: In the formula, The marginal utility gain generated by purchasing the j-th client from the i-th server. The utility gain coefficient. For the j-th client to the j-th Local data volume for each server task For the first The amount of data already purchased for each server For the j-th client to the j-th Pricing executed per server; The minimum price for the client is: In the formula, For the j-th client to the j-th The lowest price implemented per server. Complete the j-th client The total cost of each server task.
[0012] Furthermore, the price increase that brings the maximum marginal utility gain after the client pricing is: In the formula, For clients that can bring the maximum marginal utility gain The price after the price increase For the client that can bring the maximum marginal utility gain, the first The lowest price implemented per server. For pricing increments, For the first Each server purchases the marginal utility gain generated by the client that brings the greatest marginal utility gain at the lowest price that brings the client that brings the greatest marginal utility gain. For the first The server uses the j-th client Lowest price The marginal utility gain resulting from purchasing the j-th client, For client collections.
[0013] Furthermore, in scenarios where information is not publicly available, an incentive mechanism based on multi-agent reinforcement learning is adopted to model the utility optimization problem between the client and the server as a Markov decision process. Multi-agent reinforcement learning is then used to iteratively explore client pricing strategies and server purchasing strategies. The specific steps for determining the client pricing strategy and the server purchasing strategy include: Each client is treated as an independent intelligent agent. In each iteration, the client outputs a pricing strategy for each server through the policy network based on its current state. Calculate the marginal utility gain generated by purchasing each client based on the client's pricing strategy for each server; Each server selects the client with the highest marginal utility gain to make a purchase, generating the purchase strategy for that round of iteration; Based on the server's purchase strategy, calculate the immediate reward for each client in this iteration, and update and optimize its policy network through a multi-agent reinforcement learning algorithm; Once all client pricing strategies and all server purchasing strategies are made public, the client transitions from its current state to the state of the next iteration, enabling iterative exploration of client pricing strategies and server purchasing strategies until a termination state is reached.
[0014] Furthermore, the state of the client is: In the formula, For the j-th client in the th... The state in the round of iteration, For the length of the history window, For the first In the first round of iteration, the first client... Pricing for each server. For the first In the first iteration The client for the first Pricing for each server. For the first In the first round of iteration, the first client... Pricing for each server. For the first In the first iteration The client for the first Pricing for each server. For the first In the first iteration Does the server purchase the first client as a task participant? For the first In the first iteration Is the server purchased? Each client acts as a task participant; The instant reward is: In the formula, For the j-th client in the th... Instant rewards in rounds of iteration For the first One server, For a set of servers, Let the utility obtained by the j-th client on the i-th server be _____. For the first In the round of iteration, the j-th client interacts with the j-th client. Pricing for each server. For the first In the first iteration Does each server purchase the j-th client as a task participant?
[0015] Furthermore, the termination state is when the pricing of all clients has been reduced to their respective minimum prices, and the marginal utility gain from purchasing any client cannot cover the minimum cost required to purchase the data for that client.
[0016] According to another aspect of the present invention, an information scene perception multi-task federated learning incentive system based on game theory and reinforcement learning is provided, comprising: The utility acquisition module is used to acquire the utility of clients and servers participating in multi-task federated learning based on a preset multi-task federated learning system model. The scenario type determination module is used to determine the type of information scenario based on the information sharing status between the client and the server. The types of information scenarios include information disclosure scenarios and information non-disclosure scenarios. The information disclosure scenario processing module is used to model the utility optimization problem between the client and the server as a resource pricing and purchasing game in the context of information disclosure by adopting a game theory-based incentive mechanism, and to theoretically derive the Nash equilibrium solution to determine the client pricing strategy and the server purchasing strategy. The information non-disclosure scenario processing module is used to model the utility optimization problem between the client and the server as a Markov decision process in the case of information non-disclosure by adopting an incentive mechanism based on multi-agent reinforcement learning. It then uses multi-agent reinforcement learning to iteratively explore the client pricing strategy and the server purchasing strategy to determine the client pricing strategy and the server purchasing strategy. The federated learning module is used to perform formal federated learning training based on the client pricing strategy and server purchase strategy until the predetermined training rounds are reached.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention determines the type of information scenario by acquiring the utility of clients and servers participating in multi-task federated learning. In scenarios with publicly available information, a game theory-based incentive mechanism is used to model the utility optimization problem of clients and servers as a resource pricing and purchasing game. In scenarios with non-public information, a multi-agent reinforcement learning-based incentive mechanism is used to model the utility optimization problem of clients and servers as a Markov decision process. Formal federated learning training is performed based on client pricing strategies and server purchasing strategies. This solves the problem that single incentive strategies are ineffective in different information scenarios and that the overall system utility and model performance are limited, significantly improving the adaptability and overall utility of the federated learning system in different information environments.
[0018] 2. In the context of information disclosure, this invention employs a game theory-based incentive mechanism to model the utility optimization problem between clients and servers as a resource pricing and purchasing game. It also theoretically derives the Nash equilibrium solution, from the client offering the lowest price, the server selecting the client with the highest marginal utility gain, to the selected client raising its price until Nash equilibrium is reached. This solves the problems that servers may still face in information disclosure scenarios, such as high procurement costs, unstable incentive strategies, and difficulty in achieving the optimal system state. It significantly improves the predictability of system utility, strategy stability, and overall resource allocation efficiency.
[0019] 3. In scenarios where information is not publicly available, this invention employs an incentive mechanism based on multi-agent reinforcement learning. It models the utility optimization problem between clients and servers as a Markov decision process and uses multi-agent reinforcement learning to iteratively explore client pricing strategies and server purchasing strategies. This solves the problem that effective incentives cannot be achieved through centralized theoretical derivation and system utility is difficult to optimize in scenarios where information is not publicly available. It significantly improves the adaptability, robustness, and overall system utility of the incentive mechanism in complex and opaque environments. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating a multi-task federated learning incentive method for information scene perception based on game theory and reinforcement learning proposed in this invention. Figure 2 This is a flowchart illustrating the process of classifying incentives based on information scenario type. Detailed Implementation
[0021] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0022] Example 1 This embodiment provides an information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning, such as... Figure 1 As shown, it includes the following steps: S1. Obtain the utility of clients and servers participating in multi-task federated learning based on a preset multi-task federated learning system model.
[0023] The pre-defined multi-task federated learning system model comprises multiple servers and multiple clients. Each server manages an independent federated learning task, and each client holds various types of data and can participate in multiple federated learning tasks simultaneously. Based on this, client utility models and server utility models are constructed separately, and the relationship between the amount of client data and the final model performance is coupled. This ensures that every expense paid by the server is strongly correlated with its expected or actual performance improvement, thereby optimizing server utility to enhance model performance.
[0024] Client Total cost Including training costs and communication costs Two parts, represented as: The training cost is determined by the client's training latency, local data volume, and computing power. CPU training latency for performing a local update Represented as: Client CPU computation cost of performing a local update Represented as: in, Indicates client The effective capacitance coefficient, Indicates client The number of CPU cycles required to process a single data sample. Indicates client The size of the local dataset. Indicates client The computing resources allocated during the training of the local knowledge tracing model, i.e., the CPU cycle frequency used to execute the task.
[0025] After the local knowledge tracing model is trained, the client... Uploading model parameters to the server requires communication latency, which depends on the number of parameters. It can be represented as: in, This indicates the size of the uploaded model parameters (in bits). Indicates client The achievable uplink data rate, expressed according to Shannon's formula, is: in, Indicates client Communication power, Indicates the total system bandwidth. Indicates assignment to the client The percentage of resource blocks, , Indicates client Channel gain between the data center and the data center This represents the power spectral density of additive white Gaussian noise during wireless transmission. (Client) Uploading the parameters of a trained local intelligent model to the server requires communication costs that depend on upload time and power consumption. It can be represented as: in, Indicates the communication power of the participants. This indicates the communication latency of the participants.
[0026] Therefore, the client Total cost required for local training The sum of computation cost and communication cost can be expressed as: In a multitasking environment, the client The total utility is the sum of the utilities of all participating tasks, where the utility for each server task consists of the revenue from the sale of data and the local cost, and can be expressed as: In the formula, For the total utility of the j-th client, For the j-th client in the th... The utility gained on each server Total number of servers For the j-th client to the j-th Pricing for each server. For the j-th client to the j-th Local data volume for each server task Execute the first for the j-th client The local training cost of a single server task Execute the first for the j-th client The communication cost of each server task Indicates the first Does server purchase the j-th client as a task participant? ,in, Indicates the first A server purchases the j-th client as a task participant; otherwise, it does not purchase any.
[0027] server The utility The performance gains from the model can be expressed as follows, relative to the cost of purchasing clients: in, The performance gains of the model in relation to the amount of data can be expressed in detail as follows: in, This is the utility gain coefficient.
[0028] Expenses related to purchasing client software can be detailed as follows: Therefore, the server The utility This can be expressed in detail as follows: In the formula, For the first The utility of a server The utility gain coefficient. For the j-th client, For client collections.
[0029] S2. Determine the type of information scenario based on the information sharing status between the client and the server. The types of information scenarios include information disclosure scenarios and information non-disclosure scenarios.
[0030] like Figure 2 As shown, the type of information scenario is determined by the information sharing between the client and server. In an information-public scenario, the information or services provided by the server to the client can be freely and indiscriminately obtained without specific authentication or authorization. The interaction between the client and server is open and transparent, and the information is public to all requesters. In a non-public scenario, the server needs to perform strict authentication, permission verification, or session management before providing information or services to the client, and the acquisition of information is conditional and exclusive.
[0031] S3. In the context of information disclosure, an incentive mechanism based on game theory is adopted to model the utility optimization problem between the client and the server as a resource pricing and purchasing game. The Nash equilibrium solution is theoretically derived to determine the client pricing strategy and the server purchasing strategy.
[0032] In the context of information disclosure, a game-theoretic incentive mechanism is employed to model the utility optimization problem between clients and servers as a resource pricing and purchasing game. In the multi-task federated learning system within this information disclosure scenario, participating clients are considered price setters, and servers are considered buyers. The utility optimization problem for both clients and servers can be modeled as a resource pricing and purchasing game. This pricing and purchasing game aims to explore price-purchasing strategies for clients and servers that maximize their respective utilities. .make Indicates that, except for the client Pricing strategies for all other clients, based on the pricing strategies of other clients. and server Purchasing strategy Client The goal is to formulate an optimal pricing strategy to maximize its own utility. Therefore, the focus is on the pricing and buying game. It can be modeled as: in, For client collection, For a set of servers, This represents the set of pricing strategies for all clients. This represents the set of purchasing strategies for all servers. Indicates client The effect, Indicates server The utility of the client set and the server set together constitutes the participant set. For the server... The system first evaluates the marginal utility gain offered by all clients, then purchases the client that provides the maximum marginal utility gain to increase its own utility. For clients, to be purchased by the server and improve their utility, they compete by adjusting prices, forming a non-cooperative game. In Nash equilibrium, none of the participants can increase their payoff by unilaterally changing their strategies. Therefore, they all adhere to their current strategies, forming an equilibrium. When the multi-task federated learning system reaches Nash equilibrium, no client or server can increase its utility by individually adjusting its strategy. In this state, each client's strategy is the optimal response to the current strategies of other clients, thus forming a mutually constrained and relatively stable equilibrium.
[0033] The theoretical derivation of the Nash equilibrium solution and the specific steps to determine the client pricing strategy and server purchasing strategy include: In the context of information disclosure, a game theory-based incentive mechanism is adopted to model the utility optimization problem between clients and servers as a resource pricing and purchasing game. The Nash equilibrium solution is theoretically derived, and the specific steps for determining the client's pricing strategy and the server's purchasing strategy include: For a given server task, each client sets its pricing to the minimum price; The server evaluates the marginal utility gain that purchasing each client can bring based on the pricing and data volume of all clients, and selects the client that can bring the greatest marginal utility gain for purchase, which serves as the server purchasing strategy in the information disclosure scenario. Given that the client that brings the greatest marginal utility gain is purchased by the server, we increase its price and ensure that its marginal utility gain to the server is greater than that of other clients until a Nash equilibrium is reached, thus obtaining the client pricing strategy in the information disclosure scenario.
[0034] For servers The formula for calculating its marginal utility gain is defined as follows: In the formula, For the first The marginal utility gain generated by a server purchasing the j-th client The utility gain coefficient. For the j-th client to the j-th Local data volume for each server task For the first The amount of data already purchased for each server For the j-th client to the j-th Pricing for each server.
[0035] To maximize its effectiveness, the server Will choose to purchase ,in satisfy: Then, calculate about First derivative: This indicates that server utility increases as client prices decrease. Therefore, in order to be purchased by the server, clients will gradually lower their prices to the minimum through competition with other clients. Considering that clients are rational, they will not reduce their own utility to a negative value, leading to the conclusion that... For the server The lowest price is: In the formula, For the j-th client to the j-th The lowest price implemented per server. Complete the j-th client The total cost of each server task.
[0036] When the client Its price Set as the lowest price At this time, it allows the server to obtain maximum utility from the current client. In this case, each client For server The resulting marginal utility gains are respectively Assume that the maximum of these marginal utility gains is determined by the client. This occurs because the server is rational. Purchase client In order to maximize its own utility.
[0037] For the client In the server In the case of purchase, the client The goal is to maximize its own utility. Based on the client utility function formula, we obtain... Regarding price First derivative: Therefore, the client The utility Relative to price It is monotonically increasing, meaning that as the price increases, the client's utility increases. Therefore, in order to increase its own utility, the client... Will be on the server Given the premise of purchase, appropriately increasing the price will result in the client-side pricing that maximizes marginal utility gain. The increased price is: In the formula, For clients that can bring the maximum marginal utility gain The price after the price increase For the client that can bring the maximum marginal utility gain, the first The lowest price implemented per server. For pricing increments, For the first Each server purchases the marginal utility gain generated by the client that brings the greatest marginal utility gain at the lowest price that brings the client that brings the greatest marginal utility gain. For the first The server uses the j-th client Lowest price The marginal utility gain resulting from purchasing the j-th client, For client collections.
[0038] Client By raising its price To increase its utility, while ensuring its effectiveness on the server. utility gain It is the largest compared to other clients. In this case, on the server... Purchase client At the same time, they increase their utility. At this point, the client and server reach a Nash equilibrium because neither can increase their own utility by unilaterally changing their strategies.
[0039] S4. In scenarios where information is not publicly available, an incentive mechanism based on multi-agent reinforcement learning is adopted to model the utility optimization problem between the client and the server as a Markov decision process. The multi-agent reinforcement learning method is then used to iteratively explore the client pricing strategy and the server purchasing strategy to determine the client pricing strategy and the server purchasing strategy.
[0040] In scenarios where information is not publicly available, a multi-agent reinforcement learning-based incentive mechanism is employed to model the client-server utility optimization problem as a Markov decision process. Following the design principles of Markov decision processes, each client is treated as an agent in each iteration. A multi-agent Markov decision process with state, action, and reward can be formulated as follows: Client First, based on the previous The current pricing strategy is determined by analyzing the client's historical purchasing and pricing strategies. The client's status is: In the formula, For the j-th client in the th... The state in the round of iteration, For the length of the history window, For the first In the first round of iteration, the first client... Pricing for each server. For the first In the first iteration The client for the first Pricing for each server. For the first In the first round of iteration, the first client... Pricing for each server. For the first In the first iteration The client for the first Pricing for each server. For the first In the first iteration Does the server purchase the first client as a task participant? For the first In the first iteration Is the server purchased? Each client acts as a participant in the task.
[0041] Each client independently determines its pricing strategy based on its own state. Here, the action of client c is defined as... ,in .
[0042] After all clients take action, each server Calculate the utility gain generated by the client and purchase the client with the maximum utility gain. ,in After that, the client... It can know the pricing strategies of other clients and the purchasing strategies of all servers, and then calculate the instant reward. Among them, the client Instant rewards Represented as: In the formula, For the j-th client in the th... Instant rewards in rounds of iteration For the first One server, For a set of servers, For the j-th client in the th... The utility gained on each server For the first In the round of iteration, the j-th client interacts with the j-th client. Pricing for each server. For the first In the first iteration Does each server purchase the j-th client as a task participant?
[0043] After designing the states, actions, and rewards, a multi-agent reinforcement learning method is used to iteratively explore the client pricing strategy and the server purchase strategy. The specific steps to determine the client pricing strategy and the server purchase strategy include: Each client is treated as an independent intelligent agent. In each iteration, the client outputs a pricing strategy for each server through the policy network based on its current state. Calculate the marginal utility gain generated by purchasing each client based on the client's pricing strategy for each server; Each server selects the client with the highest marginal utility gain to make a purchase, generating the purchase strategy for that round of iteration; Based on the server's purchase strategy, calculate the immediate reward for each client in this iteration, and update and optimize its policy network through a multi-agent reinforcement learning algorithm; Once all client pricing strategies and all server purchasing strategies are public, the client transitions from its current state to the next iteration, iteratively exploring client pricing and server purchasing strategies until a termination state is reached. The termination state occurs when all client pricing has fallen to their respective minimum prices, and the marginal utility gain from purchasing any client cannot cover the minimum cost required to purchase that client's data.
[0044] S5. Perform formal federated learning training based on the client pricing strategy and server purchase strategy until the predetermined training rounds are reached.
[0045] The client pricing strategy and server purchase strategy determined in the previous steps serve as the basis for this round of training. The server distributes the global model to the selected clients according to the purchase strategy. The selected clients train using local data and upload model updates. The server aggregates these updates to generate a new global model. This training process is repeated, and at the end of each round, it is determined whether the preset total number of training rounds has been reached. If not, the next round of training continues based on the latest global model and a potentially updated strategy. This iterative cycle continues until the number of training rounds meets the predetermined conditions. The process then enters the final payment stage, where the server pays the corresponding rewards to the clients who participated in the training according to the finally determined client pricing strategy, completing the entire incentive and training process.
[0046] To more clearly illustrate the purpose, technical solution, and advantages of this invention, this embodiment further demonstrates the technical effectiveness of the information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning through comparative simulation experiments.
[0047] In the simulation experiment, the default settings for the system parameters are as follows: The experiment considered a multi-task federated learning system with two servers, each independently executing the AI model training task. Regarding the number of participating clients, two different client scales were considered: the first scale included 30 clients, and the second scale expanded to 50 clients. For client computing power, the CPU frequency of all client devices was uniformly set to [1.0, 2.0] GHz, and the number of CPU cycles processing a single sample was [not specified]. for Period / sample, effective switched capacitor Set as Regarding client communication parameters, the total bandwidth per server... 2MHz. Client communication power. The channel gain is within the range of [0.2, 1.0]W. Following an independent Rayleigh fading model, background noise for To represent a realistic multi-task federated learning scenario, we set up federated learning training tasks for two servers. One server is responsible for the CIFAR10 training task, representing a relatively simple image classification task. The other server is responsible for the CIFAR100 training task, representing a more complex image classification task. During the multi-task federated learning training process, each server only interacts with its recruited clients to update the artificial intelligence model.
[0048] The accuracy of the task, server utility, client utility, and the number of clients recruited by the server are used as performance comparison standards. Four benchmark methods—static incentive mechanism, random incentive mechanism, auction incentive mechanism, and data volume incentive mechanism—are selected to verify the effectiveness of the proposed incentive mechanism. In the static incentive mechanism, the client price is determined between its minimum and maximum values at the beginning of training and remains constant in subsequent operations. In the random incentive mechanism, the client price varies randomly between the minimum and maximum values acceptable to the server. In the auction incentive mechanism, the server determines the winning client and purchases its service based on the client's own cost-effectiveness. In the data volume incentive mechanism, the server purchases the client with the largest amount of data.
[0049] Compared to the baseline scheme, the incentive mechanism of this invention significantly optimizes server utility. Taking server 1 as an example, when the number of clients is 30 (i.e., N=30), the game theory-based and reinforcement learning-based schemes achieve utilities of 259.13 and 254.29, respectively. In contrast, the baseline methods based on auction, quantity, randomness, and statics achieve server utilities of 224.10, 213.78, 201.98, and 230.06, respectively. When the number of clients N is set to 50, the server utilities obtained based on game theory, multi-agent reinforcement learning, auction, data volume, randomness, and statics are 266.49, 242.96, 230.30, 205.30, 194.46, and 213.35, respectively. Compared to the baseline scheme, the game theory-based method proposed in this invention can improve server utility by 27.03%, while the reinforcement learning-based method can improve it by 19.96%.
[0050] When server 2 recruited 50 clients, the client utility values obtained based on game theory, reinforcement learning, auction mechanism, quantity allocation, random allocation, and static allocation methods were 3.91, 19.68, 27.80, 28.56, 27.21, and 17.00, respectively. It is evident that the game theory method proposed in this invention yielded the lowest client utility value. This result demonstrates that this invention constructs dynamic competition among clients. To improve their utility and be adopted by the server, clients will gradually lower their bids as expected until they approach cost, thus alleviating the information asymmetry problem between clients and the server.
[0051] The incentive mechanisms based on game theory and reinforcement learning rank first and second in terms of total utility, indicating that the incentive mechanism of this invention can significantly improve the overall utility of the multi-task federated learning system. Taking the federated learning task of server 1 as an example, when the number of clients is 50, the total utility generated by the incentive mechanisms based on game theory and reinforcement learning is 270.41 and 262.64, respectively, while the total utility generated by the auction, quantitative, random, and static methods is 258.10, 233.85, 221.67, and 230.35, respectively. It can be observed that compared with the auction, quantitative, random, and static benchmark methods, the incentive mechanism based on game theory improves the total utility by 4.55%, 13.52%, 18.02%, and 14.81%, respectively.
[0052] For the multi-task federated learning system of interest, the incentive mechanism proposed in this invention recruits a larger number of clients. When the client size of server 1 is set to 30, the number of clients recruited using game theory, reinforcement learning, auction, quantitative, random, and static incentive mechanisms are 12, 10, 9, 3, 3, and 2, respectively.
[0053] For a federated learning task with server 1 collaboration (with 30 clients that can be recruited), the highest training accuracies achieved by recruiting clients using incentive schemes based on game theory, reinforcement learning, auction mechanisms, data volume, randomized strategies, and static incentive mechanisms were 74.24%, 72.00%, 68.00%, 51.34%, 49.69%, and 56.55%, respectively. In a federated learning task with server 1 collaboration (with 50 clients that can be recruited), these incentive mechanisms achieved the highest training accuracies of 76.45%, 75.09%, 73.09%, 58.87%, 50.90%, and 49.67%, respectively. Compared to other incentive mechanism methods, the game theory method and reinforcement learning method proposed in this invention can improve the training accuracy by 26.78% and 25.42%, respectively.
[0054] For a federated learning task completed collaboratively by server 2 (with 30 clients that can be recruited), the highest training accuracies achieved by recruiting clients using game theory, reinforcement learning, auction, data volume, randomness, and static incentive mechanisms were 67.51%, 64.32%, 39.47%, and 43.90%, respectively. For a federated learning task on server 2 (with 50 clients that can be recruited), the highest training accuracies achieved by these incentive mechanisms were 69.62%, 68.21%, 59.68%, 39.66%, 36.29%, and 44.23%, respectively. Compared to other incentive mechanism methods, the game theory method and reinforcement learning method proposed in this invention can improve training accuracy by 33.33% and 31.92%, respectively.
[0055] Example 2 This embodiment provides an information scene perception multi-task federated learning incentive system based on game theory and reinforcement learning, including: The utility acquisition module is used to acquire the utility of clients and servers participating in multi-task federated learning based on a preset multi-task federated learning system model. The scenario type determination module is used to determine the type of information scenario based on the information sharing status between the client and the server. The types of information scenarios include information disclosure scenarios and information non-disclosure scenarios. The information disclosure scenario processing module is used to model the utility optimization problem between the client and the server as a resource pricing and purchasing game in the context of information disclosure by adopting a game theory-based incentive mechanism, and to theoretically derive the Nash equilibrium solution to determine the client pricing strategy and the server purchasing strategy. The information non-disclosure scenario processing module is used to model the utility optimization problem between the client and the server as a Markov decision process in the case of information non-disclosure by adopting an incentive mechanism based on multi-agent reinforcement learning. It then uses multi-agent reinforcement learning to iteratively explore the client pricing strategy and the server purchasing strategy to determine the client pricing strategy and the server purchasing strategy. The Federated Learning module is used to perform formal federated learning training based on the client pricing strategy and the server purchase strategy until the predetermined training rounds are reached, and to pay the client according to the client pricing strategy.
[0056] The rest is the same as in Example 1.
[0057] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A multi-task federated learning incentive method for information scene perception based on game theory and reinforcement learning, characterized in that, Includes the following steps: The utility of clients and servers participating in multi-task federated learning is obtained based on a pre-defined multi-task federated learning system model. The type of information scenario is determined based on the information sharing status between the client and the server. The types of information scenarios include information disclosure scenarios and information non-disclosure scenarios. In the context of information disclosure, a game theory-based incentive mechanism is adopted to model the utility optimization problem between the client and the server as a resource pricing and purchasing game. The Nash equilibrium solution is theoretically derived to determine the client pricing strategy and the server purchasing strategy. In scenarios where information is not publicly available, an incentive mechanism based on multi-agent reinforcement learning is adopted to model the utility optimization problem between clients and servers as a Markov decision process. The multi-agent reinforcement learning method is then used to iteratively explore client pricing strategies and server purchasing strategies to determine the client pricing strategies and server purchasing strategies. Formal federated learning training is performed based on the client pricing strategy and server purchase strategy until the predetermined training rounds are reached.
2. The information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning according to claim 1, characterized in that, The preset multi-task federated learning system model includes multiple servers and multiple clients. Each server manages an independent federated learning task, and each client holds multiple types of data and participates in multiple federated learning tasks simultaneously.
3. The information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning according to claim 1, characterized in that, The utility of clients participating in multi-task federated learning is as follows: In the formula, For the total utility of the j-th client, For the j-th client in the th... The utility gained on each server Total number of servers For the j-th client to the j-th Pricing for each server. For the j-th client to the j-th Local data volume for each server task Execute the first for the j-th client The local training cost of a single server task Execute the first for the j-th client The communication cost of each server task Indicates the first Does server purchase the j-th client as a task participant? ,in, Indicates the first The server purchases the j-th client as a task participant; otherwise, it does not purchase it. The utility of servers participating in multi-task federated learning is as follows: In the formula, For the first The utility of a server The utility gain coefficient. For the j-th client, For client collections.
4. The information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning according to claim 1, characterized in that, In the context of information disclosure, a game theory-based incentive mechanism is adopted to model the utility optimization problem between clients and servers as a resource pricing and purchasing game. The Nash equilibrium solution is theoretically derived, and the specific steps for determining the client's pricing strategy and the server's purchasing strategy include: For a given server task, each client sets its pricing to the minimum price; The server evaluates the marginal utility gain that purchasing each client can bring based on the pricing and data volume of all clients, and selects the client that can bring the greatest marginal utility gain for purchase, which serves as the server purchasing strategy in the information disclosure scenario. Given that the client that brings the greatest marginal utility gain is purchased by the server, we increase its price and ensure that its marginal utility gain to the server is greater than that of other clients until a Nash equilibrium is reached, thus obtaining the client pricing strategy in the information disclosure scenario.
5. The information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning according to claim 4, characterized in that, The formula for calculating the marginal utility gain is as follows: In the formula, The marginal utility gain generated by purchasing the j-th client from the i-th server. The utility gain coefficient. For the j-th client to the j-th Local data volume for each server task For the first The amount of data already purchased for each server. For the j-th client to the j-th Pricing executed per server; The minimum price for the client is: In the formula, For the j-th client to the j-th The lowest price implemented per server. Complete the j-th client The total cost of each server task.
6. The information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning according to claim 4, characterized in that, The client-side pricing increase that yields the maximum marginal utility gain is: In the formula, For clients that can bring the maximum marginal utility gain The price after the price increase For the client that can bring the maximum marginal utility gain, the first The lowest price implemented per server. For pricing increments, For the first Each server purchases the marginal utility gain generated by the client that brings the greatest marginal utility gain at the lowest price that brings the client that brings the greatest marginal utility gain. For the first The server uses the j-th client Lowest price The marginal utility gain resulting from purchasing the j-th client, For client collections.
7. The information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning according to claim 1, characterized in that, In scenarios where information is not publicly available, an incentive mechanism based on multi-agent reinforcement learning is adopted to model the utility optimization problem between clients and servers as a Markov decision process. Multi-agent reinforcement learning is then used to iteratively explore client pricing strategies and server purchasing strategies. The specific steps for determining the client pricing strategy and server purchasing strategy include: Each client is treated as an independent intelligent agent. In each iteration, the client outputs a pricing strategy for each server through the policy network based on its current state. Calculate the marginal utility gain generated by purchasing each client based on the client's pricing strategy for each server; Each server selects the client with the highest marginal utility gain to make a purchase, generating the purchase strategy for that round of iteration; Based on the server's purchase strategy, calculate the immediate reward for each client in this iteration, and update and optimize its policy network through a multi-agent reinforcement learning algorithm; Once all client pricing strategies and all server purchasing strategies are made public, the client transitions from its current state to the state of the next iteration, enabling iterative exploration of client pricing strategies and server purchasing strategies until a termination state is reached.
8. The information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning according to claim 7, characterized in that, The client's status is: In the formula, For the j-th client in the th... The state in the round of iteration, For the length of the history window, For the first In the first round of iteration, the first client... Pricing for each server. For the first In the first iteration The client for the first Pricing for each server. For the first In the first round of iteration, the first client... Pricing for each server. For the first In the first iteration The client for the first Pricing for each server. For the first In the first iteration Does the server purchase the first client as a task participant? For the first In the first iteration Is the server purchased? Each client acts as a task participant; The instant reward is: In the formula, For the j-th client in the th... Instant rewards in rounds of iteration For the first One server, For a set of servers, For the j-th client in the th... The utility gained on each server For the first In the round of iteration, the j-th client interacts with the j-th client. Pricing for each server. For the first In the first iteration Does each server purchase the j-th client as a task participant? 9. The information scene perception multi-task federated learning incentive method based on game theory and reinforcement learning according to claim 7, characterized in that, The termination state is when the pricing of all clients has been reduced to their respective minimum prices, and the marginal utility gain from purchasing any client cannot cover the minimum cost required to purchase the data for that client.
10. An information scene perception multi-task federated learning incentive system based on game theory and reinforcement learning, characterized in that, include: The utility acquisition module is used to acquire the utility of clients and servers participating in multi-task federated learning based on a preset multi-task federated learning system model. The scenario type determination module is used to determine the type of information scenario based on the information sharing status between the client and the server. The types of information scenarios include information disclosure scenarios and information non-disclosure scenarios. The information disclosure scenario processing module is used to model the utility optimization problem between the client and the server as a resource pricing and purchasing game in the context of information disclosure by adopting a game theory-based incentive mechanism, and to theoretically derive the Nash equilibrium solution to determine the client pricing strategy and the server purchasing strategy. The information non-disclosure scenario processing module is used to model the utility optimization problem between the client and the server as a Markov decision process in the case of information non-disclosure by adopting an incentive mechanism based on multi-agent reinforcement learning. It then uses multi-agent reinforcement learning to iteratively explore the client pricing strategy and the server purchasing strategy to determine the client pricing strategy and the server purchasing strategy. The federated learning module is used to perform formal federated learning training based on the client pricing strategy and server purchase strategy until the predetermined training rounds are reached.
Citation Information
Patent Citations
Pricing and incentive method and device for federal learning calculation task
CN117010533A