A federated learning incentive method, system, medium, device, terminal and application

By adopting a federated learning incentive method based on a continuous zero determinant strategy, the problem of mobile devices being unwilling to contribute high-precision data is solved, thereby improving the efficiency and optimizing social welfare of the federated learning system and ensuring fair cooperation and stable utility between devices and servers.

CN117010525BActive Publication Date: 2026-05-19ZHEJIANG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG NORMAL UNIV
Filing Date
2023-05-17
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing federated learning systems struggle to balance data privacy and system efficiency. Mobile devices are unwilling to contribute high-precision data, leading to low federated learning efficiency and reduced social welfare. Existing incentive mechanisms are unable to achieve a balance between efficiency and fairness, and the model settings are unreasonable, making it impossible to flexibly adjust reward levels.

Method used

We employ a federated learning incentive method based on a continuous zero-determinant policy to construct a system and quantify the policies of devices and servers. By controlling the expected returns through a zero-determinant policy in a continuous policy scenario, we design an incentive mechanism to maximize social welfare and ensure full cooperation between devices and servers.

Benefits of technology

It significantly improves the efficiency of federated learning, attracts mobile devices to contribute high-precision data, optimizes social welfare, achieves fairness and utility stability between devices and servers, and dynamically adjusts reward levels to promote high-quality data contributions from devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117010525B_ABST
    Figure CN117010525B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of federated learning system optimization, and discloses a federated learning incentive method, system, medium, equipment, terminal and application, constructs a federated learning system based on a continuous zero determinant strategy, quantifies the strategy of the equipment and the server, calculates the utility function and the expected utility, controls the expected income of the federated learning participating equipment and the server to meet a linear relationship by using the zero determinant strategy in the continuous strategy scene, takes the maximization of the social welfare of the federated learning as an optimal target, ensures that the utility can be maximized only in the case that the equipment and the server fully cooperate, solves the optimization problem, and completes the design of the federated learning incentive mechanism algorithm according to the found optimal strategy. The application can attract mobile equipment to participate and contribute high-precision data, significantly improves the calculation complexity and the time for completing the federated learning, and is beneficial to adaptively incentivizing the contribution of the federated learning participants, completing the federated learning model training task and designing an optimal response strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of federated learning system optimization technology, and particularly relates to a federated learning incentive method, system, medium, device, terminal and application. Background Technology

[0002] The emergence of artificial intelligence and the rapid development of 5G technology have further transformed the "Internet of Things" into the "Internet of Intelligence." However, traditional machine learning suffers from problems such as data silos and privacy security. Federated learning, as a representative of distributed machine learning, has attracted widespread attention due to its characteristic of "making data usable but not visible." In federated learning, all participants work collaboratively. Specifically, the server first publishes the model training task for federated learning. Mobile devices do not need to upload their sensitive data to the server; instead, they only need to use their data to train a shared model locally. The model is locally computed and updated by executing the training program, and finally, the server aggregates and updates the model. However, in practical applications of federated learning, considering risks such as resource consumption, privacy leaks, and break-even, data-sensitive devices are often unwilling to use all their high-precision data, or even refuse to participate in federated learning. Therefore, an important and challenging problem arises—how to design an effective incentive mechanism that not only optimizes the social welfare of the federated learning system but also attracts devices to participate in federated learning and contribute all their high-quality data.

[0003] Existing federated learning faces several challenges, such as the incentive problem of mobile devices contributing low-quality data or even refusing to participate. On the one hand, although methods such as homomorphic encryption and differential privacy improve the security of federated learning from the perspective of privacy protection, high computational costs and unknown reward policies still greatly hinder the promotion and development of federated learning. Due to incomplete information between mobile devices and servers and concerns about the balance of payments for devices, mobile devices participating in federated learning are often unwilling to contribute their high-precision data, leading to reduced efficiency and social welfare. Therefore, it is necessary to design an effective incentive mechanism to attract devices to contribute all their high-quality data and optimize the social welfare of federated learning. On the other hand, as a new paradigm for basic strategy research in game theory, the zero-determinant strategy can unilaterally control its own expected payoff to be linearly related to the opponent's expected payoff, which is expected to provide a solution to the incentive problem in federated learning. Existing federated learning incentive models based on the zero-determinant strategy should consider the following two points. First, the zero-determinant strategy is proposed based on the scenario of two discrete strategies, while the strategies of federated learning participants are generally continuous. Second, existing federated learning incentive methods are difficult to achieve a trade-off between efficiency and fairness.

[0004] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:

[0005] (1) System Efficiency. Currently, federated learning systems strictly protect user local data from leakage, transmitting only model updates. However, this also requires strict encryption before transmission. For more complex encryption systems, this means that information return transmission also requires more resources and time for decryption. The current approach is to strike a balance between data protection and system efficiency, protecting data privacy and security while maximizing the efficiency of the federated learning system, thereby achieving the goal of maximizing the efficiency of the federated learning system while ensuring data privacy and security.

[0006] (2) Incentive Mechanism. Federated learning systems require multi-party collaboration. In reality, there will inevitably be uneven distribution of computing power and resources among the parties. How to take this resource difference into account, formulate a flexible resource allocation mechanism, and design corresponding incentive mechanisms based on this difference are issues that need to be addressed not only from a technical perspective but also from a comprehensive perspective of market resources.

[0007] Due to factors such as incomplete information between mobile devices and servers and privacy sensitivities, mobile devices participating in federated learning are often unwilling to contribute their high-precision data, leading to reduced efficiency and reduced social welfare. Solving this problem requires incentive mechanisms with strong control and that encourage devices to contribute as much high-precision data as possible. Furthermore, concerns about the cost-benefit ratio of devices hinder the implementation of federated learning, and existing incentive methods struggle to strike a balance between efficiency and fairness.

[0008] (3) Unreasonable model settings. To facilitate the implementation of solutions, the designed mathematical models should be as realistic as possible. However, existing solutions designed for federated learning technology often have the following problems: For example, some works set up mobile devices to contribute their data to federated learning free of charge without considering privacy leaks, cost-benefit losses, etc., which is obviously unrealistic. Many related works also set the strategies of servers and devices in the problem model as discrete rather than continuous. That is, the server can only mechanically control a fixed reward level, and cannot flexibly adjust the payment based on the device's contribution. Or, it defines cooperation and betrayal based on the data quality and quantity of the device, instead of dynamically adjusting its contribution based on the reward level set by the server. These unreasonable problem settings often greatly affect the system efficiency of federated learning and even its actual implementation. Summary of the Invention

[0009] To address the problems existing in the prior art, this invention provides a federated learning incentive method, system, medium, device, terminal, and application, and particularly relates to a federated learning incentive method, system, medium, device, terminal, and application based on a continuous zero-determinant strategy.

[0010] This invention is implemented as follows: a federated learning incentive method, comprising: constructing a federated learning system based on a continuous zero-determinant policy; quantifying the policies of devices and servers; calculating the utility function and expected utility; using the zero-determinant policy in the continuous policy scenario to control the expected returns of participating devices and servers in the federated learning to satisfy a linear relationship; taking the maximization of social welfare in the federated learning as the optimal objective, and ensuring that utility is maximized only when devices and servers cooperate fully; solving the optimization problem, and designing the federated learning incentive mechanism algorithm based on the found optimal policy.

[0011] Furthermore, the federated learning incentive method includes the following steps:

[0012] Step 1: Construct a federated learning system based on a continuous zero-coefficient strategy and define the functions of each participating role in the mobile device and server.

[0013] Step 2: Define the strategies for participating devices and servers in federated learning, and quantify the utility function-related variables of devices and servers under different strategies;

[0014] Step 3: Define and prove that there is a dilemma in federated learning games that will harm social welfare, namely, the optimal strategy of the device is to provide incomplete, high-precision data, and the optimal strategy of the server is to set a low level of reward.

[0015] Step 4: Based on the zero determinant policy theory in the two-person continuous policy scenario, control the expected benefits of the participating devices and servers in federated learning to satisfy a linear relationship, ensuring that utility can be maximized only when the devices and servers cooperate fully.

[0016] Step 5: In a multi-player continuous strategy game scenario involving N devices and a federated learning server, maximizing social welfare is taken as the optimal objective of the incentive mechanism, and a federated learning incentive mechanism algorithm is designed based on the optimal strategy found by solving the problem.

[0017] Furthermore, the server in step one, acting as the administrator, constitutes the federated learning system and participates in the subsequent model aggregation and reward allocation process, specifically including:

[0018] (1) Issue the federated learning model training task and initialize the global model;

[0019] (2) Receive requests sent by the device and organize relevant devices to participate in pre-training for accuracy evaluation;

[0020] (3) Screen and evaluate the upper and lower bounds of the accuracy of the device in the pre-training process based on the sample data provided by the device, give a reward to the device that meets the pre-training requirements, and send the initial global model to it.

[0021] (4) Within a limited training period, after the device performs multiple rounds of local computation to converge the model parameters and gradient data to a certain set threshold, the local model sent by the receiving device is received.

[0022] (5) Aggregate all local models received in this stage and update the global model;

[0023] (6) Calculate the equipment contribution based on the relevant data from the local model and pay the remuneration accordingly;

[0024] (7) If the global model converges to the expected value, stop the iteration; otherwise, return to step (1).

[0025] Furthermore, in step one, the participants in building a hierarchical federated learning system based on a continuous zero-determinant strategy include servers and mobile devices; the servers are responsible for publishing and distributing machine learning tasks to participating devices to collaboratively train shared models; and the devices use the relevant data they possess to train federated learning models.

[0026] The entire model training process in federated learning consists of several rounds of global iterations, each of which covers multiple local iterations. Each device executes the training program locally to compute and update the model during its local iterations. During the global iteration, the server sends the trained model to the device, aggregates and updates the model after receiving multiple trained models.

[0027] Furthermore, in step two, the mobile devices, acting as managed entities, constitute the federated learning system and participate in subsequent local model training, specifically including:

[0028] (1) If the device has high-precision local data related to the broadcast federated learning task, it sends a participation request to the server.

[0029] (2) Participate in pre-training; qualified devices will receive the initial global model sent by the server.

[0030] (3) Use local data for iterative training and updating of local models;

[0031] (4) If the relevant data of the local model after training meets the threshold set by the server, then request to upload;

[0032] (5) Upload the updated local model and receive the reward paid by the server.

[0033] Furthermore, in step two, the interaction between the devices and servers involved in federated learning is modeled as a multi-player synchronous game in a continuous policy scenario. Policy definitions are given, and the utility function variables of the devices and servers under different policies are quantified, specifically including:

[0034] 1) Define the strategy for participating devices and servers in federated learning;

[0035] 2) Set constraints for the hybrid strategies of federated learning participants;

[0036] 3) Quantify the utility equations of federated learning participants.

[0037] Furthermore, in step three, the federated learning game is defined based on the interaction between the server and the device, including:

[0038] (1) The interaction between the device and the server is modeled as a multi-player synchronous repeated game, and defined as a federated learning game G={N∪{s},[l k ,h k ]∪[l,h],{u k}∪{u s}};

[0039] (2) The dilemma in federated learning games: the optimal strategy for the k-th device is l k The optimal strategy for the server is l;

[0040] (3) Introduce the zero determinant strategy.

[0041] Furthermore, step four, based on the zero-determinant strategy theory in continuous strategy scenarios, ensures that system utility is maximized only when devices and servers cooperate fully. This includes:

[0042] (1) Quantify the policy and utility functions of the devices and servers involved in federated learning, and calculate the expected utility;

[0043] (2) Using a continuous zero-determinant strategy to control the linear relationship between the expected utility of the device and the server;

[0044] (3) The optimization objective is to maximize social welfare through federated learning, and the solution is obtained by using the linear relationship controlled by the zero determinant strategy.

[0045] Furthermore, in step five, maximizing social welfare is taken as the optimal objective of the incentive mechanism. Based on the optimal strategy found through the solution, an incentive mechanism algorithm is designed to maximize the social welfare of the entire federated learning system. The following formula is used for solving:

[0046] The social benefits of the federal learning system consist of a weighted sum of two parts:

[0047] The expected utility of the k-th participating device in federated learning is:

[0048] The expected utility of the federated learning server is:

[0049] The social benefits of the federal learning system are:

[0050] Utilize to satisfy Continuous zero determinant strategy:

[0051]

[0052] Then we get:

[0053]

[0054] The objective optimization problem is as follows:

[0055] P2: maxγ

[0056]

[0057] Another object of the present invention is to provide a federated learning incentive system that applies the aforementioned federated learning incentive method, the federated learning incentive system comprising:

[0058] The system building module is used to build a federated learning system based on a continuous zero determinant policy, quantify the policies of devices and servers, and calculate the utility function and expected utility.

[0059] The linear relationship control module is used to control the expected returns of the devices and servers participating in federated learning to satisfy a linear relationship by utilizing the zero-determinant policy in continuous policy scenarios.

[0060] The utility maximization module is designed to maximize social welfare through federated learning, ensuring that social welfare is maximized only when devices and servers cooperate fully.

[0061] The incentive mechanism design module is used to solve optimization problems and design the incentive mechanism algorithm for federated learning based on the found optimal strategy, so as to maximize the social welfare of the entire federated learning system.

[0062] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the federated learning incentive method.

[0063] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the federated learning incentive method.

[0064] Another object of the present invention is to provide an information data processing terminal for implementing the aforementioned federated learning incentive system.

[0065] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:

[0066] First, this invention deploys an incentive mechanism algorithm based on a continuous zero-determinant policy in the federated learning system. By attracting mobile devices to participate in the federated learning model training task, it can control their rewards to encourage them to contribute their high-precision data, thereby improving federated learning efficiency and avoiding the predicament of low-quality data contributed by participating devices. This significantly improves computational complexity and the time required to complete the federated learning process. This invention effectively attracts mobile devices to participate and contribute their high-precision data, significantly improving computational complexity and the time required to complete the federated learning process. Furthermore, this invention facilitates adaptively incentivizing the contributions of federated learning participants, completing the federated learning model training task, designing optimal response policies, and dynamically adjusting reward levels based on device actions and policies, thereby improving the efficiency and security of federated learning.

[0067] This invention provides a federated learning incentive mechanism based on a continuous zero-determinant policy. During the iterative training of the federated learning model, the interaction between mobile devices and the server is modeled as a repeated multi-player continuous policy game. To reflect the complex relationships among multiple participants, this invention proposes an incentive mechanism based on a continuous zero-determinant policy to attract mobile devices to contribute their high-precision data, thereby optimizing social welfare and improving federated learning efficiency. Furthermore, this invention extends the theory of multi-player zero-determinant policies. The proposed incentive mechanism based on a continuous zero-determinant policy to optimize the social welfare of federated learning not only helps incentivize as many mobile devices as possible to contribute high-quality data in federated learning, helping the server improve learning efficiency and enabling social welfare to converge to a high and stable level, but also further enriches the application of zero-determinant policies in continuous policy scenarios. The incentive mechanism based on a continuous zero-determinant policy provided by this invention is simple in design and flexible in implementation, effectively attracting devices to contribute all their high-precision data, thereby improving federated learning efficiency and optimizing social welfare. This invention also provides a new method with low computational requirements, ensuring the fairness of the utility between mobile devices and the server, and the relative stability of social welfare. In other words, the incentive problem for mobile devices in federated learning is formulated as an optimization problem, in which the feasibility of the incentive mechanism is also guaranteed.

[0068] This invention, applied to the field of federated learning, focuses on addressing the incentive problem of promoting full cooperation between mobile devices and servers to contribute high-precision data. It designs an incentive algorithm based on a continuous zero-determinant strategy to optimize the social welfare level of federated learning. Experimental results show that the incentive mechanism based on the continuous zero-determinant strategy proposed in this invention is effective and reasonable, ensuring relatively stable utility for both mobile devices and servers, while significantly improving the social welfare and efficiency of the federated learning system.

[0069] Second, the incentive mechanism based on the continuous zero-determinant strategy of the present invention can incentivize as many devices as possible to participate in federated learning and contribute high-quality data, thereby optimizing social welfare and improving the efficiency of federated learning.

[0070] Third, as supplementary evidence of the inventive step of the claims of this invention, it is also reflected in the following important aspects:

[0071] (1) The expected benefits and commercial value of the technical solution of this invention after transformation are as follows:

[0072] By introducing the technical solution of this invention, it is possible to effectively attract as many devices as possible to participate in federated learning and contribute their high-precision data, thereby effectively optimizing the level of social welfare and improving the efficiency of federated learning. Specifically, the server can dynamically adjust the reward level according to the contribution of mobile devices to achieve a fair and reasonable reward distribution, controlling the server's own income and the income of mobile devices to satisfy a linear relationship, thereby promoting mobile devices to contribute their high-precision data as much as possible. Devices can dynamically adjust their data quality or data volume in conjunction with the rewards set by the server to achieve a balance between income and expenditure. In order to give full play to the incentive effect of the technical solution of this invention, the reward level set by the server and the contribution of the devices should satisfy a fair linear relationship. This corresponds to the characteristic of the zero-determinant strategy that it can unilaterally control the expected income of the opponent to satisfy a linear relationship with its own expected income. In addition, the strategy of federated learning participants is extended from discrete to continuous, thereby optimizing the level of social welfare and improving the efficiency of federated learning. While facilitating the modeling and solving of the federated learning incentive problem, it also facilitates the implementation of the designed federated learning incentive algorithm based on the continuous zero-determinant strategy, which corresponds to the effectiveness and feasibility of the technical solution of this invention.

[0073] (2) The technical solution of this invention fills a technical gap in the industry both domestically and internationally:

[0074] Existing techniques for modeling incentive problems in federated learning using game theory often consider discrete policy scenarios. This means they categorize the server's rewards and the quality of data contributed by devices into cooperation or betrayal to facilitate modeling and solving the problem using game theory. However, in actual federated learning model training, participants' strategies should be considered continuous. Mobile devices should dynamically adjust the quality or quantity of data they contribute based on the server's reward level, rather than being forced to participate and contribute all data, or, in extreme cases, not participate at all. The server should dynamically control the payoff relationship between the two parties based on the device's contribution to model training to prevent devices from "cheating" by contributing only partial or irrelevant data, rather than being forced to choose between full payment or zero payment strategies. The technical solution of this invention models the interaction between devices and the server in federated learning as an iterative game. Based on a continuous zero-determinant strategy, a federated learning incentive method is designed to optimize social welfare. In this method, devices can dynamically adjust their data quality or quantity according to the rewards set by the server to achieve a balance between income and expenditure. The server can dynamically adjust the reward level based on the contribution of mobile devices to achieve a fair and reasonable reward distribution. The method ensures a linear relationship between the server's own income and the income of mobile devices, thereby encouraging mobile devices to contribute as much high-precision data as possible. Based on theoretical analysis and experimental evaluation, by applying the continuous zero-determinant strategy, the technical solution of this invention can effectively help the server control social welfare at a high level and stably, thereby optimizing the level of social welfare and improving the efficiency of federated learning.

[0075] (3) Whether the technical solution of the present invention solves the technical problem that people have long wanted to solve but have never been able to solve successfully:

[0076] Existing technical solutions using game theory to solve the incentive problem in federated learning either contain unreasonable assumptions, such as assuming mobile devices participate in federated learning model training free of charge, or setting participants' strategies to be discrete. While this may be for the convenience of problem modeling and optimization, existing methods largely possess theoretical value but lack practical application, which clearly hinders the improvement of federated learning efficiency. Improving and optimizing federated learning efficiency and social welfare corresponds to the coordination and control between the data contributed by mobile devices and the rewards controlled by the server, which has always been a challenge in the incentive problem of federated learning. To solve this technical problem, we turn to game theory. Zero-determinant strategies, which have a strong control effect in the repeated prisoner's dilemma, can unilaterally control a linear relationship between their own expected payoff and the expected payoff of their opponents, thereby controlling the opponent's strategies and actions. Based on the zero-determinant strategy, we extend it to continuous strategy scenarios to facilitate its practical application. We model the interaction between the server and the device as a multi-player iterative game and design an incentive algorithm based on the continuous zero-determinant strategy to help the server control its expected payoff to satisfy a linear relationship with the expected payoff of the opponent. We also dynamically and fairly control the reward level according to the contribution of the mobile device. This can not only attract as many mobile devices as possible to participate in federated learning, but also promote the devices to contribute all their high-precision data, further optimize the social welfare level and effectively improve the efficiency of federated learning.

[0077] (4) Does the technical solution of the present invention overcome technical bias?

[0078] The optimization objectives of the federated learning incentive problem generally include incentive compatibility (the optimal strategy is equivalent to full cooperation among all participants), individual rationality (participants' decision-making strategies are only related to their payoffs), budget equilibrium (the payoffs of all participants exceed their expenditures), social optimization (maximizing the weighted sum of the expected payoffs of all participants), and distributive fairness (the payoffs of participating in federated learning are related to their contributions). However, due to limitations in the development and conditions of game theory, there is a trade-off between these objectives. In practical applications, we often select only the most important objectives for optimization. The zero-determinant strategy proposed by Press and Dyson in 2012 has greatly promoted the development of game theory. By adopting the zero-determinant strategy, players can unilaterally control the linear relationship between their expected payoffs and those of their opponents, thereby controlling the behavior of their opponents. However, the zero-determinant strategy was originally designed for two-player prisoner's dilemma games, where players only have two discrete strategies: cooperation and betrayal. This is clearly different from the federated learning scenario. In the technical solution of this invention, we extend it to the continuous strategy scenario, modeling the interaction between the server and the device as a multi-player iterative game. The server uses a continuous zero-determinant strategy to control the relationship between its own payoff and the device's payoff. Based on the continuous zero-determinant strategy, we designed a new incentive method to help the server dynamically adjust the reward level paid according to the device's contribution to the training of the federated learning model. The device can also dynamically adjust its data quality and data volume according to the rewards received by the server. On the basis of ensuring budget balance and fair distribution among federated learning participants, the goal is to maximize social welfare, thereby achieving social welfare maximization and improving the efficiency of federated learning. Attached Figure Description

[0079] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0080] Figure 1 This is a flowchart of the federated learning incentive method provided in an embodiment of the present invention;

[0081] Figure 2 This is a schematic diagram of a federated learning system based on a continuous zero determinant strategy provided in an embodiment of the present invention;

[0082] Figure 3 This is a comparison chart of social welfare when the total number of devices N=1, as provided in the embodiments of the present invention;

[0083] Figure 4This is a comparison chart of social welfare when the total number of devices N=10, as provided in the embodiments of the present invention;

[0084] Figure 5 This is a graph showing the changes in social welfare when the factor χ takes different values, provided in an embodiment of the present invention.

[0085] Figure 6 This is a graph showing the changes in social welfare when the factor β takes different values, provided in an embodiment of the present invention.

[0086] Figure 7 This is a comparison chart of the stable values ​​of the relative utility of federated learning participants when the device adopts the TFT strategy and the server device adopts different strategies, provided by an embodiment of the present invention.

[0087] Figure 8 This is a comparison chart of the stable values ​​of the relative utility of federated learning participants when the device adopts a WSLS strategy and the server device adopts different strategies, provided by an embodiment of the present invention.

[0088] Figure 9 This is a graph showing the change in relative device utility when all devices adopt the WSLS strategy and the server adopts different strategies, provided by an embodiment of the present invention.

[0089] Figure 10 This is a graph showing the change in the relative utility of the server when all devices adopt the WSLS strategy and the server adopts different strategies, as provided in an embodiment of the present invention. Detailed Implementation

[0090] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0091] To address the problems existing in the prior art, the present invention provides a federated learning incentive method, system, medium, device, terminal, and application. The present invention will be described in detail below with reference to the accompanying drawings.

[0092] like Figure 1 As shown, the federated learning incentive method provided in this embodiment of the invention includes the following steps:

[0093] S101, Construct a federated learning system based on a continuous zero determinant policy, quantify the policies of devices and servers, and calculate the utility function and expected utility;

[0094] S102 utilizes a zero-determinant policy in a continuous policy scenario to control the expected returns of participating devices and servers in federated learning to satisfy a linear relationship.

[0095] S103 aims to maximize social welfare through federated learning and ensures that utility is maximized only when devices and servers cooperate fully.

[0096] S104, solve the optimization problem, and design the federated learning incentive mechanism algorithm based on the found optimal strategy to maximize the social welfare of the entire federated learning system.

[0097] As a preferred embodiment, the federated learning incentive method provided by this invention includes the following steps:

[0098] S1. Construct a federated learning system based on a continuous zero-determinant strategy and define the functions of each participating role, such as mobile devices and servers.

[0099] S2 defines the policies for devices and servers, and quantifies the utility functions of devices and servers under different policies and other related variables.

[0100] S3 defines and proves that there is a dilemma in federated learning games that will harm social welfare, namely, the optimal strategy of the device is to provide incomplete, high-precision data, and the optimal strategy of the server is to set a low level of reward.

[0101] S4, based on the zero determinant policy theory in the two-person continuous policy scenario, ensures that the utility can be maximized only when the device and the server fully cooperate (i.e., choose to join federated learning and contribute all their high-precision data);

[0102] S5 proposes a multi-player continuous strategy game scenario involving N devices and a federated learning server. It takes maximizing social welfare as the optimal objective of the incentive mechanism and designs a federated learning incentive mechanism algorithm based on the optimal strategy found by solving the problem.

[0103] In order to implement real-world application scenarios, such as Figure 2 As shown, this embodiment of the invention first constructs a federated learning system based on a continuous zero-determinant policy. The federated learning system based on a continuous zero-determinant policy constructed in step S1 of this embodiment includes a server and a mobile device;

[0104] In federated learning, the server is responsible for publishing machine learning tasks and distributing them to participating devices to collaboratively train a shared model. Devices use their available data to train the federated learning model. Because the raw data used for model training is stored locally on mobile devices, rather than on cloud servers, this greatly ensures the data security and privacy of the devices.

[0105] The entire model training process in federated learning consists of several rounds of global iterations, each of which covers multiple local iterations. Each device executes the training procedure locally to compute and update the model during its local iterations. During the global iterations, the server sends the trained model to the devices and, upon receiving multiple trained models, aggregates and updates the model.

[0106] The embodiment of the present invention provides step S2, which defines the strategies for participating devices and servers in federated learning and quantifies relevant variables such as the utility functions of devices and servers under different strategies, including the following steps:

[0107] S21 defines the strategy for participating devices and servers in federated learning.

[0108] The strategy for the k-th device in the current round can be defined as x k ∈[l k h k This refers to the local data accuracy controlled by the device, which is then selected for training in federated learning. Considering the incomplete information among federated learning participants, this can be achieved through pre-training. k and h k The current round server's strategy can be defined as y∈[l,h], which is the reward level for all devices under the server's control. These variables will obviously affect the accuracy of the trained global model.

[0109] The policy function of the k-th device can be modeled using curve fitting with the following equation:

[0110]

[0111] in, Δ k This is the bulldozer distance (EMD) metric for the k-th device, often used to quantify the variability in a data distribution. A larger EMD indicates greater height variation, which can negatively impact the accuracy of local models. k This represents the size of the federated learning data associated with the k-th device. Here, c i (i∈{1, 2, 3, 4, 5, 6}) are positive curve fitting parameters. The first term π(Δ) k ) captured when Δ k The performance of the model decreases as the data size increases. The exponential term reflects that the larger the data quality and scale, the better the model accuracy. Therefore, the accuracy of the federated learning global model can be expressed as:

[0112]

[0113] in, It is the average EMD value of all devices during the training process of the federated learning model.

[0114] S22 sets constraints for the hybrid strategies of federated learning participants.

[0115] In the previous round, x′ was taken from the kth device. k Given that the server takes y′, the k-th device and the server will each take x in this round. k The conditional probability of y is defined as p k (x k ) = p k (x′1,…,x′ N ,y′;x k ) and q(y)=q(x′1,…,x′ N If ,y′;y), then the mixing strategy of the k-th device has constraints. Similarly, the constraints of the server's hybrid strategy are:

[0116] S23, the utility equation for quantifying federated learning participants.

[0117] To compensate for the model training costs of the equipment, a payment strategy is implemented. The server determines the reward paid to the k-th device based on its contribution level, where If τ and τ are constant coefficients, then the utility equation for the k-th device can be defined as:

[0118] u k =σ k r k -δ k C k (x k ),

[0119] Where, σ k ,δ k >0 is a weighting parameter that balances revenue and costs; C k (x k ) is the total cost of the k-th device, regarding x. k Monotonically increasing.

[0120] Accordingly, the utility equation for the server can be defined as:

[0121]

[0122] Where, σ s ,δ s >0 is a weighting parameter that balances model accuracy and total payout; ν is the expected revenue of the server corresponding to a unit of global model accuracy; ρ is a positive adjustment parameter; R is the server's transmission and aggregation costs.

[0123] The step S3 provided in this embodiment of the invention maps the interactions between federated learning participants to a game, and proves that there exists a dilemma in federated learning games where the optimal response of a participant is a low-level strategy. Specifically, this includes:

[0124] S31 maps the interactions between federated learning participants to a game.

[0125] Based on the interaction between the server and the device, this invention can define federated learning game as follows:

[0126] Definition 1 (Federated Learning Game): In a federated learning scenario consisting of one server and N devices, the devices can decide the quality and size of the data input for model training, while the server can control the reward level given to the devices. The interaction between the devices and the server is equivalent to a multi-player synchronous repeated game, called the federated learning game G={N∪{s},[l k ,h k ]∪[l,h],{u k}∪{u s}}, where k={1,2,…,N}.

[0127] S32, The Dilemma in Federated Learning Games.

[0128] The optimal action for a rational participant corresponds to the action that maximizes their own profit. In fact, according to the following proposition, there is a dilemma in federated learning games.

[0129] Proposition 1: In the federated learning game of Definition 1, the optimal strategy for the k-th device is l. k The optimal strategy for the server is l.

[0130] Proof: For the k-th device, according to formula u k =σ k r k -δ k C k (x k If the fixed server policy y is used, then the device's utility u k (x k ) is about x k A monotonically decreasing function. That is, regardless of the server's strategy, the optimal strategy for the k-th device is l. k Let's verify the conclusion of this invention through the following two examples. If y = h, then r k It will take the maximum value. However, considering the total cost C k (l k ) < C k (h k This invention has u k(l k )<u k (h k This leads to its optimal strategy being x. k =l k Conversely, if y = l, strategy x k =h k This will result in even greater equipment damage.

[0131] Similarly, for the server, the strategy x of fixing the k-th device... k After taking any value, the server's utility function u is now... s (y) is a function that is monotonically decreasing with respect to y. Considering r k (l)<r k (h), then u s (l)>u s (h). According to formula Only when r k Minimize server efficiency as much as possible. s Only when (y) reaches its maximum value can this be achieved, which corresponds to y = 1. In other words, regardless of the strategy adopted by the device, the optimal strategy for the server is y = 1. Q.E.D.

[0132] S33 introduces a zero determinant strategy to alleviate the dilemma in federated learning games.

[0133] As can be seen from Proposition 1, the optimal strategy for the devices is to provide incomplete, low-precision data, while the optimal strategy for the servers is to set a low level of reward. This dilemma will severely damage social welfare and even hinder federated learning. As the initiator and manager of federated learning, the servers have the motivation and ability to take preventative measures (such as incentive mechanisms) to attract the full contribution of the devices and avoid this dilemma.

[0134] However, on the one hand, the server's utility is influenced by the policies of all participants, making it difficult to unilaterally adopt the optimal strategy. On the other hand, the additional cost of preventative measures is a key challenge. Zero-determinant strategies can help adopters unilaterally control the proportional relationship between the expected utilities of the two participants, potentially resolving the federated learning dilemma.

[0135] It is important to note that traditional zero-determinant strategies are suitable for discrete-policy games with two participants, while in federated learning, participants' policies are typically continuous. Therefore, this invention focuses on analyzing continuous-policy games involving multiple participants and designs a federated learning incentive mechanism algorithm based on continuous zero-determinant (CZD) strategies to attract all contributions from the devices.

[0136] Step S4 in this embodiment of the invention, based on the zero-determinant policy theory in a continuous policy scenario, ensures that utility is maximized only when the device and server cooperate fully. Specifically, this includes:

[0137] Consider the two-person continuous policy scenario when N=1, which is a special case where there is a federated learning participant device and a server.

[0138] S41 defines the strategy and utility function of federated learning participants and calculates the expected utility.

[0139] Let the strategy equation and utility function of the k-th device be x. k and u k =u k (x k The policy equation and utility function of the federated learning server are y and u, respectively. s =u s (x k ,y).

[0140] Define the selection strategy x′ for the k-th device in the previous round. k The probability that the server chooses strategy y′ is z(x′). k ,y′). In the current round, the k-th device selects strategy x. k The probability that the server chooses strategy y is z(x) k If the steady-state vector satisfies z = z(x′, y). k ,y′)=z(x k If ,y), then the expected utility of the k-th device in each round of the game can be expressed as The expected utility of the server is

[0141] S42 uses a continuous zero determinant strategy to ensure that the payoffs of both parties are linearly related.

[0142] According to the theory of zero determinant strategy, the following lemma can be obtained:

[0143] Lemma 1: For x k ∈[l k h k ], y∈[l,h], if the server's policy satisfies and The expected utility of the device and server will then satisfy αE. k +βE s -γ = 0, where χ is a non-zero scalar, α and β are weighting factors, and γ is a parameter.

[0144] Proof: First, introduce a sufficiently small number ε. k The strategy x of the k-th devicek Divided into l k l k +ε k 、···、l k +nε k =h k This is equivalent to dividing the data used by the device for federated learning training tasks into n parts, each part being of size n. Similarly, we introduce a sufficiently small number ε to divide the server's policy y into l, l+ε, ..., l+nε=h. This is equivalent to dividing the server's configurable reward into n parts, each of size h. When n is large enough, or ε k When both ε and ε approach zero, the present invention can consider the policy equations of the device and the server to be continuous.

[0145] The conditional probability of a device can be defined as follows: The conditional probability of a server can be defined as follows: i∈{1,2,…,n}, j∈{1,2,…,n}, k∈{1,2,…,n}, in the previous round of device and server policy functions (l k +iε k ) and (l+jε), respectively. In the current round, the server's policy function is (l+kε).

[0146] Based on the policy functions of the devices and servers defined under initial conditions, the two-player iterative game can be viewed as a stochastic process with Markov chain characteristics. By introducing a stationary vector v, the state transition matrix can be defined as... as well as Where H ij It is the conditional probability (l) of the strategy the device will adopt in the current round. k +iε k The conditional probability of the server adopting strategy (l+jε) is considered, taking into account all results from the previous round. Therefore, for i∈{1,2,…,n}, j∈{1,2,…,n}, the present invention has

[0147] Therefore, the revenue matrix of the k-th device can be defined as:

[0148]

[0149] The server's revenue matrix can be defined as:

[0150]

[0151] Therefore, the expected utility of the k-th device is E′. k =vT ·U k The expected utility of the server is E′ s =v T ·U s .

[0152] set up Where I is the identity matrix, therefore

[0153] According to Cramer's rule, this invention has in The auxiliary matrix. Therefore, Each column is proportional to v. Therefore, The first column can be replaced with v. Furthermore, the present invention has...

[0154] Here, the present invention can use a determinant to represent any vector f = [f1, f2, ..., f nn ] T The dot product with a stationary vector v. This is achieved by utilizing elementary transformations and introducing a vector. The outcome will depend on a strategy with only one participant, namely:

[0155]

[0156] When f=αU k +βU s -γ1, with α and β as weighting factors, and γ being a non-zero scalar, then:

[0157] v·f=v(αU k +βU s -γ1)=αE′ k +βE′ s -γ.

[0158] When h nn =χ(αU k +βU s When -γ1), according to the determinant property, we have αE′ k +βE′ s -γ=0.

[0159] When ε k As ε approaches zero, the policy equations for the device and server are continuous, and the expected utility E′ of the device is... k Approaching E k The expected utility of the server E′ s Approaching E s , i.e. αE k +βE s -γ = 0. Q.E.D.

[0160] S43, with the goal of maximizing social welfare through federated learning, establishes and solves an optimization problem.

[0161] When the server adopts a continuous zero-determinant strategy, social welfare can always be maintained at a high and stable level, regardless of the strategy employed by the device. This objective can be transformed into the following optimization problem:

[0162]

[0163]

[0164] Based on the given conditions, this problem can be transformed into:

[0165] P1:

[0166]

[0167] If χ < 0, note the constraint q(y) ≤ 1, and assume... Then we have:

[0168]

[0169] in,

[0170] Note the constraint q(y)≥0 and assume Then we have:

[0171]

[0172] in,

[0173] Obviously γ min ≤γ max That is, max(w2) ≤ min(w1).

[0174] Since χ can take any negative number with a very small absolute value, we have:

[0175]

[0176] Similarly, if χ > 0, then:

[0177]

[0178] In summary, the maximum value of γ is as follows:

[0179]

[0180] Therefore, in order to maintain social welfare at a relatively high level, the CZD strategy adopted by the server needs to meet the following requirements.

[0181] S44, Based on the optimization problem, organize and derive the conclusions for the two-person continuous strategy scenario.

[0182] By solving the above optimization problem and setting... (i = 0, 1) yields the following theorem:

[0183] Theorem 1: When the server's policy satisfies At this time, social welfare γ can reach its maximum value.

[0184] At this point, the social welfare of the federal learning system can be unilaterally controlled at γ. max In other words, regardless of the strategy employed by the device, by applying a continuous zero-column strategy, the server can unilaterally control social welfare to the desired level.

[0185] In step S5 of this embodiment of the invention, maximizing social welfare is taken as the optimal objective of the incentive mechanism, and an incentive mechanism algorithm is designed based on the optimal strategy found by the solution to maximize the social welfare of the entire federated learning system. The following formula is used for solving:

[0186] S51, calculate the expected utility of participants in federated learning based on their policies and utility functions in a multi-person continuous policy scenario.

[0187] Consider a multi-player continuous policy scenario, where one server and N devices play a game using continuous policies in an iterative federated learning model training task.

[0188] When the strategy equation for the k-th device is x k ∈[l k h k When k = 1, 2, ..., N, and the server's policy equation is y ∈ [l, h], the utility function of the k-th device can be defined as u k =u k (x1,…,x N Given k = 1, 2, ..., N, the utility function of the server is u. s =u s (x1,…,x N ,y).

[0189] To more intuitively reveal the relationship between variables, we can assume...

[0190] Define the selection strategy x′ for the k-th device in the previous round. kThe probability that the server chooses strategy y′ is z(x′1,…,x′). N ,y′). In the current round, the k-th device selects strategy x. k The probability that the server chooses strategy y is z(x1,…,x). N If ,y), then the corresponding transformation equation can be written as:

[0191]

[0192] The state transition equation can then be written as:

[0193] z(x′1,…,x′ N ,y′;y′)M(x′1,…,x′ N ,y′;x1,…,x N ,y)=z(x1,…,x N ,y).

[0194] If the steady-state vector satisfies z = z(x1,…,x) N ,y)=z(x′1,…,x′ N If ,y′), then the expected utility of the k-th device can be written as:

[0195]

[0196] The expected utility of a server can be written as:

[0197]

[0198] S52 uses a continuous zero determinant strategy to control the returns to be linear.

[0199] Introducing a weighting factor α k If k = 1, 2, ..., N and β, then the weighted sum of the expected utilities of all game participants can be defined as the social welfare of the federated learning system, i.e. Using the zero determinant policy theory in continuous policy scenarios, the following lemma can be obtained:

[0200] Lemma 2: For x k ∈[l k h k ], k = 1, 2, ..., N, y ∈ [l, h], if the server's policy satisfies and The expected utility of the equipment and server will then be satisfied. Where χ is a non-zero scalar, and α k (k = 1, 2, ..., N) and β are weighting factors, and γ is a parameter.

[0201] Proof: If there are (N-1) devices and one server in a federated learning game, according to the proof of Lemma 1, the policy of the k-th device in the current round is l. k +kε k The conditional probability can be defined as:

[0202]

[0203] for The conditional probability that the server's current policy function is l+kε can be defined as:

[0204]

[0205] Therefore, the transition matrix M of the Markov chain N Let v be a stationary vector.

[0206]

[0207] And the stationary vector v is:

[0208] v T ·M N =v T ;

[0209] in, and It is the Hadamard product in matrix multiplication.

[0210] Let M′ N =M N -I, therefore v·M′ N =0. According to Kramer's rule, we can obtain adj(M′) N Each row of adj(M′) is linearly related to v. Therefore, adj(M′) N The first line of ) can be replaced with v.

[0211] This can be represented by a determinant. The dot product with a stationary vector v. This is achieved by utilizing elementary transformations and introducing a vector. The outcome will depend on a strategy with only one participant, namely:

[0212]

[0213] The revenue matrix of the kth device can be defined as follows: The server's revenue matrix can be defined as follows: Therefore, the expected utility of the k-th device is E′. k =v T ·U k , k∈{1,2,…,N-1}, and the expected utility of the server is E′ s=v T ·U s .

[0214] When the server adopts a policy and introduce α1, ..., α N-1 f = α1U1 + ... α, where β and γ are scalars. N-1 U N-1 +βU s When -γ1, the present invention has:

[0215] v T ·f=v T (α1U1+…+α N-1 U N-1 +βU s -γ1)=α1E′1+…+α N-1 E′ N-1 +βE′ s -γ=0.

[0216] When ε k As ε approaches zero, the policy functions of both the device and the server are continuous. Therefore, E′ k Approaching E k , k∈{1,2,…,N-1} and E′ s Approaching E s That is, α1E1+…+α N-1 E N-1 +βE s -γ=0,

[0217] Additionally, when there are N devices and one server, the policy function for the k-th device and server in the current round is (l k +kε k The conditional probabilities of (l+kε) and (l+kε) can be defined as:

[0218] Based on this definition, when the number of players in a federated learning game is (N+1), this invention has... and v T ·M N+1 =v T M N+1 It is the Markov chain transition matrix, v is the steady-state vector, and g is the Markov chain transition matrix. N =[g1,g2,…,g n ] T , It is the Kronecker product.

[0219] Let M′ N+1 =M N+1 -I, therefore v·M′ N+1 =0. According to Cramer's rule, this invention can obtain adj(M′)N+1 )M′ N+1 =det(M′ N+1 ) = 0, adj(M′ N+1 Each row of ) is proportional to v. Therefore, we can use adj(M′) N+1 Replace v with the first line of )

[0220] Furthermore, a vector is defined as in

[0221] because Through some elementary transformations, the dot product of any vector f and a stationary vector v can be expressed as: Where 1 is a 1×n N-1 A 3D vector, where all elements are 1.

[0222] Therefore, in a federated learning game with N players, this invention can be represented using a determinant. The dot product of the stationary vector v and the server's policy, where one column depends only on the server's policy. in

[0223] The revenue matrix of the kth device is: The server's revenue matrix is Therefore, the expected utility of the k-th device is E′. k =v T ·U k , k∈{1,2,…,N-1}, and the expected utility of the server is E′ s =v T ·U s .

[0224] When the server adopts a policy and introduce α1, ..., α N f = α1U1 + ... α, where β and γ are scalars. N U N +βU s When -γ1, the present invention has:

[0225] v T ·f=v T (α1U1+…+α N U N +βU s -γ1)=α1E′1+…+α N E′ N +βE′ s -γ=0.

[0226] When ε kAs ε approaches zero, the policy functions of both the device and the server are continuous. Therefore, E′ k Approaching E k , k∈{1,2,…,N} and E′ s Approaching E s That is, α1E1+…+α N E N +βE s -γ=0,

[0227] Through mathematical induction, it is proved that in an (N+1)-player federated learning game (N≥2), when the server adopts a CZD strategy, it can unilaterally control the expected utility of all participants, maintaining social welfare at a high and stable level. Therefore, Lemma 2 is proved.

[0228] S53 establishes and solves an optimization problem with the goal of maximizing social welfare through federated learning.

[0229] When the server adopts a continuous zero-determinant strategy, social welfare can always remain at a high and stable level, regardless of the device's strategy. This objective can be transformed into the following optimization problem:

[0230] P2: maxγ

[0231]

[0232] Similar to solving for class P1, it is easy to obtain:

[0233]

[0234] Therefore, in order to optimize the social welfare of federated learning, the CZD strategy adopted by the server needs to satisfy:

[0235]

[0236] S54, Based on the optimization problem, organize and derive the conclusions for the two-person continuous strategy scenario.

[0237] By solving the above optimization problem, the following theorem can be obtained:

[0238] Theorem 2: When the server's policy satisfies:

[0239]

[0240] At this time, social welfare γ can reach its maximum value:

[0241]

[0242] At this point, the social welfare of the federal learning system can be unilaterally controlled at γ. maxIn other words, regardless of the strategy adopted by the device, by applying a continuous zero-coefficient strategy, the server can unilaterally control social welfare to remain at a high and stable level.

[0243] S55 is a federated learning incentive mechanism algorithm based on continuous zero determinants (see Algorithm 1).

[0244] Algorithm 1: Federated Learning Incentive Mechanism Based on Continuous Zero-Determinant Rows

[0245]

[0246] Considering the selfishness of participating devices in federated learning and the continuity of their actual strategies, this invention designs a federated learning incentive mechanism based on the CZD policy. Algorithm 1 provides a detailed description of the incentive mechanism algorithm. Considering all x... k Both and y are continuous bounded variables. In practical applications, a water-filling algorithm with a complexity of O(N) can be used to approximate the optimal solution. Since this algorithm is equivalent to traversing all participating devices, the complexity of Algorithm 1 is O(N). The specific process is as follows:

[0247] (1) Initialization. The server publishes the federated learning model training task and initializes the global model to obtain θ. The k-th device sends a request to the server and participates in pre-training (line 1).

[0248] (2) Selection. In pre-training, if the accuracy value of the sample data of the k-th device meets the threshold Θ... k The server can assess its upper bound of accuracy, h. k and lower bound l k Then, the CZD strategy is applied to give the predicted reward level (lines 2-8).

[0249] (3) Model Training. N devices that meet the training requirements input their high-precision data into the downloaded global model. After several rounds of local iterative training, the local model (including gradients, parameters, etc.) calculated locally is trained until it meets the threshold Θ set by the server. k Then, the trained local model will be sent to the server. Otherwise, the iteration will continue until the end of one round of federated learning (lines 9–16).

[0250] (4) Model aggregation and reward allocation. In one iteration (T) k In the case of ≤Τ), the server aggregates the local models sent by all devices that meet the requirements within a finite time. Then, the server calculates the model based on the contribution of the k-th device (e.g., the local model accuracy x). k To calculate and distribute rewards r kIf the aggregated global model θ′ satisfies the pre-set expectation metric Θ by the server, the federated learning task ends. Otherwise, the next round of model iteration training will begin until it satisfies the federated learning task objective (lines 17–26).

[0251] like Figure 2 As shown, the federated learning incentive system provided in this embodiment of the invention includes:

[0252] The system building module is used to build a federated learning system based on a continuous zero determinant policy, quantify the policies of devices and servers, and calculate the utility function and expected utility.

[0253] The linear relationship control module is used to control the expected returns of the devices and servers participating in federated learning to satisfy a linear relationship by utilizing the zero-determinant policy in continuous policy scenarios.

[0254] The utility maximization module is designed to maximize social welfare through federated learning, and ensures that utility is maximized only when devices and servers cooperate fully.

[0255] The incentive mechanism design module is used to solve optimization problems and design the incentive mechanism algorithm for federated learning based on the found optimal strategy, so as to maximize the social welfare of the entire federated learning system.

[0256] As a preferred embodiment of the present invention, such as Figure 2As shown, this embodiment of the invention constructs a federated learning system based on a continuous zero determinant (CZD) strategy to implement real-world application scenarios. Specifically, in the initial stage, the server publishes a federated learning model training task and initializes a global model. Devices send a task reception request to the server and participate in pre-training. In the preparation and screening stage, during pre-training, if the accuracy of a device's sample data meets the set threshold, the server can evaluate its accuracy upper and lower bounds and then apply the CZD strategy to give the predicted reward level. N devices that meet the training requirements input their high-precision data into the downloaded global model for local iterative training. After several rounds of local iterative training, if the local training model (including gradients, parameters, etc.) obtained by the device meets the accuracy threshold set by the server, the trained local model can be sent to the server. Otherwise, local iterative training will continue until the end of a round of federated learning task. In the model aggregation and reward allocation stage, the server aggregates the local models sent by all devices that meet the conditions in a round of iteration within a limited time. Subsequently, the server calculates and allocates the reward due to each device based on its contribution (such as the accuracy of the local model) during this round of federated learning and the CZD strategy. During the acceptance phase, if the global model aggregated by algorithms such as weighted averaging meets the preset expected value, the federated learning model training task ends; otherwise, the next round of model iteration training will be restarted until the global model meets the model accuracy threshold.

[0257] A significant and equally challenging problem in federated learning systems is how to design an effective incentive mechanism that not only optimizes the social welfare of the federated learning system but also encourages all devices to contribute all their high-quality data to model training.

[0258] In steps S1-S5 of this invention, the interaction between the server and devices in the federated learning system is first modeled as a multi-player repeated game, where the strategies of the participants are continuous. To address the incentive problem in federated learning, this invention utilizes the zero-determinant strategy, which exhibits strong control in repeated prisoner's dilemma games. However, the traditional zero-determinant strategy is only applicable to two-player discrete strategy scenarios, which is not suitable for the application scenario in federated learning. Therefore, this invention extends the zero-determinant strategy to multi-player continuous strategy scenarios, enabling the server to apply a continuous zero-determinant (CZD) strategy to control the linear relationship between its expected payoff and the device's performance. The weighted sum of the expected utilities of all participants in the federated learning is defined as social welfare. Maximizing the social welfare of the federated learning system is the optimization objective, with constraints such as cooperation probability and the linear relationship controlled by the CZD strategy. This optimization problem is solved to obtain the optimal actions of the server and devices. Based on this, a federated learning incentive algorithm based on the CZD strategy is designed. The construction method and characteristic properties of the CZD strategy-based federated learning incentive mechanism are discussed and explained from the perspective of mathematical modeling and theoretical analysis. Furthermore, in the experimental evaluation section, this invention was compared with other methods in existing work and classical game theory methods, and the effectiveness and advantages of the federated learning incentive method based on the CZD strategy were verified from the perspectives of federated learning operation efficiency, model accuracy, social welfare, and fairness (relative benefits of servers and devices).

[0259] By introducing the designed CZD-based federated learning incentive mechanism into the existing federated learning system, considering the linear relationship between rewards and model training contributions, mobile devices with high-precision data are more willing to participate in model training tasks. Devices can dynamically adjust the quality and quantity of data they contribute based on the reward level set by the server. By attracting as many devices as possible to participate in federated learning, the server can use the CZD-based incentive algorithm to control the relationship between the set reward level and device contributions, achieving a fairer and more reasonable reward distribution and faster, more efficient model training. Data verification in the graphs shows that the introduction of the designed incentive method improves both federated learning efficiency and model training accuracy. This verifies the effectiveness of the technical solution of this invention and demonstrates its creativity and technical value.

[0260] Example: Experimental Evaluation

[0261] To demonstrate the performance of the proposed activation algorithm, the MNIST dataset, containing 10,000 test examples and 60,000 training examples, was used for evaluation. A scenario involving one server and 10 devices participating in federated learning was considered, and a multilayer perceptron (MLP) deep learning model was used to illustrate the performance comparison, where the MLP optimizer is stochastic gradient descent. The code for the activation mechanism algorithm was edited using PyTorch 1.2.0 in Python 3.7.3. The machine used for the simulation experiments was an 11th Gen Intel(R) Core(TM) i5-11500@2.70GHz desktop computer. To improve the confidence of the numerical simulation and eliminate errors, the corresponding data in the figures were obtained by averaging the results from 20 repeated experiments. The server employs a CZD strategy, with scalars independent of federated learning set as c1 = 0.361, c2 = 4.348, c3 = 0.001, c4 = 0.993, c5 = 0.178, and c6 = 1.743, respectively. Other parameters in the simulation experiment are shown in Table 1. The present invention will now be evaluated experimentally from the perspectives of the effectiveness of the CZD strategy, the impact of system parameters, and the fairness of the CZD strategy.

[0262] Table 1 Parameter Settings

[0263]

[0264]

[0265] Table 2 Feature Comparison

[0266] Strategy Approach Strategy Scenarios CE[1] Two-person discrete strategy scenario, multi-person discrete strategy scenario MMZD[2] Multi-person discrete strategy scenario Our CZD Two-player continuous strategy scenarios, multi-player continuous strategy scenarios

[0267] A. The effectiveness of the CZD strategy

[0268] To demonstrate the effectiveness of the CZD strategy, Table 2 first compares the characteristics of the CE and MMZD strategies with the CZD strategy. It can be noted that CE and MMZD are both based on discrete policy scenarios, while the CZD strategy focuses on continuous policy scenarios with two or more participants.

[0269] References for simulation experiment comparison:

[0270] [1] Q.Hu, S.Wang, Z.Xiong, and X.Cheng, "Nothing wasted: Full contribution enforcement in federated edge learning," IEEE Transactions on MobileComputing, doi:10.1109 / TMC.2021.3123195.

[0271] [2] J.Chen, Q.Hu, and H.Jiang, "Social welfare maximization in crosssilofederated learning," in Proc.ICASSP 2022-2022IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, Singapore: IEEE, 2022, pp.4258-4262.

[0272] Since MMZD and this invention share similar scenarios, this invention uses MMZD as a representative, and a direct performance comparison is presented in Table 3. The data in Table 4 represent model accuracy, training loss, and training time, respectively. In contrast, the method of this invention only requires 20 iterations to improve model accuracy to approximately 90%. After 10 iterations, the difference in model accuracy between the two methods is 8.29%. The training loss is reduced by 0.2059 within 150 iterations, saving approximately 400 seconds at 200 iterations. This is because when the policies of federated learning participants are extended from discrete to continuous, devices can not only choose whether to participate but also specifically control their contributions based on conditions such as income and cost. Therefore, compared to policies in discrete scenarios, CZD policies containing more complete information have greater practical application value in the model training process of federated learning.

[0273] Table 3 Performance Evaluation

[0274] Epochs MMZD Our approach 5 (65.33%,-0.9348,109s) (73.94%,-0.8424,94s) 10 (75.83%,-0.9027,205s) (84.12%,-0.8395,179s) 20 (85.00%,-0.8982,418s) (89.60%,-0.8391,355s) 50 (90.37%,-0.8702,938s) (92.76%,-0.8024,819s) 100 (86.67%,-0.8654,2222s) (93.57%,-0.7719,1751s) 150 (91.92%,-0.8598,3187s) (95.02%,-0.6539,2854s) 200 (95.94%,-0.6856,3993s) (98.33%,-0.5030,3604s)

[0275] like Figures 3-4 As shown, this embodiment of the invention evaluates the social welfare evolution of a federated learning system under different game scenarios in 200 rounds of global iteration of federated learning, where the server and devices adopt different strategies. The social welfare value is the weighted sum of the expected utilities of the server and all devices.

[0276] To confirm that the server, after adopting the CZD strategy, can encourage evolving devices to contribute all their high-precision data to federated learning, this invention compares it with classic game strategies such as TFT (Tit-for-Tat) and WSLS (Win-Save-Lose). TFT means that the player who chooses to cooperate as the initial action replicates the opponent's action from the previous round starting from the second round. WSLS means that the player cooperates in the first round, and then chooses the action for each subsequent step based on the cooperation / betrayal situation of both sides in the previous round. In federated learning games, the strategies of devices and the server can be divided into two cases: cooperation (C) and betrayal (D). For the k-th device, C corresponds to x k ∈[λ k ,h k ],in Similarly, for the server, C corresponds to y∈[μ,h], where Otherwise, their actions all correspond to betrayal (D).

[0277] Figures 3-4 This describes the change in social welfare with the number of iterations for N=1 and N=10, where the server uses CZD, TFT, and WSLS strategies, and the device uses TFT and WSLS strategies. The CZD vs TFT figures correspond to the server using CZD and the device using TFT strategies. In a continuous policy scenario with two participants, [the following is a continuation of the previous sentence, but the context is unclear]. Figure 3 The results show that when the device uses a TFT strategy and the server uses either WSLS or TFT, social welfare is unstable. If the device uses WSLS and the server uses either WSLS or TFT, social welfare will stabilize at a low level. When the server applies a CZD strategy, social welfare converges to a high and stable value regardless of the device's strategy. Figure 4 In this context, the number of devices participating in federated learning increased from 1 to 10. Figure 3 In comparison, social welfare converges to a higher stable value in all possible scenarios. Specifically, when the server employs the CZD strategy, social welfare increases from 10 to 16 as N changes. Although social welfare increases with N, it only reaches its optimal value when the server uses the CZD strategy.

[0278] B. Influence of system parameters

[0279] In CZD strategies, system parameters play a crucial role in modulating the utility relationship between devices and servers. To reveal the impact of system parameters on federated learning games, Figures 5-6The results show how social welfare changes with the number of iterations when the system parameters take different values. It can be observed that as the value of χ increases from 1 to 4, the fluctuation in social welfare also increases, but with the increase in the number of global iterations, social welfare eventually tends to stabilize before 60 rounds. Different values ​​of β have little impact on the fluctuation of social welfare, but because the values ​​of devices in social welfare vary, more global iterations are needed to control the overall social welfare and bring the entire federated learning system to a stable state. It can be seen that changes in the value of χ lead to fluctuations in social welfare, while changes in β affect the convergence speed of social welfare. However, the social welfare of federated learning will eventually stabilize at a consistent level. This is because the CZD strategy can regulate the behavior of devices, making them contribute all the high-precision data, thereby ensuring that the expected utility of all participants tends to stabilize, and then allowing social welfare to converge to the same level. Figures 3-6 The effectiveness of the federated learning incentive system based on the continuous zero determinant strategy was evaluated from the perspective of social welfare (weighted sum of expected benefits).

[0280] Fairness of C.CZD strategy

[0281] This invention implements a federated learning incentive system based on a continuous zero determinant (CZD) strategy. This system uses a CZD strategy to ensure a linear relationship between the expected returns of the server and the expected returns of the devices, encouraging all devices participating in the federated learning to contribute all their high-precision data to model training in order to maximize their individual returns. This leads to the convergence of the social welfare of the federated learning system to the optimal level. However, is this result fair to all participants in the federated learning system, namely the server and the devices? Figures 7-10 As shown, this embodiment of the invention introduces the concept of relative utility to evaluate the fairness of the CZD strategy, wherein the value of relative utility is equal to the quotient obtained by dividing its expected utility by the utility under ALLC (all-cooperative) conditions.

[0282] To demonstrate the fairness of the CZD strategy, Figures 7-8 This describes the stable values ​​of the relative utility of federated learning participants when the server employs different strategies and the device uses TFT or WSLS. The value of relative utility is equal to the quotient obtained by dividing the actual utility by the expected utility in ALLC (All-Cooperative Learning). Figure 7 and Figure 8 The data shows that only the CZD strategy can bring the relative utility of devices and servers closer to 1. The TFT and WSLS strategies result in relative utility below 1, and even reduce server revenue. Figures 9-10This shows the changes in the relative utility of federated learning participants over 200 global iterations when all devices use the WSLS strategy and the servers use different strategies. Although the relevant utility levels for the different strategies are roughly equal and fluctuate in the initial iterations, the relative utility of both the servers and devices tends to stabilize as the number of global iterations increases, and their stable state is similar to... Figures 7-8 The relative utility values ​​of federated learning participants shown in the figure are consistent, and the TFT and WSLS strategies tend to stabilize the relative utility at values ​​below 1. Only when the server adopts the CZD strategy does the relative utility of federated learning participants converge to 1.

[0283] In summary, the federated learning incentive system based on the continuous zero determinant strategy constructed in this invention incentivizes mobile devices participating in the federated learning model training task to contribute all of their high-precision data. From the perspective of the social welfare of the entire federated learning system, the benefits are relatively fair for both the server and the devices. This demonstrates the effectiveness of the federated learning incentive system based on the continuous zero determinant strategy in terms of fairness.

[0284] It should be noted that embodiments of the present invention can be implemented using hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented using hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or using software executed by various types of processors, or using a combination of the above-described hardware circuitry and software, such as firmware. The above descriptions are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A federated learning incentive method, characterized in that, Includes the following steps: Step 1: Construct a federated learning system based on a continuous zero-coefficient strategy and define the functions of each participating role in the mobile device and server. Step 2: Define the strategies for participating devices and servers in federated learning, and quantify the utility function-related variables of devices and servers under different strategies; Step 3: Define and prove that there is a dilemma in federated learning games that will harm social welfare, where the optimal strategy for the device is to provide incomplete, high-precision data, and the optimal strategy for the server is to set a lower reward level. Step 4: Based on the zero-determinant strategy theory in a two-person continuous strategy scenario, ensure that the utility is maximized only when the device and server cooperate fully. Step 5: In a multi-player continuous policy game scenario involving N devices and a federated learning server, maximizing social welfare is taken as the optimal objective of the incentive mechanism, and a federated learning incentive mechanism algorithm is designed based on the optimal strategy found by solving the problem. In step three, the federated learning game is defined based on the interaction between the server and the device, including: The interaction between devices and servers is modeled as a multi-player synchronous repeated game and defined as a federated learning game. The dilemma in federated learning games, Part 1 The best strategy for each device is The best strategy for the server is Introduce a zero-determinant strategy; Step four, based on the zero-determinant policy theory in continuous policy scenarios, ensures that utility is maximized only when the device and server cooperate fully. This includes: (1) Quantify the policy and utility functions of the devices and servers participating in federated learning, and calculate the expected utility; (2) Using a continuous zero-determinant strategy to control the expected utility of the device and the server to be linearly related; (3) The optimization objective is to maximize social welfare through federated learning, and the solution is obtained by using the linear relationship controlled by the zero determinant strategy. In step five, maximizing social welfare is taken as the optimal objective of the incentive mechanism. Based on the optimal strategy found by the solution, an incentive mechanism algorithm is designed to maximize the social welfare of the entire federated learning system. The following formula is used for solving: The social benefits of the federal learning system consist of a weighted sum of two parts: The expected utility of federated learning across all devices is: ; The expected utility of the federated learning server is: ; The social benefits of the federal learning system are: ; Utilize to satisfy Continuous zero determinant strategy: ; Then we get: ; The objective optimization problem is as follows: 。 2. The federated learning incentive method as described in claim 1, characterized in that, In step one, the server acts as the administrator, forming the federated learning system and participating in subsequent model aggregation and reward allocation processes, specifically including: (1) Issue the federated learning model training task and initialize the global model; (2) Receive requests sent by the receiving device and organize relevant devices to participate in pre-training for accuracy evaluation; (3) During pre-training, the upper and lower bounds of the accuracy are evaluated based on the sample data provided by the device, a reward is given to the device that meets the pre-training requirements, and the initial global model is sent to it; (4) After the device has performed multiple rounds of local computation to converge the model parameters and gradient data to a certain set threshold, the local model sent by the receiving device is received. (5) Aggregate all received local models in this stage and update the global model; (6) Calculate the equipment contribution based on the relevant data from the local model and pay the remuneration accordingly; (7) If the global model converges to the expected value, stop the iteration; otherwise, return to step (1). In a hierarchical federated learning system based on a continuous zero-coefficient policy, the participants include servers and mobile devices; the servers are responsible for distributing machine learning tasks to the participating devices to collaboratively train a shared model; the devices use the relevant data they possess to train the federated learning model. The entire model training process of federated learning consists of several rounds of global iterations, each global iteration covering multiple local iterations. Each device trains each training model in the local iterations; in the global iteration, the server sends the training model to the device and updates the model after receiving multiple trained models.

3. The federated learning incentive method as described in claim 1, characterized in that, In step two, the mobile devices, acting as managed entities, constitute the federated learning system and participate in subsequent local model training. Specifically, this includes: (1) If the device has high-precision local data related to the broadcast federated learning task, it sends a participation request to the server; (2) Participate in pre-training; qualified devices will receive the initial global model sent by the server. (3) Iterative training and updating of local models using local data; (4) If the relevant data of the local model after training meets the threshold set by the server, then apply for uploading; (5) Upload the updated local model and receive the reward paid by the server; This paper models the interaction between devices and servers participating in federated learning as a multi-player synchronous game in a continuous policy scenario. It provides policy definitions and quantifies the utility function variables related to devices and servers under different policies, specifically including: 1) Define the strategy for participating devices and servers in federated learning; 2) Set constraints for the hybrid strategies of federated learning participants; 3) Quantify the utility equations of federated learning participants.

4. A federated learning incentive system applying the federated learning incentive method as described in any one of claims 1 to 3, characterized in that, The federal learning incentive system includes: The system building module is used to build a federated learning system based on a continuous zero determinant policy, quantify the policies of devices and servers, and calculate the utility function and expected utility. The linear relationship control module is used to control the expected returns of the devices and servers participating in federated learning to satisfy a linear relationship using the zero-determinant policy in continuous policy scenarios. The utility maximization module is designed to maximize social welfare through federated learning, and ensures that utility is maximized only when devices and servers cooperate fully. The incentive mechanism involves modules used to solve optimization problems and to design the federated learning incentive mechanism algorithm based on the found optimal strategy, thereby maximizing the social welfare of the entire federated learning system.

5. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the federated learning incentive method as described in any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the federated learning incentive method as described in any one of claims 1 to 3.

7. An information data processing terminal, characterized in that, The information data processing terminal includes the federated learning incentive system as described in claim 4.