Federal learning dynamic excitation method and device fusing data freshness
By introducing a data freshness index and a multi-agent reinforcement learning model into federated learning, and dynamically adjusting the incentive strategy, the problem of neglecting data quality in existing technologies is solved, thereby improving model accuracy and overall returns.
Patent Information
- Application Number
- CN202511066538.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-11
AI Technical Summary
Existing federated learning incentive mechanisms primarily focus on the quantity of data when evaluating participants' contributions, neglecting factors such as data freshness or quality. This results in an incomplete evaluation of contributions, limiting overall benefits and improvements in model performance.
We introduce a data freshness index as an incentive criterion, dynamically adjust the incentive strategy through a multi-agent reinforcement learning model, use a server-side comment model to evaluate the data freshness and contribution of participants, design a reward function that combines data freshness and model accuracy improvement, and iteratively optimize the agent strategy.
It can more accurately distinguish between high-quality and low-quality contributions, prevent duplicate uploads of the same data from obtaining undue high rewards, improve the overall performance of federated learning and the enthusiasm of participants, and is particularly suitable for application scenarios where the data distribution is dynamically changing and the quality varies significantly.
Smart Images

Figure CN120930829A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically, to a method and apparatus for dynamically incentivizing federated learning that incorporates data freshness. Background Technology
[0002] Federated learning is a distributed machine learning technique that allows multiple participants to collaboratively train a model without sharing raw data, offering significant advantages in protecting user privacy and data security. A typical federated learning process includes steps such as global model initialization, local model training, parameter uploading and aggregation, and global model updating.
[0003] Existing federated learning incentive mechanisms are mainly divided into two categories: fixed-rule-based schemes and learning-based dynamic schemes. Fixed-rule-based incentive mechanisms rely on pre-defined theoretical models, such as Stackelberg games, auction theory, and contract theory. To address the need for incentive mechanisms to adapt to dynamic changes, related technologies have begun to introduce reinforcement learning methods to design incentive strategies for federated learning. Multi-Agent Reinforcement Learning (MARL) incentive mechanisms no longer pre-assume fixed game rules, but instead allow participating agents to gradually approach the optimal contribution strategy through continuous interaction with the environment. However, existing reinforcement learning-based federated incentive mechanisms primarily focus on the quantity of data when evaluating participant contributions, neglecting the freshness or quality of the data. This results in an incomplete evaluation of contributions, limiting the overall gains and improving model performance. Summary of the Invention
[0004] In view of the above problems, this disclosure provides a dynamic incentive method and apparatus for federated learning that integrates data freshness, aiming to solve the technical problem that the incentive mechanism in the prior art does not take into account data quality, resulting in limited overall benefits.
[0005] One aspect of this disclosure provides a dynamic incentive method for federated learning that incorporates data freshness, comprising: obtaining the data freshness of each participant in the federated learning, wherein data freshness characterizes the degree of difference between the training data of the participant in the current round and the historical training data; inputting the data freshness into a multi-agent reinforcement learning model, wherein each participant represents an agent, and performing the following operations: independently deciding on the batch size of the training data in the current round based on the local state information of each agent, using a preset actor model, and outputting actions that characterize the amount of training data or the proportion of resource input that each agent is willing to provide in the current round of training; evaluating the value of the actions of all agents based on the global state information of all agents, using a preset commenting model, and obtaining a global reward; calculating the local reward of each agent based on the global reward; updating the policy parameters of each agent using the local reward, and iteratively training until the multi-agent reinforcement learning model converges, wherein the policy parameters characterize the network weights of the actor model.
[0006] According to embodiments of this disclosure, obtaining the data freshness of each participant in federated learning includes: obtaining the current training data features of each participant in federated learning, and the historical training data features used by each participant in previous rounds; in response to the current training data features containing new features or the feature distribution changing compared to historical training data features, the data freshness of the current training data features is high, wherein high data freshness indicates that the contribution to model improvement is higher than a contribution threshold; in response to the current training data features not containing new features or the feature distribution not changing compared to historical training data features, the data freshness of the current training data features is low, wherein low data freshness indicates that the contribution to model improvement is lower than a contribution threshold.
[0007] According to embodiments of this disclosure, based on the local state information of each agent, the batch size of the data for this training round is independently decided using a preset actor model, and an action that can represent the amount of training data or the proportion of resource input that each agent is willing to provide in this training round is output. This includes: defining the state information of the agent, wherein the state information includes the historical data contribution, resource consumption, and global model accuracy of the participants; and outputting an action based on the local state information of each agent using the actor model corresponding to each agent, wherein the action represents the amount of training data or the proportion of resource input that each agent is willing to provide in this training round.
[0008] According to embodiments of this disclosure, the global reward is obtained by evaluating the actions of all agents based on the global state information of all agents using a preset comment model. This includes: acquiring the actions of all agents, the corresponding data freshness, and the resulting global model performance change characteristics; and calculating the global reward using the comment model based on the actions, data freshness, and global model performance change characteristics. The global reward represents the accuracy improvement parameters of the multi-agent reinforcement learning model.
[0009] According to embodiments of this disclosure, the calculation method for the global reward includes:
[0010]
[0011] Where, r i This represents the global reward, where i represents the participant, and d represents the global reward. i D represents the amount of data provided by each participant, and F represents the percentage of total data provided by all participants. i This represents the data freshness index of the participants, with α and β representing balance coefficients.
[0012] According to embodiments of this disclosure, updating the policy parameters of each agent using local rewards and iteratively training until the multi-agent reinforcement learning model converges includes: aggregating and calculating the model update parameters of each agent using a federated averaging algorithm based on the model update parameters obtained after each agent completes local training based on actions, to obtain an updated global model; sending the updated global model parameters to each agent based on the performance of the updated global model to provide an initial model for the next round of federated training; repeating the local incentive allocation process to iteratively optimize the global model and the reward policies of each agent until the global model converges or reaches a preset stopping condition, to obtain the final optimized multi-agent reinforcement learning model.
[0013] According to embodiments of this disclosure, updating the policy parameters of each agent using local rewards and iteratively training until the multi-agent reinforcement learning model converges further includes: allocating incentives corresponding to the benefits of each agent based on the local rewards calculated for each agent by the comment model; wherein the incentives and the updated global model parameters are synchronously sent to each agent, and the incentives include virtual points, tokens, or actual rewards.
[0014] Another aspect of this disclosure provides a dynamic incentive device for federated learning that incorporates data freshness, comprising: an acquisition module for acquiring the data freshness of each participant in the federated learning, wherein data freshness characterizes the degree of difference between the training data of the current round and the historical training data of the participant; inputting the data freshness into a multi-agent reinforcement learning model, wherein each participant represents an agent, and executing the following modules: a decision module for independently deciding the batch size of the training data in this round based on the local state information of each agent and using a preset actor model, and outputting actions that characterize the amount of training data or the proportion of resource input that each agent is willing to provide in this round of training; an evaluation module for evaluating the value of the actions of all agents based on the global state information of all agents and using a preset commenting model to obtain a global reward; a calculation module for calculating the local reward of each agent based on the global reward; and an iteration module for updating the policy parameters of each agent using the local reward, iteratively training until the multi-agent reinforcement learning model converges, wherein the policy parameters characterize the network weights of the actor model.
[0015] Another aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method described above.
[0016] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the methods described above.
[0017] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, are used to implement the methods described above.
[0018] Compared with the prior art, the federated learning dynamic incentive method and apparatus that integrates data freshness provided in the embodiments of this disclosure have at least the following beneficial effects:
[0019] (1) The federated learning dynamic incentive method and apparatus that integrates data freshness provided in the embodiments of this disclosure introduces the data freshness consideration index as the incentive basis in the process of dynamic adjustment of incentive strategy. This avoids the defect of traditional mechanism that only focuses on data quantity and ignores data quality. It can more accurately distinguish between high-quality contributions and low-quality contributions, prevent participants from obtaining undue high returns by repeatedly uploading the same data, encourage each participant to continuously provide new data, thereby improving the overall performance of federated learning and the enthusiasm of each participant.
[0020] (2) The federated learning dynamic incentive method and apparatus that integrates data freshness provided in this disclosure adopts a centralized training and decentralized execution (CTDE) architecture. It uses a server-side reviewer network (i.e., a review model) to evaluate the value of the joint actions of each participant and dynamically optimizes the agent's strategy by combining data volume and freshness information. After each round of training, the federated averaging (FedAvg) algorithm is used to aggregate the model update parameters, and incentives are allocated based on comprehensive rewards, forming a closed loop of model update and incentive distribution. This process iterates until the global model converges or reaches a preset indicator. This method can significantly improve the overall benefit and model accuracy of federated learning, and is particularly suitable for application scenarios where data distribution changes dynamically and data quality varies significantly. Attached Figure Description
[0021] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0022] Figure 1 A flowchart illustrating a federated learning dynamic incentive method for fusing data freshness according to an embodiment of the present disclosure is shown schematically.
[0023] Figure 2 A schematic diagram illustrating the principle of a multi-agent reinforcement learning architecture with centralized training and distributed execution according to an embodiment of the present disclosure is shown.
[0024] Figure 3 A schematic diagram illustrates the principle of a federated learning dynamic incentive method that integrates data freshness according to an embodiment of the present disclosure;
[0025] Figure 4 This schematically illustrates a structural block diagram of a federated learning dynamic incentive device that integrates data freshness according to an embodiment of the present disclosure;
[0026] Figure 5 The diagram schematically illustrates a structural block diagram of an electronic device suitable for implementing a federated learning dynamic incentive method that incorporates data freshness, according to an embodiment of the present disclosure. Detailed Implementation
[0027] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0028] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0030] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0031] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0032] Federated learning is a distributed machine learning technique that allows multiple participants to collaboratively train a model without sharing raw data, offering significant advantages in protecting user privacy and data security. A typical federated learning process includes global model initialization, local model training, parameter uploading and aggregation, and global model updates. In practice, each participant needs to provide local data, consume computing resources, and bear communication costs to participate in federated learning. When the direct benefits of participating in training are not obvious, some data owners may be hesitant to actively participate in model training. This conservative behavior can affect the overall performance of the federated learning model; therefore, effectively incentivizing participants to actively contribute data and computing power is a significant challenge facing federated learning.
[0033] Existing federal learning incentive mechanisms are mainly divided into two categories: schemes based on fixed rules and dynamic schemes based on learning.
[0034] First, incentive mechanisms based on fixed rules rely on pre-defined theoretical models, such as Stackelberg games, auction theory, and contract theory. In these methods, the system allocates incentives through predefined rules. For example, in a Stackelberg game, the global model owner, as the leader, sets the payoff function, and participating parties, as followers, adjust their data contributions according to this function. Auction theory models the incentive problem as a bidding process, allocating resources and rewards through the interaction between the auctioneer and bidders. Contract theory encourages high-quality data providers to contribute more data by signing differentiated incentive contracts with different types of participants. These methods have the advantages of clear models and high stability, but their incentive rules are fixed at the design stage, lacking the flexibility to dynamically adjust according to changes in participant behavior and the external environment. When the data contribution levels of participants in a federated learning environment fluctuate over time, the fixed-rule mechanism struggles to adapt in a timely manner, potentially leading to decreased incentive efficiency.
[0035] Second, to address the issue of incentive mechanisms needing to adapt to dynamic changes, related technologies have begun to incorporate reinforcement learning methods to design incentive strategies for federated learning. Multi-Agent Reinforcement Learning (MARL) incentive mechanisms no longer presuppose fixed game rules, but instead allow participating agents to gradually approach the optimal contribution strategy through continuous interaction with the environment. In this framework, each participant is modeled as an agent that can dynamically adjust the amount of data or resources it is willing to contribute in each training round based on its own state (e.g., the amount of data already contributed, computational and communication overhead, etc.) and feedback from the global model, in order to maximize its long-term gains and global model performance. The MARL method can continuously update the incentive strategy based on changes in participant performance during federated learning, adapting well to dynamic environments. However, existing reinforcement learning-based federated incentive mechanisms primarily focus on the quantity of data when evaluating participant contributions, neglecting data freshness or quality factors, resulting in an incomplete evaluation of contributions and limiting the improvement of overall gains and model performance. For example, if a participant submits relatively outdated data in each training round, almost identical to previously submitted data, existing methods, which only award the same reward based on the amount of data, may continue to encourage that participant to maintain the status quo. In this case, repetitive data with low freshness will not effectively improve the overall model; instead, it may negatively impact model accuracy due to a lack of diversity, resulting in wasted resources and unfair incentive allocation.
[0036] Therefore, it is necessary to provide an improved incentive mechanism for federated learning that incorporates data quality indicators such as data freshness into the evaluation system for the contributions of participating parties.
[0037] Based on this, the present disclosure provides a federated learning dynamic incentive method and apparatus that integrates data freshness, aiming to solve the technical problem of limited overall benefits caused by insufficient consideration of data quality in the incentive mechanism in the prior art.
[0038] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0039] Figure 1 A flowchart illustrating a federated learning dynamic incentive method that incorporates data freshness according to an embodiment of this disclosure is shown.
[0040] like Figure 1 As shown, the federated learning dynamic incentive method that integrates data freshness in this embodiment may include operations S1 to S5, and this method can be applied to a server.
[0041] In operation S1, the data freshness of each participant in the federated learning is obtained, where data freshness characterizes the degree of difference between the training data of the current round and the historical training data of the participant.
[0042] Input the freshness of this data into the multi-agent reinforcement learning model and perform the following operations S2~S4. Each participant represents one agent.
[0043] In operation S2, based on the local state information of each agent, the pre-set actor model is used to make independent decisions on the batch size of the training data in this round, and output actions that can represent the amount of training data or the proportion of resources that each agent is willing to provide in this round of training.
[0044] In operation S3, based on the global state information of all agents, the value of the actions of all agents is evaluated using a preset comment model to obtain a global reward.
[0045] In operation S4, the local reward for each agent is calculated based on the global reward.
[0046] In operation S5, the policy parameters of each agent are updated using local rewards, and training is iterated until the multi-agent reinforcement learning model converges. Here, the policy parameters represent the network weights of the agent model.
[0047] The federated learning dynamic incentive method that integrates data freshness provided in this disclosure introduces a data freshness metric as an incentive basis during the dynamic adjustment of the incentive strategy. This avoids the shortcomings of traditional mechanisms that only focus on data quantity while ignoring data quality. It can more accurately distinguish between high-quality and low-quality contributions, prevent participants from obtaining undue high returns by repeatedly uploading the same data, encourage participants to continuously provide new data, thereby improving the overall performance of federated learning and the enthusiasm of each participant.
[0048] According to embodiments of this disclosure, operation S1, which obtains the data freshness of each participant in federated learning, may specifically include:
[0049] Obtain the current training data features of each participant in the federated learning process, as well as the historical training data features used by each participant in previous rounds.
[0050] If the current training data features contain new features or the feature distribution changes compared to the historical training data features, then the current training data features have high data freshness. High data freshness indicates that the contribution to model improvement is higher than the contribution threshold.
[0051] If the current training data features do not contain new features or the feature distribution has not changed compared with the historical training data features, then the data freshness of the current training data features is low. Low data freshness indicates that the contribution to model improvement is lower than the contribution threshold.
[0052] In this embodiment, the server obtains the current training data features of each participant and compares them with the historical training data features used by that participant in previous rounds to calculate a "data freshness" value that represents the novelty of the data. This value can be obtained by quantifying the difference between the participant's current training data and its historical data.
[0053] Specifically, the server stores data indexes or feature representations used by participants in their past training. When new training data arrives, a difference index is calculated between the new dataset and the historical datasets. If the training data provided by a participant in this round contains a large number of new samples or a significantly different distribution compared to its previous data, its data freshness index will increase; conversely, if the currently provided data highly overlaps with historical data (e.g., repeating previous samples), the freshness index will decrease. Through this index, the system can gain a more comprehensive understanding of the value of the participant's data. Highly fresh data usually means a new contribution to improving the model (i.e., a contribution to model improvement exceeding the contribution threshold), while purely repetitive data has limited or even zero marginal contribution to the model (i.e., a contribution to model improvement below the contribution threshold).
[0054] This method ensures that the data quality of participants is taken into account by introducing data freshness evaluation into the incentive mechanism, and prevents some participants from obtaining unduly high incentives simply because the data volume is large but the content is outdated.
[0055] According to embodiments of this disclosure, operation S2, based on the local state information of each agent, independently decides the batch size of the training data for this round using a preset actor model, and outputs an action that can characterize the amount of training data or the proportion of resource input that each agent is willing to provide in this round of training. Specifically, it may include:
[0056] Define the state information of the agent, which includes the historical data contribution, resource consumption, and global model accuracy of the participants;
[0057] Based on the local state information of each agent, an action is output using the corresponding actor model of each agent. The action represents the amount of training data or the proportion of resources that each agent is willing to provide in this round of training.
[0058] In this embodiment, the next step is the reinforcement learning-based incentive decision-making phase, where the system abstracts each participant as an agent for reinforcement learning modeling. This incentive decision-making phase is implemented based on a multi-agent reinforcement learning architecture of centralized training and decentralized execution (CTDE), which is detailed in [link to documentation]. Figure 2 As shown.
[0059] Figure 2 A schematic diagram illustrating a multi-agent reinforcement learning architecture with centralized training and distributed execution according to an embodiment of the present disclosure is shown.
[0060] like Figure 2 As shown, the multi-agent reinforcement learning architecture with centralized training and distributed execution in this embodiment first defines the state information of the agents. For example, the state information of the agents includes the historical data contribution of the participants, resource consumption (such as communication and computing overhead), and global model accuracy (such as accuracy or loss information).
[0061] Then, each agent outputs an action through the actor network (Actor), i.e., the actor model, based on its state, which determines the amount of training data or the proportion of resources it is willing to provide in this round.
[0062] For example, in a mobile device scenario, an action can represent the size of the local data batch used for this round of training, or the number of training epochs, etc. The actions of all participants will work together on the federated learning environment, and accordingly, each participant will perform several steps of model training locally (the training intensity is determined by the action), resulting in new model parameter updates.
[0063] According to embodiments of this disclosure, operation S3 evaluates the value of all agents' actions using a preset comment model based on the global state information of all agents to obtain a global reward, which may specifically include:
[0064] Acquire the actions of all agents, the corresponding data freshness, and the resulting global model performance changes;
[0065] Based on the characteristics of action, data freshness, and global model performance changes, the global reward is calculated using a comment model. The global reward represents the accuracy improvement parameter of the multi-agent reinforcement learning model.
[0066] In this embodiment, as Figure 2 As shown, the server-side setup includes a critic network (Critic), also known as a commenting model, used to evaluate the joint actions of the agents. During the centralized training phase, the critic network has access to global state information, including the action choices of all participants, their corresponding data freshness values, and the resulting changes in global model performance. Based on this, the critic network calculates a global reward and the local rewards for each agent.
[0067] Specifically, the critic network introduces a data freshness-related term into its reward function: for participant i, its reward $r_i$ depends not only on the amount of data it contributes and its contribution to improving the global model accuracy, but also on its data freshness. When participant i provides data that provides new information gain to the model (high freshness), it will receive a higher reward; conversely, if its data has high redundancy (low freshness), the reward will be reduced accordingly.
[0068] For example, assume the global reward is improved by the model accuracy parameters. Given this representation, the reward function (global reward) for participant i can be designed as follows:
[0069]
[0070] Where, r i This represents the global reward, where i represents the participant, and d represents the global reward. i D represents the amount of data provided by each participant, and F represents the percentage of total data provided by all participants. i This represents the data freshness index of the participants, with α and β representing balance coefficients.
[0071] This reward design allows for a more precise measurement of each participant's true contribution to model training. It should be noted that the reward function described above is merely an example; in practical applications, an appropriate function form can be chosen based on application requirements, as long as data quality (freshness) is included in the reward calculation.
[0072] According to embodiments of this disclosure, operation S5 updates the policy parameters of each agent using local rewards, iteratively training until the multi-agent reinforcement learning model converges, specifically including:
[0073] Based on the model update parameters obtained by each agent after completing local training based on actions, the federated averaging algorithm is used to aggregate and calculate the model update parameters of each agent to obtain the updated global model.
[0074] Based on the updated global model performance, the updated global model parameters are sent to each agent to provide an initial model for the next round of federated training.
[0075] The process of allocating local incentives is repeated, and the reward strategies of the global model and each agent are iteratively optimized until the global model converges or reaches the preset stopping condition, thus obtaining the final optimized multi-agent reinforcement learning model.
[0076] In this embodiment, as Figure 2 As shown, after each round of training, the critic network calculates the reward for each agent based on the joint actions of all agents and the resulting global model effect, and uses these rewards to update the policy parameters (actor network weights) of each agent through the policy gradient algorithm.
[0077] Because of its centralized training, the critic network can consider the collaborative relationships between agents during policy updates. For example, if reducing a participant's data contribution lowers the global model accuracy, the critic will give it a lower reward, thus penalizing this behavior during policy updates. Conversely, if a participant's data significantly improves the model's performance and is novel, the critic network will encourage it to maintain or increase its contribution with a high reward. This centralized evaluation avoids the "free-rider" phenomenon that can result from agents acting independently, where a single participant attempts to rely on others' contributions to improve the model while reducing its own efforts (i.e., the "lazy agent problem"). Through the global perspective guidance provided by the critic network during the training phase, each participating agent gradually learns to optimize its own gains without compromising the overall benefit, achieving a balance between maximizing global gains and ensuring reasonable individual gains.
[0078] It's worth noting that in actual operation (execution phase), each participating agent makes independent decisions using a pre-trained policy, without needing to know the actions or states of other participants. In other words, the training process is centralized (requiring a commentator network to aggregate information for policy training), while the execution process is decentralized, with each participant contributing only the amount of data they need to determine based on their own circumstances. This "centralized training, distributed execution" architecture ensures that while achieving collaborative optimization benefits, the system maintains the privacy and low communication overhead expected of a distributed system: participants do not need to disclose their local data or model parameters; only model updates and necessary metrics during training are sent to the server for evaluation by the commentator network. The commentator network and the actor network work together to achieve policy optimization.
[0079] After the incentive decision is completed, the global model aggregation and update phase begins. Each participant, based on the actions determined in the previous phase, performs the corresponding data training task locally (e.g., running several rounds of gradient descent locally), and then uploads the calculated model parameter updates (e.g., weight differences or gradients) to the server. The server receives all the model updates from the participants and integrates them using the standard model aggregation algorithm in federated learning.
[0080] For example, the FedAvg algorithm can be used to weight the weight updates uploaded by each participant according to their local data volume to obtain new global model parameters. For example, for weight parameters... The amount of data used by each participant in this round can be determined based on the amount of data they use. Weighted summation:
[0081]
[0082] If some participants choose to contribute very little data due to low rewards, their uploaded updates will have a relatively small impact on the global model, thus adaptively reducing the interference of low-quality data on the model. After aggregation, the server will provide the next-generation global model to each participant for initializing their model for the next round of local training.
[0083] According to embodiments of this disclosure, during this process, the server also distributes corresponding incentives to the participants based on the rewards previously calculated by the commentator network. For example, incentives can take the form of virtual points, tokens, or actual rewards, depending on the application scenario. The incentive distribution process can be notified to all participants simultaneously with model deployment. Through this closed loop of model update + incentive distribution, the federated learning system enters the next cycle.
[0084] The above process continues iteratively until the training termination condition is met. Typically, a predetermined global model performance threshold or maximum number of training epochs can be set. When the global model's accuracy reaches the required level or the overall return growth slows down, the server terminates the training process and publishes the final global model and each participant's accumulated incentive value to all participants. At this point, the method of this embodiment is complete.
[0085] In summary, the federated learning dynamic incentive method that integrates data freshness provided in this disclosure adopts a multi-agent reinforcement learning architecture with centralized training and distributed execution during the training phase. A commentator network on the server side acquires global state information and evaluates the joint actions of all agents to provide value guidance, promote collaboration among agents, and avoid the "lazy agent" problem caused by lack of coordination. Each participating agent independently makes decisions based on local observations through its own actor network. During the execution phase, each agent independently decides its data contribution according to the trained strategy. Finally, a reinforcement learning reward function is designed to incorporate the freshness of the data provided by each participant in each training round into the reward calculation. When the new data provided by a participant in the current round differs significantly from its historical data, the reward value of that participant increases; if a participant repeatedly provides data similar to historical data, its reward value decreases. This approach allows for a more accurate assessment of the data contribution quality of each participant, preventing low-quality or outdated data from receiving inappropriate incentives. Furthermore, based on the value assessment and reward function of the critic network, the policy parameters of each agent are updated after each training round, optimizing the data contribution strategies of each participant round by round to approximate their optimal contribution strategies and achieve dynamic optimization of incentive strategies during federated learning.
[0086] The federated learning dynamic incentive method based on data freshness provided in this disclosure employs a centralized training and distributed execution architecture. It uses a server-side reviewer network to evaluate the value of joint actions by each participant and dynamically optimizes the agent's strategy by combining data volume and freshness information. After each training round, the federated averaging (FedAvg) algorithm aggregates the model update parameters, and incentives are allocated based on comprehensive rewards, forming a closed loop of model update and incentive distribution. This iteration continues until the global model converges or reaches a preset metric. Therefore, it can significantly improve the overall benefit and model accuracy of federated learning, and is particularly suitable for application scenarios where data distribution is dynamically changing and data quality varies significantly.
[0087] It should be noted that in practical applications, the above steps can be adjusted or extended depending on the specific task and data characteristics. For example, the measurement of data freshness can employ a weighted function based on timestamp decay to better highlight the most recently acquired data, or an auxiliary discriminant model can be trained to assess whether the data makes a new contribution to the model. Furthermore, the state in reinforcement learning can include more contextual information, such as the contribution trends of participants over several rounds, or the model convergence speed metrics in the current round, allowing the agent's policy to consider a wider range of factors. Regardless of the specific implementation variations, all fall within the scope of this disclosure.
[0088] Figure 3 A schematic diagram illustrates the principle of a federated learning dynamic incentive method that incorporates data freshness according to an embodiment of the present disclosure.
[0089] like Figure 3 As shown in the embodiments of this disclosure, the principle of the federated learning dynamic incentive method that integrates data freshness is as follows:
[0090] Initialize the federated learning environment and set up the global model.
[0091] Calculate the data freshness of each participant. For example, obtain the local data timestamps of each participant and compare the differences between current and historical data.
[0092] Dynamic incentive decision-making based on multi-agent reinforcement learning. For example, (1) agent state definition and action selection; (2) centralized training and distributed execution CTDE; (3) reward function design incorporating data freshness; (4) policy parameter update.
[0093] Global model aggregation and incentive allocation. For example, the FedAvg algorithm is used to aggregate model updates; participant benefits are evaluated and incentives are allocated.
[0094] Determine if the convergence condition is met. If the convergence condition is met, output the final global model and end the training; if the convergence condition is not met, return to "Calculate the data freshness of the participants" and repeat the above operation until the convergence condition is met.
[0095] Figure 4 A schematic block diagram of a federated learning dynamic incentive device that integrates data freshness according to an embodiment of the present disclosure is shown.
[0096] like Figure 4 As shown, the federated learning dynamic incentive device 400 that integrates data freshness according to an embodiment of this disclosure includes: an acquisition module 410, a decision module 420, an evaluation module 430, a calculation module 440, and an iteration module 450.
[0097] The acquisition module 410 is used to acquire the data freshness of each participant in federated learning, where data freshness characterizes the degree of difference between the training data of the current round and the historical training data of the participant.
[0098] The data freshness is input into a multi-agent reinforcement learning model, where each participant represents an agent, and the following modules are executed:
[0099] The decision module 420 is used to make independent decisions on the batch size of the training data in this round based on the local state information of each agent and using the preset actor model, and outputs actions that can represent the amount of training data or the proportion of resources that each agent is willing to provide in this round of training.
[0100] The evaluation module 430 is used to evaluate the value of the actions of all agents based on the global state information of all agents and using a preset comment model to obtain a global reward.
[0101] The calculation module 440 is used to calculate the local reward for each agent based on the global reward.
[0102] The iteration module 450 is used to update the policy parameters of each agent using local rewards, and iterates the training until the multi-agent reinforcement learning model converges. Here, the policy parameters represent the network weights of the agent model.
[0103] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0104] For example, any multiple of the acquisition module 410, decision module 420, evaluation module 430, calculation module 440, and iteration module 450 can be combined into one module / unit / subunit, or any one of these modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits can be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of this disclosure, at least one of the acquisition module 410, decision module 420, evaluation module 430, calculation module 440, and iteration module 450 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the acquisition module 410, decision module 420, evaluation module 430, calculation module 440 and iteration module 450 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0105] It should be noted that the dynamic incentive device part of federated learning that integrates data freshness in the embodiments of this disclosure corresponds to the dynamic incentive method part of federated learning that integrates data freshness in the embodiments of this disclosure. For a detailed description of the dynamic incentive device part of federated learning that integrates data freshness, please refer to the dynamic incentive method part of federated learning that integrates data freshness, which will not be repeated here.
[0106] Figure 5 The diagram schematically illustrates a structural block diagram of an electronic device suitable for implementing a federated learning dynamic incentive method that incorporates data freshness, according to an embodiment of the present disclosure. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0107] like Figure 5As shown, an electronic device 500 according to an embodiment of this disclosure includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this disclosure.
[0108] RAM 503 stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0109] According to embodiments of this disclosure, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The electronic device 500 may also include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.
[0110] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0111] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0112] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0113] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than ROM 502 and RAM 503.
[0114] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of this disclosure.
[0115] When the computer program is executed by the processor 501, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0116] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 509, and / or installed from a removable medium 511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0117] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0119] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A dynamic incentive method for federated learning that integrates data freshness, characterized in that, The method includes: Obtain the data freshness of each participant in federated learning, wherein the data freshness characterizes the degree of difference between the training data of the participant in the current round and the historical training data; The data freshness is input into a multi-agent reinforcement learning model, where each participant represents an agent, and the following operations are performed: Based on the local state information of each agent, the system independently decides the batch size of the training data in this round using a pre-set actor model, and outputs actions that can represent the amount of training data or the proportion of resources that each agent is willing to provide in this round of training. Based on the global state information of all agents, the value of the actions of all agents is evaluated using a pre-defined comment model to obtain a global reward. Based on the global reward, the local reward for each agent is calculated; The policy parameters of each agent are updated using the local reward, and the training is iterated until the multi-agent reinforcement learning model converges, wherein the policy parameters represent the network weights of the agent model.
2. The method according to claim 1, characterized in that, The process of obtaining the data freshness of each participant in federated learning includes: Obtain the current training data features of each participant in the federated learning process, as well as the historical training data features used by each participant in previous rounds. In response to the fact that the current training data features contain new features or the feature distribution changes compared with the historical training data features, the current training data features have high data freshness, wherein the high data freshness indicates that the contribution to model improvement is higher than the contribution threshold. If the current training data features do not contain new features or the feature distribution has not changed compared with the historical training data features, then the data freshness of the current training data features is low, wherein the low data freshness indicates that the contribution to model improvement is lower than the contribution threshold.
3. The method according to claim 1, characterized in that, The step of independently deciding the batch size of training data for this round based on the local state information of each agent and using a preset actor model, and outputting actions that represent the amount of training data or the proportion of resources that each agent is willing to provide in this round of training, includes: Define the state information of the intelligent agent, wherein the state information includes the historical data contribution, resource consumption, and global model accuracy of the participating party; Based on the local state information of each agent, an action is output using the actor model corresponding to each agent. The action represents the amount of training data or the proportion of resources that each agent is willing to provide in this round of training.
4. The method according to claim 1, characterized in that, The step of evaluating the value of all agents' actions using a preset comment model based on the global state information of all agents to obtain a global reward includes: Acquire the actions of all agents, the corresponding data freshness, and the resulting global model performance changes; Based on the action, the data freshness, and the global model performance change characteristics, a global reward is calculated using a comment model, wherein the global reward represents the accuracy improvement parameter of the multi-agent reinforcement learning model.
5. The method according to claim 4, characterized in that, The calculation method for the global reward includes: Where, r i This represents the global reward, where i represents the participant, and d represents the global reward. i D represents the amount of data provided by each participant, and F represents the percentage of total data provided by all participants. i This represents the data freshness index of the participants, with α and β representing balance coefficients.
6. The method according to claim 3, characterized in that, The step of updating the policy parameters of each agent using the local reward and iteratively training until the multi-agent reinforcement learning model converges includes: Based on the model update parameters obtained by each agent after completing local training based on the actions, the federated averaging algorithm is used to aggregate and calculate the model update parameters of each agent to obtain the updated global model. Based on the updated global model performance, the updated global model parameters are sent to each agent to provide an initial model for the next round of federated training. Repeat the local incentive allocation process, iteratively optimize the global model and the reward strategies of each agent, until the global model converges or reaches the preset stopping condition, to obtain the final optimized multi-agent reinforcement learning model.
7. The method according to claim 6, characterized in that, The step of updating the policy parameters of each agent using the local reward and iteratively training until the multi-agent reinforcement learning model converges further includes: Based on the local reward calculated for each agent by the comment model, an incentive corresponding to the agent's own gain is allocated to each agent. The incentives and updated global model parameters are synchronously sent to each agent, and the incentives include virtual points, tokens, or actual rewards.
8. A dynamic incentive device for federated learning that integrates data freshness, characterized in that, The device includes: The acquisition module is used to acquire the data freshness of each participant in federated learning, wherein the data freshness characterizes the degree of difference between the training data of the current round and the historical training data of the participant; The data freshness is input into a multi-agent reinforcement learning model, where each participant represents an agent, and the following modules are executed: The decision module is used to make independent decisions on the batch size of training data for this round based on the local state information of each agent and using the preset actor model, and output actions that can represent the amount of training data or the proportion of resources that each agent is willing to provide in this round of training. The evaluation module is used to evaluate the value of the actions of all agents based on the global state information of all agents and to obtain a global reward using a preset comment model. The calculation module is used to calculate the local reward for each agent based on the global reward; An iterative module is used to update the policy parameters of each agent using the local reward, and iteratively train until the multi-agent reinforcement learning model converges, wherein the policy parameters represent the network weights of the agent model.
9. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having executable instructions stored thereon, characterized in that, When this instruction is executed by the processor, it causes the processor to perform the method according to any one of claims 1 to 7.