A personalized federated learning method for green industrial internet of things

By employing a personalized federated learning approach, combined with reinforcement learning and incentive models, the problem of balancing resources and incentives in the green industrial Internet of Things (IIoT) is solved, enabling sustainable operation and efficient collaborative learning of equipment under renewable energy.

CN120875458BActive Publication Date: 2026-02-10JINAN UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511358709.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-02-10
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

In the green industrial Internet of Things (IIoT), traditional federated learning methods fail to effectively balance computing resource utilization, energy consumption, and participant incentives, making it difficult for devices to achieve sustainable operation and efficient collaborative learning under unstable energy supply conditions.

Method used

A personalized federated learning method is constructed by initializing the system architecture, building an energy model and an incentive model, and using a reinforcement learning framework for dynamic resource allocation and reward scheduling to optimize the resource management and incentive mechanism for participants.

Benefits of technology

It improves the overall efficiency of federated learning and participant engagement, achieves sustainability of green IIoT devices and accuracy of the global model, and optimizes energy efficiency and incentive mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875458B_ABST
    Figure CN120875458B_ABST
Patent Text Reader

Abstract

The application discloses a personalized federated learning method for green industrial Internet of Things, relates to the technical field of federated learning in industrial Internet of Things, and comprises the following steps: initializing a federated learning system architecture, wherein the federated learning system architecture is composed of an edge server and a plurality of industrial Internet of Things nodes; constructing an energy model of participants, calculating energy buffer capacity and total energy supply; constructing an incentive model based on service quality and monetary rewards, and quantifying participant preference parameters; deriving personalized scheduling parameters through feature analysis and questionnaire feedback; and realizing dynamic resource allocation and reward budget scheduling by using a reinforcement learning framework. The application can guarantee the sustainability of renewable energy driven green IIoT devices of federated learning participants in real federated learning deployment, improve the accuracy of a global model of federated learning, realize personalized resource management and reward allocation, and optimize energy efficiency and an incentive mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of federated learning technology in the Industrial Internet of Things (IIoT), and particularly relates to a personalized federated learning method for green industrial IoT. Background Technology

[0002] With the rapid development of Industrial Internet of Things (IIoT) technology, the demand for intelligent manufacturing and device interconnection is increasing, placing higher demands on data processing and analysis capabilities. Against this backdrop, Federated Learning (FL), as a distributed machine learning method, can achieve collaborative learning of data while protecting privacy, and has become one of the key technologies in the IIoT field. However, traditional federated learning research often neglects the energy efficiency and incentive mechanism design of IIoT devices in practical deployments. Especially in green IIoT environments driven by renewable energy, balancing the utilization of computing resources, energy consumption, and the enthusiasm of participants becomes a significant challenge.

[0003] In the IIoT field, federated learning typically needs to be conducted on heterogeneous edge devices, which not only participate in federated learning activities but also perform their core functions. Given the limited computing resources and energy supply of these devices, optimizing energy consumption and incentivizing participant engagement while ensuring federated learning efficiency has become a pressing issue. Although existing research has proposed various federated learning frameworks and algorithms, most employ a one-size-fits-all approach to resource management and reward allocation, failing to fully consider the time-varying characteristics and individual differences of IIoT devices across different training rounds.

[0004] Furthermore, the energy supply for IIoT devices often relies on renewable energy sources such as solar and wind power, which are characterized by uncertainty and intermittency. Therefore, achieving sustainable operation of federated learning tasks under unstable energy supply conditions presents a new challenge in the IIoT field. To address this challenge, new federated learning incentive models and scheduling techniques need to be developed to adapt to the energy characteristics and computational requirements of IIoT devices.

[0005] Currently, most federated learning research focuses on improving the global accuracy of the model, neglecting the personalized needs and incentive preferences of participants. This leads to situations where, in practical deployments, participants may be unwilling to contribute their computing resources due to a lack of sufficient incentives, or the federated learning task may be inefficient due to unreasonable resource allocation. Therefore, there is an urgent need to propose a personalized federated learning method for green industrial IoT. Summary of the Invention

[0006] To address the aforementioned technical issues, this invention proposes a personalized federated learning method for green industrial IoT, which can dynamically adjust based on each participant's resource usage patterns and incentive preferences to improve the overall efficiency of federated learning and participant engagement.

[0007] To achieve the above objectives, this invention provides a personalized federated learning method for green industrial Internet of Things, comprising:

[0008] Initialize the federated learning system architecture, which consists of edge servers and multiple industrial IoT nodes;

[0009] Construct energy models for participants and calculate energy buffer capacity and total energy supply;

[0010] Construct an incentive model based on service quality and monetary rewards, and quantify participants' preference parameters;

[0011] Personalized scheduling parameters are derived through feature analysis and questionnaire feedback;

[0012] A reinforcement learning framework is used to achieve dynamic resource allocation and reward budget scheduling.

[0013] Optionally, initializing the federated learning system architecture includes:

[0014] Configure an edge server responsible for global model aggregation; define industrial IoT nodes with energy harvesting capabilities as participants; allocate a local dataset containing feature vectors and labels to each participant; divide the scheduling time range into non-uniform time intervals; and predefine the regular task computation requirement curves for each participant.

[0015] Optionally, constructing the energy models for participants includes:

[0016] The system integrates solar and vibration energy harvesting modules; calculates the harvesting power based on instantaneous irradiance and vibration velocity; determines the charging and discharging dynamics of the energy buffer module; calculates the energy consumption for training and routine tasks respectively; and establishes energy supply and demand balance constraints.

[0017] Optionally, constructing an incentive model based on service quality and monetary rewards includes:

[0018] Service quality indicators are calculated by differentiating resource allocation; potential variables are generated by combining service quality with monetary rewards; participation is calculated based on basic willingness to participate and marginal utility; preference weight parameters are dynamically adjusted; and a cumulative reward budget control mechanism is implemented.

[0019] Optional, the derivation of personalized scheduling parameters includes:

[0020] Analyze participants' historical participation counts; construct an objective function to maximize model accuracy; implement real-time monitoring of energy consumption; verify training deadlines; and ensure that computing resource allocation does not exceed limits.

[0021] Optional, personalized scheduling includes:

[0022] Collect routine task procedure feature data; select some participants to complete the payment willingness questionnaire; construct the questionnaire participant feature matrix; optimize the preference weights using Latin hypercube sampling; and calculate the optimal weight matrix using a linear solver.

[0023] Optionally, implementing reinforcement learning includes:

[0024] Define the system state, which includes remaining energy, budget, and time; generate resource allocation actions based on the policy; generate training time intervals according to an exponential distribution; implement a preference-based greedy reward allocation; and update the policy parameters and criticism parameters.

[0025] Optional, dynamic resource allocation includes:

[0026] Maintain the remaining energy state of participants; calculate the remaining reward budget and training time; store system state transition records; and return scheduling parameters containing participation decisions, time allocation, and resource allocation.

[0027] Technical effects of the invention: The invention discloses a personalized federated learning method for green industrial Internet of Things (IIoT), which can improve the accuracy of the global federated learning model while ensuring the sustainability of renewable energy-driven green IIoT devices of federated learning participants in real-world federated learning deployments, realize personalized resource management and reward allocation, and optimize energy efficiency and incentive mechanisms. Attached Figure Description

[0028] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0029] Figure 1 This is a flowchart illustrating the personalized federated learning method for green industrial Internet of Things according to an embodiment of the present invention.

[0030] Figure 2 These are the static and dynamic program features of routine tasks in embodiments of the present invention;

[0031] Figure 3 The FL accuracy achieved by the duration allocation method in this embodiment of the invention is shown in (a) for independent and identically distributed scenarios and (b) for non-independent and identically distributed scenarios.

[0032] Figure 4This is the confusion matrix of the willingness-to-pay prediction method in this embodiment of the invention;

[0033] Figure 5 The accuracy of the FL global model obtained by five methods in the embodiments of the present invention;

[0034] Figure 6 The training time required for the embodiments of the present invention to achieve the expected accuracy;

[0035] Figure 7 This describes the utilization of computational resources in the federated learning method of this invention. Detailed Implementation

[0036] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0037] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0038] like Figure 1 As shown, this embodiment provides a personalized federated learning method for green industrial IoT, including:

[0039] Initialize the federated learning system architecture, which consists of edge servers and multiple industrial IoT nodes;

[0040] Construct energy models for participants and calculate energy buffer capacity and total energy supply;

[0041] Construct an incentive model based on service quality and monetary rewards, and quantify participants' preference parameters;

[0042] Personalized scheduling parameters are derived through feature analysis and questionnaire feedback;

[0043] A reinforcement learning framework is used to achieve dynamic resource allocation and reward budget scheduling.

[0044] Specifically, the following steps are included:

[0045] Step 1: Initialize the system architecture of the federated learning system;

[0046] Step 2: Initialize the energy model of the federated learning system;

[0047] Step 3: Initialize the incentive model of the federated learning system;

[0048] Step 4: Personalized scheduling and preference parameter derivation for federated learning participants;

[0049] Step 5: Reinforcement learning scheduling and dynamic resource allocation in the federated learning system.

[0050] Furthermore, the system architecture initialization of the federated learning system in step 1 specifically includes:

[0051] Step 101: Define the edge server and Industrial IoT nodes (participants) ,in For the set of all possible participants, Responsible for global model aggregation and scheduling decisions. For the m-th participant in the participant set, Number the participants;

[0052] Step 102: For each participant Allocate local dataset ,in for 3D feature vectors For the corresponding tags, Indicates the size of the dataset. For sample index; It is a d-dimensional real space; Let be the d-dimensional feature vector of the ϑ-th sample; For participants Local dataset;

[0053] Step 103: Divide the scheduling time range for There are non-uniform time intervals, among which The deadline for the task. The scheduling start time;

[0054] Step 104: Define the first The time interval of the round is ,in This is the end time of the k-th time interval; This represents the start time of the k-th time interval; k is the index of the time interval.

[0055] Step 105: Predefine participants using trajectory analysis technology The computational requirements curve for routine tasks ;

[0056] Step 106: Initialize total computing capacity To meet the requirements of federal learning and training resource allocation ,in, Let m be the training resource requirement of the m-th participant within the k-th time interval. The total computing capacity for the m-th participant is... Resources are actually allocated for the routine tasks of the m-th participant within the k-th time interval; resources are actually allocated for routine tasks. ,in The actual allocation of resources for the routine tasks of the m-th participant within the k-th time interval. Let m be the resource requirement for the m-th participant's regular tasks within the k-th time interval. Let m be the training resource requirement of the m-th participant within the k-th time interval. The total computing capacity for the m-th participant;

[0057] Step 107: The system architecture initialization of the federated learning system is complete. Exit.

[0058] Furthermore, the initialization of the energy model of the federated learning system in step 2 specifically includes:

[0059] Step 201: For each participant Initialize the energy generation, buffering, and consumption modules;

[0060] Step 202: Use Calculation at an instant Collection power ,in, The scheduling start time, For solar panel conversion efficiency, Instantaneous solar irradiance, For the efficiency of vibration energy harvester, For mechanical vibration velocity;

[0061] Step 203: Use Calculate the scheduling time range Total cumulative energy supply ,in This represents the buffer energy at the current time t. The scheduling start time, The end time of the scheduling. The deadline for the task. Let be the collected power at time instant t;

[0062] Step 204: Use

[0063] Calculate buffer energy ,in The buffer energy for time t+Δt, The scheduling start time, For time intervals, For maximum buffer capacity, Let be the power collected at time τ; For integration variables; The energy requirement at time t;

[0064] Step 205: Use Calculate the first Federal learning and training energy consumption ,in It is a hardware-related constant. It is the number of CPU cycles required to process one data sample. Let m be the size of the dataset for the m-th participant. Let m be the training resource requirement of the m-th participant within the k-th time interval. It is transmission power. From arrive Communication rate, Let m be the number of model parameters for the m-th participant;

[0065] Step 206: Use Calculate the energy consumption of routine tasks ,in It is a hardware-dependent constant. The actual allocation of resources for the routine tasks of the m-th participant within the k-th time interval. Let k be the end time of the k-th time interval. This represents the start time of the k-th time interval;

[0066] Step 207: Use Calculate participants Total energy consumption over all time intervals ,in The total number of time intervals. Indexed by the current time interval. As a decision-making indicator, Let m be the energy consumption for federated learning training of the m-th participant during the k-th time interval. Let m be the energy consumption of the m-th participant during the k-th time interval for routine tasks;

[0067] Step 208: The energy model initialization of the federated learning system is complete. Exit.

[0068] Furthermore, the initialization of the incentive model for the federated learning system in step 3 specifically includes:

[0069] Step 301: Use Calculate service quality under different resource allocation conditions ,in A larger value indicates less performance loss for routine tasks; Let m be the resource requirement for the m-th participant's regular tasks within the k-th time interval. The actual allocation of resources for the routine tasks of the m-th participant within the k-th time interval;

[0070] Step 302: Use Combined with service quality and monetary rewards Through latent variables Quantify the preferences of participants, among which Specific preference weights for participants, This indicates a complete bias towards service quality; This indicates a complete bias towards rewards;

[0071] Step 303: Use Calculate each participant In the Willingness to participate in the round ,in Indicates basic willingness to participate. This represents the marginal utility coefficient. Rounded to select One of the integers;

[0072] Step 304: Use Calculate the overall willingness to participate of all participants. , The total number of participants. For participants' index, The total number of time intervals. Index of time interval, As a decision-making indicator, For each participant In the Willingness to participate in the round;

[0073] Step 305: Adjust preference weights based on participant feedback and experimental data. and preference parameters and ;

[0074] Step 306: The incentive model initialization of the federated learning system is complete. Exit.

[0075] Furthermore, step 4, which involves the personalized scheduling and preference parameter derivation of federated learning participants, specifically includes:

[0076] Step 401: Use Calculate participants Participation count counter , For the current round, The previous round indices are from 1 to k-1. As a decision-making indicator;

[0077] Step 402: Define the objective function To maximize the accuracy of the global model, where For the first Learning rate of each round of federated learning To participate in the decay factor , The total number of participants. For participants' index, The total number of time intervals. The index of the current time interval. As a decision-making indicator, For participants The participation count counter in round k, Indicates the size of the dataset;

[0078] Step 403: Monitor the energy consumption of each participant to ensure it does not exceed the supply. ,in For participants Total energy consumption For participants Total energy supply For all participants, An index for participants;

[0079] Step 404: Check that all participants have completed local training before the specified deadline. ,in The total number of time intervals. The index of the current time interval. Let k be the end time of the k-th time interval. Let k be the start time of the k-th time interval. The deadline for the task;

[0080] Step 405: Ensure that the total computing resource usage of each participant does not exceed the limit. ,in As a decision-making indicator, Let m be the training resource requirement of the m-th participant during the k-th time interval. Let m be the resource requirements for the m-th participant's routine tasks during the k-th time interval. Let m be the total computing resource capacity of the m-th participant. For all participants m and all time intervals k;

[0081] Step 406: Verify whether the overall WTP of all participants is higher than that of the participants. The lower limit of the representation, ,in The total number of participants. For participants' index, The total number of time intervals. As a decision-making indicator, This is a preset lower limit for willingness to participate;

[0082] Step 407: Ensure that the total monetary rewards for all participants do not exceed the budget. , ;

[0083] Step 408: Collect key procedural feature data from the routine tasks of all federated learning participants;

[0084] Step 409: Select a subset of participants to complete the willingness to pay questionnaire;

[0085] Step 410: Transmit the collected feature data and completed payment willingness questionnaire data to the edge server;

[0086] Step 411: Construct the questionnaire participant set ,in For the nth participant in the participant set, For participants' index, The number of participants in the questionnaire. The total number of participants. For the set of all possible participants;

[0087] Step 412: Use calculate ,in , , and These represent the number of questionnaire participants on the [date / time]. Service Quality - Reward Level The following parameters are considered: QoS for routine tasks, federated learning rewards, latent variables, and WTP scores. The preference weight for the nth participant is used to balance federated learning rewards and service quality in the calculation. The impact of time;

[0088] Step 413: Initialize flags: , This is the identifier for the Nth participant;

[0089] Step 414: Initialize the preference matrix: , The preference weight of the Nth participant. The average preference weight of the Nth participant;

[0090] Step 415: Determine Is it If the condition is met, proceed to step 416; otherwise, proceed to step 421.

[0091] Step 416: Use infer and ,in For the intercept term, For the slope term, To find α and β that minimize the objective function, For participants' index, The total number of participants. The total number of service quality-reward levels, For service quality-reward levels index, Let the nth participant's WTP score be the l-th service quality-reward level. For the nth participant, the latent variables are defined at the l-th service quality-reward level.

[0092] Step 417: For arrive ,use Calculate preference weights ,in, For the optimal intercept term, This is the optimal slope term;

[0093] Step 418: Determine If the condition is met, proceed to step 420; otherwise, proceed to step 419. The average preference weight of the Nth participant. For optimal preference weights, To participate in the decay factor;

[0094] Step 419: Update The logo: Proceed to step 417, where For the nth participant in the participant set, This is the identifier for the nth participant;

[0095] Step 420: Generate a new set of preference weights and resample the preference matrix using the Latin hypercube importance sampling technique. Proceed to step 415;

[0096] Step 421: Construct the optimization problem ,in It is a preference weight matrix. It has A mapping matrix of dimensions, For feature scoring matrix;

[0097] Step 422: Use the LINGO solver to efficiently process the optimization problem and obtain the mapping matrix. ;

[0098] Step 423: Define a feature scoring matrix each of the rows exist Storage and Participants Quantified values ​​of nine features related to routine tasks;

[0099] Step 424: For arrive Parallel execution, using the open-source compiler tool Clang to obtain participants. characteristic scores ,use Calculate preference weights To obtain the optimal preference weight matrix , For a predefined weight vector;

[0100] Step 425: Return , Let be the optimal preference weight for the Mth participant;

[0101] Step 426: The personalized scheduling and preference parameter derivation of federated learning participants are completed, exit.

[0102] Furthermore, step 5, the reinforcement learning scheduling and dynamic resource allocation of the federated learning system, specifically includes:

[0103] Step 501: Define Participants Energy consumption during federal learning training and routine mission execution ,in Let be the remaining energy of the m-th participant at the start of the (k+1)-th time interval. Let m be the remaining energy of the m-th participant at the beginning of the k-th time interval. This is the participation decision indicator for the m-th participant in the k-th time interval. The energy consumed by the m-th participant during federated learning training at the k-th time interval. The energy consumed by the m-th participant when performing a routine task in the k-th time interval;

[0104] Step 502: Define the remaining reward budget ,in Let $\frac{k+1} ...$ be the remaining reward budget for the time interval $\frac{k The remaining reward budget for the k-th time interval; This is the participation decision indicator for the m-th participant in the k-th time interval; As a monetary reward;

[0105] Step 503: Define the remaining training time ,in This represents the remaining training time for the (k+1)th time interval. This represents the remaining training time for the k-th time interval. The duration of the k-th time interval;

[0106] Step 504: Initialize the round counter ;

[0107] Step 505: Initialize the parameter vector and ,in The parameter vector generated for the strategy, This is the error term for model updates;

[0108] Step 506: Determine Check if the condition is met. If it is met, proceed to step 507; otherwise, proceed to step 522.

[0109] Step 507: Policy-based Generate Action ,in This represents the state during the k-th time interval;

[0110] Step 508: Based on the conditional exponential probability density function produce ,in ,in For the parameters (rate parameters) of the exponential distribution. The duration of the k-th time interval;

[0111] Step 509: All federated learning participants in the... Local sampling is performed within a time interval, and the future resource status is reported to obtain a sample set. ,in Let m be the expected amount of routine task resources required by the m-th participant during the k-th time interval. The total number of participants. An index for participants;

[0112] Step 510: Determine the reward budget for one round ,in The percentage of the reward budget allocated for the k-th time interval. For the total reward budget;

[0113] Step 511: If Then the participants Insert into the active set In the middle, and sorted in descending order according to preference weight;

[0114] Step 512: For arrive If for Each participant ,if Then allocate rewards. Use renew , Let be the optimal weight for the i-th participant. Let i be the service quality score of the i-th participant in the k-th time interval. The remaining reward budget for the k-th time interval;

[0115] Step 513: Otherwise, allocate the remaining budget to ,Right now , ;

[0116] Step 514: Use the reward function Export Action Rewards ,in It is the accuracy of the federated learning global model. It is up to the number The total workload time of all federated learning participants in the time slot. The target level of willingness to participate;

[0117] Step 515: Use Inferring advantage, The reward for the i-th time interval. For the dominant function, For error terms, The expected reward under a given policy Ai and error term E;

[0118] Step 516: Use Update policy parameter vector ,in Let R be the objective function for optimizing the policy parameter vector. In the state The following actions This is the current policy parameter vector. In the state Take action The new strategy probability distribution For the old policy parameter vector, In the state Take action The expected advantage function, In the state Take action The probability distribution of the old strategy, Indicates the shift from the old strategy Middle of the state and action vectors The expectation of sampling, It is a penalty coefficient that balances maximizing expected advantage and minimizing the importance of differences between strategies;

[0119] Step 517: Use Update the criticism parameter vector , The update is performed by minimizing the loss between the estimated cumulative reward and the actual cumulative reward. For loss function, This represents the actual cumulative reward for the i-th time interval. This represents the state during the i-th time interval. In the state Given parameter E, the expected cumulative reward;

[0120] Step 518: For each participant ,use Update individual participant status ,in Participant As of the The remaining energy in each time slot, Let m be the expected resource requirement for routine tasks for the m-th participant during the k-th time interval;

[0121] Step 519: Use Update the status of the Federated Learning System ,in and These represent the remaining budget available for monetary rewards and training time in future rounds of federal learning training, respectively.

[0122] Step 520: Storage State Transition , This represents the state during the k-th time interval. In the state The following actions were taken. The reward for the k-th time interval. This represents the state during the (k+1)th time interval;

[0123] Step 521: Add a round counter Proceed to step 506. This is the current round counter;

[0124] Step 522: Return scheduling parameters , This is the participation decision indicator for the m-th participant in the k-th time interval. The percentage of the reward budget allocated for the k-th time interval. Let k be the duration of the k-th time interval. Let m be the training resource requirement of the m-th participant in the k-th time interval. The total number of participants. For participants' index, This represents the total number of time intervals. Indexed by the current time interval;

[0125] Step 523: The reinforcement learning scheduling and dynamic resource allocation of the federated learning system are completed. Exit.

[0126] A specific application example of this invention:

[0127] This invention constructs a prototype FL system consisting of a high-performance edge server and seven terminal devices. The edge server is equipped with an i9-10900x 10-core CPU and a GeForce RTX 3080-10G GPU, while the terminal devices include an NVIDIA Jetson AGX Xavier module, a Raspberry Pi-5 board, and a Raspberry Pi-4B board. The system uses a Transformer model as the FL training model and simulates solar power trajectories to simulate energy harvesting power.

[0128] This invention utilizes the AWE dataset, which contains six machine health conditions, including one normal condition and five fault conditions. The dataset is distributed to end devices in both IID (Independent and Identically Distributed) and non-IID (Non-Independent and Identically Distributed) environments to simulate data distribution in a real industrial environment.

[0129] This invention collects participants' willingness to pay (WTP) data through a questionnaire survey and constructs 36 QoS reward levels. The WTP prediction method is trained using three terminal devices, with the remaining devices used as test devices to evaluate the accuracy of the WTP predictions.

[0130] This invention employs a customized fast PPO method for personalized participant scheduling within a reinforcement learning framework. The algorithm optimizes reward allocation and training time for federated learning participants through greedy reward distribution and heuristic training duration partitioning.

[0131] This invention uses the LINGO solver to handle linear problems to obtain the optimal preference weight matrix. This step is to establish the correlation between participants' preferences and the project characteristics of their daily tasks.

[0132] This invention aims to improve the accuracy and sustainability of federated learning models by optimizing the management of computational resources and the allocation of rewards for participants. Specifically, this invention focuses on how, in a renewable energy-driven green IIoT environment, personalized federated learning incentive models and participant scheduling techniques can be used to adapt to the unique resource usage patterns and incentive preferences of federated learning participants. Furthermore, it explores how policy learning within a reinforcement learning framework can accelerate this process, thereby ensuring active participant engagement while enhancing the overall performance of the federated learning model.

[0133] The objective function of this invention is to maximize the accuracy of the federated learning model while satisfying the conditions of renewable energy driving forces and participant incentive preferences. The specific objective function is as follows:

[0134] ;

[0135] in, It is the learning rate of FL in the kth round. Participant The size of the local dataset, It is the total size of all distributed local datasets. Participant The participation count counter.

[0136] Figure 2 The importance of static and dynamic program features for routine tasks is demonstrated. Static features #1 (number of instruction blocks) and #2 (load / store count) have the highest importance, at 0.70 and 0.50 respectively, while dynamic feature #7 (global thread count) also has a high importance, at 0.60. This indicates that in personalized federated learning, these features are crucial for predicting participants' resource needs and incentive preferences.

[0137] Figure 3 The variation in FL precision was demonstrated using a duration allocation method. Figure 3 (a) In the IID scenario, the maximum accuracy reached 67.4%, while Figure 3(b) In the non-IID scenario, the maximum accuracy is 62.7%. This shows that the method of the present invention can maintain high model accuracy under different data distributions.

[0138] Figure 4 The confusion matrix of the willingness-to-pay (WTP) prediction method is presented. The model's prediction accuracy for willingness levels 2, 3, and 4 is 84.4%, 81.3%, and 82.7%, respectively, with an average prediction accuracy of 83.1%. This demonstrates that the WTP prediction method of this invention has high accuracy.

[0139] Figure 5 The accuracy of the global FL model obtained through five methods was compared. The method of this invention achieves 97.7% accuracy in the IID environment and 90.4% accuracy in the non-IID environment, outperforming all benchmark methods. This highlights the superior accuracy of the method of this invention when handling different data distributions.

[0140] Figure 6 The total training time required to achieve the desired accuracy level is listed. In the IID setting, the method of this invention requires 92.3 minutes, 131.2 minutes, and 175.7 minutes at 50%, 60%, and 70% accuracy, respectively, demonstrating high time efficiency. In the non-IID setting, the method of this invention requires 128.7 minutes, 180.3 minutes, and 251.2 minutes at 45%, 55%, and 60% accuracy, respectively, also demonstrating good time efficiency.

[0141] Figure 7 The average computational resource utilization of terminal devices implemented using different federated learning methods is shown. The method of this invention achieves resource utilization exceeding 90% in both IID and non-IID environments, demonstrating superior resource management capabilities. In contrast, other methods such as DTEI and BDMA achieve resource utilization below 75% in non-IID environments, further proving the effectiveness of the method of this invention.

[0142] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A personalized federated learning method for green industrial Internet of Things, characterized in that, include: Initialize the federated learning system architecture, which consists of edge servers and multiple industrial IoT nodes; Construct energy models for participants and calculate energy buffer capacity and total energy supply; Construct an incentive model based on service quality and monetary rewards, and quantify participants' preference parameters; Personalized scheduling parameters are derived through feature analysis and questionnaire feedback; A reinforcement learning framework is used to achieve dynamic resource allocation and reward budget scheduling; The energy model for the participants includes: The system integrates solar and vibration energy harvesting modules; calculates the harvesting power based on instantaneous irradiance and vibration velocity; determines the charging and discharging dynamics of the energy buffer module; calculates the energy consumption for training and routine tasks respectively; and establishes energy supply and demand balance constraints. Constructing an incentive model based on service quality and monetary rewards includes: Service quality indicators are calculated based on resource allocation differences; potential variables are generated by combining service quality and monetary rewards; participation is calculated based on basic willingness to participate and marginal utility; preference weight parameters are dynamically adjusted; and a cumulative reward budget control mechanism is implemented. The derivation of personalized scheduling parameters includes: Analyze participants' historical participation counts; construct an objective function to maximize model accuracy; implement real-time energy consumption monitoring; verify training deadlines; ensure computational resource allocation does not exceed limits. Personalized scheduling includes: Collect routine task procedure feature data; select some participants to complete the payment willingness questionnaire; construct the questionnaire participant feature matrix; optimize the preference weights using Latin hypercube sampling; and calculate the optimal weight matrix using a linear solver.

2. The personalized federated learning method for green industrial IoT as described in claim 1, characterized in that, Initializing the federated learning system architecture includes: Configure an edge server responsible for global model aggregation; define industrial IoT nodes with energy harvesting capabilities as participants; allocate a local dataset containing feature vectors and labels to each participant; divide the scheduling time range into non-uniform time intervals; and predefine the regular task computation requirement curves for each participant.

3. The personalized federated learning method for green industrial IoT as described in claim 1, characterized in that, Implementing reinforcement learning includes: Define the system state, which includes remaining energy, budget, and time; generate resource allocation actions based on the policy; generate training time intervals according to an exponential distribution; implement a preference-based greedy reward allocation; and update the policy parameters and criticism parameters.

4. The personalized federated learning method for green industrial IoT as described in claim 1, characterized in that, Dynamic resource allocation includes: Maintain the remaining energy state of participants; calculate the remaining reward budget and training time; store system state transition records; and return scheduling parameters containing participation decisions, time allocation, and resource allocation.

Citation Information

Patent Citations

  • Federal learning system and incentive method of energy block chain

    CN116579442A

  • Joint optimization method and system for participant selection and resource allocation of federated learning

    CN117560724A