Medical data federal learning sharing excitation method based on Stackelberg-Bayesian game

By combining Stackelberg–Bayesian game theory with blockchain smart contracts, the problems of information asymmetry and privacy utility in the incentive mechanism for edge medical data sharing are solved, realizing efficient and reliable incentives for medical data sharing and improving the real-time performance and privacy protection of edge computing.

CN122069076APending Publication Date: 2026-05-19YUNNAN UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUNNAN UNIVERSITY OF FINANCE AND ECONOMICS
Filing Date
2026-02-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing incentive mechanisms for sharing edge medical data suffer from problems such as idealized rational assumptions, low solution efficiency, high consensus resource consumption, mismatch with information asymmetry, and privacy-utility trade-offs. These issues result in insufficient adaptability and efficiency of the mechanisms, making it difficult to achieve efficient and reliable sharing of medical data.

Method used

We construct a shared incentive method for federated learning of medical data based on Stackelberg–Bayesian game theory. We adopt a cloud-edge-device collaborative architecture, leverage MEC edge offloading to overcome the resource bottleneck of IoMT devices, and realize decentralized end-to-end trust and automatic settlement through blockchain smart contracts. We combine Bayesian belief update mechanism and differential privacy technology to design unified and differentiated incentive strategies, and optimize computation offloading mode and privacy budget.

Benefits of technology

It enables precise incentives even with incomplete information, alleviates equipment pressure, ensures privacy and security, optimizes computational offloading and privacy protection, and improves the efficiency and credibility of medical data sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069076A_ABST
    Figure CN122069076A_ABST
Patent Text Reader

Abstract

The invention discloses a medical data federated learning sharing excitation method based on a Stackelberg-Bayesian game, and belongs to the technical field of IoT (Internet of Things) data processing. The architecture is composed of three layers of IoMT data proxy nodes DA, mobile edge computing servers MEC and alliance chain nodes CB in a collaborative mode, the DA has double identities of a data supplier DAi and a demander DAr, the MEC is responsible for cost estimation, model aggregation and training unloading, and the alliance chain nodes achieve task broadcasting, privacy evidence storage and incentive execution. According to the method, an incentive mechanism is designed based on a Stackelberg-Bayesian game framework, federated learning (FL) and a differential privacy technology are combined, credible data sharing between IoMT nodes is achieved through the steps of task chaining, cost reporting, local training and privacy protection, contribution evaluation, model aggregation iteration and the like, and meanwhile effectiveness of a data supply party and a data demand party is maximized. According to the method, the problems of privacy disclosure, trust missing and non-uniform incentive in medical networking data sharing are solved, the training efficiency and safety are improved, and the method is suitable for a medical networking multi-node collaborative data processing scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of Internet of Things and edge computing, and in particular to a medical data federated learning and sharing incentive method based on Stackelberg–Bayesian game theory, which is mainly applied to edge medical data sharing scenarios. Background Technology

[0002] The integration of blockchain and federated learning is a research hotspot in the fields of IoT and edge computing. The combination of the two can rely on the immutability and incentive mechanism of blockchain to effectively incentivize the sharing of sensitive data and build a trusted distributed collaborative environment, which has broad application prospects in medical data processing scenarios.

[0003] Traditional IoT architectures suffer from high transmission latency and weak real-time performance. Medical data must be uploaded to cloud servers for centralized processing, which not only fails to meet the real-time response requirements of medical scenarios but also easily leads to privacy risks during transmission. Edge computing technology pushes data processing tasks to the network edge, effectively shortening transmission distances and improving real-time processing. Existing research has explored integration solutions for IoT and edge computing in smart city healthcare services, providing preliminary ideas for the construction of edge medical data sharing systems.

[0004] Stackelberg games, due to their leader-follower sequential decision-making structure, are widely used in the design of incentive mechanisms for sharing edge medical data. However, existing solutions have significant technical flaws: First, they are based on the assumption of "perfect rationality," which is inconsistent with real-world scenarios, ignores the bounded rationality of participants, and has poor incentive adaptability. Second, facing the high-dimensional policy space of large-scale edge nodes, traditional algorithms have low convergence efficiency and cannot meet the real-time requirements of edge computing. Third, they are not tightly integrated with blockchain consensus algorithms, and the high-energy-consuming consensus algorithms increase the burden on resource-constrained edge devices, increase node sharing costs, and weaken participation enthusiasm.

[0005] The participants in medical IoT data sharing are rational economic agents within a decentralized federated learning ecosystem, whose decision-making requires balancing monetary gains, computational energy consumption, and privacy risks. Nodes typically employ differential privacy mechanisms to meet privacy compliance requirements, but there is an inherent contradiction between privacy protection and data utility. Non-cooperative game theory becomes an effective tool for analyzing this multi-dimensional conflict scenario. The hierarchical nature of federated learning task publishing and response is highly compatible with the single-leader-multiple-follower Stackelberg game architecture. However, the traditional Stackelberg model, based on the assumption of complete information, cannot address the information asymmetry issues caused by private information such as IoMT node computational costs, data quality, and privacy preferences in real-world medical data sharing scenarios. This can easily lead to pricing strategy deviations and affect incentive effects, necessitating the introduction of a Bayesian game framework to achieve optimal strategy formulation under incomplete information constraints.

[0006] In summary, existing incentive mechanisms for edge healthcare data sharing suffer from problems such as idealized rational assumptions, low solution efficiency, high consensus resource consumption, failure to adapt to information asymmetry, and privacy-utility trade-offs. These issues result in insufficient adaptability and efficiency, hindering the efficient and reliable sharing of medical data and failing to meet the practical application needs in the context of deep integration of the Internet of Things and edge computing. Developing an edge healthcare data sharing incentive scheme that adapts to real-world scenarios and balances incentive effectiveness, real-time performance, and low cost has become a pressing technical challenge in this field. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a medical data federated learning and sharing incentive method based on Stackelberg-Bayesian game theory. It constructs a cloud-edge-device collaborative trusted architecture, leverages MEC edge offloading to overcome IoMT device resource bottlenecks, and achieves decentralized end-to-end trust and automated settlement through blockchain smart contracts. Secondly, addressing the challenge of incomplete information—that data requesters cannot know the true type of nodes—the incentive process is modeled as a two-stage Stackelberg-Bayesian game. This model uses a Bayesian belief update mechanism to accurately characterize the private cost distribution of nodes and designs two incentive strategies: unified and differentiated. Through backward induction, the optimal pricing and closed-form solution for data contribution under Bayesian Nash equilibrium are derived. Finally, in the follower phase of the game, the computation offloading mode and differential privacy budget are jointly optimized, maximizing utility while ensuring privacy.

[0008] To achieve the above objectives, this invention provides a medical data federated learning and sharing incentive method based on Stackelberg–Bayesian game theory. The method adopts a cloud-edge-device three-layer architecture, namely IoMT data proxy node DA, mobile edge computing servers (MEC servers), and consortium blockchain node CB. The IoMT data proxy node DA simultaneously possesses the dual identities of a data provider DAi and a data demander DAr; the data demander DAr publishes task T to the consortium blockchain node CB, and the consortium blockchain node CB initializes the model parameters. The task template is generated and the on-chain task contract is broadcast to all potential data supply nodes (DAi), and the data supply nodes (DAi) perform local model training. The Mobile Edge Computing (MEC) servers are set up for each region and are used to estimate the data supply nodes. The cost function Ci is used, and the aggregated global model W and each data supply node are fed together. The contribution level and data quality are uploaded to the consortium blockchain node to provide a basis for incentive allocation; The consortium blockchain node CB is used to implement task publishing, identity verification, model parameter storage, contribution verification, smart contract execution and incentive distribution, and adopts the PBFT consensus mechanism to verify the global model; The method specifically includes the following six steps: Step 1: Task On-Chain After Initialization: When the IoMT data broker node (DA) participates in data sharing or publishes a task for the first time, it triggers a smart contract to complete identity verification and registration on the blockchain, obtaining identity information, certificates, keys, and token account addresses; Data requester Task information is published through consortium blockchain nodes. The task information includes a reward mechanism, required computing resources, and a description of data requirements. The task information is then broadcast to all data supply nodes that intend to participate in the data sharing. ; Step 2: Reporting Costs and Selecting Participating Nodes: Data Supply Nodes to be Participated in Data Sharing The report details the cost of completing this task and the data requester. Trigger the smart contract, combined with data supply nodes The cost and reputation values ​​were calculated and sorted to determine the list of data suppliers; Step 3: Announce reward measures for data requesters. Incentive policies should be developed based on the type of data sharing, the amount of data, and the actual computing costs for data providers. The data providers and data users can independently determine their actual contribution and negotiate to determine the optimal data allocation decision, so that both data providers and data users can obtain the maximum benefit. Step 4: Local Training and Privacy Protection: Data Supply Nodes Obtain the initial training model and perform local training. The training process can be executed locally or offloaded to the corresponding mobile edge computing (MEC) servers in the specified region. After training is complete, data is supplied to the nodes. The local model update is protected by a differential privacy mechanism. The processed model parameters are signed and submitted to the consortium blockchain verification node, triggering the contribution contract to complete secure notarization. Step 5: Contribution Assessment and Reward Distribution: Consortium blockchain validator nodes verify the signature validity of submitted local models and evaluate model quality according to unified verification metrics; subsequently, they calculate the contribution assessment for each data supply node according to the contribution assessment mechanism. The contribution and reputation value of the data supply nodes are updated. The reputation value is stored in the consortium blockchain reputation node, and rewards are issued simultaneously; Step 6: Model Aggregation and Iterative Training: After the consortium blockchain verification nodes complete their verification, the edge aggregation nodes use the FedAvg algorithm to aggregate the local models and generate a new global model. This global model is verified in the consortium blockchain through the PBFT consensus mechanism and then broadcast to all participating nodes. Each participating node downloads the latest global model and continues the next round of training until the model converges. During the aggregation process, a differential privacy mechanism is applied to protect the privacy of the participants' intermediate gradients.

[0009] Preferably, the IoMT data proxy nodes (DAs) constitute a set of nodes. Where N is the total number of IoMT nodes, which is also the total number of clients participating in federated learning. Holding local datasets K is a variable index representing the "Kth" node in the set. This represents the input of the j-th sample. For the corresponding prediction results, The number of samples, and each Maintaining a weighted vector To store model parameters.

[0010] Preferably, the mobile edge computing server acts as an edge aggregation node, maintaining global model parameters W. Its optimization objective is to minimize the weighted average loss of DA across all IoMT data proxy nodes, with the weighting coefficients being... Number of samples The specific optimization objective satisfies the formula: ,in It is a weight vector; This represents the number of samples, where K represents the Kth node, and the superscript N represents the total number of IoMT nodes. yes The local training loss function is given by the following formula: Where j is the index of the data sample in the local dataset, representing the nth data record possessed by that node, and the value of j ranges from 1 to... That is, the node The total number of local samples; T represents "matrix transpose".

[0011] Preferably, the update formula for the local model training in step 4 is: Where λ is the learning rate, and t represents the current federated learning communication round index, which is an integer, typically ranging from 0 (initial state to T-1) or 1 to T, where T is the preset maximum convergence round; in the formula, it is used to mark the evolution of model parameters over time; after the model training is completed... Differential privacy dynamics (DP) is applied to local parameters to prevent reconstruction attacks and model inference attacks. The specific formula is as follows: ,in Privacy Budget Decide; Local loss function, while Represents gradient, It is the gradient vector of the loss function, and its dimension is the same as that of the model parameters. same.

[0012] Preferably, the update formula for the global model aggregation in step 6 is: ,in For each data supply node Privacy-protected local model parameters For the sample size is each The number of samples is N, where N is the total number of IoMT nodes and t is the communication round index of the current federated learning.

[0013] Preferably, the incentive policy is formulated based on the Stackelberg-Bayesian game framework, and the data demand side... As a leader, the data supply node As followers, both parties aim to maximize their own utility functions; The data supply node The utility function satisfies: ,in For node i, a valid payment is made per unit. The amount of data contributed to node i. Let i be the total cost of node i; The data demander The utility function satisfies: ,in For data users to extract value from data sharing, For the total payment to node i, The amount of data contributed to node i. ;in, This is the data quality coefficient for node i. The accuracy of the data, the quality of the annotation, or the completeness of the medical records are determined by... express; It is the data contribution of node i; It is the "unit contribution reward coefficient" or "payment strategy" set by the data demander for node i.

[0014] Preferably, the data supply node Total cost ; It consists of computing energy consumption, communication energy consumption, encryption energy consumption, privacy costs, and latency costs, satisfying the formula: ; in, To calculate energy consumption, satisfy ; It is the energy consumption coefficient per node. It is the amount of contribution completed. The required computation is often approximated as follows: ; For communication energy consumption, in, For upload time, Transmission power; For encryption power consumption, , Encryption energy consumption per unit To measure the amount of data uploaded; To satisfy privacy costs , The level of privacy sensitivity depends on the privacy budget. ; To meet the delay cost, , This represents the delay cost coefficient.

[0015] Preferably, the incentive policy includes two types: a unified reward mechanism (UIM) and a differentiated reward mechanism (DIM). In the unified reward mechanism, the effective payment per unit satisfies the formula: ,in for A unified basic unit price for decision-making. For adjustment coefficients, qi represents data quality; data supply nodes The total reward satisfies the formula: ; In the aforementioned differentiated reward mechanism, the effective payment per unit is for the data demander. For each data supply node Customized exclusive unit price Data supply nodes The total reward satisfies the formula: .

[0016] Preferably, the Stackelberg-Bayesian game has a unique equilibrium solution, specifically: Follower subgame equilibrium: Given an incentive price, each data supply node... The optimal data contribution satisfies the formula: ,in The linear cost coefficient is... Here, represents the coefficient of the quadratic term related to the load, and qi represents the data quality. for A unified basic unit price for decision-making; Leader Game Equilibrium: Data Demander Based on data supply nodes To find the optimal response, formulate the optimal reward strategy r* to maximize one's expected utility.

[0017] Preferably, the method further includes a federated learning training process based on the Stackelberg-Bayesian incentive mechanism, the specific process of which is as follows: Input: Initial model Maximum number of iterations T, learning rate η, privacy parameters; Output: Final Model ; Initialization: Consortium blockchain launch task T and initial model ; For t = 1, 2, …, T, perform the following steps: Phase 1: Leader's Incentive Decision-Making: Data Demanders Estimate the follower type distribution parameters Hi and Ai, where Hi = E1 / βi and Ai = Eαi / βi; If a unified reward mechanism (UIM) is adopted, the optimal scalar reward r* = argmaxEURUIM is calculated. If a differentiated reward mechanism (DIM) is adopted, the optimal vector reward r* = argmaxEURDIM is calculated; The broadcast reward policy r* and the current global model Wt-1; Phase Two: Follower Response and Local Training Each IoMT data broker node (DAi) performs the following operations in parallel: Minimize the cost coefficient βi by selecting the mode (local training or MEC-assisted training); Calculate the optimal data contribution xi*=max{0,rqi-αi / βi}; If xi*>0, perform local training: ωit←ωit-1-η∇Fi; Apply differential privacy mechanism: ωit←ωit+N0,σ2I; Upload the processed local model parameters ωit and contribution proof; Aggregation and reward distribution; Mobile edge computing servers / consortium blockchains verify the contributions and data quality of each node (qi). The FedAvg algorithm is used to aggregate the global model: Wt←DkDtotalωkt; Smart contract execution rewards payments, paying Pi to active data provider nodes (DAi); If the model accuracy converges, terminate the iteration; Return the final model Wt.

[0018] Table 1 Simulation parameter settings Compared with the prior art, the present invention has the following beneficial effects: (1) Establishing a Bayesian game incentive mechanism and equilibrium solution under incomplete information: To address the challenge of incomplete information where service demanders cannot know the true cost and data quality of nodes, we modeled the incentive process as a two-stage Stackelberg-Bayesian game. By introducing a Bayesian belief update mechanism, we designed a unified incentive mechanism suitable for low-complexity scenarios and a differentiated incentive mechanism suitable for high-performance scenarios. Using backward induction, we derived the unique Bayesian Nash equilibrium of the game and provided a closed-form solution for optimal pricing and data contribution, effectively solving the problem of precise incentives under information asymmetry.

[0019] (2) A cloud-edge-device collaborative federated learning architecture with differential privacy enhancement was proposed: a hierarchical system integrating consortium blockchain and mobile edge computing was constructed. On the one hand, MEC computing power offloading and differential privacy technology were used to alleviate device pressure and ensure parameter security; on the other hand, decentralized model verification and smart contract reward distribution were achieved by relying on blockchain, thus constructing a dual guarantee of "privacy computing + trusted service".

[0020] (3) Joint optimization of computation offloading and privacy budget is achieved: In view of the coupling characteristics of physical resources and privacy protection, the computation offloading mode and differential privacy budget are jointly decided in the follower stage of the game. Attached Figure Description

[0021] Figure 1 System architecture diagram of this invention; Figure 2 Global convergence performance under different differential privacy budgets; Figure 3 a. Existence of Nash equilibrium; 4b. Convergence of Stackelberg-Bayesian game mechanism in dynamic interactions; Figure 4 Privacy Budget Adaptive excitation Leader's expected utility Inter-coupling relationship; Figure 5 MEC collaborative computing mechanism optimizes the physical layer performance. Detailed Implementation

[0022] The present invention will now be described in further detail with reference to specific embodiments.

[0023] This invention proposes a shared incentive method for federated learning of medical data based on Stackelberg–Bayesian game theory to support trusted data sharing and incentive allocation among IoMT nodes. The system consists of a three-layer architecture: IoMT data agent nodes (DAs), mobile edge computing servers (MEC servers), and consortium blockchain nodes (CBs). As a medical internet terminal device, the DA can act as a data provider. It can also serve the data demand side. . Issue tasks to consortium blockchain nodes Initialize model parameters for nodes The task template triggers an on-chain task contract broadcast to all potential data supply nodes. Each node performs local model training. The MEC servers in each region estimate the cost function for each node. And the aggregated global model w, contribution, and data quality will be combined. Uploaded to consortium blockchain nodes for incentive distribution, specifically as follows: Figure 1 As shown.

[0024] Assuming each participant in the system possesses a local dataset and can participate in data sharing activities, the workflow steps for data sharing are as follows: Step 1 Initialization and Task On-Chain: When the terminal DA participates in data sharing or publishes a task for the first time, it triggers a smart contract for identity verification, completes registration on the blockchain, and obtains identity information, certificates, keys, and token account addresses. Data requesters publish task information through blockchain nodes, including the reward mechanism, required computing resources, and a description of data requirements, which is then broadcast to various data providers intending to participate in the sharing.

[0025] Step 2: Report Cost Participation Node Selection: Data owners who intend to participate in data sharing report the cost of completing the task. The task issuer needs to comprehensively consider the data owners' costs and reputation values, which will trigger the smart contract to calculate and sort the list of data suppliers.

[0026] Step 3: Issue incentive measures: The data requester provides incentive policies based on the type and amount of data shared and the actual computing cost, while the data provider decides its own actual contribution to negotiate the optimal data allocation decision, so that both the data provider and the data requester can obtain the maximum benefit.

[0027] Step 4: Local Training. The data supply node obtains the initial training model for local training, either locally or offloaded to the MEC server. After training, the node updates the local model, applies a differential privacy mechanism to protect the model's privacy, signs the update, submits it to the blockchain verification node, and triggers the contribution contract, achieving secure on-chain parameter verification and notarization.

[0028] Step 5: Evaluate Contributions and Issue Rewards: Verify the signature validity of the submitted local model and evaluate its quality according to unified verification metrics. Then, calculate the contribution and reputation value of each node based on the contribution evaluation mechanism and record them for reward issuance. Finally, update the data owner's reputation value and store it in the blockchain reputation node for node selection in the next task.

[0029] Step 6: Model Aggregation and Iterative Training: After validation, the aggregation nodes use algorithms such as FedAvg to aggregate the local models and generate a new global model. Once this model reaches consensus in the blockchain, it is broadcast to all nodes. Nodes download the latest global model to continue the next round of training. Differential privacy mechanisms are applied during the aggregation process to protect the privacy of intermediate gradients of participants until the model converges.

[0030] The model was established as follows: Preliminary preparations for federal learning All medical network nodes constitute a set of data proxy nodes, that is ,in Holding local datasets .in, This represents the input of the j-th sample. For the corresponding prediction, Define a weight vector for the number of samples. To maintain model parameters. Referring to method

[34] , we will define the local training loss function as shown in formula (1): (1) The MEC server, acting as an edge aggregation node, maintains the global model parameters W. Its optimization objective is to minimize the weighted average loss of all DAs, as shown in formula (2), where... As a weight, it reflects the importance of the sample's contribution.

[0031] (2) The entire federated learning process includes the following four steps, iterating repeatedly in the system to prevent model convergence.

[0032] Step 1: Global Model Deployment: At the start of training, the blockchain node triggers a smart contract to issue an initialization task, recording the task requirements and model parameter template. The aggregation node will then obtain the initial global model from the blockchain. The broadcast is sent to all participating DA devices.

[0033] Step 2: Local model training: Each Perform model updates using local data: (3) in To determine the learning rate and adapt to the heterogeneous computing power of IoMT nodes, DA can be trained either directly by the local gateway or offloaded to the MEC server for training.

[0034] Step 3: Differential privacy protection

[35] : After the model training is completed, Apply dynamic programming (DP) to local parameters to prevent reconstruction attacks and model inference attacks: (4) in Privacy Budget The model parameters after DP processing are determined. After being signed, it is submitted to the blockchain verification node to achieve tamper-proof evidence storage.

[0035] Step 4: Global Model Aggregation: After successful verification, the aggregation node performs a global model update according to the FedAvg algorithm

[36] : (5) After the update Once uploaded to the blockchain and verified through the PBFT consensus mechanism, the data is broadcast to all nodes, and then the next round of training begins.

[0036] Game modeling Utility of Cost-utility function of data suppliers The utility function is defined as: (6) For node i, a valid payment per unit. The definitions are referenced in

[25]

[37]

[38] , and will not be discussed in detail here; Let the total cost of node i be defined as follows; (7) First, the energy consumption of computation performed by the reference

[25] node locally or with MEC assistance is expressed as follows: (8) in, It is the energy consumption coefficient per node. It is the amount of contribution completed. The required computation is often approximated as follows: Communication energy consumption is: (9) in, For upload time, The transmission power is [value]. Encryption power consumption can be written as [value]. (10) in, Encryption energy consumption per unit This refers to the amount of data uploaded. Here, the privacy cost is defined as: (11) The level of privacy sensitivity depends on the privacy budget. The delay cost is expressed as: (12) therefore, The utility function can be written as: (13) Considering the increasing data volume As the cost increases, marginal utility decreases, and the computational load on node devices grows non-linearly. Similar to the method in

[39] , the cost function is modeled as a convex function as shown in Formula 1, where Including linear cost coefficients such as those related to communication and privacy, These are the coefficients of the quadratic term related to the load.

[0037] (14) For ease of differentiation and analysis, it is written as: (15) The derivative obtained in the first section is: (16) Then obtain the follower in the given Through other nodes To achieve the best response under resource allocation, Bayesian settings need to be configured accordingly. The component that depends on others' strategies or types takes the expected value. If differentiation is adopted... Then replace accordingly. , (17) (2) Utility of Data demander's payment function for Its core objective is to maximize its net profit by exchanging high-quality medical data models for payment incentives while meeting budget constraints

[38] . Specifically, Gain value from contributions and pay the total payment Therefore, its utility function is: (18) To characterize the differentiated importance of contributions from different nodes in the task, a commonly used linear quality-weighted value function is adopted: (19) in for The weights for the importance of data at different nodes are therefore: (20) The goal is to determine the best response to the Follower. Prediction, selection This maximizes the expected net return. Within the Stackelberg-Bayesian game framework, Since the specific response of a node cannot be known precisely, it is necessary to base it on the optimal response of the follower. Perform Bayesian predictions and develop incentive strategies to maximize expected utility.

[0038] Here, referring to

[39] , two reward mechanisms are given: Uniform Incentive Mechanism (UIM) and Discriminatory Incentive Mechanism (DIM). In the case of incomplete information, in order to ensure the fairness of incentives and reduce the complexity of mechanism implementation, Since it's impossible to customize prices for each node, a unified reward mechanism is established, with a standardized reward definition as shown in Formula 1. (twenty one) in, (twenty two) for A unified basic unit price for decision-making. The adjustment coefficient is used to provide additional compensation for high-quality data. Therefore, the total reward of node i is as shown in Formula 1. At this time, when facing heterogeneous nodes, the Leader will weigh the pricing in order to obtain the maximum benefit.

[0039] (twenty three) In the case of complete information, if Having complete knowledge of the cost function and type of each node, first-level price discrimination can be implemented, assigning a unique optimal unit reward to each node i. Therefore, the differentiated reward is determined by… Specify as: (twenty four) Therefore, the total compensation is: (25) Although it is difficult to implement in real-world privacy protection scenarios, the system utility under this mechanism constitutes the upper bound of theoretical performance and will be used as a benchmark scheme for quantitative analysis of the change in incentive efficiency of the Bayesian mechanism under conditions of incomplete information.

[0040] Under a unified reward system, if the Followers' response is Therefore, the optimization of the leader can be expressed as a Bayesian optimization problem: (26) If there is a budget constraint, it can be expressed as: (27) For differentiated rewards, the leader can set rewards for each individual. Then we have: (28) At this point, each Follower pair The responses are independent, and the problem can be decomposed into N subproblems.

[0041] (3) Optimization Objectives The primary goal of both data demanders and data providers is to maximize their respective utility functions throughout the entire FL training and data contribution process. Firstly, each data provider receives a reward price... Then, determine your optimal data contribution. To optimize its own utility, we simplify the utility function of Ffollowers as follows: (29) So, the data provider The optimization problem is: (30) The leader is responsible for publishing the optimization goals for reward prices. However, they did not know the details of each individual. Real type Knowing only its prior distribution, its optimization problem is: (31) Among them, the Leader is subject to budget B and the system incentive cap. constraint, It is the followers who are rewarding The optimal Bayesian response is obtained by merging the two levels, which can be rewritten as a standard Stackelberg optimization problem: (32) Game equilibrium analysis This section will use backward induction to determine the equilibrium of the Stackelberg-Bayesian game, which consists of two phases: the second phase is at a given price. At that time, each data provider is based on its private type Select the optimal contribution amount The cost parameters of each node are unknown to each other, and the supply side constitutes a Bayesian game. In the first stage, the data requester pre-selects the optimal reward strategy to maximize its expected utility, while also considering the optimal response of all followers in the second stage. Next, we will prove that there exists a unique Stackelberg-Bayesian equilibrium (SBE) in this game.

[0042] (1) Second-stage solution: Follower subgame In the second phase, given the incentive price vector r (a scalar r under the unified mechanism) published by the data demander, each data supplier... Aiming to select the optimal data contribution To maximize its own utility. First, we give the definition of Nash equilibrium (NE).

[0043] Definition 1 (Nash equilibrium, NE): We will combine the strategies of the data provider. A Nash equilibrium of this subgame is defined if and only if, for any node and any feasible strategy, the inequality is satisfied. in This represents the combination of equilibrium strategies for nodes other than i. In this invention, there is no direct resource competition between nodes; therefore, the problem is an independent optimal response problem for each node.

[0044] Theorem 1: For a given incentive price, there exists a unique Nash equilibrium strategy in the follower subgame. .

[0045] Proof: The utility function defined in Section 3.3 , It is about For a continuously differentiable function, taking its first-order partial derivative gives: (33) Finding the second derivative gives: (34) Due to the cost function The second derivative is always less than zero, therefore It is a strictly concave function. There exists a unique maximum point, let the first derivative be zero and consider the nonnegativity constraint. The unique closed-form optimal response strategy is obtained as follows: (35) Q.E.D.

[0046] (2) First-stage solution: Leader game In the first phase, the data demand side Anticipating the follower's response strategy Maximize its Bayesian expected utility by designing incentive mechanisms. The following discussion will focus on differentiated incentives and uniform incentives.

[0047] Case 1: In the case of a differentiated incentive mechanism, the decision variables are a vector. The follower's optimal response obtained in the second stage. Substitute Given the objective function, we can obtain the unconstrained optimization problem: (36) Case 2: Under the unified incentive mechanism, the decision variable is a scalar r, and the objective function is... (37) (3) Proof of the existence and uniqueness of SBE Theorem 2: The proposed Stackelberg-Bayesian game has a unique Stackelberg-Bayesian equilibrium under both differentiated and uniform incentive mechanisms.

[0048] Proof: We will prove them separately below. The concavity and convexity of the expected utility function under both mechanisms.

[0049] Proof 1 uses a Hessian matrix to prove the case of differentiated incentive mechanisms: Let Since the parameters of each node are independent, the utility function has an additive separability property, and it needs to be calculated. The Hessian matrix of the decision vector r To determine its concavity / convexity. Find the objective function. The partial derivatives are: (38) in, , The desired parameters are used. The second-order partial derivatives are calculated to determine the matrix elements. The Hessian matrix is ​​constructed as follows: (39) From this, we can obtain the Hessian matrix. diagonal matrix (40) From the above formula, we can see that And data quality Since all elements on the diagonal are negative, according to the negative definite matrix criterion theorem, It is a strictly negative definite matrix, therefore, It is a strictly multidimensional concave function with respect to r, and has a unique global maximum point. .

[0050] Proof 2 uses the second derivative to prove the case of a unified incentive mechanism: Let Find its second derivative with respect to the scalar r: (41) in , Therefore, the above expression is always less than zero, that is: (42) Therefore, this expected utility function is also a strictly concave function with respect to r, and has a unique maximum point. .

[0051] In summary, considering the uniqueness of the follower equilibrium in Theorem 1, the entire Stackelberg-Bayesian game has a unique equilibrium solution. Setting its first derivative to zero, we have: Thus, the closed-form solution described above can be obtained. Q.E.D.

[0052] (4) Algorithm design In Algorithm 1, we designed a FL training process for medical data that combines privacy protection and energy efficiency optimization. This process incorporates a Stackelberg-Bayesian game mechanism, and the training process iterates continuously until the global model converges. As shown in lines 3-10, each round of FL communication t begins with... The first phase of the game involves the leader estimating the cost distribution of the followers and calculating the optimal Bayesian reward strategy to maximize expected utility. Subsequently, as described in lines 11-20, the second phase of the game unfolds among the followers. This phase is the core of the device-side optimization, where each agent determines the optimal data contribution by minimizing costs through pattern selection. Furthermore, privacy-preserving local training is performed by injecting Gaussian noise. Finally, lines 21-28 define the process for secure aggregation and incentive implementation. The consortium blockchain is responsible for verifying data quality, aggregating the global model through FedAvg, and automatically executing reward payments via smart contracts.

[0053] The above description is only a part of the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the protection scope of the present invention.

Claims

1. A shared incentive method for federated learning of medical data based on Stackelberg-Bayesian game theory, characterized by: The method adopts a three-layer architecture of cloud-edge-device collaboration, namely IoMT data proxy node DA, mobile edge computing servers MEC servers, and consortium blockchain node CB; The IoMT data proxy node DA simultaneously possesses the dual identities of a data provider (DAi) and a data demander (DAr); the data demander (DAr) publishes task T to the consortium blockchain node CB, and the consortium blockchain node CB initializes the model parameters. The task template is generated and the on-chain task contract is broadcast to all potential data supply nodes (DAi), and the data supply nodes (DAi) perform local model training. The Mobile Edge Computing (MEC) servers are set up for each region and are used to estimate the data supply nodes. The cost function Ci is used, and the aggregated global model W and each data supply node are fed together. The contribution level and data quality are uploaded to the consortium blockchain node to provide a basis for incentive allocation; The consortium blockchain node CB is used to implement task publishing, identity verification, model parameter storage, contribution verification, smart contract execution and incentive distribution, and adopts the PBFT consensus mechanism to verify the global model; The method specifically includes the following six steps: Step 1: Task On-Chain After Initialization: When the IoMT data broker node (DA) participates in data sharing or publishes a task for the first time, it triggers a smart contract to complete identity verification and registration on the blockchain, obtaining identity information, certificates, keys, and token account addresses; Data requester Task information is published through consortium blockchain nodes. The task information includes a reward mechanism, required computing resources, and a description of data requirements. The task information is then broadcast to all data supply nodes that intend to participate in the data sharing. ; Step 2: Reporting Costs and Selecting Participating Nodes: Data Supply Nodes to be Participated in Data Sharing The report details the cost of completing this task and the data requester. Trigger the smart contract, combined with data supply nodes The cost and reputation values ​​were calculated and sorted to determine the list of data suppliers; Step 3: Announce reward measures for data requesters. Incentive policies should be developed based on the type of data sharing, the amount of data, and the actual computing costs for data providers. The data providers and data users can independently determine their actual contribution and negotiate to determine the optimal data allocation decision, so that both data providers and data users can obtain the maximum benefit. Step 4: Local Training and Privacy Protection: Data Supply Nodes Obtain the initial training model and perform local training. The training process can be executed locally or offloaded to the corresponding mobile edge computing (MEC) servers in the region. After training is complete, the data is supplied to the nodes. The local model update is protected by a differential privacy mechanism. The processed model parameters are signed and submitted to the consortium blockchain verification node, triggering the contribution contract to complete secure notarization. Step 5: Contribution Assessment and Reward Distribution: Consortium blockchain validator nodes verify the signature validity of submitted local models and evaluate model quality according to unified verification metrics; subsequently, they calculate the contribution assessment for each data supply node according to the contribution assessment mechanism. The contribution and reputation value of the data supply nodes are updated. The reputation value is stored in the consortium blockchain reputation node, and rewards are issued simultaneously; Step 6: Model Aggregation and Iterative Training: After the consortium blockchain verification nodes complete their verification, the edge aggregation nodes use the FedAvg algorithm to aggregate the local models and generate a new global model. This global model is verified in the consortium blockchain through the PBFT consensus mechanism and then broadcast to all participating nodes. Each participating node downloads the latest global model and continues the next round of training until the model converges. During the aggregation process, a differential privacy mechanism is applied to protect the privacy of the participants' intermediate gradients.

2. The medical data federated learning sharing incentive method based on Stackelberg-Bayesian game theory according to claim 1, characterized in that, The IoMT data proxy nodes (DA) constitute a set of nodes. Where N is the total number of IoMT nodes, i.e., the total number of clients participating in federated learning, each Holding local datasets K is a variable index representing the "Kth" node in the set. This represents the input of the j-th sample. For the corresponding prediction results, The number of samples, and each Maintaining a weighted vector To store model parameters.

3. The medical data federated learning sharing incentive method based on Stackelberg-Bayesian game theory according to claim 1, characterized in that, The mobile edge computing server, acting as an edge aggregation node, maintains global model parameters W. Its optimization objective is to minimize the weighted average loss of DA across all IoMT data proxy nodes, with the weighting coefficients being... Number of samples The specific optimization objective satisfies the formula: ,in It is a weight vector; This represents the number of samples, where K represents the Kth node, and the superscript N represents the total number of IoMT nodes. yes The local training loss function is given by the following formula: ; Where j is the index of the data sample in the local dataset, representing the nth data record possessed by that node, and the value of j ranges from 1 to... That is, the node The total number of local samples; T represents "matrix transpose".

4. The medical data federated learning and sharing incentive method based on Stackelberg-Bayesian game theory according to claim 1, characterized in that, The update formula for the local model training mentioned in step 4 is: Where λ is the learning rate, and t represents the current federated learning communication round index, which is an integer, typically ranging from 0 (initial state to T-1) or 1 to T, where T is the preset maximum convergence round; in the formula, it is used to mark the evolution of model parameters over time; after the model training is completed... Differential privacy dynamics (DP) is applied to local parameters to prevent reconstruction attacks and model inference attacks. The specific formula is as follows: ,in Privacy Budget Decide; Local loss function, while Represents gradient, It is the gradient vector of the loss function, and its dimension is the same as that of the model parameters. same.

5. The medical data federated learning sharing incentive method based on Stackelberg-Bayesian game theory according to claim 1, characterized in that, The update formula for global model aggregation in step 6 is: ,in For each data supply node Privacy-protected local model parameters For the sample size is each The number of samples is N, where N is the total number of IoMT nodes and t is the communication round index of the current federated learning.

6. The medical data federated learning sharing incentive method based on Stackelberg-Bayesian game theory according to claim 1, characterized in that, The incentive policy is formulated based on the Stackelberg-Bayesian game theory framework, with data demanders... As a leader, the data supply node As followers, both parties aim to maximize their own utility functions; The data supply node The utility function satisfies: ,in For node i, a valid payment is made per unit. The amount of data contributed to node i. Let i be the total cost of node i; The data demander The utility function satisfies: ,in For data users to extract value from data sharing, For the total payment to node i, The amount of data contributed to node i. ;in, This is the data quality coefficient for node i. The accuracy of the data, the quality of the annotation, or the completeness of the medical records are determined by... express; It is the data contribution of node i; It is the "unit contribution reward coefficient" or "payment strategy" set by the data demander for node i.

7. The medical data federated learning sharing incentive method based on Stackelberg-Bayesian game theory according to claim 6, characterized in that, The data supply node Total cost ; It consists of computing energy consumption, communication energy consumption, encryption energy consumption, privacy costs, and latency costs, satisfying the formula: ; in, To calculate energy consumption, satisfy ; It is the energy consumption coefficient per node. It is the amount of contribution completed. The required computation is often approximated as follows: ; For communication energy consumption, in, For upload time, Transmission power; For encryption power consumption, , Encryption energy consumption per unit For the amount of data uploaded; To satisfy privacy costs , The level of privacy sensitivity depends on the privacy budget. ; To meet the delay cost, , This represents the delay cost coefficient.

8. The medical data federated learning sharing incentive method based on Stackelberg-Bayesian game theory according to claim 1, characterized in that, The incentive policies include two types: a unified reward mechanism (UIM) and a differentiated reward mechanism (DIM). In the unified reward mechanism, the effective payment per unit satisfies the formula: ,in for A unified basic unit price for decision-making. For adjustment coefficients, qi represents data quality; data supply nodes The total reward satisfies the formula: ; In the aforementioned differentiated reward mechanism, the effective payment per unit is for the data demander. For each data supply node Customized exclusive unit price Data supply nodes The total reward satisfies the formula: .

9. The medical data federated learning sharing incentive method based on Stackelberg-Bayesian game theory according to claim 11, characterized in that, The Stackelberg-Bayesian game has a unique equilibrium solution, specifically: Follower subgame equilibrium: Given an incentive price, each data supply node... The optimal data contribution satisfies the formula: ,in The linear cost coefficient is... Here, represents the coefficient of the quadratic term related to the load, and qi represents the data quality. for A unified basic unit price for decision-making; Leader Game Equilibrium: Data Demander Based on data supply nodes To find the optimal response, formulate the optimal reward strategy r∗ to maximize one's expected utility.

10. The medical data federated learning sharing incentive method based on Stackelberg-Bayesian game theory according to claim 1, characterized in that, The method also includes a federated learning training process based on the Stackelberg-Bayesian incentive mechanism, the specific process of which is as follows: Input: Initial model Maximum number of iterations T, learning rate η, privacy parameters; Output: Final Model ; Initialization: Consortium blockchain launch task T and initial model ; For t = 1, 2, …, T, perform the following steps: Phase 1: Leader's Incentive Decision-Making: Data Demanders Estimate the follower type distribution parameters Hi and Ai, where Hi = E1 / βi and Ai = Eαi / βi; If a unified reward mechanism (UIM) is adopted, the optimal scalar reward r* = argmaxEURUIM is calculated. If a differentiated reward mechanism (DIM) is adopted, the optimal vector reward r* = argmaxEURDIM is calculated; The broadcast reward policy r* and the current global model Wt-1; Phase Two: Follower Response and Local Training Each IoMT data broker node (DAi) performs the following operations in parallel: Minimize the cost coefficient βi by selecting the mode (local training or MEC-assisted training); Calculate the optimal data contribution xi*=max{0,rqi−αi / βi}; If xi*>0, perform local training: ωit←ωit-1-η∇Fi; Apply differential privacy mechanism: ωit←ωit+N0,σ2I; Upload the processed local model parameters ωit and contribution proof; Aggregation and reward distribution; Mobile edge computing servers / consortium blockchains verify the contributions and data quality of each node (qi). The FedAvg algorithm is used to aggregate the global model: Wt←DkDtotalωkt; Smart contract execution rewards payments, paying Pi to active data provider nodes (DAi); If the model accuracy converges, terminate the iteration; Return the final model Wt.