Wireless computing power network federal split learning incentive method

CN122621952APending Publication Date: 2026-08-21GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610414166.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-31
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

在这种机制下,模型训练任务不再被强制绑定在数据本地,而是允许数据节点将计算密集型的训练任务,在保障原始数据隐私的前提下(如仅传输中间层激活值),卸载给邻近的强算力节点执行,从而实现“数据所有权”与“计算执行权”的有效分离与解耦在上述的联邦多任务学习场景中,网络中的节点参与到训练当中必然会占用其计算和通信资源,如果没有合理的激励机制,会导致节点不愿意参与到训练当中,导致训练任务失败

Benefits of technology

本发明通过构建面向无线算力网络的联邦分裂学习双边协同架构,创新性地引入数据采购与算力外包的双边合同市场机制,利用深度强化学习实现复杂动态环境下的合同参数自适应生成,有效解决了资源受限节点的训练参与难题,实现了数据资源与计算资源的最优配置,显著提升了无线算力网络环境下联邦多任务学习的参与意愿与训练效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122621952A_ABST
    Figure CN122621952A_ABST
Patent Text Reader

Abstract

The application discloses a wireless computing power network federal split learning incentive method, which is applied to a task publisher unmanned aerial vehicle, and the method comprises the following steps: acquiring state information of a current network environment, including the states of data nodes and computing power nodes and communication link states; inputting the state information into a pre-trained deep reinforcement learning model, wherein the model outputs actions containing a data contribution contract, a computing power service contract and a bandwidth allocation scheme; publishing the contracts to corresponding nodes to encourage the nodes to participate in federal split learning training; and receiving and aggregating model parameters uploaded by the nodes participating in training according to the contracts to update a global model. Through the design of a bilateral contract mechanism, in combination with a task offloading mode of federal split learning and a proximal policy optimization algorithm, the technical problems of resource and data mismatch in a computing power network, low willingness of nodes to participate and difficulty in adaptive incentive strategies in a dynamic environment are solved, and the method is suitable for efficient federal multi-task learning scenes in a wireless computing power network environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of federated learning technology, and in particular to a method for incentivizing federated split learning in wireless computing networks. Background Technology

[0002] With the explosive growth of 6G communication, the Internet of Things (IoT), and artificial intelligence technologies, massive computing demands and data resources have emerged in edge network environments. Traditional cloud computing models, due to their high transmission latency and bandwidth pressure, are no longer sufficient to meet the real-time requirements of applications such as autonomous driving, smart cities, and industrial automation. Against this backdrop, Wireless Computing Network (WCPN) has emerged as a new network architecture. By deeply integrating multi-dimensional heterogeneous computing resources from the cloud, edge, and endpoint, it achieves virtualization and ubiquitous scheduling of computing resources, aiming to break down physical device boundaries and provide flexible, unified, and efficient computing services for the entire network. However, the openness and distributed nature of WCPN nodes present serious challenges to data privacy and security trust in practical deployments. Because computing nodes typically belong to different stakeholders, widespread "data silos" have emerged to protect user privacy and trade secrets, making traditional centralized machine learning methods unfeasible. To address data privacy issues, Federated Learning, a distributed paradigm where "the model moves while the data remains stationary," has been introduced into edge computing. However, traditional Federated Learning usually assumes that all participating nodes jointly train a globally shared model, an assumption that shows limitations in the highly heterogeneous environment of WCPN. In real-world computing networks, the physical environments of different edge nodes, the distribution of collected data (Non-IID), and the specific tasks they need to perform often differ significantly. Forcing all nodes to train the same model not only leads to a decline in the performance of local personalized tasks but also causes a huge waste of communication resources. Therefore, evolving the paradigm into Federated Multi-Task Learning (FMTL) is particularly necessary. It allows nodes to collaboratively share general knowledge while optimizing models for their specific tasks, thus better adapting to the heterogeneity and task diversity of ubiquitous devices in WCPN.

[0003] Among the many derivative forms of wireless computing networks, unmanned aerial vehicles (UAVs), with their high mobility, line-of-sight transmission advantages, and flexible deployment capabilities, are gradually transforming from simple data relays into key hubs for building integrated air-ground computing networks. This provides a new spatial dimension for the deployment of federated multi-task learning. In this more complex three-dimensional network architecture, assigning UAVs the role of "mobile aggregators and publishers" is particularly crucial. That is, UAVs directly act as task initiators and local parameter aggregation centers near the ground, dynamically forming collaborative learning clusters based on the real-time needs of ground users. However, if the traditional centralized federated learning topology continues to be used, relying on a single ground base station or central UAV to coordinate all tasks, it will be inadequate when facing the inherent high dynamism, frequent topology changes, and limited wireless backhaul bandwidth of UAV networks. The communication bottlenecks of the central node, the risk of single-point failures, and the high latency caused by long-distance transmission will severely restrict the convergence efficiency and real-time performance of multi-task learning. Therefore, pushing federated multi-task learning further towards a "decentralized" paradigm becomes an inevitable choice to solve the above bottlenecks.

[0004] Furthermore, in practical ubiquitous sensing networks, nodes holding high-value data are often limited by extremely limited battery capacity and chip computing power; meanwhile, there are also a large number of nodes with abundant computing power but lacking specific task data. If the rigid principle of "whoever has the data trains" in traditional federated learning is adhered to, these computing-limited data nodes will be forced to withdraw because they cannot bear the local training load, resulting in the loss of key samples and severely restricting the generalization performance of the global model. To break this deadlock and achieve optimal allocation of data and computing resources, a task offloading method based on split learning is introduced. Under this mechanism, model training tasks are no longer forcibly bound to the local data, but are allowed to offload computationally intensive training tasks to nearby high-computing-power nodes, while ensuring the privacy of the original data (such as only transmitting intermediate layer activation values), thereby achieving effective separation and decoupling of "data ownership" and "computation execution rights." In the above federated multi-task learning scenario, the participation of nodes in the network in training will inevitably occupy their computing and communication resources. Without a reasonable incentive mechanism, nodes will be unwilling to participate in training, leading to training task failure. However, this tripartite collaboration model significantly increases the complexity of incentive mechanism design: as the initiator and ultimate beneficiary of the mission, the drone not only needs to incentivize data owners to contribute high-quality data, but also needs to incentivize computing power providers to expend energy to execute others' training tasks. Drones face the challenge of dual information asymmetry—they cannot ascertain the true quality of the data held by data nodes, nor can they ascertain the true computational cost and energy consumption of computing power nodes. Therefore, how to design separate contracts for data nodes and computing power nodes based on their different attributes in an incomplete information environment to incentivize different nodes to participate in training tasks is a pressing problem. Summary of the Invention

[0005] Based on this, the objective of this invention is to address the aforementioned technical problems by providing a wireless computing power network federated split learning incentive method that can design contracts for different attributes of data nodes and computing power nodes in an incomplete information environment to incentivize different nodes to participate in training tasks and improve training efficiency.

[0006] To achieve the above-mentioned objectives, this application provides a method for incentivizing federated split learning in wireless computing networks, comprising: In response to the triggering of the federated learning task, the status information of the current UAV and its corresponding data nodes and computing power nodes constituted subnet network environment is obtained. The status information includes the status of data nodes, the status of computing power nodes and the status of communication links. The state information is input into a pre-trained deep reinforcement learning model, and the model outputs a set of actions, including a data contribution contract generated for at least one data node, a computing power service contract generated for at least one computing power node, and bandwidth resources allocated to the data node and the computing power node. The generated data contribution contract and computing power service contract are published to the corresponding nodes to incentivize the corresponding nodes to participate in federated split learning training and upload model parameters after the training iteration is completed. Receive and aggregate model parameters uploaded by the data nodes and computing power nodes that participate in training according to the corresponding contracts, in order to update the global model.

[0007] Preferably, the data contribution contract is represented as follows: Where Rn is the incentive reward provided for the nth data node, Dn is the amount of data it is required to contribute, and DN represents the total number of data nodes. The design satisfies individual rationality constraints and incentive compatibility constraints to ensure that data nodes select a contract tailored to their own type. The computing power service contract is represented as follows: , where Rm is the incentive reward provided for the m-th computing power node, αm is the task offloading policy allocated to the computing power node, fm is the CPU processing power required to be provided, and CN represents the total computing power nodes; The design of the data contribution contract and computing power service contract satisfies individual rationality constraints and incentive compatibility constraints to ensure that nodes select a contract tailored to their own type. Individual rationality constraints mean that the net utility obtained by each node after selecting a contract designed for it must be greater than or equal to zero. Incentive compatibility constraints mean that the utility obtained by the cluster when selecting a contract of its own type must be greater than or equal to the utility it could obtain by disguising itself as other energy consumption types and selecting other contracts.

[0008] Preferably, the federal split learning training process includes: The global model parameters are sent to the data nodes participating in the training, and the data nodes are instructed to determine the splitting layer of the model according to their own preferences, dividing the model into a head network, a middle network, and a tail network. The data node is instructed to train the head network using local data and to offload the shredded data generated by the computation to the computing power node matched according to the task offloading strategy. The computing node is instructed to receive the shredded data, complete the forward and backward propagation of the intermediate network, and send the calculated gradient information back to the data node. The data nodes are instructed to complete the backpropagation of the tail network and update the model parameters based on the returned gradient information.

[0009] Preferably, the optimization objective of the federated split learning training is to maximize the total utility of the UAV, and the optimization objective is expressed as:

[0010] Where p represents the drone; N is the number of data nodes; and M is the number of computing nodes. This represents the bandwidth vector allocated to data node n; This represents the bandwidth vector allocated to computing node m; This indicates that data node n provides the amount of data D. n The utility of the publisher; This represents the utility of the publisher when the computing node m executes the corresponding task; The total energy consumption of the drone is represented by T; T represents the number of local iterations. This represents the additional communication energy consumed by the data node in receiving and transmitting the model to the drone and unloading the model to the computing node; C n、 C k λ represents the energy consumption per unit of data in a data node, where n and k represent different data nodes in the network; λ represents the cost sensitivity coefficient. This indicates the specific energy consumption margin for each computing node; This represents the probability that each node in the data node belongs to a certain type; This represents the probability that each node in the computing power nodes belongs to a certain type; R represents the energy consumption of a computing node. max This represents the maximum value of the incentive capability provided by the node. f max This represents the maximum CPU processing power of the computing node; , This indicates the weights of data nodes and computing nodes.

[0011] Preferably, the states of the data nodes, computing nodes, and communication links respectively include: state information of individual rational constraints and incentive compatibility constraints of the data nodes and computing nodes, as well as the signal-to-noise ratio between the nodes and the UAV.

[0012] Preferably, the aggregation of model parameters uploaded by nodes participating in training according to the contract includes: Receive the updated parameters of the head network and tail network of the data nodes in the subnet, as well as the updated parameters of the intermediate network of the computing power nodes; Based on the amount of data contributed by each data node, all received model parameters are weighted and aggregated to update the global model of the current subnet.

[0013] Preferably, the method further includes: Decentralized federated multi-task learning is performed with other drone subnets, specifically: receiving shared layer model parameters published by other drones, calculating the similarity between the local shared layer model and the received other shared layer models; determining aggregation weights based on the similarity, and using the aggregation weights to aggregate the shared layer models of other drones to update the local shared layer model.

[0014] Preferably, the pre-trained deep reinforcement learning model is a proximal policy optimization algorithm model; the training process of the model includes: Construct an intelligent agent, using a drone as the intelligent agent, and set up a state space, action space, and reward function; The agent is in the current state s t Below, based on the action network, output action a. t and obtain the action to be performed, a. t The next state s t+1 and reward r t ; Experience (s) t , a t , r t , s t+1 Store in the experience pool; Experience is sampled from the experience pool, and the parameters of the action network and the critic network are updated using the dominance function and the pruning function to maximize the long-term cumulative reward until the reward value converges, thus obtaining a trained model.

[0015] Preferably, the state space is represented as:

[0016] in, Let N represent the individual rationality constraints and incentive compatibility constraints of N data nodes at time t, respectively. , ; Let represent the individual rationality constraints and incentive compatibility constraints of the N computing nodes at time t, respectively. , ; This represents the signal-to-noise ratio from the data node to the drone at time t; This represents the signal-to-noise ratio from the computing node to the drone at time t; The action space is specifically represented as follows:

[0017] in, This represents the set of data contribution contracts generated by the UAV for all data nodes at time t, where This represents the data contribution contract corresponding to each node; This represents the set of computing power service contracts generated by the drone for all computing power nodes at time t, where This represents the computing power service contract corresponding to each node; This represents the bandwidth allocation vector that the UAV assigns to all data nodes at time t. This represents the bandwidth vector allocated by the drone to all computing nodes at time t.

[0018] Preferably, the reward function is set as follows: If the action a t If the generated contract and bandwidth allocation scheme satisfy the preset individual rationality constraints, incentive compatibility constraints, total incentive budget constraints, and node processing capacity constraints, then the reward function value is the total utility of the UAV calculated according to the contract and bandwidth allocation scheme. If the action a t If any of the constraints are not met, the reward function value is a preset penalty value.

[0019] Compared with the prior art, the beneficial effects of this invention are: Beneficial effects of this invention: This invention constructs a federated split-learning bilateral collaborative architecture for wireless computing networks, innovatively introducing a bilateral contract market mechanism for data procurement and computing power outsourcing. It utilizes deep reinforcement learning to achieve adaptive generation of contract parameters in complex dynamic environments, effectively solving the training participation problem of resource-constrained nodes, realizing the optimal allocation of data and computing resources, and significantly improving the willingness to participate and the training efficiency of federated multi-task learning in wireless computing network environments. Attached Figure Description

[0020] Figure 1 A flowchart illustrating the steps of a federated split learning incentive method for wireless computing networks; Figure 2This is a schematic diagram of a decentralized federated multitasking learning scenario in a wireless computing network, as shown in one embodiment. Figure 3 This is a flowchart illustrating the task unloading process based on federated split learning. Figure 4 A flowchart illustrating the near-end strategy optimization method; Figure 5 This is a flowchart illustrating the steps involved in contract design and bandwidth allocation based on a near-end strategy optimization method. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention. The following embodiments are used to illustrate the invention but are not intended to limit its scope.

[0022] Example 1 Embodiment 1 of this application provides a federated split learning incentive method for wireless computing power networks, such as... Figure 1 As shown, it includes the following steps: S1: In response to the triggering of the federated learning task, obtain the status information of the current UAV and its corresponding data nodes and computing power nodes forming the subnet network environment. The status information includes the status of data nodes, the status of computing power nodes and the status of communication links. S2: Input the state information into a pre-trained deep reinforcement learning model, and the model outputs a set of actions, including a data contribution contract generated for at least one data node, a computing power service contract generated for at least one computing power node, and bandwidth resources allocated to the data node and the computing power node. S3: Publish the generated data contribution contract and computing power service contract to the corresponding nodes to incentivize the corresponding nodes to participate in federated split learning training and upload model parameters after the training iteration is completed; S4: Receive and aggregate the model parameters uploaded by the data nodes and computing power nodes that participate in the training according to the corresponding contracts, in order to update the global model.

[0023] The following is a detailed explanation of the steps described above: 1. Federal Multitasking Learning: like Figure 2As shown, the decentralized federated multi-task learning in the computing power network designed in this invention mainly consists of drones (task publishers), data nodes, and computing power nodes. We divide the network of drones and their corresponding data and computing power nodes into subnets. Each subnet is considered an independent federated learning task. Within a subnet, the drone designs a contract and passes the model to the nodes for multiple local training iterations. Then, the drone collects the trained models and aggregates them. Between different subnets, drones can share general knowledge to achieve decentralized federated multi-task learning. Each subnet repeats the above steps until the global accuracy reaches a certain threshold, finally training a final model that balances individual performance with global knowledge and has strong generalization ability.

[0024] Based on the scenario above, the computing power network contains P subnetworks, each P ∈ P containing multiple data nodes and computing power nodes. A drone P generates a federated training task, requiring N data nodes and M computing power nodes within its subnetwork to participate in the training. The drone designs two contracts and publishes them to the subnetworks. Clients within the subnetworks that possess data autonomously choose whether to participate in the training based on their own data type. Specifically, we assume that each client n ∈ N possesses a set of local training data, denoted as D. n = {( x n1 , y n1 ), …, ( x n D n , y n D n The data size is represented by D. n express.

[0025] (1) Training of the local model Assuming data nodes choose to participate in model training, their computationally intensive tasks are offloaded through federated splitting due to limitations in their own computing power. The global model parameters for each drone are... At the start of each global iteration, the drone passes the global model to the data nodes that received the contract. The data nodes train their own models using their data and select appropriate split layers based on their own circumstances, then offload the task. Different data nodes have different preferences for split layers. Taking node n∈N as an example, let's assume its split layer preference is... Furthermore, the intermediate steps of the model are offloaded to computing nodes for training, specifically the model... Divided into , and ,in For the head network, For intermediate networks, This is the tail network; considering that the purpose of the tail network is to prevent label leakage, its segmentation layer size is fixed, that is... It only affects the size of the head network and the intermediate network. Therefore, for node n, the model training process is as follows: (1.1) Forward propagation phase: Data nodes train based on their own data Model, and intermediate activation values The training process, which is passed to the computing power nodes connected to it, can be represented as follows:

[0026] Among them, the function This represents the forward computation process of the head neural network; This represents the calculated activation value; This represents batch training data; the computing node receives it. Then, it is input into the intermediate part of the model. The forward propagation is performed to obtain the activation values. The information is sent to the client, and the training process is represented as follows:

[0027] Among them, the function This represents the forward computation process of the intermediate neural network; similarly, the computing nodes will activate values. After being sent to the data nodes, the data nodes train the tail model and obtain the mapping values ​​of the tail network. And calculate its loss. The training process is as follows:

[0028] Among them, the function This represents the forward computation process of the tail neural network; Represented as the loss value, This indicates the process of calculating the loss.

[0029] (1.2) Backpropagation phase The tail network calculates gradient information based on the calculated loss value, as well as the gradient of the loss with respect to the input, and passes it to the intermediate network.

[0030]

[0031] express gradient information, This represents the transmission gradient of the intermediate network. Represented as backpropagation calculation of the tail network, based on the ladder information, the formula for updating the tail network parameters of the data nodes is: , This represents the learning rate. express relatively The gradient.

[0032] The computing nodes receive gradient data from the split points. The algorithm calculates and updates the gradient changes of the intermediate network, and simultaneously transmits the calculated gradient information back to the data nodes so that they can update the gradient of the head network.

[0033]

[0034] express gradient information, This represents the transmission gradient of the head network. This represents the backpropagation calculation in the intermediate network. Based on the gradient information, the parameters of the computing nodes are updated as follows: . express relatively The gradient.

[0035] Data nodes receive gradient data Then, the local gradient is calculated and the model parameters are updated according to the chain rule:

[0036] in Represented as gradient information, This is represented as backpropagation computation in the head network, where data nodes update their parameters based on the calculated gradient information: .

[0037] (2) Global model aggregation This completes one batch of training. Data nodes and computing nodes repeat this process until a certain number of training rounds have been completed. Afterward, the data nodes and computing nodes pass their model parameters to the publisher (UAV). The publisher receives the models from both and performs aggregation and updates, thus completing one round of global iterative training. The publisher's aggregation formula is:

[0038] in, It represents the total amount of data in all subnetworks.

[0039] Furthermore, for different UAVs, due to the differences in data between their subnets, directly aggregating models from different subnets can actually degrade model performance. Therefore, we consider treating different subnets as training different but related federated learning tasks. We further divide the UAV model into a shared layer and a personalized layer; UAVs from different subnets achieve federated multi-task learning by aggregating the shared parts of other UAV models (shared layers), thereby improving the overall model performance. Specifically, UAV p first receives the shared model from other UAVs. The aggregation weights are calculated by using the correlation between the shared model and its own model, and the specific formula is as follows:

[0040] in The similarity weights of different drones can be expressed by the following formula:

[0041] It is a hyperparameter used to adjust the sensitivity of similarity. Cosine similarity is used to measure the degree of similarity between the model parameters of two drones.

[0042] 2. Contract Design: To incentivize different types of data nodes and computing power nodes to participate in training, task publishers design differentiated contracts for different types of nodes. For data nodes, the designed contract... This is used to incentivize data node n to participate in the federated training task. R... n D is the incentive reward provided for the nth data node. n This refers to the required amount of data to be contributed; DN represents the total number of data nodes. For computing power nodes, the design contract... This is used to incentivize computing node m to provide computing power services. Where R... m This is the incentive reward provided for the m-th computing node, α. m It is the task offloading strategy assigned to this computing node. f mThis requires the CPU processing power provided, with CN representing the total computing power nodes. The contract design process must satisfy Individual Rationality (IR) and Incentive Compatibility (IC) constraints. The IR constraint states that after selecting a contract designed for it, each node must obtain a net utility greater than or equal to zero. Otherwise, the node will have no incentive to participate in training. The IC constraint states that when the cluster selects a contract of its own type, the utility it obtains must be greater than or equal to the utility it could obtain by "disguising" itself as other energy consumption types and selecting other contracts. This constraint ensures that each node honestly selects a contract tailored to its needs, preventing "adverse selection."

[0043] like Figure 3 The diagram illustrates the task training process based on split learning. Data nodes determine the split layer based on their privacy requirements and assessment of their own computing resources. And offload the intermediate computing tasks to the computing nodes, and through fixed This method prevents node label leakage. Throughout the training process, data nodes and computing nodes train and update different parts of the network by transmitting shredded data and gradient data. This process offloads the heavy computational burden from the data nodes while ensuring their data privacy, thereby reducing their training load and improving the publisher's effective data utilization.

[0044] (1) Energy consumption model: After selecting a contract, each node needs to consume its own computing resources to participate in task training. The energy consumption of the cluster is divided into local computing energy consumption, cluster communication energy consumption, and drone energy consumption.

[0045] (1.1) Data node energy consumption: Considering that the splitting layer of the model affects node privacy and computational burden, the choice of splitting layer should be different for each data node. Therefore, the splitting layers of the head network and intermediate network of data node n are represented as follows: The split layer of the tail network is represented as The energy consumption of data nodes is divided into training energy consumption and communication energy consumption, and the latter is compared with the split layer. The relevant energy consumption is expressed as and The following formula is defined:

[0046] This represents the computational cost per unit of data. Indicates the time required to represent a unit of data; Represented as data nodes relative to the split layer The computational load is an increasing function of it. The computational cost per unit of data in the tail network is considered. Since the tail network primarily serves to prevent tag leakage, the tail network size for each data node is fixed and standardized. This indicates its split point. Among them, , This represents the amount of computation required per unit of data for the entire network. This is represented as the CPU computing power of a node; Represented as hardware coefficients.

[0047] The energy consumption and latency of communication are defined by the following formulas:

[0048] in, This is expressed as the energy consumption per unit of data communication between data nodes and computing nodes; This indicates the communication time for the corresponding unit of data; where, This represents the transmission and reception power from the data node to the computing node. This represents the transmission and reception rates between data nodes and computing nodes. Let n be the bandwidth and signal-to-noise ratio from node n to node m. Let m be the bandwidth and signal-to-noise ratio from node m to node n; This represents the size of the shredded data produced per unit of data in the head network. This represents the size of the unit data gradient information of the tail network. This represents the size of the shredded data passed to the tail network unit. This represents the magnitude of the gradient information passed to the head network unit data.

[0049] In addition, the data node requires additional communication energy to receive and transmit the model to the drone and to unload the model to the computing node, which can be represented as:

[0050] in, Indicates additional transmission power consumption; Indicates additional transmission time; This is expressed as the amount of data required by the transmission model W. These represent the transmission volumes for the corresponding header, middle, and tail networks, respectively. These represent the data node's receiving power and transmitting power to the UAV, respectively. This represents the data transmission rate from the data node to the drone. The bandwidth allocated to the node for the drone. Its corresponding signal-to-noise ratio; Let p be the downlink transmission rate of the drone. For the bandwidth of the drone, For the corresponding signal-to-noise ratio; according to the formula above, using Represented as split layer Energy consumption per unit of data:

[0051] Therefore, when the amount of data contributed by the data node is D n When the total energy consumption and total delay are equal, they can be expressed as:

[0052] in, Indicates total energy consumption; Indicates the total latency; T represents the number of local iterations. This is expressed as the energy consumption per unit of data when the split layer is 1. This represents the energy consumption of a data node receiving training data.

[0053] (1.2) Energy consumption of computing nodes: Computing nodes can receive tasks unloaded from other data nodes. The computational workload of the central network of data node n is defined as follows: ,in Assume the node's unloading strategy is as follows: Considering the unique characteristic of split learning training requiring sequential execution, it is assumed that each data node can only offload tasks to a single computing node. .

[0054] Since the server's computing power requirements are based on the corresponding data contract, the computing load of the computing nodes will also change with the corresponding offload rate. Assume that the offload requirement received by computing node m is... ,in, Let the unloading decision of data node n on computing node m be represented by the corresponding computational cost. ,use This indicates that as the amount of data and The changes in the amount of unloaded tasks affect the computation time and energy consumption of the computing nodes.

[0055] The computing task size of the computing power node is... . This represents the CPU processing power of that node; This is represented as a hardware coefficient. Clearly, for different unloading decisions... The workload of each computing node will vary, and its computational energy consumption will increase as the amount of unloaded tasks increases. Furthermore, at the start of each iteration, the computing node may receive models unloaded from data nodes, and after completing the training task, it needs to transmit the model to the drone. This transmission consumes the computing node's communication energy, which can be expressed as follows:

[0056] Indicates communication energy consumption. To receive the received power of other data nodes, This refers to the transmission power. This represents the total communication energy consumption of data nodes and computing nodes. Similarly, with... The changes in these parameters will also affect the total communication energy consumption of data nodes and computing nodes, specifically as follows:

[0057] in, This is expressed as the corresponding total transmission energy consumption; This represents the total transmission delay. Since the nodes use orthogonal frequency division multiple access (OFDMA) for transmission, the maximum value of the delay is taken.

[0058] In addition, the communication time and energy consumption of the computing node transmitting the trained intermediate model to the UAV are as follows:

[0059] in, Indicates transmission energy consumption; Indicates transmission delay; This represents the amount of data required by the transmission model. This represents the transmission rate from the computing node to the drone. The bandwidth allocated to the node for the drone. Its corresponding signal-to-noise ratio; This represents the transmission power from the computing node to the drone; according to the formula above, the total energy consumption and total latency of the computing node are respectively:

[0060] (1.3) Energy consumption of UAVs: The communication energy consumed by the drone in transmitting the model to the data node and the receiving node can be expressed as:

[0061] in, Indicates transmission power consumption. This refers to downlink transmission power; The energy consumption is for uplink reception. Simultaneously, the energy consumption for communication with the computing node is:

[0062] in, Indicates the received transmission power consumption. For the transmission power received by the computing node, The time of receipt.

[0063] The total communication energy consumption is:

[0064] In addition, the drone aggregation model also needs to generate aggregation energy consumption. And the hovering energy consumption that drones generate during training. The two are represented as follows:

[0065] in, This indicates the CPU computing power of the drone. Its hardware parameters; The hovering energy consumption required per unit time. The corresponding hovering time is related to the training time of the data node and the computing power node, and its expression is:

[0066] in, This represents the total latency of the computing nodes. This represents the computation latency of the data node.

[0067] This represents the additional transmission latency of the data node. The latency varies between different nodes, so the maximum value is taken.

[0068] The total energy consumption of the drone is:

[0069] (2) Utility model: (2.1) Data node utility For data nodes, drones can incentivize them to provide data through contract design; for computing power nodes, drones can provide incentives to incentivize them to provide computing power. For drones, contracts can be designed based on the types of computing power nodes and data nodes. Since each data node is part of the split layer... People have different preferences, and the energy consumption per unit of data processing varies for split layers with different preferences. Therefore, C is used. nThis represents the energy consumption per unit of data from a data node, and is used to indicate the node's willingness to participate. Clearly, lower energy consumption indicates a greater willingness to participate in training and provide data. This can be represented as:

[0070] The Task Publisher (UAV) cannot directly know the specific intentions of each data node in the split layer. And its corresponding energy consumption per unit of data C n However, the probability of each cluster belonging to a certain type can be inferred empirically. This represents the probability that each node belongs to a certain type. .

[0071] For data node n, the publisher designs a contract. ,in This represents the incentive provided by the publisher to data node n. The amount of data that the publisher requests the data nodes to provide.

[0072] The data nodes select the corresponding contract and execute training; their utility can be represented as follows:

[0073] Set the amount of data node n to provide At that time, the publisher's corresponding level of satisfaction is expressed as The publisher's utility can then be expressed as:

[0074] (2.2) Utility of computing nodes: For computing nodes, the User Availability (UAV) can design another set of contracts to incentivize their participation in training. The willingness of a computing node to participate can be expressed as the relative amount of its remaining energy. Specifically, the greater the remaining energy of a computing node, the more willing it is to participate in training. Therefore, we define the type of computing node based on its remaining energy consumption as follows:

[0075] Similarly, the User-Agent Publisher (UAV) cannot directly know the specific energy reserve of each computing node. However, the probability of each cluster belonging to a certain type can be inferred empirically. This represents the probability that each node belongs to a certain type. Since all the computational tasks that a computing node needs to handle are the same as those generated by the data node's contract, to avoid designing computing node contracts that exceed or fall short of the required computing power, the computational load is directly represented by an unloaded object. Therefore, the computing node's contract can be represented as... ,in This represents the reward offered by the publisher to the computing power nodes, while This represents the task that is offloaded to the computing node. ,in This indicates whether to offload the computing task to the computing node. This refers to the processing power required from the nodes. As the offloading strategy changes, the computational load on the computing nodes will also change.

[0076] The publisher provides a contract The corresponding computing power node executes the task, and the utility of the corresponding computing power node is represented as follows:

[0077] in, , which represents the energy consumption of a computing node; This is expressed as the corresponding communication power consumption; Expressed as relative remaining energy consumption, and using To represent cost sensitivity factors, This is expressed as a cost sensitivity coefficient. A higher sensitivity coefficient indicates that computing nodes are less willing to lose energy, thus requiring more compensation. When the amount of data provided by data node n is set, the publisher's satisfaction formula can be expressed as: , This indicates that the higher the workload processed by the computing nodes, the higher the utility of the publisher. However, as the workload increases, the computing nodes also require more time. Adjustment The need to improve utility. Let be the corresponding weighting coefficient. Then the publisher's utility can be expressed as:

[0078] 3. Optimization Objectives: In summary, at the start of training, the UAV first employs two incentive mechanisms and transmits them to the data nodes and computing nodes within the subnet. Then, different nodes choose whether to participate in and complete the task based on their own resource availability. In this round of training, depending on the subnet selection, the current optimization objective is to maximize the UAV's total utility, which can be expressed as:

[0079] Where p represents the drone; N is the number of data nodes; and M is the number of computing nodes. This represents the bandwidth vector allocated to data node n; This represents the bandwidth vector allocated to computing node m; This indicates that data node n provides the amount of data D. n The utility of the publisher; This represents the utility of the publisher when the computing node m executes the corresponding task; The total energy consumption of the drone is represented by T; T represents the number of local iterations. This represents the additional communication energy consumed by the data node in receiving and transmitting the model to the drone and unloading the model to the computing node; C n、 C k λ represents the energy consumption per unit of data in a data node, where n and k represent different data nodes in the network; λ represents the cost sensitivity coefficient. This indicates the specific energy consumption margin for each computing node; This represents the probability that each node in the data node belongs to a certain type; This represents the probability that each node in the computing power nodes belongs to a certain type; R represents the energy consumption of a computing node. max This represents the maximum value of the incentive capability provided by the node. f max This represents the maximum CPU processing power of the computing node; , This indicates the weights of data nodes and computing nodes.

[0080] Among them, constraints C1 and C3 represent individual rationality constraints for data nodes or computing power nodes, indicating that the utility of each node cannot be less than zero when selecting a corresponding contract. Constraints C2 and C4 represent incentive compatibility constraints for data nodes or computing power nodes, indicating that when a node selects a contract of its own type, its utility should be greater than the utility of the contract it selects when disguised as another type. Constraint C5 indicates that the total incentive provided by the drone cannot exceed the maximum incentive it can provide. C6 represents the limitation of the CPU processing power of the computing power node. C7 represents the weights and limitations of data nodes and computing power nodes.

[0081] Example 2 This application's Embodiment 2, based on Embodiment 1, further illustrates the pre-trained deep reinforcement learning model in Embodiment 1, as follows: In this embodiment, the pre-trained deep reinforcement learning model is a near-end policy optimization algorithm model, as detailed below: The Proximal Policy Optimization (PPO) algorithm mainly consists of an actor network and a critic network. The actor network generates an action policy π. In time slot t, the actor network sets the system state... As input, action As output. For any given system state. Define strategy π as a mapping For the critic network, the evaluation is based on the value of the entire state, for a given state. As input, its output is The purpose of the critic network is to provide an accurate and stable global baseline for calculating the advantage function, thereby guiding the agent's action generation in a better direction. This can be defined as a mapping: .

[0082] In this system, the drone, acting as the task issuer, triggers the algorithm based on its task requirements. At time t, the agent's action network receives the system state. and generate action. At the same time, obtain the next state. and rewards Then, the experience of each agent is stored in an experience pool. During each training phase, samples are selected from the experience pool to update the parameters of the action network and the critic network until the long-term reward is maximized. Based on the above theory, the state space of reinforcement learning in this embodiment is defined as:

[0083] in, Let N represent the individual rationality constraints and incentive compatibility constraints of N data nodes at time t, respectively. , ; Let represent the individual rationality constraints and incentive compatibility constraints of the N computing nodes at time t, respectively. , ; This represents the signal-to-noise ratio from the data node to the drone at time t; This represents the signal-to-noise ratio from the computing node to the drone at time t; Action Space: Based on its corresponding state space, the UAV needs to ensure that data nodes and computing nodes meet their own utility requirements and are allocated within the limited bandwidth, thereby maximizing the utility of the task issuer. Therefore, its corresponding action space can be represented as:

[0084] in, This represents the set of data contribution contracts generated by the UAV for all data nodes at time t, where This represents the data contribution contract corresponding to each node; This represents the set of computing power service contracts generated by the drone for all computing power nodes at time t, where This represents the computing power service contract corresponding to each node; This represents the bandwidth allocation vector that the UAV assigns to all data nodes at time t. This represents the bandwidth vector allocated by the drone to all computing nodes at time t.

[0085] Reward Function: Since reinforcement learning requires a certain exploration time at the beginning, and the UAV needs to consider the impact of corresponding constraints when designing contracts and allocating bandwidth, the reward function is set in two forms: one is a penalty for the generated action space not satisfying the constraints, and the other is a reward design for satisfying the constraints. According to our set optimization objective, if the generated action space satisfies C1... Given the constraints of C7, its reward function can be expressed as: When an action does not meet the constraints, we apply a preset penalty value. U is used as the corresponding penalty. Therefore, the reward function can be set as follows:

[0086] Following the iterative steps described above, in each iteration, the agent stores experience in an experience pool until certain conditions are met. At this point, the agent retrieves experience from the pool to update the Actor network and the Critic network. The update of the Critic network can be represented as... in, The advantage function describes the advantage of a decision relative to other feasible decisions, and can be expressed as: , It is a discount factor. The single-step time error can be expressed as: An update to an Actor can be represented as: ,in The update rate of the agent's parameters is represented by the following calculation method: That is, the ratio of the probabilities of the new and old strategies. It is the dominant function. For the clipping function, Limiting to confidence intervals This ensures the adaptability and stability of the action network. The parameters of both networks are updated by optimizing these two loss functions, allowing the optimal policy to be learned, thus completing the training process. Finally, a contract is generated based on the agent's performance, and bandwidth is allocated.

[0087] Example 3 This embodiment 3 further explains the steps of contract design and bandwidth allocation based on the near-end policy optimization method of embodiment 2, such as... Figure 5 As shown, it includes: Step 1: Construct the agent network environment and initialize the system environment information, including agent initialization information and publisher utility functions.

[0088] Step 2: Construct the neural network, including the action network and the critic network, and initialize the neural network parameters, including weights, biases, learning rate, number of layers, etc.

[0089] Step 3: The agent inputs the observed state into its action network model to train the neural network and obtain the system action, namely the contract designed by the agent and the bandwidth allocation.

[0090] Step 4: The agent inputs actions into the environment to simulate the node selection process. Each node makes the most rational choice based on its own state. The agent evaluates the optimal choice of the node to calculate utility. This utility is used as the reward value for the action, and the agent obtains the next state from the environment. The parameters of the neural network are trained based on the reward value and state information.

[0091]

[0092] Step 5: Use the rewards and states obtained in the above steps to train the neural network until the reward value stabilizes. Based on the training task of the agent, train to obtain the optimal contract design and the optimal allocation scheme of bandwidth resources.

[0093] In summary, this invention provides a federated split-learning incentive method for wireless computing networks. By constructing a bilateral collaborative architecture for federated split-learning in wireless computing networks, it innovatively introduces a bilateral contract market mechanism for data procurement and computing power outsourcing. It utilizes deep reinforcement learning to achieve adaptive generation of contract parameters in complex dynamic environments, effectively solving the training participation problem of resource-constrained nodes, achieving optimal allocation of data and computing resources, significantly improving the willingness to participate in federated multi-task learning and training efficiency in wireless computing network environments, and significantly enhancing the robustness and convergence efficiency of the incentive mechanism in dynamic network environments.

[0094] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A federated learning incentive method for wireless computing networks, characterized in that, Includes the following steps: In response to the triggering of the federated learning task, the status information of the current UAV and its corresponding data nodes and computing power nodes constituted subnet network environment is obtained. The status information includes the status of data nodes, the status of computing power nodes and the status of communication links. The state information is input into a pre-trained deep reinforcement learning model, and the model outputs a set of actions, including a data contribution contract generated for at least one data node, a computing power service contract generated for at least one computing power node, and bandwidth resources allocated to the data node and the computing power node. The generated data contribution contract and computing power service contract are published to the corresponding nodes to incentivize the corresponding nodes to participate in federated split learning training and upload model parameters after the training iteration is completed. Receive and aggregate model parameters uploaded by the data nodes and computing power nodes that participate in training according to the corresponding contracts, in order to update the global model.

2. The method according to claim 1, characterized in that, The data contribution contract is represented as follows: , where R n D is the incentive reward provided for the nth data node. n It refers to the amount of data that it is required to contribute, where DN represents the total number of data nodes. The design satisfies individual rationality constraints and incentive compatibility constraints to ensure that data nodes select a customized contract based on their own type. The computing power service contract is represented as follows: , where R m This is the incentive reward provided for the m-th computing node, α. m It is the task offloading strategy assigned to this computing node. f m It requires the CPU processing power it provides; CN represents the total computing power nodes. The design of the data contribution contract and computing power service contract satisfies individual rationality constraints and incentive compatibility constraints to ensure that nodes select a contract tailored to their own type. Individual rationality constraints mean that the net utility obtained by each node after selecting a contract designed for it must be greater than or equal to zero. Incentive compatibility constraints mean that the utility obtained by the cluster when selecting a contract of its own type must be greater than or equal to the utility it could obtain by disguising itself as other energy consumption types and selecting other contracts.

3. The method according to claim 2, characterized in that, The process of federal secession learning and training includes: The global model parameters are sent to the data nodes participating in the training, and the data nodes are instructed to determine the splitting layers of the model according to their own preferences, dividing the model into head network, middle network and tail network; The data node is instructed to train the head network using local data and to offload the shredded data generated by the computation to the computing power node matched according to the task offloading strategy; The computing node is instructed to receive the shredded data, complete the forward and backward propagation of the intermediate network, and send the calculated gradient information back to the data node. The data nodes are instructed to complete the backpropagation of the tail network and update the model parameters based on the returned gradient information.

4. The method according to claim 3, characterized in that, The optimization objective of the federated split learning training is to maximize the total utility of the UAV, and the optimization objective is expressed as: Where p represents the drone; N is the number of data nodes; and M is the number of computing nodes. This represents the bandwidth vector allocated to data node n; This represents the bandwidth vector allocated to computing node m; This indicates that data node n provides the amount of data D. n The utility of the publisher; This represents the utility of the publisher when the computing node m executes the corresponding task; The total energy consumption of the drone is represented by T; T represents the number of local iterations. This represents the additional communication energy consumed by the data node in receiving and transmitting the model to the drone and unloading the model to the computing node; C n、 C k λ represents the energy consumption per unit of data in a data node, where n and k represent different data nodes in the network; λ represents the cost sensitivity coefficient. This indicates the specific energy consumption margin for each computing node; This represents the probability that each node in the data node belongs to a certain type; This represents the probability that each node in the computing power nodes belongs to a certain type; R represents the energy consumption of a computing node. max This represents the maximum value of the incentive capability provided by the node. f max This represents the maximum CPU processing power of the computing node; , This indicates the weights of data nodes and computing nodes.

5. The method according to claim 2, characterized in that, The states of the data nodes, computing nodes, and communication links include: state information of individual rational constraints and incentive compatibility constraints of the data nodes and computing nodes, as well as the signal-to-noise ratio between the nodes and the UAV.

6. The method according to claim 1, characterized in that, The aggregation includes model parameters uploaded by nodes participating in training according to the contract, including: Receive the updated parameters of the head network and tail network of the data nodes in the subnet, as well as the updated parameters of the intermediate network of the computing power nodes; Based on the amount of data contributed by each data node, all received model parameters are weighted and aggregated to update the global model of the current subnet.

7. The method according to claim 1, characterized in that, The method further includes: Decentralized federated multi-task learning is performed with other drone subnets, specifically: receiving shared layer model parameters published by other drones, calculating the similarity between the local shared layer model and the received other shared layer models; determining aggregation weights based on the similarity, and using the aggregation weights to aggregate the shared layer models of other drones to update the local shared layer model.

8. The method according to claim 1, characterized in that, The pre-trained deep reinforcement learning model is a proximal policy optimization algorithm model; the training process of the model includes: Construct an intelligent agent, using a drone as the intelligent agent, and set up a state space, action space, and reward function; The agent is in the current state s t Below, based on the action network, output action a. t and obtain the action to be performed, a. t The next state s t+1 and reward r t ; Experience (s) t , a t , r t , s t+1 Store in the experience pool; Experience is sampled from the experience pool, and the parameters of the action network and the critic network are updated using the dominance function and the pruning function to maximize the long-term cumulative reward until the reward value converges, thus obtaining a trained model.

9. The method according to claim 8, characterized in that, The state space is represented as follows: in, Let N represent the individual rationality constraints and incentive compatibility constraints of N data nodes at time t, respectively. , ; Let represent the individual rationality constraints and incentive compatibility constraints of the N computing nodes at time t, respectively. , ; This represents the signal-to-noise ratio from the data node to the drone at time t; This represents the signal-to-noise ratio from the computing node to the drone at time t; The action space is specifically represented as follows: in, This represents the set of data contribution contracts generated by the UAV for all data nodes at time t, where This represents the data contribution contract corresponding to each node; This represents the set of computing power service contracts generated by the drone for all computing power nodes at time t, where This represents the computing power service contract corresponding to each node; This represents the bandwidth allocation vector that the UAV assigns to all data nodes at time t. This represents the bandwidth vector allocated by the drone to all computing nodes at time t.

10. The method according to claim 8, characterized in that, The reward function is set as follows: If the action a t If the generated contract and bandwidth allocation scheme satisfy the preset individual rationality constraints, incentive compatibility constraints, total incentive budget constraints, and node processing capacity constraints, then the reward function value is the total utility of the UAV calculated according to the contract and bandwidth allocation scheme. If the action a t If any of the constraints are not met, the reward function value is a preset penalty value.