A Federated Learning Method Based on Partial Customer Participation and Power Constraints

CN115549962BActive Publication Date: 2026-08-11YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0006]2、传输精度:无线链路上的通信可能会受到不良信道条件的影响,FL收敛不可避免地会受到传输错误的影响

Benefits of technology

[0037] Based on the above technical solutions and the technical problems solved, please analyze the advantages and positive effects of the technical solution to be protected by this invention from the following aspects:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115549962B_ABST
    Figure CN115549962B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of communication technology and discloses a federated learning method based on partial client participation and power constraints. The method includes: employing biased client selection, setting client addressability and continuously reporting to a central server, allowing the central server to understand the state of each client during iteration; and allocating power across communication time slots by adding a total power constraint, enabling devices to transmit their local updates to the server. This invention addresses the problems of existing technologies by reducing communication costs while maintaining model convergence performance. The proposed design can mitigate the negative impacts caused by imperfect channel conditions. Convergence of the learning process is slowed by power control, while biased client selection accelerates convergence, balancing model aggregation. This invention can guarantee fast convergence under harsh wireless conditions with low signal-to-noise ratios, and is cost-effective in terms of time and energy savings in wireless FL systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology, and in particular relates to a federated learning method based on partial customer participation and power constraints. Background Technology

[0002] Currently, with the development of modern communication networks and transmission technologies, the application of highly maneuverable drones in distributed tasks is rapidly increasing. Drones can serve as both flying base stations and fast, flexible data acquisition, computation, and wireless communication devices, making them very promising terminals capable of enabling many distributed applications. However, drones are also highly resource-constrained, especially in terms of power and bandwidth. Traditional deep learning-based deployments are cloud-centric, requiring raw data to be uploaded to centralized servers, which imposes heavy network communication overhead on the wireless system. If the communication module shares the same battery as the drone, flight time will be reduced by 16%, and if GPS and other sensors are included, flight time can easily be reduced by more than 20%. Furthermore, reliable and efficient data sharing in dynamic communication links becomes another bottleneck. In addition, centralized schemes may introduce unacceptable latency for real-time applications (such as instance decision-making tasks). Moreover, using drones to collect information from the sky raises critical issues of information privacy and trust. Such aerial footage may contain sensitive frames with privacy concerns.

[0003] Federated learning (FL) algorithms have attracted increasing interest in the field of wireless communications. Research has demonstrated that FL schemes enable wireless devices to collaboratively learn shared models without sharing data. Due to limited available power and bandwidth, and the ubiquitous nature of communication and computing services, the FL concept offers a workaround for multi-UAV wireless networks. Communication in wireless FL networks remains a critical bottleneck in distributed networks, as it can be slow and unstable due to the limited availability of resources such as bandwidth and power. This situation is exacerbated when devices suffer from higher latency, lower throughput, and intermittent poor connectivity. The communication challenges and practical problems in FL systems have spurred the development of new methods to reduce overall power consumption, average client-server communication, and model parameters to determine sufficient accuracy for communication.

[0004] To understand the model convergence of FL, the architecture of FL has been a major focus of theoretical research. However, most of these works concentrate on single aspects of FL optimization, such as device selection schemes, model compression, or model convergence analysis. In wireless FL systems, the optimization problem mainly involves convergence speed and resource efficiency, as well as potentially unreliable or unpredictable client-server communication and limited resource availability. To address the wireless communication overhead of FL algorithms, this research identifies three noteworthy trends for investigation in optimizing distributed learning over wireless networks.

[0005] 1. Limited Bandwidth: Limited wireless channel capacity can lead to severe congestion of the air interface, resulting in unwanted latency and even failure to converge. Furthermore, poor channel conditions limit device availability. Data compression before transmission is one promising solution for reducing power consumption and latency. Additionally, scheduling devices for aggregation can help address the problem of excessive communication congestion.

[0006] 2. Transmission Accuracy: Communication on wireless links can be affected by poor channel conditions, and FL convergence is inevitably affected by transmission errors. To maintain the accuracy of the model, in addition to considering channel bandwidth bottlenecks, robustness to channel effects should also be taken into account to establish a communication model and design an efficient FL framework.

[0007] 3. Power Constraints: Communication and computation consume significant amounts of energy, limiting flight time. One approach to power consumption is to reduce communication frequency; however, model convergence is typically affected by reduced communication. Less frequent communication also leads to slower convergence and suboptimal performance, ultimately increasing overall power consumption due to the need for more local training time. The UAV's transmit power should be limited to extend battery life and meet requirements regarding interference with nearby in-band communications.

[0008] Based on the above analysis, the problems and shortcomings of the existing technology are: the current communication cost is high and the convergence speed is slow. Summary of the Invention

[0009] To address the problems existing in the prior art, this invention provides a federated learning method based on partial client participation and power constraints. This method can overcome the resource limitations of UAVs in terms of power and bandwidth, and better realize the federated learning system of high-mobility wireless communication networks.

[0010] This invention is implemented as follows: a federated learning method based on partial client participation and power constraints, wherein the federated learning method based on partial client participation and power constraints includes:

[0011] Step 1: Employ biased client selection, configure clients to be addressable and continuously report to the central server, so that the central server can understand the status of each client during the iteration process;

[0012] Step two involves allocating power across communication time slots by adding a total power constraint, which enables the device to transmit local updates of the device to the server.

[0013] Furthermore, step one includes:

[0014] (1) Clients connect to the server and seek to find a common model parameter minimization empirical loss function; select the top M clients with the largest losses to perform local updates;

[0015] (2) Adopt a customer selection strategy with linear incremental subsets: Set Select the client with the highest local loss among the top-M values ​​of the current global model for global update; define the client selection policy function π. tm Map the global model w(t) to the selected customer set S(π). tm ,w).

[0016] Furthermore, the empirical loss function is as follows:

[0017]

[0018] Where F(w) represents the empirical loss function; K represents the number of clients; w * The model parameters are represented by f(w,ξ); f(w,ξ) represents the model parameters from the local dataset B. k The loss function for a randomly selected batch of samples ξ in the model parameters w; p k F represents the proportion of data from the k-th client; k (w) represents the local loss function of customer k; F * =min w F(w),

[0019] Furthermore, the federated learning method based on partial customer participation and power constraints also includes:

[0020] The client selects the following update rules for the iteration t of FedAvg:

[0021]

[0022] Among them, w m (t,E) represents the local model parameters of client m at time t. Let ξ(t,E) represent the stochastic gradient on the batch. This represents the global model parameters.

[0023] Furthermore, the step of selecting the top M clients with the greatest losses to perform local updates includes:

[0024] The devices continuously send their local loss information to the central server, which constantly tracks the availability and training status of each responding device. Client selection is made by the server based on the device's latest status and channel status information, and is broadcast to the devices before aggregation. The latest loss value of unselected clients is set to 1. Simultaneously, the selected devices upload their corresponding model parameters.

[0025]

[0026] in and This represents the global model parameters, M represents the size of the active client set, and b represents the size of the mini-batch.

[0027] Furthermore, in step two, power allocation across communication time slots is performed by adding a total power constraint, enabling the device to transmit local updates to the server, which includes:

[0028] 1) Model the system that connects the device and the parameter server via wireless fading MAC and use OFDM for transmission;

[0029] 2) In distributed SGD or DSGD, the server receives superimposed signals from its neighbors: In the t-th round of communication in the DSGD algorithm, local updates are sent to the server via radio fading MAC, using s sub-channels with N time slots; the sub-channels... arrive Assigned to device m∈[M]; for transmit power Equipment m * The capacity of the parallel wireless Gaussian channel between (t) and the server is determined using the following formula:

[0030]

[0031] in, This represents the power allocated to device m in communication round t;

[0032] 3) Power allocation p on sub-channels m [s] adapts to the corresponding channel coefficient h m [s] Gradient aggregation via AirComp: The server aggregates messages according to certain aggregation rules to obtain new global model updates.

[0033] Another object of the present invention is to provide a wireless federated learning communication system that implements the federated learning method based on partial client participation and power constraints.

[0034] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the federated learning method based on partial client participation and power constraints.

[0035] Another object of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the federated learning method based on partial client participation and power constraints.

[0036] Another objective of the present invention is to provide an information data processing terminal for implementing the wireless federated learning communication system.

[0037] Based on the above technical solutions and the technical problems solved, please analyze the advantages and positive effects of the technical solution to be protected by this invention from the following aspects:

[0038] First, addressing the technical problems existing in the prior art and the difficulty in solving them, this paper closely analyzes, in conjunction with the technical solution to be protected by this invention and the results and data obtained during the research and development process, how the technical solution of this invention solves the technical problems, and the inventive technical effects brought about by solving these problems. The specific description is as follows:

[0039] This invention addresses the problems of existing technologies by reducing communication costs while maintaining model convergence performance. The proposed FL scheme employs a biased client participation strategy and then considers a linearly decreasing total power constraint for each communication round to minimize required resources, thereby avoiding over-congestion and ensuring good model convergence for Distributed SGD (DSGD). Experiments show that the proposed design can mitigate the negative impacts caused by imperfect channel conditions. While convergence is slowed by power control, biased client selection accelerates convergence, balancing model aggregation. This invention guarantees fast convergence under harsh wireless conditions with low signal-to-noise ratios, offering cost-effectiveness in terms of time and energy savings for wireless FL systems.

[0040] Second, considering the technical solution as a whole or from a product perspective, the technical effects and advantages of the technical solution to be protected by this invention are specifically described as follows:

[0041] This invention proposes a flow-through (FL) system suitable for highly mobile wireless communication networks. This system enables partial client participation and power constraints to reduce communication costs and avoid exhaustive channel bandwidth, while ensuring convergence speed and exhibiting good noise and attack resistance. In this invention, clients are designed to be addressable and continuously report to the server, allowing the central server to understand the status of each client during iterations. Compared to state-of-the-art FL algorithms that only consider ideal communication environments, the method proposed in this invention theoretically and numerically requires fewer communication rounds and incurs less communication overhead. This invention can be easily applied to aircraft with robust hardware support.

[0042] Third, as supplementary evidence of the inventive step of the claims of this invention, it is also reflected in the following important aspects:

[0043] The expected benefits and commercial value of the technical solution of this invention after transformation are as follows:

[0044] With the further development of big data, prioritizing data privacy and security has become a global trend. Every public data breach attracts significant attention from the media and the public. Federated learning, as an emerging technology, can keep data locally, ensuring privacy is not compromised and regulations are not violated. Simultaneously, multiple participants can collaborate on data to build a virtual shared model and a system that benefits all.

[0045] However, high-speed mobile systems like drones are highly resource-constrained, especially in terms of power and bandwidth. Traditional federated learning-based deployments are cloud-centric, requiring raw data to be uploaded to a centralized server, which imposes heavy network communication overhead on the wireless system. The solution designed in this invention can significantly reduce drone resource consumption through partial customer participation and power constraints, while ensuring high-quality data transmission, and holds promise for commercial deployment to save drone resource consumption. Attached Figure Description

[0046] Figure 1 This is a flowchart of a federated learning method based on partial customer participation and power constraints provided in an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram illustrating the training accuracy of a biased customer selection strategy provided in an embodiment of the present invention.

[0048] Figure 3 This is a schematic diagram of the training loss of a biased customer selection strategy provided in an embodiment of the present invention.

[0049] Figure 4 This is a schematic diagram illustrating the training accuracy of the federated learning system provided in an embodiment of the present invention.

[0050] Figure 5 This is a schematic diagram of a wireless federated learning communication system provided in an embodiment of the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0052] I. Explanation and Description of Embodiments. To enable those skilled in the art to fully understand how the present invention is specifically implemented, this section provides an explanation and description of the embodiments that expand upon the technical solutions of the claims.

[0053] like Figure 1 As shown, the federated learning method based on partial customer participation and power constraints provided in this embodiment of the invention includes:

[0054] S101 employs biased client selection, sets up client addressability and continuously reports to the central server, allowing the central server to understand the status of each client during the iteration process;

[0055] S102, by adding a total power constraint to perform power allocation across communication time slots, the device ground transmits the local update of the device to the server.

[0056] This invention proposes a novel FL optimization paradigm to address the client-to-server communication bottleneck through a power allocation scheme. The proposed FL solution uses a partial client selection strategy and power constraints to reduce overall communication overhead. This invention differs from standard federated learning in two main aspects. First, devices continuously send their state to a central node to make them addressable, facilitating client policy selection. Second, power constraints are taken into account. Compared to state-of-the-art methods in wireless setups, this invention ensures faster convergence and lower power consumption.

[0057] Step 1: Consider a simple communication system consisting of a transmitter, a receiver, and a wireless channel, as shown in the figure. To simplify the system and address the limitations of available computational resources, detailed lemmas and pseudocode are provided to demonstrate the method provided by the embodiments of the present invention under certain assumptions, and numerical experiments were conducted on a lightweight model.

[0058] Step 2: Biased Customer Selection

[0059] Compared to an unbiased client selection strategy, favoring clients with higher local losses can improve convergence speed. Due to the nature of gradient descent, if the current local loss is high, the working device is likely to still have a high local loss in the next epoch. Therefore, in this invention, clients are made addressable so that the central server knows that clients with high local losses have not yet converged to their local optimum.

[0060] (1) Consider a cross-device setup with a total of K clients, where clients connect to the server and seek to find the model parameter w together. * To minimize the empirical loss function F(w):

[0061]

[0062] Where f(w,ξ) represents the value from the local dataset B k The loss function p is the randomly selected sample batch ξ in the model parameters w. k F is the data proportion of the kth client. k (w) is the local loss function for customer k. In fact, the optimal vector w... * and w k It can be very diverse; in this case, we define F. * =min w F(w), To ensure that the set of active clients remains constant in each local iteration, the update rule for the iteration t of FedAvg, chosen by the client, is written as:

[0063]

[0064] Where w m (t,E) represents the local model parameters of client m at time t. It is the stochastic gradient on batch ξ(t,E). This represents the global model parameters.

[0065] (2) In each round, the top M clients with the largest losses are selected to perform local updates. Devices continuously send their local losses to the central server, so the server constantly tracks the availability and training status of each responding device. Client selection is made by the server based on the device's latest state and Channel State Information (CSI) (optional, known to the server), and is broadcast to the devices before aggregation. For unselected clients, their latest loss value is recorded as 1. The selected devices then upload their model parameters accordingly.

[0066]

[0067] in and This represents the global model parameters, M is the size of the active client set, and b represents the size of the mini-batch.

[0068] (3) This invention proposes a customer selection strategy with linear incremental subsets. To reduce the transmission cost... m (t) Regarding the cost to the server, one intuitive approach is to use a steeper convergence curve, which means the number of non-zero entries in the model's parameter updates will decrease rapidly, making w m (t) is a very sparse matrix. A sparse matrix only needs to send its non-zero entries. For each iteration on the server side, this embodiment of the invention sets... Select the client with the highest local loss among the top-M values ​​for global model updates. Define the client selection policy function π. tmIt maps the global model w(t) to the selected customer set S(π). tm (w). This embodiment of the invention assumes that the number of responding clients is always greater than the maximum capacity of the wireless channel.

[0069] Consider selecting a subset S with a dynamic number of clients. t The strategy aims to ensure that the wireless FL converges to a fixed point of the objective function while maintaining transmission within the constraints of the shared channel capacity. The convergence speed may increase with increasing M. However, if... If the ratio is less than 0:3, the performance gain will decrease. A larger M can also lead to a worse convergence ratio because transmission interference introduces errors into the global model aggregation. Therefore, more communication and computation are required, which is energy inefficient. Below, we will discuss improving communication efficiency through a power allocation scheme.

[0070] Step 3: Total Power Constraint

[0071] Using FL in a wireless network requires overcoming the power limitations of mobile devices. The total time spent by the drone uploading parameter updates of size Z is... Where τ is the global iteration for completing FL training, P t Indicates the transmission power. This represents the scaling factor, which includes considerations such as bandwidth. m This represents the channel power gain. The transmission power consumption of each drone is... The performance of the wireless FL algorithm degrades significantly in noisy, fading channels, leading to increased model switching overhead and even convergence failure. The goal of this step is to allocate power across communication slots so that devices can transmit their local updates to the server as accurately as possible, achieving fast convergence within a limited power budget. Consider an optimization problem to minimize the computational distortion of the signal summation, where each participating device has a power constraint to utilize the signal superposition characteristics of the MAC.

[0072] (1) This embodiment of the invention models a system connecting a device and a parameter server via a wireless fading MAC and uses OFDM for transmission. During each communication block, each sensor simultaneously scales its signal via the MAC. The data is transmitted linearly to the receiver. For uplink transmission Tx, embodiments of the invention assume that the wireless channel independently follows a Rayleigh distribution for each device k. Therefore, in the following analysis, the channel power gain follows an exponential distribution with a unit mean. The wireless noise is n 2 ~CN(0,σ 2 I). For the Rx scaling factor Considering that base station power constraints are more stringent than those of equipment, reflecting downlink communication... This is considered ideal. Assuming the channel coefficients are constant within a communication block, during which each device should upload the entire update message of dimension d to the central server. On the server side, the real part of the received signal is isolated because the phase shift of each channel has been compensated.

[0073] (2) In Distributed SGD (DSGD), the server receives superimposed signals from its neighbors. In the t-th round of communication in the DSGD algorithm, the local update is sent to the server via radio fading MAC, with s sub-channels having N time slots. The server receives the channel output y in the N-th time slot of iteration t. n (t) is considered an intermediate product of the model of the embodiments of the present invention, and is given by the following formula:

[0074]

[0075] The i-th binary entry This indicates whether device m was selected in this time slot. The transmission power from device m to the server follows a complex Gaussian distribution CN(0; 2), where pnm(t) represents the transmit power following a power control strategy to be specified later. Represents the channel coefficient, transmitted signal It is obtained through model encoding and normalization, which facilitates power control, n 2 (t) represents the noise vector.

[0076] If the CSI is known at each iteration t, then at iteration t, select the number m of devices whose gain exceeds the power-off threshold and construct an activity set A(t) of size d.

[0077]

[0078] Where s represents the number of parallel Gaussian channels. In this case, each device can achieve channel inversion by multiplying its local model by its inverse channel coefficient. For digital DSGD, this embodiment of the invention considers N=1. The i-th entry of the transmission from device m(t) is given by the following equation:

[0079]

[0080] Considering symmetry, embodiments of the present invention will use sub-channels arrive Assigned to device m∈[M]. For transmit power... Equipment m * The capacity of the parallel wireless Gaussian channel between (t) and the server is given by the following formula:

[0081]

[0082] in This is the power allocated to device m during communication round t. For small values ​​of s, the upper limit of capacity is quite large.

[0083] (3) In order to reduce the transmission w m One intuitive approach to reducing the cost to the server (t) is to balance the bit depth of quantization to achieve a steeper convergence curve. In earlier training phases, large model gradients are quantized with a smaller bit depth representing floating-point numbers to save on communication costs from each device to the server. Ideally, after several rounds of communication, the information content of the gradient decreases as the number of non-zero entries in the model updates parameters decreases. Sparse matrices only need to send the values ​​of their non-zero entries and reuse previously unchanged gradients. In this case, more bit depth should be allocated to the client in time slots to maintain the accuracy of non-zero entries.

[0084] In practice, each device m∈[M] is constrained by the average power budget P0. Power allocation p on the sub-channels... m [s] adapts to the corresponding channel coefficient h m [s] Gradient aggregation is implemented via AirComp. The transmission power constraints for each individual device are as follows:

[0085]

[0086] Where i∈[s] represents a sub-channel. Because the channel coefficients are distributed identically across the sub-channels, the power constraint for each channel can be reduced to:

[0087]

[0088] Because the subchannel only inverts when its gain exceeds a threshold, otherwise it may encounter deep attenuation. Therefore, device m * The transmission power (t) on the i-th sub-channel is as follows:

[0089]

[0090] Where p0 is the scaling factor for average transmit power control. Embodiments of the present invention are shown in the following analysis. Since the server is unaware of the variance of gradient estimates across different devices, power is distributed evenly among the devices. The maximum transmit power of each device in the long run is defined as... Assume that device m is assigned * The power of each device in iteration τ is equal, and the upper limit of broadband capacity is satisfied. t Send its processed signal, where P t satisfy:

[0091]

[0092] Power limiting can suppress interference from nearby communication systems. In this embodiment of the invention, P is represented as a linearly decreasing function because a power-limited AirComp system helps extend battery life.

[0093]

[0094] Where `max_iter` represents the predefined maximum number of communication rounds, and `num_iter` represents the current number of communication rounds. The precise value of `p0` can be obtained through:

[0095]

[0096] in It's an exponential integral function. Recall that the channel gain v m (t) follows an exponential distribution with a unit mean. Therefore, the power constraint is... Given that, under linear power constraints, convergence should be slightly faster than the theoretically achievable average power distribution.

[0097] (4) For the model aggregation step, the server aggregates messages according to certain aggregation rules to obtain new global model updates. Unlike standard FL methods with one-step average aggregation, wireless FL systems require more robust aggregation rules because naive average aggregation rules are vulnerable to model or data attacks. Geometric median aggregation rules provide good convergence guarantees for DSGD, especially when some devices are attacked, making it an ideal tool for wireless FL systems. To find the geometric median of a set of points in z, the classic Weiszfeld algorithm is applied to this problem. Given points z (t) ∈{w m For the set {(t), m∈[M]}, their l1 mean is minimized. point z * , where |.|| denotes the Euclidean norm. The power constraint still holds. The original problem is formulated as follows:

[0098]

[0099] in It is ω m The weight of (t)

[0100]

[0101] Where v is a smoothing factor to avoid minimum distances. The objective function aims to find the vector {ω} that minimizes the distance to all updated messages. m (t),m∈[M]} is defined as:

[0102]

[0103]

[0104] The above equation can be solved as a convex optimization problem using the Weiszfeld algorithm. For simplicity, the channel coefficients are set to constants during the communication block, and the phase shift of each sub-channel can be compensated at the transmitter. Let ω... m (t) is the initial point, therefore z = ω in each global aggregation. m (t), m∈[M]. The update message is constructed as follows: Where W represents ω m The magnitude of (t). Scale factor Shared among active devices, as a complement to the original ω m The remedy for (t) is to address the different average power expectations. and To satisfy the average transmit power constraint, the channel inversion at device m is applied as follows:

[0105]

[0106] Among them, κ m (t) represents updating information, x' m (t) represents x m Channel inversion of (t), This represents the power scaling factor. In the inner loop of the Weiszfeld algorithm for geometric median aggregation via AirComp, the received signal z from the server to the participating devices should satisfy amplitude alignment. The power scaling factor ρ for all devices... m All are equal; for convenience, we will replace them with ρ0.

[0107] II. Application Examples. To demonstrate the inventiveness and technical value of the present invention, this section provides application examples of the claimed technical solutions applied to specific products or related technologies.

[0108] The federated learning method based on partial client participation and power constraints provided in this invention is applied to a high-speed mobile unmanned aerial vehicle (UAV) communication system. This system requires communication between the UAV and ground clients, and the UAV is in a state of high-speed movement. To overcome the resource limitations of UAVs, applying partial client participation and power constraints according to this invention can achieve an efficient federated learning system.

[0109] The federated learning method based on partial client participation and power constraints provided in this embodiment of the invention is applied to a computer device, the computer device including a memory and a processor, the memory storing a computer program, and when the computer program is executed by the processor, the processor performs the steps of the federated learning method based on partial client participation and power constraints.

[0110] The federated learning method based on partial client participation and power constraints provided in this embodiment of the invention is applied to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor performs the steps of the federated learning method based on partial client participation and power constraints.

[0111] The federated learning method based on partial client participation and power constraints provided in this embodiment of the invention is applied to an information data processing terminal, which is used to implement the wireless federated learning communication system.

[0112] III. Evidence of the Relevant Effects of the Embodiments. The embodiments of the present invention have achieved some positive effects during research and development or use, and indeed possess significant advantages compared to existing technologies. The following description, in conjunction with data, charts, and other materials from the experimental process, illustrates these advantages.

[0113] Compared to existing federated learning methods, the federated learning method based on partial client participation and power constraints provided in this invention reduces communication costs while maintaining model convergence performance. The proposed FL scheme employs a biased client participation strategy and then considers a linearly decreasing total power constraint for each communication round to minimize required resources, thereby avoiding over-congestion and ensuring good model convergence for Distributed SGD (DSGD). Experiments show that the proposed design can address the negative impacts caused by imperfect channel conditions. Convergence of the learning process is slowed by power control, while biased client selection accelerates convergence, balancing model aggregation. This invention guarantees fast convergence under harsh wireless conditions with low signal-to-noise ratios, offering cost-effectiveness in terms of time and energy savings for wireless FL systems.

[0114] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0115] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A federated learning method based on partial customer participation and power constraints, characterized in that, The federated learning method based on partial customer participation and power constraints includes: Step 1: Employ biased client selection, configure clients to be addressable and continuously report to the central server, so that the central server can understand the status of each client during the iteration process; Step 2: Power allocation across communication time slots is performed by adding a total power constraint, causing the device to transmit the device's local update to the server; Step one includes: (1) The client connects to the server and seeks to find the model parameters that minimize the empirical loss function; the top M clients with the largest losses are selected to perform local updates; (2) Adopt a customer selection strategy with linear incremental subsets: Set Select the client with the highest local loss among the top-M values ​​for the current global model and perform a global update; define the client selection policy function. , global model Mapped to selected customer set ; The empirical loss function is as follows: ; in, The empirical loss function is represented by K; the number of clients is represented by K. Indicates model parameters; Indicates from local dataset and model parameters Randomly selected sample batches The loss function; This represents the proportion of data from the k-th client; Let the local loss function for customer k be denoted as . .

2. The federated learning method based on partial client participation and power constraints as described in claim 1, characterized in that, The federated learning method based on partial customer participation and power constraints also includes: The client selects the following update rules for the iteration t of FedAvg: in, This represents the local model parameters of client m at time t. Indicates batch stochastic gradient on, Represents global model parameters. The learning rate represents the rate of change over time. Local epoch count The change controls the gradient descent step size.

3. The federated learning method based on partial client participation and power constraints as described in claim 1, characterized in that, The process of selecting the top M clients with the greatest losses to perform local updates includes: The devices continuously send their local loss information to the central server, which constantly tracks the availability and training status of each responding device. Client selection is made by the server based on the device's latest status and channel status information, and is broadcast to the devices before aggregation. The latest loss value of unselected clients is set to 1. Simultaneously, the selected devices upload their corresponding model parameters. ; in and Here, M represents the size of the active client set, and b represents the size of the mini-batch. Let represent the global learning rate in the t-th communication round.

4. The federated learning method based on partial client participation and power constraints as described in claim 1, characterized in that, In step two, power allocation across communication time slots is performed by adding a total power constraint, enabling the device to transmit local updates to the server, which includes: 1) Model the system that connects the device and the parameter server via wireless fading MAC and use OFDM for transmission; 2) In distributed SGD or DSGD, the server receives superimposed signals from its neighbors: In the t-th round of communication in the DSGD algorithm, local updates are sent to the server via radio fading MAC, using s sub-channels with N time slots; the sub-channels... arrive Assigned to device Regarding transmission power ,equipment The capacity of the parallel wireless Gaussian channel between the server and the server is determined using the following formula: ; in, This represents the power allocated to device m in communication round t; 3) Power allocation on sub-channels Adapt to the corresponding channel coefficients Gradient aggregation can be performed via AirComp: The server aggregates messages according to certain aggregation rules to obtain new global model updates.

5. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the federated learning method based on partial client participation and power constraints as described in any one of claims 1-4.

6. A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the federated learning method based on partial client participation and power constraints as described in any one of claims 1-4.

7. An information data processing terminal, characterized in that, The information data processing terminal includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the steps of the federated learning method based on partial client participation and power constraints as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Resource allocation method and system for multitask federated learning in 5G network

    CN113900796A

  • System and method for distributed learning of wireless edge dynamics

    CN114930347A