A fully asynchronous decentralized hierarchical federated learning method and system
Patent Information
- Application Number
- CN202411437653.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-10-15
AI Technical Summary
同步式训练的过程中,由于节点硬件算力的异构性,训练速度可能会由于等待某个慢速节点而被拖慢
[0038]本发明实施例的技术方案,通过各上层服务器初始化生成服务器模型,并将所述服务器模型下发至通信范围内的各客户端设备;客户端设备根据接收到的所述服务器模型进行全异步客户端本地训练,更新生成客户端模型,并将所述客户端模型上传至筛选的目标上层服务器;目标上层服务器根据所述客户端模型进行全异步服务器本地模型聚合,更新生成服务器模型;目标上层服务器在确定接收的客户端模型数量满足通信条件时,在筛选的目标邻居服务器中拉取模型,全异步进行服务器模型聚合更新;目标上层服务器将当前的所述服务器模型下发至对应的客户端设备;返回客户端设备根据接收到的所述服务器模型进行全异步客户端本地训练,更新生成客户端模型步骤,直至各上层服务器的服务器模型满足收敛条件,得到全局模型,解决了传统联邦学习中模型训练精度差且单个服务器潜在的单节点故障问题,以及由慢速客户端设备带来的问题,通过客户端设备进行模型更新可保护用户隐私数据;通过客户端设备与服务器均进行模型更新可提高模型训练精度;通过客户端筛选目标上层服务器可增加网络动态性,客户端设备能及时与服务器通信进行模型训练,无需绑定在某个服务器上,防止单个服务器拥塞或故障;通过目标上层服务器筛选目标邻居服务器进行服务器模型聚合,无需在上层添加中央服务器即去中心化,可避免中央服务器的网络拥塞及单节点故障情况,且无需向其余服务器广播数据,减少网络通信量,避免资源消耗。
Smart Images

Figure CN119358610B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a fully asynchronous decentralized hierarchical federated learning method and system. Background Technology
[0002] With the development of data-driven artificial intelligence (AI) technology, its wide-ranging applications have entered various fields such as image processing and intelligent push notifications. One of the core aspects of AI technology is the data required to train models. However, as data privacy protections become increasingly stringent in sectors such as finance and healthcare, the collection of the large amounts of data needed for model training has become a major bottleneck. Therefore, federated learning has emerged as a way to perform data analysis and model training without exposing users' raw data.
[0003] Due to the development of mobile computing and Internet of Things (IoT) technologies and the improvement of terminal hardware performance in recent years, the data sources for many emerging application scenarios have gradually shifted from storage in cloud data centers or single data management facilities to a wider range of terminal devices (such as personal mobile phones, home appliances, and sensors). Because of differences in geographical environment, application scenarios, user habits, and other factors, the local data generated by these edge nodes is usually non-independent and identically distributed (Non-IID), which can lead to a decrease in the accuracy of model training.
[0004] Traditional federated learning often employs a synchronous central server architecture. During synchronous training, the heterogeneity of node hardware computing power can slow down the training process by waiting for a slow node. Furthermore, the architecture of a central server collecting or aggregating local models is not only more prone to network congestion and communication bottlenecks when there are many connected nodes, but also susceptible to single points of failure that could lead to training termination.
[0005] Therefore, there is an urgent need to provide a fully asynchronous decentralized federated learning method to improve model training accuracy while protecting user privacy data and avoiding training interruption caused by network congestion and single-node failure. Summary of the Invention
[0006] This invention provides a fully asynchronous decentralized hierarchical federated learning method and system to improve model accuracy and save network costs.
[0007] According to one aspect of the present invention, a fully asynchronous decentralized hierarchical federated learning method is provided, the method comprising:
[0008] Each upper-layer server initializes and generates a server model, and then distributes the server model to each client device within the communication range;
[0009] The client device performs fully asynchronous local training based on the received server model, updates and generates the client model, and uploads the client model to the selected target upper-layer server;
[0010] The target upper-layer server performs fully asynchronous local model aggregation based on the client model and updates and generates the server model.
[0011] When the target upper-layer server determines that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs server model aggregation and update in a fully asynchronous manner.
[0012] The target upper-layer server distributes the current server model to the corresponding client device;
[0013] The client device then performs fully asynchronous local training based on the received server model, updating and generating the client model, until the server models of each upper-layer server meet the convergence condition, thus obtaining the global model.
[0014] According to another aspect of the present invention, a fully asynchronous decentralized hierarchical federated learning system is provided, the hierarchical federated learning system comprising multiple upper-layer servers and multiple lower-layer client devices; wherein:
[0015] The upper-layer server is used to initialize and generate the server model, and then distribute the server model to each client device within the communication range.
[0016] The client device is used to perform fully asynchronous local training based on the received server model, update and generate the client model, and upload the client model to the selected target upper-layer server.
[0017] The target upper-layer server is used to perform fully asynchronous server local model aggregation based on the client model and update the generated server model; when it is determined that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs fully asynchronous server model aggregation and update; and sends the current server model to the corresponding client device.
[0018] The client device and the target upper-layer server repeatedly update the model until the server model of each upper-layer server meets the convergence condition, thus obtaining the global model.
[0019] According to another aspect of the present invention, a fully asynchronous decentralized hierarchical federated learning method is provided, the method being applied to an upper-layer server, the method comprising:
[0020] Initialize and generate a server model, and then distribute the server model to each client device within the communication range;
[0021] Obtain the upload request from the client device, obtain the client model, and perform fully asynchronous server-local model aggregation based on the client model to update and generate the server model;
[0022] Once it is determined that the number of received client models meets the communication conditions, models are retrieved from the selected target neighbor servers, and server model aggregation and updates are performed asynchronously.
[0023] The current server model is distributed to the corresponding client device; and the upload request of the client device is returned, and the client model acquisition step is repeated until the server models of each upper-layer server meet the convergence condition, thus obtaining the global model.
[0024] According to another aspect of the present invention, a fully asynchronous decentralized hierarchical federated learning method is provided, the method being applied to a client device, the method comprising:
[0025] Obtain the server model sent by the upper-layer server;
[0026] Based on the received server model, perform fully asynchronous local training on the client, update and generate the client model, and upload the client model to the selected target upper-layer server;
[0027] Obtain the server model sent by the target upper-layer server; and return the steps of performing fully asynchronous local training on the client based on the received server model, updating and generating the client model, until the server model is obtained as the global model.
[0028] According to another aspect of the present invention, a server is provided, the server comprising:
[0029] At least one processor; and
[0030] A memory communicatively connected to the at least one processor; wherein,
[0031] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the fully asynchronous decentralized hierarchical federated learning method according to any embodiment of the present invention.
[0032] According to another aspect of the present invention, a client device is provided, the client device comprising:
[0033] At least one processor; and
[0034] A memory communicatively connected to the at least one processor; wherein,
[0035] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the fully asynchronous decentralized hierarchical federated learning method according to any embodiment of the present invention.
[0036] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the fully asynchronous decentralized hierarchical federated learning method described in any embodiment of the present invention.
[0037] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the fully asynchronous decentralized hierarchical federated learning method described in any embodiment of the present invention.
[0038] The technical solution of this invention involves initializing and generating a server model on each upper-layer server, and then distributing the server model to each client device within the communication range. The client devices perform fully asynchronous local training based on the received server model, update and generate a client model, and upload the client model to a selected target upper-layer server. The target upper-layer server performs fully asynchronous local model aggregation based on the client model, and updates and generates a server model. When the target upper-layer server determines that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs fully asynchronous server model aggregation and update. The target upper-layer server then distributes the current server model to the corresponding client device. Finally, the client device performs fully asynchronous local training based on the received server model to update and generate a client model. The process continues until the server models of each upper-layer server meet the convergence condition, resulting in a global model. This solves the problems of poor model training accuracy and potential single-node failures of individual servers in traditional federated learning, as well as the problems caused by slow client devices. Updating the model through client devices protects user privacy data; updating the model through both client devices and servers improves model training accuracy; filtering target upper-layer servers through the client increases network dynamism, allowing client devices to communicate with servers in a timely manner for model training without being bound to a specific server, preventing congestion or failure of a single server; and filtering target neighbor servers through target upper-layer servers for server model aggregation eliminates the need for a central server, thus achieving decentralization and avoiding network congestion and single-node failures of the central server. Furthermore, it eliminates the need to broadcast data to other servers, reducing network communication and avoiding resource consumption.
[0039] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart of a fully asynchronous decentralized hierarchical federated learning method provided in Embodiment 1 of the present invention;
[0042] Figure 2 This is a schematic diagram of a fully asynchronous decentralized hierarchical federated learning system structure provided in Embodiment 1 of the present invention;
[0043] Figure 3 This is a flowchart of a fully asynchronous decentralized hierarchical federated learning method provided in Embodiment 2 of the present invention;
[0044] Figure 4 This is a flowchart of a fully asynchronous decentralized hierarchical federated learning method provided in Embodiment 4 of the present invention;
[0045] Figure 5 This is a flowchart of a fully asynchronous decentralized hierarchical federated learning method provided in Embodiment 5 of the present invention;
[0046] Figure 6 This is a schematic diagram of the structure of a server or client device that implements the fully asynchronous decentralized hierarchical federated learning method of this invention. Detailed Implementation
[0047] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0048] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0049] Example 1
[0050] Figure 1 This is a flowchart of a fully asynchronous decentralized hierarchical federated learning method according to Embodiment 1 of the present invention. This embodiment is applicable to model training when protecting user privacy data. The method can be executed by a fully asynchronous decentralized hierarchical federated learning system.
[0051] Figure 2 This is a schematic diagram of a fully asynchronous decentralized hierarchical federated learning system architecture provided in Embodiment 1 of the present invention. Figure 2 As shown, the system consists of multiple upper-layer servers and multiple lower-layer client devices. The connection status between the upper-layer servers and the lower-layer client devices is not fixed. Client devices select a target upper-layer server within their communication range to upload their client models, and the target upper-layer server sends a server model to the corresponding client device. The connection status between the upper-layer servers is also not fixed. The target upper-layer server selects a target neighbor server from its neighbor servers for server model aggregation and updates. Upper-layer servers can be edge servers. Server models can be maintained within the upper-layer servers. A server model is a model generated by aggregating the server's local model with models uploaded by clients or models pulled from neighbor servers. Client devices can be mobile phones, home appliances, cameras, computers, and game consoles, etc. Client devices can maintain their own client models. A client model can be a model obtained by receiving a model sent from an upper-layer server and training it on local data.
[0052] like Figure 1 As shown, the method includes:
[0053] Step 110: Each upper-layer server initializes and generates a server model, and then distributes the server model to each client device within the communication range.
[0054] Multiple upper-layer servers can initialize the server model. These servers can then distribute the server model to client devices within the communication range. Each upper-layer server can maintain a version number for the server model, facilitating subsequent model aggregation. When distributing the server model to client devices, the upper-layer server can simultaneously distribute the server model's version number. Each upper-layer server can increment the version number by 1 when a server model is aggregated. For example, each upper-layer server can maintain a timestamp S. sid This represents the server model version number. As models aggregate from upper-layer servers, S... sid Add 1 to the update. Client devices can maintain server models and version numbers issued by each connectable upper-layer server. For example, client devices can maintain timestamp vectors. Among them, s x This is the server model version number sent from the upper-layer server x to the client device.
[0055] Step 120: The client device performs fully asynchronous local training based on the received server model, updates and generates the client model, and uploads the client model to the selected target upper-layer server.
[0056] Each client device can perform fully asynchronous model training based on local user data from the received server model, and update and generate the client model. When uploading the client model to the upper-layer server, the client device can filter among the upper-layer servers to obtain the target upper-layer server. The client device then uploads the client model to the selected target upper-layer server. The target upper-layer server determined by each client device can be different.
[0057] The client device can query upper-layer servers to determine which upper-layer servers are connectable. There can be multiple connectable upper-layer servers. The client device can then filter for a target upper-layer server from these multiple connectable servers. There are several ways to filter. For example, the client device can query the bandwidth offered by each connectable upper-layer server. Alternatively, the client device can calculate the upload speed of each connectable upper-layer server. The client device can then filter for a target upper-layer server based on bandwidth and / or upload speed.
[0058] In an optional embodiment of the present invention, uploading the client model to the selected target upper-layer server includes: the client device querying connectable upper-layer servers; the connectable upper-layer servers determining the bandwidth allocated to the client device based on the client device's connection query; and the client device asynchronously selecting target upper-layer servers based on the bandwidth and uploading the client model to the target upper-layer server.
[0059] Specifically, any upper-layer server within the communication range of the client device can be considered a candidate upper-layer server. If all candidate upper-layer servers are communicating with neighboring servers (i.e., performing model aggregation between servers), then there are currently no available upper-layer servers, and the client device enters a polling waiting state. The client device can query every Δt time slots. If a candidate upper-layer server is not communicating with a neighboring server, the client device can query that candidate upper-layer server and designate it as a connectable upper-layer server.
[0060] Connectable upper-layer servers can allocate bandwidth to client devices based on their own bandwidth margin. There are several allocation methods. For example, a connectable upper-layer server can allocate a preset percentage of its bandwidth margin to the client device. This preset percentage can be between 60% and 95%. Alternatively, a connectable upper-layer server can calculate the minimum bandwidth required for the client device to upload client models and allocate bandwidth to the client device based on both the bandwidth margin and the minimum bandwidth. For example, bandwidth greater than or equal to the minimum bandwidth and less than or equal to the bandwidth margin can be allocated to the client device.
[0061] In this embodiment of the invention, in order to allocate a more suitable bandwidth to the client device, the upper-layer server may optionally determine the bandwidth allocated to the client device based on the client device's connection query. This includes: the upper-layer server determining the current wireless channel capacity, minimum upload time, and the amount of data uploaded by the client device based on the client device's connection query; the upper-layer server determining the minimum bandwidth required by the client device based on the channel capacity, minimum upload time, and amount of data uploaded; and the upper-layer server determining the bandwidth allocated to the client device when the bandwidth margin is greater than the minimum bandwidth.
[0062] Among them, the connectable upper-layer servers can be based on Shannon's formula. Determine the channel capacity. Where P... cid Let g be the transmit power of the client device, g be the wireless channel gain, and N0 be the noise power. g can be decomposed into small-scale fading with large fluctuations but relatively small impacts and large-scale fading with small fluctuations but relatively large impacts. N0 can be the Gaussian white noise power. Minimum upload time t min This can be predefined in the system. The amount of data uploaded by the client device, 'a'. cid This can be determined by the client device. The connectable upper-layer server can be determined according to the formula. Determine the minimum bandwidth required by the client device.
[0063] The connectable upper-layer server can determine its bandwidth margin B based on its own parameters. remain If the available bandwidth exceeds the minimum bandwidth, the connectable upper-layer server can allocate bandwidth to the client device. Otherwise, the connectable upper-layer server can directly inform the client device that bandwidth is unavailable. When the connectable upper-layer server is available, various methods can be used to allocate bandwidth to the client device.
[0064] Optionally, in this embodiment of the invention, when the bandwidth margin of the connectable upper-layer server is greater than the minimum bandwidth, the determination of the bandwidth allocated to the client device includes: the connectable upper-layer server determining a server update threshold based on the number of times each client device uploads client model updates within the time interval between two communications with the target connected server; the connectable upper-layer server determining at least two recording windows based on the server update threshold; and within each recording window, determining a historical average update time interval based on the arrival time interval of client model uploads from the client device; the connectable upper-layer server determining an update frequency change trend based on each historical average update time interval; the connectable upper-layer server determining a server contention factor based on the number of successful connections and the number of connection failures of the client device within the target recording window; and the connectable upper-layer server determining the bandwidth allocated to the client device based on the update frequency change trend, the server contention factor, the minimum bandwidth, and the bandwidth margin.
[0065] Among them, the server update threshold u thd This represents the number of times the client model is updated by the current server during two communications with neighboring servers. Upper-layer servers can maintain the arrival time intervals of client model updates from each client device as historical information. For example, historical information can be represented as his = {h1, h2, ..., h...} y}. h y This represents the time interval between the current reception of a client model and the last reception. y represents the size of the recording window. y can be determined based on the server update threshold u. thd .
[0066] For example, each upper-layer server can maintain two recording windows, one of which is a long window T. wndL One is the short-term window T wndS Long-term window T wndL The length can be used to update the server threshold u thd Short-term window T wndS The length can be less than the server update threshold u thd For example, short-term window T wndS The length can be For the long window T wndLIn the historical information, y can be the server update threshold u. thd For a short-term window T wndS In historical information, y can be Historical information his = {h1, h2, ..., h y This records the window length and the latest client model update time interval. When the data record is full, the first data entry can be deleted, and new data can be added to the end.
[0067] The upper-layer server can determine the historical average update interval for each recording window based on historical messages within that window. For example, the exponential moving average (EMA) can be used to determine the historical average update interval for each recording window based on historical messages within that window. Therefore, the trend of update frequency changes can be determined based on these historical average update intervals.
[0068] For example, a long-term window T wndL The historical average update interval is h meanL Short-term window T wndS The historical average update interval is h meanS The update frequency trend can be represented as trend=h meanS -h meanL By using long-term and short-term windows to determine the update frequency trend, the latest trend can be captured while taking into account the long-term average time interval, thus improving the reliability of the target upper-layer server determination.
[0069] In this embodiment of the invention, a short-term window can be used as the target recording window period. The upper-layer server can determine the server contention factor based on the number of successful and failed connections from client devices within the target recording window period. For example, the server contention factor can be determined using a formula... Determined. Wherein, t fail The number of client device connection failures during the target recording window, t success This records the number of successful client device connections within the target recording window. If an upper-layer server receives a query from a client device and can provide at least the minimum bandwidth, but is not selected by the client device, it is considered a failed connection between the upper-layer server and the client device.
[0070] There are many ways for the upper-layer server to determine the amount of bandwidth allocated to client devices based on update frequency trends, server contention factors, minimum bandwidth, and bandwidth margin. One exemplary method could be that the upper-layer server uses formula B. alloc =B min +γ(B remain -Bmin The sigmoid(trend×β) function determines the amount of bandwidth allocated to the client device. Here, γ is a control coefficient for the amount of additional bandwidth allocated, and it can take values within the interval [0,1]. When allocating bandwidth, the sigmoid function can be used to map trend×β to the interval [0,1].
[0071] This invention determines the bandwidth allocated to client devices based on update frequency trends, server competition factors, minimum bandwidth, and bandwidth margin. Multiple influencing factors are considered during bandwidth allocation. Besides the minimum bandwidth required for model uploads, additional bandwidth is allocated to ensure reliable model uploads. This additional bandwidth allocation considers factors such as additional bandwidth control coefficients, server competitiveness, and model upload time intervals. This allows for the selection of a better upper-layer server for client devices while providing appropriate bandwidth, thereby improving model training speed and accuracy.
[0072] After the client device obtains the bandwidth available from each connectable upper-layer server, it can filter and determine the target upper-layer server based on the bandwidth available. For example, the client device can choose the upper-layer server that can provide a larger bandwidth as the target upper-layer server. Alternatively, the client device can further infer the upload speed based on the bandwidth allocated to each connectable upper-layer server. Then, the client device selects the target upper-layer server based on the upload speed of each connectable upper-layer server.
[0073] In an optional embodiment of the present invention, the client device asynchronously filters target upper-layer servers based on bandwidth size, including: the client device asynchronously infers the upload rate of each connectable upper-layer server based on bandwidth size and the channel capacity of the current wireless channel; the client device asynchronously filters the connectable upper-layer server with the optimal upload rate as the target upper-layer server based on the upload rate.
[0074] The product of the available bandwidth and channel capacity of the connectable upper-layer server can be used as the upload rate of the connectable upper-layer server. The client device can then select the upper-layer server with the optimal upload rate as the target upper-layer server. The client device can then upload its client model to the target upper-layer server.
[0075] Step 130: The target upper-layer server performs fully asynchronous local model aggregation based on the client model and updates and generates the server model.
[0076] There are several ways for the target upper-layer server to perform model aggregation. For example, the target upper-layer server can average the current server model with the client model and update the generated server model. Alternatively, the target upper-layer server can set weights to perform a weighted average of the current server model and the client model and update the generated server model, thus avoiding global model offset issues caused by the different training speeds of the models on various client devices.
[0077] Step 140: When the target upper-layer server determines that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs server model aggregation and update asynchronously.
[0078] The requirement that the number of received client models meets the communication conditions for the target upper-layer server can be defined as the number of received client models being greater than or equal to a preset reception threshold. The target upper-layer server can have multiple communicable neighbor servers. The target upper-layer server can filter among these neighbor servers to determine the target neighbor server. There are several filtering methods. For example, the target upper-layer server can filter based on the version number of the server models in each neighbor server. Alternatively, it can filter based on the difference in the amount of data between the target upper-layer server and each neighbor server. Another example is filtering based on the link transmission rate of each neighbor server. Yet another example is filtering based on a combination of factors, including version number, data volume difference, and link transmission rate, to determine the target neighbor server. There can be one or more target neighbor servers.
[0079] The target upper-layer server can pull the server models of the target neighbor servers and perform model aggregation updates. There are various aggregation methods, such as average aggregation or weighted average aggregation. By selecting target neighbor servers through filtering, there is no need to broadcast to other servers, avoiding the large amount of information exchange caused by differences in control nodes and reducing network communication consumption.
[0080] Step 150: The target upper-layer server distributes the current server model to the corresponding client device.
[0081] If the target upper-layer server meets the communication requirements, the server model obtained by aggregating the models with the target neighbor servers can be sent to the corresponding client device. If the target upper-layer server does not meet the communication requirements, the server model obtained by aggregating the models with the client models can be sent to the corresponding client device.
[0082] By using a fully asynchronous decentralized model aggregation method, we can avoid model training failures caused by single server failures, avoid global model offsets caused by slow updates from a certain client device, avoid network congestion, and also make reasonable allocation of wireless resources.
[0083] Step 160: Return to step 120 until the server models of each upper-layer server meet the convergence condition, and obtain the global model.
[0084] If each upper-layer server can achieve the target accuracy on the potential connected client devices, it indicates that the server models have converged and the system training can be stopped to obtain the global model.
[0085] The technical solution of this embodiment involves each upper-layer server initializing and generating a server model, and then distributing the server model to each client device within the communication range. The client devices perform fully asynchronous local training based on the received server model, update and generate a new client model, and upload the client model to the selected target upper-layer server. The target upper-layer server performs fully asynchronous local model aggregation based on the client model, and updates and generates a new server model. When the target upper-layer server determines that the number of received client models meets the communication conditions, it retrieves models from the selected target neighbor servers and performs fully asynchronous server model aggregation and update. The target upper-layer server then distributes the current server model to the corresponding client device. The process of the client device performing fully asynchronous local training based on the received server model and updating and generating a new client model continues until all upper-layer servers... The server model meets the convergence condition to obtain the global model, which solves the problems of poor model training accuracy and potential single-node failure of a single server in traditional federated learning, as well as the problems caused by slow client devices. Updating the model through client devices can protect user privacy data; updating the model through both client devices and servers can improve model training accuracy; filtering target upper-layer servers through the client can increase network dynamism, and client devices can communicate with the server in a timely manner to train the model without being bound to a single server, preventing congestion or failure of a single server; filtering target neighbor servers through target upper-layer servers for server model aggregation does not require adding a central server at the upper layer, thus decentralizing the process and avoiding network congestion and single-node failure of the central server. It also eliminates the need to broadcast data to other servers, reducing network communication volume and avoiding resource consumption.
[0086] Example 2
[0087] Figure 3 This is a flowchart of a fully asynchronous decentralized hierarchical federated learning method according to Embodiment 2 of the present invention. This embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 3 As shown, the method includes:
[0088] Step 310: Each upper-layer server initializes and generates a server model, and then distributes the server model to each client device within the communication range.
[0089] Step 320: The client device performs fully asynchronous local training based on the received server model and updates and generates the client model.
[0090] Step 330: The client device queries the connectable upper-layer servers.
[0091] Step 340: The connectable upper-layer server determines the bandwidth allocated to the client device based on the client device's connection query.
[0092] In an optional embodiment of the present invention, the connectable upper-layer server determines the bandwidth allocated to the client device based on the connection query of the client device, including: the connectable upper-layer server determining the channel capacity of the current wireless channel, the minimum upload time, and the amount of data uploaded by the client device based on the connection query of the client device; the connectable upper-layer server determining the minimum bandwidth required by the client device based on the channel capacity, the minimum upload time, and the amount of data uploaded; and the connectable upper-layer server determining the bandwidth allocated to the client device when the bandwidth margin is greater than the minimum bandwidth.
[0093] Based on the above implementation, optionally, when the bandwidth margin is greater than the minimum bandwidth, the connectable upper-layer server determines the bandwidth allocated to the client device, including: the connectable upper-layer server determining a server update threshold based on the number of times each client device uploads client model updates within the time interval between two communications with the target connected server; the connectable upper-layer server determining at least two recording windows based on the server update threshold; and within each recording window, determining a historical average update time interval based on the arrival time interval of client model uploads from the client device; the connectable upper-layer server determining the update frequency change trend based on each historical average update time interval; the connectable upper-layer server determining a server contention factor based on the number of successful connections and the number of connection failures of the client device within the target recording window; and the connectable upper-layer server determining the bandwidth allocated to the client device based on the update frequency change trend, the server contention factor, the minimum bandwidth, and the bandwidth margin.
[0094] Step 350: The client device asynchronously estimates the upload rate to each connectable upper-layer server based on the bandwidth and the current wireless channel capacity.
[0095] Step 360: The client device asynchronously selects the upstream server with the best upload speed as the target upstream server based on the upload speed.
[0096] Step 370: The client device asynchronously uploads the client model to the selected target upper-layer server.
[0097] Step 380: The target upper-layer server performs fully asynchronous local model aggregation based on the client model and updates and generates the server model.
[0098] In an optional embodiment of the present invention, the target upper-layer server performs fully asynchronous server-local model aggregation based on the client model and updates the generated server model, including: the target upper-layer server determines the model aggregation weight based on the version number of the current server model and the version number of the client model; the target upper-layer server performs fully asynchronous server-local model aggregation based on the current server model, the client model and the aggregation weight, updates the generated server model and updates the version number of the server model.
[0099] For example, the model aggregation weights can be expressed by the formula α = G(S) sid -s i +1) Determined. Wherein, the function G(x) is any reasonably reasonable decreasing function, S sid s represents the version number of the current server model in the target upper-layer server. i This refers to the version number of the client model received by the target upper-layer server. The target upper-layer server can use the formula... Update the generated server model. Where S... sid +1 represents the version number of the updated server model. For the current server model in the target upper-layer server, The client model received by the target upper-layer server. To update the generated server model.
[0100] In this embodiment of the invention, when the server performs aggregation of server model and client model, the model aggregation weight is set considering the version number difference between the models. Model aggregation is performed according to the model aggregation weight. This can avoid the impact of client model version uploaded by client device being too low on the global model. At the same time, it can take such slow client model into account in the global model generation, thereby avoiding the global model offset problem caused by the lagging model of slow node and the frequent model update of fast node.
[0101] Step 390: When the target upper-layer server determines that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs server model aggregation and update in a fully asynchronous manner.
[0102] In an optional embodiment of the present invention, when the target upper-layer server determines that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs server model aggregation updates asynchronously. This includes: when the target upper-layer server determines that the number of received client models meets the communication conditions, it queries the connectable neighbor servers to obtain the neighbor update counts of each client device update model maintained by the neighbor servers; the target upper-layer server performs difference statistics based on its own update counts of each client device update model and the obtained neighbor update counts to obtain the difference in the amount of data that the target upper-layer server performs compared to the neighbor servers, and determines the link transmission rate of each neighbor server; the target upper-layer server selects and determines the target neighbor server from the connectable neighbor servers based on the data difference and the link transmission rate.
[0103] If the number of client models received by the target upper-layer server does not reach the preset threshold, the server model, aggregated with the client models, can be directly sent to the corresponding client device. Here, the corresponding client device can specifically refer to the device that uploaded the client models to the target upper-layer server. If the number of client models received by the target upper-layer server reaches the preset threshold, target neighbor servers can be filtered, and after aggregating server models with the target neighbor servers, the resulting server model can be sent to the corresponding client device.
[0104] When filtering target neighbor servers, the target upper-layer server can first query the connected neighbor servers. The neighbor servers can maintain the neighbor update counts for each client device's update model. For example, each server can maintain a set of update counts. Where y represents the number of client devices that have uploaded the client model to this server, and element i y This includes the identifier of the corresponding client device and the cumulative number of updates successfully uploaded by the client device to this server. The target upper-layer server can maintain its own set of update counts, and neighboring servers can maintain a set of neighboring update counts, both of which are related to the aforementioned update count set. Similarly, when the target upper-layer server queries the neighboring server, it can obtain the set of neighbor update counts maintained by the neighboring server.
[0105] The target upper-layer server can calculate the difference between its own update count set and the neighbor's update count set to determine the amount of data exceeding that of the neighboring servers. For example, based on its own update count set and the neighbor's update count set, the target upper-layer server can determine the difference in the number of client model uploads from each client device on the target upper-layer server and the neighboring servers. The target upper-layer service can then use this upload count difference to determine the amount of data exceeding that of the neighboring servers.
[0106] For example, based on the set of update counts, the target upper-layer server can determine that client device m uploads the client model to the target upper-layer server twice and to the neighboring server n three times. Therefore, the target upper-layer server can determine, based on the set of update counts, that the difference in data volume required for model aggregation with the neighboring server n is the amount of data required for one more client device m to upload the client model. Based on this example, the target upper-layer server can determine the difference in data volume compared to the neighboring server.
[0107] Assuming there is only one client m in the system, and its data size is a, then the data size of the neighbor server n relative to the target upper-level server is (3-2)×a. If a neighbor server has client device updates that the target upper-level server has not obtained, then the number of uploads from that client device to the target upper-level server can be considered 0. Before the target upper-level server determines the final selected neighbor server, it only transmits statistical counts, using the total statistical count values of all neighbor servers, such as (3-2)×a above, to generate the data size difference weight value. The data size transmitted after the final selection of the target neighbor server is the size of the parameters of a single model, which the target upper-level server can obtain based on its own model size and does not need to determine.
[0108] The target upper-layer server can also determine the link transmission rate of each neighboring server. For example, each upper-layer server can determine the link transmission rate through network operator planning strategies. These network operator planning strategies include, but are not limited to, Quality of Service (QoS) policies and network communication protocols.
[0109] After determining the data volume difference and link transmission rate on the target upper-layer server, target neighbor servers can be selected based on these factors. For example, a better neighbor server can be selected as the target neighbor server based on the average or weighted average of the data volume difference and link transmission rate. There can be one or more target neighbor servers. The number of target neighbor servers can be determined based on network resource consumption. For example, when network overhead is high, only one target neighbor server can be selected, reducing the amount of information exchanged during model aggregation.
[0110] To select a more suitable target neighbor server, based on the above implementation method, optionally, the target upper-layer server filters and determines the target neighbor server from among the connectable neighbor servers according to the data volume difference and link transmission rate. This includes: the target upper-layer server standardizing each data volume difference and link transmission rate to obtain standard values for the data volume difference and link transmission rate of each connectable neighbor server; the target upper-layer server determining the neighbor selection probability of each connectable neighbor server based on the standard values for the data volume difference and link transmission rate of each connectable neighbor server, as well as a preset importance weight value between the data volume difference and link transmission rate; and the target upper-layer server filtering and determining the target neighbor server from among the connectable neighbor servers according to the neighbor selection probability.
[0111] The Max-Min method can be used to standardize both the data volume difference and the link transmission rate. The formula for the Max-Min method is: In the formula, x can be the difference in data volume or the link transmission rate. min x is the minimum value of the corresponding parameter. max X is the maximum value of the corresponding parameter. norm These are the standard values for the corresponding parameters.
[0112] For example, the target upper-layer server can be configured according to the formula. Determine the neighbor selection probability for each connectable neighbor server. Assume the target upper-layer server has `total` connectable neighbor servers, p i Let μ be the probability of selecting the i-th neighbor server, and μ be the preset importance weight between the data volume difference and the link transmission rate. i Speed is the standard value of the data volume difference between the i-th neighboring server and the i-th neighboring server. i Let be the standard value of the link transmission rate of the i-th neighboring server.
[0113] The target upper-layer server can select the neighbor server with the highest selection probability from each neighbor as the target neighbor server. By selecting the target neighbor server based on the neighbor selection probability, a trade-off can be struck between the difference in data volume and the link transmission rate, allowing for the selection of a better server model for aggregation, thereby improving the model training speed and accuracy.
[0114] After identifying and selecting target neighbor servers, the target upper-layer server can retrieve server models from these neighbor servers. The target upper-layer server can then aggregate its own maintained server models and the retrieved server models to update and generate a new server model. The aggregation method for the two server models can be average aggregation or weighted average aggregation, among others. Directly using average aggregation can balance the client models, avoiding the problem of poor global model adaptability caused by imbalances in client devices within the communication range of each server.
[0115] Step 3100: The target upper-layer server distributes the current server model to the corresponding client device.
[0116] Specifically, when the target upper-layer server communicates with neighboring servers, it can send the server model obtained by aggregating the model with the target neighboring servers to the client device. When the target upper-layer server is not communicating with neighboring servers, it can directly send the server model generated by aggregating the model with the client model to the client device.
[0117] Step 3110, return to step 320, until the server models of each upper-layer server meet the convergence condition, and the global model is obtained.
[0118] The technical solution of this invention involves initializing and generating server models on each upper-layer server and distributing these models to client devices within the communication range. The client devices perform fully asynchronous local training based on the received server models, updating and generating client models. The client devices query connectable upper-layer servers. The connectable upper-layer servers determine the bandwidth allocated to the client devices based on the connection query. The client devices asynchronously estimate the upload rate to each connectable upper-layer server based on the bandwidth and the current wireless channel capacity. Based on the upload rate, the client devices asynchronously select the connectable upper-layer server with the optimal rate as the target upper-layer server. The client devices asynchronously upload the client models to the selected target upper-layer server. The target upper-layer server performs fully asynchronous local model aggregation based on the client models, updating and generating server models. When the target upper-layer server determines that the number of received client models meets the communication requirements, it retrieves models from the selected target neighbor servers and asynchronously performs server model aggregation and updates. The upper-layer server distributes the current server model to the corresponding client device. The client device then performs fully asynchronous local training based on the received server model, updating and generating the client model until the server models of each upper-layer server meet the convergence condition, resulting in a global model. This solves the problems of poor model training accuracy and potential single-node failure of a single server in traditional federated learning, as well as the problems caused by slow client devices. Updating the model through the client device protects user privacy data. Updating the model through both the client device and the server improves model training accuracy. Filtering target upper-layer servers through the client increases network dynamism, allowing client devices to communicate with the server in a timely manner for model training without being bound to a specific server, preventing congestion or failure of a single server. Filtering target neighbor servers through the target upper-layer server for server model aggregation eliminates the need for a central server at the upper layer, thus decentralizing the process and avoiding network congestion and single-node failure of the central server. Furthermore, it eliminates the need to broadcast data to other servers, reducing network communication and resource consumption. Communication between client devices and servers, as well as between servers, is asynchronous. Client devices dynamically select the server to communicate with based on network conditions, avoiding excessive time waste that may occur from waiting for a slow device.
[0119] In the fully asynchronous decentralized hierarchical federated learning method provided in this embodiment of the invention, the local model training of each client device, the uploading of client models, the receiving of client models by each upper-layer server for model aggregation or the distribution of server models, and the aggregation of server models between upper-layer servers are all executed asynchronously. There is no need to wait for slow nodes, and the global model offset problem caused by the deviation between the training progress of fast and slow nodes can be avoided.
[0120] Example 3
[0121] This invention provides a fully asynchronous, decentralized, hierarchical federated learning system. For example... Figure 2 As shown, the hierarchical federated learning system consists of multiple upper-layer servers and multiple lower-layer client devices. Among them:
[0122] The upper-layer server initializes and generates the server model, then distributes it to client devices within the communication range. Client devices perform fully asynchronous local training based on the received server model, update and generate their own client models, and upload these client models to the selected target upper-layer server. The target upper-layer server performs fully asynchronous local model aggregation based on the client models, updating and generating its own server model. When the number of received client models meets the communication requirements, it retrieves models from the selected target neighbor servers and performs fully asynchronous server model aggregation and updates. The current server model is then distributed to the corresponding client device. The client devices and the target upper-layer server repeatedly update their models until the server models of each upper-layer server meet the convergence criteria, resulting in a global model.
[0123] Optionally, the client device is used to query connectable upper-layer servers; the connectable upper-layer servers are used to determine the bandwidth allocated to the client device based on the client device's connection query; the client device is used to asynchronously filter target upper-layer servers based on bandwidth size and upload the client model to the target upper-layer server.
[0124] Optional, connectable upper-layer server, specifically used to determine the current wireless channel capacity, minimum upload time, and the amount of data uploaded by the client device based on the client device's connection query; determine the minimum bandwidth required by the client device based on the channel capacity, minimum upload time, and amount of data uploaded; and determine the bandwidth allocated to the client device when the bandwidth margin is greater than the minimum bandwidth.
[0125] Optionally, the connectable upper-layer server is further specifically used to determine a server update threshold based on the number of times each client device uploads client model updates within the time interval between two communications with the target connected server; determine at least two recording windows based on the server update threshold; and within each recording window, determine the historical average update time interval based on the arrival time interval of the client device uploading client model; determine the update frequency change trend based on each historical average update time interval; determine the server contention factor based on the number of successful connections and the number of connection failures of client devices within the target recording window; and determine the bandwidth allocated to the client device based on the update frequency change trend, the server contention factor, the minimum bandwidth, and the bandwidth margin.
[0126] Optionally, the client device is used to asynchronously estimate the upload rate to each connectable upper-layer server based on the bandwidth and the current wireless channel capacity; and asynchronously select the connectable upper-layer server with the best upload rate as the target upper-layer server.
[0127] Optionally, the target upper-layer server is specifically used to determine the model aggregation weight based on the version number of the current server model and the version number of the client model; based on the current server model, client model, and aggregation weight, it performs fully asynchronous local server model aggregation, updates and generates the server model, and updates the version number of the server model.
[0128] Optionally, the target upper-layer server is specifically used to query connectable neighbor servers when the number of received client models meets the communication conditions, obtain the neighbor update counts of each client device update model maintained by the neighbor servers; perform difference statistics based on the self-update counts of each client device update model maintained by itself and the obtained neighbor update counts to obtain the data volume difference required for model aggregation with the neighbor servers, and determine the link transmission rate of each neighbor server; and select the target neighbor server from among the connectable neighbor servers based on the data volume difference and the link transmission rate.
[0129] Optionally, the target upper-layer server is further specifically used to standardize the differences in data volume and the link transmission rate to obtain standard values for the differences in data volume and the link transmission rate of each connectable neighbor server; based on the standard values for the differences in data volume and the link transmission rate of each connectable neighbor server, as well as the preset importance weight values between the differences in data volume and the link transmission rate, the neighbor selection probability of each connectable neighbor server is determined; based on the neighbor selection probability, the target neighbor server is selected from among the connectable neighbor servers.
[0130] The fully asynchronous decentralized hierarchical federated learning system provided in this embodiment of the invention can execute the fully asynchronous decentralized hierarchical federated learning method provided in any embodiment of the invention, and achieve the same technical effect as the fully asynchronous decentralized hierarchical federated learning method.
[0131] Example 4
[0132] Figure 4 This is a flowchart of a fully asynchronous decentralized hierarchical federated learning method according to Embodiment 4 of the present invention. This embodiment is applicable to model training when protecting user privacy data. This method can be executed by a server in a fully asynchronous decentralized hierarchical federated learning system. Figure 4 As shown, the method includes:
[0133] Step 410: Initialize and generate the server model, and distribute the server model to each client device within the communication range.
[0134] Step 420: Obtain the upload request from the client device, obtain the client model, and perform fully asynchronous local server model aggregation based on the client model to update and generate the server model.
[0135] Optionally, before obtaining the upload request from the client device and obtaining the client model, the method further includes: determining the bandwidth allocated to the client device based on the connection query of the client device, and distributing the bandwidth to the client device so that the client device can select the target upper-layer server for uploading the client model based on the bandwidth.
[0136] Optionally, based on the connection query of the client device, the bandwidth allocated to the client device is determined, including: determining the channel capacity, minimum upload time, and upload data volume of the current wireless channel based on the connection query of the client device; determining the minimum bandwidth required by the client device based on the channel capacity, minimum upload time, and upload data volume; and determining the bandwidth allocated to the client device when the bandwidth margin is greater than the minimum bandwidth.
[0137] Optionally, when the bandwidth margin is greater than the minimum bandwidth, the bandwidth allocated to the client device is determined, including: determining a server update threshold based on the number of times each client device uploads client model updates within the time interval between two communications with the target server; determining at least two recording windows based on the server update threshold; determining the historical average update interval within each recording window based on the arrival time interval of the client model uploaded by the client device; determining the update frequency change trend based on each historical average update interval; determining the server contention factor based on the number of successful connections and the number of connection failures of the client device within the target recording window; and determining the bandwidth allocated to the client device based on the update frequency change trend, the server contention factor, the minimum bandwidth, and the bandwidth margin.
[0138] The client device selects a target upper-layer server based on bandwidth capacity for uploading the client model. The target upper-layer server receives the client model and performs fully asynchronous local model aggregation.
[0139] Optionally, perform fully asynchronous server-local model aggregation based on the client model and update the generated server model, including: determining the model aggregation weight based on the version number of the current server model and the version number of the client model; performing fully asynchronous server-local model aggregation based on the current server model, the client model, and the aggregation weight, updating the generated server model, and updating the version number of the server model.
[0140] Step 430: When it is determined that the number of received client models meets the communication conditions, models are pulled from the selected target neighbor servers, and server model aggregation and update are performed asynchronously.
[0141] Optionally, when it is determined that the number of received client models meets the communication conditions, models are retrieved from the selected target neighbor servers, and server model aggregation and updates are performed asynchronously. This includes: when it is determined that the number of received client models meets the communication conditions, querying the connectable neighbor servers to obtain the neighbor update counts of each client device's updated model maintained by the neighbor servers; based on the self-update counts of each client device's updated model maintained by itself and the obtained neighbor update counts, performing difference statistics to obtain the data volume difference required for model aggregation with the neighbor servers, and determining the link transmission rate of each neighbor server; and based on the data volume difference and the link transmission rate, selecting and determining the target neighbor server from among the connectable neighbor servers.
[0142] Optionally, the target neighbor server is selected from among the connectable neighbor servers based on the data volume difference and the link transmission rate, including: standardizing each data volume difference and the link transmission rate to obtain standard values for the data volume difference and the link transmission rate of each connectable neighbor server; determining the neighbor selection probability of each connectable neighbor server based on the standard values for the data volume difference and the link transmission rate of each connectable neighbor server, as well as a preset importance weight value between the data volume difference and the link transmission rate; and selecting the target neighbor server from among the connectable neighbor servers based on the neighbor selection probability.
[0143] Then, it can pull the server model of the target neighbor server, conduct inter-server communication, aggregate the server model maintained by itself with the pulled server model, and update and generate the server model and its corresponding version number.
[0144] Step 440: Distribute the current server model to the corresponding client device; and return the upload request of the client device to obtain the client model step, until the server models of each upper-layer server meet the convergence condition to obtain the global model.
[0145] The fully asynchronous decentralized hierarchical federated learning method provided in this invention solves the problems of poor model training accuracy and potential single-node failure of a single server in traditional federated learning, as well as the problems caused by slow client devices. It can protect user privacy data, improve model training accuracy, avoid network congestion and single-node failure, and reduce network communication and resource consumption by eliminating the need to broadcast data to other servers.
[0146] Example 5
[0147] Figure 5 This is a flowchart of a fully asynchronous decentralized hierarchical federated learning method according to Embodiment 5 of the present invention. This embodiment is applicable to model training when protecting user privacy data. This method can be executed by a client device in a fully asynchronous decentralized hierarchical federated learning system. Figure 5 As shown, the method includes:
[0148] Step 510: Obtain the server model sent by the upper-layer server.
[0149] Step 520: Perform fully asynchronous local training on the client based on the received server model, update and generate the client model, and upload the client model to the selected target upper-layer server.
[0150] Optionally, the client model can be uploaded to the selected target upper-layer server, including: querying the connectable upper-layer server; obtaining the bandwidth allocated by the connectable upper-layer server; filtering the target upper-layer server based on the bandwidth in a fully asynchronous manner, and uploading the client model to the target upper-layer server.
[0151] Optionally, the target upper-layer server can be selected asynchronously based on bandwidth, including: asynchronously estimating the upload rate of each connectable upper-layer server based on bandwidth and the current wireless channel capacity; and asynchronously selecting the connectable upper-layer server with the best upload rate as the target upper-layer server.
[0152] The client device can upload its client model to the target upper-layer server and then wait for the target upper-layer server to send the server model. When uploading the client model, the client device can simultaneously upload the client model's version number. This client model version number can be the version number of the server model used by the client device when training and generating the client model.
[0153] Step 530: Obtain the server model sent by the target upper-layer server; and return to perform fully asynchronous local training on the client based on the received server model, update the client model generation steps, until the server model is obtained as the global model.
[0154] The fully asynchronous decentralized hierarchical federated learning method provided in this invention solves the problems of poor model training accuracy and potential single-node failure of a single server in traditional federated learning, as well as the problems caused by slow client devices. It can protect user privacy data, improve model training accuracy, and avoid network congestion and single-node failure.
[0155] Example 6
[0156] Figure 6A schematic diagram of a server or client device 10 that can be used to implement embodiments of the present invention is shown. The server or client device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The server or client device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0157] like Figure 6 As shown, the server or client device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the server or client device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0158] Multiple components in server or client device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows server or client device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0159] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a fully asynchronous decentralized hierarchical federated learning approach.
[0160] In some embodiments, the fully asynchronous decentralized hierarchical federated learning method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on server or client device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the fully asynchronous decentralized hierarchical federated learning method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the fully asynchronous decentralized hierarchical federated learning method by any other suitable means (e.g., by means of firmware).
[0161] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0162] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0163] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0164] To provide interaction with a user, the systems and techniques described herein can be implemented on a server or client device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the server or client device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0165] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0166] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0167] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0168] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A fully asynchronous decentralized hierarchical federated learning method, characterized in that, include: Each upper-layer server initializes and generates a server model, and then distributes the server model to each client device within the communication range; The client device performs fully asynchronous local training based on the received server model, updates and generates a client model, and uploads the client model to the selected target upper-layer server; wherein, the connection status between the upper-layer server and the lower-layer client device is not fixed; The target upper-layer server performs fully asynchronous local model aggregation based on the client model and updates and generates the server model. When the target upper-layer server determines that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs server model aggregation and update in a fully asynchronous manner; the connection status between each upper-layer server is not fixed. The target upper-layer server distributes the current server model to the corresponding client device; The client device then performs fully asynchronous local training based on the received server model, updating and generating the client model, until the server models of each upper-layer server meet the convergence condition, thus obtaining the global model.
2. The method according to claim 1, characterized in that, Uploading the client model to the selected target upper-layer server includes: The client device queries the available upper-layer servers; The connectable upper-layer server determines the amount of bandwidth allocated to the client device based on the client device's connection query. The client device asynchronously filters the target upper-layer server based on the bandwidth size and uploads the client model to the target upper-layer server.
3. The method according to claim 2, characterized in that, The connectable upper-layer server determines the bandwidth allocated to the client device based on the client device's connection query, including: The connectable upper-layer server determines the current wireless channel capacity, minimum upload time, and the amount of data the client device can upload based on the connection query of the client device; The connectable upper-layer server determines the minimum bandwidth required by the client device based on the channel capacity, minimum upload time, and amount of data to be uploaded; When the available bandwidth of the upper-layer server is greater than the minimum bandwidth, the server determines the amount of bandwidth to allocate to the client device.
4. The method according to claim 3, characterized in that, When the available bandwidth from the upper-layer server is greater than the minimum bandwidth, the server determines the bandwidth allocated to the client device, including: The connectable upper-layer server determines the server update threshold based on the number of times each client device uploads updates to the client model within the time interval between two communications with the target server. The connectable upper-layer server determines at least two recording windows based on the server update threshold; and within each recording window, it determines the historical average update interval based on the arrival time interval of the client model uploaded by the client device. The connectable upper-layer server determines the trend of update frequency changes based on the historical average update time intervals; The connectable upper-layer server determines the server contention factor based on the number of successful and failed connections of the client device during the target record window period; The connectable upper-layer server determines the bandwidth allocated to the client device based on the update frequency change trend, the server contention factor, the minimum bandwidth, and the bandwidth margin.
5. The method according to any one of claims 2-4, characterized in that, The client device asynchronously filters target upper-layer servers based on the bandwidth size, including: The client device asynchronously estimates the upload rate to each connectable upper-layer server based on the bandwidth and the current wireless channel capacity. The client device asynchronously selects the upstream server with the best connection speed as the target upstream server based on the upload speed.
6. The method according to claim 1, characterized in that, The target upper-layer server performs fully asynchronous server-local model aggregation based on the client model, and updates and generates the server model, including: The target upper-layer server determines the model aggregation weight based on the version number of the current server model and the version number of the client model; The target upper-layer server performs fully asynchronous local model aggregation based on the current server model, client model, and the aggregation weight, updates and generates the server model, and updates the version number of the server model.
7. The method according to claim 1, characterized in that, When the target upper-layer server determines that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs server model aggregation and update asynchronously, including: When the target upper-layer server determines that the number of received client models meets the communication conditions, it queries the connectable neighbor servers to obtain the neighbor update counts of each client device update model maintained by the neighbor servers. The target upper-layer server calculates the difference between the target upper-layer server and the neighboring servers by performing differential statistics based on the number of updates of the update models of each client device maintained by itself and the number of updates of the neighbors, and determines the link transmission rate of each neighboring server. The target upper-layer server selects the target neighbor server from among the connectable neighbor servers based on the difference in data volume and the link transmission rate.
8. The method according to claim 7, characterized in that, The target upper-layer server determines the target neighbor server from among the connectable neighbor servers based on the data volume difference and the link transmission rate, including: The target upper-layer server standardizes the data volume difference and the link transmission rate of each of the aforementioned data volume differences and the link transmission rate of each connectable neighbor server. The target upper-layer server determines the neighbor selection probability of each connectable neighbor server based on the standard value of the difference in data volume between each connectable neighbor server, the standard value of the link transmission rate, and the preset importance weight value between the difference in data volume and the link transmission rate. The target upper-layer server selects the target neighbor server from among the connectable neighbor servers based on the selection probability of each neighbor.
9. A fully asynchronous decentralized hierarchical federated learning system, characterized in that, The hierarchical federated learning system consists of multiple upper-layer servers and multiple lower-layer client devices; wherein: The upper-layer server is used to initialize and generate the server model, and then distribute the server model to each client device within the communication range. The client device is used to perform fully asynchronous local training based on the received server model, update and generate the client model, and upload the client model to the selected target upper-layer server; wherein, the connection status between the upper-layer server and the lower-layer client device is not fixed; The target upper-layer server is used to perform fully asynchronous local server model aggregation based on the client model and update the generated server model; when it is determined that the number of received client models meets the communication conditions, it pulls models from the selected target neighbor servers and performs fully asynchronous server model aggregation and update; and sends the current server model to the corresponding client device; wherein, the connection status between each upper-layer server is not fixed. The client device and the target upper-layer server repeatedly update the model until the server model of each upper-layer server meets the convergence condition, thus obtaining the global model.
10. A fully asynchronous decentralized hierarchical federated learning method, characterized in that, The method is applied to an upper-layer server, and the method includes: Initialize and generate a server model, and then distribute the server model to each client device within the communication range; The process involves obtaining an upload request from a client device, obtaining a fully asynchronous local training process performed by the client device based on the received server model, updating and uploading the client model, and then performing fully asynchronous local server model aggregation based on the client model to update and generate a server model. When uploading the client model, the client device uploads it to a selected target upper-layer server. The connection status between the upper-layer server and the lower-layer client device is not fixed. Once the number of received client models meets the communication requirements, models are retrieved from the selected target neighbor servers, and server model aggregation and updates are performed asynchronously; the connection status between each upper-layer server is not fixed. The current server model is sent to the corresponding client device; and the upload request from the client device is returned. The client device performs fully asynchronous local training based on the received server model, updates the generated and uploaded client model, and repeats this process until the server models of each upper-layer server meet the convergence condition to obtain the global model.
11. A fully asynchronous decentralized hierarchical federated learning method, characterized in that, The method is applied to a client device, and the method includes: Obtain the server model sent by the upper-layer server; The client-side model is trained asynchronously based on the received server model, updated and generated, and then uploaded to the selected target upper-layer server. The connection status between the upper-layer server and the lower-layer client device is not fixed; the connection status between the upper-layer servers is also not fixed. The server model sent by the target upper-layer server is obtained through the following steps: The target upper-layer server performs fully asynchronous local server model aggregation based on the client model and updates and generates the server model; when it is determined that the number of received client models meets the communication conditions, the model is pulled from the selected target neighbor servers and the server model is aggregated and updated in a fully asynchronous manner; the current server model is then sent to the corresponding client device; wherein, the connection status between each upper-layer server is not fixed. Return to the step of performing fully asynchronous local training on the client based on the received server model, updating and generating the client model, until the server model is obtained as the global model.
Citation Information
Patent Citations
Model training method and device and electronic equipment
CN116957107A
Distributed federated learning-oriented equipment scheduling method and device, medium and equipment
CN117014962A