A multifunctional drone-assisted asynchronous cluster personalized federated learning method
Through the asynchronous cluster personalized federated learning method assisted by multifunctional drone, deep reinforcement learning is used to optimize the drone flight path, solving the problems of device data heterogeneity and communication efficiency in traditional federated learning, and achieving efficient model training and communication efficiency improvement.
Patent Information
- Application Number
- CN202411558130.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-11-04
AI Technical Summary
There are problems with device data heterogeneity and communication efficiency in traditional federated learning. Although asynchronous federated learning improves training speed, it brings about the impact of model outdatedness. The existing drone assistance solutions have failed to effectively solve the challenges of model outdatedness and communication efficiency at the group level.
Multifunctional drones are adopted as edge servers, combined with deep reinforcement learning, optimize the drone's flight path, design dual-layer parallel optimization problems, and reduce the outdated model at the group level and improve communication efficiency through LoS communication channels and asynchronous training mechanisms.
It realizes more efficient model training in data heterogeneous scenarios, reduces the outdated model at the group level, improves communication efficiency and resource utilization efficiency, and adapts to diversified task requirements.
Smart Images

Figure CN119440048B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless communication technology and relates to a multifunctional drone-assisted asynchronous cluster personalized federated learning method. Background Art
[0002] Federated learning (FL) is an emerging distributed machine learning paradigm that has attracted widespread attention from industry and academia for its ability to train global models for all participating devices while protecting user data privacy. However, in practical applications, traditional FL faces challenges in device data heterogeneity and communication efficiency.
[0003] Data heterogeneity stems from differences in device attributes, preferences, and data collection patterns. To address this challenge, personalized federated learning (PFL) was proposed. It aims to develop personalized solutions for each device, address data variability, and improve model performance. Communication efficiency is a challenge because traditional federated learning requires waiting for all devices to complete training and upload their models before proceeding to the next round of global training. This forces faster-training devices to wait for slower ones, wasting time. To address this issue, asynchronous federated learning (AFL) was developed. AFL allows devices to immediately advance to the next round after completing training, significantly reducing training time. However, AFL also introduces the problem of model staleness. Devices upload models at different times, causing some devices to update based on outdated global models, impacting model aggregation and communication efficiency. Existing research has primarily addressed this issue indirectly by adjusting model update strategies, rather than directly improving communication efficiency.
[0004] Due to their high maneuverability and LoS communication links, drones are widely used in fields such as wireless communications, cargo delivery, data collection, and edge computing. To address these issues, drones can be used to improve communication efficiency between users and servers, accelerate model transmission, and reduce communication latency. However, existing drone-assisted AFL solutions do not address how to reduce the impact of model staleness. To improve the performance of AFL in heterogeneous data scenarios, APFL has been proposed. However, these studies primarily focus on designing personalized solutions and improving communication efficiency, lack research on the impact of model staleness, and most remain at the individual level, ignoring the challenges of group data heterogeneity. Furthermore, with the development of science and technology and the maturity of drone technology, single-function drones have limitations in multi-tasking, resource utilization, and cost-effectiveness, making it difficult to meet diverse needs. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide an asynchronous cluster personalized federated learning method assisted by a multifunctional UAV, aiming to utilize the high maneuverability of the multifunctional UAV to improve the communication efficiency of the FL system and reduce the average staleness of the model at the group level.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A multifunctional drone-assisted asynchronous cluster personalized federated learning method includes the following steps:
[0008] S1: Establish a multifunctional UAV-assisted FL network system model, which includes a multifunctional UAV equipped with communication, computing modules and a cargo hold as the edge server of the network system, a ground base station and the server forming a sub-server, and user devices associated with the same sub-server forming a group;
[0009] S2: Establish a multifunctional UAV model that includes logistics transportation and assisting ground user groups for FL training. At the same time, establish a LoS wireless communication channel model between the UAV and the sub-server and a simplified wireless communication model between the sub-server and the associated user group;
[0010] S3: Combined with S2, a model of downlink communication delay, model training time, and uplink communication delay between the multifunctional UAV and the sub-server is established;
[0011] S4: Establish a federated learning mechanism to asynchronously train personalized models for user groups, targeting data heterogeneity at the user group level. Based on a two-layer parallel optimization problem where the inner layer solves the personalized model and the outer layer solves the global model, the personalized model is trained for each group using asynchronous two-layer parallel optimization between groups and synchronous federated averaging within the group.
[0012] S5: Under the constraints of cargo delivery, drone flight, and training time slots, the staleness of the swarm model is used as the optimization variable to construct an optimization problem and optimize the flight path of the multifunctional drone;
[0013] S6: Based on the relationship between model staleness and the number of times the user device completes the training task, the optimization variables in S5 are modified while keeping the constraints unchanged. The transformed optimization problem is an integer nonlinear optimization problem. A dynamic optimization algorithm for drone paths based on deep reinforcement learning (DRL) is proposed.
[0014] Furthermore, in S1, a multifunctional UAV-assisted FL network system model is established, including the following processes:
[0015] The FL network system model includes a UAV equipped with communication, computing modules and cargo compartment, M ground base stations and sub-servers consisting of servers, and N users. The set of sub-servers is represented as The set of users is represented as The ML task is performed in a FL network assisted by a multifunctional UAV. The multifunctional UAV acts as a mobile base station and edge server, connected to the sub-server on the ground through a wireless channel. The multifunctional UAV maintains a safe height H sf , fly from the starting point to the destination, and deliver the goods to O receiving points, The coordinates are in Represents the collection of these receiving points; each sub-server are all located on the horizontal plane, and their coordinates are Each user Equipped with an antenna located on a horizontal plane; each sub-server Serving a group of users through wireless channels, this group of users constitutes a user group in Use a binary random variable X m,n ∈{0,1} represents the users in group m Status, X m,n =1 means user device n can participate in model training, X m,n =0 means that user device n cannot participate in model training; define the user in group m The probability of participating in model training is Group User devices in The dataset used is represented as The total data set of group m is then expressed as And the data distribution between different groups m is non-IID at the group level.
[0016] Furthermore, in S2, a multifunctional UAV model including logistics transportation and assisting ground user groups in FL training is established, including the following processes:
[0017] The flight time of the multifunctional UAV is divided into K time slots, using represents the index of the flight time slot, the length of each flight time slot is τ; the maximum flight distance of each time slot of the multifunctional UAV is D max , the horizontal projection of the flight area is Since the cargo delivery task has success or failure, a binary random variable is used Indicates the delivery point The delivery status of goods, including Indicates that the goods have been delivered successfully. Indicates that the goods have not been delivered. Assuming that the drone descends vertically from the delivery point o during delivery, and the time spent in the middle is not included in the flight time slot, the multifunctional drone needs to complete the delivery of all goods before the end of the flight. The delivery status of the goods in the Kth flight time slot should meet the following constraints:
[0018]
[0019] The flight path length of the multi-function UAV is limited by its own battery capacity. Assume that the maximum flight path length of the UAV is Q max time slot; the drone is in The horizontal projection coordinates of the time slot are Indicates that the trajectory is expressed as express.
[0020] Furthermore, in S2, a LoS wireless communication channel model between the multifunctional UAV and the sub-server is established, which includes the following processes:
[0021] In each time slot k, the sub-server accesses the multifunctional UAV through orthogonal frequency division multiplexing, with a total bandwidth of B, and each sub-server gets an average bandwidth of B m =B / N; Assuming that the wireless communication channel between the multifunctional UAV and the sub-server is a LoS channel and the channel gain at 1 meter is β0, the channel gain between the multifunctional UAV and the sub-server m is as follows:
[0022]
[0023] Furthermore, in S2, the wireless communication model between the sub-server and the associated user group includes the following process:
[0024] Assume that in each group iteration, the sum of the downlink communication delay between the sub-server and the user, the user's local update time, and the uplink communication delay is less than τ e .
[0025] Furthermore, in S3, a downlink communication delay, model training time, and uplink communication delay model between the multifunctional UAV and the sub-server are established, which includes the following processes:
[0026] Use p u [k] represents the transmission power of the multi-functional UAV, p m [k] represents the transmission power of the sub-server, N0 represents the noise power, and combined with S2, the downlink communication delay, model training time, and uplink communication delay between the multi-functional UAV and the sub-server are modeled. The model includes the following sub-steps:
[0027] S31: Establish a downlink communication delay model, the content of which is as follows:
[0028] At the beginning of each training time slot k, the multi-function UAV broadcasts the latest global model to all sub-servers; assuming the broadcast channel bandwidth is B broadcast , then in the training time slot k, the downlink communication data rate between the multifunctional UAV and the sub-server m is:
[0029]
[0030] Assuming that the number of parameters of the ML model is s, where each parameter is stored as a b-bit floating-point number, the downlink communication delay is:
[0031]
[0032] S32: Establish a model training time model, the content is as follows:
[0033] Assume that the time required for each group iteration is τ e , then in the training time slot k, the model training time required for each global iteration of the idle user group and the associated sub-server m is:
[0034]
[0035] Where E is the number of group iterations;
[0036] S33: Establish an uplink communication delay model, the content of which is as follows:
[0037] In training time slot k, the uplink communication rate of sub-server m is:
[0038]
[0039] The uplink communication delay is:
[0040]
[0041] Furthermore, in S4, a federated learning mechanism is established to asynchronously train personalized models for user groups based on data heterogeneity at the user group level, including the following steps:
[0042] S41: Establish a two-level parallel optimization problem for asynchronous personalized federated learning assisted by multifunctional drones. The content is as follows:
[0043] Establish the following optimization problem P1, aiming to find its optimal solution w * :
[0044] P1:
[0045] Among them, u m represents the personalized model of group m, w represents the global model, and F m(w) is defined as the following optimization problem P2:
[0046] P2:
[0047] Among them, λ is a regularization parameter used to control the personalized model u m The difference from the global model w; f m (u m ) represents the training loss of all users in group m The sum of f n (u m ) represents the personalized model u m On a local privacy dataset The loss function on , and F m The definition of (w) is related to the Moreau envelope;
[0048] In the two-level parallel optimization problem represented by optimization problems P1 and P2, P1 is the outer optimization problem for solving the global model w, and P2 is the outer optimization problem for solving the personalized model u. m The inner optimization problem of ,these two optimization problems are decoupled;
[0049] S42: The group model and the personalized model are synchronously updated to obtain the personalized model and the group model, the contents of which are as follows:
[0050] Use e∈E to represent the index of E group iterations, where is the index set of E group iterations; at the beginning of each group iteration, sub-server m broadcasts the latest group model to all members in group m who can participate in training The user n participating in the training is calculated by performing gradient descent on the optimization problem P2 When all users participating in the training have completed the training, the sub-server m aggregates the user models of all users participating in the training To generate a personalized model; after the e∈Eth group iteration, the group model is updated by performing gradient descent on the optimization problem P1; finally, after E group iterations, the obtained group model is uploaded to the multi-functional UAV, completing the group training task after one training time slot update;
[0051] S43: Global model asynchronous update, the content is as follows:
[0052] Before performing asynchronous update of the global model, the staleness of the group model uploaded by the sub-server is first analyzed. In the example, we use a binary random variable π m [k]∈{0,1} represents a group Whether the training task is completed in this time slot, the training task status of the group is updated on the multifunctional drone, where π m [k] = 1 means that group m has completed this task and is ready for subsequent training tasks, and π m [k]=0 means other situations; sequence Contains group m in all training time slots The state of; the set of user groups that can participate in model training in training time slot k Including all π m [k-1] = 1 group; use γ m [k] represents the group model uploaded by sub-server m in training time slot k If the group model does not exist, then its corresponding γ m [k] does not exist, group model Obsolescence m [k] is updated as follows:
[0053]
[0054] Among them, k m is the index of the global model used by the collaborative training model of sub-server m and the associated user group m, N / A indicates staleness γ m [k] does not exist, it is not difficult to find that if π m [k]=1, then γ m [k] represents the sequence π m The time slot k is the same as the previous π m = the number of zeros between time slots of 1; if we define 0×N / A=0, and use the integer ν k ≥0 represents the sequence π m The number of trailing zeros gives the following formula:
[0055]
[0056] Combined with the above formula, the average staleness of the local model uploaded by group m is Satisfies the following inequality:
[0057]
[0058] The larger the value of The smaller it is, that is, the more times group m completes the training task, the smaller the average staleness of the local model uploaded by group m; combining the downlink communication delay, model training time and uplink communication delay model in S3 to update the group training task completion state sequence π m ; Finally, the global model is obtained by aggregation update through the following formula:
[0059]
[0060] Among them, g(γ m [k]) is a dynamic adaptive discount factor used to reduce the impact of stale group models on the global model. The function g(·) is defined as g(x) = (x + 1) -a , a is a positive constant; parameter β>0 is used to control the model aggregation to the global model w k The smoothing coefficient of the influence degree.
[0061] Furthermore, in S5, under the constraints of the number of cargo delivery, drone flight, and training time slots, an optimization problem of minimizing the staleness of the group model is constructed to optimize the flight path of the multifunctional drone, including the following process:
[0062] When the multifunctional UAV is in different locations, the uplink and downlink communication delays between it and the sub-server are different, that is, the time required for the group to complete a training task is inconsistent. A multifunctional UAV path optimization problem that minimizes the average staleness of the group model is constructed. By optimizing the flight path of the multifunctional UAV, the average staleness of all group models is minimized while completing the cargo delivery task. The mean
[0063] Furthermore, in said S6, according to the average obsolescence of the model in S4 and the number of group training tasks The relationship between the optimization variables will be minimized, that is, the average staleness of all group models The solution is modified to maximize the number of completed group training tasks while keeping the constraints of the original optimization problem unchanged.
[0064] Furthermore, in S6, a UAV path dynamic optimization algorithm based on deep reinforcement learning (DRL) is proposed to solve the modified optimization problem, which includes the following steps:
[0065] A1: To implement the DRL method, it is necessary to simplify and discretize the flight actions of the multi-functional UAV and divide the flight area of the multi-functional UAV into Modeled as a A square grid map, where the size of each grid is D max , flight area in The multi-function UAV can only fly in four directions: "East", "South", "West" and "North", and the flight distance each time is D max , that is, each time slot can only fly from one grid to four adjacent grids;
[0066] A2: Establish the DRL model action space;
[0067] A3: Establish a state space consisting of environmental information and multifunctional drone state, where the environmental information includes the starting point coordinate q start , end point coordinate q end and multi-function drones to the border The state of the multifunctional UAV includes the position q[k] of the current time slot, the number of time slots flown k, the position two time slots ago, and the cargo delivery status.
[0068] A3: Construct the reward function of the DRL model to ensure that the multi-functional drone completes the delivery of all goods and the path length is within the specified range. The reward function should include the round reward R episode , Safety Reward R safe , Mobile Reward R move , Group training task reward R π and cargo delivery incentives;
[0069] A4: The DQN algorithm is used to solve the flight path of the multi-functional UAV. In each round of each iteration, there are three parts: the interaction between the multi-functional UAV and the environment, experience replay, and the DQN algorithm with the target Q network. The third part of the process will be iterated multiple times until the end of the round or the maximum number of iterations is reached.
[0070] The beneficial effects of the present invention are:
[0071] (1) This invention breaks away from the inherent individual-level personalized design and proposes a group-level PFL solution that is more suitable for practical task requirements. At the same time, traditional PFL solutions do not consider reducing the obsolescence of user training or the impact of obsolescence. Therefore, from this perspective, this invention proposes an asynchronous cluster personalized federated learning method that takes into account communication efficiency and data heterogeneity.
[0072] (2) In the asynchronous cluster personalized federated learning method assisted by a multifunctional UAV proposed in this invention, a multifunctional UAV is used to assist the ground user group and the sub-server base station in training the FL network. Compared with fixed cloud servers, multifunctional UAVs are more flexible and can adapt to changing communication channels. At the same time, they can also complete multiple tasks and improve resource utilization efficiency.
[0073] (3) In the multifunctional drone-assisted asynchronous cluster personalized federated learning method proposed in the present invention, the proposed path optimization algorithm designs a reward function composed of multiple parts, among which the group training task reward is mainly used to motivate drones to find paths with lower staleness, and reduce model staleness by improving the average uplink communication rate between sub-services, thereby effectively improving communication efficiency.
[0074] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0076] Figure 1 A detailed flowchart of the personalized federated learning method for asynchronous clusters assisted by multifunctional drones;
[0077] Figure 2 A practical application scenario diagram of the asynchronous cluster personalized federated learning method assisted by multifunctional drones;
[0078] Figure 3 Grid graphs of the UAV flight paths generated by two comparative path optimization algorithms and the algorithm proposed in this invention;
[0079] Figure 4 for Figure 3 Scatter plot of the staleness of paths generated by the comparison algorithm in and the algorithm proposed in the present invention;
[0080] Figure 5 As the time slot increases, the drone moves along Figure 3 The average uplink transmission rate of users when flying over the paths generated by the comparison algorithm and the algorithm proposed in the present invention;
[0081] Figure 6 for Figure 3 The reward curves of the three path optimization algorithms;
[0082] Figure 7 Average test accuracy of the proposed method and the comparative method under different device availability when the number of devices is fixed, the device group division is fixed, and the device data is non-independent and identically distributed (non-IID). DETAILED DESCRIPTION
[0083] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the features in the following embodiments and embodiments can be combined with each other without conflict.
[0084] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0085] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0086] like Figure 1 and Figure 2 As shown in FIG, a multifunctional drone-assisted asynchronous cluster personalized federated learning method includes the following steps:
[0087] S1: Establish a multifunctional UAV-assisted FL network system model, which includes a multifunctional UAV equipped with communication, computing modules and a cargo hold as the edge server of the network system, a ground base station and the server forming a sub-server, and user devices associated with the same sub-server forming a group;
[0088] S2: Establish a multifunctional UAV model that includes logistics transportation and assisting ground user groups for FL training. At the same time, establish a LoS wireless communication channel model between the UAV and the sub-server and a simplified wireless communication model between the sub-server and the associated user group;
[0089] S3: Combined with S2, a model of downlink communication delay, model training time, and uplink communication delay between the multifunctional UAV and the sub-server is established;
[0090] S4: Establish a federated learning mechanism to asynchronously train personalized models for user groups, targeting data heterogeneity at the user group level. This mechanism is based on a two-layer parallel optimization problem: an inner layer solves the personalized model, and an outer layer solves the global model. It utilizes asynchronous two-layer parallel optimization between groups and a synchronous federated averaging mechanism within groups to train personalized models for each group.
[0091] S5: Under the constraints of cargo delivery, drone flight, and training time slots, the staleness of the swarm model is used as the optimization variable to construct an optimization problem and optimize the flight path of the multifunctional drone;
[0092] S6: Based on the relationship between model staleness and the number of times the user device completes training tasks, the optimization variables in S5 are modified, while keeping the constraints unchanged. The resulting optimization problem is an integer nonlinear optimization problem. This step will analyze its structural characteristics and propose a dynamic optimization algorithm for drone paths based on deep reinforcement learning (DRL).
[0093] Specifically, a multifunctional UAV-assisted FL network system model is established in S1, including the following processes:
[0094] The network consists of a UAV equipped with communication, computing modules and a cargo hold, M ground base stations and sub-servers consisting of servers, and N users. The set of sub-servers is represented as The set of users is represented as The ML task is performed in a FL network assisted by a multifunctional UAV. The multifunctional UAV acts as a mobile base station and edge server, connected to the sub-server on the ground through a wireless channel. The multifunctional UAV maintains a safe height H sf , fly from the starting point to the destination, and deliver the goods to O receiving points, The coordinates are in Represents the collection of these receiving points; each sub-server are all located on the horizontal plane, and their coordinates are Each user Equipped with an antenna located on a horizontal plane; each sub-server Serving a group of users through wireless channels, this group of users constitutes a user group in Due to unstable network connection, device calls or low battery, users in each group are not always available to participate in model training. A binary random variable X m,n ∈{0,1} represents the users in group m The state, specifically, X m,n=1 means user device n can participate in model training, X m,n =0 means that user device n cannot participate in model training; define the user in group m The probability of participating in model training is User devices in group m The dataset used is represented as Thus the total data set of group m can be expressed as And the data distribution between different groups m is non-independent and identically distributed (non-IID) at the group level;
[0095] Specifically, a multifunctional UAV model is established in S2, which includes logistics transportation and assisting ground user groups in FL training. The process includes the following:
[0096] The flight time of the multifunctional UAV is divided into K time slots, using represents the index of the flight time slot, the length of each flight time slot is τ; the maximum flight distance of each time slot of the multifunctional UAV is D max , the horizontal projection of the flight area is Since the cargo delivery task has success or failure, a binary random variable is used Indicates the delivery point The delivery status of goods, including Indicates that the goods have been delivered successfully. Indicates that the goods have not been delivered. Assuming that the drone descends vertically from the delivery point o during delivery, and the time spent in the middle is not included in the flight time slot, the multifunctional drone needs to complete the delivery of all goods before the end of the flight. The delivery status of the goods in the Kth flight time slot should meet the following constraints:
[0097]
[0098] Because the flight path length of the multi-function UAV is limited by its own battery capacity, it is assumed that the UAV can fly at most Q max time slot; the drone is in The horizontal projection coordinates of the time slot are Indicates that the trajectory is expressed as The flight constraints of the multifunctional UAV can be expressed as follows:
[0099]
[0100] C3:q start =q[0]
[0101] C4:q end =q[K]
[0102]
[0103] Among them, q in constraints C3 and C4 start and q end denote the projection coordinates of the starting point and the end point respectively; assuming that the ML model is trained to deliver O min The kth delivery point after the goods start >0 flight slots start and last K train Flight time slots; define the drone delivery time slot min The time slot of the goods at the receiving point is k min , the index set of training time slots is expressed as In the training time slot The idle group set is
[0104] Specifically, S2 establishes a LoS wireless communication channel model between the multifunctional UAV and the sub-server, which includes the following processes:
[0105] In each time slot k, the sub-server accesses the multifunctional UAV through orthogonal frequency division multiplexing, with a total bandwidth of B, and each sub-server gets an average bandwidth of B m =B / N; Assuming that the wireless communication channel between the multifunctional UAV and the sub-server is a LoS channel and the channel gain at 1 meter is β0, the channel gain between the multifunctional UAV and the sub-server m is as follows:
[0106]
[0107] Specifically, the wireless communication model between the S2 neutron server and the associated user group includes the following processes:
[0108] Since the communication distance between the sub-server and the associated user group is shorter than that between the multi-function drone and the sub-server, this step will simplify the wireless communication model between the sub-server and the associated user group; it is assumed that in each group iteration, the sum of the downlink communication delay between the sub-server and the user, the user local update time and the uplink communication delay is less than τ e ;
[0109] Specifically, S3 establishes the downlink communication delay, model training time, and uplink communication delay models between the multifunctional drone and the sub-server, which includes the following processes:
[0110] This step uses p u [k] represents the transmission power of the multi-functional UAV, p m [k] represents the transmission power of the sub-server, N0 represents the noise power, and combined with S2, the downlink communication delay, model training time, and uplink communication delay between the multi-functional UAV and the sub-server are modeled. The model includes the following sub-steps:
[0111] S31: Establish a downlink communication delay model, the content of which is as follows:
[0112] At the beginning of each training time slot k, the multi-function UAV broadcasts the latest global model to all sub-servers; assuming the broadcast channel bandwidth is B broadcast , then in the training time slot k, the downlink communication data rate between the multifunctional UAV and the sub-server m is:
[0113]
[0114] Assuming that the number of parameters of the ML model is s, where each parameter is stored as a b-bit floating-point number, the downlink communication delay is:
[0115]
[0116] S32: Establish a model training time model, the content is as follows:
[0117] Assume that the time required for each group iteration is τ e , then in the training time slot k, the model training time required for each global iteration of the idle user group and the associated sub-server m is:
[0118]
[0119] Where E is the number of group iterations;
[0120] S33: Establish an uplink communication delay model, the content of which is as follows:
[0121] In training time slot k, the uplink communication rate of sub-server m is:
[0122]
[0123] The uplink communication delay is:
[0124]
[0125] Specifically, S4 establishes a federated learning mechanism that asynchronously trains personalized models for user groups based on data heterogeneity at the user group level. This includes the following steps:
[0126] S41: Establish a two-level parallel optimization problem for asynchronous personalized federated learning assisted by multifunctional drones. The specific content is as follows:
[0127] Establish the following optimization problem P1, aiming to find its optimal solution w * :
[0128] P1:
[0129] Among them, u m represents the personalized model of group m, w represents the global model, and F m (w) is defined as the following optimization problem P2:
[0130] P2:
[0131] Among them, λ is a regularization parameter used to control the personalized model u m Difference from the global model w. f m (u m ) represents the training loss of all users in group m The sum of f n (u m ) represents the personalized model u m On a local privacy dataset The loss function on , and F m The definition of (w) is related to the Moreau envelope;
[0132] In the two-level parallel optimization problem represented by optimization problems P1 and P2, P1 is the outer optimization problem for solving the global model w, and P2 is the outer optimization problem for solving the personalized model u. m The inner optimization problem of ,these two optimization problems are decoupled;
[0133] S42: The group model and the personalized model are synchronously updated to obtain a personalized model and a group model. The specific contents are as follows:
[0134] Use e∈E to represent the index of E group iterations, where is the index set of E group iterations; definition Indicates the The global model after the updated training time slot, represents the updated group model for the kth training slot, represents the personalized model after the kth training time slot update, The group model after the e-th group iteration of training time slots is expressed as Before the start of the e∈E group iteration, initialize When the group iteration starts, the child server m will update the latest group model Broadcast to all participants in group m The user n participating in the training calculates the user model by performing H steps of gradient descent As shown in the following formula:
[0135]
[0136] in, ηt represents the step length, f n (·) represents the user loss function; the sub-server m then aggregates the user models of all users in the associated group that can participate in training through the following formula To generate a personalized model:
[0137]
[0138] After the e∈Eth group iteration, the group model is calculated by performing gradient descent on the optimization problem P1 to update the group model, as shown in the following formula:
[0139]
[0140] Get the updated group model; finally, after E group iterations, get the personalized model and group models And the group model Upload to a multi-function drone;
[0141] S43: Global model asynchronous update, the specific content is as follows:
[0142] Before performing asynchronous update of the global model, the staleness of the group model uploaded by the sub-server is first analyzed. In the example, we use a binary random variable π m [k]∈{0,1} represents a group Whether the training task is completed in this time slot, the training task status of the group is updated on the multifunctional drone, where π m [k] = 1 means that group m has completed this task and is ready for subsequent training tasks, and π m [k]=0 means other situations; sequence Contains group m in all training time slots The state of; the set of user groups that can participate in model training in training time slot k Including all π m [k-1] = 1 group; use γ m [k] represents the group model uploaded by sub-server m in training time slot k If the group model does not exist, then its corresponding γ m [k] does not exist, group model Obsolescence m [k] is updated as follows:
[0143]
[0144] Here, we define π m[k star -1]=1 and k m is the index of the global model used by the collaborative training model of sub-server m and the associated user group m, N / A indicates staleness γ m [k] does not exist, it is not difficult to find that if π m [k]=1, then γ m [k] represents the sequence π m The time slot k is the same as the previous π m = the number of zeros between time slots of 1; if we define 0×N / A=0, and use the integer ν k ≥0 represents the sequence π m The number of trailing zeros can be obtained as follows:
[0145]
[0146] Combined with the above formula, the average staleness of the local model uploaded by group m is Satisfies the following inequality:
[0147]
[0148] From the above formula, we can observe that The larger the value of The smaller it is, that is, the more times group m completes training tasks, the smaller the average staleness of the local model uploaded by sub-server m; combining the downlink communication delay, model training time and uplink communication delay model in S3 to update the group training task completion state sequence π m , including the following:
[0149] Assume that the length τ of a time slot k is much larger than the downlink communication delay, use Indicates the time required for user group m to complete model training in training time slot k. Updated by the following formula:
[0150]
[0151] Among them, the definition Indicates that user group m completes model training in training time slot k, and sub-server m can upload the obtained group model to the multi-function drone, or sub-server m is uploading the group model (if π m [k-1]=0 and );use Indicates the remaining data amount that sub-server m needs to upload in time slot k. In training time slot k, Updated by the following formula:
[0152]
[0153] Among them, the definition represents the uplink communication rate of the sub-server m in the kth training time slot. Then the sub-server m completes the training task in the training time slot k. Therefore, the task state π of user n in the kth training time slot is n [k] is updated as follows:
[0154]
[0155] Finally, the global model is obtained by aggregation update:
[0156]
[0157] Among them, g(γ m [k]) is a dynamic adaptive discount factor used to reduce the impact of stale group models on the global model. The function g(·) is defined as g(x) = (x + 1) -a , a is a positive constant; parameter β>0 is used to control the model aggregation to the global model w k Smoothing coefficient of the influence degree;
[0158] Specifically, S5 constructs an optimization problem to minimize the staleness of the swarm model under the constraints of the number of cargo delivery, drone flight, and training time slots to optimize the flight path of multi-functional drones. The process includes the following:
[0159] When the multifunctional drone is in different locations, the uplink and downlink communication delays between it and the sub-server are different, that is, the time required for the group to complete a training task is inconsistent. Therefore, the flight path of the multifunctional drone will not only affect the training task status of the group, but also affect the staleness of the group model. This step constructs a multifunctional drone path optimization problem that minimizes the average staleness of the group model. By optimizing the flight path of the multifunctional drone, the average staleness of all group models is minimized while completing the cargo delivery task. The mean The specific optimization problem is as follows:
[0160]
[0161]
[0162]
[0163] Among them, constraints C1, C2, C3, C4, and C5 are the flight constraints of the multifunctional UAV given in S2. is the cargo delivery constraint given in S2, constraint Indicates that the multifunctional UAV is in time slot kstart The remaining flight path length must be greater than the number of training time slots K train ;
[0164] Specifically, S6 specifically includes modifying the optimization variables of the optimization problem established in S5. Since the objective function of the optimization problem established in step 5 is an implicit function of the optimization variables and cannot be solved directly, the average staleness of the model in S4 and the number of group training tasks are used to modify the optimization variables. The relationship between the optimization variables (average staleness of all group models) will be minimized. ) is modified to maximize the number of group training task completions while keeping the constraints of the original optimization problem unchanged. The modified optimization problem is as follows:
[0165]
[0166]
[0167]
[0168] Specifically, S6 specifically includes proposing a UAV path dynamic optimization algorithm based on deep reinforcement learning (DRL) to solve the modified optimization problem, where the optimization variables are mapped to the action space, the constraints are mapped to the state space, and the reward function is constructed according to the optimization objectives and constraints. Specifically, it includes the following steps:
[0169] A1: To implement the DRL method, it is necessary to simplify and discretize the flight actions of the multi-functional UAV and divide the flight area of the multi-functional UAV into Modeled as a A square grid map, where the size of each grid is D max , flight area in The multi-function UAV can only fly in four directions: "East", "South", "West" and "North", and the flight distance each time is D max , that is, each time slot can only fly from one grid to the four adjacent grids, so the flight area The bounds of are given by the set defined below:
[0170]
[0171] If the next position of the multi-function drone is Then the multi-function drone will hover in its original position and receive punishment from the environment;
[0172] A2: Establish the DRL model action space, which includes the following:
[0173] In the DRL model, the action space only includes the flight direction of the multifunctional UAV, so the action space is represented by the set It is defined as follows:
[0174]
[0175] A3: Establish a state space consisting of environmental information and multifunctional drone state, where the environmental information includes the starting point coordinate q start , end point coordinate q end and multi-function drones to the border The state of the multifunctional UAV includes the position q[k] of the current time slot, the number of time slots flown k, the position two time slots ago, and the cargo delivery status Therefore, the specific state space can be set It is defined as follows:
[0176]
[0177] in,
[0178] A3: Construct the reward function of the DRL model to ensure that the multi-functional drone completes the delivery of all goods and the path length is within the specified range. within, among them Defined as The reward function should include the round reward R episode , Safety Reward R safe , Mobile Reward R move , Group training task reward R π and cargo delivery reward R o , so the specific reward function R(k) is obtained as follows:
[0179] R(k)=α1R episode [k]+α2R safe [k]+α3R move [k]+α4R π [k]+α5R o [k]
[0180] Among them, α1, α2, α3, α4 and α5 are all positive constants; the round reward R episode [k] is related to the cargo delivery constraints and drone flight constraints in the optimization problem P4. In the last time slot of the round, the round reward can be calculated as follows:
[0181]
[0182] Among them, r success Is a large positive number. Other time slots R episode[k] = 0; safety reward R safe [k] is related to the UAV flight constraints in the optimization problem P4. When the next position of the multifunctional UAV is on the boundary, that is, Then the multi-function drone will hover in its original position and be punished safe [k]<0; otherwise, p safe [k]=0, the safety reward can be calculated as follows:
[0183]
[0184] Mobile Rewards R move [k] is related to the UAV flight constraints in the optimization problem P4, which is given by a negative heuristic and a non-positive constant Composition, that is The former is when the path length exceeds Q max The latter is the penalty for returning to the position two time slots ago, that is, where p move <0. The mobile bonus can be calculated as follows:
[0185]
[0186] where p time Is a very small negative number; Group training task reward where r n [k] is the reward for each base station’s task completion status change, which corresponds to the objective function of the optimization problem P4. During the training period, the group training task reward can be calculated as follows:
[0187]
[0188] Note that when When r m [k] = 0; Goods delivery reward R o [k] is related to the cargo delivery constraints in the optimization problem P4, which consists of three parts, namely Among them, the first It can be obtained by the following formula:
[0189]
[0190] where p o is a negative constant. If the multifunctional UAV completes the cargo delivery before the Kth time slot, that is, but otherwise, Item 2 It is used to reward the multifunctional UAV for completing the delivery of goods to a receiving point. If the multifunctional UAV completes the delivery task in the kth training time slot, is a positive constant r o , otherwise 0; the last item is a negative heuristic that encourages multi-purpose drones to fly over delivery points:
[0191]
[0192] Where ω is a positive constant;
[0193] A4: Use the DQN algorithm to solve the flight path of the multifunctional drone. In each round of each iteration, there are three parts: the interaction between the multifunctional drone and the environment, experience playback, and the DQN algorithm with the target Q network. The multifunctional drone inputs the current state s[k] into the Q network to obtain the action a with the largest Q value. max [k], in order to explore the map, the multi-function drone's action has a certain probability (exploration degree) of being a random action a random [k]; Select a in the multi-function drone max [k] or a random [k] is used as the action a[k], and the next state s[k+1] and reward r[k] are obtained by interacting with the environment according to s[k] and a[k]. The experience (s[k], a[k], r[k], s[k+1]) obtained from the current interaction is stored in the experience pool. After the number of samples stored in the experience pool reaches the set value, the current parameters of the Q network and the target Q network are defined as w now and Start updating the Q network and Q target network, including the following steps:
[0194] a.1: Perform forward propagation on the Q network and obtain:
[0195] a.2: Perform forward propagation on the target Q network and obtain
[0196] a.3: Calculate TD target Where γ is the discount rate;
[0197] a.4: Calculate TD error
[0198] a.5: Backpropagate the Q network to get the gradient
[0199] a.6: Update the Q network parameters by gradient descent:
[0200]
[0201] a.7: After a certain number of iterations, use the Q network parameters to update the target Q network parameters;
[0202] The update process of the above Q network and target Q network will be iterated multiple times until the end of the round or the maximum number of iterations is reached.
[0203] Furthermore, the present invention uses a multifunctional drone-assisted asynchronous cluster personalized federated learning method to perform simulation analysis:
[0204] Figure 3 The figure shows the flight path grid diagram of the multifunctional UAV under two comparison algorithms and the algorithm proposed in the present invention. Among them, the comparison algorithm 1 is the random path algorithm, and the comparison algorithm 2 is the shortest path algorithm. Since the coordinates of the starting point, the end point and the delivery point are symmetrical about the diagonal line of the flight area, the shortest path algorithm corresponds to two different paths, indicating the different order of passing through the delivery points. It should be noted that the random path algorithm and the shortest path algorithm are selected as the comparison algorithms of the algorithm proposed in the present invention because the latest existing UAV path optimization algorithm is not suitable for the system constructed by the present invention. Except for the reward function, the other simulation parameters of the several algorithms are the same as those of the proposed algorithm. Figure 3 The following points can be seen:
[0205] (1) The flight paths found by the proposed algorithm are more concentrated near the sub-server than those found by the random path algorithm. The reason is that the reward function of the proposed algorithm includes the group training task reward, and this reward is related to the communication distance between the multifunctional UAV and the sub-server. The shorter the communication distance between the multifunctional UAV and the sub-server, the lower the staleness of the group model uploaded by the sub-server, and the greater the group training task reward.
[0206] (2) Multifunctional UAVs in O min = kth time after delivery to 2 receiving points start = 20 flight slots to start cluster asynchronous personalized federated learning training. During this period, the distances between the two paths and the sub-servers are very different. “Path 1” is closer to the sub-server than “Path 2”.
[0207] Figure 4 This is a scatter plot of the average model staleness corresponding to the four paths generated by the proposed algorithm and the two comparison algorithms. If the algorithm finds a path that completes the cargo delivery task and reaches the destination in a certain round, there is a horizontal axis representing the round; Figure 4 It can be seen that the average staleness of the paths found by the algorithm proposed in the present invention after convergence is concentrated between 0.15 and 0.2, which is significantly lower than the average staleness of the other two comparison algorithms.
[0208] Figure 5The figure shows the change of the average uplink communication rate of users under the flight paths generated by the algorithm proposed in this invention and the two comparison algorithms as the time slot increases. It can be seen that in the training time slot, the path corresponding to the algorithm of this invention significantly improves the average uplink communication rate of users. Since the average uplink communication rate is closely related to the communication distance from the drone to the sub-server, combined with Figure 3 As a result, the flight path of the algorithm of the present invention is more concentrated near the sub-server, so the communication distance between the drone and the sub-server is shorter, which makes the average uplink communication rate of the user higher.
[0209] Figure 6 The reward curves corresponding to the four paths generated by the algorithm proposed in this invention and the two comparison algorithms are shown in Figure 2. Figure 5 It can be seen that the convergence reward value of the algorithm proposed in the present invention is significantly higher than that of the two comparison algorithms. The reason is also that the reward function in the path optimization algorithm proposed in the present invention includes the group training task reward.
[0210] Figure 7 The performance of a multifunctional drone-assisted asynchronous cluster personalized federated learning method in a group-level Non-IID data setting is demonstrated; Figure 6 The first comparison solution is the PFL solution for solving the Non-IID data at the individual level, and the second comparison solution is the non-personalized solution. Figure 6 It can be seen that both the proposed solution and the individual-level PFL solution outperform the non-personalized solution, indicating that the personalized algorithm has a significant advantage under the non-IID data setting. The model performance of the proposed solution under this data setting is significantly better than the two comparison solutions. However, when the probability of user participation in training decreases, the performance of the proposed solution is close to that of the individual-level PFL solution because the amount of data available for training in the group is very close to that of the user.
[0211] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A multifunctional drone-assisted asynchronous cluster personalized federated learning method, characterized by: The following steps are involved: S1: Establish a multifunctional UAV-assisted FL network system model, which includes a multifunctional UAV equipped with communication, computing modules and a cargo hold as the edge server of the network system, a ground base station and the server forming a sub-server, and user devices associated with the same sub-server forming a group; S2: Establish a multifunctional UAV model that includes logistics transportation and assisting ground user groups for FL training. At the same time, establish a LoS wireless communication channel model between the UAV and the sub-server and a simplified wireless communication model between the sub-server and the associated user group; S3: Combined with S2, a model of downlink communication delay, model training time, and uplink communication delay between the multifunctional UAV and the sub-server is established; S4: Establish a federated learning mechanism to asynchronously train personalized models for user groups based on data heterogeneity at the user group level; Based on a two-layer parallel optimization problem where the inner layer solves the personalized model and the outer layer solves the global model, the personalized model is trained for each group using asynchronous two-layer parallel optimization between groups and synchronous federated averaging within the group. S5: Under the constraints of cargo delivery, drone flight, and training time slots, the staleness of the swarm model is used as the optimization variable to construct an optimization problem and optimize the flight path of the multifunctional drone; S6: Based on the relationship between model staleness and the number of times the user device completes the training task, the optimization variables in S5 are modified while keeping the constraints unchanged. The transformed optimization problem is an integer nonlinear optimization problem. A dynamic optimization algorithm for drone paths based on deep reinforcement learning (DRL) is proposed.
2. The method for asynchronous cluster personalized federated learning assisted by a multifunctional drone according to claim 1, characterized in that: In S1, a multifunctional UAV-assisted FL network system model is established, including the following processes: The FL network system model includes a UAV equipped with communication, computing modules and cargo compartment, M ground base stations and sub-servers consisting of servers, and N users. The set of sub-servers is represented as The set of users is represented as The ML task is performed in a FL network assisted by a multifunctional UAV. The multifunctional UAV acts as a mobile base station and edge server, connected to the sub-server on the ground through a wireless channel. The multifunctional UAV maintains a safe height H sf , fly from the starting point to the destination, and deliver the goods to O receiving points, The coordinates are in Represents the collection of these receiving points; each sub-server are all located on the horizontal plane, and their coordinates are Each user Equipped with an antenna located on a horizontal plane; each sub-server Serving a group of users through wireless channels, this group of users constitutes a user group in Use a binary random variable X m,n ∈{0,1} represents the users in group m Status, X m,n =1 means user device n can participate in model training, X m,n =0 means that user device n cannot participate in model training; define the user in group m The probability of participating in model training is Group User devices in The dataset used is represented as The total data set of group m is then expressed as And the data distribution between different groups m is non-IID at the group level.
3. The multifunctional drone-assisted asynchronous cluster personalized federated learning method according to claim 2, characterized in that: In S2, a multifunctional UAV model is established that includes logistics transportation and assists ground user groups in FL training, including the following processes: The flight time of the multifunctional UAV is divided into K time slots, using represents the index of the flight time slot, the length of each flight time slot is τ; the maximum flight distance of each time slot of the multifunctional UAV is D max , the horizontal projection of the flight area is Since the cargo delivery task has success or failure, a binary random variable is used Indicates the delivery point The delivery status of goods, including Indicates that the goods have been delivered successfully. Indicates that the goods have not been delivered. Assuming that the drone descends vertically from the delivery point o during delivery, and the time spent in the middle is not included in the flight time slot, the multifunctional drone needs to complete the delivery of all goods before the end of the flight. The delivery status of the goods in the Kth flight time slot should meet the following constraints: The flight path length of the multi-function UAV is limited by its own battery capacity. Assume that the maximum flight path length of the UAV is Q max time slot; the drone is in The horizontal projection coordinates of the time slot are Indicates that the trajectory is expressed as express.
4. The method for asynchronous cluster personalized federated learning assisted by a multifunctional drone according to claim 3, characterized in that: In S2, a LoS wireless communication channel model is established between the multifunctional UAV and the sub-server, which includes the following processes: In each time slot k, the sub-server accesses the multifunctional UAV through orthogonal frequency division multiplexing, with a total bandwidth of B, and each sub-server gets an average bandwidth of B m =B / N; Assuming that the wireless communication channel between the multifunctional UAV and the sub-server is a LoS channel and the channel gain at 1 meter is β0, the channel gain between the multifunctional UAV and the sub-server m is as follows:
5. The method of asynchronous cluster personalized federated learning assisted by a multifunctional drone according to claim 4, characterized in that: In S2, the wireless communication model between the sub-server and the associated user group includes the following processes: Assume that in each group iteration, the sum of the downlink communication delay between the sub-server and the user, the user's local update time, and the uplink communication delay is less than τ e .
6. The method of asynchronous cluster personalized federated learning assisted by a multifunctional drone according to claim 5, characterized in that: In S3, a downlink communication delay, model training time, and uplink communication delay model between the multifunctional UAV and the sub-server are established, which includes the following processes: Use p u [k] represents the transmission power of the multi-functional UAV, p m [k] represents the transmission power of the sub-server, N0 represents the noise power, and combined with S2, the downlink communication delay, model training time, and uplink communication delay between the multi-functional UAV and the sub-server are modeled. The model includes the following sub-steps: S31: Establish a downlink communication delay model, the content of which is as follows: At the beginning of each training time slot k, the multi-function UAV broadcasts the latest global model to all sub-servers; assuming the broadcast channel bandwidth is B broadcast , then in the training time slot k, the downlink communication data rate between the multifunctional UAV and the sub-server m is: Assuming that the number of parameters of the ML model is s, where each parameter is stored as a b-bit floating-point number, the downlink communication delay is: S32: Establish a model training time model, the content is as follows: Assume that the time required for each group iteration is τ e , then in the training time slot k, the model training time required for each global iteration of the idle user group and the associated sub-server m is: Where E is the number of group iterations; S33: Establish an uplink communication delay model, the content of which is as follows: In training time slot k, the uplink communication rate of sub-server m is: The uplink communication delay is:
7. The multifunctional drone-assisted asynchronous cluster personalized federated learning method according to claim 6, characterized in that: In S4, a federated learning mechanism is established to asynchronously train personalized models for user groups based on data heterogeneity at the user group level, including the following steps: S41: Establish a two-level parallel optimization problem for asynchronous personalized federated learning assisted by multifunctional drones. The content is as follows: Establish the following optimization problem P1, aiming to find its optimal solution w * : P1: Among them, u m represents the personalized model of group m, w represents the global model, and F m (w) is defined as the following optimization problem P2: P2: Among them, λ is a regularization parameter used to control the personalized model u m The difference from the global model w; f m (u m ) represents the training loss of all users in group m The sum of f n (u m ) represents the personalized model u m On a local privacy dataset The loss function on , and F m The definition of (w) is related to the Moreau envelope; In the two-level parallel optimization problem represented by optimization problems P1 and P2, P1 is the outer optimization problem for solving the global model w, and P2 is the outer optimization problem for solving the personalized model u. m The inner optimization problem of ,these two optimization problems are decoupled; S42: The group model and the personalized model are synchronously updated to obtain the personalized model and the group model, the contents of which are as follows: Use e∈E to represent the index of E group iterations, where is the index set of E group iterations; at the beginning of each group iteration, sub-server m broadcasts the latest group model to all members in group m who can participate in training The user n participating in the training is calculated by performing gradient descent on the optimization problem P2 When all users participating in the training have completed the training, the sub-server m aggregates the user models of all users participating in the training To generate a personalized model; after the e∈Eth group iteration, the group model is updated by performing gradient descent on the optimization problem P1; finally, after E group iterations, the obtained group model is uploaded to the multi-functional UAV, completing the group training task after one training time slot update; S43: Global model asynchronous update, the content is as follows: Before performing asynchronous update of the global model, the staleness of the group model uploaded by the sub-server is first analyzed. In the example, we use a binary random variable π m [k]∈{0,1} represents a group Whether the training task is completed in this time slot, the training task status of the group is updated on the multifunctional drone, where π m [k] = 1 means that group m has completed this task and is ready for subsequent training tasks, and π m [k]=0 means other situations; sequence Contains group m in all training time slots The state of; the set of user groups that can participate in model training in training time slot k Including all π m [k-1] = 1 group; use γ m [k] represents the group model uploaded by sub-server m in training time slot k If the group model does not exist, then its corresponding γ m [k] does not exist, group model Obsolescence m [k] is updated as follows: Among them, k m is the index of the global model used by the collaborative training model of sub-server m and the associated user group m, N / A indicates staleness γ m [k] does not exist, it is not difficult to find that if π m [k]=1, then γ m [k] represents the sequence π m The time slot k is the same as the previous π m = the number of zeros between time slots of 1; if we define 0×N / A=0, and use the integer ν k ≥0 represents the sequence π m The number of trailing zeros gives the following formula: Combined with the above formula, the average staleness of the local model uploaded by group m is Satisfies the following inequality: The larger the value of The smaller it is, that is, the more times group m completes the training task, the smaller the average staleness of the local model uploaded by group m; combining the downlink communication delay, model training time and uplink communication delay model in S3 to update the group training task completion state sequence π m ; Finally, the global model is obtained by aggregation update through the following formula: Among them, g(γ m [k]) is a dynamic adaptive discount factor used to reduce the impact of stale group models on the global model. The function g(·) is defined as g(x) = (x + 1) -a , a is a positive constant; parameter β>0 is used to control the model aggregation to the global model w k The smoothing coefficient of the influence degree.
8. The multifunctional drone-assisted asynchronous cluster personalized federated learning method according to claim 7, characterized in that: In S5, under the constraints of the number of cargo delivery, drone flight, and training time slots, an optimization problem of minimizing the staleness of the group model is constructed to optimize the flight path of the multifunctional drone, including the following process: When the multi-function UAV is in different locations, the uplink and downlink communication delays between it and the sub-server are different, that is, the time required for the group to complete a training task is inconsistent; Construct a multifunctional UAV path optimization problem that minimizes the average staleness of the group model. By optimizing the flight path of the multifunctional UAV, the average staleness of all group models is minimized while completing the cargo delivery task. The mean 9. The multifunctional drone-assisted asynchronous cluster personalized federated learning method according to claim 8, characterized in that: In S6, according to the average staleness of the model in S4 and the number of group training tasks The relationship between the optimization variables will be minimized, that is, the average staleness of all group models The solution is modified to maximize the number of completed group training tasks while keeping the constraints of the original optimization problem unchanged.
10. The multifunctional drone-assisted asynchronous cluster personalized federated learning method according to claim 9, characterized in that: In S6, a UAV path dynamic optimization algorithm based on deep reinforcement learning (DRL) is proposed to solve the modified optimization problem, which includes the following steps: A1: To implement the DRL method, it is necessary to simplify and discretize the flight actions of the multi-functional UAV and divide the flight area of the multi-functional UAV into Modeled as a A square grid map, where the size of each grid is D max , flight area in The multi-function UAV can only fly in four directions: "east", "south", "west" and "north", and the flight distance each time is D max , that is, each time slot can only fly from one grid to four adjacent grids; A2: Establish the DRL model action space; A3: Establish a state space consisting of environmental information and multifunctional drone state, where the environmental information includes the starting point coordinate q start , end point coordinate q end and multi-function drones to the border The state of the multifunctional UAV includes the position q[k] of the current time slot, the number of time slots flown k, the position two time slots ago, and the cargo delivery status. A3: Construct the reward function of the DRL model to ensure that the multi-functional drone completes the delivery of all goods and the path length is within the specified range. The reward function should include the round reward R episode , Safety Reward R safe , Mobile Reward R move , Group training task reward R π and cargo delivery incentives; A4: The DQN algorithm is used to solve the flight path of the multi-functional UAV. In each round of each iteration, there are three parts: the interaction between the multi-functional UAV and the environment, experience replay, and the DQN algorithm with the target Q network. The third part of the process will be iterated multiple times until the end of the round or the maximum number of iterations is reached.
Citation Information
Patent Citations
Unmanned aerial vehicle auxiliary relay method for emergency scene
CN117792481A
Multi-rotor unmanned aerial vehicle cluster training neural network model by applying federated learning framework
CN117873171A