Terminal selection method, model training method, device and system
Patent Information
- Application Number
- CN202280100467.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-05-06
AI Technical Summary
In federated learning, the existing technology lacks an effective terminal selection solution, which makes it difficult to balance the energy consumption, learning accuracy and convergence speed of edge terminals. Especially in non-independent and identically distributed data scenarios, the local updates of terminals deviate from each other, resulting in The global model quality is not high.
By comprehensively considering factors such as terminal energy consumption and data quality, the deep deterministic policy gradient DDPG algorithm is used to dynamically select the most suitable terminal from multiple terminals to participate in model training, establish an objective function and solve it under constraints to optimize the terminal Selection and resource allocation.
It achieves rapid model convergence while minimizing energy consumption overhead, improves the efficiency and accuracy of federated learning, and balances the energy consumption and learning accuracy of edge terminals.
Smart Images

Figure CN119948501A_ABST
Abstract
Description
Terminal selection method, model training method, device and system Technical Field
[0001] The present disclosure relates to the field of mobile communication technology, and in particular to a terminal selection method, a model training method, a device, and a system. Background Art
[0002] With the continuous evolution of mobile network communication technology, various application scenarios are increasingly demanding higher network communication efficiency. Artificial Intelligence (AI) technology has achieved continuous breakthroughs in the communications field, bringing richer application experiences to users. In AI technology, terminals can use their local data to train and update the Federated Learning (FL) models required by the server. However, there is currently no effective solution for the problem of biased selection of appropriate terminals in each global iteration.
[0003] Summary of the Invention
[0004] The present disclosure proposes a terminal selection method, a model training method, an apparatus, and a system, which provide a method that comprehensively considers factors such as terminal energy consumption and data quality, thereby biasedly selecting the most suitable terminal in each round of global iteration to balance the energy consumption of edge terminals and the learning accuracy and convergence speed of FL.
[0005] A first aspect embodiment of the present disclosure provides a terminal selection method, which is executed by a network device, and includes: receiving model training parameters sent by at least one terminal, wherein the at least one terminal is selected by the network device from multiple terminals for performing first model training on a global model issued by the network device, and the model training parameters are obtained after the at least one terminal performs the first model training on the global model; based on the model training parameters, determining data quality parameters and energy consumption parameters of the terminals participating in the second model training; and selecting the terminals participating in the second model training from multiple terminals according to the data quality parameters and energy consumption parameters.
[0006] In some embodiments of the present disclosure, the model training parameters include the model weight of the first model training of at least one terminal, and determining the data quality parameters of the terminals participating in the second model training includes: determining the aggregated model weight of the first model training based on the model weight and the local data size of at least one terminal; determining the weight divergence based on the difference between the model weight and the aggregated model weight, and using the weight divergence as the data quality parameter.
[0007] In some embodiments of the present disclosure, the model training parameters include the local data distribution value of the first model training of at least one terminal, and determining the data quality parameters of the terminal participating in the second model training also includes: weighted merging the weight divergence and the local data distribution value, determining the quality score value, and using the quality score value as the data quality parameter.
[0008] In some embodiments of the present disclosure, the model training parameters include the local training energy consumption of the first model training of at least one terminal, and determining the energy consumption parameters of the terminal participating in the second model training includes: determining the energy consumption parameters based on the local training energy consumption and the transmission energy consumption of the model training parameters uploaded by at least one terminal.
[0009] In some embodiments of the present disclosure, selecting a terminal to participate in the second model training from multiple terminals based on data quality parameters and energy consumption parameters includes: establishing an objective function based on the data quality parameters and energy consumption parameters; establishing constraints, the constraints including at least one of time constraints, energy consumption constraints, bandwidth constraints, terminal processor cycle frequency constraints, transmission power constraints, and terminal selection decision constraints; solving the objective function under the constraints according to the deep deterministic policy gradient DDPG algorithm to select the terminal to participate in the second model training from multiple terminals.
[0010] In some embodiments of the present disclosure, according to the DDPG algorithm, solving the objective function to select a terminal to participate in the second model training from multiple terminals includes: representing the objective function with a Markov quadruple, wherein the Markov quadruple includes a state space, an action space, a reward function, and a transition probability; establishing a policy gradient and a loss function under a policy-evaluation network, and determining the optimal state-action pair to select a terminal to participate in the second model training from multiple terminals.
[0011] In some embodiments of the present disclosure, the state space includes data quality parameters, uplink channel gain, channel bandwidth from the terminal to the network device, and the remaining energy of the terminal; the action space includes terminal selection parameters and the transmission power of the selected terminal; the reward function is the ratio of the sum of the data quality parameters of the terminals participating in the second model training to the energy consumption; and the transition probability is the state transition probability.
[0012] The second aspect of the present disclosure provides a model training method, which is executed by a network device, and includes: creating an initial global model using public data; receiving resource information sent by multiple terminals; selecting an initial training terminal from the multiple terminals based on the resource information; sending the initial global model to the initial training terminal, and receiving the first model training parameters uploaded by the initial training terminal after training the initial global model; performing model aggregation to generate a global model for first model training; determining data quality parameters and energy consumption parameters of the first terminal participating in the first model training according to the first model training parameters, and selecting the first terminal from the first terminal according to the data quality parameters and the energy consumption parameters. Determine the first terminal for the first model training from the multiple terminals, and send the global model used for the first model training to the first terminal for training; receive the second model training parameters uploaded by the first terminal after the global model training; perform model aggregation to generate a global model for the second model training; determine the data quality parameters and energy consumption parameters of the second terminal participating in the second model training based on the second model training parameters, and determine the second terminal for the second model training from the multiple terminals based on the data quality parameters and the energy consumption parameters, and send the global model used for the second model training to the second terminal for training until the model accuracy converges.
[0013] An embodiment of the third aspect of the present disclosure provides a terminal selection device, which includes: a receiving module for receiving model training parameters sent by at least one terminal, wherein the at least one terminal is a terminal selected by a network device from multiple terminals for performing a first model training on a global model issued by the network device, and the model training parameters are obtained after the at least one terminal performs the first model training on the global model; a determination module for determining data quality parameters and energy consumption parameters of terminals participating in the second model training based on the model training parameters; and a selection module for selecting terminals participating in the second model training from multiple terminals based on the data quality parameters and energy consumption parameters.
[0014] The fourth aspect of the present disclosure provides a model training device, which includes: a creation module for creating an initial global model using public data; a transceiver module for receiving resource information sent by multiple terminals; a selection module for selecting an initial training terminal from the multiple terminals based on the resource information; the transceiver module is also used to: send the initial global model to the initial training terminal, and receive the first model training parameters uploaded by the initial training terminal after training the initial global model; an aggregation module is used to perform model aggregation to generate a global model for first model training; the selection module is also used to: determine the data quality parameters and energy consumption parameters of the first terminal participating in the first model training according to the first model training parameters, and determine the data quality parameters and energy consumption parameters of the first terminal participating in the first model training according to the data quality parameters and the energy consumption parameters , determine the first terminal for performing the first model training from the multiple terminals; the transceiver module is also used to: send the global model used for the first model training to the first terminal for training, and receive the second model training parameters uploaded by the first terminal after training the global model; the aggregation module is also used to: perform model aggregation to generate a global model for the second model training; the selection module is also used to: determine the data quality parameters and energy consumption parameters of the second terminal participating in the second model training according to the second model training parameters, determine the second terminal for performing the second model training from the multiple terminals according to the data quality parameters and the energy consumption parameters, and send the global model used for the second model training to the second terminal for training until the model accuracy converges.
[0015] The fifth aspect embodiment of the present disclosure provides a communication device, which includes: a transceiver; a memory; and a processor, which is connected to the transceiver and the memory respectively, and is configured to control the wireless signal reception and transmission of the transceiver by executing computer-executable instructions on the memory, and can implement any method of the above-mentioned first aspect embodiment or second aspect embodiment.
[0016] The sixth embodiment of the present disclosure provides a computer storage medium, wherein the computer storage medium stores computer-executable instructions; after the computer-executable instructions are executed by a processor, the method described in the first or second embodiment of the present disclosure can be implemented.
[0017] According to the terminal selection method disclosed herein, a network device receives model training parameters sent by at least one terminal, wherein the at least one terminal is selected by the network device from multiple terminals for performing first model training on a global model sent by the network device, and the model training parameters are obtained after at least one terminal performs first model training on the global model; based on the model training parameters, the data quality parameters and energy consumption parameters of the terminals participating in the second model training are determined; based on the data quality parameters and energy consumption parameters, the terminals participating in the second model training are selected from multiple terminals. The solution disclosed herein comprehensively considers factors such as terminal energy consumption and data quality, thereby biasedly selecting the most suitable terminal in each round of global iteration to balance the energy consumption of the edge terminal and the learning accuracy and convergence speed of the FL.
[0018] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0020] FIG1 is a schematic diagram of a flow chart of a terminal selection method according to an embodiment of the present disclosure;
[0021] FIG2 is a schematic diagram of a flow chart of a transmission configuration method according to an embodiment of the present disclosure;
[0022] FIG3 is a flow chart of a model training method according to an embodiment of the present disclosure;
[0023] FIG4 is a schematic block diagram of a terminal selection device according to an embodiment of the present disclosure;
[0024] FIG5 is a schematic block diagram of a model training device according to an embodiment of the present disclosure;
[0025] FIG6 is a schematic structural diagram of a communication device according to an embodiment of the present disclosure;
[0026] FIG7 is a schematic structural diagram of a chip provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0028] With the continuous evolution of mobile network communication technology, from the fifth generation (5G) to the sixth generation (6G), the interconnectedness of everything is a defining trend. AI will become a core technology in future communications. Typical application scenarios for 6G and AI overlap by over 80%, demonstrating a deep integration of the two. Furthermore, the massive coverage of 6G networks will provide a ubiquitous platform for AI, addressing the significant pain point of a lack of carriers and channels for AI implementation and significantly promoting the development and prosperity of the AI industry.
[0029] The Internet of Everything (IoE) enables all terminal devices to function as intelligent entities for analysis and computing, forming an intelligent network. Mobile phones, laptops, sensors, and other terminal devices generate large amounts of heterogeneous data locally. If this data can be effectively utilized by the network, it will contribute significantly to intelligent analysis and resource optimization. In traditional cloud-server-centric computing, user data is typically sent to a central server for computation and storage. This approach not only raises numerous security and privacy issues, but also incurs significant energy consumption and latency overhead during transmission, contradicting the principles of 6G green communications. It also consumes significant bandwidth and hinders the normal operation of other task queues in the network. With the prevalence of mobile edge computing and federated learning technologies, and the increased computing power of terminal devices, more computing tasks are being offloaded to edge servers or local users. This intelligent network architecture offers greater possibilities for orchestration and deployment.
[0030] In federated learning, devices use their local data to train and update the machine learning (ML) model required by the server. Instead of sending the original data, the device sends the updated model parameters to the server for aggregation. The server then sends the aggregated model to the device, repeating this process until the model converges.
[0031] However, FL inherently suffers from low classification accuracy in non-IID scenarios due to biased data distribution across devices. While various approaches have been proposed to address this issue, such as personalized local loss functions and model distillation, they lack a comprehensive consideration of device data distribution, local model contribution, and communication overhead, which are crucial for determining the quality of global model aggregation. Differing training data distributions cause local updates across devices to deviate from each other, resulting in a biased aggregated global model. Furthermore, lagging devices with limited computational power or poor channel conditions significantly slow model convergence. Due to scarce bandwidth resources and limited energy budgets for devices, it is impractical to have all devices participate in every round of training. Instead, only a subset of clients are required to participate in each training round. Compared to randomly selecting participating clients for each round, it is crucial to comprehensively consider factors such as client energy consumption, power allocation, and client data quality to selectively select the most suitable clients for each global iteration, balancing edge device energy consumption with FL learning accuracy and convergence speed.
[0032] Therefore, the present disclosure aims to solve the problem that there is no corresponding terminal (or client) selection solution in the relevant technology, that is, how to dynamically select the participating terminals in each round according to the model update situation during the training process to balance the data distribution and reduce the impact of terminal non-iid data on the model aggregation effect.
[0033] The purpose of this invention is to comprehensively consider the energy consumption and latency generated by each terminal during local computation and model transmission during FL training, as well as the local data quality of the terminal, to analyze and select the set of terminals participating in each round of training and the corresponding power allocation scheme, thereby achieving flexible and efficient federated learning training, while minimizing energy consumption while achieving rapid model convergence.
[0034] To this end, the present disclosure proposes a terminal selection method, a model training method, an apparatus, and a system, which provide a method that comprehensively considers factors such as terminal energy consumption and data quality, thereby biasedly selecting the most suitable terminal in each round of global iteration to balance the energy consumption of edge terminals and the learning accuracy and convergence speed of FL.
[0035] It can be understood that the solution provided in the present disclosure can be used for the fifth generation mobile communication technology (Fifth Generation, 5G) and its subsequent communication technologies, such as the fifth generation mobile communication technology evolution (5G-advanced), the sixth generation mobile communication technology (Sixth Generation, 6G), etc., and is not limited in the present disclosure.
[0036] The solution provided in this application is described in detail below with reference to the accompanying drawings.
[0037] Figure 1 shows a schematic flow chart of a terminal selection method according to an embodiment of the present disclosure. As shown in Figure 1 , the method is executed by a network device, which may be a server. As shown in Figure 1 , the method includes the following steps.
[0038] S101: Receive model training parameters sent by at least one terminal.
[0039] It should be understood that the solution described in the present disclosure is for multiple rounds of model training and iteration processes in the FL process. For each round of training, the server can dynamically select the terminals participating in the training in this round based on the terminal conditions selected by the server in the previous round, taking into account the energy consumption and data quality of each terminal. Therefore, in an embodiment of the present disclosure, at least one terminal is a terminal selected by a network device from multiple terminals for performing the first model training on the global model issued by the network device, wherein the first model training is the previous round of model training. In other words, for the current round of training process, the model training parameters received by the server may be those sent by the terminal selected by the server in the previous round of training.
[0040] In the embodiments of the present disclosure, model training parameters are obtained after at least one terminal performs a first model training on the global model. These model training parameters can include parameters such as data quality score, CPU frequency, channel gain, and battery power. As long as the parameters used for model training and iteration fall within the scope of this disclosure, they are not limited in this disclosure.
[0041] S102: Determine data quality parameters and energy consumption parameters of terminals participating in the second model training based on the model training parameters.
[0042] In the embodiments of the present disclosure, data quality parameters can be used to assess the local data quality of a terminal, and energy consumption parameters can be used to assess the energy consumption of information transmission between the terminal and the server, as well as the energy consumption of local model training on the terminal. This embodiment does not limit the types and forms of data quality parameters and energy consumption parameters.
[0043] In the embodiments of the present disclosure, the second model training is the current round of model training, that is, the current round of model training, which is a relative concept to the first model training, where "first" and "second" are only used as names to distinguish and do not represent limitations on the present disclosure.
[0044] S103: Select a terminal to participate in the second model training from multiple terminals based on the data quality parameter and the energy consumption parameter.
[0045] In the present disclosure, the server can select terminals suitable for participating in the second model training based on the determined data quality parameters and energy consumption parameters, and then, during the subsequent model training process, send the global model to the selected terminals for the current round of training. After the current round of training is completed, the next round of training can be carried out until the model converges. In the next round of training, the server can again dynamically select terminals for the next round of training based on the model training parameters fed back by the terminals until the final model converges.
[0046] In summary, according to the terminal selection method provided by the embodiment of the present disclosure, the network device receives model training parameters sent by at least one terminal, wherein at least one terminal is selected by the network device from multiple terminals for performing first model training on the global model sent by the network device, and the model training parameters are obtained after at least one terminal performs the first model training on the global model; based on the model training parameters, the data quality parameters and energy consumption parameters of the terminals participating in the second model training are determined; based on the data quality parameters and energy consumption parameters, the terminals participating in the second model training are selected from multiple terminals. The solution of the present disclosure comprehensively considers factors such as terminal energy consumption and data quality, thereby biasedly selecting the most suitable terminal in each round of global iteration to balance the energy consumption of the edge terminal and the learning accuracy and convergence speed of the FL.
[0047] Figure 2 shows a flow chart of a transmission configuration method according to an embodiment of the present disclosure. Based on the embodiment shown in Figure 1 , as shown in Figure 2 , the method may include the following steps.
[0048] S201: Receive model training parameters sent by at least one terminal.
[0049] It should be understood that for each round of training in the iterative process, the server may directly request some or all terminals to upload their resource information, such as data quality score, CPU frequency, channel gain, battery power, etc., for the server's initial modeling process (or initialization). The server selects terminals to participate in the first round of training based on the resource information of each terminal. For the second and subsequent training processes, the server selects terminals to participate in the current round of training based on the model training parameters obtained after training the global model, which are fed back by the terminals selected in the previous round.
[0050] The other principles of the above step 201 are the same as those of step 101 and will not be repeated here.
[0051] In the embodiment of the present disclosure, the model training parameters include the model weights of the first model training of at least one terminal. The process of determining the data quality parameters in step S102 shown in FIG1 is explained in detail below through steps S202-204.
[0052] First, assuming that the data size of each terminal is the same, the data heterogeneity of the terminal is understood as category imbalance, in which each device does not follow a common data distribution, that is, the data distribution in the device is non-iid. This data heterogeneity will greatly weaken the aggregation quality of the model, resulting in low classification accuracy. In this disclosure, this difference is referred to as the skewness of the data distribution. Since the test accuracy is determined by the training weight, there is a certain correlation between the weight update of each terminal and its data distribution. Therefore, it is necessary to select devices with uniform data distribution to participate in training in each round as much as possible, but at the same time, in order to ensure fairness and ensure that each terminal has the opportunity to participate in training, it is more necessary to pay attention to the label distribution of the overall data and the weight of the model after aggregation and averaging, and make timely adjustments.
[0053] Therefore, the entire FL system needs to provide feedback on the overall training process through aggregated updates from the server, and infer information through the aggregated model weight parameter results, so as to more effectively select terminals participating in future training.
[0054] S202: Determine an aggregated model weight of the first model training according to a model weight of the first model training of at least one terminal and a local data size of at least one terminal.
[0055] It should be understood that each round of training will go through the process of server aggregation model → model distribution → terminal training model → terminal reporting model training parameters → server selection terminal → server aggregation model → model distribution... After each round of model aggregation, the server can determine the direction of model weight update based on the aggregation results, and then calculate the aggregated model weight based on the model weight uploaded by each terminal in the previous round and the local data size, as shown in the following formula:
[0056]
[0057] Where n represents the nth terminal, where n∈[1,N], N is a positive integer, and t represents the tth round of training. is the model weight of the first model training reported by the terminal (that is, it can be understood as the weight value updated after the last round of training of terminal n), D n is the local data size of the terminal, w Global is the aggregate model weight (i.e., it can be understood as the result of averaging the updated weights trained with data from each terminal).
[0058] S203: Determine the weight divergence according to the difference between the model weight and the aggregated model weight.
[0059] In the present disclosure, the server determines the weight divergence (WD) based on the difference between the model weight and the aggregate model weight, and uses the weight divergence as a data quality parameter for terminal selection, as shown in the following formula:
[0060]
[0061] in, Identifies the weight divergence of terminal n in the tth round of training. To a certain extent, it can reflect the direction in which the model should be corrected, so that the server can select the terminal to participate in the current round of training. Therefore, the data quality parameter described in this disclosure can be weight divergence. Specifically, if the current aggregate model weight is significantly different from the local weight of a terminal, that is, If the terminal has never participated in training before, the local model weight is set to 0. If it has participated in training many times before, the current data quality is judged by the weight of the most recent round of update. To represent the weight of the most recent (previous) update.
[0062] Through this interaction, feedback can be achieved between the server and the terminal, and the terminals participating in each round of training can be selected by fully considering the data quality of each terminal, thereby achieving faster model convergence.
[0063] On the basis of the above embodiment, as an optional or preferred embodiment, in the present disclosure, the model training parameters include a local data distribution value of the first model training of at least one terminal.
[0064] S204: Perform weighted combination on the weight divergence and the local data distribution value to determine a quality score value.
[0065] In order to improve the fairness of the probability of all terminals being selected and to make more use of terminals with special data distribution to participate in training, this disclosure adopts a parameter item to reduce the selection probability of terminals with limited contribution to the model. This parameter item is determined according to the weight divergence of the previous round of model and the local data distribution. This disclosure calls this parameter item the quality score value. Indicates that the quality score value Owei data quality parameter is used to measure the data quality of each terminal in the current round.
[0066] It is understandable that, when the number of samples of each terminal is the same, the more evenly the data sample labels of the terminals participating in each round are distributed, the smaller the loss function of the training result. In other words, when the terminals have similar data, their gradients are also similar. Therefore It should be proportional to the local data distribution (EMD value) of the data distribution of terminal n, where EMD can be expressed by the following formula:
[0067]
[0068] Among them, C represents the number of tag types, represents the proportion of the i-th data type of terminal n, p y=i represents the global proportion of the amount of data of type i. In addition, for the sake of customer privacy and security, the server cannot directly obtain the user's local data distribution. Therefore, in this disclosure, the terminal only needs to send the EMD value of the local data distribution to the server in each round to protect its local data while informing the server of its distribution status.
[0069] Taking all the above factors into consideration, after receiving the EMD value of each terminal, the server compares it with the EMD value of the previous round of training. Perform weighted merging to obtain quality score values
[0070]
[0071] Among them, P const is a constant term used to prevent the fractional value from being negative. is a scaling factor used to balance the unit differences between the terms.
[0072] Assuming the local data label distribution of each terminal remains unchanged, after each model aggregation is completed and before the next local training begins, the model weight parameters of each terminal after the previous training round are recorded in the server for update, facilitating the comparison of data quality scores of each terminal in the next round. This process is iterated multiple times throughout the training process.
[0073] In other words, the data quality parameter described in the present disclosure may be a quality score value, The higher the value, the more useful the terminal n is for the model update and convergence of the current round.
[0074] Therefore, the present disclosure selects terminals by weighted merging of quality score values, which can reduce the probability of terminals participating in training multiple times and give terminals with extreme data distribution an opportunity to participate.
[0075] In an embodiment of the present disclosure, the model training parameters include local training energy consumption of at least one terminal for first model training. The process of determining the energy consumption parameters in step S102 shown in FIG1 is explained in detail below through step S205.
[0076] S205: Determine energy consumption parameters according to local training energy consumption and transmission energy consumption of at least one terminal uploading model training parameters.
[0077] In order to maximize the ratio between the quality fraction of the terminal and the energy consumption in the overall FL process, the present disclosure uses E g Represents the total energy loss of all terminals participating in the training in the tth round of training:
[0078]
[0079] Among them, a n,t Represents a binary indicator. When terminal n is selected by the edge server in round t, then a n,t =1; otherwise, a n,t =0.
[0080] Since terminals typically have limited energy budgets and bandwidth, only a subset of them participate in each round of FL training. The energy consumption of each terminal consists of two aspects: the energy consumed by transmitting model parameters from the terminal to the edge server, and the energy consumed by local model calculations on the terminal.
[0081] 1) Model calculation energy consumption
[0082] For the terminal n selected in round t, it incurs energy consumption due to local training and uploading local updates to the edge server through the wireless channel. For each terminal n, let represents its local training energy consumption in the tth round, which depends on its computing architecture, hardware and dataset. n represents the number of CPU cycles required for the nth terminal to execute one data sample, which is known a priori. Assuming that all data samples have the same data size, i.e., the number of bits, let Dn be the amount of data for terminal n, then the number of CPU cycles required for terminal n to run a local epoch is c n D n .U t,n represents the number of local training rounds of terminal n in the tth global iteration. n represents the CPU cycle frequency of terminal n, β n is the effective capacitance coefficient of the computing chip of terminal n. Then the local training energy consumption formula of terminal n in the tth round of global iteration is as follows:
[0083]
[0084] 2) Model parameter transmission energy consumption
[0085] After local training is completed, an Orthogonal Frequency Division Multiple Access (OFDMA) scheme is used for local model upload, with a total bandwidth of B, in round t, where the system bandwidth is divided into n subchannels equal to the number of selected terminals. OFDMA is a multiple access scheme based on OFDM. In OFDMA, different users are assigned different subcarriers, so multiple users can transmit their data simultaneously.
[0086] set up is the bandwidth allocation ratio of terminal n in the tth round of global iteration, and the sub-channel bandwidth allocated to the nth device is make For the bandwidth allocation of the entire terminal, it must meet If terminal n is selected in round t, the present disclosure requires that at least a minimum bandwidth ratio bmin be allocated to terminal n, i.e. This is because actual systems cannot allocate arbitrarily small bandwidth to a single terminal due to the limited resource block size.
[0087] According to Shannon's theorem, the achievable transmission rate of terminal n is defined as:
[0088]
[0089] Where B is the bandwidth, N0 is the power spectral density of Gaussian white noise, is the transmission power of terminal n in round t, G n is the channel gain between terminal n and the central server.
[0090] Model parameter w n and gradient The data size is s n Assuming that the size of Sn is constant, the time for each terminal to transmit the local model update in round t is:
[0091]
[0092] Then the energy consumption of terminal n during the t-th round of model uploading is:
[0093]
[0094] Then the total energy consumption of terminal n is:
[0095]
[0096] The specific process of step S103 in the embodiment shown in FIG. 1 is explained below through steps S206 - S208 .
[0097] S206: Establish an objective function based on the data quality parameters and the energy consumption parameters.
[0098] In this disclosure, the joint task scheduling and resource allocation optimization problem is modeled, and the objective function is represented by P1:
[0099]
[0100] Among them, when it iterates to the Tth round, the model converges.
[0101] S207 , establishing constraint conditions, where the constraint conditions include at least one of a time constraint, an energy consumption constraint, a bandwidth constraint, a terminal processor cycle frequency constraint, a transmission power constraint, and a terminal selection decision constraint.
[0102] The time constraints of this disclosure are:
[0103]
[0104] That is, the time taken by the terminal to complete model training and uploading in each iteration must be less than or equal to the deadline Tmax.
[0105] The energy consumption constraint is:
[0106]
[0107] That is, it is ensured that the energy consumed by the terminal n participating in each round of training is less than the battery capacity.
[0108] The bandwidth constraint is:
[0109]
[0110] That is, it is ensured that the total bandwidth of the selected terminals should be within the bandwidth capacity.
[0111] The terminal processor cycle frequency constraint is:
[0112]
[0113] That is, it satisfies the CPU cycle frequency range of each terminal.
[0114] The transmission power constraint is:
[0115]
[0116] That is, the transmission power range of each terminal in the tth round is satisfied.
[0117] The terminal selection decision constraints are:
[0118]
[0119] That is, the decision of terminal selection is a binary variable. When terminal n is selected by the edge server in round t, then a n,t =1; otherwise, a n,t =0.
[0120] S208 , solving the objective function under constraints according to the deep deterministic policy gradient (DDPG) algorithm to select a terminal to participate in the second model training from multiple terminals.
[0121] Specifically, this step includes: representing the objective function with a Markov quadruple, where the Markov quadruple includes a state space, an action space, a reward function, and a transition probability; establishing a policy gradient and a loss function under a policy-evaluation network, and determining the optimal state-action pair to select a terminal participating in the second model training from multiple terminals.
[0122] The state space includes data quality parameters, uplink channel gain, channel bandwidth from the terminal to the network device, and the remaining energy of the terminal; the action space includes terminal selection parameters and the transmission power of the selected terminal; the reward function is the ratio of the sum of the data quality parameters of the terminals participating in the second model training to the energy consumption; the transition probability is the state transition probability.
[0123] The following is a detailed explanation of this algorithm.
[0124] Since P1 is a collaboration problem between multiple terminal devices, it is a MINLP problem (mixed integer nonlinear programming problem), and it is difficult to obtain its optimal solution within the polynomial time complexity. Therefore, an online optimization method, namely the deep reinforcement learning method, is used to solve it. Since the action space consists of two factors, one is terminal selection and the other is power allocation, it contains both continuous actions and discrete actions. The traditional value-based Q-learning algorithm cannot solve this problem well. Therefore, the present disclosure adopts the DDPG algorithm, which is a deterministic strategy method that can effectively deal with continuous action decisions. Specifically, the present disclosure models the formulated problem as a Markov decision process and defines its state space and action space. Then the reward function is designed to represent the ratio of the data quality score and energy consumption corresponding to each round of specific actions.
[0125] The terminal selection and resource allocation problem considered above can be expressed as an MDP process (Markov decision making). An MDP can be represented by a 4-tuple (S, A, P, R), where S is the state space, A is the action space, and P is the state transition probability, which means that the state S is the action space. t Transition to the next state S t+1 The probability of , R is the reward function.
[0126] State space S: According to problem P1, the state space of MDP contains the data quality score, the channel gain of the uplink, the channel bandwidth from the terminal to the server, and the residual energy of the IoT device. The network state S in round t t The definition is as follows:
[0127] S t ={(PD n,t ,G n,t ,b n,t ,E n,t ),n=1,...,N}. (18)
[0128] Action space A: The action of MDP is the terminal selection α of FL n,t , and the transmission power P of the selected device n,t 。 Where Α∈{(α n,t ,P n,t ),n=1,...,N,α n,t ∈{0,1} and
[0129] Reward function R: Let R(S t+1 |S t ,A t ) indicates that in state S t Next, take action A t ∈Α is the timely reward received. The reward is defined as the ratio of the sum of the data quality scores of the devices selected by FL in the current round to the energy consumption, that is:
[0130]
[0131] Indicates that from the current state S t Through action A t Transition to the next state S t+1 The rewards received.
[0132] Then, the present disclosure uses the evaluation strategy μ to evaluate the selected action, where μ is the mapping from each state to action. The goal of the present disclosure is to maximize the expected total reward, that is, the action value function Q(S t ,A t |θ μ ):
[0133]
[0134] Where γ∈[0,1] is the discount factor for future states. According to the Bellman formula, a state-action pair (S t ,A t ) and subsequent state-action pairs (S t' ,A t') can be expressed as:
[0135]
[0136] Optimal Action It can be expressed as:
[0137]
[0138] in Provides the maximum FL target ratio for the current state.
[0139] Specifically, this paper uses the DDPG algorithm to find a near-optimal solution to problem P1. The DDPG algorithm employs an actor-critic network structure. The server in the FL, acting as the algorithm's decision maker, is trained to optimize actions in a continuous action space, specifically the selection of terminal devices and power allocation, to continuously increase the system's reward.
[0140] DDPG uses the policy gradient method to map the network state to a specific action. μ ) deterministically maps states to a specific continuous action, and the Critic network is used to approximate the actor-value function. The DDPG algorithm uses an experience replay mechanism to use a register B to store the previous action taken by the server, the current state, the current action, the reward value, and the next state information at each training step. Each time a mini-batch of data is sampled from the experience pool to train the actor-critic network. The result of each action consists of two parts, as mentioned above, namely the terminal selection of FL and the transmission power of the selected device. Α∈{(α n,t ,P n,t ),n=1,...,N}, where α n,t ∈{0,1} and
[0141] After obtaining each action result, since the terminal selection action output from the policy network is a continuous value, and the specific terminal selection policy is a binary value 0 or 1, the present disclosure converts the continuous value into a binary value by setting a threshold.
[0142]
[0143] For continuous terminal selection and power allocation, the goal is to find the optimal strategy π θ To maximize the expectation of the reward function The policy μ can be updated by taking the gradient of the expected return with respect to the network parameters θ, and the actor's policy gradient can be calculated using the chain rule.
[0144]
[0145] This is the parameter θ of the actor network from the starting distribution J μ The gradient of the expected return is calculated and averaged over the sampled mini-batch.
[0146] In addition, the optimal action value function can be obtained through the Critic network Q(S t ,A t |θ Q ) to approximate. In order to update Q(S t ,A t |θ Q ), the Critic network adjusts the parameter θ Q To minimize the current target value and Q(S t ,A t |θ Q ), and MSE is used to measure this error.
[0147]
[0148] For the target network, a soft update method is used to update it. The learning rate τ is introduced to update the old target network parameters θ Q ,θ μ and the new corresponding network parameters θ Q' ,θ μ' Taking a weighted average and then assigning it to the target network ensures a smoother network learning process and reduces overestimation to a certain extent.
[0149] θ Q' ←τθ Q +(1-τ)θ Q'
[0150] θ μ' ←τθ μ +(1-τ)θ μ' (26)
[0151] In addition, during the action phase of training the Actor network output, noise ε needs to be added to the action to enable the agent to have exploration capabilities during the training phase.
[0152] It should be understood that in each round of global iteration, the total delay of terminal n mainly includes the local model training time and parameter result upload time of terminal n. Since the downlink bandwidth is much larger than the uplink bandwidth, the time it takes for the server to send the model to the terminal can be ignored. and They represent the model local training time and result upload time of terminal n in round t respectively.
[0153] The calculation formula is:
[0154]
[0155] As mentioned above, Since each terminal performs local computation in parallel and starts transmitting model parameters once the computation is completed, the time consumed in the t-th round of FL follows the laggard effect, that is:
[0156]
[0157] Therefore, the present disclosure uses an algorithm to select a terminal, which can be summarized as the following process.
[0158]
[0159]
[0160] Figure 3 shows a flow chart of a model training method according to an embodiment of the present disclosure. The method can be executed by a network device, specifically, the network device can be a server.
[0161] As shown in FIG3 , the method may include the following steps.
[0162] S301, creating an initial global model using public data.
[0163] In this disclosure, the server can create an initial global model using public data. Public data can be understood as the θ10CIFA2 dataset. For example, a public dataset of 60,000 images can include a training set of 50,000 images and a test set of 10,000 images, which can be distributed across multiple terminals.
[0164] S302: Receive resource information sent by multiple terminals.
[0165] In the present disclosure, the service may initiate a resource request: the server requests each terminal to upload its resource information (data quality score, CPU frequency, channel gain, battery power, etc.) to the server.
[0166] S303: Select an initial training terminal from the multiple terminals based on the resource information.
[0167] That is, the server can perform joint optimization terminal selection: based on the information provided by each terminal, the server selects several clients within the total bandwidth B to participate in model training.
[0168] It should be understood that the initial training terminal is a terminal that performs model training for the first time after initialization.
[0169] S304: Send the initial global model to the initial training terminal, and receive first model training parameters uploaded by the initial training terminal after training the initial global model.
[0170] After the selection is completed, the server sends the global model to the selected terminal. The selected terminal then uses its local data to train the shared model and uploads the new model parameters to the server through the channel.
[0171] S305: Perform model aggregation to generate a global model for first model training.
[0172] The server uses the Fedavg method to aggregate the uploaded model parameters and generate a new model.
[0173] S306: Determine the data quality parameters and energy consumption parameters of the first terminal participating in the first model training based on the first model training parameters; determine the first terminal for performing the first model training from the multiple terminals based on the data quality parameters and the energy consumption parameters; and send the global model used for the first model training to the first terminal for training.
[0174] S307: Receive second model training parameters uploaded by the first terminal after training the global model.
[0175] S308, performing model aggregation to generate a global model for second model training;
[0176] S309: Determine the data quality parameters and energy consumption parameters of the second terminal participating in the second model training based on the second model training parameters; determine the second terminal for the second model training from the multiple terminals based on the data quality parameters and the energy consumption parameters; and send the global model used for the second model training to the second terminal for training until the model accuracy converges.
[0177] The above steps are repeated until the model converges.
[0178] Therefore, based on the above terminal selection algorithm, the overall federated learning training process is as follows:
[0179]
[0180] In summary, the model training method provided in this disclosure specifically addresses the heterogeneous distribution of device data in federated learning systems by designing a strategy for terminal selection and resource allocation within a federated learning model. To reduce planning time complexity, the latest deep reinforcement learning algorithm, DDPG, is used to generate actions for each round, including terminal selection and power allocation. This allows the FL system to not only maximize model accuracy and accelerate convergence within a limited timeframe, but also comprehensively considers multiple factors, enabling the system to meet model accuracy requirements while minimizing energy consumption within the constraints of remaining battery power and transmission bandwidth on terminal devices, in line with the development of green communications.
[0181] The embodiments provided above in this application introduce the methods provided in the embodiments of this application. To implement the various functions of the methods provided in the embodiments of this application, the network device may include a hardware structure and a software module, and implement the aforementioned functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. A particular function of the aforementioned functions may be implemented in the form of a hardware structure, a software module, or a hardware structure plus a software module.
[0182] Corresponding to the terminal selection methods provided in the above-mentioned embodiments, the present disclosure also provides a terminal selection device. Since the terminal selection device provided in the embodiment of the present disclosure corresponds to the terminal selection methods provided in the above-mentioned embodiments, the implementation method of the terminal selection method is also applicable to the terminal selection device provided in this embodiment and will not be described in detail in this embodiment.
[0183] Figure 4 is a structural diagram of a terminal selection device 400 provided in an embodiment of the present disclosure. As shown in Figure 4, the device 400 may include: a receiving module 410, used to receive model training parameters sent by at least one terminal, wherein the at least one terminal is a terminal selected by a network device from multiple terminals for performing a first model training on a global model issued by the network device, and the model training parameters are obtained after the at least one terminal performs the first model training on the global model; a determination module 420, used to determine the data quality parameters and energy consumption parameters of the terminals participating in the second model training based on the model training parameters; a selection module 430, used to select the terminals participating in the second model training from multiple terminals based on the data quality parameters and energy consumption parameters.
[0184] In some embodiments, the model training parameters include the model weight of the first model training of at least one terminal, and the determination module 420 is used to: determine the aggregated model weight of the first model training based on the model weight and the local data size of at least one terminal; determine the weight divergence based on the difference between the model weight and the aggregated model weight, and the weight divergence is used as a data quality parameter.
[0185] In some embodiments, the model training parameters include the local data distribution value of the first model training of at least one terminal, and the determination module 420 is used to: weightedly combine the weight divergence and the local data distribution value to determine the quality score value, and the quality score value is used as the data quality parameter.
[0186] In some embodiments, the model training parameters include local training energy consumption of the first model training of at least one terminal, and the determination module 420 is used to determine the energy consumption parameters based on the local training energy consumption and the transmission energy consumption of uploading the model training parameters by at least one terminal.
[0187] In some embodiments, the selection module 430 is used to: establish an objective function based on data quality parameters and energy consumption parameters; establish constraints, which include at least one of time constraints, energy consumption constraints, bandwidth constraints, terminal processor cycle frequency constraints, transmission power constraints, and terminal selection decision constraints; solve the objective function under the constraints according to the deep deterministic policy gradient DDPG algorithm to select terminals participating in the second model training from multiple terminals.
[0188] In some embodiments, the selection module 430 is used to: represent the objective function with a Markov quadruple, where the Markov quadruple includes a state space, an action space, a reward function, and a transition probability; establish a policy gradient and a loss function under a policy-evaluation network, and determine the optimal state-action pair to select a terminal participating in the second model training from multiple terminals.
[0189] In some embodiments, the state space includes data quality parameters, uplink channel gain, channel bandwidth from the terminal to the network device, and the remaining energy of the terminal; the action space includes terminal selection parameters and the transmission power of the selected terminal; the reward function is the ratio of the sum of the data quality parameters of the terminals participating in the second model training to the energy consumption; the transition probability is the state transition probability.
[0190] In summary, according to the terminal selection device provided by the embodiment of the present disclosure, the network device receives model training parameters sent by at least one terminal, wherein at least one terminal is selected by the network device from multiple terminals for performing the first model training on the global model sent by the network device, and the model training parameters are obtained after the at least one terminal performs the first model training on the global model; based on the model training parameters, the data quality parameters and energy consumption parameters of the terminals participating in the second model training are determined; based on the data quality parameters and energy consumption parameters, the terminals participating in the second model training are selected from multiple terminals. The solution of the present disclosure comprehensively considers factors such as terminal energy consumption and data quality, thereby biasedly selecting the most suitable terminal in each round of global iteration to balance the energy consumption of the edge terminal and the learning accuracy and convergence speed of the FL.
[0191] Corresponding to the terminal selection methods provided in the above-mentioned embodiments, the present disclosure also provides a terminal selection device. Since the terminal selection device provided in the embodiment of the present disclosure corresponds to the terminal selection methods provided in the above-mentioned embodiments, the implementation method of the terminal selection method is also applicable to the terminal selection device provided in this embodiment and will not be described in detail in this embodiment.
[0192] Figure 5 is a structural diagram of a model training device 500 provided by an embodiment of the present disclosure. As shown in Figure 5, the device 500 may include: a creation module 510 for creating an initial global model using public data; a transceiver module 520 for receiving resource information sent by multiple terminals; a selection module 530 for selecting an initial training terminal from the multiple terminals based on the resource information; the transceiver module 520 is also used to: send the initial global model to the initial training terminal, and receive the first model training parameters uploaded by the initial training terminal after training the initial global model; an aggregation module 540 is used to perform model aggregation to generate a global model for first model training; the selection module 530 is also used to: determine the data quality parameters and energy consumption parameters of the first terminal participating in the first model training according to the first model training parameters, and select the first terminal according to the first model training parameters. The data quality parameters and the energy consumption parameters are used to determine the first terminal for the first model training from the multiple terminals; the transceiver module 520 is also used to: send the global model used for the first model training to the first terminal for training, and receive the second model training parameters uploaded by the first terminal after the global model training; the aggregation module 540 is also used to: perform model aggregation to generate a global model for the second model training; the selection module 530 is also used to: determine the data quality parameters and energy consumption parameters of the second terminal participating in the second model training according to the second model training parameters, determine the second terminal for the second model training from the multiple terminals according to the data quality parameters and the energy consumption parameters, and send the global model used for the second model training to the second terminal for training until the model accuracy converges.
[0193] In summary, the model training device provided by the embodiments of the present disclosure specifically addresses the heterogeneous distribution of device data in FL systems by designing a strategy for terminal selection and resource allocation within a federated learning model. To reduce the time complexity of planning, the latest deep reinforcement learning algorithm, DDPG, is used to generate actions for each round, including terminal selection and power allocation. This enables the FL system to not only maximize model accuracy and accelerate convergence within a limited timeframe, but also comprehensively considers multiple factors, enabling the system to meet model accuracy requirements while minimizing energy consumption within the constraints of the remaining battery power and transmission bandwidth of terminal devices, in line with the development of green communications.
[0194] The present application provides a communication system, including: a network device and a terminal, wherein: the network device is configured to execute the terminal selection method or model training method shown in the embodiments of Figures 1-3 of the present disclosure.
[0195] Please refer to Figure 6, which is a schematic diagram of the structure of a communication device 600 provided in an embodiment of the present application. Communication device 600 can be a network device, a user device, a chip, a chip system, or a processor that supports the network device to implement the above method, or a chip, a chip system, or a processor that supports the user device to implement the above method. This device can be used to implement the method described in the above method embodiment. For details, please refer to the description of the above method embodiment.
[0196] The communication device 600 may include one or more processors 601. The processor 601 may be a general-purpose processor or a dedicated processor. For example, it may be a baseband processor or a central processing unit. The baseband processor may be used to process communication protocols and communication data, and the central processing unit may be used to control the communication device (e.g., a base station, a baseband chip, a terminal device, a terminal device chip, a DU or CU, etc.), execute computer programs, and process computer program data.
[0197] Optionally, the communication device 600 may further include one or more memories 602, on which a computer program 604 may be stored. The processor 601 executes the computer program 604 to enable the communication device 600 to perform the method described in the above method embodiment. Optionally, the memory 602 may also store data. The communication device 600 and the memory 602 may be provided separately or integrated together.
[0198] Optionally, the communication device 600 may further include a transceiver 605 and an antenna 606. The transceiver 605 may be referred to as a transceiver unit, a transceiver, or a transceiver circuit, and is configured to implement transceiver functions. The transceiver 605 may include a receiver and a transmitter. The receiver may be referred to as a receiver or a receiving circuit, and is configured to implement a receiving function; the transmitter may be referred to as a transmitter or a transmitting circuit, and is configured to implement a transmitting function.
[0199] Optionally, the communication device 600 may further include one or more interface circuits 607. The interface circuit 607 is configured to receive code instructions and transmit the code instructions to the processor 601. The processor 601 executes the code instructions to enable the communication device 600 to perform the method described in the above method embodiment.
[0200] In one implementation, processor 601 may include a transceiver for implementing receiving and transmitting functions. For example, the transceiver may be a transceiver circuit, an interface, or an interface circuit. The transceiver circuit, interface, or interface circuit for implementing the receiving and transmitting functions may be separate or integrated. The transceiver circuit, interface, or interface circuit may be used for reading and writing code / data, or may be used for transmitting or delivering signals.
[0201] In one implementation, processor 601 may store a computer program 603. Computer program 603, when executed on processor 601, enables communication device 600 to perform the method described in the above method embodiment. Computer program 603 may be embedded in processor 601, in which case processor 601 may be implemented by hardware.
[0202] In one implementation, the communication device 600 may include a circuit that can implement the functions of sending, receiving, or communicating in the aforementioned method embodiments. The processor and transceiver described in this application can be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit RFIC, a mixed signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (NMetal-Oxide-Semiconductor, NMOS), P-type metal oxide semiconductor (Positive Channel Metal Oxide Semiconductor, PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.
[0203] The communication device described in the above embodiments may be a network device or a user device, but the scope of the communication device described in this application is not limited thereto, and the structure of the communication device may not be limited to FIG6 . The communication device may be an independent device or may be part of a larger device. For example, the communication device may be:
[0204] (1) An independent integrated circuit (IC), or chip, or chip system or subsystem;
[0205] (2) a collection of one or more ICs, optionally including a storage component for storing data and computer programs;
[0206] (3) ASIC, such as modem;
[0207] (4) Modules that can be embedded in other devices;
[0208] (5) Receivers, terminal devices, intelligent terminal devices, cellular phones, wireless devices, handheld devices, mobile units, vehicle-mounted devices, network devices, cloud devices, artificial intelligence devices, etc.;
[0209] (6)Others, etc.
[0210] If the communication device can be a chip or a chip system, please refer to the schematic diagram of the chip structure shown in Figure 7. The chip shown in Figure 7 includes a processor 701 and an interface 702. The number of processors 701 can be one or more, and the number of interfaces 702 can be multiple.
[0211] Optionally, the chip further includes a memory 703, which is used to store necessary computer programs and data.
[0212] Those skilled in the art will also appreciate that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present application.
[0213] The present application also provides a readable storage medium having instructions stored thereon, which implement the functions of any of the above method embodiments when executed by a computer.
[0214] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs. When the computer program is loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer program can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. Available media may be magnetic media (eg, floppy disks, hard disks, tapes), optical media (eg, high-density digital video discs (DVDs)), or semiconductor media (eg, solid state disks (SSDs)).
[0215] Those skilled in the art will understand that the various numerical numbers such as first and second involved in this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application, and also indicate the order of precedence.
[0216] In this application, at least one can also be described as one or more, and multiple can be two, three, four or more, which is not limited in this application. In the embodiments of this application, for a technical feature, the technical features in the technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", and there is no order of precedence or size between the technical features described by "first", "second", "third", "A", "B", "C" and "D".
[0217] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0218] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0219] A computer system may include terminals and servers. The terminals and servers are generally remote from each other and typically interact through a communication network. The relationship of a terminal and server arises through computer programs running on the respective computers and having a terminal-server relationship to each other.
[0220] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0221] Furthermore, it should be understood that the various embodiments of the present application may be implemented individually or in combination with other embodiments where the solution permits.
[0222] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0223] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0224] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A terminal selection method, characterized in that: The method is performed by a network device, and includes: Receiving model training parameters sent by at least one terminal, wherein the at least one terminal is a terminal selected by the network device from multiple terminals for performing a first model training on a global model sent by the network device, and the model training parameters are obtained by the at least one terminal after performing the first model training on the global model; Determining data quality parameters and energy consumption parameters of terminals participating in second model training based on the model training parameters; A terminal participating in second model training is selected from the multiple terminals according to the data quality parameter and the energy consumption parameter.
2. The method according to claim 1, characterized in that The model training parameters include a model weight for the at least one terminal to perform the first model training, and the data quality parameters for the terminals participating in the second model training are determined to include: Determining an aggregated model weight for first model training based on the model weight and the local data size of the at least one terminal; A weight divergence is determined according to a difference between the model weight and the aggregated model weight, and the weight divergence is used as the data quality parameter.
3. The method according to claim 2, characterized in that The model training parameters include a local data distribution value of the at least one terminal for the first model training, and the data quality parameters of the terminals participating in the second model training are further determined: The weight divergence and the local data distribution value are weighted and combined to determine a quality score value, and the quality score value is used as the data quality parameter.
4. The method according to any one of claims 1 to 3, characterized in that The model training parameters include local training energy consumption of the at least one terminal for the first model training, and the energy consumption parameters of the terminals participating in the second model training are determined to include: The energy consumption parameter is determined according to the local training energy consumption and the transmission energy consumption of the model training parameter uploaded by the at least one terminal.
5. The method according to any one of claims 1 to 4, characterized in that The selecting, from the plurality of terminals, a terminal to participate in the second model training according to the data quality parameter and the energy consumption parameter comprises: Establishing an objective function according to the data quality parameter and the energy consumption parameter; Establishing constraint conditions, wherein the constraint conditions include at least one of a time constraint, an energy consumption constraint, a bandwidth constraint, a terminal processor cycle frequency constraint, a transmission power constraint, and a terminal selection decision constraint; According to the deep deterministic policy gradient DDPG algorithm, the objective function is solved under the constraints to select a terminal participating in the second model training from the multiple terminals.
6. The method according to claim 5, characterized in that Solving the objective function under the constraint conditions according to the DDPG algorithm to select a terminal to participate in the second model training from the multiple terminals includes: Representing the objective function as a Markov quadruple, wherein the Markov quadruple includes a state space, an action space, a reward function, and a transition probability; A policy gradient and a loss function are established under a policy-evaluation network to determine an optimal state-action pair to select a terminal participating in the second model training from the multiple terminals.
7. The method according to claim 6, characterized in that The state space includes the data quality parameter, the channel gain of the uplink, the channel bandwidth from the terminal to the network device, and the remaining energy of the terminal; the action space includes the terminal selection parameter and the transmission power of the selected terminal; the reward function is the ratio of the sum of the data quality parameters of the terminals participating in the second model training to the energy consumption; the transition probability is the state transition probability.
8. A model training method, characterized in that: The method is performed by a network device, and includes: Create an initial global model using public data; Receive resource information sent by multiple terminals; selecting an initial training terminal from the plurality of terminals based on the resource information; Sending the initial global model to the initial training terminal, and receiving the first model training parameters uploaded by the initial training terminal after training the initial global model; performing model aggregation to generate a global model for training the first model; Determining, based on the first model training parameters, a data quality parameter and an energy consumption parameter of a first terminal participating in the first model training; determining, based on the data quality parameter and the energy consumption parameter, a first terminal to perform the first model training from the multiple terminals; and delivering a global model for the first model training to the first terminal for training; Receiving second model training parameters uploaded by the first terminal after training the global model; performing model aggregation to generate a global model for training a second model; According to the second model training parameters, the data quality parameters and energy consumption parameters of the second terminal participating in the second model training are determined; according to the data quality parameters and the energy consumption parameters, the second terminal for performing the second model training is determined from the multiple terminals; the global model used for the second model training is sent to the second terminal for training until the model accuracy converges.
9. A terminal selection device, characterized in that: The device comprises: a receiving module, configured to receive model training parameters sent by at least one terminal, wherein the at least one terminal is a terminal selected by the network device from multiple terminals for performing a first model training on the global model sent by the network device, and the model training parameters are obtained by the at least one terminal after performing the first model training on the global model; A determination module, configured to determine data quality parameters and energy consumption parameters of terminals participating in the second model training based on the model training parameters; A selection module is used to select a terminal participating in the second model training from the multiple terminals based on the data quality parameter and the energy consumption parameter.
10. A model training device, characterized in that: The device comprises: Creation module for creating an initial global model using public data; A transceiver module, configured to receive resource information sent by multiple terminals; A selection module, configured to select an initial training terminal from the plurality of terminals based on the resource information; The transceiver module is further configured to: send the initial global model to the initial training terminal, and receive the first model training parameters uploaded by the initial training terminal after training the initial global model; an aggregation module, configured to perform model aggregation to generate a global model for training the first model; The selection module is further configured to: determine, based on the first model training parameters, a data quality parameter and an energy consumption parameter of a first terminal participating in the first model training; and determine, based on the data quality parameter and the energy consumption parameter, a first terminal to perform the first model training from the multiple terminals; The transceiver module is further configured to: send the global model used for training the first model to the first terminal for training, and receive the second model training parameters uploaded by the first terminal after training the global model; The aggregation module is further configured to: perform model aggregation to generate a global model for second model training; The selection module is also used to: determine the data quality parameters and energy consumption parameters of the second terminal participating in the second model training based on the second model training parameters; determine the second terminal for performing the second model training from the multiple terminals based on the data quality parameters and the energy consumption parameters; and send the global model used for the second model training to the second terminal for training until the model accuracy converges.
11. A communication device, wherein: include: transceiver; Memory; A processor is connected to the transceiver and the memory respectively, and is configured to control the wireless signal reception and transmission of the transceiver by executing computer-executable instructions on the memory, and can implement the method according to any one of claims 1 to 8.
12. A computer storage medium, wherein: The computer storage medium stores computer-executable instructions; after the computer-executable instructions are executed by the processor, the method according to any one of claims 1 to 8 can be implemented.