Model training methods, information recommendation methods and devices, media, and electronic equipment

By training a reinforcement learning model on the client side and uploading the parameters to the cloud for aggregation, and combining the results with user feedback for iterative federated learning, the problems of inaccurate data plan recommendations and poor security are solved, achieving efficient and secure model training and recommendation.

CN117196721BActive Publication Date: 2025-10-28CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311274229.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-10-28
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

In existing technologies, data plan recommendations are inaccurate and have poor security because centralized machine learning frameworks require a large amount of network resources to transmit user data, resulting in low processing efficiency and an inability to protect user privacy.

Method used

We employ a reinforcement learning model combined with federated learning. The model is trained on the client side and the parameters are uploaded to the cloud server for aggregation, avoiding the direct upload of training data. We use user feedback results to train and iterate the model, and select appropriate candidate clients for federated learning to accelerate training.

Benefits of technology

It improves model training efficiency and security, protects user privacy, makes the model more realistic, and enhances the accuracy and reliability of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117196721B_ABST
    Figure CN117196721B_ABST
Patent Text Reader

Abstract

This disclosure relates to a model training method, an information recommendation method and apparatus, and an electronic device, belonging to the field of computer technology. The method includes: acquiring a reinforcement learning model; training the reinforcement learning model and sending the trained reinforcement learning model to a cloud server to obtain a latest reinforcement learning model; fitting the user's traffic usage status to the latest reinforcement learning model to obtain an output action; and generating sample data based on the traffic usage status, the output action, the feedback result corresponding to the output action, and the next traffic usage status; and performing iterative federated learning on the latest reinforcement learning model using the sample data until the model converges to obtain a target reinforcement learning model for determining information recommendation suggestions. This disclosure can improve the accuracy of the model, thereby improving the precision of information recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a model training method, an information recommendation method, a model training device, an information recommendation device, a computer-readable storage medium, and an electronic device. Background Technology

[0002] Currently, with the continuous development of hardware and software, the scenarios in which users consume data are becoming increasingly diverse. In different scenarios, users have different data needs. Therefore, it is necessary to accurately identify users' current data needs in order to recommend corresponding data plans and provide more suitable services.

[0003] In related technologies, a centralized machine learning framework is generally used for traffic recommendation. However, in this approach, the large number of users on the operator's platform means that transmitting training data to the operator's data center consumes significant network resources, resulting in poor processing efficiency. Furthermore, uploading all user information to the server compromises user privacy and security. Summary of the Invention

[0004] The purpose of this disclosure is to provide a model training method, an information recommendation method, a model training device, an information recommendation device, a computer-readable storage medium, and an electronic device, thereby overcoming, to at least a certain extent, the problem of inaccurate data plan recommendations caused by the limitations and defects of related technologies.

[0005] According to one aspect of this disclosure, a model training method is provided, comprising: acquiring a reinforcement learning model; training the reinforcement learning model and sending the trained reinforcement learning model to a cloud server to obtain a latest reinforcement learning model; fitting the user's traffic usage status to the latest reinforcement learning model to obtain an output action, and generating sample data based on the traffic usage status, the output action, the feedback result corresponding to the output action, and the next traffic usage status; and using the sample data to perform iterative federated learning on the latest reinforcement learning model until the model converges to obtain a target reinforcement learning model for determining information recommendation suggestions.

[0006] In one exemplary embodiment of this disclosure, training the reinforcement learning model and sending the trained reinforcement learning model to a cloud server to obtain the latest reinforcement learning model includes: training the reinforcement learning model based on the client and sending the model parameters of the trained reinforcement learning model to the cloud server; obtaining a global model obtained by fusing the trained reinforcement learning model through the cloud server, and training based on the global model to obtain the latest reinforcement learning model corresponding to each client.

[0007] In one exemplary embodiment of this disclosure, the step of training the reinforcement learning model based on the client includes: inputting the historical traffic usage status corresponding to the historical time slot into the reinforcement learning model to obtain the output action corresponding to the historical traffic usage status and the next traffic usage status; determining the reward for the output action corresponding to the historical traffic usage status, and adjusting the parameters of the reinforcement learning model according to the reward.

[0008] In one exemplary embodiment of this disclosure, generating sample data based on the traffic usage status, the output action, the feedback result corresponding to the output action, and the next traffic usage status includes: determining a reward based on the feedback result and cost information, and determining the next traffic usage status; storing the traffic usage status, the output action, the reward, and the next traffic usage status into the client's sample pool as sample data for the client.

[0009] In one exemplary embodiment of this disclosure, the step of using the sample data to perform iterative federated learning on the latest reinforcement learning model until the model converges to determine the target reinforcement learning model includes: if the number of sample data is greater than a preset value, determining candidate clients to participate in federated learning; performing global federated learning based on the candidate clients, and continuing to select candidate clients to participate in federated learning to obtain reselected candidate clients; performing multiple rounds of iterative global federated learning based on the reselected candidate clients to obtain the latest global model; and training the model based on the convergence of the latest global model.

[0010] In one exemplary embodiment of this disclosure, determining the candidate clients to participate in federated learning includes: randomly selecting multiple clients to participate in federated learning in a preset round of federated learning; predicting the predicted training time of each client based on historical training time, and determining the probability of each client being selected in the target round of training based on the predicted training time of each client; randomly sampling all clients based on the probability to determine the candidate clients to participate in the current round of federated learning.

[0011] In one exemplary embodiment of this disclosure, the step of training the model based on the convergence of the latest global model includes: if the latest global model converges, ending the model training and using the latest global model as the target reinforcement learning model; if the latest global model does not converge, re-obtaining the output action and generating sample data based on the user's traffic usage status until the model converges, so as to obtain the target reinforcement learning model.

[0012] According to one aspect of this disclosure, an information recommendation method is provided, comprising: obtaining a user's current data usage status; inputting the current data usage status into a target reinforcement learning model to obtain information recommendation suggestions for the user; wherein the target reinforcement learning model is trained according to any one of the model training methods described above.

[0013] According to one aspect of this disclosure, a model training apparatus is provided, comprising: a reinforcement learning module for acquiring a reinforcement learning model, training the reinforcement learning model, and sending the trained reinforcement learning model to a cloud server to obtain a latest reinforcement learning model; a fitting module for fitting a user's traffic usage state to obtain an output action based on the latest reinforcement learning model, and generating sample data based on the traffic usage state, the output action, the feedback result corresponding to the output action, and the next traffic usage state; and a federated learning module for iteratively federating the latest reinforcement learning model using the sample data until the model converges to obtain a target reinforcement learning model for determining information recommendation suggestions.

[0014] According to one aspect of this disclosure, an information recommendation device is provided, comprising: a status acquisition module for acquiring a user's current data usage status; and an information recommendation determination module for inputting the current data usage status into a target reinforcement learning model to obtain information recommendation suggestions for the user; wherein the target reinforcement learning model is trained according to any one of the model training methods described above.

[0015] According to one aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the model training method or the information recommendation method described in any one of the preceding claims.

[0016] According to one aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the model training method or the information recommendation method described above by executing the executable instructions.

[0017] The technical solution provided in this disclosure has several advantages. First, it allows for direct training of the reinforcement learning model on the client side to obtain the latest model. During training, there's no need to upload training data to a server; instead, the trained model is directly sent to a cloud server, reducing network resource consumption, improving resource utilization, and increasing training efficiency. Second, the elimination of the need to upload training data avoids privacy breaches and enhances security. Third, by incorporating user feedback on output actions during training, the model becomes more realistic, improving accuracy and reliability.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0020] Figure 1 The flowchart illustrating a model training method in an embodiment of this disclosure is shown schematically.

[0021] Figure 2 A schematic diagram of a federated learning system in an embodiment of this disclosure is shown.

[0022] Figure 3 This illustration shows a schematic diagram of training a reinforcement learning model based on a client in an embodiment of this disclosure.

[0023] Figure 4 The schematic diagram illustrates the process of selecting candidate clients in an embodiment of this disclosure.

[0024] Figure 5 The diagram illustrates the overall process of model training according to an embodiment of the present disclosure.

[0025] Figure 6 The illustration shows a schematic diagram of the information recommendation process in an embodiment of this disclosure.

[0026] Figure 7 The schematic diagram illustrates a block diagram of a model training apparatus in an embodiment of this disclosure.

[0027] Figure 8 The diagram illustrates a block diagram of an information recommendation device in an embodiment of this disclosure.

[0028] Figure 9 A schematic block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0030] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0031] In some embodiments of this disclosure, data usage recommendation can mainly include the following steps: obtaining a target time series, which includes data usage data of a target user over several consecutive days; determining the target user's expected data usage for the remaining time of the month based on the target time series; determining the target user's expected total data usage for the month based on the expected data usage; determining the similarity between the target user and each data plan based on the expected total data usage and the number of days the target user used data in the month; and determining the target data plan to be recommended to the target user based on the similarity. This allows for reasonable prediction of expected data usage based on the target time series. Furthermore, determining the similarity between the target user and the data plan based on the expected total data usage and the number of days the target user used data in the month allows for more accurate matching of target data plans.

[0032] The aforementioned technical solution employs supervised centralized machine learning, requiring the prior collection of a large amount of user traffic usage data as input and traffic usage data for the remaining time of the month as output to train a usable model. This necessitates significant time and resources for collecting and transmitting user data, and the use of numerous high-performance GPUs for model training. Furthermore, the large-scale transmission of user traffic usage data may lead to security issues and compromise user data privacy.

[0033] To address the aforementioned technical issues, this disclosure provides a model training method that primarily employs joint reinforcement learning and federated learning for model training, and then uses the trained model to perform traffic recommendations. This method allows for simultaneous data collection and model training, while simultaneously providing recommendations to users. Furthermore, it refines and adjusts the model based on user feedback, thereby improving model training efficiency and data utilization efficiency, and ultimately making the model more accurate.

[0034] Next, refer to Figure 1 The diagram illustrates the model training method in the embodiments of this disclosure.

[0035] In step S110, a reinforcement learning model is obtained, the reinforcement learning model is trained, and the trained reinforcement learning model is sent to the cloud server to obtain the latest reinforcement learning model.

[0036] In this embodiment of the disclosure, since users' data usage requirements may vary significantly across different application scenarios, such as the different data requirements at home and while traveling, the different data requirements for browsing text and watching videos, and the different data requirements for specific application categories, a reinforcement learning model can be trained using a combination of reinforcement learning and federated learning processes to recommend data usage to clients in order to improve the accuracy of data plan recommendations.

[0037] Reinforcement learning (RL) is a machine learning approach whose main process includes: an agent continuously interacting with the environment and learning from its feedback. At each decision point, the state of the environment serves as the agent's input, and the agent outputs an appropriate action based on the state. After the agent executes an action, the environment undergoes a state transition and provides feedback to the agent. The agent learns from the environment's feedback whether the chosen action is appropriate and makes adjustments, gradually achieving the target effect. In reinforcement learning, the agent is not trained on a pre-defined labeled dataset but learns through interaction with the environment, thereby improving performance. In this embodiment, reinforcement learning uses a reward function to determine the merits of performing an action in a given state. The learning process involves changing the strategy for executing actions through reward signals, ultimately forming a strategy that maximizes the reward. Reinforcement learning mainly consists of an agent, environment, state, action, and reward. After the agent performs an action, the environment transitions to a new state, providing a reward signal (positive or negative) for this new state. Subsequently, the agent executes a new action according to a certain strategy based on the new state and the reward feedback from the environment. The above process describes how intelligent agents and the environment interact through states, actions, and rewards.

[0038] The hierarchical federated learning system in this embodiment is divided into three layers: client, edge server, and cloud server. (See reference...) Figure 2 As shown, a cloud server refers to a server with abundant computing and network resources. An edge server refers to a server with a certain computing capacity, such as a WiFi access point or base station. There can be one or more edge servers. A client refers to a user device that uses data traffic, such as a smartphone, tablet, smart robot, or other device capable of using data traffic. There can also be one or more clients. Specifically, cloud servers can communicate with edge servers, edge servers can communicate with both cloud servers and clients, and clients can communicate with edge servers.

[0039] Each client can have a corresponding reinforcement learning model, and different clients can train models with the cloud server through a federated learning process. Federated learning can address the data privacy issues in traditional centralized machine learning, enabling multiple clients to collaboratively train models without sharing data. During federated learning model training, each client only needs to transmit the model parameters or gradients trained iteratively on its local data to the cloud server, which then aggregates the models from all clients. Finally, the global model on the cloud server is broadcast to the clients for the next iteration.

[0040] For example, the cloud server can initialize a global model (i.e., global RL) and distribute it to all edge servers. The edge servers then distribute the global model to clients connected to that edge server as reinforcement learning models for those clients.

[0041] Based on this, each client has a local reinforcement learning model. Each client can train its own reinforcement learning model and send the model parameters to the cloud server. The cloud server can then combine the model parameters uploaded by each client to determine new model parameters, thus achieving model aggregation to obtain a global model. Furthermore, the cloud server can distribute the global model to the clients, and the clients can deploy the new model parameters onto their models, resulting in the latest reinforcement learning model for each client. Training can then continue based on this latest model, and this training process can be iterated repeatedly until the model converges or other conditions are met. Model aggregation can be performed using any of the following methods: simple averaging, weighted averaging, federated averaging, or a hybrid approach.

[0042] refer to Figure 3 As shown, client 301 can train model 1 and upload parameter 1 of model 1 to cloud server 304; client 302 can train model 2 and upload parameter 2 of model 2 to cloud server; client 303 can train model 1 and upload parameter 3 of model 3 to cloud server. The cloud server can aggregate the parameters of model 1, model 2, and model 3 to obtain global model 4, and distribute the parameters of global model 4 to each client to obtain the latest reinforcement learning model.

[0043] Next, the process of training the reinforcement learning model on the client side will be described. In this embodiment, the reinforcement learning model can be defined. Specifically, the input to the reinforcement learning model can be used to describe the user's data usage characteristics; the input can specifically be the user's data usage status. The output of the reinforcement learning model can be information recommendation suggestions for the user. These suggestions can be recommendations for data usage information, which can be data plans, i.e., each data plan corresponds to an action. It should be noted that the information recommendation suggestions can also be for call charges, etc.; here, data plan recommendations are used as an example for explanation.

[0044] In some embodiments, the traffic usage status here can be used to describe the user's traffic usage characteristics over the past k time slots. The traffic usage status can include one or more of traffic usage status and transmission rate. Traffic usage status can include traffic usage and remaining available traffic; transmission rate can include uplink rate and downlink rate, etc. Specifically, the input to the reinforcement learning model can be expressed as formula (1):

[0045]

[0046] in, It represents the traffic usage of a certain application i in time slot t-1, and N represents the number of applications that are running. This represents the user's uplink rate in time slot t-1. This represents the user's downlink rate in time slot t-1, l t-1 This indicates the remaining available traffic for the user in time slot t-1.

[0047] After determining the input to the reinforcement learning model, it can be fed into the model for fitting to obtain the corresponding output action. The output action can be a recommendation for a user's data plan; that is, each plan corresponds to one action, and the output can be represented as... in, This indicates whether package i is recommended, and M represents all package types. Each value in the output can be 0 or 1. If a package's value is 1, it is recommended; if a package's value is 0, it is not recommended. It should be noted that... This indicates that no data plan is recommended. Carriers can recommend corresponding data plans to users based on the data plan value of 1 in the output action.

[0048] In addition, the reward R of the reinforcement learning model can be determined. t The reward can be used to train the model. The reward for the reinforcement learning model can be determined by subtracting the cost incurred by the operator from the revenue gained by the operator from recommending the data plan. Based on this, the reward for the reinforcement learning model can be expressed by formula (2):

[0049] R t =p t -h t Formula (2)

[0050] Where, p t Used to represent the revenue that an operator gains from recommending data plans, h t This is used to represent the costs incurred by operators due to recommending data plans.

[0051] It should be noted that the reward can be determined based on the user's feedback regarding the output action, which can be either purchasing the recommended plan or not purchasing it. Therefore, the reward can be determined based on the feedback and the costs incurred by the operator due to recommending the plan.

[0052] After defining a reinforcement learning model, the client can train the model to improve its accuracy and reliability. In some embodiments, deep reinforcement learning algorithms can be used to train the model. These algorithms can include any one of DQN (Deep Q Network), DDQN (Double Deep Q Network), and DDPG (deep deterministic policy gradient). This section uses DQN as an example.

[0053] During reinforcement learning model training, the environmental state is first perceived, which can be the historical traffic usage state. The historical traffic usage state represents the traffic usage characteristics of a user over k historical time slots. This historical traffic usage state is input into the reinforcement learning model to obtain the corresponding output action. This output action, when applied to the environment, causes a change in the environmental state, transitioning the historical traffic usage state to the next traffic usage state. Simultaneously, a feedback signal is generated, which provides a positive or negative reward for the next traffic usage state. Furthermore, the reward can be used to adjust the policy for executing the action. The policy refers to the mapping from environmental state to action; therefore, it can be considered that the parameters of the reinforcement learning model are adjusted through the reward to generate the reinforcement learning model that maximizes the reward, which is then used as the reinforcement learning model trained on each client. The reinforcement learning models trained on each client can be the same or different, depending on the actual data.

[0054] Next, the model parameters of the reinforcement learning model trained by each client can be sent to the cloud server, allowing the cloud server to perform aggregate training on the models to obtain a global model. Furthermore, the cloud server can distribute the global model to clients for further training, resulting in the latest reinforcement learning model. In this process, although the amount of data on each client is limited, data sample sharing is achieved through aggregation. Each client cannot see the data of other clients, thus improving privacy.

[0055] Next, continue to refer to Figure 1 As shown, in step S120, the user's traffic usage status is fitted according to the latest reinforcement learning model to obtain the output action, and the feedback result corresponding to the output action is determined.

[0056] In this embodiment, after adjusting the client's local reinforcement learning model based on the global model to obtain the latest reinforcement learning model, the user's traffic usage status can be input into the latest reinforcement learning model for fitting to obtain the corresponding output action. Simultaneously, feedback results corresponding to this output action can also be obtained. These feedback results can be used to indicate the user's execution status of the output action, such as whether the recommended package was purchased or not.

[0057] For example, the traffic usage status S t The input is fed into the latest reinforcement learning model to obtain the output action A. t Each value in the output action indicates whether a specific data plan has been selected. Furthermore, it can transition to the next data usage state S. t+1 At the same time, user feedback on the recommended plans will be obtained, such as whether they have decided to purchase them. Furthermore, the operator's revenue from the recommended plans can be determined based on this feedback, and a reward R can be determined based on this revenue and the operator's costs associated with recommending the plans. t .

[0058] Based on this, traffic usage status, output actions, corresponding feedback results, and the next traffic usage status can be combined to generate sample data. Sample data can be represented as (S... t A t ,R t ,S t+1 Furthermore, sample data can be stored in the client's sample pool for model training. This process incorporates actual user feedback during model training, thereby improving model accuracy.

[0059] Continue to refer to Figure 1 As shown, in step S130, the latest reinforcement learning model is iteratively federated learning is performed using the sample data until the model converges, so as to obtain the target reinforcement learning model used to determine information recommendation suggestions.

[0060] In this embodiment of the disclosure, after determining the sample data, sample data collection and model aggregation can be performed simultaneously. For example, sample data collection can be achieved by iteratively fitting based on the user's traffic usage status to generate new sample data. Simultaneously, a reward can be determined based on the traffic usage status, output action, the feedback result corresponding to the output action, and the next traffic usage status. This reward is then used to update the reinforcement learning model, which is then sent to a cloud server for aggregation to obtain the latest reinforcement learning model.

[0061] Next, the decision to continue model training or redefine the output action can be made based on whether the quantity of sample data in the sample pool meets preset conditions. During model training, local model training can first be performed using sample data, i.e., training the latest reinforcement learning model locally on the client. Further, candidate clients can be selected to participate in federated learning, and the latest reinforcement learning model trained on the candidate clients is iteratively trained through the federated learning process to obtain the final target reinforcement learning model. The target reinforcement learning model can be used to determine traffic recommendation suggestions for each user.

[0062] When training a model using sample data, the sufficiency of the sample data can be determined by checking if the number of sample data points exceeds a preset value. If the sample data is sufficient, the preset condition can be considered met. The size of the preset value can be determined based on actual needs. If the number of sample data points is less than the preset value, the sample data collection process continues.

[0063] If the number of samples in the sample pool is greater than or equal to a preset value, local model training can be performed based on the sample data in the sample pool, i.e., training the latest reinforcement learning model. Further, candidate clients can be selected to participate in federated learning, so that the models trained on the latest reinforcement learning models of these candidate clients can participate in federated learning. Specifically, multiple rounds of global federated learning can be iteratively executed to obtain the latest global model, and the convergence of the latest global model is used to determine whether to terminate the model training process. During the iterative execution of multiple rounds of global federated learning, to improve training efficiency, candidate clients can be selected to participate in federated learning, thereby iteratively performing multiple rounds of global federated learning based on the determined candidate clients to obtain the target reinforcement learning model.

[0064] In this embodiment, due to the large number of users of the operator and the significant differences in communication conditions among different users' clients, and since the training speed of the federated learning model generally depends on the slowest client, if all clients meeting the conditions for participating in HFL are involved in training, the HFL training process will be extremely slow. To solve the above problem, this embodiment uses a client sampling algorithm to select suitable candidate clients to participate in federated learning training based on their communication conditions, thereby accelerating the model training speed. Specifically, candidate clients participating in each round of federated learning training can be selected based on their communication conditions.

[0065] In some embodiments, reference Figure 4 As shown, selecting a candidate client may include the following steps:

[0066] In step S410, multiple clients are randomly selected to participate in federated learning in a preset round of federated learning;

[0067] In step S420, the predicted training time for each client is predicted based on the historical training time, and the probability of each client being selected in the target round of training is determined based on the predicted training time of each client.

[0068] In step S430, all clients are sampled according to the probability to determine candidate clients to participate in the current round.

[0069] In this embodiment of the disclosure, the sampling algorithm may specifically include the following steps: during a preset round of training, multiple clients may be randomly selected to participate in federated learning. The preset rounds refer to the initial multiple rounds of training, and the number of preset rounds is specifically set according to actual needs.

[0070] Furthermore, the training time of each client can be estimated based on historical training times. Specifically, the predicted training time of a client can be determined by the ratio of the logical processing result of whether each client participates in federated learning and the historical training time to the ratio of whether all clients participate in federated learning. The predicted training time of a client can be determined by formula (3):

[0071]

[0072] in This indicates whether client i participates in federated learning in round τ. The probability of each client being selected in round t can be determined by the predicted training time of each client. For example, the predicted training time can be used as an exponent for power operation, and the result of the power operation can be logically processed to obtain the probability that each client is selected in the target round. The target round can be the t-th round, that is, any round after the preset round. This probability can be expressed as formula (4):

[0073]

[0074] After determining the probability of each client being selected in the target round (round t), all clients can be sampled a predetermined number of times based on the aforementioned probabilities to determine candidate clients participating in the current round. The predetermined number of samplings can be 10 or 20, etc. If the same client is obtained from multiple samplings, that client is ignored, and only clients that were not sampled repeatedly are retained. The number of candidate clients can be one or more, and the candidate clients corresponding to different rounds can be different, depending on the actual probabilities. In this embodiment, clients are screened based on their communication conditions to obtain candidate clients participating in the current round of federated learning. By selecting a subset of suitable candidate clients to participate in federated learning training, the process of requiring all clients to participate in federated learning is avoided, thus improving the model training speed.

[0075] After identifying the candidate clients for the current round, global federated learning can be performed based on these candidates to obtain the global model for that round. After obtaining the global model for the current round, the next round can be used as the current round. The client sampling algorithm is then used to select candidate clients for the next round of federated learning, resulting in newly selected candidate clients. Another round of global federated learning is then performed based on these newly selected candidate clients, continuing until a predetermined number of rounds of federated learning have been completed, thus achieving multi-round iterative global federated learning. The predetermined number of rounds can be L rounds, which can be determined based on actual conditions or preset. After completing the predetermined number of rounds of federated learning based on the candidate clients, the latest global model is obtained.

[0076] In some embodiments, each global federated learning process based on candidate clients mainly includes the following steps:

[0077] The latest reinforcement learning model is trained based on the sample data from the candidate clients, and the trained reinforcement learning model is sent to the cloud server to obtain the global model for the current round. The cloud server distributes the global model for the current round to the candidate clients, and the candidate clients continue to train the global model for the current round to achieve one round of global federated learning. Next, the next round can be used as the current round to select candidate clients again based on the communication conditions of the clients, thereby continuing to complete one round of global federated learning. This process is repeated for multiple rounds until the number of rounds reaches a preset number, and the model obtained from multiple rounds of iterative global federated learning is used as the latest global model.

[0078] After L rounds of global federated learning are completed and the latest global model is obtained, it will be synchronously updated to all clients that have enabled real-time recommendation services. The latest global model deployed on each client will provide recommendations based on the user's current traffic usage. The operator will push these recommendations to the user and receive feedback, generating new sample data for subsequent training. Once the total amount of new sample data has accumulated to a certain level, the next round of global federated learning (HFL) will be conducted.

[0079] After obtaining the latest global model, the cloud server can distribute the latest global model to all edge servers, and the edge servers can distribute the model to the clients connected to them.

[0080] Furthermore, the convergence of the latest global model can be determined, and the model training process can be terminated based on the convergence of the latest global model.

[0081] For example, if the latest global model converges, the entire model training process ends. If the latest global model does not converge, the output action is re-obtained based on the user's traffic usage status, sample data is generated, and multiple rounds of global federated learning are performed based on the sample data until the latest global model converges. That is, steps S120-S140 are re-executed until the latest global model converges, and the converged latest global model is distributed as the target reinforcement learning model to the client and all edge servers.

[0082] In this embodiment of the disclosure, for hierarchical federated learning (HFL) architecture, after the client trains its local reinforcement learning model, the model parameters of the reinforcement learning model are uploaded to the cloud server for aggregation. This eliminates the need to upload the model to the cloud server for aggregation, thus avoiding network congestion on the cloud server and addressing the shortcomings of centralized cloud computing architectures when aggregating FL models.

[0083] Figure 5 The flowchart illustrating the model training process is shown in the image. Figure 5 As shown, the main steps include:

[0084] In step S502, the cloud server initializes the global model; the cloud server distributes the global model to the edge server, and the edge server distributes the model to the client; wherein, the global model refers to the global reinforcement learning model;

[0085] In step S504, the latest reinforcement learning model locally on the client is based on the client's state S. t Get output action A t Based on the output actions, the operator recommends corresponding packages to the user;

[0086] In step S506, the reward R is calculated based on the user's feedback on the recommended package and the cost of the recommended package. t And then transition to the next state S. t+1 The sample data (S) t A t ,R t ,S t+1 The sample pool for storing clients;

[0087] In step S508, it is determined whether there is enough new sample data in the sample pool; if yes, proceed to step S510 and simultaneously execute step S504; if no, proceed to step S504.

[0088] In step S510, local model training is performed; that is, the latest reinforcement learning model is trained based on the sample data.

[0089] In step S512, it is determined whether the local model has converged; if it has converged, proceed to step S514; if it has not converged, proceed to step S510 to continue execution.

[0090] In step S514, candidate clients for participating in the current round of federated learning are selected based on the client sampling algorithm;

[0091] In step S516, a round of global federated learning is performed based on the candidate client, and the federated learning is iterated for a preset number of rounds to obtain the latest global model;

[0092] In step S518, the latest global model is distributed to all edge servers and clients;

[0093] In step S520, it is determined whether the latest global model has converged; if yes, the entire model training process ends; if no, it returns to step S504 to continue execution.

[0094] The technical solution in this disclosure, through processes such as RL model definition, model training and client sampling, model inference and sample collection, avoids consuming large amounts of resources to transmit training data, while protecting user traffic privacy information, thus improving security and privacy.

[0095] The input to the reinforcement learning model describes the user's traffic usage characteristics across multiple time slots, including traffic usage patterns and transmission rates. The output is a recommendation for a data plan for the user, and the reward is the operator's revenue minus the costs incurred due to the recommended plan. The method in this embodiment combines reinforcement learning and federated learning models through sampling, allowing both to run simultaneously and ultimately converge. This approach enables simultaneous data collection and model training while providing recommendations to users, and then improving the model based on user feedback. It incorporates user feedback into model training, improving training efficiency and data utilization, enhancing model realism, and better reflecting real-world scenarios.

[0096] In this embodiment, considering client communication conditions, a client sampling algorithm is used to select suitable candidate clients to participate in federated learning training, thereby accelerating model training. By selecting candidate clients, the differences in communication conditions among a massive number of clients are taken into account. Therefore, candidate clients with better communication conditions can be selected to participate in federated learning, thereby reducing communication time and the time required to train the model through federated learning.

[0097] This disclosure also provides an information recommendation method, referring to... Figure 6 As shown, the main steps include:

[0098] In step S610, the user's current data usage status is obtained;

[0099] In step S620, the current traffic usage status is input into the target reinforcement learning model to obtain information recommendation suggestions for the user.

[0100] In this embodiment of the disclosure, for each user's client, the current traffic usage status can be obtained. The current traffic usage status can be any scenario's traffic usage status, and can include one or more of traffic usage information and transmission rate. Specifically, traffic usage information can include traffic usage and remaining available traffic; transmission rate can include uplink rate and downlink rate, etc.

[0101] Next, the current data usage status can be input into the target reinforcement learning model to obtain the output action. This output action can be an information recommendation suggestion for the user. The information recommendation suggestion can be the same or different for each user. The information recommendation suggestion can be a data plan recommendation, a call plan recommendation, etc., depending on the application scenario. Here, we will use a data plan recommendation suggestion as an example for explanation. The output action corresponds to one action for each plan, and the output action can be represented as follows: in, This indicates whether package i is recommended, and M represents all package types. Each value in the output can be 0 or 1. If a package's value is 1, it is recommended; if a package's value is 0, it is not recommended. It should be noted that... This indicates that no data plans are recommended. Operators can use actions output by a target reinforcement learning model to recommend appropriate data plans to users. For example, if... If the value is 1, it means that the third data plan is recommended to the client. That is, a data plan with a value of 1 in the output action can be recommended to the user.

[0102] In this embodiment of the disclosure, the current data usage status of a user can be predicted based on a trained target reinforcement learning model to accurately predict the corresponding output action and obtain a recommendation for a specific data plan, thereby determining the recommended data plan to be given to the user. Because the model incorporates user feedback information, the accuracy of the recommended data plan can be improved.

[0103] This disclosure also provides a model training apparatus. (See reference...) Figure 7 As shown, the model training device 700 may include:

[0104] The reinforcement learning module 701 is used to acquire a reinforcement learning model and train the reinforcement learning model to obtain the latest reinforcement learning model.

[0105] The fitting module 702 is used to fit the user's traffic usage status according to the latest reinforcement learning model to obtain the output action, and generate sample data based on the traffic usage status, the output action, the feedback result corresponding to the output action, and the next traffic usage status.

[0106] The federated learning module 703 is used to perform iterative federated learning on the latest reinforcement learning model using the sample data until the model converges, so as to obtain a target reinforcement learning model for determining information recommendation suggestions.

[0107] In one exemplary embodiment of this disclosure, training the reinforcement learning model and sending the trained reinforcement learning model to a cloud server to obtain the latest reinforcement learning model includes: training the reinforcement learning model based on the client and sending the model parameters of the trained reinforcement learning model to the cloud server; obtaining a global model obtained by fusing the trained reinforcement learning model through the cloud server, and training based on the global model to obtain the latest reinforcement learning model corresponding to each client.

[0108] In one exemplary embodiment of this disclosure, the step of training the reinforcement learning model based on the client includes: inputting the historical traffic usage status corresponding to the historical time slot into the reinforcement learning model to obtain the output action corresponding to the historical traffic usage status and the next traffic usage status; determining the reward for the output action corresponding to the historical traffic usage status, and adjusting the parameters of the reinforcement learning model according to the reward.

[0109] In one exemplary embodiment of this disclosure, generating sample data based on the traffic usage status, the output action, the feedback result corresponding to the output action, and the next traffic usage status includes: determining a reward based on the feedback result and cost information, and determining the next traffic usage status; storing the traffic usage status, the output action, the reward, and the next traffic usage status into the client's sample pool as sample data for the client.

[0110] In one exemplary embodiment of this disclosure, the step of using the sample data to perform iterative federated learning on the latest reinforcement learning model until the model converges to determine the target reinforcement learning model includes: if the number of sample data is greater than a preset value, determining candidate clients to participate in federated learning; performing global federated learning based on the candidate clients, and continuing to select candidate clients to participate in federated learning to obtain reselected candidate clients; performing multiple rounds of iterative global federated learning based on the reselected candidate clients to obtain the latest global model; and training the model based on the convergence of the latest global model.

[0111] In one exemplary embodiment of this disclosure, determining the candidate clients to participate in federated learning includes: randomly selecting multiple clients to participate in federated learning in a preset round of federated learning; predicting the predicted training time of each client based on historical training time, and determining the probability of each client being selected in the target round of training based on the predicted training time of each client; randomly sampling all clients based on the probability to determine the candidate clients to participate in the current round of federated learning.

[0112] In one exemplary embodiment of this disclosure, the step of training the model based on the convergence of the latest global model includes: if the latest global model converges, ending the model training and using the latest global model as the target reinforcement learning model; if the latest global model does not converge, re-obtaining the output action and generating sample data based on the user's traffic usage status until the model converges, so as to obtain the target reinforcement learning model.

[0113] This disclosure also provides an information recommendation device. (See reference) Figure 8 As shown, the information recommendation device 800 may include:

[0114] The status acquisition module 801 is used to acquire the user's current traffic usage status;

[0115] The information recommendation determination module 802 is used to input the current traffic usage status into the target reinforcement learning model to obtain information recommendation suggestions for the user.

[0116] It should be noted that the specific details of each module in the above-mentioned model training device and information recommendation device have been described in detail in the corresponding methods, so they will not be repeated here.

[0117] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0118] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0119] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.

[0120] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."

[0121] The following reference Figure 9 To describe an electronic device 900 according to such an embodiment of the present disclosure. Figure 9 The electronic device 900 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0122] like Figure 9 As shown, the electronic device 900 is manifested in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: at least one processing unit 910, at least one storage unit 920, a bus 930 connecting different system components (including storage unit 920 and processing unit 910), and a display unit 940.

[0123] The storage unit stores program code that can be executed by the processing unit 910, causing the processing unit 910 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 910 can perform actions such as... Figure 1 The steps are shown in the figure.

[0124] Storage unit 920 may include readable media in the form of volatile storage units, such as random access memory (RAM) 9201 and / or cache memory 9202, and may further include read-only memory (ROM) 9203.

[0125] Storage unit 920 may also include a program / utility 9204 having a set (at least one) program module 9205, such program module 9205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0126] Bus 930 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0127] Electronic device 900 can also communicate with one or more external devices 1000 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 900, and / or with any device that enables electronic device 900 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 950. Furthermore, electronic device 900 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 960. As shown, network adapter 960 communicates with other modules of electronic device 900 via bus 930. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0128] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or electronic device, etc.) to execute the methods according to the embodiments of this disclosure.

[0129] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible implementations, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of this disclosure described in the "Exemplary Methods" section above.

[0130] The program product for implementing the above-described method according to embodiments of the present disclosure may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0131] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0133] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0134] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0135] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0136] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention described herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not invented by this disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

Claims

1. A model training method, characterized in that, include: Obtain a reinforcement learning model, train the reinforcement learning model on the client, and send the trained reinforcement learning model to the cloud server. Obtain a global model obtained by fusing the trained reinforcement learning model through the cloud server, and train based on the global model to obtain the latest reinforcement learning model corresponding to each client. The latest reinforcement learning model is used to fit the user's data usage status to obtain the output action, and sample data is generated based on the data usage status, the output action, the feedback result corresponding to the output action, and the next data usage status. The latest reinforcement learning model is iteratively federated learning using the sample data until the model converges, so as to obtain the target reinforcement learning model used to determine information recommendation suggestions; The step of iteratively federating the latest reinforcement learning model using the sample data until the model converges to obtain the target reinforcement learning model for determining information recommendation suggestions includes: If the number of sample data is greater than a preset value, candidate clients for participating in federated learning are determined; Global federated learning is performed based on the candidate clients, and candidate clients participating in federated learning are selected again to obtain reselected candidate clients. Multi-round iterative global federated learning is implemented based on the reselected candidate clients to obtain the latest global model. Model training is performed based on the convergence of the latest global model.

2. The model training method according to claim 1, characterized in that, The training of the reinforcement learning model based on the client includes: Input the historical traffic usage status corresponding to the historical time slot into the reinforcement learning model to obtain the output action corresponding to the historical traffic usage status, as well as the next traffic usage status; Determine the reward for the output action corresponding to the historical traffic usage status, and adjust the parameters of the reinforcement learning model based on the reward.

3. The model training method according to claim 1, characterized in that, The process of generating sample data based on the traffic usage status, the output action, the feedback result corresponding to the output action, and the next traffic usage status includes: The reward is determined based on the feedback results and cost information, and the next traffic usage status is determined. The traffic usage status, output action, reward, and next traffic usage status are stored in the client's sample pool as sample data for the client.

4. The model training method according to claim 1, characterized in that, The process of identifying candidate clients to participate in federated learning includes: In the preset round of federated learning, multiple clients are randomly selected to participate in the federated learning process; The predicted training time for each client is predicted based on the historical training time, and the probability of each client being selected in the target round of training is determined based on the predicted training time of each client. Based on the stated probability, all clients are randomly sampled to determine candidate clients to participate in the current round of federated learning.

5. The model training method according to claim 1, characterized in that, The step of training the model based on the convergence of the latest global model includes: If the latest global model converges, the model training ends, and the latest global model is used as the target reinforcement learning model. If the latest global model fails to converge, the output action is re-obtained based on the user's traffic usage status, and sample data is generated until the model converges, thus obtaining the target reinforcement learning model.

6. An information recommendation method, characterized in that, include: Get the user's current data usage status; The current traffic usage status is input into the target reinforcement learning model to obtain information recommendation suggestions for the user; wherein the target reinforcement learning model is trained by the model training method according to any one of claims 1-5.

7. A model training device, characterized in that, include: The reinforcement learning module is used to obtain a reinforcement learning model, train the reinforcement learning model on the client, send the trained reinforcement learning model to the cloud server, obtain a global model obtained by fusing the trained reinforcement learning model through the cloud server, and train on the global model to obtain the latest reinforcement learning model for each client. The fitting module is used to fit the user's traffic usage status to obtain the output action based on the latest reinforcement learning model, and to generate sample data based on the traffic usage status, the output action, the feedback result corresponding to the output action, and the next traffic usage status. The federated learning module is used to iteratively federate the latest reinforcement learning model using the sample data until the model converges, so as to obtain the target reinforcement learning model used to determine information recommendation suggestions. The step of iteratively federating the latest reinforcement learning model using the sample data until the model converges to obtain the target reinforcement learning model for determining information recommendation suggestions includes: If the number of sample data is greater than a preset value, candidate clients for participating in federated learning are determined; Global federated learning is performed based on the candidate clients, and candidate clients participating in federated learning are selected again to obtain reselected candidate clients. Multi-round iterative global federated learning is implemented based on the reselected candidate clients to obtain the latest global model. Model training is performed based on the convergence of the latest global model.

8. An information recommendation device, characterized in that, include: The status acquisition module is used to obtain the user's current data usage status; The information recommendation determination module is used to input the current traffic usage status into a target reinforcement learning model to obtain information recommendation suggestions for the user; wherein the target reinforcement learning model is trained by the model training method according to any one of claims 1-5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the model training method according to any one of claims 1-5 or the information recommendation method according to claim 6.

10. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the model training method of any one of claims 1 to 5 or the information recommendation method of claim 6 by executing the executable instructions.

Citation Information

Patent Citations

  • Federal learning-based sequence recommendation method and system

    CN114595396A