Personalized Recommendation Global Model Training Method

Through the methods of reinforcement learning and federal knowledge distillation, the personalized recommendation strategy is dynamically optimized, and the data privacy, accuracy and efficiency of the personalized recommendation system are solved, and cross-platform data collaboration and efficient personalized recommendation are achieved.

CN120235213BActive Publication Date: 2025-07-29CHONGQING TELECOMM PLAN & DESIGN INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510724337.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-29
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing personalized recommendation systems have challenges in data privacy protection, recommendation accuracy and system efficiency, especially in heterogeneous data environments where global model generalization capabilities are reduced and communication costs are high.

Method used

The method of reinforcement learning and federal knowledge distillation is adopted, and personalized recommendation strategies are dynamically optimized through knowledge sharing between global teacher model and client student model, and the communication overhead is reduced using sparse technology, and recommendation strategies are adjusted through reinforcement learning to improve accuracy and efficiency.

Benefits of technology

On the premise of protecting user privacy, cross-platform data collaboration is realized, recommendation accuracy and user satisfaction are improved, communication overhead is reduced, and overall computing efficiency of the system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235213B_ABST
    Figure CN120235213B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for training a personalized recommendation global model, belonging to the field of artificial intelligence technology. The method includes: initializing personalized weights and personalized regularization parameters; determining personalized recommendation scores; determining the total loss of the current training batch; calculating the gradient of the total loss with respect to the trainable parameters in the personalized student model; updating the trainable parameters of the personalized student model; adjusting the personalized weights, personalized regularization parameters, and reward function weights; updating the reinforcement learning reward function value based on the reward function weights; maximizing the reward based on the determined optimal policy parameters; aggregating each reward, and updating the global weights and global regularization parameters. The model trained by the model training method based on reinforcement learning and federated knowledge distillation has higher recommendation accuracy, and the communication overhead is lower when training the global model. In addition, by training the global model with this method, privacy leakage can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of artificial intelligence, and particularly relates to a method for training a personalized recommendation global model. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, personalized recommendation systems have become an important tool for major Internet platforms to improve user experience and commercial value. Personalized recommendation aims to provide users with the most relevant content or products based on users' historical behaviors, interest preferences, and environmental characteristics, thereby improving user satisfaction and platform revenue. In a recommendation system, the core problem is how to improve the accuracy of recommendations and system efficiency while meeting user needs.

[0003] Traditional recommendation algorithms, such as collaborative filtering, matrix factorization, and deep learning, usually rely on centralized data storage and training, but face many challenges in data privacy protection, computational efficiency, and cross-platform collaboration. For example, data barriers between Internet enterprises make it difficult for recommendation systems to fully utilize users' cross-platform behavioral data. At the same time, centralized data storage is also prone to the risk of privacy leakage. Therefore, how to efficiently train a personalized recommendation model without data leaving the local has become an urgent problem to be solved.

[0004] In related technologies, federated learning, as a distributed machine learning method, provides an effective solution to solve the above problems. Federated learning allows multiple clients, such as different user devices or institutions, to train personalized recommendation models locally and only transmit model parameters to the server instead of the original data, thereby protecting user privacy to a certain extent. However, federated learning still faces challenges in a heterogeneous data environment. The data distributions on different user devices vary greatly, which may lead to a decline in the generalization ability of the global model. In addition, the communication cost of federated learning is high. How to reduce the transmission overhead and improve the training efficiency is also an important issue faced by researchers. Summary of the Invention

[0005] The present disclosure proposes a method for training a personalized recommendation global model to solve the problems of low recommendation accuracy of traditional recommendation models, large communication overhead, and easy leakage of data privacy when training a recommendation global model.

[0006] A method for training a personalized recommendation global model includes: initializing a global item embedding , a global teacher model , a global weight , and a global regularization parameter , where the global item embedding is sparsified to obtain a sparsified global item embedding ; through the global teacher model The predicted value is calculated , and the soft label ; The sparsified global item embedding , and the soft label , the global weight , and the global regularization parameter are sent to the client. Among them, the steps executed on the client include: using the global weight , and the global regularization parameter to be respectively assigned to the personalized weight , and the personalized regularization parameter ; According to the user embedding , the local item embedding , and the sparsified global item embedding , the personalized recommendation score is determined ; According to the personalized recommendation score , the user's true score , the soft label , the reinforcement learning reward function value , and the local hyperparameters , the total loss of the current training batch is determined ; Calculate the gradient of the total loss with respect to the trainable parameters in the personalized student model; Use the selected optimization algorithm and gradient to update the trainable parameters of the personalized student model ; Use the set reinforcement learning agent to adjust the personalized weight , the personalized regularization parameter , the reward function weight , , ; Based on the reward function weight , update the reinforcement learning reward function value ; Based on the determined optimal policy parameters , maximize ; Aggregate the uploaded by each client, and update the global weight .

[0007] In some embodiments, the soft label calculated by the global teacher model includes: According to the formula: , calculate the soft label , where represents the predicted value calculated by the global teacher model , A temperature parameter for controlling the softening degree of the knowledge distillation signal.

[0008] In some embodiments, according to the user embedding , local item embedding , sparsified global item embedding , to determine the personalized recommendation score , including: according to the formula: , calculate the personalized recommendation score .

[0009] In some embodiments, according to the personalized recommendation score , user true score , soft label , reinforcement learning reward function value , local hyperparameter , to determine the total loss of the current training batch , including: according to the formula: , determine the total loss ; where represents the reinforcement learning reward function value; , represents the cross-entropy loss of the true score; , represents the distillation loss of the soft label.

[0010] In some embodiments, calculating the gradient of the total loss with respect to the trainable parameters in the personalized student model, including: according to the formula: , determine the gradient , where represents the gradient of the hard loss with respect to the student model parameters, represents the gradient of the soft loss with respect to the student model parameters, represents the gradient of the reward with respect to the student model parameters.

[0011] In some embodiments, using the selected optimization algorithm and gradient to update the trainable parameters of the personalized student model, including: according to the formula: , update , where represents the local learning rate of the student model.

[0012] In some embodiments, using the set reinforcement learning agent to adjust the personalized weight , personalized regularization parameter , reward function weight , , , including: defining a state space as , where represents the user's behavior data, represents the user's historical click-through rate, represents the user's device type, represents the current personalized weight of the recommendation system, represents the current personalized regularization parameter of the recommendation system, represents the reward function weight; defining an action space as ; according to the current state and the probability distribution output by the policy network , adopting an action ; according to the formula: , calculating the dynamic step size ; according to the step size , updating , , , , .

[0013] In some embodiments, updating the reinforcement learning reward function value , , based on the reward function weight , including: according to the formula: , where represents the click-through rate, represents the normalized discounted cumulative gain, represents the user's stay time on the recommended content.

[0014] In some embodiments, maximizing based on the determined optimal policy parameter , including: according to the formula: , determining the gradient of ; according to the formula: , updating until the optimal policy parameter is found.

[0015] In some embodiments, aggregating the uploaded by each client and updating the global weight and the global regularization parameter , including: according to the formula: , determining the aggregated global average reward ; according to the formula: , update , where represents the learning rate of the global weight; according to the formula: , update , where represents the learning rate of the global regularization parameter.

[0016] The above steps do not clearly describe the update mechanism of the global teacher model during the federated learning iteration process. In this disclosure, it should be assumed that it is pre-trained and kept static before the start of federated learning. In this disclosure, the update of the global model is as follows: The server aggregates the predicted soft labels of the client student models for the public goods set, and then retrains the teacher model using the aggregated soft labels so that it can absorb the collective knowledge of the clients and remain up-to-date.

[0017] Global teacher model update: After the end of each federated training round or after a preset number of rounds, the server side can initiate the update process of the global teacher model so that it can learn the collective knowledge of the client models and the latest changes in data publication. This disclosure updates the global teacher model by aggregating the knowledge of the client student models .

[0018] First, client prediction result upload: Each client participating in federated learning after completing local training, uses its current student model to predict a preset public goods set (where ) or the global item sample set specified by the server. For each item in the set, the student model will calculate the corresponding prediction . The client converts these logits into soft labels by applying the same temperature parameter as the teacher model, as follows: , where is a vector representing the prediction probability distribution of the student model for the item in the set . The client uploads the set of soft labels generated by these student models to the server.

[0019] Server aggregates student model knowledge: The server collects the soft labels of the student models uploaded by all participating clients . For each item , the server aggregates the soft labels of the student models received, for example, calculates the average value to obtain the aggregated student soft labels , as follows: . This aggregation result represents the collective prediction distribution or collective knowledge of all current client models for the item .

[0020] Global teacher model retraining: The server uses as the training target to retrain the global teacher model . The retraining goal of the teacher model is to minimize the difference between its soft label output and the aggregated student soft labels . This is usually achieved by minimizing the knowledge distillation loss between them, such as cross-entropy loss or KL divergence. Taking the cross-entropy loss between the aggregated student soft labels and the soft labels output by the teacher model as an example, the objective function for updating the teacher model can be defined as: , where is the soft label generated by the global teacher model for the item , obtained by applying Softmax with a temperature of to its logits. The server uses an optimization algorithm to minimize to update the parameters of the global teacher model . The parameter update formula can be expressed as: , where represents the learning rate of the teacher model.

[0021] Update the teacher model for the next round of distillation: The model after retraining in the previous step will be used as the new global teacher model for the knowledge distillation link in the next round of federated training, that is, for calculating new soft labels to be sent to the client, and so on in a loop until the model converges.

[0022] As above, the present disclosure provides a method for training a global model for personalized recommendation based on reinforcement learning and federated knowledge distillation. This method combines the privacy protection ability of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization ability of reinforcement learning, aiming to solve the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0023] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without data leaving the local area, effectively balancing global generality and local personalization needs. At the same time, the client uses reinforcement learning to dynamically optimize the personalized recommendation strategy. For example, it dynamically adjusts the weights of personalization and generalization as well as the weights of the reward function, enabling it to adaptively adjust the recommendation behavior according to the real-time feedback of users, so as to obtain and execute a better recommendation strategy in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0024] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, this disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the policy adjustment feedback (reward signals) of the client instead of transmitting a large number of model parameters or raw data. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0026] With reference to the drawings, the present disclosure can be more clearly understood according to the following detailed description.

[0027] Figure 1 is a flowchart showing a method for training a global model for personalized recommendation according to some embodiments of the present disclosure.

[0028] Figure 2 is an architecture diagram showing the implementation of this method according to some embodiments of the present disclosure.

[0029] Figure 3 is a schematic diagram showing the accuracy of personalized recommendation implemented by applying different methods according to some embodiments of the present disclosure.

[0030] Figure 4 is a schematic diagram showing the recommendation time and communication overhead of personalized recommendation implemented by applying different methods according to some embodiments of the present disclosure.

[0031] Figure 5 is a block diagram showing an apparatus for training a global model for personalized recommendation according to some embodiments of the present disclosure.

[0032] Figure 6 is a block diagram showing an apparatus for training a global model for personalized recommendation according to other embodiments of the present disclosure.

[0033] Figure 7 is a block diagram showing a computer system for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0035] Meanwhile, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship.

[0036] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation to the present disclosure, its application, or its use.

[0037] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and devices should be regarded as part of the specification.

[0038] In all the examples shown and discussed herein, any specific values should be understood as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0039] It should be noted that: Like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, further discussion thereof is not required in subsequent drawings.

[0040] Currently, with the rapid development of artificial intelligence and big data technologies, personalized recommendation systems have become an important tool for major Internet platforms to enhance user experience and commercial value. Personalized recommendation aims to provide users with the most relevant content or products based on users' historical behaviors, interest preferences, and environmental characteristics, thereby improving user satisfaction and platform revenue. In a recommendation system, the core issue is how to improve the accuracy of recommendations and system efficiency while meeting user needs.

[0041] Traditional recommendation algorithms, such as collaborative filtering, matrix factorization, and deep learning, usually rely on centralized data storage and training, but there are many challenges in data privacy protection, computational efficiency, and cross-platform collaboration. For example, data barriers between Internet enterprises make it difficult for recommendation systems to fully utilize users' cross-platform behavioral data. At the same time, centralized data storage is also prone to the risk of privacy leakage. Therefore, how to efficiently train a personalized recommendation model without data leaving the local area has become an urgent problem to be solved.

[0042] In the related art, as a distributed machine learning method, federated learning provides an effective solution to solve the above problems. Federated learning allows multiple clients, such as different user devices or institutions, to locally train personalized recommendation models and only transmit model parameters to the server instead of the original data, thus protecting user privacy to a certain extent. However, federated learning still faces challenges in heterogeneous data environments. The data distributions on different user devices vary greatly, which may lead to a decline in the generalization ability of the global model. In addition, the communication cost of federated learning is relatively high. How to reduce the transmission overhead and improve the training efficiency is also an important issue facing researchers.

[0043] In view of this, the present disclosure proposes a method for training a global model for personalized recommendation. This method combines the privacy protection ability of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization ability of reinforcement learning, aiming to solve the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0044] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without the data leaving the local, effectively balancing global generality and local personalized needs. At the same time, the client uses reinforcement learning to dynamically optimize the personalized recommendation strategy. For example, it dynamically adjusts the weights of personalization and generalization as well as the weights of the reward function, enabling it to adaptively adjust the recommendation behavior according to the real-time feedback of users, so as to obtain and execute better recommendation strategies in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0045] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, the present disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the policy adjustment feedback (reward signals) of the clients instead of transmitting a large number of model parameters or original data.

[0046] The inventive concept of the present disclosure lies in integrating three technologies to dynamically optimize the recommendation strategy, improve the recommendation accuracy and efficiency while protecting user privacy. First, this method introduces reinforcement learning to dynamically optimize the personalized recommendation strategy of the client. Each client is regarded as an RL agent, which takes the user's historical behavior, device information, and the current state of the recommendation system, including the personalized weight and personalized regularization parameter, as the state S. The action A of the RL agent is to adjust the key policy parameters that affect the behavior of the recommendation model, such as the personalized weight and personalized regularization strength. By presenting the recommendation results to the user and collecting the user's interaction feedback, such as clicks and dwell time, the system calculates the reward, which quantifies the quality of the current recommendation strategy and user satisfaction. The RL agent aims to maximize the cumulative reward by learning an optimal strategy, that is, continuously adjusting the policy parameters according to the user feedback, so that the recommendation system can adapt to the user's real-time preferences and context changes, thereby maximizing the user's long-term satisfaction.

[0047] Secondly, this method adopts federated knowledge distillation to achieve secure and efficient cross-platform collaborative learning and knowledge sharing. The server maintains a global teacher model, which learns general and structured item knowledge and user behavior patterns through aggregated or globally visible information. In each federated round, the server sends the soft labels generated by the teacher model to the clients. The personalized recommendation model of the client, as the student model, not only learns to fit the hard target of the user's true rating (through the hard loss, that is, the loss based on the user's true rating) when training locally with its private user data. The hard target refers to the user's true rating data, which is the supervised information that the model needs to directly fit, and the hard loss is the loss calculated according to the hard target, that is, the cross-entropy loss , but also learns to fit the soft labels provided by the teacher model (through the soft loss, that is, the loss calculated based on the guidance information output by the global teacher model). The soft labels are the outputs in the form of probability distributions generated by the global teacher model, which contain the knowledge of the teacher model and are used to guide the learning of the student model. The soft loss is the distillation loss calculated according to the soft labels provided by the teacher model, that is . This way enables the client student model to benefit from the global knowledge of the teacher model and enhances the generalization ability of the model, which is particularly important for users with sparse data or new items. Since only the knowledge of the teacher model (soft labels) is shared among clients through the server, without directly exchanging the original user data or local model parameters (the core training process is completed locally), the user's privacy is effectively protected.

[0048] In this disclosure, there is a close interaction and correlation between the reinforcement learning-optimized policy and the knowledge shared by federated knowledge distillation, specifically manifested as follows: Reinforcement learning dynamically regulates the influence degree of knowledge distillation: The reinforcement learning agent directly controls the balance between the hard loss and the soft loss (derived from knowledge distillation) in the training of the client model by optimizing the personalized weights. This means that the system can intelligently decide whether to rely more on personalized local data learning or more on the knowledge of the global teacher model according to the specific situation of the user and the reward feedback obtained from reinforcement learning.

[0049] Knowledge distillation provides a generalization basis for reinforcement learning: Knowledge distillation enables the client student model to possess the basic generalization ability, which provides a stable starting point for the reinforcement learning agent to explore better personalized strategies. A student model with strong generalization ability performs more robustly when facing new items or user changes, enabling the reinforcement learning agent to learn and adjust more effectively.

[0050] The reinforcement learning feedback optimizes the goal of knowledge distillation: By aggregating the knowledge of the student model trained under the optimization of reinforcement learning on the client side, such as predicting soft labels, the server can periodically update the global teacher model. This means that the global knowledge learned by the teacher model is not static, but absorbs the effective patterns explored by the client during the personalized reinforcement learning process, forming a closed-loop iterative optimization of the federated teacher model.

[0051] By achieving a better balance between global knowledge sharing and personalized learning, while effectively reducing the overall communication cost. It mainly includes the following aspects: First, the design of a two-way personalized mechanism, which optimizes the personalized recommendation strategy by combining reinforcement learning and federated learning; Second, the optimization of the dynamic regularization learning strategy, which dynamically adjusts the regularization strategy according to the user behavior feedback to improve the stability and generalization ability of the model; Then, the design of global item sparsification, which optimizes the embedding representation of global items through sparsification technology to improve the accuracy and efficiency of the recommendation effect; Finally, the design of the privacy protection mechanism, which realizes cross-platform data collaboration and knowledge sharing using federated knowledge distillation on the premise of ensuring data privacy. These designs work together to promote the improvement of the efficiency, accuracy, and privacy protection ability of the personalized recommendation system.

[0052] Such as Figure 2As shown in the figure, this architecture mainly includes a server and each client. The server is responsible for aggregating and updating the reward function values uploaded from each client, and saving the embedded features of the global project and the shared knowledge. The client is responsible for training and updating the personalized recommendation model for the user data stored locally, and optimizing the recommendation strategy through reinforcement learning. The client and the server work together through the way of federated learning to ensure the protection of data privacy, and share the global knowledge through the knowledge distillation technology, so as to improve the accuracy and efficiency of the recommendation system.

[0053] First, for the initial users and items in the recommendation system, definitions are made. Among them, is the globally shared information, that is, the item embeddings that all users can use; represents the -th user's local item embedding; the local item embedding can be the embedding of the items interacted locally by the user, or a sparse matrix or low-rank matrix corresponding to the global item embedding and reflecting the personalized differences. Its specific implementation aims to balance the personalized representation ability and the local storage and calculation overhead. is the rating matrix of all users for items. Among them, represents the number of users, represents the number of items, represents the -th user's rating. Assuming that each client only contains the information of one user, therefore, also represents the rating of the client. Assuming , it means that the user has rated the item, and is further used to mark the rated items in the rating matrix . The main role of is to define the training sample set with the true ratings of users used when calculating the hard loss

[0054] Figure 1 is a flowchart showing the method for training the global model of personalized recommendation according to some embodiments of the present disclosure. As Figure 1 shown, the method for training the global model of personalized recommendation includes steps S110 to S140.

[0055] In step S110, initialize the global item embedding , the global teacher model , the global weight , the global regularization parameter , among which, the global item embedding is sparsified to obtain the sparsified global item embedding .

[0056] Server initialization, including global item embedding , global teacher model , global weights , global regularization parameters .

[0057] While the server is initializing, each client also needs to initialize, including user embedding and local item embedding , set the reinforcement learning agent (RL-Agent), and initialize local policy parameters , , , , , local parameters , are initialized to the global parameters sent by the server , .

[0058] After initializing the global item embedding , sparsify it. For example, based on item popularity or preset rules, some dimensions in or the embeddings of specific items can be set to zero or the number of non-zero elements can be restricted, so as to obtain the sparsified global item embedding . This sparsification process aims to reduce the storage and transmission overhead of the global item embedding and may improve the efficiency and generalization ability of the model by introducing inductive bias. In subsequent steps, the client will use this sparsified global item embedding .

[0059] In step S120, the predicted value is calculated through the global teacher model , soft label .

[0060] The soft label is calculated through the global teacher model , including: According to the formula: , the soft label is calculated, where represents the predicted value calculated by the global teacher model , represents the temperature parameter used to control the softening degree of the knowledge distillation signal. The temperature parameter is usually a hyperparameter greater than 0, and its specific value can be set according to experience or determined by experimental tuning. A larger value will produce a smoother probability distribution, and a smaller Values result in a smoother probability distribution, while smaller values (close to 0) make the probability distribution sharper, approaching the one-hot form.

[0061] The temperature parameter is a positive number used to control the soft labels generated by the global teacher model during the knowledge distillation process. Generally, as a hyperparameter, it can be determined by cross-validation, grid search, or experimental tuning on the validation set within a preset range to obtain the optimal personalized student model training results. Larger values make the probability distribution of the soft labels smoother, helping the student model learn the potential relationships between classes; smaller values (when = 1, it is the standard Softmax) make the probability distribution closer to the original prediction confidence of the teacher model. In this disclosure, selecting an appropriate value helps balance the degree of the student model's absorption of the teacher model's knowledge and its ability to fit hard targets.

[0062] is the predicted value calculated by the teacher model; is the temperature parameter used to control the softening degree of the knowledge distillation signal. When it is larger, the distribution is smoother.

[0063] The server sends the sparsified global item embedding and the soft labels to the client to guide personalized training;

[0064] The server sends 、 as the initial values for the client's reinforcement learning.

[0065] In step S130, the sparsified global item embedding 、the soft labels 、the global weight 、the global regularization parameter are sent to the client. Among them, the steps executed on the client include: assigning the global weight 、the global regularization parameter to the personalized weight 、the personalized regularization parameter respectively; determining the personalized recommendation score according to the user embedding 、the local item embedding 、the sparsified global item embedding ; Determine the total loss of the current training batch based on personalized recommendation scores , user true scores , soft labels , reinforcement learning reward function values , local hyperparameters , and determine the total loss of the current training batch ; Calculate the total loss with respect to the trainable parameters in the personalized student model ; Use the selected optimization algorithm and the gradient to update the trainable parameters of the personalized student model ; Use the set reinforcement learning agent to adjust the personalized weights , personalized regularization parameters , reward function weights , , ; Based on the reward function weights , , , update the reinforcement learning reward function values ; Based on the determined optimal policy parameters , maximize ;

[0066] Initialize the local reward function weights , , For example, allocate the weights evenly.

[0067] Calculate the personalized recommendation scores as follows: .

[0068] Locally train the personalized student model to optimize the objective function ; On the client side, locally use the received sparsified global item embeddings , teacher model soft labels , client-side local data (including user embeddings , local item embeddings , user true scores ), the current local personalized weights and personalized regularization parameters to train the personalized student model. The main trainable parameters of the personalized student model include user embeddings, local item embeddings, and other possible parameters in the model, such as including additional neural network layers. The training process is achieved by minimizing the following objective function : ; where represents the cross-entropy loss of the true scores, and its calculation formula is: ; Denotes the distillation loss of the teacher model's soft labels, and its calculation formula is: . is the reward function value obtained by the client through reinforcement learning. The specific training and update steps are as follows:

[0069] Calculate the current loss . According to the formula , use the calculated predicted score , the user's true score , the teacher model's soft label , and the local hyperparameters , , and the client reward (function value) to calculate the total loss in the current training batch or the current state .

[0070] Calculate the model parameter gradients. Calculate the gradient of the total loss function with respect to all trainable parameters in the personalized student model. This gradient is a vector composed of the partial derivatives of the loss function with respect to each parameter. Specifically, the gradient is expressed as: , where represents the gradient of the hard loss with respect to the student model parameters, represents the gradient of the soft loss with respect to the student model parameters, represents the gradient of the reward with respect to the student model parameters. This term is usually related to the policy gradient or related techniques in reinforcement learning. These gradients can be calculated through the backpropagation algorithm.

[0071] Update the model parameters. Use a selected optimization algorithm, such as stochastic gradient descent, Adam, etc., and the calculated gradients to update each trainable parameter of the personalized student model. The parameter update formula is expressed as: , where represents the local learning rate of the student model.

[0072] Local iterative training. On the user dataset stored locally on the client, repeat the above steps for several training epochs or multiple mini-batch iterations. In this local training stage, the personalized weights and the personalized regularization parameter are usually fixed to the values at the start of this federated learning or after the most recent adjustment by the RL agent in this round. They affect the composition and optimization direction of the loss function, but their updates are completed through the reinforcement learning policy.

[0073] The client uses reinforcement learning to optimize the recommendation strategy (including dynamic parameter updates). The reinforcement learning agent (RL-Agent) of the client is responsible for dynamically adjusting the local personalized weights , personalized regularization parameters , and reward function weights , , to maximize the long-term recommendation reward.

[0074] Construct the state space. The state space S is composed of the user's historical interaction data and the feedback information of the recommendation system, mainly including: the user's behavior data , that is, the item ID sequence of the user's last interactions; the user's historical click-through rate , the user's device type , the current local personalized weights of the recommendation system , regularization parameters , and reward function weights . The final state space is defined as follows: .

[0075] Construct the action space. The action space is mainly the change amount of the policy parameters that the recommendation system can adjust. The RL agent selects an action to adjust the local personalized weights , personalized regularization parameters and reward function weights . The final formed action space is as follows: .

[0076] Sampling action execution and parameter update. The RL agent samples an action according to the current state S (including user behavior data, device type, current personalized weights, regularization parameters, and reward function weights) and the probability distribution output by the policy network (referring to a policy network with parameters that outputs the probability distribution on the action space ). The client calculates the dynamic step size based on the obtained reward , for example . Then, use this step size to scale the sampled action and update the local hyperparameters . Then, use this step size to scale the sampled action and update the local hyperparameters , and reward function weights : ; ; ; .

[0077] Design the reward function. The reward function is mainly used to measure the recommendation quality and is defined as follows: , represents the click-through rate, represents the normalized discounted cumulative gain, represents the user's stay time on the recommended content, , , are the weight coefficients of the reward function dynamically adjusted locally on the client side and are used to control the contribution of different metrics to the reward.

[0078] Calculate the policy gradient. By adopting the policy gradient method, find the optimal policy parameters , to maximize the expected reward. Calculate the policy gradient as follows: . According to the formula , update , until the optimal policy parameters are found, where at the initial stage of training a larger value, such as 0.1, is adopted and gradually decays later.

[0079] The client uploads the reward. The client uploads the calculated reward function value (reward ) to the server side while protecting user privacy. The weight of the reward function dynamically adjusted locally on the client side , , is not uploaded to the server.

[0080] In step S140, aggregate the rewards uploaded by each client and update the global weights , the global regularization parameter .

[0081] The server side aggregates the reward function values uploaded by the clients and updates the global personalized weights and global regularization parameters.

[0082] The server side collects the rewards uploaded by all active clients participating in the federated learning , and calculates the aggregated global average reward : , represents the total number of active clients participating in the current round of federated learning.

[0083] Update the global weights. Use the aggregated global average reward to update the global weights , that is: , where Represents the learning rate of the global weight, with an initial value of 0.01 to avoid oscillations caused by overly large update amplitudes.

[0084] Update the global regularization parameter. Use the aggregated global average reward to update the global regularization parameter , Represents the learning rate of the global regularization parameter. and have the same value. Use to update The purpose is that when the overall recommendation effect is good (aggregated reward is high), is lower, resulting in the regularization parameter decreasing, allowing the global model to fit the data more precisely; when the recommendation effect is poor (aggregated reward is low), is higher, resulting in the regularization parameter increasing, enhancing the generalization ability or diversity of the model to cope with the overall performance decline.

[0085] The server sends the updated global parameters. The server-side sends the updated global weight and global regularization parameter to the client as the initial value for the next round of reinforcement learning.

[0086] Enter the next round of training until the model converges.

[0087] The above steps do not clearly describe the update mechanism of the global teacher model during the federated learning iteration process and should be assumed to be pre-trained and remain static before the start of federated learning in this disclosure.

[0088] In this disclosure, the update of the global model can be further discussed.

[0089] Global teacher model update. At the end of each federated training round or after a preset number of rounds, the server-side can initiate the update process of the global teacher model so that it can learn the collective knowledge of the client models and the latest changes in the data distribution. This disclosure uses the method of aggregating the knowledge of the client student models to update the global teacher model .

[0090] Client prediction result upload. Each client participating in federated learning after completing its local training, uses its current student model to predict a preset set of public items or a global set of item samples specified by the server for prediction. For each item in the set, the student model calculates the corresponding prediction . The student model calculates the corresponding prediction logits, which are the raw prediction values output by the personalized student model for the item before passing through the final activation function (e.g., the Sigmoid function used to calculate the recommendation score ). The client converts these logits into soft labels by applying the same temperature parameter as the teacher model as follows: , where, andis a vector representing the predicted probability distribution of the item by the student model over the set. The client uploads this set of soft labels generated by the student model to the server side.

[0091] The server aggregates the knowledge of the student models. The server aggregates all the soft labels of the student models uploaded by the clients. For each item in the set, the server aggregates the received soft labels of the student models, e.g., by calculating the average, to obtain the aggregated student soft label as follows: , which represents the collective prediction distribution or collective knowledge of all current client models for the item.

[0092] Retraining the global teacher model. The server uses the obtained aggregated student labels as the training target to retrain the global teacher model . The retraining objective of the teacher model is to minimize the difference between its soft label output and the aggregated student soft label . This is typically achieved by minimizing the knowledge distillation loss between them, such as cross-entropy loss or KL divergence. Taking the cross-entropy loss between the aggregated student soft label and the soft label output by the teacher model as an example, the objective function for updating the teacher model can be defined as: , where, represents the soft label generated by the global teacher model for the item obtained by applying Softmax with temperature to its logits. The server uses an optimization algorithm, e.g., gradient descent, to minimize and updates the parameters of the global teacher model . The parameter update formula can be expressed as , where, is the trainable parameter of the global teacher model . Represents the learning rate of the teacher model, is the target loss function for retraining the teacher model defined above, is the target loss function with respect to the parameter gradient.

[0093] Update the teacher model for the next round of distillation. The model after retraining will be used as the new global teacher model for the knowledge distillation session in the next round of federated training, and so on in a loop until the global model converges.

[0094] In summary, the present disclosure provides a method for training a global model for personalized recommendation based on reinforcement learning and federated knowledge distillation. This method integrates the privacy protection ability of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization ability of reinforcement learning, aiming to address the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0095] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without data leaving the local, effectively balancing global generality and local personalization needs. At the same time, the client uses reinforcement learning to dynamically optimize the personalized recommendation strategy. For example, dynamically adjusting the weights of personalization and generalization as well as the reward function weight, enabling it to adaptively adjust the recommendation behavior according to the user's real-time feedback, so as to obtain and execute a better recommendation strategy in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0096] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, the present disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the policy adjustment feedback (reward signals) of the client instead of transmitting a large number of model parameters or raw data.

[0097] The present disclosure achieves high precision, high efficiency, and strong adaptability of the personalized recommendation system (globally) on the premise of ensuring the privacy and security of user data.

[0098] As Figure 3 and Figure 4 shown, in order to verify the recommendation accuracy, recommendation time, and communication overhead of the method proposed in the present disclosure under different datasets, the following experiments were conducted in this solution.

[0099] This solution experimentally validates the effectiveness of the proposed global personalized recommendation model based on reinforcement learning and federated knowledge distillation. The experiment selected multiple public recommendation system datasets, such as MovieLens, Amazon, and Last.fm, covering different user interaction scenarios, and divided the data into training and test sets.

[0100] MovieLens is a dataset for movie recommendation systems maintained by the GroupLens research team at the University of Minnesota. It contains user ratings, movie metadata, user attributes, and other information. The Amazon dataset primarily includes user consumption behavior data from the e-commerce platform, covering product reviews, purchase history, and user attributes. Last.fm is a dataset from a music streaming platform that records users' music playback behavior and social interactions.

[0101] The comparison methods include traditional collaborative filtering (CF) recommendation methods, deep learning (DNN)-based recommendation methods, recommendation methods using only federated learning (FL), recommendation methods using only reinforcement learning (RL), and the proposed recommendation method based on reinforcement learning and federated knowledge distillation (RL-FedKD).

[0102] CF: Collaborative Filtering, corresponding Chinese: collaborative filtering; DNN: Deep Neural Detwork, corresponding Chinese: deep neural network; FL: Federated Learning, corresponding Chinese: federated learning; RL: Reinforcement Learning, corresponding Chinese: reinforcement learning; RL-FedKD: Reinforcement Learning-Federated Knowledge Distillation, corresponding Chinese: reinforcement learning combined with federated knowledge distillation.

[0103] The experiment uses indicators such as recommendation accuracy, recommendation time (computational time of different methods) and communication overhead (data transmission volume between the client and the server) to evaluate the performance of different recommendation methods.

[0104] The results show that the (personalized) recommendation method based on reinforcement learning and federated knowledge distillation performs excellently in multiple aspects. First, in terms of recommendation accuracy, RL-FedKD achieves the highest precision rates on the MovieLens, Amazon, and Last.fm datasets respectively, outperforming traditional CF and DNN, with an improvement range of 20% - 30%, indicating that the method of the present disclosure can capture user interests more accurately and improve the personalized recommendation effect. Second, in terms of recommendation time, the computational time consumption of the method of the present disclosure is approximately 0.95 seconds, which is 21% less than that of DNN and 9.5% less than that of RL, indicating that while optimizing the recommendation accuracy, the response speed of the recommendation is improved. Finally, in terms of communication overhead, the method of the present disclosure is approximately 52MB, which is 35% less than FL and 25% less than RL, effectively reducing the cost of cross-platform data exchange.

[0105] As described above, while ensuring data privacy and improving the recommendation effect, RL-FedKD reduces the consumption of communication bandwidth and optimizes the operation efficiency of the global model.

[0106] In the solution of the present disclosure, the recommendation strategy is optimized through reinforcement learning, enabling the system to dynamically adjust the recommended content, thereby improving the recommendation accuracy and reducing the search space for ineffective recommendations. The performance of the recommendation model is tested in different dataset scenarios, providing strong data support for the optimization of subsequent recommendation algorithms. The experimental results show that the method of the present disclosure is superior to traditional recommendation methods in terms of recommendation accuracy, recommendation time, and communication overhead, and can improve the overall performance of the global model (personalized recommendation global model) while ensuring data privacy.

[0107] Figure 5 is a block diagram showing a personalized recommendation global model training device according to some embodiments of the present disclosure. As Figure 5 shown, the personalized recommendation global model 500 includes an initialization module 510, a calculation module 520, a reinforcement learning reward function value update module 530, and a weight and regularization parameter update module 540.

[0108] The initialization module 510 is configured to initialize the global item embedding , the global teacher model , the global weight , and the global regularization parameter . Among them, the global item embedding is sparsified to obtain the sparsified global item embedding ;

[0109] The calculation module 520 is configured to calculate the predicted value and the soft label through the global teacher model ;

[0110] The reinforcement learning reward function value update module 530 is configured to send the sparsified global item embedding , soft labels , global weights , global regularization parameters to the client. Among them, the steps executed on the client include: assigning the global weights , global regularization parameters to the personalized weights , personalized regularization parameters respectively; determining the personalized recommendation score according to the user embedding , local item embedding , sparsified global item embedding ; determining the total loss of the current training batch according to the personalized recommendation score , user true score , soft labels , reinforcement learning reward function value , local hyperparameters ; calculating the gradient of the total loss with respect to the trainable parameters in the personalized student model; updating the trainable parameters of the personalized student model using the selected optimization algorithm and gradient; adjusting the personalized weights , personalized regularization parameters , reward function weights , , using the set reinforcement learning agent; updating the reinforcement learning reward function value , , based on the reward function weights ; maximizing based on the determined optimal policy parameters ;

[0111] The weight and regularization parameter update module 540 is configured to aggregate the uploaded by each client and update the global weights , global regularization parameters .

[0112] In the device of the present disclosure embodiment, the privacy protection ability of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization ability of reinforcement learning are integrated to address the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0113] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without the data leaving the local device, effectively balancing global generality and local personalization requirements. At the same time, the client uses reinforcement learning to dynamically optimize the personalized recommendation strategy. For example, it dynamically adjusts the weights of personalization and generalization as well as the weights of the reward function, enabling it to adaptively adjust the recommendation behavior according to the user's real-time feedback. Thus, it can obtain and execute a better recommendation strategy in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0114] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, this disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the client's policy adjustment feedback (reward signals) instead of transmitting a large number of model parameters or raw data.

[0115] Figure 6 is a block diagram showing a personalized recommendation global model training device according to some other embodiments of the present disclosure. As Figure 6 shown, the personalized recommendation global model 600 includes a memory 610; and a processor coupled to the memory 610. The memory 610 is used to store instructions for executing the corresponding embodiments of the personalized recommendation global model training method. The processor 620 is configured to execute the personalized recommendation global model training method in any of the embodiments of the present disclosure based on the instructions stored in the memory 610.

[0116] Figure 7 is a block diagram showing a computer system for implementing some embodiments of the present disclosure. As Figure 7 shown, the computer system 700 may be embodied in the form of a general-purpose computing device. The computer system 700 includes a memory 710, a processor 720, and a bus 730 connecting different system components.

[0117] The memory 710 may include, for example, a system memory and a non-volatile storage medium. The system memory stores, for example, an operating system, application programs, a boot loader (BootLoader), and other programs. The system memory may include a volatile storage medium, such as random access memory (RAM) and / or a cache memory. The non-volatile storage medium stores, for example, instructions for executing at least one of the corresponding embodiments of the personalized recommendation global model training method. The non-volatile storage medium includes, but is not limited to, disk memories, optical memories, flash memories, etc.

[0118] The processor 720 can be implemented in the form of a general - purpose processor, a digital signal processor (DSP), an application - specific integrated circuit (ASIC), a field - programmable gate array (FPGA), or other programmable logic devices, discrete hardware components such as discrete gates or transistors. Correspondingly, each of the modules such as the initialization module, the calculation module, the sending module, the aggregation module, and the update module can be implemented by a central processing unit (CPU) running instructions that execute the corresponding steps in the memory, or can be implemented by dedicated circuitry that executes the corresponding steps.

[0119] The bus 730 can use any of a variety of bus structures. For example, the bus structure includes, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, and a Peripheral Component Interconnect (PCI) bus.

[0120] The computer system 700 may also include an input / output interface 740, a network interface 750, a storage interface 760, etc. These interfaces 740, 750, 760 can be connected to the memory 710 and the processor 720 via the bus 730. The input / output interface 740 can provide a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 750 provides a connection interface for various networking devices. The storage interface 760 provides a connection interface for external storage devices such as a floppy disk, a USB flash drive, and an SD card.

[0121] Here, various aspects of the present disclosure are described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments of the present disclosure. It should be understood that each box of the flowcharts and / or block diagrams, and combinations of the boxes, can be implemented by computer - readable program instructions.

[0122] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable devices to produce a machine, such that the device implemented by executing the instructions by the processor performs the functions specified in one or more boxes of the flowcharts and / or block diagrams.

[0123] These computer - readable program instructions can also be stored in a computer - readable memory, and these instructions cause the computer to work in a specific manner, thereby producing a manufactured article including instructions for implementing the functions specified in one or more boxes of the flowcharts and / or block diagrams.

[0124] The present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0125] The present disclosure provides a method for training a global model for personalized recommendation based on reinforcement learning and federated knowledge distillation. This method combines the privacy protection capabilities of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization capabilities of reinforcement learning, aiming to address the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0126] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without the data leaving the local, effectively balancing global generality and local personalized needs. At the same time, the client dynamically optimizes the personalized recommendation policy using reinforcement learning. For example, it dynamically adjusts the weights of personalization and generalization as well as the weights of the reward function, enabling it to adaptively adjust the recommendation behavior according to the user's real-time feedback, thereby obtaining and executing a better recommendation policy in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0127] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, the present disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the policy adjustment feedback (reward signals) of the clients instead of transmitting a large number of model parameters or raw data.

[0128] So far, the method for training a global model for personalized recommendation according to the present disclosure has been described in detail. To avoid obscuring the concept of the present disclosure, some details well known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

[0129] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration purposes and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A global model training method for personalized recommendation, characterized in that The method includes: Initialize global item embeddings , global teacher model , global weights , global regularization parameter , where the global item embeddings are sparsified to obtain sparsified global item embeddings ; Through the global teacher model The predicted value is calculated and the soft label ; Send the sparsified global item embedding , soft labels , global weights , global regularization parameters to the client. The steps performed on the client include: Assign the global weights , global regularization parameters to the personalized weights , personalized regularization parameters respectively; Determine the personalized recommendation score based on the user embedding , local item embedding , sparsified global item embedding ; Determine the total loss of the current training batch based on the personalized recommendation score , user true score , soft labels , reinforcement learning reward function value , local hyperparameters ; Calculate the gradient of the total loss with respect to the trainable parameters in the personalized student model; Update the trainable parameters of the personalized student model using the selected optimization algorithm and gradient; Adjust the personalized weights , personalized regularization parameters , reward function weights , , using the set reinforcement learning agent; Update the reinforcement learning reward function value , , based on the reward function weights ; Maximize based on the determined optimal policy parameters ; Aggregate the data uploaded by each client and update the global weights and the global regularization parameter .

2. The personalized recommendation global model training method according to claim 1, wherein The soft labels calculated by the global teacher model include: ​ According to the formula: , the soft label is calculated, where represents the predicted value calculated by the global teacher model , and represents the temperature parameter used to control the softening degree of the knowledge distillation signal.

3. The personalized recommendation global model training method according to claim 2, characterized in that The said user-embedded , locally item-embedded , sparsified globally item-embedded , to determine the personalized recommendation score , including: According to the formula: , the personalized recommendation score is calculated as .

4. The personalized recommendation global model training method according to claim 3, wherein According to the personalized recommendation score , the user's true score , soft labels , the value of the reinforcement learning reward function , local hyperparameters , determine the total loss of the current training batch , including: According to the formula: , determine the total loss ; where represents the value of the reinforcement learning reward function; , represents the cross-entropy loss of the true score; , represents the distillation loss of the soft label.

5. The personalized recommendation global model training method according to claim 4, characterized in that The total loss calculated Regarding the trainable parameters in the personalized student model The gradients of which include: According to the formula: , determine the gradient , where represents the gradient of the hard loss with respect to the student model parameters, represents the gradient of the soft loss with respect to the student model parameters, represents the gradient of the reward with respect to the student model parameters.

6. The personalized recommendation global model training method according to claim 5, wherein Updating the trainable parameters of the personalized student model by using the selected optimization algorithm and gradient , including: According to the formula: , update , where represents the local learning rate of the student model.

7. The personalized recommendation global model training method according to claim 6, wherein Adjusting personalized weights using a set reinforcement learning agent , personalized regularization parameters , reward function weights , , , including: Define the state space be , where represents the user's behavior data, represents the user's historical click-through rate, represents the user's device type, represents the current personalized weight of the recommendation system, represents the current personalized regularization parameter of the recommendation system, represents the reward function weight; Define the action space For ; According to the current state and the policy network output the probability distribution and adopt an action ; According to the formula: , calculate the dynamic step size ; According to the step size , update , , , , .

8. The personalized recommendation global model training method according to claim 7, wherein The reward function weight-based , , , update the reinforcement learning reward function value , including: According to the formula: , where represents the click-through rate, represents the normalized discounted cumulative gain, represents the user's stay time on the recommended content.

9. The personalized recommendation global model training method according to claim 8, wherein Based on the determined optimal policy parameters , maximize , including: According to the formula: , determine gradient of ; According to the formula: , update until the optimal policy parameters are found.

10. The personalized recommendation global model training method according to claim 1, characterized in that Aggregate the uploaded by each client and update the global weight and the global regularization parameter , including: According to the formula: , determine the globally averaged reward after aggregation ; According to the formula: , update , where represents the learning rate of the global weight; According to the formula: , update , where represents the learning rate of the global regularization parameter.

Citation Information

Patent Citations

  • Content recommendation model training method, content recommendation method and related equipment

    CN113360777A

  • Object recommendation method and device and storage medium

    CN116992162A