Personalized recommendation global model training method

By introducing reinforcement learning and federal knowledge distillation technologies into personalized recommendation systems, the challenges of traditional recommendation models in data privacy, computing efficiency and cross-platform collaboration are solved, efficient and accurate personalized recommendations are achieved, and communication overhead is reduced.

CN120235213AActive Publication Date: 2025-07-01CHONGQING TELECOMM PLAN & DESIGN INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510724337.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Traditional recommendation models have challenges in data privacy protection, computing efficiency and cross-platform collaboration, especially in heterogeneous data environments, where the generalization ability of the global model has decreased and the communication cost is high.

Method used

A personalized recommendation global model training method is proposed, combining reinforcement learning and federal knowledge distillation, and knowledge sharing is shared by sparse global item embedding and soft labels, and personalized recommendation strategies are dynamically optimized to reduce communication overhead.

Benefits of technology

Without revealing user privacy, we will improve recommendation accuracy and system efficiency, significantly reduce communication overhead, and enhance the generalization ability of the model and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235213A_ABST
    Figure CN120235213A_ABST
Patent Text Reader

Abstract

The invention relates to a personalized recommendation global model training method, and belongs to the technical field of artificial intelligence. The method comprises the following steps: initializing a personalized weight and a personalized regularization parameter; determining a personalized recommendation score; determining the total loss of the current training batch; calculating the gradient of the total loss about trainable parameters in the personalized student model; updating trainable parameters of the personalized student model; adjusting a personalized weight, a personalized regularization parameter and a reward function weight; updating a reinforcement learning reward function value based on the reward function weight; based on the determined optimal strategy parameter, maximizing the reward; and aggregating each reward, and updating global weight and global regularization parameters. The model obtained through training of the model training method based on reinforcement learning and federal knowledge distillation is higher in recommendation precision, communication overhead is low when the global model is trained, and in addition, privacy leakage can be avoided when the global model is trained through the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the field of artificial intelligence technology, and particularly relates to a method for training a personalized recommendation global model. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, personalized recommendation systems have become an important tool for major Internet platforms to enhance user experience and commercial value. Personalized recommendation aims to provide users with the most relevant content or products based on users' historical behaviors, interest preferences, and environmental characteristics, thereby improving user satisfaction and platform revenue. In a recommendation system, the core problem is how to improve the accuracy of recommendations and system efficiency while meeting user needs.

[0003] Traditional recommendation algorithms, such as collaborative filtering, matrix factorization, and deep learning, usually rely on centralized data storage and training, but face many challenges in data privacy protection, computational efficiency, and cross-platform collaboration. For example, data barriers between Internet enterprises make it difficult for recommendation systems to fully utilize users' cross-platform behavioral data. At the same time, centralized data storage is also prone to the risk of privacy leakage. Therefore, how to efficiently train a personalized recommendation model without data leaving the local area has become an urgent problem to be solved.

[0004] In related technologies, federated learning, as a distributed machine learning method, provides an effective solution to solve the above problems. Federated learning allows multiple clients, such as different user devices or institutions, to locally train a personalized recommendation model and only transmit model parameters to the server instead of the original data, thereby protecting user privacy to a certain extent. However, federated learning still faces challenges in a heterogeneous data environment. The data distributions on different user devices vary greatly, which may lead to a decline in the generalization ability of the global model. In addition, the communication cost of federated learning is relatively high. How to reduce the transmission overhead and improve the training efficiency is also an important issue facing researchers. Summary of the Invention

[0005] The present disclosure proposes a method for training a personalized recommendation global model to solve the problems of low recommendation accuracy of traditional recommendation models, large communication overhead, and easy leakage of data privacy when training a recommendation global model.

[0006] A method for training a personalized recommendation global model includes: initializing a global item embedding , a global teacher model , a global weight , and a global regularization parameter , where the global item embedding is sparsified to obtain a sparsified global item embedding ; through the global teacher model The predicted value is calculated , and the soft label ; The sparsified global item embedding , and the soft label , the global weight , and the global regularization parameter are sent to the client. Among them, the steps executed on the client include: using the global weight , and the global regularization parameter to be assigned to the personalized weight , and the personalized regularization parameter respectively; According to the user embedding , the local item embedding , and the sparsified global item embedding , the personalized recommendation score is determined; According to the personalized recommendation score , the user's true score , the soft label , the reinforcement learning reward function value , and the local hyperparameters , the total loss of the current training batch is determined; Calculate the total loss with respect to the trainable parameters in the personalized student model; Use the selected optimization algorithm and gradient to update the trainable parameters of the personalized student model; Use the set reinforcement learning agent to adjust the personalized weight , the personalized regularization parameter , the reward function weight , , ; Based on the reward function weight , , , update the reinforcement learning reward function value ; Based on the determined optimal policy parameters , maximize ; Aggregate the uploaded by each client, and update the global weight , and the global regularization parameter .

[0007] In some embodiments, the soft label calculated by the global teacher model includes: According to the formula: , calculate the soft label , where represents the predicted value calculated by the global teacher model , A temperature parameter for controlling the softening degree of the knowledge distillation signal.

[0008] In some embodiments, according to the user embedding , local item embedding , sparsified global item embedding , to determine the personalized recommendation score , including: according to the formula: , calculate the personalized recommendation score .

[0009] In some embodiments, according to the personalized recommendation score , user true score , soft label , reinforcement learning reward function value , local hyperparameter , to determine the total loss of the current training batch , including: according to the formula: , determine the total loss ; where represents the reinforcement learning reward function value; , represents the cross-entropy loss of the true score; , represents the distillation loss of the soft label.

[0010] In some embodiments, calculating the gradient of the total loss with respect to the trainable parameters in the personalized student model, including: according to the formula: , determine the gradient , where represents the gradient of the hard loss with respect to the student model parameters, represents the gradient of the soft loss with respect to the student model parameters, represents the gradient of the reward with respect to the student model parameters.

[0011] In some embodiments, using the selected optimization algorithm and gradient to update the trainable parameters of the personalized student model, including: according to the formula: , update , where represents the local learning rate of the student model.

[0012] In some embodiments, using the set reinforcement learning agent to adjust the personalized weight , personalized regularization parameter , reward function weight , , , including: defining a state space as , where represents the user's behavior data, represents the user's historical click-through rate, represents the user's device type, represents the current personalized weight of the recommendation system, represents the current personalized regularization parameter of the recommendation system, represents the reward function weight; defining an action space as ; according to the current state and the probability distribution output by the policy network , taking an action ; according to the formula: , calculating the dynamic step size ; according to the step size , updating , , , , .

[0013] In some embodiments, updating the reinforcement learning reward function value , , based on the reward function weight , including: according to the formula: , where represents the click-through rate, represents the normalized discounted cumulative gain, represents the user's stay time on the recommended content.

[0014] In some embodiments, maximizing based on the determined optimal policy parameters , including: according to the formula: , determining the gradient of ; according to the formula: , updating until the optimal policy parameters are found.

[0015] In some embodiments, aggregating the uploaded by each client and updating the global weight , the global regularization parameter , including: according to the formula: , determining the aggregated global average reward ; according to the formula: , update , where represents the learning rate of the global weight; according to the formula: , update , where represents the learning rate of the global regularization parameter.

[0016] The above steps do not clearly describe the update mechanism of the global teacher model during the federated learning iteration process. In this disclosure, it should be assumed that it is pre-trained and remains static before the start of federated learning. In this disclosure, the update of the global model is as follows: The server aggregates the predicted soft labels of the client student models for the public goods set, and then retrains the teacher model using the aggregated soft labels so that it can absorb the collective knowledge of the clients and remain up-to-date.

[0017] Global teacher model update: After the end of each round of federated training or after a preset number of rounds, the server side can initiate the update process of the global teacher model so that it can learn the collective knowledge of the client models and the latest changes in data publication. This disclosure updates the global teacher model by aggregating the knowledge of the client student models .

[0018] First, client prediction result upload: Each client participating in federated learning after completing local training, uses its current student model , to predict a preset public goods set (where ) or the global item sample set specified by the server. For each item in the set , the student model will calculate the corresponding prediction . The client converts these logits into soft labels by applying the same temperature parameter as the teacher model, as follows: , where is a vector representing the prediction probability distribution of the student model for the item in the set . The client uploads the set of soft labels generated by these student models to the server.

[0019] Server aggregates student model knowledge: The server collects the soft labels of the student models uploaded by all participating clients . For each item in the set ​The server aggregates the soft labels of the student models received, for example, by calculating the average value, to obtain the aggregated student soft labels , as follows: . This aggregation result represents the collective prediction distribution or collective knowledge of all current client models for the item .

[0020] Global teacher model retraining: The server uses as the training objective to retrain the global teacher model . The retraining objective of the teacher model is to minimize the difference between its soft label output and the aggregated student soft labels . This is usually achieved by minimizing the knowledge distillation loss between them, such as cross-entropy loss or KL divergence, to minimize the cross-entropy loss between the aggregated student soft labels and the soft labels output by the teacher model . Taking the cross-entropy loss between the aggregated student soft labels and the soft labels output by the teacher model as an example, the objective function for updating the teacher model can be defined as: , where is the soft label generated by the global teacher model for the item, obtained by applying Softmax with temperature to its logits. The server uses an optimization algorithm to minimize to update the parameters of the global teacher model . The parameter update formula can be expressed as: , where represents the learning rate of the teacher model

[0021] Update the teacher model for the next round of distillation: The model after retraining in the previous step will be used as the new global teacher model for the knowledge distillation link in the next round of federated training, that is, for calculating new soft labels to be sent to the clients, and so on in a loop until the model converges

[0022] As described above, the present disclosure provides a method for training a global model for personalized recommendation based on reinforcement learning and federated knowledge distillation. This method integrates the privacy protection ability of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization ability of reinforcement learning, aiming to solve the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency

[0023] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without the data leaving the local area, effectively balancing global generality and local personalization requirements. At the same time, the client uses reinforcement learning to dynamically optimize the personalized recommendation strategy. For example, it dynamically adjusts the weights of personalization and generalization as well as the weights of the reward function, enabling it to adaptively adjust the recommendation behavior according to the real-time feedback of users, so as to obtain and execute a better recommendation strategy in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0024] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, this disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the policy adjustment feedback (reward signals) of the clients instead of transmitting a large number of model parameters or raw data. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.

[0026] With reference to the drawings, the present disclosure can be more clearly understood from the following detailed description.

[0027] Figure 1 is a flowchart showing a method for training a global model for personalized recommendation according to some embodiments of the present disclosure.

[0028] Figure 2 is an architecture diagram showing the implementation of this method according to some embodiments of the present disclosure.

[0029] Figure 3 is a diagram showing the accuracy of personalized recommendation achieved by applying different methods according to some embodiments of the present disclosure.

[0030] Figure 4 is a diagram showing the recommendation time and communication overhead of personalized recommendation achieved by applying different methods according to some embodiments of the present disclosure.

[0031] Figure 5 is a block diagram showing a device for training a global model for personalized recommendation according to some embodiments of the present disclosure.

[0032] Figure 6 is a block diagram showing a device for training a global model for personalized recommendation according to some other embodiments of the present disclosure.

[0033] Figure 7 is a block diagram showing a computer system for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0035] At the same time, it should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0036] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation on the present disclosure, its application, or its utility.

[0037] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.

[0038] In all the examples shown and discussed here, any specific value should be understood as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0039] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0040] Currently, with the rapid development of artificial intelligence and big data technologies, personalized recommendation systems have become an important tool for major Internet platforms to enhance user experience and commercial value. Personalized recommendation aims to provide users with the most relevant content or products based on users' historical behaviors, interest preferences, and environmental characteristics, thereby improving user satisfaction and platform revenue. In a recommendation system, the core problem is how to improve the accuracy of recommendations and system efficiency while meeting user needs.

[0041] Traditional recommendation algorithms, such as collaborative filtering, matrix factorization, and deep learning, usually rely on centralized data storage and training, but there are many challenges in data privacy protection, computational efficiency, and cross-platform collaboration. For example, data barriers between Internet enterprises make it difficult for recommendation systems to fully utilize users' cross-platform behavior data. At the same time, centralized data storage also easily brings the risk of privacy leakage. Therefore, how to efficiently train a personalized recommendation model without data leaving the local area has become an urgent problem to be solved.

[0042] In the related art, as a distributed machine learning method, federated learning provides an effective solution to address the above problems. Federated learning allows multiple clients, such as different user devices or institutions, to locally train personalized recommendation models and only transmit model parameters to the server instead of the original data, thereby protecting user privacy to a certain extent. However, federated learning still faces challenges in heterogeneous data environments, where the data distributions on different user devices vary greatly, which may lead to a decline in the generalization ability of the global model. In addition, the communication cost of federated learning is relatively high, and how to reduce the transmission overhead and improve the training efficiency is also an important issue facing researchers.

[0043] In view of this, the present disclosure proposes a method for training a global model for personalized recommendation, which integrates the privacy protection ability of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization ability of reinforcement learning, aiming to address the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0044] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without the data leaving the local, effectively balancing global generality and local personalized requirements. At the same time, the client dynamically optimizes the personalized recommendation strategy using reinforcement learning. For example, it dynamically adjusts the weights of personalization and generalization as well as the weights of the reward function, enabling it to adaptively adjust the recommendation behavior according to the real-time feedback of users, so as to obtain and execute better recommendation strategies in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0045] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, the present disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the policy adjustment feedback (reward signals) of the clients instead of transmitting a large amount of model parameters or original data.

[0046] The inventive concept of the present disclosure lies in integrating three technologies to dynamically optimize the recommendation strategy, improve the recommendation accuracy and efficiency while protecting user privacy. First, this method introduces reinforcement learning to dynamically optimize the personalized recommendation strategy of the client. Each client is regarded as an RL agent, which takes the user's historical behavior, device information, and the current state of the recommendation system, including personalized weights and personalized regularization parameters, as the state S. The action A of the RL agent is to adjust the key policy parameters that affect the behavior of the recommendation model, such as personalized weights and personalized regularization strength. By presenting the recommendation results to the user and collecting the user's interaction feedback, such as clicks and dwell time, the system calculates the reward, which quantifies the quality of the current recommendation strategy and user satisfaction. The RL agent aims to maximize the cumulative reward by learning an optimal strategy, that is, continuously adjusting the policy parameters according to the user feedback, so that the recommendation system can adapt to the user's real-time preferences and situational changes, thereby maximizing the user's long-term satisfaction.

[0047] Second, this method uses federated knowledge distillation to achieve secure and efficient cross-platform collaborative learning and knowledge sharing. The server maintains a global teacher model, which learns general and structured item knowledge and user behavior patterns through aggregated or globally visible information. In each federated round, the server sends the soft labels generated by the teacher model to the clients. The personalized recommendation model of the client, as a student model, not only learns to fit the hard target of the user's true rating (through the hard loss, that is, the loss based on the user's true rating) when training locally with its private user data. The hard target refers to the user's true rating data, which is the supervised information that the model needs to directly fit, and the hard loss is the loss calculated according to the hard target, that is, the cross-entropy loss , but also learns to fit the soft labels provided by the teacher model (through the soft loss, that is, the loss calculated based on the guidance information output by the global teacher model). The soft labels are the outputs in the form of probability distributions generated by the global teacher model, which contain the knowledge of the teacher model and are used to guide the learning of the student model. The soft loss is the distillation loss calculated according to the soft labels provided by the teacher model, that is . This way enables the client student model to benefit from the global knowledge of the teacher model and enhances the generalization ability of the model, which is particularly important for users with sparse data or new items. Since the clients only share the knowledge (soft labels) of the teacher model through the server and do not need to directly exchange the original user data or local model parameters (the core training process is completed locally), the user privacy is effectively protected.

[0048] In this disclosure, there is a close interaction and correlation between the reinforcement learning-optimized policy and the knowledge shared by federated knowledge distillation, which is specifically manifested as follows: Reinforcement learning dynamically regulates the influence degree of knowledge distillation: The reinforcement learning agent directly controls the balance between the hard loss and the soft loss (derived from knowledge distillation) in the client model training by optimizing the personalized weights. This means that the system can intelligently decide whether to rely more on personalized local data learning or more on the knowledge of the global teacher model according to the specific situation of the user and the reward feedback obtained by reinforcement learning.

[0049] Knowledge distillation provides a generalization basis for reinforcement learning: Knowledge distillation enables the client student model to possess basic generalization ability, which provides a stable starting point for the reinforcement learning agent to explore better personalized strategies. A student model with strong generalization ability performs more robustly when facing new items or user changes, enabling the reinforcement learning agent to learn and adjust more effectively.

[0050] The reinforcement learning feedback optimizes the goal of knowledge distillation: By aggregating the knowledge of the student model trained under the optimization of reinforcement learning by the client, such as predicting soft labels, the server can periodically update the global teacher model . This means that the global knowledge learned by the teacher model is not static, but absorbs the effective patterns explored by the client in the personalized reinforcement learning process, forming a closed-loop iterative optimization of a federated teacher model.

[0051] By achieving a better balance between global knowledge sharing and personalized learning, while effectively reducing the overall communication cost. It mainly includes the following aspects: First, the design of a two-way personalized mechanism, which optimizes the personalized recommendation strategy by combining reinforcement learning and federated learning; Second, the optimization of the dynamic regularization learning strategy, which dynamically adjusts the regularization strategy according to the user behavior feedback to improve the stability and generalization ability of the model; Then, the design of global item sparsification, which optimizes the embedding representation of global items through sparsification technology to improve the accuracy and efficiency of the recommendation effect; Finally, the design of a privacy protection mechanism, which realizes cross-platform data collaboration and knowledge sharing by using federated knowledge distillation on the premise of ensuring data privacy. These designs work together to promote the improvement of the efficiency, accuracy, and privacy protection ability of the personalized recommendation system.

[0052] Such as Figure 2As shown in the figure, this architecture mainly includes a server and various clients. The server is responsible for aggregating and updating the reward function values uploaded from each client, and saving the embedded features of the global project and the shared knowledge. The client is responsible for training and updating the personalized recommendation model for the user data stored locally, and optimizing the recommendation strategy through reinforcement learning. The client and the server cooperate through the method of federated learning to ensure the protection of data privacy, and share the global knowledge through the knowledge distillation technology, so as to improve the accuracy and efficiency of the recommendation system.

[0053] First, for the initial users and items in the recommendation system, definitions are made. Among them, is the globally shared information, that is, the item embedding that all users can use; represents the th user's local item embedding; the local item embedding can be the embedding of the items interacted locally by the user, or a sparse matrix or low-rank matrix corresponding to the global item embedding and reflecting the personalized differences. Its specific implementation aims to balance the personalized representation ability and the local storage and computing overhead. is the rating matrix of all users for items. Among them, represents the number of users, represents the number of items, represents the th user's rating. Assuming that each client only contains the information of one user, therefore, also represents the rating of the client. Assuming , it means that the user has rated the item, and is used to mark the rated items in the rating matrix . The main role of is to define the set of training samples with the user's true ratings used when calculating the hard loss

[0054] Figure 1 is a flowchart showing a method for training a global model for personalized recommendation according to some embodiments of the present disclosure. As Figure 1 shown, the method for training a global model for personalized recommendation includes steps S110 to S140.

[0055] In step S110, initialize the global item embedding , the global teacher model , the global weight , the global regularization parameter , among which, sparsify the global item embedding to obtain the sparsified global item embedding .

[0056] Server initialization, including global item embedding and global teacher model and global weights and global regularization parameters .

[0057] While the server is initializing, each client also needs to be initialized, including user embedding and local item embedding , set the reinforcement learning agent (RL-Agent), and initialize local policy parameters , , , , , local parameters , are initialized to the global parameters sent by the server , .

[0058] After initializing the global item embedding , sparsification processing is performed on it. For example, based on item popularity or preset rules, some dimensions in or the embeddings of specific items can be set to zero or the number of non-zero elements can be restricted, so as to obtain the sparsified global item embedding . This sparsification processing aims to reduce the storage and transmission overhead of the global item embedding, and may improve the efficiency and generalization ability of the model by introducing inductive bias. In subsequent steps, the client will use this sparsified global item embedding .

[0059] In step S120, the predicted value and soft label are calculated through the global teacher model .

[0060] The soft label is calculated through the global teacher model , including: According to the formula: , the soft label is calculated, where represents the predicted value calculated by the global teacher model , represents the temperature parameter used to control the softening degree of the knowledge distillation signal. The temperature parameter is usually a hyperparameter greater than 0, and its specific value can be set according to experience or determined by experimental tuning. A larger value will generate a smoother probability distribution, and a smaller Values result in a smoother probability distribution, while smaller values (close to 0) make the probability distribution sharper, approaching the one-hot form.

[0061] The temperature parameter is a positive number used to control the soft labels generated by the global teacher model during the knowledge distillation process. Generally, as a hyperparameter, it can be determined by cross-validation, grid search, or experimental tuning on the validation set within a preset range to obtain the optimal personalized student model training results. Larger values make the probability distribution of the soft labels smoother, helping the student model learn the potential relationships between classes; smaller values (when = 1, it is the standard Softmax) make the probability distribution closer to the original prediction confidence of the teacher model. In this disclosure, choosing an appropriate value helps balance the degree to which the student model absorbs the knowledge of the teacher model and its ability to fit hard targets.

[0062] is the predicted value calculated by the teacher model; is the temperature parameter used to control the softening degree of the knowledge distillation signal, when it is larger, the distribution is smoother.

[0063] The server sends the sparsified global item embeddings and the soft labels to the client to guide personalized training; The server sends 、 as the initial values for the client's reinforcement learning.

[0064] In step S130, the sparsified global item embeddings 、soft labels 、global weights 、global regularization parameters are sent to the client. Among them, the steps executed on the client include: assigning the global weights 、global regularization parameters to the personalized weights 、personalized regularization parameters respectively; determining the personalized recommendation score based on the user embedding 、local item embedding 、sparsified global item embedding ; according to the personalized recommendation score , user's true rating , soft label , reinforcement learning reward function value , local hyperparameters , determine the total loss of the current training batch ; calculate the total loss with respect to the trainable parameters in the personalized student model ; use the selected optimization algorithm and the gradient to update the trainable parameters of the personalized student model ; use the set reinforcement learning agent to adjust the personalized weights , personalized regularization parameter , reward function weight , , ; based on the reward function weight , , , update the reinforcement learning reward function value ; based on the determined optimal policy parameters , maximize ; local reward function weight , , are initialized, for example, by evenly distributing the weights.

[0065] Calculate the personalized recommendation rating as follows: .

[0066] Locally train the personalized student model to optimize the objective function ; on the client side locally, use the received sparsified global item embeddings , teacher model soft labels , client-side local data (including user embeddings , local item embeddings , user's true rating ), the current local personalized weights and personalized regularization parameters , train the personalized student model. The main trainable parameters of the personalized student model include user embeddings, local item embeddings, and other possible parameters in the model, such as including additional neural network layers. The training process is achieved by minimizing the following objective function : ; where represents the cross-entropy loss of the true rating, and its calculation formula is: ; represents the distillation loss of the teacher model soft labels, and its calculation formula is: 。 is the reward function value obtained by the client through reinforcement learning. The specific training and update steps are as follows: Calculate the current loss 。According to the formula ,use the calculated predicted score 、the user's true score 、the teacher model's soft label 、and the local hyperparameters of the current round 、 、 and the client reward (function value) to calculate the total loss in the current training batch or the current state 。

[0067] Calculate the model parameter gradients. Calculate the total loss function with respect to all trainable parameters in the personalized student model. This gradient is a vector composed of the partial derivatives of the loss function with respect to each parameter. Specifically, the gradient is expressed as: where, represents the gradient of the hard loss with respect to the student model parameters, represents the gradient of the soft loss with respect to the student model parameters, represents the gradient of the reward with respect to the student model parameters, which is usually related to the policy gradient or related techniques in reinforcement learning. These gradients can be calculated through the backpropagation algorithm.

[0068] Update the model parameters. Use a selected optimization algorithm, such as stochastic gradient descent, Adam, etc., and the calculated gradients to update each trainable parameter of the personalized student model 。The parameter update formula is expressed as: where, represents the local learning rate of the student model.

[0069] Local iterative training. Repeat the above steps for several training epochs or multiple mini-batch iterations on the user dataset stored locally on the client. In this local training phase, the personalized weights and the personalized regularization parameter are usually fixed to the values at the start of this federated learning or after the most recent adjustment by the RL agent in this round. They affect the composition and optimization direction of the loss function, but their updates are completed through the reinforcement learning policy.

[0070] The client uses reinforcement learning to optimize the recommendation strategy (including Parameter dynamic update). The reinforcement learning agent (RL-Agent) on the client side is responsible for dynamically adjusting the local personalized weights , the personalized regularization parameter , and the reward function weight , , , to maximize the long-term recommendation reward.

[0071] Construct the state space. The state space S is composed of the user's historical interaction data and the feedback information of the recommendation system, mainly including: the user's behavior data , that is, the item ID sequence of the user's last interactions; the user's historical click-through rate , the user's device type , the current local personalized weight of the recommendation system , the regularization parameter , and the reward function weight . The final state space is defined as follows: .

[0072] Construct the action space. The action space is mainly the change amount of the policy parameters that the recommendation system can adjust. The RL agent selects an action to adjust the local personalized weight , the personalized regularization parameter and the reward function weight . The final formed action space is as follows: .

[0073] Sampling of action execution and parameter update. The RL agent samples an action according to the probability distribution output by the current state S (including user behavior data, device type, current personalized weight, regularization parameter, and reward function weight) and the policy network (referring to a policy network with parameters , which outputs the probability distribution on the action space based on the current state ). The client calculates the dynamic step size using the obtained reward , for example . Then, use this step size to scale the sampled action and update the local hyperparameters , and the reward function weight : ; ; ; .

[0074] Design the reward function. The reward function is mainly used to measure the recommendation quality and is defined as follows: , represents the click-through rate, represents the normalized discounted cumulative gain, represents the user's stay time of the recommended content, , , are the reward function weight coefficients dynamically adjusted locally on the client side, used to control the contribution of different metrics to the reward.

[0075] Calculate the policy gradient. By adopting the policy gradient method, find the optimal policy parameters , to maximize the expected reward. Calculate the policy gradient as follows: . According to the formula , update , until the optimal policy parameters are found, where at the initial stage of training a larger value is adopted, such as 0.1, and it gradually decays in the later stage.

[0076] The client uploads the reward. The client uploads the calculated reward function value (reward ) to the server side, protecting user privacy. The reward function weights , , dynamically adjusted locally on the client side are not uploaded to the server.

[0077] In step S140, aggregate the rewards uploaded by each client and update the global weights , the global regularization parameter .

[0078] The server side aggregates the reward function values uploaded by the client and updates the global personalized weights and the global regularization parameter.

[0079] The server side collects the rewards uploaded by all active clients participating in the federated learning , and calculates the aggregated global average reward : , represents the total number of active clients participating in the current round of federated learning.

[0080] Update the global weights. Use the aggregated global average reward to update the global weights , that is: , where represents the learning rate of the global weights, the initial value is 0.01, to avoid large update amplitudes causing oscillations.

[0081] Update the global regularization parameter. Use the aggregated global average reward to update the global regularization parameter , denotes the learning rate of the global regularization parameter. and have the same value. Use to update The purpose is that when the overall recommendation effect is good (aggregated reward is high), is lower, resulting in the regularization parameter decreasing, allowing the global model to fit the data more precisely; when the recommendation effect is poor (aggregated reward is low), is higher, resulting in the regularization parameter increasing, enhancing the generalization ability or diversity of the model to cope with the overall performance decline.

[0082] The server sends the updated global parameters. The server side sends the updated global weights and global regularization parameters to the client as the initial values for the next round of reinforcement learning.

[0083] Enter the next round of training until the model converges.

[0084] The above steps do not clearly describe the update mechanism of the global teacher model during the federated learning iteration process. In this disclosure, it should be assumed that it is pre-trained and remains static before the start of federated learning.

[0085] In this disclosure, the update of the global model can be further discussed.

[0086] Update of the global teacher model. At the end of each federated training round or after a preset number of rounds, the server side can initiate the update process of the global teacher model so that it can learn the collective knowledge of the client models and the latest changes in the data distribution. This disclosure uses the method of aggregating the knowledge of the client student models to update the global teacher model . Upload of client prediction results. Each client participating in federated learning after completing its local training, uses its current student model to predict a preset set of common items or a set of global item samples specified by the server. For each item in the set, the student model calculates the corresponding prediction . The student model calculates the corresponding prediction logits, which are the personalized student models for the item The raw prediction values before the output passes through the final activation function (e.g., the Sigmoid function used to calculate the recommendation score). The client converts these logits into soft labels by applying the same temperature parameter as the teacher model , as follows: where, is a vector representing the predicted probability distribution of the item by the student model over the set. The client uploads this set of soft labels generated by the student model to the server. Server aggregates student model knowledge. The server aggregates all the soft labels of the student models uploaded by the clients. For each item in the set, the server aggregates the received soft labels of the student models, e.g., by calculating the average, to obtain the aggregated student soft labels, as follows:

[0087] This aggregated result represents the collective prediction distribution or collective knowledge of all current client models for the item.

[0088] Global teacher model retraining. The server uses the obtained aggregated student labels as the training target to retrain the global teacher model . The retraining objective of the teacher model is to minimize the difference between its soft label output and the aggregated student soft labels . This is usually achieved by minimizing the knowledge distillation loss between them, such as cross-entropy loss or KL divergence. Taking the cross-entropy loss between the aggregated student soft labels and the soft labels output by the teacher model as an example, the objective function for updating the teacher model can be defined as: where, represents the soft label generated by the global teacher model for the item , obtained by applying the Softmax with temperature to its logits. The server uses an optimization algorithm, e.g., gradient descent, to minimize and update the parameters of the global teacher model . The parameter update formula can be expressed as where, are the trainable parameters of the global teacher model , represents the learning rate of the teacher model, is the objective loss function for the retraining of the teacher model defined above, is the gradient of the objective loss function with respect to the parameter .​

[0089] Update the teacher model for the next round of distillation. The retrained model will serve as the new global teacher model , which is used for the knowledge distillation process in the next round of federated training. This cycle continues until the global model converges.

[0090] In summary, the present disclosure provides a method for training a global model for personalized recommendation based on reinforcement learning and federated knowledge distillation. This method combines the privacy protection capabilities of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization capabilities of reinforcement learning, aiming to address the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0091] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without data leaving the local, effectively balancing global generality and local personalization needs. At the same time, the client uses reinforcement learning to dynamically optimize the personalized recommendation strategy. For example, it dynamically adjusts the weights of personalization and generalization as well as the weights of the reward function, enabling it to adaptively adjust the recommendation behavior according to the user's real-time feedback. Thus, it can obtain and execute better recommendation strategies in different user behaviors and cross-platform data environments, significantly improving recommendation accuracy and user satisfaction.

[0092] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, the present disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the client's policy adjustment feedback (reward signals) instead of transmitting a large number of model parameters or raw data.

[0093] Under the premise of ensuring the privacy and security of user data, the present disclosure achieves high accuracy, high efficiency, and strong adaptability of the personalized recommendation system (globally).

[0094] As Figure 3 and Figure 4 shown, in order to verify the recommendation accuracy, recommendation time, and communication overhead of the method proposed in the present disclosure under different datasets, the following experiments were conducted in this solution.

[0095] This solution experimentally verified the actual effects of the proposed global model for personalized recommendation based on reinforcement learning and federated knowledge distillation. The experiments selected multiple publicly available recommendation system datasets, such as MovieLens, Amazon, and Last.fm, covering different user interaction scenarios, and divided the data into training sets and test sets.

[0096] MovieLens is a movie recommendation system research dataset maintained by the GroupLens research team at the University of Minnesota. It contains information such as user ratings of movies, movie metadata, and user attributes. The Amazon dataset mainly contains user consumption behavior datasets on e-commerce platforms, covering product reviews, purchase records, user attributes, etc. Last.fm is a dataset for a music streaming platform that records users' music playback behaviors and social interactions.

[0097] The comparison methods include traditional collaborative filtering (CF) recommendation methods, deep learning (DNN) based recommendation methods, recommendation methods using only federated learning (FL), recommendation methods using only reinforcement learning (RL), and the proposed recommendation method based on reinforcement learning and federated knowledge distillation (RL-FedKD).

[0098] CF: Collaborative Filtering, corresponding Chinese: collaborative filtering; DNN: Deep Neural Detwork, corresponding Chinese: deep neural network; FL: Federated Learning, corresponding Chinese: federated learning; RL: Reinforcement Learning, corresponding Chinese: reinforcement learning; RL-FedKD: Reinforcement Learning-Federated Knowledge Distillation, corresponding Chinese: reinforcement learning combined with federated knowledge distillation.

[0099] The experiment uses indicators such as recommendation accuracy, recommendation time (computational time of different methods) and communication overhead (data transmission volume between the client and the server) to evaluate the performance of different recommendation methods.

[0100] The results show that the (personalized) recommendation method based on reinforcement learning and federated knowledge distillation performs well in many aspects. First, in terms of recommendation accuracy, RL-FedKD achieves the highest accuracy on MovieLens, Amazon, and Last.fm datasets, respectively, which is better than traditional CF and DNN, with an improvement of 20% to 30%, indicating that the method disclosed in this paper can more accurately capture user interests and improve personalized recommendation effects. Secondly, in terms of recommendation time, the calculation time of the method disclosed in this paper is about 0.95 seconds, which is 21% less than DNN and 9.5% less than RL, indicating that while optimizing the recommendation accuracy, the response speed of the recommendation is improved. Finally, in terms of communication overhead, the method disclosed in this paper is about 52MB, which is 35% less than FL and 25% less than RL, effectively reducing the cost of cross-platform data exchange.

[0101] As described above, while ensuring data privacy and improving the recommendation effect, RL-FedKD reduces the consumption of communication bandwidth and optimizes the operation efficiency of the global model.

[0102] In the solution of the present disclosure, the recommendation strategy is optimized through reinforcement learning, enabling the system to dynamically adjust the recommended content, thereby improving the accuracy of the recommendation and reducing the search space for ineffective recommendations. The performance of the recommendation model is tested in different dataset scenarios, providing strong data support for the optimization of subsequent recommendation algorithms. The experimental results show that the method of the present disclosure is superior to traditional recommendation methods in terms of recommendation accuracy, recommendation time, and communication overhead, and can improve the overall performance of the global model (personalized recommendation global model) while ensuring data privacy.

[0103] Figure 5 It is a block diagram showing a personalized recommendation global model training device according to some embodiments of the present disclosure. As Figure 5 shown, the personalized recommendation global model 500 includes an initialization module 510, a calculation module 520, a reinforcement learning reward function value update module 530, and a weight and regularization parameter update module 540.

[0104] The initialization module 510 is configured to initialize the global item embedding , the global teacher model , the global weight , the global regularization parameter , wherein the global item embedding is sparsified to obtain the sparsified global item embedding ; The calculation module 520 is configured to calculate the predicted value , the soft label through the global teacher model ; The reinforcement learning reward function value update module 530 is configured to send the sparsified global item embedding , the soft label , the global weight , the global regularization parameter to the client, wherein the steps executed on the client include: assigning the global weight , the global regularization parameter to the personalized weight , the personalized regularization parameter respectively; determining the personalized recommendation score according to the user embedding , the local item embedding , the sparsified global item embedding ; according to the personalized recommendation score , User's true rating , Soft label , Reinforcement learning reward function value , Local hyperparameters , Determine the total loss of the current training batch ; Calculate the total loss with respect to the trainable parameters in the personalized student model ; Update the trainable parameters of the personalized student model using the selected optimization algorithm and gradient ; Adjust the personalized weights using the set reinforcement learning agent , Personalized regularization parameter , Reward function weight , , ; Based on the reward function weight , , , Update the reinforcement learning reward function value ; Based on the determined optimal policy parameters , Maximize ; The weight and regularization parameter update module 540 is configured to aggregate the uploaded by each client and update the global weights , Global regularization parameter .

[0105] In the device of the present disclosure embodiment, the privacy protection ability of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization ability of reinforcement learning are integrated, aiming to solve the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0106] Through federated knowledge distillation, this method can achieve secure knowledge sharing between clients and servers without data leaving the local, effectively balancing global generality and local personalized needs. At the same time, the client uses reinforcement learning to dynamically optimize the personalized recommendation strategy. For example, dynamically adjusting the weights of personalization and generalization and the reward function weight, enabling it to adaptively adjust the recommendation behavior according to the user's real-time feedback, so as to obtain and execute a better recommendation strategy in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0107] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, the present disclosure effectively reduces the communication overhead and improves the overall computational efficiency of the system by federally aggregating the policy adjustment feedback (reward signals) of clients instead of transmitting a large number of model parameters or raw data.

[0108] Figure 6is a block diagram showing a personalized recommendation global model training apparatus according to other embodiments of the present disclosure. As Figure 6 shown, the personalized recommendation global model 600 includes a memory 610; and a processor coupled to the memory 610. The memory 610 is used to store instructions for implementing corresponding embodiments of the personalized recommendation global model training method. The processor 620 is configured to execute the personalized recommendation global model training method in any of the embodiments of the present disclosure based on the instructions stored in the memory 610.

[0109] Figure 7 is a block diagram showing a computer system for implementing some embodiments of the present disclosure. As Figure 7 shown, the computer system 700 may be embodied in the form of a general-purpose computing device. The computer system 700 includes a memory 710, a processor 720, and a bus 730 connecting different system components.

[0110] The memory 710 may include, for example, a system memory, a non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs. The system memory may include a volatile storage medium, such as a random access memory (RAM) and / or a cache memory. The non-volatile storage medium stores, for example, instructions for implementing corresponding embodiments of at least one of the personalized recommendation global model training methods. The non-volatile storage medium includes, but is not limited to, a disk memory, an optical memory, a flash memory, etc.

[0111] The processor 720 may be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete hardware components such as discrete gates or transistors. Correspondingly, each of the modules such as an initialization module, a calculation module, a sending module, an aggregation module, and an update module may be implemented by a central processing unit (CPU) running instructions for executing corresponding steps in the memory, or may be implemented by a dedicated circuit for executing the corresponding steps.

[0112] The bus 730 may use any of a variety of bus structures. For example, the bus structure includes, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus.

[0113] The computer system 700 may further include an input / output interface 740, a network interface 750, a storage interface 760, etc. These interfaces 740, 750, 760, the memory 710, and the processor 720 may be connected through a bus 730. The input / output interface 740 may provide a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 750 provides a connection interface for various networking devices. The storage interface 760 provides a connection interface for external storage devices such as a floppy disk, a USB flash drive, and an SD card.

[0114] Here, various aspects of the present disclosure have been described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments of the present disclosure. It should be understood that each box of the flowcharts and / or block diagrams, and combinations of the boxes, can be implemented by computer-readable program instructions.

[0115] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable devices to generate a machine, such that the device implemented by executing the instructions by the processor performs the functions specified in one or more boxes of the flowcharts and / or block diagrams.

[0116] These computer-readable program instructions can also be stored in a computer-readable memory, and these instructions cause the computer to work in a specific manner, thereby generating a manufactured article including instructions for implementing the functions specified in one or more boxes of the flowcharts and / or block diagrams.

[0117] The present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0118] The present disclosure provides a method for training a personalized recommendation global model based on reinforcement learning and federated knowledge distillation. This method integrates the privacy protection ability of federated learning, the global knowledge sharing mechanism of knowledge distillation, and the dynamic policy optimization ability of reinforcement learning, aiming to address the challenges faced by existing recommendation systems in terms of data privacy, recommendation accuracy, and system efficiency.

[0119] Through federated knowledge distillation, this method can achieve secure knowledge sharing between the client and the server without the data leaving the local, effectively balancing global generality and local personalized needs. At the same time, the client uses reinforcement learning to dynamically optimize the personalized recommendation strategy. For example, it dynamically adjusts the weights of personalization and generalization as well as the weights of the reward function, enabling it to adaptively adjust the recommendation behavior according to the real-time feedback of users, so as to obtain and execute a better recommendation strategy in different user behaviors and cross-platform data environments, significantly improving the recommendation accuracy and user satisfaction.

[0120] Compared with traditional centralized training or federated learning methods that only rely on model parameter aggregation, the present disclosure effectively reduces communication overhead and improves the overall computational efficiency of the system by federally aggregating the policy adjustment feedback (reward signals) of clients instead of transmitting a large number of model parameters or raw data.

[0121] So far, the method for training a global model for personalized recommendation according to the present disclosure has been described in detail. To avoid obscuring the concept of the present disclosure, some details well known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

[0122] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A method for training a global model for personalized recommendation, characterized in that, The method includes: Initialize global item embeddings , global teacher model , global weights , global regularization parameter , where the global item embeddings are sparsified to obtain sparsified global item embeddings ; Through the global teacher model The predicted value is calculated and the soft label ; Send the sparsified global item embeddings , soft labels , global weights , global regularization parameters to the client, where the steps executed on the client include: Assign the global weights , global regularization parameters to the personalized weights , personalized regularization parameters respectively; Determine the personalized recommendation score according to the user embeddings , local item embeddings , sparsified global item embeddings ; Determine the total loss of the current training batch according to the personalized recommendation score , user true score , soft labels , reinforcement learning reward function value , local hyperparameters ; Calculate the gradient of the total loss with respect to the trainable parameters in the personalized student model; Update the trainable parameters of the personalized student model using the selected optimization algorithm and gradient; Adjust the personalized weights , personalized regularization parameters , reward function weights , , using the set reinforcement learning agent; Update the reinforcement learning reward function value , , based on the reward function weights ; Maximize based on the determined optimal policy parameters ; Aggregate the data uploaded by each client and update the global weights and the global regularization parameter .

2. The personalized recommendation global model training method according to claim 1, wherein The soft labels calculated by the global teacher model include: ​ According to the formula: , the soft label is calculated, where represents the predicted value calculated by the global teacher model , and represents the temperature parameter used to control the softening degree of the knowledge distillation signal.

3. The personalized recommendation global model training method according to claim 2, wherein The above-mentioned user-embedded , locally item-embedded , sparsified globally item-embedded , to determine personalized recommendation scores , including: According to the formula: , the personalized recommendation score is calculated as .

4. The personalized recommendation global model training method according to claim 3, wherein According to the personalized recommendation score , the user's true score , soft labels , the value of the reinforcement learning reward function , local hyperparameters , determine the total loss of the current training batch , including: According to the formula: , determine the total loss ; where represents the value of the reinforcement learning reward function; , represents the cross-entropy loss of the true score; , represents the distillation loss of the soft label.

5. The personalized recommendation global model training method according to claim 4, wherein The calculated total loss Regarding the trainable parameters in the personalized student model, including: According to the formula: , determine the gradient , where represents the gradient of the hard loss with respect to the student model parameters, represents the gradient of the soft loss with respect to the student model parameters, represents the gradient of the reward with respect to the student model parameters.

6. The personalized recommendation global model training method according to claim 5, wherein Updating the trainable parameters of the personalized student model by using the selected optimization algorithm and gradient , including: According to the formula: , update , where represents the local learning rate of the student model.

7. The personalized recommendation global model training method according to claim 6, wherein Adjusting the personalized weights using the set reinforcement learning agent , the personalized regularization parameter , the reward function weight , , , including: Define the state space is , where represents the user's behavior data, represents the user's historical click-through rate, represents the user device type, represents the current personalized weight of the recommendation system, represents the current personalized regularization parameter of the recommendation system, represents the reward function weight; Define the action space For ; According to the current state and the policy network output the probability distribution, and adopt an action ; According to the formula: , calculate the dynamic step size ; According to the step size , update , , , , .

8. The personalized recommendation global model training method according to claim 7, characterized in that Based on the reward function weights , , , update the value of the reinforcement learning reward function , including: According to the formula: , where represents the click-through rate, represents the normalized discounted cumulative gain, represents the user's stay time on the recommended content.

9. The personalized recommendation global model training method according to claim 8, characterized in that Based on the determined optimal policy parameters , maximize , including: According to the formula: , determine gradient of ; According to the formula: , update until the optimal policy parameters are found.

10. The personalized recommendation global model training method according to claim 1, characterized in that Aggregating the uploaded by each client and updating the global weights and the global regularization parameter , including: According to the formula: , determine the globally averaged reward after aggregation ; According to the formula: , update , where represents the learning rate of the global weight; According to the formula: , update , where represents the learning rate of the global regularization parameter.

Citation Information

Patent Citations

  • Content recommendation model training method, content recommendation method and related equipment

    CN113360777A

  • Object recommendation method and device and storage medium

    CN116992162A

  • Method for knowledge distillation and model genertation

    US20230351203A1