Communication-Efficient Federated Matrix Factorization Recommendation Method

By combining federated learning and matrix decomposition models in the recommendation system, a federated matrix decomposition recommendation method with efficient communication is designed, which solves the problems of privacy leakage and poor data dispersion in traditional recommendation systems, and achieves efficient and secure personalized recommendations.

CN118332165BActive Publication Date: 2025-06-13SOUTHWESTERN UNIV OF FINANCE & ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410479102.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-21
Publication Date
2025-06-13
Estimated Expiration
2044-04-21

AI Technical Summary

Technical Problem

Traditional centralized recommendation systems have the risk of privacy leakage when processing user data, and the data is poorly distributed, making it difficult to effectively utilize user distributed data.

Method used

Combining federated learning technology with matrix decomposition recommendation model is used to design a communication-efficient federated matrix decomposition recommendation method. By building a federated matrix recommendation framework, building a proxy objective function, and solving parameter update formulas using alternating direction multipliers (ADMM), we can realize model training and recommendation without sharing the original data.

Benefits of technology

On the premise of ensuring the recommendation effect, reduce training time and speed up the algorithm convergence speed, thereby reducing communication costs, protecting user privacy, and effectively utilizing distributed data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118332165B_ABST
    Figure CN118332165B_ABST
Patent Text Reader

Abstract

The present invention discloses a federated matrix factorization recommendation method with high communication efficiency, including: constructing a federated matrix recommendation framework to generate a federated matrix factorization recommendation algorithm; constructing a surrogate objective function and using the surrogate objective function to replace the global objective function; using the Alternating Direction Method of Multipliers (ADMM) to solve the parameter update formula in the surrogate objective function; designing a communication-efficient federated matrix factorization recommendation algorithm, including designing a central server algorithm and a client algorithm; the present invention combines federated learning with a recommendation system, which can solve the problems of privacy and data dispersion to a certain extent and improve the effect of personalized recommendation at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of federated learning, and particularly to a communication-efficient federated matrix factorization recommendation method. Background Art

[0002] A recommendation system is a technology that provides personalized recommendations by analyzing user behavior and personal preferences. Traditional centralized recommendation systems first collect users' behavioral data, such as website browsing records, search histories, etc., and centrally summarize these data. Then, on the server side, common recommendation algorithms, such as collaborative filtering, content filtering, or deep learning models, are used to centrally train the summarized user behavioral data. Finally, the trained model is deployed to an online recommendation system, and personalized recommendations are provided for users through this model. The specific process is as Figure 1 shown. In traditional recommendation systems, the client sends user data to the central server for model training, which may involve the risk of user privacy leakage. Summary of the Invention

[0003] To solve the problems existing in the prior art, the purpose of the present invention is to provide a communication-efficient federated matrix factorization recommendation method. The combination of federated learning and the recommendation system in the present invention can, to a certain extent, solve the problems of privacy and data dispersion, and at the same time improve the effect of personalized recommendations.

[0004] To achieve the above object, the technical solution adopted by the present invention is: A communication-efficient federated matrix factorization recommendation method, comprising the following steps:

[0005] Step 1: Construct a federated matrix recommendation framework and generate a federated matrix factorization recommendation algorithm;

[0006] Step 2: Construct a surrogate objective function and use the surrogate objective function to replace the global objective function;

[0007] Step 3: Use the alternating direction method of multipliers (ADMM) to solve the parameter update formula in the surrogate objective function;

[0008] Step 4: Design a communication-efficient federated matrix factorization recommendation algorithm, including designing a central server algorithm and a client algorithm.

[0009] As a further improvement of the present invention, the specific content of the said Step 1 is as follows:

[0010] Suppose there are N clients participating in federated learning. There are a total of m users in the clients. The client set is C = {C 1 , C 2 , ···, C N}, and the client C i has m iA user, all clients have the same n items, and the data sets owned by each client are \(R = \{R 1 ,R 2 ,\cdots,R N \}\), where \(R i \) is the local data owned by client \(C i \), and \(R i \) is an \(m i \times n\)-dimensional user-item rating matrix. The global objective function is defined as:

[0011]

[0012] Among them, represents the weight of each client, and F i (\theta)\) represents the local objective function of client \(C i \). The formula for \(F i \)(\theta)\) is as follows:

[0013]

[0014] Among them, \(l\) is the loss function, and \(h(x j ;\theta)\) represents the prediction of the instance \(x j \) in the data set by the model with parameter \(\theta\);

[0015] In federated learning, it is usually hoped to obtain a global model \(\theta'\) with optimal performance. Then the goal of the central server is to optimize the minimum objective function \(F R \)(\theta)\), and then obtain a set of model parameters \(\theta'\) that makes its optimization effect the best as the finally trained global model parameters. Its formulaic representation is:

[0016]

[0017] Decompose the rating matrix \(R u×n \) to obtain two matrices \(P u×k \) and \(Q k×n \). Among them, \(P u×k \) is the user feature matrix, and \(Q k×n \) is the item feature matrix, and \(k\) is the dimension of the latent vector;

[0018] For the prediction formula of the rating i of user \(u\) on a certain item \(n\) in client \(C \) is as follows:

[0019]

[0020] Among them, \(\mu\) is the global mean bias value set according to the global rating data, and \(p iu represents client \(C iThe user feature matrix P i The u-th row user feature vector of denotes p iu vector transpose, q n denotes the n-th row item feature vector of Q;

[0021] The global objective function is defined as:

[0022]

[0023] where λ is the regularization coefficient, and λ(||p iu || 2 +||q n || 2 ) is the regularization term;

[0024] The federated matrix factorization recommendation algorithm generates an m i ×n-dimensional predicted rating matrix R' i for client C without directly sharing its own dataset among all clients. Then, based on the predicted rating matrix, suitable items are recommended for the users of client C. The federated matrix factorization recommendation algorithm first randomly generates an m i ×k-dimensional user feature matrix P i for client C at the client side, randomly generates an n×k-dimensional global item feature matrix Q at the recommendation server, and then generates an accurate predicted rating matrix by minimizing the global objective function, and further makes recommendations for users. i randomly generates an m i ×k-dimensional user feature matrix P i at client C, randomly generates an n×k-dimensional global item feature matrix Q at the recommendation server, and then generates an accurate predicted rating matrix by minimizing the global objective function, and further makes recommendations for users.

[0025] As a further improvement of the present invention, step 2 is specifically as follows:

[0026] Construct a surrogate objective function to replace the global objective function F R (θ); performing a Taylor expansion on F R (θ) can obtain an infinite series:

[0027]

[0028] where, is an arbitrary initial estimate of θ, is a computable constant, and <·,·> represents the inner product; next, a small sample dataset is randomly selected from each client's local dataset, a sample objective function is constructed, and the high-order derivative part in the sample objective function is used to replace the high-order derivative part of the global objective function. The specific steps are as follows:

[0029] Step 2.1, set a sampling ratio α, and randomly select a dataset of size r from each client's local dataseti = αR i is a subset of, constituting the sampling sample set r = {r 1 , r 2 , ···, r N}. Assume that r is much smaller than R, i.e., and when R → ∞, r → ∞; The objective function based on the sampling samples is expressed as:

[0030]

[0031] Perform Taylor expansion on F r (θ) and transform it into the following form:

[0032]

[0033] Step 2.2: Use the high-order derivative part of the objective function based on the sampling samples to replace the high-order derivative in the global objective function That is, let:

[0034]

[0035] Substitute Equation (9) into Equation (6), and the approximation of the global objective function F R (θ) can be obtained:

[0036]

[0037] Ignore the constant term therein and simplify to obtain a communication-efficient surrogate objective function:

[0038]

[0039] Solving for the optimal parameters is expressed as minimizing the global objective function, i.e.:

[0040]

[0041] As a further improvement of the present invention, the specific steps of step 3 are as follows:

[0042] Use the Alternating Direction Method of Multipliers (ADMM) to solve Equation (12), and decompose the minimization of the surrogate objective function into the minimization of each client's local objective function; Let Rewrite Equation (12) into the following form:

[0043]

[0044] where represents the weight of each client, and represents the i-th client based on the sampling sample r iThe local objective function; Equation (13) is further expressed as the following global consistency optimization problem:

[0045]

[0046] where θ i represents the local model parameters of the i-th client; under the constraint θ i = θ, i ∈ [1, 2, ···, N], Equation (14) is equivalent to Equation (13); transform Equation (14) into the form required by the ADMM algorithm:

[0047]

[0048] where B represents the N×N identity matrix, represents the set of all θ i , that is The specific definition of

[0049]

[0050] The augmented Lagrangian function corresponding to Equation (15) is as follows:

[0051]

[0052] where ρ > 0 is the penalty parameter, and the dual variable u ∈ [u 1 , u 2 , ···, u N ; Next, the update rules of the ADMM algorithm are given:

[0053]

[0054]

[0055]

[0056] For Equation (18), this problem is only related to the client C i , and it is rewritten in the following form:

[0057]

[0058] The linearized ADMM algorithm is used to simplify the solution of Equation (21); for at using the second-order Taylor expansion, the expansion of can be obtained:

[0059]

[0060] where η is the penalty term coefficient; then The update formula can be written as:

[0061]

[0062] Equation (23) is an approximation of Equation (21), and an approximate solution of Equation (21) can be obtained by solving Equation (23):

[0063]

[0064] For Equation (19), it is a quadratic function of θ. Taking the partial derivative of it and setting the partial derivative equal to zero can obtain the minimum point, that is:

[0065]

[0066] As a further improvement of the present invention, step 4 is specifically as follows:

[0067] Let Before each round of federated training starts, first update the surrogate objective function, that is Subsequently, update the parameters according to the formula;

[0068] The federated matrix factorization algorithm contains two parameters, namely the user feature matrix P i and the item feature matrix Q. The client initializes the user feature matrix P locally i , and the central server initializes the global item feature matrix Q and sends it to the client; after receiving the global item feature matrix Q, the client updates the user feature matrix P i and the item feature matrix Q, and only uploads the item feature matrix to the central server; for the update of the user feature vector, the gradient descent method is used to update locally at the client, and the update formula is as follows:

[0069]

[0070]

[0071] Let γ = 2α, Then:

[0072] p iu ← p iu + γ(e un q n - λ piu )#(28)

[0073] For the update of the item feature vector, after local update training by each client, it is sent to the central server for aggregation; according to Equation (24) and the learning rate γ, the update formula of the client item feature vector is obtained:

[0074]

[0075] After the t-th round of communication, the formula for the central server to update the project feature vector is shown in Equation (30):

[0076]

[0077] The specific algorithm of the central server is as follows:

[0078] The central server initializes the global project feature matrix Q. In any iteration of round t ∈ T, the central server sends the global project feature vector to the client The client participating in the federated training receives the global project feature vector After that, calculate Each client then performs local updates according to the client algorithm and uploads the parameters to the central server; after receiving the parameters from each client, the central server aggregates them and derives the iterative formula for the global project feature vector according to Equation (30), as shown in Equation (31):

[0079]

[0080] The specific algorithm of the client is as follows:

[0081] The client participating in the federated training randomly samples a sample r according to the sampling ratio α i = αR i , and receives the global project feature vector before the start of each round of global iteration Then calculate Derive the update formula for the user feature vector p according to Equation (28), as shown in Equation (32), update the project feature vector q according to Equation (29) iu , update the dual variable u according to Equation (33) in , according to Equation (33); after local training for E times, obtain the parameters p i , q iu , and u in , and then send q i , u in , and R i (θ i ) to the central server t-1 .

[0082]

[0083]

[0084] The beneficial effects of the present invention are:

[0085] The present invention applies the federated learning algorithm to the recommendation system, combines this algorithm with the matrix factorization recommendation model, and designs a communication-efficient federated matrix factorization recommendation algorithm CSOFMF. This method can reduce the training time, accelerate the algorithm convergence speed, and thus reduce the communication cost on the premise of ensuring the recommendation effect of the model. Description of the Drawings

[0086] Figure 1 It is a flowchart of a centralized recommendation system;

[0087] Figure 2 It is a schematic diagram of matrix factorization in an embodiment of the present invention;

[0088] Figure 3 It is a schematic diagram of the average MSE of the client model on the Moivelens-1M dataset in an embodiment of the present invention;

[0089] Figure 4 It is a schematic diagram of the average MSE of the client model on the Epinions dataset in an embodiment of the present invention. Detailed Embodiment

[0090] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0091] Embodiment

[0092] A communication-efficient federated matrix factorization recommendation method includes:

[0093] 1. Federated Recommendation Framework:

[0094] For protecting the privacy and security of users in federated learning, all participating parties train the dataset locally and only transmit the parameters obtained from model training to the server for aggregation. In the federated recommendation system, assume that there are N clients participating in federated learning. There are m users in total on the clients. The client set is C = {C 1 , C 2 , ···, C N}, where the client C i has m i users. All clients have the same n items. The data sets owned by each client are R = {R 1 , R 2 , ···, R N}, where R i is the local data owned by the client C i , and R i is an m i ×n-dimensional user-item rating matrix. The global objective function is defined as:

[0095]

[0096] Among them, represents the weight of each client, and F i (θ) represents the local objective function of client C i . The formula for F i (θ) is as follows:

[0097]

[0098] Among them, l is the loss function, and h(x j ; θ) represents the prediction of the model with parameter θ for the instance x j in the dataset.

[0099] In federated learning, it is usually desired to obtain a global model θ' with optimal performance. Therefore, the goal of the central server is to optimize the minimum objective function F R (θ), and then obtain a set of model parameters θ' that gives the best optimization effect as the globally trained model parameters. Its formulaic representation is:

[0100]

[0101] The federated matrix factorization algorithm is a widely used algorithm in federated recommendation systems. Using the method of matrix factorization can, to a certain extent, make up for the deficiency of the collaborative filtering model in dealing with sparse matrices. The matrix factorization model converts a high-dimensional rating matrix into the product of low-dimensional feature matrices, mainly including two parts of parameters. One part of the parameters represents the latent vector of users, and the other part represents the latent vector of items. The specific method is as Figure 2 shown.

[0102] Decompose the rating matrix R u×n to obtain two matrices P u×k and Q k×n . Among them, P u×k is the user feature matrix, Q k×n is the item feature matrix, k is the dimension of the latent vector, and the size of k determines the strength of the latent vector's expressive ability. Usually, a relatively small value is selected. The user feature vector reflects the user's interests, while the item feature vector reflects the characteristics of the item. The inner product of the two vectors represents the degree of the user's preference for the item. For the rating i of user u for a certain item n in client C , the prediction formula is as follows:

[0103]

[0104] Among them, μ is the global mean bias value set according to the global rating data, and p iu represents client C iThe user feature matrix P i The u-th row user feature vector of denotes p iu vector transpose, q n denotes the n-th row item feature vector of Q.

[0105] The global objective function is defined as:

[0106]

[0107] where λ is the regularization coefficient, and λ(||p iu || 2 +||q n || 2 ) is the regularization term to avoid overfitting the observed data. The federated matrix factorization recommendation algorithm generates an m i ×n dimensional predicted rating matrix R' i for client C i without directly sharing its own dataset among all clients, and then recommends suitable items for the users of client C i . The federated matrix factorization recommendation algorithm first randomly generates an m i ×k dimensional user feature matrix P i at client C i , randomly generates an n×k dimensional global item feature matrix Q at the recommendation server, and then generates an accurate predicted rating matrix by minimizing the global objective function, and further makes recommendations for users.

[0108] 2. Construction of the surrogate objective function:

[0109] Construct a surrogate objective function to replace the global objective function F R (θ). Performing a Taylor expansion on F R (θ) can obtain an infinite series as follows:

[0110]

[0111] where, is an arbitrary initial estimate of θ, is a computable constant, and <·,·> represents the inner product. Next, randomly extract a small sample dataset from each client's local dataset, construct a sample objective function, and use the high-order derivative part in the sample objective function to replace the high-order derivative part of the global objective function. The specific steps are as follows:

[0112] Step 1: Set a sampling ratio α, and randomly extract a size of r i =αR iA subset that forms the sampling sample set r = {r 1 , r 2 , ···, r N}. Assume that r is much smaller than R, i.e., and when R → ∞, r → ∞. The objective function based on the sampling samples can be expressed as:

[0113]

[0114] Perform a Taylor expansion on F r (θ) and transform it into the following form:

[0115]

[0116] Step 2: Use the high-order derivative part of the objective function based on the sampling samples to replace the high-order derivative in the global objective function That is, let:

[0117]

[0118] Substitute Equation (9) into Equation (6), and the approximation of the global objective function F R (θ) can be obtained:

[0119]

[0120] Ignore the constant term and simplify it to obtain the communication-efficient surrogate objective function:

[0121]

[0122] Solving for the optimal parameters is expressed as minimizing the global objective function, i.e.:

[0123]

[0124] 3. Alternating direction method of multipliers:

[0125] Use the alternating direction method of multipliers (ADMM) to solve Problem (12). The advantage of the ADMM algorithm is that it can decompose a large-scale optimization problem into multiple small-scale sub-problems and can solve these sub-problems distributively. Therefore, minimizing the surrogate objective function can be decomposed into minimizing the local objective functions of each client. Let Rewrite Problem (12) in the following form:

[0126]

[0127] where represents the weights of each client, and represents the i-th client based on the sampling sample ri The local objective function. Problem (13) can be further expressed as the following global consistency optimization problem:

[0128]

[0129] where θ i represents the local model parameters of the i-th client. Under the constraint θ i = θ, i ∈ [1, 2, ···, N], problem (14) is equivalent to problem (13). Transform problem (14) into the form required by the ADMM algorithm:

[0130]

[0131] where B represents the N×N identity matrix, represents the set of all θ i , that is The specific definition of is:

[0132]

[0133] The augmented Lagrangian function corresponding to problem (15) is as follows:

[0134]

[0135] where ρ > 0 is the penalty parameter, and the dual variable u ∈ [u 1 , u 2 , ···, u N . Next, the update rules of the ADMM algorithm are given:

[0136]

[0137]

[0138]

[0139] For problem (18), this problem is only related to client C i , and it can be rewritten in the following form:

[0140]

[0141] For problem (21), a closed-form solution is usually not allowed, which will increase the computational cost. To accelerate the local computation speed of the client, the linearized ADMM algorithm is used to simplify and solve this formula. For at using the second-order Taylor expansion, the expansion of can be obtained:

[0142]

[0143] where η is the penalty term coefficient. Then The update formula of

[0144]

[0145] Equation (23) is an approximation of Equation (21). By solving Equation (23), an inexact solution of Equation (21) can be obtained:

[0146]

[0147] For problem (19), this is a quadratic function of θ. By taking the partial derivative of it and setting the partial derivative equal to zero, the minimum point can be obtained, that is:

[0148]

[0149] 4. Design of the CSOFMF algorithm:

[0150] According to the above derivation, the update formula of the parameters of the federated matrix factorization model can be obtained. Let Before each round of federated training, first update the surrogate objective function, that is Subsequently, update the parameters according to the formula.

[0151] The federated matrix factorization algorithm contains two parameters, namely the user feature matrix P i and the item feature matrix Q. The client initializes the user feature matrix P locally i , and the central server initializes the global item feature matrix Q and sends it to the client. After receiving the global item feature matrix Q, the client updates the user feature matrix P i and the item feature matrix Q locally, but only uploads the item feature matrix to the central server. Since the user feature vector involves user privacy data, for the update of the user feature vector, the gradient descent method is used to update it locally on the client side. The update formula is as follows:

[0152]

[0153]

[0154] Let γ = 2α, Then

[0155] p iu ← p iu + γ(e un q n - λp iu ) #(28

[0156] For the update of the item feature vectors, after local update and training on each client, they are sent to the central server for aggregation. According to Equation (24) and the learning rate γ, the update formula for the client item feature vectors can be obtained, as shown in Equation (29):

[0157]

[0158] After the t-th round of communication, the formula for the central server to update the item feature vectors is as shown in Equation (30):

[0159]

[0160] (1) Central Server Algorithm Design

[0161] The training process of the central server in the CSOFMF algorithm can be mainly divided into the following steps:

[0162] Step 1: The central server randomly initializes the global item feature matrix Q.

[0163] Step 2: The central server sends the global item feature matrix to the clients.

[0164] Step 3: The clients randomly sample the sample data according to the sampling ratio.

[0165] Step 4: The clients perform local training according to Algorithm 4 and upload the parameters.

[0166] Step 5: The central server receives the parameters uploaded by the clients.

[0167] Step 6: Aggregate the client parameters and update the global item feature vectors according to the formula.

[0168] Step 7: Repeat Steps 2 to 6 until the global communication round reaches the specified value.

[0169] The pseudo-code of the central server's algorithm is shown in Algorithm 1. The central server initializes the global item feature matrix Q. In any iteration of t ∈ T, the central server sends the global item feature vectors to the clients The clients participating in the federated training receive the global item feature vectors After that, calculate Each client then performs local updates according to Algorithm 4 and uploads the parameters to the central server. After the central server receives the parameters of each client, it aggregates them. According to Equation (4.30), the iterative formula for the global item feature vectors can be derived, as shown in Equation (31):

[0170]

[0171] Table 1 Pseudo-code of the CSOFMF Server-side Algorithm

[0172]

[0173]

[0174] (2) Client-side Algorithm Design

[0175] The client training process of the CSOFMF algorithm can be mainly divided into the following steps:

[0176] Step 1: Load the dataset, determine the local update times E and the sampling ratio α.

[0177] Step 2: The client randomly extracts a sample set of size r i = αR i from the local dataset.

[0178] Step 3: Receive the global item feature vector and calculate R i (θ t-1 ).

[0179] Step 4: Update the local model parameters according to the learning rate γ until the current training round reaches E.

[0180] Step 5: Send the parameters q in , u i and R i (θ t-1 ) to the central server.

[0181] The pseudo-code of the client algorithm is shown in Algorithm 2. The clients participating in the federated training randomly extract samples r i = αR i according to the sampling ratio α. Before each global iteration, receive the global item feature vector and then calculate According to Equation (28), the update formula of the user feature vector p iu can be derived, as shown in Equation (32). Update the item feature vector q in according to Equation (29), and update the dual variable u i . After local training E times, obtain the parameters p iu , q in and u i , and then send q in , u i and R i (θ t-1 ) to the central server.

[0182]

[0183]

[0184] Table 2 Pseudo-code of the CSOFMF Client Algorithm

[0185]

[0186]

[0187] The following further illustrates this embodiment through experiments:

[0188] To verify the feasibility and effectiveness of the CSOFMF algorithm, an experiment was designed to compare the recommendation system based on the traditional federated learning algorithm FedAvg with the recommendation system based on the CSOFMF algorithm. The specific experimental settings are as follows.

[0189] 1. Experimental Environment:

[0190] The experiment simulated a federated learning environment with 10 clients and a central server. All experiments were run on a single computer. The software and hardware environments for system operation and testing are shown in Table 3.

[0191] Table 3 Physical Machine System Hardware and Software Environment

[0192]

[0193] 2. Experimental Data:

[0194] To verify the feasibility of the CSOFMF algorithm, the Movielens-1M dataset and the Epinions dataset were selected for a series of simulation experiments. The basic information of the two datasets is shown in Table 4.

[0195] Table 4 Experimental Datasets

[0196]

[0197] The Movielens-1M dataset was collected by the GroupLens research group at the University of Minnesota and is a publicly available movie rating dataset for testing recommendation systems. The ratings are represented as integers from 1 to 5, where 5 might indicate that the user really likes the movie and 1 might indicate that the user dislikes it. The Movielens-1M dataset contains 1 million rating records of 3,900 movies by 6,040 users. The Movielens-1M dataset also includes detailed information about the movies, such as the movie titles, genres, release years, etc. This enables researchers to conduct in-depth research in aspects such as evaluating the performance of recommendation algorithms, performing user behavior analysis, and exploring the field of personalized recommendations. Due to its large scale and rich content, the Movielens-1M dataset has become one of the important benchmark datasets in the fields of recommendation systems and machine learning.

[0198] The Epinions dataset is a large-scale online community dataset covering user-generated content and social network information, and is widely used in research fields such as recommendation systems, social network analysis, and information mining. This dataset consists of evaluations and relevance information generated by users on the Epinions.com platform, including users' evaluations of products, services, and other users, as well as the social relationships between users. The dataset contains evaluations of hundreds of thousands of products by more than 5,000 users, and the evaluations use an integer range from 1 to 5. Each user has evaluated at least 20 products, providing researchers with a large amount of user behavior data.

[0199] Table 5 Data distribution of each client in the Movielens-1M dataset

[0200]

[0201] Table 6 Data distribution of each client after partitioning the Epinions dataset

[0202]

[0203]

[0204] The federated learning framework of this embodiment is a horizontal federated learning framework, which is used to simulate the personalized needs of different users when the projects or items are the same. For effective training and evaluation, it is assumed that this data is distributed among 10 clients, and the overall dataset is randomly divided into 10 parts from the perspective of users. 70% of the data of each client is randomly selected as training data, and the remaining 30% is used as test data. The data volumes of each client in the Movielens-1M dataset and the Epinions dataset after partitioning are shown in Table 5 and Table 6.

[0205] 3. Evaluation metrics:

[0206] The simulation experiment selects the mean square error and the algorithm training time as evaluation metrics.

[0207] (1) Mean Square Error (MSE)

[0208] The mean square error is a metric for measuring the deviation between the true value and the model prediction value, and is mostly used in regression tasks. Its formula is shown in Equation (4.34):

[0209]

[0210] Among them, is the predicted output, y i is the true data label, and m is the number of data samples participating in the calculation.

[0211] (2) Training Time (Time)

[0212] To reflect the improvement of the algorithm in communication, the training time is used as an indicator to evaluate the performance and efficiency of the model. The shorter the training time, the better the algorithm effect.

[0213] 4. Comparative experiment:

[0214] To verify the performance of the efficient communication federated recommendation algorithm CSOFMF proposed in this embodiment, the traditional federated learning algorithm FedAvg is combined with the matrix factorization model for comparison. The client updates the user feature vector p iu and the item gradient vector g i locally. The update formulas for p iu and g i are shown in Equations (35) and (36) respectively:

[0215] p iu ← p iu + γ(e ui q n - λp iu ) #(35)

[0216] g i = γ(e ui p iu - λq n ) #(36)

[0217] After the client performs E rounds of updates, it sends the item gradient vector g i to the central server. After receiving the gradient, the central server aggregates it and then updates the item feature vector q i according to Equation (37). Finally, the client generates the prediction score matrix R based on the user feature vector p iu and the item feature vector q n ​i , provide personalized recommendations for users according to the scoring matrix.

[0218]

[0219] 5. Parameter setting:

[0220] The simulation experiment sets different hyperparameters for the Moivelens-1M dataset and the Epinions dataset respectively. The specific parameter settings are shown in Table 7.

[0221] Table 7 Hyperparameter settings for the simulation experiment

[0222]

[0223] 5. Experimental results and analysis:

[0224] The simulation experiment compares the proposed efficient communication federated recommendation algorithm CSOFMF with the recommendation algorithm based on FedAvg in this embodiment. To avoid accidental errors, all experiments are repeated 5 times, and the average value of the 5 experiments is taken as the final experimental result.

[0225] To verify the recommendation effect of the recommendation system based on CSOFMF, this embodiment conducts a control experiment using two methods on the Moivelens-1M dataset and the Epinions dataset. Set the sampling ratio α = 0.005. For the Moivelens-1M dataset, set the global communication round T to 100, the local update times E to 20, the learning rate γ to 0.001, the regularization coefficient λ to 0.01, and the penalty parameter ρ to 1. For the Epinions dataset, set the global communication round T to 50, the local update times E to 20, the learning rate γ to 0.005, the regularization coefficient λ to 0.01, and the penalty parameter ρ to 1. The experimental results are as Figure 3 and Figure 4 shown. The abscissa in the figure represents the number of communications, and the ordinate represents the average mean square error of 10 client models.

[0226] It can be seen from the experimental results that whether it is on the Moivelens-1M dataset or the Epinions dataset, the algorithm proposed in this embodiment can achieve better model accuracy within fewer communication rounds compared with the baseline method, and the average mean square error of the client model is lower. This shows that the algorithm proposed in this embodiment has a faster convergence speed and can achieve an ideal performance level while significantly reducing the communication time.

[0227] After various training models reach the specified number of iterations, the average accuracy (Accuracy) and training time (Time, in hours) of the client models are shown in Table 8.

[0228] Table 8 Average accuracy and training time of client models under various training modes

[0229]

[0230] As can be seen from Table 8, in the Moivelens-1M dataset, the average accuracy of the client model of the recommendation system based on FedAvg is 57.34%, and the training time is 2.132 hours. While the accuracy of the final model of the recommendation system based on CSOFMF is 58.11%, and the training time is 1.773 hours. Compared with FedAvg, the average accuracy of the client model based on the CSOFMF algorithm has increased by 1.34%, and the training time has been reduced by 16.84%. In the Epinions dataset, the average accuracy of the client model of the recommendation system based on FedAvg is 43.89%, and the training time is 1.692 hours. While the average accuracy of the client model of the recommendation system based on CSOFMF is 44.19%, and the training time is 1.388 hours. Compared with FedAvg, the accuracy of the final model of the recommendation system based on CSOFMF has increased by 0.68%, and the training time has been reduced by 17.97%. The experimental results verify that the recommendation system based on the CSOFMF algorithm can reduce the training time, accelerate the algorithm convergence speed, thereby reducing the communication cost, and the accuracy of the client model has been improved.

[0231] To further analyze the performance of the CSOFMF algorithm, the average mean square error of the client model when FedAvg is close to the convergence state is used as the target mean square error. According to Figure 3 and Figure 4 's experimental results, it can be found that the average mean square error of the client model in the Moivelens-1M dataset is approximately 0.85 when convergence is reached, and the average mean square error of the client model in the Epinions dataset is approximately 1.30 when convergence is reached. To compare the convergence speeds of different algorithms, the target mean square errors of the Moivelens-1M dataset and the Epinions dataset are set to 0.85 and 1.30 respectively. The global communication rounds (T) and training time (Time, in hours) required for various training modes to reach the target mean square error are shown in Table 9.

[0232] Table 9 Communication rounds and training time required for various training modes to reach the target mean square error

[0233]

[0234] As can be seen from Table 9, the CSOFMF algorithm can achieve the target mean squared error of each dataset with fewer communication rounds, and the training time is also significantly reduced. Specifically, when reaching the target mean squared error, for the Moivelens-1M dataset, FedAvg needs to communicate 84 times with a training time of 1.789 hours, while CSOFMF needs to communicate 67 times with a training time of 1.192 hours. Compared with FedAvg, when the client model achieves the same recommendation effect, the communication times of CSOFMF are reduced by 20.24%, and the training time is reduced by 33.37%. For the Epinions dataset, FedAvg needs to communicate 44 times with a training time of 1.489 hours, while CSOFMF only needs to communicate 32 times with a training time of 0.971 hours. Compared with FedAvg, when the client model achieves the same recommendation effect, the communication times of CSOFMF are reduced by 27.27%, and the training time is reduced by 34.79%. In federated learning, the more communication rounds mean higher communication costs. For some clients with limited budgets, they may not be able to conduct multiple rounds of communication, which will have a negative impact on the global model. When achieving the same model effect, the CSOFMF algorithm can not only reduce the communication rounds but also significantly reduce the training time. The experimental results fully demonstrate the feasibility and effectiveness of the new method.

[0235] The above-described embodiments only represent the specific implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A communication efficient federated matrix factorization recommendation method, characterized in that: The following steps are involved: Step 1: Construct a federated matrix recommendation framework and generate a federated matrix decomposition recommendation algorithm; The step 1 is specifically as follows: All participants train the data set locally and only pass the parameters obtained from model training to the server for aggregation. Assume that there are N clients participating in federated learning, and there are m clients in total. The client set is C = {C1, C2, ···, C N }, Client C i Own m i users, all clients have the same n projects, and the data set owned by each client is R = {R1, R2, ···, R N }, where R i It is client C i Owned local data, R i is m i ×n-dimensional user-item rating matrix, the global objective function is defined as: in, represents the weight of each client, and F i (θ) represents the client C i The local objective function, F i The formula for (θ) is as follows: Among them, l is the loss function, h(x j ; θ) represents the model with parameter θ for instance x in the dataset j predictions; In federated learning, we want to obtain a global model θ' with optimal performance. The goal of the central server is to optimize the minimum objective function F R (θ), and then obtain a set of model parameters θ' that makes it the best optimization effect as the global model parameters obtained by the final training, which is formulated as: The rating matrix R u×n Decompose it and get P u×k and Q k×n Two matrices, where P u×k is the user feature matrix, Q k×n is the item feature matrix, k is the dimension of the latent vector; For client C i The rating of user u on item n The prediction formula is as follows: Among them, μ is the global average bias value set according to the global scoring data, p iu Represents client C i The user feature matrix P i The u-th row user feature vector of Indicates p iu Vector transpose, q n represents the feature vector of the nth row item of Q; The global objective function is defined as: Among them, λ is the regularization coefficient, λ(||p iu || 2 +||q n || 2 ) is the regularization term; The federated matrix decomposition recommendation algorithm is based on the premise that all clients do not directly share their own data sets. i Generate m i ×n-dimensional predicted rating matrix R' i , and then according to the prediction score matrix for client C i Recommend suitable items to users; the federated matrix decomposition recommendation algorithm first i Randomly generate m i ×k-dimensional user feature matrix P i , randomly generates a global item feature matrix Q of n×k dimensions on the recommendation server, and then generates an accurate prediction score matrix by minimizing the global objective function, and then makes recommendations for users; Step 2: construct a proxy objective function, and use the proxy objective function to replace the global objective function; Step 3, use the alternating direction multiplier method ADMM to solve the parameter update formula in the agent objective function; Step 4: Design a federated matrix decomposition recommendation algorithm, including designing a central server algorithm and designing a client algorithm.

2. The communication efficient federated matrix decomposition recommendation method according to claim 1, characterized in that: The step 2 is specifically as follows: Constructing a surrogate objective function To replace the global objective function F R (θ); for F R Taylor expansion of (θ) can yield an infinite series: in, is an arbitrary initial guess for θ, is a calculable constant, and <·,·> represents the inner product. Next, a small sample data set is randomly selected from the local data set of each client, a sample objective function is constructed, and the high-order derivative part of the sample objective function is used to replace the high-order derivative part of the global objective function. The specific steps are as follows: Step 2.1: Set a sampling ratio α and randomly extract a sample of size r from each client's local data set. i =αR i The subset of the sample set r = {r1, r2, ···, r N }; Assume that r is much smaller than R, that is And when R→∞, r→∞; the objective function based on the sampling sample is expressed as: F r (θ) is transformed into the following form by Taylor expansion: Step 2.2: Use higher-order derivatives of the objective function based on the sampled samples Replacing higher-order derivatives in the global objective function That is: Substituting equation (9) into equation (6), we can get the global objective function F R Approximation of (θ): Ignoring the constant term and simplifying, we get the communication efficient proxy objective function: Solving the optimal parameters is expressed as minimizing the global objective function, that is:

3. The communication efficient federated matrix decomposition recommendation method according to claim 2, characterized in that: The step 3 is as follows: The alternating direction multiplier method ADMM is used to solve equation (12), and the minimization of the proxy objective function is decomposed into minimizing the local objective function of each client; let Rewrite equation (12) as follows: in, represents the weight of each client, and Indicates that the i-th client is based on the sample r i The local objective function of ; Formula (13) is further expressed as the following global consistency optimization problem: Among them, θ i represents the local model parameters of the i-th client; under the constraint θ i =θ,i∈[1,2,···,N], formula (14) is equivalent to formula (13); transform formula (14) into the form required by the ADMM algorithm: Among them, B represents the N-block identity matrix, represents all θ i The collection of The specific definition of is: The augmented Lagrangian function corresponding to formula (15) is as follows: Among them, ρ>0 is the penalty parameter, and the dual variable u∈[u1,u2,···,u N ]; Next, the update rules of the ADMM algorithm are given: For formula (18), this problem is only related to client C i Rewrite it as follows: The linearized ADMM algorithm is used to simplify the solution of equation (21); exist Using the second-order Taylor expansion, we can get Extensions: Where η is the penalty term coefficient; The update formula can be written as: Formula (23) is an approximation of formula (21). By solving formula (23), we can get an inexact solution of formula (21): For formula (19), it is a quadratic function about θ. By taking its partial derivative and setting it equal to zero, we can find the minimum point, that is:

4. The communication efficient federated matrix decomposition recommendation method according to claim 3, characterized in that: The step 4 is specifically as follows: make Before each round of federated training begins, the proxy objective function is first updated, namely Then update the parameters according to the formula; The federated matrix decomposition algorithm contains two parameters, namely, the user feature matrix P i And the project feature matrix Q, the client initializes the user feature matrix P locally i , the central server initializes the global item feature matrix Q and sends it to the client; the client receives the global item feature matrix Q and updates the user feature matrix P locally i And the project feature matrix Q, only upload the project feature matrix to the central server; for the update of the user feature vector, the gradient descent method is used to update it locally on the client, and the update formula is as follows: Let γ = 2α, Then: p iu ←p iu +γ(e un q n -λp iu )#(28) For the update of the project feature vector, each client sends the local update training to the central server for aggregation. According to formula (24) and the learning rate γ, the update formula of the client project feature vector is obtained: After the tth round of communication, the formula for the central server to update the item feature vector is shown in formula (30): The central server algorithm is as follows: The central server initializes the global item feature matrix Q. In any iteration of t∈T, the central server sends the global item feature vector Clients participating in federated training receive the global project feature vector Afterwards, calculate Each client then performs local updates according to the client algorithm and uploads the parameters to the central server. After receiving the parameters of each client, the central server aggregates them and derives the iterative formula of the global project feature vector according to formula (30), as shown in formula (31): The client algorithm is as follows: The client participating in the federated training randomly selects samples r according to the sampling ratio α i =αR i , before each global iteration begins, receive the global project feature vector Then calculate According to formula (28), the user feature vector p is derived iu The update formula of is shown in formula (32). According to formula (29), the item feature vector q is updated in , update the dual variable u according to formula (33) i ; After local training E times, the parameter p is obtained iu ,q in and u i , then q in 、u i and R i (θ t-1 ) is sent to the central server;

Citation Information

Patent Citations

  • Personalized social recommendation method based on federated matrix decomposition

    CN115374953A

  • Recommendation method and system based on improved federated matrix decomposition

    CN116304321A