A cross-region and cross-rating collaborative filtering recommendation method and system
By using the Funk-SVD model and deep regression network to extract the hidden vectors of users and projects and construct a constrained matrix decomposition model in the collaborative filtering recommendation system, the problem of sparse data and differences in score density affecting the recommendation effect is solved, and more accurate and efficient recommendation results are achieved.
Patent Information
- Application Number
- CN202210021494.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-10
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-01-10
AI Technical Summary
The existing collaborative filtering recommendation algorithm is difficult to effectively capture user preferences when user feedback data is sparse, resulting in poor recommendation results. Especially when there is heterogeneity between rating scores and binary scores, the transfer learning model cannot effectively utilize the difference in score density, affecting the accuracy of score predictions.
By forming the user-item scoring data of the target domain and the source domain, the matrix decomposition is performed based on the Funk-SVD model, the hidden vectors of the user and the project are extracted. Then, the hidden vector mapping relationship between active users and popular projects is learned using deep regression network, generalize to inactive users and non-popular projects, and build a constrained matrix decomposition model for scoring prediction.
Through fine-grained scoring prediction strategies, we can improve the performance of recommendations, effectively utilize data from inactive users and non-popular projects, improve the accuracy of mapping relationship modeling, avoid negative transfer phenomena in transfer learning, and improve the recommendation effect.
Smart Images

Figure CN114329233B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of collaborative filtering recommendation methods, and in particular, relates to a cross-region and cross-scoring collaborative filtering recommendation method and system. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] The traditional collaborative filtering recommendation algorithm is an important means to solve the problem of information overload in the era of big data. The main idea of the algorithm is to learn user preferences based on historical user feedback data, provide personalized services to users, and improve user satisfaction and platform business revenue. However, when user feedback data is very sparse, the collaborative filtering algorithm often cannot effectively capture user preferences. Data sparsity will cause serious overfitting of the recommendation algorithm, affecting the performance of the recommendation algorithm. This data sparsity phenomenon is often more obvious when the user feedback data is the 1-5 rating information that most recommendation algorithms rely on. In order to alleviate the problem of data sparsity, the idea of transfer learning is applied to the recommendation system to extract and transfer knowledge from the information in the source domain to the target domain to improve the recommendation effect of the target domain.
[0004] Migrating the user's information on dense binary ratings to the target domain can alleviate the problem of data sparsity in the target domain and effectively improve the recommendation effect in the target domain. Although there are some collaborative filtering recommendation models based on transfer learning in other scenarios, on the one hand, these models cannot well consider the heterogeneity between grade ratings and binary ratings. Directly migrating binary ratings to grade ratings may lead to negative knowledge transfer, resulting in poor recommendation effects. On the other hand, the regions in the target domain composed of rating data of different users on different items have different numerical rating densities. High-density regions have rich feedback information and less reliance on source domain information, while low-density regions have scarce feedback information and greater reliance on source domain information. Existing models often assume that all regional numerical ratings in the recommendation system are relatively sparse, and adopt a consistent rating prediction strategy for different regions, ignoring the impact of rating density on the accuracy of solving user and item latent vectors, resulting in inaccurate rating prediction in sparse rating regions. Summary of the invention
[0005] In order to solve at least one technical problem existing in the above-mentioned background technology, the present invention provides a cross-regional and cross-rating collaborative filtering recommendation method and system, which respectively composes the target domain and source domain user-project rating data into a target domain rating matrix and a source domain rating matrix, sorts the users and projects in the target domain rating matrix according to the number of ratings, divides all users into active users and inactive users according to the threshold, and divides all projects into popular projects and non-popular projects. Then, the target domain and source domain rating matrices are respectively matrix decomposed based on the Funk-SVD model to extract the latent vectors of users and projects in the target domain and the source domain. Secondly, for active users and popular projects, a deep regression network based on self-teaching learning is constructed to learn the mapping relationship between the user latent vector and the project latent vector corresponding to the two ratings in the target domain and the source domain. Then, the mapping relationship between the latent vectors of active users and popular projects is generalized to the inactive users and non-popular projects in the target domain, and the latent vectors of inactive users and non-popular projects in the auxiliary domain are used to derive their latent vectors in the target domain. Finally, the latent vectors of inactive users and non-popular items in the target domain are used as constraints to solve the restricted matrix decomposition model and give the corresponding recommendation results.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] A first aspect of the present invention provides a cross-region and cross-rating collaborative filtering recommendation method, comprising the following steps:
[0008] Obtain user-item rating data from the target domain and source domain, and obtain the target domain rating matrix and the source domain rating matrix after preprocessing;
[0009] Divide all users in the target domain rating matrix and the source domain rating matrix into active users and inactive users, and divide all projects into popular projects and non-popular projects;
[0010] Decompose the target domain rating matrix and the source domain rating matrix based on the latent semantic model, and extract the user latent vector and item latent vector in the target domain and the source domain;
[0011] For active users and popular items, based on the trained deep regression network, we learn the mapping relationship between the user latent vector and the item latent vector corresponding to the target domain and the source domain under the two rating systems.
[0012] The mapping relationship between the user latent vectors and item latent vectors of active users and popular items is used to obtain the features of inactive users and non-popular items in the target domain;
[0013] According to the characteristics of inactive users and non-popular items in the target domain, a restricted matrix decomposition model is constructed to predict the score of any user for any item, and the item with the highest predicted score is selected as the user's recommendation result.
[0014] A second aspect of the present invention provides a cross-region and cross-rating collaborative filtering recommendation system, comprising:
[0015] The data preprocessing module is configured to: obtain user-item rating data of the target domain and the source domain, and obtain a target domain rating matrix and a source domain rating matrix after preprocessing;
[0016] Divide all users in the target domain rating matrix and the source domain rating matrix into active users and inactive users, and divide all projects into popular projects and non-popular projects;
[0017] The feature extraction module is configured to: decompose the target domain scoring matrix and the source domain scoring matrix based on the latent semantic model, and extract the user latent vector and the item latent vector in the target domain and the source domain;
[0018] For active users and popular items, based on the trained deep regression network, we learn the mapping relationship between the user latent vector and the item latent vector corresponding to the target domain and the source domain under the two rating systems.
[0019] The mapping relationship between the user latent vectors and item latent vectors of active users and popular items is used to obtain the features of inactive users and non-popular items in the target domain;
[0020] The recommendation acquisition module is configured to: construct a restricted matrix decomposition model based on the characteristics of inactive users and non-popular items in the target domain, predict the score of any user for any item, and select the item with the highest predicted score as the user's recommendation result.
[0021] A third aspect of the present invention provides a computer-readable storage medium.
[0022] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the cross-region and cross-rating collaborative filtering recommendation method as described above.
[0023] A fourth aspect of the present invention provides a computer device.
[0024] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the cross-region and cross-rating collaborative filtering recommendation method as described above are implemented.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] The present invention defines active users and inactive users, popular projects and non-popular projects, divides the target domain rating matrix into four areas with different densities according to the active users and inactive users, popular projects and non-popular projects, and adopts different rating prediction strategies for areas with different rating densities in the rating matrix to perform fine-grained and accurate recommendations and improve the recommendation performance. A deep regression network based on self-teaching learning is proposed to learn the mapping relationship between the latent vectors corresponding to active users and popular projects in the target domain and the auxiliary domain, which can make full use of a large amount of unsupervised data related to inactive users and non-popular projects to improve the accuracy of mapping relationship modeling.
[0027] The present invention proposes a constrained matrix decomposition model to effectively fuse the sparse numerical scores in the target domain and the binary scores in the auxiliary domain, and effectively avoid the negative transfer phenomenon in transfer learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0029] Figure 1 It is a flow chart of the cross-region and cross-rating collaborative filtering recommendation method;
[0030] Figure 2 It is a large sparse rating matrix consisting of the rating data of all users on all items;
[0031] Figure 3 It is a schematic diagram of data arrangement;
[0032] Figure 4 This is a schematic diagram of regression model pre-training;
[0033] Figure 5 It is the regression model fine-tuning box diagram; DETAILED DESCRIPTION
[0034] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0035] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0036] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0037] Terminology explanation:
[0038] Cross-region: regions with high score density and regions with low score density in the score matrix.
[0039] Cross-rating: user's 1-5 rating and user's 0-1 (like / dislike, thumbs up / down) binary rating.
[0040] For recommendation platforms with two rating formats (for example, the movieplot movie website contains two rating formats: 1-10 rating and like / dislike binary rating), users are more inclined to make simple 1, 0 binary ratings (such as like / dislike, thumbs up / down) than ratings (1-5, 1-10 ratings). Therefore, in addition to the sparse 5-point numerical ratings (target domain), recommendation platforms often contain relatively rich 1, 0 binary ratings (source domain), that is, users' binary ratings are more dense than users' ratings. Since ratings have more fine-grained rating rules and can better reflect users' preference characteristics, using binary rating data (source domain) to assist ratings (target domain) for rating prediction can obtain more accurate user characteristics and generate more targeted recommendations compared to directly using binary rating data for rating prediction. In addition, it is observed that the number of ratings for users and projects in life often shows a long-tail distribution in distribution. Even if most users have a relatively small number of ratings, there are still a small number of users with more ratings. Similarly, even if a large number of non-popular projects have only a few ratings, there are still a small number of popular projects with more ratings. For example, in the famous public dataset MovieLens, it is easy to find a rating subset consisting of 100 active users and 200 popular movies, which has a relatively high rating density. Figure 2 As shown in FIG. 1 , in a sparse large rating matrix composed of rating data of all users on all items, there is still a dense small rating matrix composed of rating data of a small number of users with relatively high ratings on popular items.
[0041] Embodiment 1
[0042] like Figure 1 As shown, this embodiment provides a cross-region and cross-rating collaborative filtering recommendation method, including the following steps:
[0043] Step 1: Obtain user-item rating data in the target domain and source domain;
[0044] Step 2: The user-item rating data of the target domain and the source domain are combined into a target domain rating matrix and a source domain rating matrix;
[0045] Step 3: Sort the users and items in the target domain rating matrix according to the number of ratings; divide all users into active users and inactive users according to the threshold, and divide all items into popular items and non-popular items;
[0046] Step 4: Based on the latent semantic Funk-SVD model, the target domain and source domain rating matrices are decomposed to extract the latent vectors of users and items in the target domain and source domain;
[0047] Step 5: For active users and popular items, a deep regression network based on self-teaching learning is constructed to learn the mapping relationship between the user latent vector and the item latent vector corresponding to the two ratings in the target domain and the source domain respectively;
[0048] Step 6: The mapping relationship between the user latent vectors and the item latent vectors of active users and popular items is obtained from the target domain and generalized to the inactive users and non-popular items in the target domain. The latent vectors of inactive users and non-popular items in the auxiliary domain are used to derive their latent vectors in the target domain.
[0049] Step 7: Based on the latent vectors of inactive users and non-popular items in the target domain, a restricted matrix factorization model is constructed to predict the score of any user for any item, and the item with the highest predicted score is selected as the user's recommendation result.
[0050] In step 2, if Figure 3 As shown in the figure, it is organized according to the cross-region recommendation scenario. (5) is the target domain data, R (2) is the auxiliary domain data, R (5) and R (2) Share the same user set U and item set I.
[0051] Among them, R (5) A 5-point rating matrix (1-5 points) can be used, R (2) A binary (1 / 0, i.e. like / dislike) rating matrix may be used.
[0052] Among them, the number of ratings in each row of the rating matrix is the number of ratings received by the user, and the number of ratings in each column of the rating matrix is the number of ratings received by the project.
[0053] In step 3, for the target domain rating matrix, users are sorted according to the number of user ratings, and users are arranged from top to bottom in the rating matrix in descending order of the number of ratings. Similarly, according to the number of project ratings, projects are arranged from left to right in the rating matrix in descending order of the number of ratings.
[0054] In this way, users with more scores are concentrated on the top of the scoring matrix, and projects with more scores are concentrated on the left side of the scoring matrix; users and projects in the source domain are arranged in the same order as in the target domain, and there is a one-to-one correspondence between users and projects in the target domain and the source domain.
[0055] like Figure 3 As shown, in order to facilitate the observation of R (5) and R (2) , we use the front and back slices to represent them respectively. (5) middle and Represents the set of active users and inactive users respectively. and represent the set of popular items and non-popular items respectively. So a (i) ,b (i) ,c (i) ,d (i) (i=5,2) represents the scoring areas consisting of active users and popular items, active users and non-popular items, inactive users and popular items, and inactive users and non-popular items in the target domain and the auxiliary domain, respectively.
[0056] Usually active users provide more ratings than inactive users, and popular items receive more ratings than non-popular items. Therefore, density (a (i) ) is relatively high, and density(d (i) ) <density(b (i) or c (i) ) <density(a (i) ), note density(b (i) ) and density(c (i) ) usually does not have an obvious size relationship, where i = 5, 2. In addition, compared with more complex numerical ratings, all users tend to be more inclined to give a binary rating of 1, 0. Therefore, compared with R (5) , it can be considered that R (2) Even d (2) All have high scoring density and meet density(R ( 5)<<density(d (2) )<density(R (2) ).
[0057] Since the rating density of different data sets is different, and active users and popular items are relative concepts, a clear definition is needed for how to divide active users and popular items. In the following, users and items are divided into active users and inactive users, popular items and non-popular items based on the number of ratings of users and items, so as to make more targeted recommendations.
[0058] The definitions of active users and inactive users are as follows:
[0059] For any user u∈U={u 1 ,u 2 ,…,u m}, let d u represents the number of ratings of user u in the target domain (i.e., the number of all items rated by user u). Users are sorted from large to small according to the number of ratings, and the top μ 1 % of users are active users, and the remaining users are inactive users; where μ 1 is a pre-set parameter, called the user activity threshold, μ 1 The optimal value of is determined through experiments.
[0060] The definitions of popular and non-popular projects are as follows:
[0061] For any item i∈I={i 1 ,i 2 ,…,i n}, let d i represents the number of ratings of target domain item i (i.e., the number of all users who have rated item i). The items are sorted from large to small according to the number of ratings, and the top μ 2 % of the items are regarded as popular items, and the remaining items are regarded as non-popular items; 2 It is called the item popularity threshold.
[0062] In step 4, the target domain and source domain rating matrices are decomposed based on the latent semantic Funk-SVD model to extract the latent vectors of users and items in the target domain and source domain. Specifically, the following steps are performed:
[0063] (1) Decompose the target domain rating matrix to extract the user latent vector p u and the item latent vector q i ;
[0064] Stochastic gradient descent is used to solve the following optimization problem to obtain the user hidden vector p corresponding to the rating matrix u and the item latent vector q i :
[0065]
[0066] Among them, D represents the score r ui is a set of (u,i) pairs, and λ is the regularization coefficient.
[0067] To avoid overfitting, we use cross-validation to determine the appropriate value of λ.
[0068] The stochastic gradient descent iteration formula is:
[0069] q i ←q i +γ(e ui p u -λq i )
[0070] p u ←p u +γ(e ui q i -λp u )
[0071] in, γ is the learning rate. Specifically, this real-time example uses and Represents the 5-point numerical scoring matrix R (5) Decomposed user and item latent vectors.
[0072] (2) Perform matrix decomposition on the source domain rating matrix to extract the user latent vector p u and the item latent vector q i ;
[0073] Since the auxiliary domain 0-1 rating prediction is more suitable to be regarded as a classification problem with 0, 1 labels, rather than a numerical rating regression problem.
[0074] This embodiment adopts an improved Funk-SVD model to extract auxiliary domain latent vector features, that is, a cross entropy loss function is used instead of a least squares loss function as the loss function of the model, thereby converting the regression problem into a classification problem.
[0075] Specifically, the following optimization problem is solved for auxiliary domain latent vector feature extraction:
[0076]
[0077] In the formula, D 0 Represents the score r on the auxiliary domain ui The corresponding (u,i) pair set, λ is the regularization coefficient.
[0078] Use stochastic gradient descent to solve the above optimization problem. The iterative formula is as follows:
[0079]
[0080]
[0081] Specifically, this embodiment uses and Represents the binary rating matrix R (2) Decomposed user and item latent vectors.
[0082] In step 5, for active users and popular items, a deep regression network based on self-teaching learning is constructed to learn the mapping relationship between the user latent vector and the item latent vector corresponding to the two ratings in the target domain and the source domain respectively; including:
[0083] The deep regression model is trained using the latent vectors of active users in the target domain and the source domain to construct the mapping relationship F between the latent vectors of active users in the source domain and the latent vectors in the target domain. 1 .
[0084] Similarly, the latent vectors of popular projects in the target domain and source domain are used to train the deep regression model, and the mapping relationship F between the source domain latent vector and the target domain latent vector of popular projects is constructed. 2 .
[0085] Since there are relatively rich ratings related to active users and popular items, it is helpful to solve relatively accurate latent vector features. In this embodiment, latent vector features are first calculated for active users and popular items, and then the latent vector mapping relationship corresponding to the two ratings of active users and popular items is modeled.
[0086] make and Represents the 5-point rating matrix R (5) The corresponding active user u a and popular projects p The hidden vector of and Represents the binary rating matrix R (2) The corresponding latent vectors of active users and popular items.
[0087] It is worth noting that in this embodiment, the scoring matrix R (5) and R (2) Perform matrix decomposition on the whole, rather than just on the regions associated with active users and popular items. (5) and a (2) The corresponding scoring sub-matrix R(a (5) ) and R(a (2) ) for decomposition.
[0088] Due to R (5) and R (2) R(a (5)) and R(a (2) ) has more rating information, so the rating matrix R (5) and R (2) Performing matrix decomposition as a whole can obtain more accurate latent vector features.
[0089] Based on the acquired latent vector features of active users and by As input, As output, a deep regression network is constructed to learn the mapping relationship F between them. 1 ;
[0090] The same principle is used to learn the two latent vector mapping relationships F corresponding to popular items 2 .
[0091] However, since the number of active users and popular projects is often small, directly constructing a deep regression network is not ideal.
[0092] Taking the latent vector mapping relationship modeling of active users as an example, considering that there are still a large number of inactive users on the recommendation platform, their latent vector features share the same feature space with the latent vector features of active users. In order to further improve the accuracy of mapping relationship modeling, this embodiment includes the following steps when modeling the mapping relationship:
[0093] First, we use the latent vector features of a large number of inactive users Used as unsupervised training data to train stacked denoising autoencoders (SDAE) to obtain low-dimensional high-level representations of latent vector features;
[0094] For example, let x denote the original training data, and x plus Gaussian noise is transformed into After encoding by the encoder, the low-dimensional feature representation y is obtained, and the formula is as follows:
[0095]
[0096] Where W and b represent the encoder weight matrix and bias vector respectively, and S represents the ReLu activation function. Passing y through the decoder to obtain the reconstructed data of the input data is expressed as:
[0097] z=g(y)=S(W′y+b′)
[0098] Where z is the reconstructed data, W′ and b′ represent the decoder weight matrix and bias vector respectively.
[0099] The loss function is:
[0100]
[0101] Where M represents the number of samples. Multiple denoising autoencoders (DAE) are stacked together to obtain a stacked denoising autoencoder. The stacked denoising autoencoder is trained using the unsupervised feature data corresponding to inactive users to obtain a low-dimensional high-level representation of the latent vector features. Figure 4 As shown in the figure, (a) learning is performed layer by layer; (b) multiple layers of denoising autoencoders are spliced; (c) the entire unsupervised data set is used to fine-tune the weights using the BP algorithm.
[0102] Then, a layer of linear regression units is added to the encoding layer to build a deep regression network, and a small amount of supervised training data corresponding to active users is used. Train the deep regression network to model the mapping relationship.
[0103] The process of fine-tuning the regression model is as follows Figure 5 As shown, the linear regression unit does not contain any activation function and only calculates the weighted sum of each input unit.
[0104] The loss function is defined as follows:
[0105]
[0106] in Is an active user a Based on R (5) The latent vector obtained by matrix decomposition, is the latent vector predicted by the deep regression network, where For active users u a Based on R (2) Hidden vector obtained by matrix decomposition.
[0107] like Figure 5 As shown, during the deep regression network training process, the Figure 4 The final weights (W′) of the encoder in the trained SDAE 1 ,W′ 2 ,W′ 3 ) Initialize the weights of the encoder in the deep regression network and randomly initialize the weights W′ of the outermost linear regression unit 4 Then, the BP algorithm is used to learn all the weights of the deep regression network to obtain the final deep regression network, that is, the mapping relationship F 1 The same method can be used to model the mapping relationship between the two latent vectors corresponding to popular items F 2 .
[0108] In step 6, the mapping relationship between the user latent vector and the item latent vector is used to obtain the features of inactive items and non-popular items in the target domain; including:
[0109] Map the latent vectors corresponding to active users and popular items to F 1 and F 2 Expand to the entire target domain;
[0110] The latent factor vector of inactive users in the source domain is relatively accurate Through the mapping relationship F 1 Get inactive user u ina The latent factor vector in the target domain Right now
[0111] Similarly, the latent factor vectors of non-popular items that are more accurate in the source domain are Through the mapping relationship F 2 Get non-popular items unp The latent factor vector in the target domain Right now
[0112] In step 7, based on the latent vectors of inactive users and non-popular items in the target domain, the constrained matrix factorization model construction process includes:
[0113] make is the numerical matrix R (5) The rating of item i by user u in is the hidden vector of any user u finally solved by the cross-region and cross-rating collaborative filtering model in this paper, is the hidden vector of any item i that is finally solved. For active users u a Based on the rating matrix R (5) The latent vector obtained by decomposition is For popular projects p Based on R (5) The latent vector obtained by decomposition.
[0114] For each region of the target domain with different rating densities, we obtain the final user and item latent vectors of the target domain by solving the following optimization problem, and realize the transfer of knowledge from the rating-dense regions of the auxiliary domain and the target domain to the rating-non-dense regions of the target domain:
[0115]
[0116] where λ 1 ,λ 2 are two regularization coefficients,
[0117] This embodiment uses stochastic gradient descent to solve the optimization problem, and the iterative formula is as follows:
[0118]
[0119]
[0120] in γ represents the learning rate.
[0121] In the above optimization problem, we use Constrain the latent vectors of active users and inactive users in the target domain. If u is an active user, then That is, based on the active user u based on the rating matrix R (5) The latent vector obtained by decomposition is used as a constraint. If u is an inactive user, then That is, the latent vector of the inactive user u based on the mapping relationship is used as a constraint. Constrain the latent vectors of popular items and non-popular items in the target domain. If i is a popular item, then That is, based on popular projects i (5) The latent vector obtained by decomposition is used as a constraint. If i is a non-hot item, then That is, the latent vector obtained based on the mapping relationship of the non-hot item i is used as a constraint. Therefore, this embodiment realizes personalized knowledge transfer for different regions of the target domain by solving the above optimization problem, and the above matrix decomposition method with added constraints is called a restricted matrix decomposition method.
[0122] According to the latent factor vector of any user u obtained by solving and the latent factor vector for any item i Predict user u's rating for item i, that is According to the predicted ratings of the target users for the predicted items, the top-N items with the highest predicted ratings are selected as the recommendation list for the user.
[0123] Embodiment 2
[0124] This embodiment provides a cross-region and cross-rating collaborative filtering recommendation system, including:
[0125] The data preprocessing module is configured to: obtain user-item rating data of the target domain and the source domain, and obtain a target domain rating matrix and a source domain rating matrix after preprocessing;
[0126] Divide all users in the target domain rating matrix and the source domain rating matrix into active users and inactive users, and divide all projects into popular projects and non-popular projects;
[0127] The feature extraction module is configured to: decompose the target domain scoring matrix and the source domain scoring matrix based on the latent semantic model, and extract the user latent vector and the item latent vector in the target domain and the source domain;
[0128] For active users and popular items, based on the trained deep regression network, we learn the mapping relationship between the user latent vector and the item latent vector corresponding to the target domain and the source domain under the two rating systems.
[0129] The mapping relationship between the user latent vectors and item latent vectors of active users and popular items is used to obtain the features of inactive users and non-popular items in the target domain;
[0130] The recommendation acquisition module is configured to: construct a restricted matrix decomposition model based on the characteristics of inactive users and non-popular items in the target domain, predict the score of any user for any item, and select the item with the highest predicted score as the user's recommendation result.
[0131] Embodiment 3
[0132] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the cross-region and cross-rating collaborative filtering recommendation method described above are implemented.
[0133] Embodiment 4
[0134] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the cross-region and cross-rating collaborative filtering recommendation method described above are implemented.
[0135] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.
[0136] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0137] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0139] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0140] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A cross-region and cross-rating collaborative filtering recommendation method, It is characterized in that include: Obtain user-item rating data from the target domain and source domain, and obtain the target domain rating matrix and the source domain rating matrix after preprocessing; Divide all users in the target domain rating matrix and the source domain rating matrix into active users and inactive users, and divide all projects into popular projects and non-popular projects; Decompose the target domain rating matrix and the source domain rating matrix based on the latent semantic model, and extract the user latent vector and item latent vector in the target domain and the source domain; For active users and popular items, based on the trained deep regression network, we learn the mapping relationship between the user latent vector and the item latent vector corresponding to the target domain and the source domain under the two rating systems. In the deep regression network training process, the final weights of the encoder in the trained SDAE are used to initialize the weights of the encoder in the deep regression network, the weights of the outermost linear regression units are randomly initialized, and then all the weights of the deep regression network are learned using the BP algorithm; The mapping relationship between the user latent vectors and item latent vectors of active users and popular items is used to obtain the features of inactive users and non-popular items in the target domain; According to the characteristics of inactive users and non-popular items in the target domain, a restricted matrix decomposition model is constructed to predict the score of any user for any item, and the item with the highest predicted score is selected as the user's recommendation result; The construction process of the restricted matrix decomposition model is as follows: in, For users in the 5-point rating matrix About Project Rating, For any project The hidden vector of For any user The hidden vector of are two regularization coefficients, is the latent vector constraint condition for the popular items and non-popular items in the target domain, is the latent vector constraint for active and inactive users in the target domain, , , For popular projects, For non-popular projects, For active users, For inactive users, and Represents a 5-point numerical scoring matrix The decomposed user and item latent vectors are Hidden vector representing the non-popular items in the source domain Through the mapping relationship Get unpopular items The latent vector in the target domain, The latent vector representing the inactive user in the source domain Through the mapping relationship Get the latent vector of inactive users in the target domain, Represents the mapping relationship between the source domain latent vector and the target domain latent vector of active users, It represents the mapping relationship between the source domain latent vector and the target domain latent vector of the popular project. The source domain uses binary ratings, and the target domain uses 5-point numerical ratings.
2. A cross-region and cross-rating collaborative filtering recommendation method as claimed in claim 1, It is characterized in that In the target domain rating matrix and the source domain rating matrix, the number of ratings in each row is the number of ratings received by the user, and the number of ratings in each column is the number of ratings received by the project.
3. A cross-region and cross-rating collaborative filtering recommendation method as claimed in claim 1, It is characterized in that For the target domain rating matrix, users are sorted according to the number of their ratings, and users are arranged from top to bottom in the rating matrix in descending order of the number of ratings. According to the number of project ratings, projects are arranged from left to right in the rating matrix in descending order of the number of ratings. Users and projects in the source domain are arranged in the same order as in the target domain, and users and projects in the target domain and source domain correspond one to one.
4. A cross-region and cross-rating collaborative filtering recommendation method as claimed in claim 1, It is characterized in that Extracting user latent vectors and item latent vectors in the source domain includes: using an improved Funk-SVD model to extract auxiliary domain latent vector features, using a cross entropy loss function instead of a least squares loss function as the loss function of the model, and converting the regression problem into a classification problem.
5. The cross-region and cross-rating collaborative filtering recommendation method according to claim 1, It is characterized in that The inactive user and non-hot item features in the target domain obtained by using the mapping relationship between the user latent vector and the item latent vector of the active user and the hot item include: Expand the latent vector mapping relationship between active users and popular items to the entire target domain; The latent factor vectors of inactive users and non-popular items in the source domain are obtained through the latent vector mapping relationship to obtain the latent factor vectors of inactive users and non-popular items in the target domain.
6. A cross-regional and cross-rating collaborative filtering recommendation system, It is characterized in that The cross-region and cross-rating collaborative filtering recommendation method according to any one of claims 1 to 5 comprises: The data preprocessing module is configured to: obtain user-item rating data of the target domain and the source domain, and obtain a target domain rating matrix and a source domain rating matrix after preprocessing; Divide all users in the target domain rating matrix and the source domain rating matrix into active users and inactive users, and divide all projects into popular projects and non-popular projects; The feature extraction module is configured to: decompose the target domain scoring matrix and the source domain scoring matrix based on the latent semantic model, and extract the user latent vector and the item latent vector in the target domain and the source domain; For active users and popular items, based on the trained deep regression network, we learn the mapping relationship between the user latent vector and the item latent vector corresponding to the target domain and the source domain under the two rating systems. The mapping relationship between the user latent vectors and item latent vectors of active users and popular items is used to obtain the features of inactive users and non-popular items in the target domain; The recommendation acquisition module is configured to: construct a restricted matrix decomposition model based on the characteristics of inactive users and non-popular items in the target domain, predict the score of any user for any item, and select the item with the highest predicted score as the user's recommendation result.
7. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the steps in the cross-region and cross-rating collaborative filtering recommendation method according to any one of claims 1 to 5 are implemented.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps in the cross-region and cross-rating collaborative filtering recommendation method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method for optimizing collaborative filtering recommendation system by aggregation
CN102129462A
Collaborative filtering-based optimization method
CN108038629A