Hybrid collaborative filtering recommendation method based on user clustering and product clustering
By introducing the combination of user clustering, product clustering, noise reduction processing and fusion factor α in the collaborative filtering recommendation algorithm, the cold start and sparse matrix problems are solved, and multi-dimensional accurate recommendation is achieved, improving the accuracy and comprehensiveness of recommendations.
Patent Information
- Application Number
- CN202210437699.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-04-25
AI Technical Summary
The existing collaborative filtering recommendation algorithm has poor recommendations for new users and new products during the cold start stage, and is sensitive to sparse matrix, affecting the accuracy of recommendations, and it is difficult to explore the relationship between products and user personalized preferences.
A hybrid collaborative filtering recommendation algorithm based on user clustering and product clustering is adopted to construct a preference matrix for user attributes to product attributes by introducing external purchase records, and achieve cold start; an improved K-Means dual clustering and noise reduction algorithm is used to improve clustering accuracy; a fusion factor α is introduced to generate prediction evaluations based on user and product clustering; an Apriori association rule analysis and ID3 empowerment analysis are used to obtain association rules and personalized recommendation sets, and a multi-dimensional accurate recommendation is achieved by combining Top-n recommendation sets.
Accurate recommendations to users during the cold start stage, improving clustering accuracy and recommendation accuracy, better tapping into user personalized preferences and product relationships, and improving the scientificity and comprehensiveness of recommendations.
Smart Images

Figure CN114741603B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of collaborative filtering recommendation algorithms and provides a recommendation method of hybrid collaborative filtering based on user clustering and commodity clustering. Background Art
[0002] With the rapid development of the Internet, recommendation technology is increasingly used in various fields of the Internet. As the types and quantities of data increase, how to process and analyze the data obtained by e-commerce platforms and accurately recommend the products that users are most likely to buy to users is a challenging problem.
[0003] At present, all major Internet platforms have implemented recommendation functions to varying degrees. In order to ensure the accuracy of recommendations, researchers have proposed many different types of recommendation algorithms, such as content-based recommendation algorithms, association rule-based recommendation algorithms, collaborative filtering-based recommendation algorithms, and machine learning model-based recommendation algorithms. Among them, collaborative filtering-based recommendation algorithms are widely used in systems with user ratings, such as large e-commerce platforms. Currently, collaborative filtering algorithms are divided into two types: user-based recommendations and item-based recommendations.
[0004] The recommendation algorithm based on collaborative filtering has significant advantages in the application of recommendation system. It does not require strict modeling of users or products, and can learn from others' experience in the recommendation process, which can well mine users' potential intended products. However, the collaborative filtering algorithm has obvious defects, which can be roughly summarized into four aspects: First, the core of the algorithm is based on historical data, and there is a "cold start" problem for new products entering the system and newly registered users. Secondly, the algorithm is sensitive to noise. If the noise points cannot be identified, the calculation accuracy will be affected. Thirdly, in actual situations, users' evaluation of products is sparse, which will cause problems such as slow calculation speed and low accuracy. Finally, during the use of the system, the algorithm is not flexible enough, it is difficult to use historical data to mine the relationship between products, and it is impossible to analyze users' personalized preferences, making it difficult to complete accurate recommendations for users.
[0005] To improve the above problems, some methods use the mean assignment method to alleviate the sparsity of the matrix. This method alleviates the problem of matrix sparsity to a certain extent. However, this method has two problems. In the user evaluation assignment stage, although it reflects a certain real situation, it also fabricates a large number of unreliable evaluations, which has a negative effect on the subsequent recommendation results. Secondly, in the clustering stage, the use of this method will affect the clustering effect, and then affect the recommendation results. In addition, the use of this method lacks the rationality of the result interpretation. Therefore, solving the high sparsity of the evaluation matrix and improving the prediction quality and prediction accuracy is a topic worthy of research.
[0006] In summary, the present invention proposes a hybrid collaborative filtering recommendation algorithm based on user clustering and product clustering. In the cold start phase, external purchase records are introduced to obtain the preference matrix of user attributes to product attributes to achieve preliminary recommendations; dual clustering of users and products solves the problem of high sparsity of user evaluation matrix, and uses noise reduction algorithm to improve clustering effect; introduces fusion factor α, combines user and product clustering, and obtains Top-n recommendation set; uses product-based collaborative filtering to obtain association rule recommendation set, and then combines ID3 classification method to weight personalized user preferences to obtain personalized user recommendation set, and achieves multi-dimensional and accurate recommendations for users. Summary of the invention
[0007] In order to make up for the deficiencies of various recommendation algorithms, the present invention proposes a hybrid collaborative filtering recommendation algorithm based on user clustering and commodity clustering for the Internet e-commerce platform. First, in order to solve the cold start problem of the recommendation algorithm, the algorithm introduces external purchase records, and uses the pseudo-inverse matrix to construct a preference matrix of user attributes to commodity attributes. Combined with the user and commodity information in the system, the cold start of the recommendation algorithm can be achieved. Subsequently, an improved K-Means dual clustering model is designed, and the normalized cosine theorem is used to cluster users and commodities respectively, and then the noise reduction algorithm is used to accurately cluster the results, and the fusion factor α is introduced to generate the final prediction score and obtain the Top-n recommendation set. Finally, the Apriori association rule analysis is used to mine the association rules of the transaction records to obtain the association rule recommendation set, and then combined with ID3, the user's personalized shopping preferences are empowered to obtain the user's personalized recommendation set, and the three recommendation sets are combined to achieve multi-dimensional and accurate recommendations for users. In summary, the present invention proposes a hybrid collaborative filtering recommendation algorithm based on user clustering and commodity clustering. This method solves the cold start problem in the initial stage of recommendation by generating a preference matrix of user attributes to product attributes. Then, by performing double clustering of users and products and denoising, the problem of sparse matrix is solved, the clustering accuracy is improved, and the Top-n recommendation set is obtained. Then, association rule analysis is performed on the purchase records, combined with ID3 weighted analysis, to obtain the user's association rule recommendation set and personalized recommendation set, and reliable multi-dimensional recommendation is achieved by combining the Top-n recommendation set.
[0008] In order to achieve the above object, the technical solution adopted by the present invention is:
[0009] A hybrid collaborative filtering recommendation method based on user clustering and product clustering includes the following steps:
[0010] Step 1: Reference external purchase records to obtain the evaluation matrix of user attributes to product attributes;
[0011] By using external purchase records, we can obtain the evaluation matrix C of user attributes to product attributes. The evaluation matrix is shown in formula (1) or formula (2):
[0012] U e ×C×P e =E e (1)
[0013]
[0014] Among them, the subscript e indicates the external purchase record content; U e represents the user attribute matrix, which is the user ID of the external data source and is listed as user attributes; C represents the evaluation matrix of user attributes on product attributes, which is the user attributes and is listed as product attributes; P e Represents the product attribute matrix, with rows being product attributes and columns being product IDs from external data sources; e Represents the user-item evaluation matrix. Indicates U e The inverse of a matrix, Indicates P e The inverse matrix of the matrix. Formula (1) can be equivalent to formula (2).
[0015] In practical applications, U e With P e The matrix is usually a non-square matrix, which cannot be inverted and cannot be obtained. and Therefore, to realize the above formula, the present invention uses the left inverse matrix (3) and the right inverse matrix (4):
[0016] H L =(H T ×H) -1 ×H T (3)
[0017] H R =H T ×(H T ×H) -1 (4)
[0018] Where H is an m-row n-column matrix: When m ≥ n, H has a left inverse matrix H L , the calculation formula is (3); when m≤n, H has a right inverse matrix H R , the calculation formula is (4). Therefore, when U e When the number of rows ≥ the number of columns, we can substitute into formula (3) to find When P e When the number of rows ≤ the number of columns, we can substitute into formula (4) to find U e With P eThe selection of obviously meets the above conditions. Therefore, the C matrix can be obtained through formula (2). This matrix is universal and can be combined with the user attribute matrix and product attribute matrix in the system to complete the cold start.
[0019] Step 2: Combine the evaluation matrix of user attributes to product attributes, the user attribute matrix in the system, and the product attribute matrix to obtain the initial user evaluation matrix, select the Top-n, and complete the cold start; the details are as follows:
[0020] In order to solve the problem of being unable to recommend in the cold start phase, the present invention introduces external purchase records to obtain the evaluation matrix C of user attributes to product attributes, and combines it with the existing user information and product information in the system to obtain the user evaluation matrix, which can be expressed as formula (5):
[0021] U i ×C×P i =E i (5)
[0022] Among them, U i represents the user attribute matrix of this system, with the behavior being user ID and the columns being user attributes; i It represents the commodity attribute matrix of this system, with rows as commodity attributes and columns as commodity IDs; i Represents the initial user product evaluation matrix of this system.
[0023] For E i In the system, the ratings of each user on the products are sorted, and the n products with the highest ratings are selected for preliminary recommendation. In this way, the top-n recommendations for users can be achieved in the cold start phase, completing the cold start. After the system accumulates a certain amount of user, product, and evaluation data, the subsequent steps are carried out.
[0024] Step 3: Use the existing user attribute matrix and product attribute matrix in the system to cluster users and products respectively, determine the initial cluster center, and perform clustering iterations until the clusters are stable and unchanged; the details are as follows:
[0025] 3.1) The present invention uses the cosine similarity theorem to calculate the similarity between users and between products. The formula can be expressed as (6):
[0026]
[0027] In this example, taking the calculation of user similarity as an example, i and j represent the attribute vectors of different users, sim(i,j) represents the similarity between the attributes of user i and user j; t represents an attribute in the user attribute category T, R i,t and R j,t They represent the evaluation of user i and j on attribute t respectively. The calculation method of product similarity is the same as above.
[0028] 3.2) In K-Means clustering, cluster centers are used to cluster similar users or commodities. The quality of the initial cluster center selection can significantly affect the number of iterations. In the present invention, k points that are the least similar to each other are selected as the initial cluster centers. Specifically: First, this method randomly selects a user or commodity attribute vector as the first initial cluster center, and then traverses all user or commodity attribute vectors, selects the user or commodity attribute vector that is the least similar to the existing cluster center as the next initial cluster center, and repeats the above steps. Finally, all initial cluster centers are selected. In the cluster center initialization stage, this method selects cluster centers with large differences, which can greatly reduce the number of iterations in the clustering iteration stage and improve the algorithm performance.
[0029] 3.3) Clusters are containers for user or product vectors gathered at the cluster center, storing user or product vectors with high similarity. In order to stabilize the clustering results, the present invention calculates the mean of all vectors in the cluster as the new cluster center and compares it with the original cluster center. This method can be expressed as formula (7):
[0030]
[0031] Among them, K old is the original cluster center, K new is the new cluster center, n is the number of vectors in cluster B, and b is the vector in cluster B.
[0032] First, this method uses formula (7) to calculate the similarity between the vector and each initial cluster center; then, the vector is assigned to the cluster where the most similar cluster center is located; finally, the mean of the sum of all vectors in the cluster is calculated as the new cluster center K. new :If K new =K old , it indicates that the cluster is stable, otherwise K new Replace K old , repeat the above steps until all clusters are stable.
[0033] Step 4: Noise reduction processing: vectors close to multiple cluster centers are identified as noise points and denoised; the details are as follows:
[0034] There are noise points in the cluster that cannot be accurately clustered. Specifically, the noise point is very similar to two or more cluster centers. If it is added to any cluster, it cannot be accurately predicted. At the same time, the existence of noise points will affect the prediction of other vectors in the cluster, so the processing of noise points is very important for the accuracy of the recommendation algorithm. The present invention compares the similarity between the vector and different cluster centers to determine whether it is a noise point, which can be expressed as formula (8):
[0035]
[0036] Among them, k i Represents the cluster center of the cluster to which vector i belongs; K represents the cluster center of the cluster to which vector i belongs;i The cluster center set other than k l represents the cluster center in K; sim(k i ,i) represents the cluster center k of the vector i and its cluster i Similarity; sim(k l ,i) represents the vector i and cluster center k l Similarity; D i,l represents sim(k i ,i) and sim(k l ,i) The absolute value of the difference; β represents the threshold.
[0037] If there is D i,j ≤β, the target vector i has a high similarity with multiple cluster centers and should be identified as a noise point; if l ∈K exists D i,l ≥β, then vector i is not identified as a noise point. Then the locations of all noise points are recorded. Furthermore, no prediction is performed in the subsequent prediction evaluation to achieve the effect of noise reduction.
[0038] Step 5: Obtain user prediction evaluation and product prediction evaluation, introduce fusion factor α, obtain the final prediction evaluation, and obtain the Top-n recommendation set; the details are as follows:
[0039] So far, this method has completed user clustering and product clustering. According to the clustering results, the neighbor set N of users and products can be found. The neighbor set N stores other user or product attribute vectors that are highly similar to the user or product. When making predictions and evaluations, predictions will be made based on the real evaluations of neighbor users or products in the neighbor set.
[0040] The user-product evaluation matrix r is introduced. r contains the real evaluation of users on the products and the unknown evaluation of users on the products. The Top-n recommendation set can be obtained by combining the predicted evaluation of the unknown evaluation with the real evaluation.
[0041] Next, we make prediction and evaluation based on the clustering results and neighbor sets. The prediction and evaluation based on user clustering can be expressed as formula (9), and the prediction and evaluation based on product clustering can be expressed as formula (10):
[0042]
[0043] Among them, P i,u represents the predicted evaluation of user i on product u based on user clustering; represents the mean of user i’s evaluation; N i represents the neighbor set of user i; r j,u represents user j’s evaluation of product u; represents the mean of user j's evaluation; sim(i,j) represents the similarity between the attributes of user i and user j. This prediction method eliminates the error caused by the different evaluation scales among users, making the prediction more accurate and reliable.
[0044]
[0045] Among them, Q i,u N represents the predicted evaluation of user i on product u based on product clustering; u represents the neighbor set of commodity u; r i,v represents the evaluation of user i on product v; sim(v,u) represents the similarity between the attributes of product v and product u.
[0046] First, the present invention obtains the prediction evaluation based on user clustering and commodity clustering as shown in formula (9) and formula (10); then, the fusion factor α is introduced, and the two prediction evaluations are combined with the fusion factor α to jointly generate the final prediction evaluation; finally, the final prediction evaluation is used to replace the unknown evaluation in the user evaluation matrix r, and combined with the existing real evaluation, the n commodities with the highest evaluation (Top-n) are selected for each user to generate the Top-n recommendation set M; the fusion formula is expressed as (11):
[0047] F i,u =α×Q i,u +(1-α)×P i,u (11)
[0048] Among them, F i,u Represents the predicted evaluation of user i on product u; α is the fusion factor. This method is more accurate than the prediction based on user clustering or product clustering alone, where α∈[0,1].
[0049] Furthermore, after repeated experiments, we found that when the fusion factor α is 0.5, the best prediction effect can be achieved.
[0050] Step 6: Perform Apriori association rule mining on the transaction data and obtain a set of recommended products with association rules based on the user's purchase records. The details are as follows:
[0051] Using clustering alone for prediction has a single evaluation angle. Therefore, the present invention introduces transaction data, which stores all the product IDs purchased by a single user within a certain period of time, and performs Apriori association rule mining on the transaction data to mine the hidden relationships between the products.
[0052] In the present invention, this method can be used to find the association between commodities. The association rules are measured using two dimensions: support and confidence. In the present invention, support is used to indicate the probability that commodity item set A and commodity item set B are purchased at the same time; confidence is used to indicate the probability that commodity item set B is purchased after commodity item set A is purchased, which can be expressed as formula (12):
[0053]
[0054] Therefore, in association rule mining, the potential associations between products can be found by selecting appropriate support and confidence. Specifically:
[0055] First, the purchase data is preprocessed to extract all the products purchased by a single user as a purchase set. In the present invention, the support of association rule mining is set to δ∈(0,1), and the confidence is set to μ∈(0,1). Rules with a probability of association occurring above δ and a probability of associated purchase above μ are mined. In the present invention, δ=0.01 and μ=0.5 are selected. Then, the present invention will retrieve the user's purchase record. If the user has a purchase record and has purchased products with association rules, the n products with the highest confidence will be selected to generate a product recommendation set N based on association rules for the user. If the user does not have a purchase record, the present invention will perform an association rule search based on the Top-n products, select the n products with the highest confidence, and generate a product recommendation set N based on association rules for the user.
[0056] Step 7: Analyze user purchase records using ID3, perform weighted analysis to obtain the user's preference weights for product attributes, and obtain the user's personalized recommendation set. Combine it with the association rule recommendation set and the Top-n recommendation set to obtain the final recommendation list. The details are as follows:
[0057] In recommendation algorithms, due to differences in preferences among different users, the recommendation results are often not ideal. Therefore, analyzing a single user's personalized preferences for product attributes based on their purchase records is particularly important for improving recommendation accuracy.
[0058] The present invention uses the ID3 classification algorithm to find out the most important commodity features for users. First, the information entropy is used to measure the uncertainty of H(S) users in purchasing commodities. The smaller the information entropy, the smaller the uncertainty of the data sample and the purer the data, which can be expressed as formula (13):
[0059] H(S)=∑p i log(p i ) (13)
[0060] Among them, p i represents the probability of the i-th category, and S represents the data set.
[0061] Then, the conditional entropy H(S|A) is used to examine the change in entropy after division according to a certain attribute feature. The conditional entropy (feature expectation) can be used to represent the uncertainty of a certain feature. The smaller the feature expectation, the smaller the data uncertainty of the feature, which can be expressed as formula (14):
[0062]
[0063] Among them, A represents the attribute features of the data set S; c j Represented as the number of samples of category j in feature A; c j The ratio c represents the probability of all categories of category j in feature A; p ij is the probability that the j-th category of feature A accounts for the different categories of the data set S.
[0064] Then, using the information gain G A (S) measures the contribution of attribute feature A to reducing the entropy of data set S. The larger the information gain, the more suitable it is for classifying S, which means that a certain feature is more important, which can be expressed as formula (15):
[0065] G A (S) = H (S) - H (S|A) (15)
[0066] Select the maximum information gain G A The feature A of (S) is the parent node, and formulas (13)-(15) are repeated until G A (S) = 0. At this time, the ID3 decision tree is constructed and the most important product features for the user have been screened out.
[0067] Finally, the user order data is segmented by attributes, and the proportion of the most important product features to the user in their purchase records is calculated. K , and get the weight W of each attribute i , which can be expressed as formula (16):
[0068]
[0069] Therefore, the present invention can obtain the user's preference weights for the detailed attributes of the product, and then generate the user's ratings for different products, and select the n products with the highest user ratings to generate a personalized recommendation set G, which together with the association rule recommendation set and the Top-n recommendation set form the final recommendation list X, which can be expressed as formula (17) for accurate recommendation.
[0070] M+N+G=X (17)
[0071] The product IDs stored in the Top-n recommendation set M, the association rule recommendation set N, and the personalized recommendation set G are stored in the recommendation list to complete accurate recommendations for users.
[0072] The beneficial effects of the present invention are:
[0073] The present invention is aimed at the Internet e-commerce platform, designs a hybrid collaborative filtering recommendation algorithm based on user clustering and commodity clustering, designs an attribute preference matrix, and realizes the cold start of system recommendation; then the clustering results are subjected to noise reduction processing to make the clustering more accurate, and the fusion factor α is introduced to obtain the Top-n recommendation set. Finally, the association rules between commodities and the user's preference weights for commodity attributes are analyzed according to the purchase records, and the association rule recommendation set and the user personalized recommendation set are obtained. The user recommendation list is obtained in combination with the Top-n recommendation set to complete the multi-dimensional accurate recommendation. The present invention can make more accurate recommendations to users in the cold start stage without actual purchase data. In addition, the dual clustering algorithm and clustering noise reduction processing are used in clustering, so that the recommendation algorithm has a significant improvement in recommendation accuracy compared with the traditional recommendation algorithm. In addition, the present invention uses association rule analysis and combines ID3 for weighting to realize the recommendation of association rules between commodities and user personality, and the recommendation results are more scientific and comprehensive. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 This is a framework diagram of the hybrid collaborative filtering recommendation algorithm based on user clustering and product clustering.
[0075] Figure 2 This is a flow chart of the method proposed by the present invention. DETAILED DESCRIPTION
[0076] The present invention will be further described below with reference to the accompanying drawings.
[0077] Figure 1 The framework diagram of the hybrid collaborative filtering recommendation algorithm based on user clustering and product clustering. First, in the cold start phase, the present invention introduces external purchase records and obtains the corresponding matrix of user attributes and product attributes, combines the existing user product information in the system, obtains the user evaluation matrix, and then makes preliminary recommendations to the user. Secondly, after obtaining the user's real purchase record, the present invention will cluster the user and the product separately, and then identify the noise points in the clustering, process them, accurately cluster the results, introduce the fusion factor α, and jointly complete the user evaluation matrix according to the different predicted evaluations obtained by user clustering and product clustering to obtain Top-n. Finally, the present invention combines Apriori association rule analysis and ID3 analysis weighting analysis to jointly generate a recommendation list.
[0078] The specific steps are as follows:
[0079] Step 1: Reference external purchase records to obtain the evaluation matrix of user attributes to product attributes;
[0080] By using external purchase records, we can obtain the evaluation matrix C of user attributes to product attributes. The evaluation matrix is shown in formula (1) or formula (2):
[0081] U e ×C×P e =E e (1)
[0082]
[0083] Among them, the subscript e indicates the external purchase record content; U e represents the user attribute matrix, which is the user ID of the external data source and is listed as user attributes; C represents the evaluation matrix of user attributes on product attributes, which is the user attributes and is listed as product attributes; P e Represents the product attribute matrix, with rows being product attributes and columns being product IDs from external data sources; e Represents the user-product evaluation matrix. Indicates that U e The inverse of a matrix, Indicates P e The inverse matrix of the matrix. Formula (1) can be equivalent to formula (2).
[0084] In practical applications, U e With P e The matrix is usually a non-square matrix, which cannot be inverted and cannot be obtained. and Therefore, to realize the above formula, the present invention uses the left inverse matrix (3) and the right inverse matrix (4):
[0085] H L =(H T ×H) -1 ×H T (3)
[0086] H R =H T ×(H T ×H) -1 (4)
[0087] Where H is an m-row n-column matrix: When m ≥ n, H has a left inverse matrix H L , the calculation formula is (3); when m≤n, H has a right inverse matrix H R , the calculation formula is (4). Therefore, when U e When the number of rows ≥ the number of columns, we can substitute into formula (3) to find When P eWhen the number of rows ≤ the number of columns, we can substitute into formula (4) to find U e With P e The selection of obviously meets the above conditions. Therefore, the C matrix can be obtained through formula (2). This matrix is universal and can be combined with the user attribute matrix and product attribute matrix in the system to complete the cold start.
[0088] Step 2: Combine the evaluation matrix of user attributes to product attributes, the user attribute matrix in the system, and the product attribute matrix to obtain the initial user evaluation matrix, select the Top-n, and complete the cold start; the details are as follows:
[0089] In order to solve the problem of being unable to recommend in the cold start phase, the present invention introduces external purchase records to obtain the evaluation matrix C of user attributes to product attributes, and combines it with the existing user information and product information in the system to obtain the user evaluation matrix, which can be expressed as formula (5):
[0090] U i ×C×P i =E i (5)
[0091] Among them, U i represents the user attribute matrix of this system, with the behavior being user ID and the columns being user attributes; i It represents the commodity attribute matrix of this system, with rows as commodity attributes and columns as commodity IDs; i Represents the initial user product evaluation matrix of this system.
[0092] For E i In the system, the ratings of each user on the products are sorted, and the n products with the highest ratings are selected for preliminary recommendation. In this way, the top-n recommendations for users can be achieved in the cold start phase, completing the cold start. After the system accumulates a certain amount of user, product, and evaluation data, the subsequent steps are carried out.
[0093] Step 3: Use the existing user attribute matrix and product attribute matrix in the system to cluster users and products respectively, determine the initial cluster center, and perform clustering iterations until the clusters are stable and unchanged; the details are as follows:
[0094] 3.1) The present invention uses the cosine similarity theorem to calculate the similarity between users and between products. The formula can be expressed as (6):
[0095]
[0096] In this example, taking the calculation of user similarity as an example, i and j represent the attribute vectors of different users, sim(i,j) represents the similarity between the attributes of user i and user j; t represents an attribute in the user attribute category T, R i,t and Rj,t They represent the evaluation of user i and j on attribute t respectively. The calculation method of product similarity is the same as above.
[0097] 3.2) In K-Means clustering, cluster centers are used to cluster similar users or commodities. The quality of the initial cluster center selection can significantly affect the number of iterations. In the present invention, k points that are the least similar to each other are selected as the initial cluster centers. Specifically: First, this method randomly selects a user or commodity attribute vector as the first initial cluster center, and then traverses all user or commodity attribute vectors, selects the user or commodity attribute vector that is the least similar to the existing cluster center as the next initial cluster center, and repeats the above steps. Finally, all initial cluster centers are selected. In the cluster center initialization stage, this method selects cluster centers with large differences, which can greatly reduce the number of iterations in the clustering iteration stage and improve the algorithm performance.
[0098] 3.3) Clusters are containers for user or product vectors gathered at the cluster center, storing user or product vectors with high similarity. In order to stabilize the clustering results, the present invention calculates the mean of all vectors in the cluster as the new cluster center and compares it with the original cluster center. This method can be expressed as formula (7):
[0099]
[0100] Among them, K old is the original cluster center, K new is the new cluster center, n is the number of vectors in cluster B, and b is the vector in cluster B.
[0101] First, this method uses formula (7) to calculate the similarity between the vector and each initial cluster center; then, the vector is assigned to the cluster where the most similar cluster center is located; finally, the mean of the sum of all vectors in the cluster is calculated as the new cluster center K. new :If K new =K old , it indicates that the cluster is stable, otherwise K new Replace K old , repeat the above steps until all clusters are stable.
[0102] Step 4: Noise reduction processing: vectors close to multiple cluster centers are identified as noise points and denoised; the details are as follows:
[0103] There are noise points in the cluster that cannot be accurately clustered. Specifically, the noise point is very similar to two or more cluster centers. If it is added to any cluster, it cannot be accurately predicted. At the same time, the existence of noise points will affect the prediction of other vectors in the cluster, so the processing of noise points is very important for the accuracy of the recommendation algorithm. The present invention compares the similarity between the vector and different cluster centers to determine whether it is a noise point, which can be expressed as formula (8):
[0104]
[0105] Among them, k i Represents the cluster center of the cluster to which vector i belongs; K represents the cluster center of the cluster to which vector i belongs; i The cluster center set other than k l represents the cluster center in K; sim(k i ,i) represents the cluster center k of the vector i and its cluster i Similarity; sim(k l ,i) represents the vector i and cluster center k l Similarity; D i,l represents sim(k i ,i) and sim(k l ,i) The absolute value of the difference; β represents the threshold.
[0106] If there is D i,j ≤β, the target vector i has a high similarity with multiple cluster centers and should be identified as a noise point; if l ∈K exists D i,l ≥β, then vector i is not identified as a noise point. Then the locations of all noise points are recorded. Furthermore, no prediction is performed in the subsequent prediction evaluation to achieve the effect of noise reduction.
[0107] Step 5: Obtain user prediction evaluation and product prediction evaluation, introduce fusion factor α, obtain the final prediction evaluation, and obtain the Top-n recommendation set; the details are as follows:
[0108] So far, this method has completed user clustering and product clustering. According to the clustering results, the neighbor set N of users and products can be found. The neighbor set N stores other user or product attribute vectors that are highly similar to the user or product. When making predictions and evaluations, predictions will be made based on the real evaluations of neighbor users or products in the neighbor set.
[0109] The user-product evaluation matrix r is introduced. r contains the real evaluation of users on the products and the unknown evaluation of users on the products. The Top-n recommendation set can be obtained by combining the predicted evaluation of the unknown evaluation with the real evaluation.
[0110] Next, we make prediction and evaluation based on the clustering results and neighbor sets. The prediction and evaluation based on user clustering can be expressed as formula (9), and the prediction and evaluation based on product clustering can be expressed as formula (10):
[0111]
[0112] Among them, P i,u represents the predicted evaluation of user i on product u based on user clustering; represents the mean of user i’s evaluation; N i represents the neighbor set of user i; rj,u represents user j’s evaluation of product u; represents the mean of user j's evaluation; sim(i,j) represents the similarity between the attributes of user i and user j. This prediction method eliminates the error caused by the different evaluation scales among users, making the prediction more accurate and reliable.
[0113]
[0114] Among them, Q i,u N represents the predicted evaluation of user i on product u based on product clustering; u represents the neighbor set of commodity u; r i,v represents the evaluation of user i on product v; sim(v,u) represents the similarity between the attributes of product v and product u.
[0115] First, the present invention obtains the prediction evaluation based on user clustering and commodity clustering as shown in formula (9) and formula (10); then, the fusion factor α is introduced, and the two prediction evaluations are combined with the fusion factor α to jointly generate the final prediction evaluation; finally, the final prediction evaluation is used to replace the unknown evaluation in the user evaluation matrix r, and combined with the existing real evaluation, the n commodities with the highest evaluation (Top-n) are selected for each user to generate the Top-n recommendation set M; the fusion formula is expressed as (11):
[0116] F i,u =α×Q i,u +(1-α)×P i,u (11)
[0117] Among them, F i,u Represents the predicted evaluation of user i on product u; α is the fusion factor. This method is more accurate than the prediction based on user clustering or product clustering alone, where α∈[0,1].
[0118] Furthermore, after repeated experiments, we found that when the fusion factor α is 0.5, the best prediction effect can be achieved.
[0119] Step 6: Perform Apriori association rule mining on the transaction data and obtain a set of recommended products with association rules based on the user's purchase records. The details are as follows:
[0120] Using clustering alone for prediction has a single evaluation angle. Therefore, the present invention introduces transaction data, which stores all the product IDs purchased by a single user within a certain period of time, and performs Apriori association rule mining on the transaction data to mine the hidden relationships between the products.
[0121] In the present invention, this method can be used to find the association between commodities. The association rules are measured using two dimensions: support and confidence. In the present invention, support is used to indicate the probability that commodity item set A and commodity item set B are purchased at the same time; confidence is used to indicate the probability that commodity item set B is purchased after commodity item set A is purchased, which can be expressed as formula (12):
[0122]
[0123] Therefore, in association rule mining, the potential associations between products can be found by selecting appropriate support and confidence. Specifically:
[0124] First, the purchase data is preprocessed to extract all the products purchased by a single user as a purchase set. In the present invention, the support of association rule mining is set to δ∈(0,1), and the confidence is set to μ∈(0,1). Rules with a probability of association occurring above δ and a probability of associated purchase above μ are mined. In this embodiment, δ=0.01 and μ=0.5 are selected. Then, the present invention will retrieve the user's purchase record. If the user has a purchase record and has purchased products with association rules, the n products with the highest confidence will be selected to generate a product recommendation set N based on association rules for the user. If the user does not have a purchase record, the present invention will perform an association rule search based on the Top-n products, select the n products with the highest confidence, and generate a product recommendation set N based on association rules for the user.
[0125] Step 7: Analyze user purchase records using ID3, perform weighted analysis to obtain the user's preference weights for product attributes, and obtain the user's personalized recommendation set. Combine it with the association rule recommendation set and the Top-n recommendation set to obtain the final recommendation list. The details are as follows:
[0126] In recommendation algorithms, due to differences in preferences among different users, the recommendation results are often not ideal. Therefore, analyzing a single user's personalized preferences for product attributes based on their purchase records is particularly important for improving recommendation accuracy.
[0127] The present invention uses the ID3 classification algorithm to find out the most important commodity features for users. First, the information entropy is used to measure the uncertainty of H(S) users in purchasing commodities. The smaller the information entropy, the smaller the uncertainty of the data sample and the purer the data, which can be expressed as formula (13):
[0128] H(S)=∑p i log(p i ) (13)
[0129] Among them, p i represents the probability of the i-th category, and S represents the data set.
[0130] Then, the conditional entropy H(S|A) is used to examine the change in entropy after division according to a certain attribute feature. The conditional entropy (feature expectation) can be used to represent the uncertainty of a certain feature. The smaller the feature expectation, the smaller the data uncertainty of the feature, which can be expressed as formula (14):
[0131]
[0132] Among them, A represents the attribute features of the data set S; c j Represented as the number of samples of category j in feature A; c j The ratio c represents the probability of all categories of category j in feature A; p ij is the probability that the j-th category of feature A accounts for the different categories of the data set S.
[0133] Then, using the information gain G A (S) measures the contribution of attribute feature A to reducing the entropy of data set S. The larger the information gain, the more suitable it is for classifying S, which means that a certain feature is more important, which can be expressed as formula (15):
[0134] G A (S) = H (S) - H (S|A) (15)
[0135] Select the maximum information gain G A The feature A of (S) is the parent node, and formulas (13)-(15) are repeated until G A (S) = 0. At this time, the ID3 decision tree is constructed and the most important product features for the user have been screened out.
[0136] Finally, the user order data is segmented by attributes, and the proportion of the most important product features to the user in their purchase records is calculated. K , and get the weight W of each attribute i , which can be expressed as formula (16):
[0137]
[0138] Therefore, the present invention can obtain the user's preference weights for the detailed attributes of the product, and then generate the user's ratings for different products, and select the n products with the highest user ratings to generate a personalized recommendation set G, which together with the association rule recommendation set and the Top-n recommendation set form the final recommendation list X, which can be expressed as formula (17) for accurate recommendation.
[0139] M+N+G=X (17)
[0140] The product IDs stored in the Top-n recommendation set M, the association rule recommendation set N, and the personalized recommendation set G are stored in the recommendation list to complete accurate recommendations for users.
[0141] Method flow description:
[0142] The overall process of the present invention is divided into three parts: a cold start stage, an improved K-Means double clustering stage, and a stage of combining multiple methods to obtain a comprehensive recommendation list. First, external purchase records are introduced to obtain a user attribute-product attribute correspondence matrix, and the initial user evaluation matrix is generated by combining the existing user information and product information in the system to make preliminary recommendations. Then, K-means double clustering is performed on users and products, and the clustering accuracy is improved using a noise reduction algorithm. The fusion factor α is introduced, and the user clustering and product clustering are combined to generate the final predicted score to obtain Top-n. Finally, the transaction data is analyzed, and Apriori association rule analysis and ID3 weighted analysis are performed on existing orders. The final recommendation list is combined with Top-n to achieve multi-dimensional and accurate recommendations for users. The specific process is as follows: Figure 2 As shown. The innovation of the present invention is to use hybrid recommendation including improved k-means (user attribute dimension), association rule mining (product association dimension), and personalized recommendation set (user personalized dimension) to achieve multi-dimensional recommendation for users. Among them, steps one and two complete the cold start, and then steps three, four, and five complete the improved k-means recommendation, step six association rule recommendation, and step seven combines the ID3 algorithm to use information entropy to complete user personalized recommendation. The three recommendation methods are independent of each other and together constitute a hybrid recommendation method. Finally, through step seven, the user's preference weights for different product attributes of the product can be obtained, thereby achieving personalized recommendation for the user.
[0143] The above-described embodiments merely express the implementation methods of the present invention, but they cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A hybrid collaborative filtering recommendation method based on user clustering and product clustering, characterized in that: The following steps are involved: Step 1: Reference external purchase records to obtain the evaluation matrix of user attributes to product attributes; By using external purchase records, we can obtain the evaluation matrix C of user attributes to product attributes. The evaluation matrix is shown in formula (1) or formula (2): The e ×C×P e =E e (1) Among them, the subscript e indicates the external purchase record content; U e represents the user attribute matrix, which is the user ID of the external data source and is listed as user attributes; C represents the evaluation matrix of user attributes on product attributes, which is the user attributes and is listed as product attributes; P e Represents the product attribute matrix, with rows being product attributes and columns being product IDs from external data sources; e Represents the user product evaluation matrix; Indicates U e The inverse of a matrix, Indicates P e The inverse matrix of the matrix; Formula (1) can be equivalent to Formula (2); Step 2: Combine the evaluation matrix of user attributes to product attributes, the user attribute matrix in the system, and the product attribute matrix to obtain the initial user evaluation matrix, select the Top-n, and complete the cold start; the details are as follows: The external purchase records are introduced to obtain the evaluation matrix C of user attributes to product attributes to solve the problem of being unable to recommend in the cold start phase. It is combined with the existing user information and product information in the system to obtain the user evaluation matrix, which is expressed as formula (5): The i ×C×P i =E i (5) Among them, U i represents the user attribute matrix of this system, with the behavior being user ID and the columns being user attributes; i It represents the commodity attribute matrix of this system, with rows as commodity attributes and columns as commodity IDs; i Represents the initial user product evaluation matrix of this system; For E i Each user sorts the ratings of the products, selects the n products with the highest ratings, and makes preliminary recommendations; Step 3: Use the existing user attribute matrix and product attribute matrix in the system to cluster users and products respectively, determine the initial cluster center, and perform clustering iterations until the clusters are stable and unchanged; Step 4: Noise reduction processing: vectors close to multiple cluster centers are identified as noise points and denoised; Step 5: Obtain user prediction evaluation and product prediction evaluation, introduce fusion factor α, obtain the final prediction evaluation, and obtain the Top-n recommendation set; The two prediction evaluations are combined with the fusion factor α to generate the final prediction evaluation. The fusion formula is expressed as (11): F i,u =α×Q i,u +(1-a)×P i,u (11) Among them, F i,u represents the predicted evaluation of user i on product u; α is the fusion factor, α∈[0,1]; Q i,u represents the predicted evaluation of user i on product u based on product clustering; P i,u represents the predicted evaluation of user i on product u based on user clustering; Step 6: Perform Apriori association rule mining on the transaction data and obtain a product recommendation set with association rules based on the user's purchase records; Step 7: Use ID3 to analyze user purchase records, perform weighted analysis to obtain the user's preference weights for product attributes, and obtain the user's personalized recommendation set. Combine it with the association rule recommendation set and the Top-n recommendation set to obtain the final recommendation list.
2. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: In step 3.1), the similarity between users and between products is expressed by formula (6): In this example, taking the calculation of user similarity as an example, i and j represent the attribute vectors of different users, sim(i,j) represents the similarity between the attributes of user i and user j; t represents an attribute in the user attribute category T, R i,t and R j,t They represent the evaluation of user i and j on attribute t respectively; the calculation method of product similarity is the same as above.
3. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: The specific steps for selecting the initial cluster center in step 3.2 are as follows: first, randomly select a user or product attribute vector as the first initial cluster center, then traverse all user or product attribute vectors, select the user or product attribute vector that is least similar to the existing cluster center as the next initial cluster center, repeat the above steps, and finally, select all initial cluster centers.
4. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: The specific steps of the noise reduction process in step 4 are as follows: Compare the similarity between the vector and different cluster centers to determine whether it is a noise point, which can be expressed as formula (8): Among them, k i represents the cluster center of the cluster to which vector i belongs; K represents the cluster center of the cluster to which vector i belongs; i The cluster center set outside k l represents the cluster center in K; sim(k i ,i) represents the cluster center k of the vector i and its cluster i Similarity; sim(k l ,i) represents the vector i and cluster center k l Similarity; D i,l represents sim(k i ,i) and sim(k l ,i) Absolute value of the difference; β represents the threshold; If there is D i,j ≤β, the target vector i has a high similarity with multiple cluster centers and should be identified as a noise point; if l ∈K exists D i,l ≥β, then vector i is not identified as a noise point; then the locations of all noise points are recorded, and further, no prediction is performed in the subsequent prediction evaluation to achieve the effect of noise reduction.
5. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: In the step 5, the fusion factor α is set to 0.
5.
6. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: In step 1, the left inverse matrix (3) and the right inverse matrix (4) shown below are used to implement the above formula: H L =(H T ×H) -1 ×H T (3) H R =H T ×(H T ×H) -1 (4) Where H is an m-row n-column matrix: When m ≥ n, H has a left inverse matrix H L , the calculation formula is (3); when m≤n, H has a right inverse matrix H R , the calculation formula is (4); therefore, when U e When the number of rows ≥ the number of columns, substitute into formula (3) to obtain When P e When the number of rows ≤ the number of columns, substitute into formula (4) to obtain U e With P e The selection of obviously meets the above conditions; therefore, the C matrix can be obtained through formula (2), which is universal and can be combined with the user attribute matrix and product attribute matrix in the system to complete the cold start.
7. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: The step 3 is as follows: 3.1) Use the cosine similarity theorem to calculate the similarity between users and products. 3.2) In K-Means clustering, similar users or products are clustered using cluster centers; Select the k points that are most dissimilar to each other as the initial cluster centers; 3.3) A cluster is a container for user or product vectors gathered at the cluster center, storing user or product vectors with high similarity. To stabilize the clustering results, the mean of all vectors in the cluster is calculated as the new cluster center and compared with the original cluster center. This method is expressed as formula (7): Among them, K old is the original cluster center, K new is the new cluster center, n is the number of vectors in cluster B, and b is the vector in cluster B; First, the similarity between the vector and each initial cluster center is calculated using formula (7); then, the vector is assigned to the cluster where the most similar cluster center is located; finally, the mean of the sum of all vectors in the cluster is calculated as the new cluster center K. new :If K new =K old , it indicates that the cluster is stable, otherwise K new Replace K old , repeat the above steps until all clusters are stable.
8. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: The step 5 is specifically as follows: based on the clustering results of the completed user clustering and product clustering, a neighbor set N of the user and the product can be found. The neighbor set N stores attribute vectors of other users or products that are highly similar to the user or product. When performing prediction and evaluation, prediction will be performed based on the real evaluations of the neighbor users or products in the neighbor set; Introduce the user product evaluation matrix r, which contains the user's real evaluation of the product and the unknown evaluation of the product. The Top-n recommendation set can be obtained by combining the predicted evaluation of the unknown evaluation with the real evaluation. According to the clustering results and neighbor sets, the prediction evaluation based on user clustering can be expressed as formula (9), and the prediction evaluation based on product clustering can be expressed as (10): Among them, P i,u represents the predicted evaluation of user i on product u based on user clustering; represents the mean of user i’s evaluation; N i represents the neighbor set of user i; r j,u represents user j’s evaluation of product u; represents the mean of user j’s evaluation; sim(i,j) represents the similarity between the attributes of user i and user j; Among them, Q i,u N represents the predicted evaluation of user i on product u based on product clustering; u represents the neighbor set of commodity u; r i,v represents the evaluation of user i on product v; sim(v,u) represents the similarity between the attributes of product v and product u.
9. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: The step 6 is specifically as follows: Transaction data is introduced, which stores all the product IDs purchased by a single user within a certain period of time, and Apriori association rule mining is performed on the transaction data to mine the hidden relationship between products; the association between products is found; association rules are measured using two dimensions: support and confidence. The support is used to represent the probability that product item set A and product item set B are purchased at the same time, and the confidence is used to represent the probability that product item set B is purchased after product item set A is purchased. It can be expressed as formula (12): In association rule mining, potential associations between products can be found by selecting appropriate support and confidence. Specifically: First, the purchase data is preprocessed to extract all the products purchased by a single user as a purchase set; the support of association rule mining is set to δ∈(0,1) and the confidence is set to μ∈(0,1); rules with an association probability greater than δ and an associated purchase probability greater than μ are mined; then, the user's purchase record is retrieved. If the user has a purchase record and has purchased products with association rules, the n products with the highest confidence are selected to generate a product recommendation set N based on association rules for the user; If the user has no purchase record, an association rule search is performed based on the Top-n products, and the n products with the highest confidence are selected to generate a product recommendation set N based on association rules for the user.
10. The hybrid collaborative filtering recommendation method based on user clustering and product clustering according to claim 1, characterized in that: The step 7 is specifically as follows: The ID3 classification algorithm is used to find the most important product features for users. First, the information entropy is used to measure the uncertainty of H(S) users in purchasing products. The smaller the information entropy, the smaller the uncertainty of the data sample and the purer the data, which can be expressed as formula (13): H(S)=∑p i log(p i )(13) Among them, p i represents the probability of the i-th category, and S represents the data set; Then, the conditional entropy H(S|A) is used to examine the change in entropy after division according to a certain attribute feature. The conditional entropy, i.e., the feature expectation, can be used to represent the uncertainty of a certain feature. The smaller the feature expectation, the smaller the data uncertainty of the feature, which can be expressed as formula (14): Among them, A represents the attribute features of the data set S; c j Represented as the number of samples of category j in feature A; c j The ratio c represents the probability of all categories of category j in feature A; p ij is the probability that the jth category of feature A accounts for different categories of the data set S; Then, using the information gain G A (S) measures the contribution of attribute feature A to reducing the entropy of data set S. The larger the information gain, the more suitable it is for classifying S, which means that a certain feature is more important, expressed as formula (15): G A (S)=H(S)-H(S|A)(15) Select the maximum information gain G A The feature A of (S) is the parent node, and formulas (13)-(15) are repeated until G A (S) = 0, at this time, the ID3 decision tree is built, and the most important product features for users have been screened out; Finally, the user order data is segmented by attributes, and the proportion of the most important product features to the user in their purchase records is calculated. K , and get the weight W of each attribute i , expressed as formula (16): In this way, the user's preference weights for the detailed attributes of the product can be obtained, and then the user's ratings for different products can be generated. The n products with the highest user ratings are selected to generate a personalized recommendation set G, which is combined with the association rule recommendation set and the Top-n recommendation set to form the final recommendation list X, which can be expressed as formula (17) for accurate recommendation. M+N+G=X(17) Store the product IDs in the Top-n recommendation set M, the association rule recommendation set N, and the personalized recommendation set G into the recommendation list. Complete accurate recommendations for users.
Citation Information
Patent Citations
Personalized recommendation method based on combination of content and collaborative filtering
CN108334592A
Clustering-based personalized shopping guide system
CN111612583A