Product Recommendation Method, Device and Equipment Based on Content and Collaborative Filtering
By adopting a product recommendation method based on content and collaborative filtering in the medical product library, the problem that users find it difficult to find products that meet personalized preferences in a large number of products is solved, and high-precision and personalized product recommendations are achieved.
Patent Information
- Application Number
- CN202210435260.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-24
AI Technical Summary
In the medical product library, it is difficult for users to accurately search for products that meet personalized preferences. The prior art cannot meet users' personalized needs through keyword fuzzy search and historical visits sorting.
The product recommendation method based on content and collaborative filtering is adopted, and the user's product query text is obtained, the product similarity is calculated using TF-IDF algorithm and word segmentation technology, the identification feature word set and the description feature word set are extracted, and the weighted similarity value is calculated to form the first recommendation result; at the same time, the neighbor users with high correlation with user behavior are determined based on the collaborative filtering algorithm, the user's score for historical behavior is calculated to form the second recommendation result, and finally the target product recommendation result is determined in combination.
It improves the accuracy and personalization of product recommendations and can more accurately meet user needs.
Smart Images

Figure CN114610859B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technologies, and in particular, to a product recommendation method, apparatus, and device based on content and collaborative filtering. Background Art
[0002] In the process of the development of information technology applications in the medical industry, the medical product library includes a large amount of complex data, and different users have different preferences. Therefore, it is difficult to accurately search for products that meet the user's needs from a large amount of data.
[0003] Currently, fuzzy search is performed through keywords, and the obtained fuzzy search results are sorted in descending order of their historical access times and recommended to users. However, on the one hand, the accuracy of the products queried by the method of fuzzy query through keywords is not high, and on the other hand, the recommendation in the order of access times cannot meet the user's personalized preferences. Summary of the Invention
[0004] In view of this, the present application provides a product recommendation method, apparatus, and device based on content and collaborative filtering, which relates to the field of Internet technologies and can solve the problems of low accuracy when users search for target products among a large number of products and inability to meet the user's personalized preferences.
[0005] According to one aspect of the present application, there is provided a product recommendation method based on content and collaborative filtering, the method including:
[0006] Obtain a product query text sent by a query user for a target product, calculate a first similarity between the product query text and a preset product text by using a preset word segmentation technology and the TF-IDF algorithm, and determine a recommended product text corresponding to the preset product text with the first similarity greater than a preset similarity threshold;
[0007] Extract an identification feature word set and a description feature word set of the recommended product text, calculate a weighted similarity value between the recommended product text and the product query text according to the identification feature word set and the description feature word set, and form a first recommendation result in descending order of the weighted similarity value;
[0008] Determine neighbor users whose relevance to the query user's behavior is higher than a first preset threshold, and a historical behavior set of the neighbor users for products, calculate a score of the query user for the historical behavior set based on a collaborative filtering algorithm, and determine a second recommendation result according to the score;
[0009] Determine a target product recommendation result according to the first recommendation result and the second recommendation result.
[0010] According to another aspect of the present application, there is provided a product recommendation device based on content and collaborative filtering, and the device includes:
[0011] A screening module, configured to obtain a product query text sent by a query user for a target product, calculate a first similarity between the product query text and a preset product text by using a preset word segmentation technique and the TF-IDF algorithm, and determine a recommended product text corresponding to the preset product text with the first similarity greater than a preset similarity threshold;
[0012] A first recommendation module, configured to obtain an identification feature word set and a description feature word set of the recommended product text, calculate a weighted similarity value between the recommended product text and the product query text according to the identification feature word set and the description feature word set, and form a first recommendation result in descending order of the weighted similarity value;
[0013] A second recommendation module, configured to determine neighbor users with a behavior correlation higher than a first preset threshold with the query user, and a historical behavior set of the neighbor users for products, calculate a score of the query user for the historical behavior set based on a collaborative filtering algorithm, and determine a second recommendation result according to the score;
[0014] A determination module, configured to determine a target product recommendation result according to the first recommendation result and the second recommendation result.
[0015] According to yet another aspect of the present application, there is provided a non-volatile readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned product recommendation method based on content and collaborative filtering is implemented.
[0016] According to still another aspect of the present application, there is provided a computer device, including a non-volatile readable storage medium, a processor, and a computer program stored on the non-volatile readable storage medium and executable on the processor, and when the processor executes the program, the above-mentioned product recommendation method based on content and collaborative filtering is implemented.
[0017] With the above technical solution, the present application discloses a product recommendation method, device and equipment based on content and collaborative filtering. First, the present application obtains a product query text sent by a query user for a target product, calculates a first similarity between the product query text and a preset product text by using a preset word segmentation technology and the TF-IDF algorithm, and determines the preset product text with the corresponding first similarity greater than a preset similarity threshold as a recommended product text; further, extracts an identification feature word set and a description feature word set of the recommended product text, calculates a weighted similarity value between the recommended product text and the product query text according to the identification feature word set and the description feature word set, and forms a first recommendation result in the order of the weighted similarity value from large to small; in addition, determines neighbor users with a correlation higher than a first preset threshold with the query user's behavior, and a historical behavior set of the neighbor users for the product, calculates the query user's score for the historical behavior set based on the collaborative filtering algorithm, and determines a second recommendation result according to the score; finally, determines a target product recommendation result according to the first recommendation result and the second recommendation result. Through the technical solution in the present application, a first recommendation result for the target product is obtained starting from the product query text, a second recommendation result for the target product is obtained from the perspective of neighbor users with a high correlation with the query user's behavior, and then the first recommendation result and the second recommendation result are used together to determine the target product recommendation result, and recommendations are made for the query user through multiple dimensions comprehensively, with high recommendation accuracy and meeting the personalized needs of the query user.
[0018] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0020] Figure 1 Shows a schematic flowchart of a product recommendation method based on content and collaborative filtering provided by an embodiment of the present application;
[0021] Figure 2 Shows a schematic flowchart of another product recommendation method based on content and collaborative filtering provided by an embodiment of the present application;
[0022] Figure 3 Shows a schematic structural diagram of a product recommendation device based on content and collaborative filtering provided by an embodiment of the present application;
[0023] Figure 4The structural schematic diagram of another product recommendation device based on content and collaborative filtering provided by the embodiments of the present application is shown. Detailed implementation manners
[0024] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0025] In view of the current problems, the embodiments of the present application provide a product recommendation method based on content and collaborative filtering, as Figure 1 shown, the method includes:
[0026] 101. Obtain a product query text sent by a query user for a target product, calculate a first similarity between the product query text and a preset product text by using a preset word segmentation technique and the TF-IDF algorithm, and determine the preset product text corresponding to the first similarity greater than a preset similarity threshold as a recommended product text.
[0027] Among them, the target product is the product required by the query user, the product query text is the product manual, etc. of the target product, and the preset product text exists in the product database and is used to match the product query text, so as to determine the preset product text that meets the condition that the first similarity is greater than the preset similarity threshold. Among them, the preset product text can also be a product manual, etc.
[0028] For this embodiment, the preset word segmentation technique can be any existing word segmentation technique, such as the CRF word segmenter, the IKAnalyzer word segmenter, etc. The product query text is segmented by the preset word segmentation technique to obtain a product query feature word set including at least one product query feature word, and the product query feature vector corresponding to the product query feature word set is calculated by the TF-IDF algorithm. Similarly, the preset product text is segmented by the preset word segmentation technique to obtain a preset product feature word set including at least one preset product feature word, and the preset product feature vector corresponding to the preset product feature word set is calculated by the TF-IDF algorithm.
[0029] Among them, TF-IDF is a commonly used information weighting technique, which is widely applied in the fields of information retrieval and data mining. The TF-IDF value can be used to evaluate whether a certain feature word in a text is a keyword of the text. The larger the TF-IDF value, the more important the feature word is to the text, that is, the feature word is a keyword of the text. Just because a certain feature word has a high word frequency in the text does not mean it is a keyword of the text. Therefore, the TF-IDF value is the product of the word frequency TF of a certain feature word in the text and the inverse document frequency IDF corresponding to the feature word. For example, the most common feature words such as "of, is, in" are given the smallest IDF, and the less common feature words such as "flu, virus" are given a larger IDF. Therefore, the product query feature vector calculated by the TF-IDF algorithm includes the TF-IDF value corresponding to each product query feature word, and the preset product feature vector calculated by the TF-IDF algorithm includes the TF-IDF value corresponding to each preset product feature word.
[0030] Further, the first similarity between the product query feature vector and the preset product feature vector is calculated using a preset similarity calculation formula, and the preset similarity calculation formula may include a cosine calculation formula, which is described as follows:
[0031] In the formula, is the product query feature vector, is the preset product feature vector, is the first similarity.
[0032] After calculating the first similarity between the product query feature vector and each preset product feature vector, the preset product texts with the corresponding first similarity greater than the preset similarity threshold are screened out and determined as recommended product texts, and a preliminary screening is carried out from a large number of preset product texts, so as to further match the preliminarily screened recommended product texts with the product query text, improving the efficiency and accuracy of product recommendation.
[0033] 102. Extract the identification feature word set and the description feature word set of the recommended product text, calculate the weighted similarity value between the recommended product text and the product query text according to the identification feature word set and the description feature word set, and form a first recommendation result in the order from large to small of the weighted similarity value.
[0034] Among them, the recommended product feature words after word segmentation of the recommended product text are grouped according to whether they belong to identifying words or descriptive words. The recommended product feature words belonging to identifying words constitute the identifying feature word set, and the recommended product feature words belonging to descriptive words constitute the descriptive feature word set. For example, if the first feature word of the recommended product text 1 belongs to a descriptive word, then the first feature word is classified into the descriptive feature word set; if the second feature word of the recommended product text 1 belongs to an identifying word, then the second feature word is classified into the identifying feature word set... Therefore, the identifying feature word set of the recommended product text 1 is {the second feature word, the third feature word, the fifth feature word...}, and the descriptive feature word set of the recommended product text 1 is {the first feature word, the fourth feature word, the sixth feature word...}. Among them, identifying words are such as "flu, virus", and descriptive words are such as "of, in, is". The beneficial effect of such classification is to improve the accuracy of the first recommendation result by assigning a smaller proportion to the descriptive words of each recommended product text and a larger proportion to the identifying words of each recommended product text.
[0035] For this embodiment, calculating the weighted similarity value between the recommended product text and the product query text based on the identifying feature word set and the descriptive feature word set is further to assign a smaller proportion to the descriptive feature word set and a larger proportion to the identifying feature word set, and calculate the weighted similarity value between each recommended product text and the product query text again, so as to sort multiple recommended product texts to obtain the first recommendation result. The steps of the embodiment for calculating the weighted similarity value may include: calculating the first intersection of the identifying feature word set and the product query feature word set, and the second intersection of the descriptive feature word set and the product query feature word set. Calculating the third weight value of the identifying feature word set relative to the product query feature word set according to the first intersection, and calculating the fourth weight value of the descriptive feature word set relative to the product query feature word set according to the second intersection, and using a preset coefficient to weight the third weight value and the fourth weight value to obtain the weighted similarity value between the recommended product text and the product query text.
[0036] The first recommendation result includes each recommended product text and the corresponding weighted similarity value. Sort the recommended product texts in descending order of the weighted similarity value, and the recommended product text corresponding to the largest weighted similarity value ranks first in the first recommendation result.
[0037] 103. Determine neighbor users whose relevance to the query user behavior is higher than the first preset threshold, and the historical behavior set of the neighbor users for the product. Calculate the score of the query user for the historical behavior set based on the collaborative filtering algorithm, and determine the second recommendation result according to the score.
[0038] Among them, the collaborative filtering algorithm discovers the preference bias of the query user based on the mining of the historical behavior data of the query user and predicts the products required by the query user. The main implementation methods include: calculating neighbor users who have common needs with the query user and making recommendations based on the historical behavior data of these neighbor users. The collaborative filtering algorithm can still help the query user make recommendations when there is no product query text or the product query text is not precise enough.
[0039] For this embodiment, the neighbor users who have common needs with the query user are reflected in the neighbor users whose behavior correlation with the query user is higher than the first preset threshold. Specifically, according to the preset correlation coefficient calculation formula, such as the cosine calculation formula, calculate the correlation coefficient between the query user and other users, compare the calculated correlation coefficient with the first preset threshold, and take the other users with a correlation coefficient higher than the first preset threshold as neighbor users.
[0040] The historical behavior set of neighbor users for products includes: the set of behaviors to be predicted, and the adjacent behavior set of the set of behaviors to be predicted. Among them, the set of behaviors to be predicted exists in the behavior set of neighbor users but does not exist in the behavior set of the query user. The determination of the adjacent behavior set of the set of behaviors to be predicted specifically includes: calculating the second similarity between the set of behaviors to be predicted and other behavior sets according to the k-nearest neighbor algorithm or the k-means algorithm, and determining the other behavior sets with a second similarity greater than the second preset threshold as the adjacent behavior set. Finally, calculate the score of the query user for the historical behavior set based on the collaborative filtering algorithm, specifically including: calculating the first score of the query user for the set of behaviors to be predicted based on the user-based collaborative filtering algorithm, calculating the second score of the query user for the adjacent behavior set based on the item-based collaborative filtering algorithm. Finally, calculate the weighted sum of the first score and the second score to obtain the score of the query user for the historical behavior set, and determine the second recommendation result according to the score.
[0041] The second recommendation result includes the ranking of the scores and the products corresponding to the scores. The one ranked first in the second recommendation result is the maximum score and the product corresponding to the maximum score.
[0042] 104. Determine the target product recommendation result according to the first recommendation result and the second recommendation result.
[0043] For this embodiment, as a preferred implementation method, the first recommendation result and the second recommendation result can be weighted and calculated to obtain the target product recommendation result. Among them, the first recommendation result calculated according to the preset word segmentation technology and the TF-IDF algorithm is based on the text content, and the second recommendation result calculated according to the collaborative filtering algorithm is based on the user behavior. By weighting, the recommendation results obtained from these two dimensions are combined to comprehensively obtain the target product recommendation result, which is more accurate than the recommendation result obtained under a single dimension.
[0044] The present application discloses a product recommendation method, device and equipment based on content and collaborative filtering. First, a product query text sent by a query user for a target product is obtained, and a first similarity between the product query text and a preset product text is calculated by using a preset word segmentation technology and the TF-IDF algorithm. The preset product text corresponding to the first similarity greater than a preset similarity threshold is determined as a recommended product text. Further, an identification feature word set and a description feature word set of the recommended product text are extracted, a weighted similarity value between the recommended product text and the product query text is calculated according to the identification feature word set and the description feature word set, and a first recommendation result is formed in the order of the weighted similarity value from large to small. In addition, neighbor users with a correlation higher than a first preset threshold with the query user's behavior are determined, and a historical behavior set of the neighbor users for the product is determined. A score of the query user for the historical behavior set is calculated based on the collaborative filtering algorithm, and a second recommendation result is determined according to the score. Finally, a target product recommendation result is determined according to the first recommendation result and the second recommendation result. Through the technical solution in the present application, a first recommendation result for the target product is obtained starting from the product query text, a second recommendation result for the target product is obtained from the perspective of neighbor users with a high correlation with the query user's behavior, and then the first recommendation result and the second recommendation result are used together to determine the target product recommendation result. Recommendations are made for the query user comprehensively from multiple dimensions, with high recommendation accuracy and meeting the personalized needs of the query user.
[0045] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another product recommendation method based on content and collaborative filtering is provided, as Figure 2 shown. This method includes:
[0046] 201. Obtain a product query text sent by a query user for a target product, perform word segmentation processing on the product query text according to a preset word segmentation technology to obtain a product query feature word set, and perform word segmentation processing on the preset product text to obtain a preset product feature word set.
[0047] For this embodiment, in a specific application scenario, the product query text may include: product manuals, medical diagnosis books, etc. The product query text is subjected to word segmentation processing by using a preset word segmentation technology. Specifically: the product query feature word set {product query feature word 1, product query feature word 2, product query feature word 3...}. For example, the medical diagnosis book includes a detailed description text of the condition. After word segmentation processing, the product query feature word set {oral ulcer, recurrent attack, facial symmetry, recurrent aphthous ulcer...} is obtained. Similarly, each preset product text is subjected to word segmentation processing to obtain a corresponding preset product feature word set {preset product feature word 1, preset product feature word 2, preset product feature word 3...}. For example, the preset product feature word set 1 {recurrent oral ulcer, herpetic oral ulcer, suppressing immune function...}.
[0048] 202. Calculate the product query feature vector corresponding to the product query feature word set and the preset product feature vector corresponding to the preset product feature word set by using the TF-IDF algorithm.
[0049] For this embodiment, as a preferred implementation manner, it may include: calculating the first weight value of each product query feature word in the product query feature word set for the product query text and the second weight value of each preset product feature word in the preset product feature word set for the preset product text; constructing a product query feature vector including the product query feature word and the corresponding first weight value, and constructing a preset product feature vector including the preset product feature word and the corresponding second weight value.
[0050] Specifically, the TF-IDF algorithm includes term frequency tf calculation and inverse document frequency idf calculation. Further, multiplying the term frequency tf by the inverse document frequency idf obtains the weight value of the feature word for the text.
[0051] The term frequency of the product query feature word represents the number of times the product query feature word appears in the product query feature word set. Since each product query text has different lengths, each term frequency is normalized. Therefore, in the term frequency calculation formula, it is divided by ∑ k n k,j .
[0052] The term frequency calculation formula is described as:
[0053] where i represents the product query feature word, j represents the product query feature word set, tf i,j represents the term frequency of i in the set j, n i,j represents the number of times i appears in the set j, ∑ k n k,j represents the total number of times all words in the set j appear.
[0054] The inverse document frequency calculation formula is described as:
[0055] where i represents the product query feature word, j represents the product query feature word set, idf i represents the inverse document frequency of i in the set j, |D| represents the total number of product texts in the database, |{j:t j ∈d j}| represents the number of product texts in which the product query feature word i appears. |{j:t j ∈d j}| is smaller, the IDF value is larger, and the text discrimination effect of the product query feature word i is better. On the contrary, the smaller the IDF value, the worse the text discrimination effect of the product query feature word.
[0056] The calculation formula of TF-IDF value is described as: tfidf i,j = tf i,j × idf i
[0057] Among them, tfidf i,j is the first weight value of the product query feature word i.
[0058] The specific implementation process of calculating the second weight value of each preset product feature word using the TF-IDF algorithm can refer to the process of calculating the first weight value of each product query feature word using the TF-IDF algorithm.
[0059] Construct a product query feature vector: {(product query feature word 1, the first weight value of product query feature word 1), (product query feature word 2, the first weight value of product query feature word 2), (product query feature word 3, the first weight value of product query feature word 3)...}.
[0060] Similarly, construct a preset product feature vector: {(preset product feature word 1, the second weight value of preset product feature word 1), (preset product feature word 2, the second weight value of preset product feature word 2), (preset product feature word 3, the second weight value of preset product feature word 3)...}.
[0061] 203. Calculate the first similarity between the product query feature vector and the preset product feature vector using a preset similarity calculation formula, and determine the preset product text with the corresponding first similarity greater than the preset similarity threshold as the recommended product text.
[0062] For this embodiment, the preset similarity calculation formula may include a cosine calculation formula, and the cosine calculation formula is described as:
[0063] In the formula, is the product query feature vector, is the preset product feature vector, is the first similarity.
[0064] After calculating the first similarity between the product query feature vector and each preset product feature vector, screen out the preset product text with the corresponding first similarity greater than the preset similarity threshold and determine it as the recommended product text, and conduct a preliminary screening from a large number of preset product texts, so as to further match the preliminarily screened recommended product text with the product query text, improving the efficiency and accuracy of product recommendation.
[0065] 204. Extract the identification feature word set and description feature word set of the recommended product text, and calculate the first intersection of the identification feature word set and the product query feature word set, as well as the second intersection of the description feature word set and the product query feature word set.
[0066] For this embodiment, the recommended product feature words obtained after word segmentation of the recommended product text can be grouped according to whether they belong to identifying words or descriptive words. The recommended product feature word set is divided into an identifying feature word set and a descriptive feature word set. Calculate the first intersection of the identifying feature word set and the product query feature word set. Taking the recommended product text 1 as an example, if the second feature word in the identifying feature word set has no intersection with the product query feature word set, then the second feature word is not included in the first intersection. If the third feature word in the identifying feature word set has an intersection with the product query feature word set, then the third feature word is included in the first intersection. Therefore, the first intersection is {the third feature word...}. Similarly, calculate the second intersection of the descriptive feature word set and the product query feature word set.
[0067] 205. Calculate the third weight value of the identifying feature word set relative to the product query feature word set according to the first intersection, and calculate the fourth weight value of the descriptive feature word set relative to the product query feature word set according to the second intersection.
[0068] For this embodiment, the calculation process of the third weight value is: tfidf w = ∑ t∈n tf t,w × idf t,w
[0069] where tfidf w represents the third weight value of the identifying feature word set w relative to the product query feature word set j, and n represents the first intersection of the identifying feature word set w and the product query feature word set j. tf t,w represents the TF value of the identifying feature word t in w, and idf t,w represents the IDF value of the identifying feature word t in w.
[0070] The calculation process of the fourth weight value is: tfidf v = ∑ t∈m tf t,v × idf t,v
[0071] where tfidf v represents the fourth weight value of the descriptive feature word set v relative to the product query feature word set j, and m represents the second intersection of the descriptive feature word set v and the product query feature word set j. tf t,v represents the TF value of the descriptive feature word t in v, and idf t,v represents the IDF value of the feature word t in v.
[0072] 206. Use a preset coefficient to weight the third weight value and the fourth weight value to obtain a weighted similarity value between the recommended product text and the product query text, and form a first recommendation result in descending order of the weighted similarity value.
[0073] For this embodiment, the calculation process of the weighted similarity value between each recommended product text and the product query text is: C = λtfidf w +(1 - λ)tfidf v
[0074] where λ is a preset coefficient, C is the weighted similarity value, and tfidf w represents the third weight value, and tfidf v represents the fourth weight value.
[0075] The first recommended result includes the weighted similarity value and the recommended product text corresponding to the weighted similarity value. Among them, the one ranked first in the first recommended result is the largest weighted similarity value and the recommended product text corresponding to the largest weighted similarity value. For example, the first recommended result includes: {(Triamcinolone Acetonide Ointment, 0.5), (Domiphen Bromide Buccal Tablets, 0.3), (Lidocaine Gel, 0.15)...}.
[0076] The effect of using the preset coefficient to weight the third weight value and the fourth weight value to obtain the weighted similarity value between the recommended product text and the product query text is as follows: by assigning a smaller coefficient (1 - λ) to the descriptive words of each recommended product text and a larger coefficient λ to the identifying words of each recommended product text, the interference of the descriptive words on the first recommended result is reduced, thereby improving the accuracy of the first recommended result.
[0077] Making recommendations based on the product query text of the query user. On the one hand, it has nothing to do with the personal data of the query user, so there are no cold start and new user problems. On the other hand, each preset product text has the possibility of being recommended, regardless of the storage time and order of the preset product text information, so there are no new project problems. Finally, compared with the method of directly performing fuzzy queries in the database using keywords, this method searches based on content text that is closer to the needs of the query user, so the first recommended result is more accurate.
[0078] 207. Determine neighbor users whose relevance to the query user's behavior is higher than the first preset threshold, and the historical behavior set of the neighbor users for products.
[0079] Among them, the historical behavior set includes the set of behaviors to be predicted where the behavior sets of the neighbor users are different from the behavior set of the query user, and the adjacent behavior set adjacent to the set of behaviors to be predicted.
[0080] For this embodiment, as a preferred implementation, the steps of determining neighbor users whose relevance to the query user behavior is higher than the first preset threshold include: calculating the correlation coefficient between the query user and other users using a preset correlation coefficient calculation formula, and determining other users with a correlation coefficient greater than the first preset threshold as neighbor users. Specifically, the preset correlation coefficient calculation formula can be expressed as:
[0081]
[0082] The set of products co-called, r u,p and r b,p respectively represent the historical call times of u and b for the co-called product p, and respectively represent the average historical call times of u and b for the products in the I u ∩I b set.
[0083] The larger the value of S(u,b), the greater the correlation coefficient between u and b. The value range of S(u,b) is generally [-1, +1]. Determine other users b with S(u,b) greater than the preset first threshold as neighbor users h.
[0084] The specific steps of determining the set of behaviors to be predicted where the set of behaviors of neighbor users is different from the set of behaviors of the query user include: determining the difference between the set of behaviors of neighbor users and the set of behaviors of the query user as the set of behaviors to be predicted. Specifically, the set of behaviors to be predicted exists in the set of behaviors of neighbor users but does not exist in the set of behaviors of the query user.
[0085] The specific steps of determining the adjacent behavior set adjacent to the set of behaviors to be predicted include: calculating the second similarity between the set of behaviors to be predicted and other behavior sets according to the k-nearest neighbor algorithm or the k-means algorithm, and determining other behavior sets with a second similarity greater than the second preset threshold as the adjacent behavior set. Specifically, the k-nearest neighbor algorithm or the k-means algorithm can refer to the prior art and will not be elaborated here.
[0086] 208. Calculate the first score of the query user's behavior towards the set to be predicted based on the user-based collaborative filtering algorithm.
[0087] The main idea of the user-based collaborative filtering algorithm is: find neighbor users whose relevance to the query user behavior is higher than the first preset threshold. Based on the similarity between the historical behaviors of the query user and neighbor users, therefore, historical behaviors that neighbor users have had but the query user has not had may be historical behaviors that the query user will have. Among them, historical behaviors that neighbor users have had but the query user has not had are reflected in the set to be predicted in this embodiment.
[0088] For this embodiment, the first scoring calculation formula is described as:
[0089] Among them, p(u, i) represents the first score of the query user u for the product i in the set of behaviors to be predicted. represents the average number of calls of the query user u for the product, s(u, h) represents the correlation coefficient between u and h, and p h,q represents the score of the neighbor user h for i, and n is the number of neighbor users.
[0090] 209. Calculate the second score of the query user for the adjacent behavior set based on the item-based collaborative filtering algorithm, and determine the second recommendation result according to the first score and / or the second score.
[0091] Among them, the adjacent behavior set is the adjacent behavior set adjacent to the set of behaviors to be predicted. For this embodiment, the second scoring calculation formula is described as:
[0092]
[0093] Among them, p(u, i) represents the second score of the query user u for the product i in the set to be predicted. represents the average number of calls of the neighbor user h for i, and i k is the adjacent behavior set. is the average number of calls of the neighbor user h for the adjacent behavior set i k of, and s(i, i k ) represents the second similarity between i and i k . represents the number of times the query user u calls the adjacent behavior set i k .
[0094] As an implementation manner, the second recommendation result may only include the first score, or may only include the second score, or may perform a weighted calculation on the first score and the second score. The role of weighted calculating the first score and the second score is: combining two collaborative filtering algorithms improves the accuracy of the second recommendation result.
[0095] The calculation formula for weighting the first score and the second score is described as: P = μp(u, i)+(1 - μ)p(u, i)
[0096] Among them, μ is the weighting coefficient and P is the score.
[0097] Sort the products corresponding to the adjacent behavior sets from largest to smallest according to the scores to obtain the second recommendation result. Specifically, the second recommendation result is: Product 1, score is 16; Product 7, score is 10; Product 5, score is 3...
[0098] Further, before sorting the products corresponding to adjacent behavior sets from large to small according to the scores, it also includes: deleting products with scores less than or equal to 0. Because query users will not search for products with scores less than 0, only products with scores greater than 0 are saved.
[0099] 210. Determine the target product recommendation result according to the first recommendation result and the second recommendation result.
[0100] Before weighting the first recommendation result obtained in step 206 and the second recommendation result obtained in step 209, it also includes: using normalization to compress the scores into the range of (0, 1).
[0101] For example, the second recommendation result is: Product 1, score before normalization is 16; Product 7, score before normalization is 10; Product 5, score before normalization is 3...
[0102] The second recommendation result after normalization is: Product 1, score is 0.53, Product 7, score is 0.33, Product 5, score is 0.1...
[0103] For this embodiment, as a preferred implementation manner, calculate the union products of the products corresponding to the first recommendation result and the products corresponding to the second recommendation result; use a preset third coefficient to weight the weighted similarity value and score of the union products to obtain the target recommendation value; sort the target recommendation values in descending order to obtain the target product recommendation result.
[0104] For example, the first recommendation result calculates the weighted similarity value corresponding to the product, and the second recommendation result calculates the score corresponding to the product. For example, the weighted similarity value of Product 1 in the first recommendation result is 0.5, and the score of Product 1 in the second recommendation result is 0.53. Then the target recommendation value is 0.5 * preset third coefficient + 0.53 * (1 - preset third coefficient). If the weighted similarity value of Product 2 in the first recommendation result is 0.1, but Product 2 is not in the second recommendation result, then the target recommendation value is 0.1 * preset third coefficient + 0 * (1 - preset third coefficient).
[0105] The present application discloses a product recommendation method, device and equipment based on content and collaborative filtering. First, the present application obtains a product query text sent by a query user for a target product, calculates a first similarity between the product query text and a preset product text by using a preset word segmentation technology and the TF-IDF algorithm, and determines the preset product text corresponding to the first similarity greater than a preset similarity threshold as a recommended product text; further, extracts an identification feature word set and a description feature word set of the recommended product text, calculates a weighted similarity value between the recommended product text and the product query text according to the identification feature word set and the description feature word set, and forms a first recommendation result in the order of the weighted similarity value from large to small; in addition, determines neighbor users whose correlation with the query user's behavior is higher than a first preset threshold, and a historical behavior set of the neighbor users for the product, calculates the query user's score for the historical behavior set based on the collaborative filtering algorithm, and determines a second recommendation result according to the score; finally, determines a target product recommendation result according to the first recommendation result and the second recommendation result. Through the technical solution in the present application, a first recommendation result for the target product is obtained starting from the product query text, a second recommendation result for the target product is obtained from the perspective of neighbor users with high correlation with the query user's behavior, and then the first recommendation result and the second recommendation result are used together to determine the target product recommendation result, which is recommended to the query user comprehensively from multiple dimensions, with high recommendation accuracy and meeting the personalized needs of the query user.
[0106] Further, as Figure 1 and Figure 2 a specific implementation of the method shown, an embodiment of the present application provides a product recommendation device based on content and collaborative filtering, as Figure 3 shown, the device includes: a screening module 31, a first recommendation module 32, a second recommendation module 33, and a determination module 34;
[0107] The screening module 31 is configured to obtain a product query text sent by a query user for a target product, calculate a first similarity between the product query text and a preset product text by using a preset word segmentation technology and the TF-IDF algorithm, and determine the preset product text corresponding to the first similarity greater than a preset similarity threshold as a recommended product text;
[0108] The first recommendation module 32 is configured to extract an identification feature word set and a description feature word set of the recommended product text, calculate a weighted similarity value between the recommended product text and the product query text according to the identification feature word set and the description feature word set, and form a first recommendation result in the order of the weighted similarity value from large to small;
[0109] The second recommendation module 33 can be used to determine neighbor users whose relevance to the query user behavior is higher than a first preset threshold, and the historical behavior set of the neighbor users for the product, calculate the score of the query user for the historical behavior set based on the collaborative filtering algorithm, and determine the second recommendation result according to the score;
[0110] The determination module 34 can be used to determine the target product recommendation result according to the first recommendation result and the second recommendation result.
[0111] In a specific application scenario, in order to calculate the first similarity between the product query text and the preset product text by using the preset word segmentation technology and the TF-IDF algorithm, as Figure 4 shown, the screening module 31 may specifically include: a word segmentation unit 311, a first calculation unit 312, and a second calculation unit 313;
[0112] The word segmentation unit 311 can be used to perform word segmentation processing on the product query text according to the preset word segmentation technology to obtain a product query feature word set, and perform word segmentation processing on the preset product text to obtain a preset product feature word set;
[0113] The first calculation unit 312 can be used to calculate the product query feature vector corresponding to the product query feature word set and the preset product feature vector corresponding to the preset product feature word set by using the TF-IDF algorithm;
[0114] The second calculation unit 313 can be used to calculate the first similarity between the product query feature vector and the preset product feature vector by using a preset similarity calculation formula.
[0115] Correspondingly, in order to calculate the product query feature vector corresponding to the product query feature word set and the preset product feature vector corresponding to the preset product feature word set by using the TF-IDF algorithm, the first calculation unit 312 can specifically be used to calculate the first weight value of each product query feature word in the product query feature word set for the product query text and the second weight value of each preset product feature word in the preset product feature word set for the preset product text; construct a product query feature vector including the product query feature word and the corresponding first weight value, and construct a preset product feature vector including the preset product feature word and the corresponding second weight value.
[0116] In a specific application scenario, calculate the weighted similarity value between the recommended product text and the product query text according to the identification feature word set and the description feature word set, as Figure 4 shown, the first recommendation module 32 may specifically include: an intersection unit 321, a weight unit 322, and a first weighting unit 323;
[0117] An intersection unit 321, which can be used to calculate a first intersection of the identified feature word set and the product query feature word set, and a second intersection of the description feature word set and the product query feature word set;
[0118] A weight unit 322, which can be used to calculate a third weight value of the identified feature word set relative to the product query feature word set according to the first intersection, and a fourth weight value of the description feature word set relative to the product query feature word set according to the second intersection;
[0119] A first weighting unit 323, which can be used to weight the third weight value and the fourth weight value by a preset coefficient to obtain a weighted similarity value between the recommended product text and the product query text.
[0120] In a specific application scenario, the historical behavior set includes a to-be-predicted behavior set in which the behavior sets of neighbor users are different from those of the query user, and an adjacent behavior set adjacent to the to-be-predicted behavior set. In order to determine neighbor users whose behavior relevance to the query user is higher than a first preset threshold, and the historical behavior sets of neighbor users for products, as Figure 4 shown, the second recommendation module 33 may specifically include: a first screening unit 331, a first determination unit 332, and a second screening unit 333;
[0121] The first screening unit 331 can be used to calculate the correlation coefficient between the query user and other users by using a preset correlation coefficient calculation formula, and determine other users with the correlation coefficient greater than the first preset threshold as neighbor users;
[0122] The first determination unit 332 can be used to determine the difference between the behavior set of the neighbor user and the behavior set of the query user as the to-be-predicted behavior set;
[0123] The second screening unit 333 can be used to calculate a second similarity between the to-be-predicted behavior set and other behavior sets according to the k-nearest neighbor algorithm or the k-means algorithm, and determine other behavior sets with the second similarity greater than the second preset threshold as adjacent behavior sets.
[0124] In a specific application scenario, calculate the score of the query user for the historical behavior set based on the collaborative filtering algorithm, and determine the second recommendation result according to the score, as Figure 4 shown, the second recommendation module 33 may specifically further include: a first scoring unit 334, a second scoring unit 335, and a second screening unit 336;
[0125] The first scoring unit 334 can be used to calculate a first score of the query user for the to-be-predicted behavior set based on the user-based collaborative filtering algorithm;
[0126] The second scoring unit 335 can be used to calculate a second score of the query user for the adjacent behavior set based on the item-based collaborative filtering algorithm;
[0127] A second determination unit 336, which can be used to determine a second recommendation result according to the first score and / or the second score.
[0128] In a specific application scenario, a target product recommendation result is determined according to the first recommendation result and the second recommendation result. For example, Figure 4 as shown, the determination module 34 specifically may include: a union unit 341, a second weighting unit 342, and a recommendation unit 343;
[0129] The union unit 341 can be used to calculate the union product of the products corresponding to the first recommendation result and the products corresponding to the second recommendation result;
[0130] The second weighting unit 342 can be used to weight the weighted similarity value and the score of the union product by using a preset third coefficient to obtain a target recommendation value;
[0131] The recommendation unit 343 can be used to sort the target recommendation values in descending order to obtain a target product recommendation result.
[0132] It should be noted that for other corresponding descriptions of each functional unit involved in a product recommendation device based on content and collaborative filtering provided in this embodiment, reference can be made to Figures 1 to 2 the corresponding description, which will not be elaborated here.
[0133] Based on the method as Figures 1 to 2 shown above, correspondingly, this embodiment further provides a storage medium. The storage medium can specifically be volatile or non-volatile, and has computer-readable instructions stored thereon. When the readable instructions are executed by a processor, the above-mentioned method for product recommendation based on content and collaborative filtering as Figures 1 to 2 shown is implemented.
[0134] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product. The software product can be stored in a storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of this application.
[0135] Based on the method as Figures 1 to 2 shown above and Figure 3 、 Figure 4 the virtual device embodiment shown, in order to achieve the above purpose, this embodiment further provides a computer device. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the above-mentioned method for product recommendation based on content and collaborative filtering as Figures 1 to 2 shown.
[0136] Optionally, the computer device may further include a user interface, a network interface, a camera, a Radio Frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen and an input unit such as a keyboard. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0137] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not limit the physical device, and it may include more or fewer components, or combine certain components, or have different component arrangements.
[0138] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or by hardware.
[0140] By applying the technical solution of the present application, compared with the current existing technologies, the present application discloses a product recommendation method, device and equipment based on content and collaborative filtering. The present application first obtains a product query text sent by a query user for a target product, calculates a first similarity between the product query text and a preset product text by using a preset word segmentation technology and the TF-IDF algorithm, and determines the preset product text with the corresponding first similarity greater than a preset similarity threshold as a recommended product text; further, extracts an identification feature word set and a description feature word set of the recommended product text, calculates a weighted similarity value between the recommended product text and the product query text according to the identification feature word set and the description feature word set, and forms a first recommendation result in the order from large to small of the weighted similarity values; in addition, determines neighbor users with a correlation higher than a first preset threshold with the behavior of the query user, and a historical behavior set of the neighbor users for the product, calculates a score of the query user for the historical behavior set based on the collaborative filtering algorithm, and determines a second recommendation result according to the score; finally, determines a target product recommendation result according to the first recommendation result and the second recommendation result. Through the technical solution in the present application, a first recommendation result for the target product is obtained starting from the product query text, a second recommendation result for the target product is obtained from the perspective of neighbor users with a high correlation with the behavior of the query user, and then the first recommendation result and the second recommendation result are used together to determine the target product recommendation result, and recommendations are made for the query user through multiple dimensions comprehensively, with high recommendation accuracy and meeting the personalized needs of the query user.
[0141] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed to be located in one or more devices different from the present implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.
[0142] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenarios. The above-disclosed are only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.
Claims
1. A product recommendation method based on content and collaborative filtering, characterized in that, Including: Obtain the product query text sent by the query user for the target product, perform word segmentation on the product query text according to the preset word segmentation technology to obtain a product query feature word set, perform word segmentation on the preset product text to obtain a preset product feature word set, and based on the product query feature word set and the preset product feature word set, use the TF-IDF algorithm to calculate the first similarity between the product query text and the preset product text, and determine the preset product text corresponding to the first similarity greater than the preset similarity threshold as the recommended product text; Extract the identification feature word set and the description feature word set of the recommended product text, calculate the first intersection of the identification feature word set and the product query feature word set, and the second intersection of the description feature word set and the product query feature word set, calculate the weighted similarity value between the recommended product text and the product query text based on the first intersection and the second intersection, and form the first recommendation result in the order of the weighted similarity value from large to small; Determine neighbor users whose relevance to the query user's behavior is higher than the first preset threshold, and the historical behavior set of the neighbor users for products, calculate the first score and the second score of the query user for the historical behavior set based on the collaborative filtering algorithm, and determine the second recommendation result according to the first score and / or the second score; Calculate the union products of the products corresponding to the first recommendation result and the products corresponding to the second recommendation result, and determine the target product recommendation result based on the union products.
2. The method according to claim 1, characterized in that, The step of calculating the first similarity between the product query text and the preset product text by using the TF-IDF algorithm based on the product query feature word set and the preset product feature word set includes: Use the TF-IDF algorithm to calculate the product query feature vector corresponding to the product query feature word set and the preset product feature vector corresponding to the preset product feature word set; Use the preset similarity calculation formula to calculate the first similarity between the product query feature vector and the preset product feature vector.
3. The method according to claim 2, characterized in that, The step of calculating the product query feature vector corresponding to the product query feature word set and the preset product feature vector corresponding to the preset product feature word set by using the TF-IDF algorithm includes: Calculate the first weight value of each product query feature word in the product query feature word set for the product query text and the second weight value of each preset product feature word in the preset product feature word set for the preset product text; Construct a product query feature vector including the product query feature word and the corresponding first weight value, and construct a preset product feature vector including the preset product feature word and the corresponding second weight value.
4. The method according to claim 1, characterized in that, The step of calculating the weighted similarity value between the recommended product text and the product query text based on the first intersection and the second intersection includes: Calculate the third weight value of the identification feature word set relative to the product query feature word set according to the first intersection, and calculate the fourth weight value of the description feature word set relative to the product query feature word set according to the second intersection; Weight the third weight value and the fourth weight value using a preset coefficient to obtain the weighted similarity value between the recommended product text and the product query text.
5. The method according to claim 1, characterized in that The historical behavior set includes a to-be-predicted behavior set in which the behavior sets of the neighbor users are different from those of the query user, and an adjacent behavior set adjacent to the to-be-predicted behavior set. Determining neighbor users whose behavior correlation with the query user is higher than a first preset threshold, and the historical behavior set of the neighbor users for products includes: Calculate the correlation coefficient between the query user and other users using a preset correlation coefficient calculation formula, and determine other users with the correlation coefficient greater than the first preset threshold as neighbor users; Determine the difference between the behavior set of the neighbor users and the behavior set of the query user as the to-be-predicted behavior set; Calculate the second similarity between the to-be-predicted behavior set and other behavior sets according to the k-nearest neighbor algorithm or the k-means algorithm, and determine other behavior sets with the second similarity greater than a second preset threshold as the adjacent behavior set.
6. The method according to claim 5, wherein The calculating the first score and the second score of the query user for the historical behavior set based on the collaborative filtering algorithm includes: Calculate the first score of the query user for the to-be-predicted behavior set based on the user-based collaborative filtering algorithm; Calculate the second score of the query user for the adjacent behavior set based on the item-based collaborative filtering algorithm.
7. The method according to claim 1, wherein The determining the target product recommendation result based on the union product includes: Weight the weighted similarity value and the score of the union product using a preset third coefficient to obtain the target recommendation value; Sort the target product recommendation results in descending order according to the target recommendation value.
8. A product recommendation device based on content and collaborative filtering, characterized in that, including: A screening module, configured to obtain a product query text sent by a query user for a target product, perform word segmentation processing on the product query text according to a preset word segmentation technique to obtain a product query feature word set, perform word segmentation processing on a preset product text to obtain a preset product feature word set, calculate the first similarity between the product query text and the preset product text based on the product query feature word set and the preset product feature word set using the TF-IDF algorithm, and determine the preset product text corresponding to the first similarity greater than a preset similarity threshold as the recommended product text; A first recommendation module, configured to extract an identification feature word set and a description feature word set of the recommended product text, calculate a first intersection of the identification feature word set and the product query feature word set, and a second intersection of the description feature word set and the product query feature word set, calculate the weighted similarity value between the recommended product text and the product query text based on the first intersection and the second intersection, and form a first recommendation result in descending order according to the weighted similarity value; A second recommendation module, configured to determine neighbor users whose behavior correlation with the query user is higher than a first preset threshold, and the historical behavior set of the neighbor users for products, calculate the first score and the second score of the query user for the historical behavior set based on the collaborative filtering algorithm, and determine a second recommendation result according to the first score and / or the second score; A determination module, configured to calculate a union product of the product corresponding to the first recommendation result and the product corresponding to the second recommendation result, and determine a target product recommendation result based on the union product.
9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the content- and collaborative-filtering-based product recommendation method according to any one of claims 1 to 7.
10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the content- and collaborative-filtering-based product recommendation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Web service credible hybrid recommendation method considering timeliness
CN110390058A
Article similarity recommendation method and device
CN111061957A