Self-adaptive recommendation method and system based on Hellinger distance
By using the Hellinger distance in the recommendation system to calculate the distribution difference between user historical data and recommended results, and adjusting the objective function according to user activity, the asymmetry problem of existing recommendation calibration algorithms and the problem of decreasing user recommendation accuracy is solved, and higher recommendation accuracy and diversity are achieved.
Patent Information
- Application Number
- CN202510175506.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
The existing recommendation calibration algorithms have asymmetry problems, which makes model interpretation difficult, and for users with less historical data, the use of traditional methods will lead to a decrease in recommendation accuracy.
Adaptive recommendation method based on Hellinger distance is adopted, and the distribution difference between user historical data and recommended results is calculated, and the objective function is adjusted to adapt to user activity, thereby generating more accurate and diverse recommendation results.
This method can effectively avoid asymmetry problems, improve the accuracy of the recommendation system, and improve the recommendation effect for users with fewer historical data while maintaining the diversity of recommendation results.
Smart Images

Figure CN120104872A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a recommendation algorithm, and in particular to an adaptive recommendation method and system based on Hellinger distance. Background Art
[0002] With the continuous development of information technology, more and more information is available on the Internet, and the era of information explosion has arrived. Faced with complex information, people may find it difficult to find the information they need, or spend too much time to find useful information, so recommendation algorithms came into being. Recommendation algorithms are designed to help users reduce the time and energy wasted by browsing a large amount of invalid data. The recommendation system analyzes the user's historical behavior and predicts the user's needs with the help of recommendation algorithms, so as to recommend content that may be of interest to the user.
[0003] In 2018, Steck first proposed the concept of recommendation calibration. According to the distribution of items in the user's historical data set, the original recommendation list is reordered to make the genre distribution of the recommendation list as consistent as possible with the user's original interest distribution. The calibration effect is measured by KL divergence, which effectively solves the interest bias problem that has appeared in previous recommendation systems. However, the existing recommendation calibration algorithms have the following problems: (1) There is an asymmetry problem when using KL divergence to measure distribution differences, that is, the calculation results of C_KL(p,q) and C_KL(q,p) are not equivalent, and C is the calibration metric, which leads to ambiguity in the interpretation of the model. (2) In the objective function of the traditional recommendation calibration model, calibration is regarded as a key attribute of the recommendation list, and a large negative number λ is used to increase the penalty for calibration. However, for users with less historical data, their historical records cannot well reflect their true interest preferences. Using the formula will lead to a decrease in the accuracy of recommendations. Summary of the invention
[0004] Purpose of the invention: In view of the above problems, the present invention proposes an adaptive recommendation method and system based on Hellinger distance, which can improve the accuracy of the recommendation system and maintain the diversity of recommendation results.
[0005] Technical solution: The technical solution adopted by the present invention is an adaptive recommendation method and system based on Hellinger distance, including: using an adaptive recommendation model based on Hellinger distance to output recommendation results; the adaptive recommendation model includes:
[0006] For any item i in the user data set, calculate the category distribution p(g j |u);
[0007] Calculate the recommended result category distribution q(g j |u);
[0008] The Hellinger distance is used to calculate p(g j |u) and q(g j |u) The difference between the two distributions C HD (p,q);
[0009] According to the difference between the two distributions C HD (p,q) calculates the adaptive objective function.
[0010] Calculate the category distribution p(g j |u), the calculation formula is:
[0011]
[0012] In the formula, p(g j |u) is the category distribution of user u in its historical dataset H, ω ui is the weight of item i, H represents the historical data set, p(g j |i) represents the category g in item i j The proportion.
[0013] Calculate the recommended result category distribution q(g j |u), the calculation formula is:
[0014]
[0015] In the formula, q(g j |u) is the recommended result category distribution of user u in the user dataset, ω r(i) is the ranking position of item i in the recommendation, p(g j |i) represents the category g in item i j The proportion.
[0016] Category g in item i j The proportion p(g j |i), the calculation formula is:
[0017]
[0018] In the formula, |G i | represents the total number of categories to which item i belongs.
[0019] The Hellinger distance is used to calculate p(g j |u) and q(g j |u) The difference between the two distributions C HD (p,q), the calculation formula is:
[0020]
[0021] In the formula, C HD (p,q) is p(g j |u) and q(g j |u) the difference between the two distributions, g j is a category, and G represents the set of item categories.
[0022] According to the difference between the two distributions C HD (p, q) calculates the adaptive objective function, and the calculation formula is:
[0023]
[0024] In the formula, I * is the objective function, P u represents the total number of records of user u in the historical data set, P MAX is the maximum value of the total number of records, s(I) is the score of item i predicted by the recommendation system, and C HD (p,q(I)) means p(g j |u) and q(g j |u) is the difference between the two distributions, argmax means finding the variable value that makes the following formula reach the maximum value, I represents the set of items, and |I|=n means the number of elements in the set is n.
[0025] The present invention proposes an adaptive recommendation system based on Hellinger distance, comprising a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the adaptive recommendation method based on Hellinger distance is implemented.
[0026] The present invention provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the adaptive recommendation method based on the Hellinger distance when executing the computer program.
[0027] The present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the adaptive recommendation method based on the Hellinger distance is implemented.
[0028] The present invention provides a computer program product, comprising a computer program and / or instructions, which implement the adaptive recommendation method based on Hellinger distance when executed by a processor.
[0029] Beneficial effects: Compared with the prior art, the present invention uses the Hellinger distance to re-measure the differences between distributions in the field of recommendation algorithms, and adjusts the objective function according to the parameters adaptively generated based on the user's activity to make recommendations. Using the Hellinger distance to re-measure the differences between distributions can avoid the ambiguity problem caused by asymmetry. Using a dynamic, adaptively generated value to modify the objective function based on the activity of each user can better make recommendations for users with less historical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flow chart of the adaptive recommendation method based on Hellinger distance described in the present invention. DETAILED DESCRIPTION
[0031] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0032] The adaptive recommendation method based on Hellinger distance of the present invention has a flow chart as shown below: Figure 1 Taking the MovieLens 1M dataset as an example, Tables 1 and 2 show the user rating data and movie category data in the MovieLens 1M dataset. The user rating data and movie category data are used as the input of the model.
[0033] Table 1 Display of some user rating datasets in the MovieLens 1M dataset
[0034]
[0035]
[0036] Table 2. Movie category dataset display in MovieLens 1M dataset
[0037]
[0038] First, for any item (i.e., movie) i in the dataset, a category (i.e., movie category) g in item i j The proportion is as follows:
[0039]
[0040] Among them, |G i | represents the total number of categories to which item i belongs, g j represents a certain category. When item i belongs to only one category, p(g j |i)=1; when item i belongs to multiple categories |G i |, all categories are treated equally and divided equally.
[0041] On this basis, user u’s category distribution p(g j |u) is defined as shown in formula (2):
[0042]
[0043] Among them, p(g j |u) is the category distribution of a user u in the MovieLens 1M dataset on its historical dataset H, ω ui It is the weight of an item i in the user's historical data set, which is related to the degree of interaction between the user and item i.
[0044] For a user u in the MovieLens 1M dataset, the category distribution p(g j |u) definition, we can calculate the user's rating for each item in his historical dataset.
[0045] Next, in the generated recommendation list I, the category distribution q(g j |u) is defined as shown in formula (3):
[0046]
[0047] Among them, q(g j |u) is the recommended result category distribution of a user u in the MovieLens 1M dataset, ω r(i) It is based on the ranking position of item i in the recommendation. Relevant ranking indicators include mean reciprocity rating (MRR), etc. I is the recommendation list generated by the system for the user.
[0048] So we get p(g j |u) and q(g j |u), the Hellinger distance is used to measure the difference between the two distributions. The Hellinger distance is symmetrical. No matter which of the two is used as the reference distribution, the calculation results are the same and will not affect the final H(p,q). The calculation formula is shown in formula (4). The recommended calibration model based on the Hellinger distance is referred to as C HD Model.
[0049]
[0050] Where p(g j |u) represents the item category g in the historical item set H of a user u in the MovieLens 1M dataset. j The distribution of q(g j|u) represents the item category g in the recommendation list dataset I generated by the system for the user j distribution, G represents the set of item categories, each g j Represents a category, traverse each category in p(g j |u) and q(g j |u) between the two. C HD (p,q) is used to calculate p(g j |u) and q(g j |u) The difference between two distributions. The Hellinger distance uses the square roots of two probability distributions and takes the square root of the sum of their differences. The Hellinger distance is symmetric and can avoid ambiguity problems caused by asymmetry.
[0051] Above we get a new difference measure C HD (p,q), next we construct the adaptive recommendation calibration model objective function.
[0052] Traditional C KL The difference between the category distribution of user historical data in the model and the category distribution of user recommendation results is as follows, referred to as C KL (p,q):
[0053]
[0054] In the formula, C KL (p,q) represents the traditional C KL The distribution difference of the model, p and q are two distributions, p(g j |u) represents the item category g in the historical item set H of a user u in the MovieLens 1M dataset. j The distribution of q(g j |u) is the category distribution of recommendation results for a user u in the MovieLens 1M dataset.
[0055] Traditional C KL The objective function in the model is as follows:
[0056] I * = argmax I,|I|=n (1-λ)·s(I)-λ·C KL (p,q(I)) (6)
[0057] Among them, I * is the objective function, s(I) is the score of item i predicted by the recommendation system, s(I) = ∑ i∈I s(i). C KL For calibration metrics, better calibration means lower C KL, so a negative coefficient is used for it. When the calibration is not ideal, C KL is larger, thus reducing the value of the objective function. λ is the calibration metric coefficient. In addition, considering that calibration is a key attribute of the recommendation list, a larger λ value is adopted, and λ∈[0,1]. Here λ is a fixed value set in advance in each experiment.
[0058] (2) For users with different levels of activity, the ratio between the accuracy and calibration of the recommendation system should not be a fixed value. HD The model designs an adaptive objective function and makes the following changes:
[0059]
[0060] Where P u represents the total number of records in the historical data set of user u. The system first traverses the total number of records for each user in the data set, and then finds P MAX is the largest one among them. MAX =max u∈U P u The fourth power of the ratio of the two can be used to reflect the activity level of user u. The size of this ratio can reflect the activity level of the user.
[0061] Considering P MAX There may be a large number of common user history records P u Relative P MAX Too small, so Develop the fourth power to reduce the impact of extreme cases. The more historical records a user u has, the more active the user is. The closer the ratio is to 1, the higher the calibration degree C HD The impact on the objective function will also be greater. On the contrary, for users with less historical data, the ratio will be close to 0, which will be more biased towards the quality of the recommendation system itself.
[0062] Formula (7) is adaptive, and the model will adjust the weights according to the user's activity. Specifically, for users with more records in the historical data set, more attention should be paid to the quality of calibration to ensure consistency with the categories in the historical data set. Therefore, the model will pay more attention to the proportion of calibration metrics in the objective function. For those users with less historical data and less active, their historical data is sparse and cannot accurately reflect their true interests and hobbies. At this time, the model will pay more attention to the recommendation system's own predicted score for the user, reducing the proportion and influence of its calibration. This adaptability can achieve a balance between different types of users, improving the overall performance of the recommendation system and user satisfaction. As shown in Table 3.
[0063] Table 3 C of the present inventionHD The model is the weights adaptively generated for some users in the MovieLens 1M dataset
[0064]
[0065]
[0066] The weights in the table are the weights in formula (7)
[0067] Finally, the system calculates the predicted score for each user based on this objective function, and then uses a greedy algorithm to solve the problem of selecting the optimal recommendation set I from n items. * Problem, the greedy optimization algorithm starts with an empty set, each iteration adds a new item to it, and then removes it from the original set.
[0068] Data Analysis:
[0069] All experiments in this article are implemented in Python on a PC with Microsoft Windows 10, AMD Ryzen 55600 6-Core, 16.0GB memory, and 1TB hard drive.
[0070] The dataset used in this paper is the MovieLens 1M dataset, and the movies without category information in the dataset are deleted. The MovieLens 1M dataset consists of 6040 users' rating data on 3900 movies (a total of 1000209 data), which contains users' ratings of different movies and the category information of each movie. The ratings in the dataset are between 1 and 5, and each movie in the dataset belongs to at least one category and may belong to multiple categories. In the experiment of this paper, the dataset is divided into 80% for training set and 20% for test set.
[0071] In the experiment, a user-based collaborative filtering algorithm is designed to generate an original recommendation list for users. Considering the performance of the model, C KL The λ in the model is set to 0.6, and the number of the user’s neighbors n is set to 10 for the experiment. HD Model Comparison C KL The model is tested on Precision and nDCG. The larger the values of these two indicators are, the better the performance is. The calculation formulas for these two indicators are as follows:
[0072] Precision measures the accuracy of the system's predictions for users by calculating the percentage of correctly predicted items in the sample. The specific formula is as follows:
[0073]
[0074] Where |U| represents the number of sets of all users in the system, S(u) represents the set of recommendation list items generated by the system for the user, and T(u) represents the set of interaction record items of the user on the test set.
[0075] Normalized Cumulative Discount Gain (nDCG) is an algorithm used to measure and evaluate search results. The specific formula is as follows:
[0076]
[0077]
[0078] Where reli refers to whether item i appears in the test set, if it appears, it is 1, if not, it is 0. u It is proposed that when a highly relevant item appears at a lower position in the recommendation list, a penalty should be imposed on the evaluation score. The penalty ratio is related to the logarithm of the item's position. u This is the most ideal sorting, and the ratio of the two is used to indicate how close the current result is to the most ideal result.
[0079] In order to further increase persuasiveness, this paper will conduct multiple experiments on different recommendation list lengths K, setting the list lengths to 5, 10, 15, 20, and 25. The specific experimental results are shown in the table below.
[0080] Table 3 Experimental results on MovieLens 1M dataset
[0081]
[0082] As can be seen from Table 1, C HD Model Comparison C kL Model, when the length of the recommendation list k is 5, C KL The performance of the model is better than C HD Model, when the length of the recommendation list k is 10, the C HD The model is only slightly worse than C KL model, and they are almost the same in terms of nDCG. When the length of the recommendation list k gradually increases to 15, 20, and 25, the C HD The model performance is better than the initial C KL Model. It is observed that when K = 20, C HD The model is better than C in nDCG KL The model has improved by about 0.54%.
[0083] In one embodiment, an adaptive recommendation system based on Hellinger distance is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the adaptive recommendation method based on Hellinger distance when executing the computer program.
[0084] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the adaptive recommendation method based on Hellinger distance when executing the computer program.
[0085] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the adaptive recommendation method based on Hellinger distance is implemented.
[0086] In one embodiment, a computer program product is provided, including a computer program and / or instructions, which implement the adaptive recommendation method based on Hellinger distance when executed by a processor.
[0087] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0088] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0089] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0090] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
Claims
1. An adaptive recommendation method based on Hellinger distance, characterized in that: include: Adopt an adaptive recommendation model based on Hellinger distance to output recommendation results; The adaptive recommendation model includes: For any item i in the user data set, calculate the category distribution p(g j |u); Calculate the recommended result category distribution q(g j |u); The Hellinger distance is used to calculate p(g j |u) and q(g j |u) The difference between the two distributions C HD (p,q); According to the difference between the two distributions C HD (p,q) calculates the adaptive objective function.
2. The adaptive recommendation method based on Hellinger distance according to claim 1, characterized in that: Calculate the category distribution p(g j |u), the calculation formula is: In the formula, p(g j |u) is the category distribution of user u in its historical dataset H, ω ui is the weight of item i, H represents the historical data set, p(g j |i) represents the category g in item i j The proportion.
3. The adaptive recommendation method based on Hellinger distance according to claim 1, characterized in that: Calculate the recommended result category distribution q(g j |u), the calculation formula is: In the formula, q(g j |u) is the recommended result category distribution of user u in the user dataset, ω r(i) is the ranking position of item i in the recommendation, p(g j |i) represents the category g in item i j The proportion.
4. The adaptive recommendation method based on Hellinger distance according to claim 2 or 3, characterized in that: Category g in item i j The proportion p(g j |i), the calculation formula is: In the formula, |G i | represents the total number of categories to which item i belongs.
5. The adaptive recommendation method based on Hellinger distance according to claim 1, characterized in that: The Hellinger distance is used to calculate p(g j |u) and q(g j |u) The difference between the two distributions C HD (p,q), the calculation formula is: In the formula, C HD (p,q) is p(g j |u) and q(g j |u) the difference between the two distributions, g j is a category, and G represents the set of item categories.
6. The adaptive recommendation method based on Hellinger distance according to claim 1, characterized in that: According to the difference between the two distributions C HD (p, q) calculates the adaptive objective function, and the calculation formula is: In the formula, I * is the objective function, P u represents the total number of records of user u in the historical data set, P MAX is the maximum value of the total number of records, s(I) is the score of item i predicted by the recommendation system, and C HD (p,q(I)) means p(g j |u) and q(g j |u) is the difference between two distributions, argmax means finding the value of the variable when the function reaches its maximum value, I means the set of items, and |I|=n means the number of elements in the set is n.
7. An adaptive recommendation system based on Hellinger distance, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the adaptive recommendation method based on Hellinger distance according to any one of claims 1 to 6 is implemented.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the adaptive recommendation method based on Hellinger distance according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the adaptive recommendation method based on Hellinger distance according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program and / or instructions, characterized in that: When the computer program and / or the instruction is executed by a processor, the adaptive recommendation method based on Hellinger distance described in any one of claims 1 to 6 is implemented.