A method, apparatus, computing device, and storage medium for recommended content
Through multiple rounds of selection and sorting, we can break up similar content and achieve diversity and efficient recommendation, and solve the problems of single and redundant content in the recommendation system, improving user experience and system performance.
Patent Information
- Application Number
- CN202010671209.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-07-13
AI Technical Summary
In the existing recommendation system, the recommended content is too single, which makes users bored, lack novelty, and redundant recommendations lead to low user satisfaction.
Through multiple rounds of selection and sorting, the recommended content is reordered in diversity according to the recommended contribution predicted value, ensuring that content with high recommendation returns ranks first, breaking up the number of similar content, realizing cross-space arrangements of multiple categories, and improving content diversity and overall recommendation returns.
It improves the diversity and effectiveness of recommended content, provides surprising and surprising recommendation experience, improves user stickiness and retention, and ensures maximum overall recommendation benefits.
Smart Images

Figure CN111782957B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, computing device, and storage medium for recommending content. Background Art
[0002] With the rapid development of the Internet, users have access to massive amounts of content more and more conveniently, but this can easily lead to information overload. Take recommendation scenarios as an example. In scenarios such as information recommendation and short video recommendation, there is a lot of content to be recommended. If redundant recommendations are made without purpose, user satisfaction will be low.
[0003] To this end, the recommendation system can filter the content before recommending it to users. Related technologies generally match recommendations to users based on their preferences. For example, if a user likes to watch beauty videos, the recommendation system will recommend a large number of beauty videos to the user. However, the recommendation method based on user preferences can easily form an information cocoon. In the long run, the content recommended by this recommendation method is single. As a result, the same user is recommended content of the same or similar categories, and the content is similar and redundant, which can easily cause user boredom.
[0004] Therefore, how to reduce the recommendation redundancy of recommended content is an issue that needs to be considered. Summary of the invention
[0005] The embodiments of the present application provide a method, apparatus, computing device, and storage medium for recommending content, which are used to increase the diversity and effectiveness of content recommended by a recommendation system, reduce recommendation redundancy, and improve the recommendation performance of the recommendation system.
[0006] In one aspect, a method for recommending content is provided, the method comprising:
[0007] Obtain at least two categories of content to be recommended for the target user, each category including at least one content to be recommended;
[0008] Multiple rounds of selection are performed on the content to be recommended in the at least two categories based on the predicted recommendation contribution values, and the recommendation order of the content selected in each round is determined according to the number of selection rounds until all the content to be recommended is sorted; wherein, in each round of selection: the predicted recommendation contribution value of each candidate content in the current round relative to the set of categories to which it belongs is determined, and at least one candidate content whose predicted recommendation contribution value meets a set condition is determined from all the candidate content in the current round as the content to be selected in the current round, and the selected content is removed from the candidate content and added to the set of categories to which it belongs; wherein the candidate content in each round of selection has not been selected before the current round of selection;
[0009] Content is recommended to the target user according to the ranking of each content to be recommended in the completed recommendation order.
[0010] On the one hand, a device for recommending content is provided. The device includes:
[0011] An acquisition module, configured to acquire at least two categories of content to be recommended for a target user, where each category includes at least one piece of content to be recommended;
[0012] A sorting module, configured to perform multiple rounds of selection on the at least two categories of content to be recommended according to the predicted recommendation contribution value, and determine the recommendation order of the content selected in each round according to the selection round until all the content to be recommended is sorted; wherein, in each round of selection: respectively determine the predicted recommendation contribution value of each candidate content in this round relative to the set of categories to which it belongs, and determine at least one candidate content whose predicted recommendation contribution value meets the set conditions from all the candidate content in this round as the content selected in this round, and remove the selected content from the candidate content and correspondingly add it to the set of categories to which it belongs, where the candidate content in each round of selection has not been selected before this round of selection;
[0013] A recommendation module, configured to perform content recommendation to the target user according to the sorting of each piece of content to be recommended whose recommendation order has been completed.
[0014] Optionally, the acquisition module is configured to:
[0015] Acquire the list of content to be recommended corresponding to the target user, where the content to be recommended in the list of content to be recommended is recalled according to the user feature data of the target user, and the recommendation order of the content to be recommended in the list of content to be recommended is related to the matching degree between each piece of content to be recommended and the user feature data;
[0016] Select at least two categories of content to be recommended starting from the content to be recommended with a high matching degree according to the matching degree of each piece of content to be recommended in the list of content to be recommended.
[0017] Optionally, the sorting module is configured to:
[0018] Select the candidate content with the largest predicted recommendation contribution value from all the candidate content in this round as the content selected in this round.
[0019] Optionally, the sorting module is configured to:
[0020] Determine a predetermined number of candidate content starting from the largest predicted recommendation contribution value in descending order of the predicted recommendation contribution value, where the predetermined number is an integer greater than or equal to 2;
[0021] Determine the content selected in this round from the predetermined number of candidate content.
[0022] Optionally, the sorting module is configured to:
[0023] In the order from the largest to the smallest predicted recommendation contribution value, if the difference between the predicted recommendation contribution values of two adjacent candidate contents among the predetermined number of candidate contents is less than the difference threshold, the content selected in this round is determined from the predetermined number of candidate contents.
[0024] Optionally, the sorting module is configured to:
[0025] Determine all of the predetermined number of candidate contents as the content selected in this round; or,
[0026] Select candidate contents that meet the priority sorting condition from the predetermined number of candidate contents as the content selected in this round.
[0027] Optionally, the sorting module is configured to:
[0028] Use candidate contents among the predetermined number of candidate contents that are different in category from the content selected in the previous round as the content selected in this round; or,
[0029] Determine a target category in which the number of selected contents included in at least two categories before this round of selection meets the set quantity limit, and use candidate contents belonging to the target category among the predetermined number of candidate contents as the content selected in this round.
[0030] Optionally, the sorting module is configured to:
[0031] Use the candidate content with the highest original sorting among the predetermined number of candidate contents as the content selected in this round, where the original sorting of the candidate content refers to the sorting in the order from the highest to the lowest matching degree between the candidate content and the user feature data of the target user.
[0032] Optionally, the sorting module is configured to:
[0033] Calculate the first total predicted recommendation contribution value of all selected contents in the category set to which each candidate content belongs through a submodular function;
[0034] Calculate the second total predicted recommendation contribution value of each candidate content and all selected contents in the category set to which it belongs through the submodular function;
[0035] Determine the predicted recommendation contribution increment of the second total predicted recommendation contribution value of each candidate content relative to the first total predicted recommendation contribution value as the predicted recommendation contribution value of the candidate content.
[0036] Optionally, the sorting module is configured to:
[0037] Arrange the recommended order for the content selected in each round in such a way that the earlier the selection round, the higher the recommended order of the selected content.
[0038] On the one hand, a computing device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method for recommending content as described above are implemented.
[0039] In one aspect, a storage medium is provided, wherein the storage medium stores computer-executable instructions, wherein the computer-executable instructions are used to enable a computer to execute the steps included in the above-mentioned method for recommending content.
[0040] In one aspect, a computer program product comprising instructions is provided. When the computer program product is run on a computer, the computer is caused to execute the steps of the method for recommending content described in the various possible implementations above.
[0041] In an embodiment of the present application, multiple categories of content to be recommended are first obtained, and then multiple rounds of selection are performed on all the content to be recommended based on the recommendation contribution prediction values of the content to be recommended, and the recommendation order of the selected content in each round is arranged according to the selection rounds until all the content to be recommended is sorted, and further content recommendations are made to the user based on the ranking of each content to be recommended that has been arranged in the recommendation order.
[0042] In each round of selection, the recommendation contribution prediction value of each candidate content in this round relative to the category set to which it belongs is first determined, and then one or more candidate content whose recommendation contribution prediction value meets the set conditions is selected from all candidate content in this round as the content selected in this round and then participates in the sorting. This is equivalent to selecting the content with the recommendation contribution prediction value that meets the set conditions relative to the category set to which it belongs from all categories of content to be recommended and sorting them first, that is, putting the content with high recommendation benefits in front, so as to ensure that the total recommendation benefits of all content to be recommended are the largest. At the same time, due to multiple rounds of selection, each round of selection determines the content to be selected in each round based on the recommendation benefit prediction value of each content to be recommended. This is equivalent to sorting all categories of content in a scattered manner, thereby achieving the purpose of suppressing the number of similar content, so that content of multiple categories is arranged as dispersedly and cross-staggered as possible, increasing the category diversity of the recommended content, and thus achieving the purpose of maximizing the total recommendation benefits of all content to be recommended.
[0043] After reordering the recommended content, content from various categories is broken up as much as possible and then arranged in a cross-wise manner. This allows for a variety of recommended content to be recommended sequentially according to the reordered order, thereby increasing the diversity of recommended content. This provides users with a surprising and surprising recommendation experience, which can, to a certain extent, tap into users' potential preferences, thereby increasing user stickiness and retention. Furthermore, the more prominent content is, the greater its contribution to the total recommended content. Therefore, sequential recommendations can also maximize the overall recommendation benefit of all recommended content, thereby ensuring the accuracy and effectiveness of recommendations.
[0044] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0046] Figure 1 A schematic diagram of an application scenario applicable to an embodiment of the present application;
[0047] Figure 2 A flowchart of a method for recommending content in an embodiment of the present application;
[0048] Figure 3 This is a schematic diagram of at least two categories of content to be recommended obtained in an embodiment of the present application;
[0049] Figure 4 This is a schematic diagram of the first round of selection in an embodiment of the present application;
[0050] Figure 5a A schematic diagram of selecting candidate content whose predicted contribution increment meets set conditions as selected content in an embodiment of the present application;
[0051] Figure 5b Another schematic diagram of selecting candidate content whose predicted contribution increment meets set conditions as selected content in an embodiment of the present application;
[0052] Figure 5c Another schematic diagram of selecting candidate content whose predicted contribution increment meets set conditions as selected content in an embodiment of the present application;
[0053] Figure 6Schematic diagram of diminishing marginal benefit;
[0054] Figure 7 Another flowchart of the method for recommended content in the embodiments of the present application;
[0055] Figure 8 Schematic diagram of the structure of the device for recommended content in the embodiments of the present application;
[0056] Figure 9 Schematic diagram of the structure of the computing device in the embodiments of the present application. Detailed implementation manners
[0057] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the application. Without conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily combined with each other. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0058] The terms "first" and "second" in the specification, claims and the above drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in the present application may represent at least two, for example, it may be two, three or more, and the embodiments of the present application do not make limitations.
[0059] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after without special explanation.
[0060] The following explains some terms involved in this article to facilitate the understanding of those skilled in the art.
[0061] 1. A recommendation system is an information filtering system that contains a large amount of content to be recommended. It filters and screens content for each user and can recommend suitable content for the corresponding user. Depending on the type of recommended content, the recommendation system can include, for example, an information recommendation system, a short video recommendation system, a music recommendation system, and so on.
[0062] 2. Content: In the embodiments of this application, the content mainly refers to the content recommended to users through the recommendation system, which can include information, short videos, music, news, books, application programs, and other types.
[0063] 3. Diminishing marginal benefit means that, under other unchanged conditions, if an input factor is continuously increased in equal amounts, after increasing to a certain output value, the increment of the provided product will decrease, that is, the marginal output of the variable factor will decrease. When the total quantity of a certain item consumed by a consumer becomes larger and larger, the utility obtained from consuming the last additional unit of the item (i.e., the marginal utility) usually shows a phenomenon of decreasing more and more (diminishing). This phenomenon is called diminishing marginal benefit, or diminishing marginal effect, or diminishing marginal utility.
[0064] Popular understanding: At the beginning, the benefit value is very high. The later it gets, the less the benefit value becomes. Expressed in mathematical terms: x is the independent variable, y is the dependent variable, y changes with the change of x. As the value of x increases, the increase in y is continuously decreasing.
[0065] 4. Submodular function, also known as "sub - modular function" or "sub - additive function". A submodular function has submodularity, which is also called "sub - modularity" or "sub - additive property". It is a formal description of the phenomenon of diminishing marginal benefit. Therefore, submodular functions are suitable for describing and calculating tasks involving diminishing marginal benefit.
[0066] The following introduces the design concept of this application.
[0067] As mentioned above, when the recommendation system in the related technology recommends content to users, it mostly focuses on how to improve the accuracy of recommendations. In these recommendation mechanisms, generally, users are matched and recommended according to their preferences. Over time, the content recommended in this way is single, and it is easy to form an information cocoon. As a result, the content categories recommended to the same user are the same or similar, the content is similar and redundant, which easily causes user boredom. That is to say, the related technology ignores the diversity of recommendation results when recommending content, has the problem of redundant recommendations, and the effectiveness of recommendations is relatively low.
[0068] Specifically, the recommendation systems in the related technologies generally only focus on recall and ranking during the recommendation process, and there is no handling of how to improve diversity to solve the problem of recommendation redundancy of the recommended content. For example, after receiving a user's recommendation request, the recommendation system can recall some content to be recommended from a recommendation pool including a large amount of content according to the user's own user data. That is to say, the content to be recommended recalled is suitable for the user's own user data. Simply understood, it is what the user is likely to like. Further, the content to be recommended is scored and ranked according to the degree of matching between each recalled content to be recommended and the user's own user data. The probability of the user clicking on each content to be recommended can be calculated based on the user's own user data. The probability of clicking on the content can be simply referred to as the click-through rate, and the click-through rate can represent the degree of preference of the user for the content to be recommended. After obtaining the click-through rates of each content to be recommended, each content to be recommended can be scored according to the click-through rate to obtain the score value of each content to be recommended, and then all the content to be recommended can be ranked in descending order of the score value to obtain the final list of content to be recommended. The higher the score value, the greater the probability that the user will click on the corresponding content to be recommended. Therefore, when recommending content to the user, the content can be recommended to the user in descending order of the score value.
[0069] For example, the user's label is "likes entertainment-type information". Based on the recommendation scheme in the related technologies, most of the information with a higher score value, that is, ranked at the top, is entertainment-type information. If the information is recommended in descending order of the score value, most of the information recommended to the user is entertainment-type information. There may be content duplication among multiple pieces of information of the same type (i.e., entertainment type). In this way, the user is pushed the same type of information for a long time, lacking novelty and surprise, which may cause the user to be bored, resulting in the attenuation of the user's interest and the gradual decrease of the viscosity to the recommendation system.
[0070] In view of this, the inventors of this application have considered that by increasing the diversity of recommended content, it is possible to provide users with a novel and surprising recommendation experience to a certain extent. According to the principle of diminishing marginal returns, in each recommendation, the more items of the same type in the recommended items (such as information or short videos), the greater the total revenue, but the revenue growth rate decreases with each additional item of the same type. In other words, as the number of recommended items of the same type increases, the total revenue gradually increases, but as the number of items of the same type increases, the growth rate of the total revenue (i.e., the revenue increment) gradually decreases. In other words, the more items of the same type there are, the slower the trend of increasing the total revenue of items of the same type. Therefore, in order to maximize the total revenue, it is necessary to increase the revenue increment, which requires that the same type of items should not be recommended too much, but the diversity of item types should be increased, and the diversity of item types should be used to promote faster revenue growth. In other words, by increasing the diversity of item types, the revenue increment can be increased as much as possible, thereby maximizing the total revenue.
[0071] Applying the above-mentioned principle of diminishing marginal returns to the recommendation scenario in the embodiment of the present application, in order to maximize the total recommendation revenue of the recommendation system, the total recommendation revenue here can be understood as the total click-through rate of users on all recommended content, that is, the numerical quantification of the total interest in all recommended content. It is necessary to put content with high recommendation revenue in front, while suppressing the number of similar content. Based on this, the recommendation scheme in the embodiment of the present application is proposed as follows: first obtain multiple categories of content to be recommended, and then perform multiple rounds of selection on all content to be recommended based on the predicted value of the recommendation contribution of each content to be recommended, and arrange the recommendation order for the selected content selected in each round according to the selection round, until all content to be recommended is sorted, and further recommend content to the user based on the ranking of each content to be recommended that has been arranged in the recommendation order. In each round of selection, the recommendation contribution prediction value of each candidate content in this round relative to the category set to which it belongs is first determined, and then one or more candidate contents whose recommendation contribution prediction value meets the set conditions are selected from all candidate contents in this round as the content selected in this round and then participate in the sorting. This is equivalent to selecting the content whose recommendation contribution prediction value meets the set conditions (for example, the largest recommendation contribution prediction value) for the category set to which it belongs from all categories of recommended content and prioritizing it, that is, putting the content with high recommendation benefits in front, so as to ensure that the total recommendation benefit of all content to be recommended recommended by the recommendation system to the user is the largest. At the same time, due to multiple rounds of selection, each round of selection determines the content selected in each round based on the recommendation contribution prediction value of each content to be recommended. This is equivalent to sorting all categories of content in a scattered manner, thereby achieving the purpose of suppressing the number of similar content, making the content of multiple categories as dispersed and cross-staggered as possible, increasing the category diversity of the recommended content, thereby achieving the purpose of maximizing the total recommendation benefit of all content to be recommended.
[0072] After re - sorting the content to be recommended through the above - mentioned solution, content of multiple categories is cross - arranged after being scattered as much as possible. During the process of recommending in sequence according to the re - sorted recommendation order, content of multiple categories can all be recommended, achieving the purpose of improving the diversity of the recommended content. In this way, a surprising and pleasant recommendation experience can be provided for users, which can, to a certain extent, explore the potential preferences of users, thereby enhancing user viscosity and increasing the user retention rate. Moreover, the content that is more forward in the order has a relatively larger recommendation contribution value for all the content to be recommended. Therefore, recommending in order can also ensure that the overall recommendation benefit of all the content to be recommended is maximized, thereby ensuring the accuracy and effectiveness of the recommendation.
[0073] It should be noted that the diversity involved in the embodiments of the present application refers to the diversity of the categories to which the recommended content belongs. For example, the content with a higher recommendation order belongs to multiple categories, which can ensure that multiple types of content are recommended to users; in addition, the diversity designed in the embodiments of the present application also refers to the diversification of the category distribution during content recommendation, that is, content of multiple categories is cross - arranged at intervals. For example, the content ranked 1st and 2nd in the recommendation order is entertainment - related information, the content ranked 3rd is military - related information, the content ranked 4th and 5th is technology - related information, the content ranked 6th and 7th is entertainment - related information, the content ranked 8th is technology - related information, and so on. In this way, the diversity of content recommendation is improved, and at the same time, the accuracy and effectiveness of the recommendation can be ensured.
[0074] The method for recommending content provided by the embodiments of the present application can be applied to various recommendation scenarios, such as scenarios for recommending videos, scenarios for recommending information, scenarios for recommending e - books, scenarios for recommending applications, and so on.
[0075] To better understand the technical solution provided by the embodiments of the present application, the following briefly introduces the application scenarios applicable to the technical solution provided by the embodiments of the present application. It should be noted that the following introduced application scenarios are only used to illustrate the embodiments of the present application rather than to limit them. In specific implementation, the technical solution provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0076] Please refer to Figure 1 , Figure 1 which is an application scenario applicable to the method for recommending content in the embodiments of the present application. The method for recommending content can be applied to a content recommendation system. As Figure 1As shown, in this application scenario, there are multiple terminal devices (such as terminal device 101, terminal device 102, and terminal device 103) and a server 104. The server 104 can be a server serving a content recommendation platform, such as a news recommendation server, a short video recommendation server, and so on. Each terminal device is communicatively connected to the server 104. Each terminal device can send content to the server 104 to publish it on the content recommendation platform served by the server 104. At the same time, it can also receive the published content that has been published on the content recommendation platform sent by the server 104.
[0077] Taking terminal device 102 and the short video recommendation scenario as an example, terminal device 102 corresponds to user 2. User 2 can operate terminal device 102 to send a short video recommendation request to server 104. Further, server 104 can use the method for recommending content in the embodiments of the present application to perform diversity ranking on the short videos to be recommended, and then recommend the short videos to terminal device 102 for user 2 to watch according to the recommendation order of the short videos after diversity ranking, so that user 2 can view as many types of short videos as possible, increasing novelty. In this way, the potential preferences of user 2 can also be explored to a certain extent, increasing the surprise level of content recommendation.
[0078] Taking terminal device 101 and the news recommendation scenario as another example, terminal device 101 corresponds to user 1. User 1 can operate terminal device 101 to send a news recommendation request to server 104. Further, server 104 can use the method for recommending content in the embodiments of the present application to perform diversity ranking on the news to be recommended, and then recommend the news to terminal device 101 for user 1 to watch according to the recommendation order of the news after diversity ranking, so that user 1 can view as many types of news as possible, increasing novelty. In this way, the potential preferences of user 1 can also be explored to a certain extent, increasing the surprise level of content recommendation.
[0079] Among them, server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Terminal devices 101, 102, and 103 can be smartphones, tablet computers, laptop computers, desktop computers, smart TVs, smart wearable devices, etc., but are not limited thereto.
[0080] To further illustrate the technical solutions provided by the embodiments of the present application, the following will provide a detailed description in conjunction with the accompanying drawings and specific implementation manners. Although the embodiments of the present application provide method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or non-creative labor. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present application. When the method is actually processed or executed by the device, it can be executed in the method order shown in the embodiments or drawings or executed in parallel.
[0081] The present application provides a method for recommending content. This method can be executed by a recommendation server in a recommendation system, for example, executed by Figure 1 the server 104 therein. The method for recommending content provided by the embodiments of the present application is as Figure 2 shown, Figure 2 and the flowchart shown is described as follows.
[0082] Step 201: Obtain at least two categories of content to be recommended for the target user, and each category includes at least one piece of content to be recommended.
[0083] Before step 201, the recommendation server may receive a recommendation request initiated by the client, and the recommendation request may carry a user identifier, such as the account identifier of the application account logged in by the user. The target user is determined through this account identifier, and then it is determined that content needs to be recommended to the target user, such as recommending news or short videos and other content. The content can be classified. For example, it can be classified according to the classification attribute information of the content. Taking news as an example, it can be divided into categories such as technology, entertainment, sports, military, society, finance, etc. according to the field to which the news belongs. Thus, further, at least two categories of content corresponding to the target user can be obtained. In the embodiments of the present application, the content that can be recommended to the user is called content to be recommended. That is to say, multiple categories of content to be recommended corresponding to the target user are obtained, and each category includes at least one piece of content to be recommended.
[0084] Corresponding to the target user may mean that these content to be recommended match the target user, which can be achieved through the information filtering function of the recommendation system. For example, it can be matched according to the user characteristic data of the target user, and content with a certain degree of match with the user characteristic data is recalled from the massive content pool maintained by the recommendation system. That is to say, the user's preferences can be determined through the user characteristic data, and then content that matches the user's preferences is selected. These selected content can be saved in the list of content to be recommended, and each piece of content to be recommended in the list of content to be recommended has its own recommendation order. And the recommendation order of each piece of content to be recommended is related to its degree of match with the user characteristic data. Specifically, it can be determined according to the level of its match with the user characteristic data. The higher the degree of match, the larger the score value of the score, indicating that the probability that the user likes it and clicks is greater, and the corresponding recommendation order is more forward.
[0085] Among them, the user characteristic data may include at least one of user attribute data and user operation data. User attribute data is data reflecting the basic attributes of the user. For example, it may include data such as age, gender, geographical location, education level, job, and preference tags. User operation data is data related to the user's usual operations on the recommended content. For example, data related to operations such as the browsing duration, browsing times, whether to comment, whether to like, whether to forward, and whether to download the recommended content. Both the user attribute data and the user operation data can reflect the user's interest and preference for the recommended content to a certain extent. Therefore, recalling the content to be recommended from the content library through the user characteristic data and sorting the recommendation order according to the level of match can make the recommended content all be content that the user is likely to be interested in, thereby improving the accuracy of content recommendation.
[0086] In the embodiment of the present application, the at least two categories of content to be recommended obtained in step 201 can be the content selected from the list of content to be recommended recalled and sorted according to the user characteristic data above. In this way, if re-sorted later, it is equivalent to re-performing diversity sorting on these content to be recommended. This is a diversity re-rank recommendation based on ensuring the accuracy of the recommendation to try to improve the effectiveness of the recommendation. Or, other methods can also be used to obtain the aforementioned at least two categories of content to be recommended, and the embodiment of the present application does not make any restrictions.
[0087] Among them, when selecting at least two categories of content to be recommended from the above list of content to be recommended, the total number of content to be recommended in the list of content to be recommended can be determined first. Further, then compare the total number with a set threshold (for example, the first threshold: 150), and select at least two categories of content to be recommended for re-sorting from the list of content to be recommended according to the comparison result.
[0088] When the total quantity is greater than or equal to the first threshold, multiple to-be-recommended contents belonging to multiple categories can be selected therefrom. This is equivalent to performing subsequent diversity reordering on a part of the to-be-recommended contents, which can reduce the computational amount generated due to reordering during the recommendation process, avoid recommendation latency and lag as much as possible, and improve the recommendation efficiency.
[0089] When the total quantity is less than the aforementioned first threshold, all the to-be-recommended contents in the to-be-recommended list can be selected as the contents to be reordered. It can be understood that if there are less than two categories of to-be-recommended contents in the to-be-recommended list at this time, some other types of contents can also be recalled from the content pool according to the user feature data to supplement at least two categories of to-be-recommended contents, so as to meet the requirements for multiple categories; or, some associated contents of other categories can also be selected from the to-be-recommended lists of users associated with the target user (such as the social friends or other family members of the target user) to supplement at least two categories. Because the life circles of the target user and the users associated with him are generally the same or close, and the contents he is interested in and concerned about may also have certain similarities, the contents obtained from the associated users are very likely to be of interest to the user himself. By this way, the accuracy of the recommendation is also improved.
[0090] When selecting multiple to-be-recommended contents belonging to at least two categories from the aforementioned to-be-recommended list, the matching degrees of the to-be-recommended contents in the to-be-recommended list can be used, and the selection can start from the to-be-recommended contents with high matching degrees. For example, some contents are respectively selected from the to-be-recommended contents with high matching degrees for each category to participate in the subsequent diversity reordering, or multiple types of contents with matching degrees greater than a set threshold (such as the second threshold) are selected to participate in the subsequent diversity reordering, or the to-be-recommended contents with the highest matching degrees in the to-be-recommended list are selected to start, and the to-be-recommended contents with consecutive sorting and belonging to at least multiple categories are selected to participate in the subsequent diversity reordering, and so on. During the selection process, the quantity can also be restricted. For example, finally only 100 to-be-recommended contents need to be selected to participate in the subsequent diversity reordering, so that a comprehensive selection can be made in combination with the aforementioned matching degree and the quantity requirement here.
[0091] Taking the recommendation of information as an example, assume that 100 to-be-recommended contents are obtained, and these 100 to-be-recommended contents respectively belong to four types: technology, entertainment, sports, and military, and these 100 to-be-recommended contents already have a recommendation order. For example, the existing recommendation sorting is called the original sorting. Please refer to Figure 3 As shown, in the original sorting, assume that the recommendation order of content 1, content 2, content 3,..., content 100 gradually becomes later. In Figure 3 the arrow indicates that the later the recommendation order is.
[0092] Step 202: Perform multiple rounds of selection on the content to be recommended for at least two categories to determine the content selected in each round; wherein, in each round of selection: respectively determine the predicted recommendation contribution value of each candidate content in this round relative to the set of categories to which it belongs, and determine at least one candidate content whose predicted recommendation contribution value meets the set conditions from all the candidate contents in this round as the content selected in this round, and remove the selected content from the candidate contents in this round and add it to the set of categories to which it belongs correspondingly.
[0093] Step 203: Determine the recommendation order of the content selected in each round according to the selection round.
[0094] In each round of selection, the content to be recommended participating in this round of selection is called the candidate content, and the content to be recommended that needs to be sorted in this round selected in this round is called the selected content in this round.
[0095] The predicted recommendation contribution value of the content to be recommended is a quantified value of the predicted recommendation benefit after the recommended content is recommended. The larger the predicted recommendation contribution value, the greater the probability that the content to be recommended will be clicked by the user after being recommended. That is to say, it indicates that the recommended benefit generated after the content to be recommended is recommended is higher, that is, the greater the contribution to the overall recommendation effect. The recommendation benefit of the content to be recommended can be characterized as the recommendation effect generated after the content to be recommended is recommended. The better the recommendation effect, the greater the corresponding recommendation benefit. Taking the content to be recommended as information as an example, the recommendation effect can be characterized by factors such as the click-through rate of the user and the operation degree of the content to be recommended (such as the number of clicks to view, browsing duration, whether to like, whether to forward and comment, etc.).
[0096] The predicted recommendation contribution value of each candidate content relative to the set of categories to which the candidate content belongs can be calculated in the following way: first calculate the total predicted recommendation contribution value of all the selected contents in the set of categories to which the candidate content belongs, for example, denoted as the first total predicted recommendation contribution value, then calculate the total predicted recommendation contribution value of the candidate content and all the selected contents in the set of categories to which it belongs, for example, denoted as the second total predicted recommendation contribution value, and finally determine the difference between the second total recommended contribution value corresponding to the candidate content and the first total recommended contribution value as the predicted recommendation contribution value of the candidate content itself. That is to say, the predicted recommendation contribution increment of the candidate content relative to all the selected contents in the set of categories to which it belongs can be determined as the predicted recommendation contribution value of the candidate content itself.
[0097] In the embodiments of the present application, the principle of diminishing marginal benefit is used to achieve the diversity reordering of the content to be recommended, and the submodular function is just suitable for dealing with tasks such as diminishing marginal benefit. Therefore, in the embodiments of the present application, the predicted recommendation contribution value of each candidate content can be calculated through the submodular function.
[0098] The definition of the submodular function is as follows:
[0099] For a set S of a certain type of item, its subset Items x, y ∈ S:
[0100] f(A ∪ {x}) - f(A) ≥ f(A ∪ {x, y}) - f(A ∪ {y}) (Equation 1)
[0101] Any function that conforms to the definition of Equation 1 can be considered a submodular function. For example:
[0102] log(1 + ∑ i score(a i )) (Equation 2)
[0103] In Equation 2, A is a subset of the category set C i and a i is an element in A. For example, each piece of technology information in a collection of technology information. The benefit of a i is denoted as score(a i ). Taking information as an example, the benefit is usually the estimated click-through rate, the estimated conversion rate, etc.
[0104] For the content to be recommended selected from multiple categories, assuming the number of categories is C. For example, information can be classified into categories such as military, entertainment, technology, sports, etc. For each of these C category sets, there can be their respective submodel functions Then the submodular benefit of A is expressed as follows in Equation 3:
[0105]
[0106] The total submodular benefit of these C sets is as follows:
[0107]
[0108] The problem is transformed into:
[0109]
[0110] The sub-optimal solution of the submodular function, that is, the approximate optimal solution, can be obtained through, for example, the greedy algorithm or other algorithms. This approximate optimal solution indicates that the overall benefit of the multiple contents ranked in the front is greater. For example, if 5 contents need to be returned from 100 contents, there are many combination methods, and each combination method has an overall benefit. Then, through this approximate optimal solution, the overall benefit of the combination of the top 5 contents (the 5 contents ranked at the very front) after reordering using the submodular function can be ensured to be the largest.
[0111] For easy understanding, the following is an example.
[0112] Such as Figure 3As shown, assume that the at least two categories of recommended content to be diversely reordered obtained in step 201 are content 1, content 2, content 3, …, content 100, and the original order of each piece of recommended content gradually moves backward in ascending order of the number, and these 100 pieces of content belong to four categories: military, entertainment, technology, and sports.
[0113] The following is to make multiple rounds of selections for these 100 pieces of content to determine the selected content in each round of selection, that is, to obtain the content selected in each round of selection, and then sort the recommended order of the selected content in each round according to the selection round.
[0114] The first round of selection:
[0115] Traverse these 100 pieces of content, and the predicted value of the recommendation contribution of each candidate content can be calculated using the submodular function, that is, the submodular gain of each candidate content can be calculated.
[0116] Before the first round of selection for these 100 pieces of content, there is no content included in each category set, that is, each category set is empty, as Figure 3 shown, there is no content in the rectangular boxes representing the 4 category sets. In the first round of selection, there is no selected content in each category set currently, so the total predicted value of the recommendation contribution of each category set is 0 currently, that is, the submodular gain of each category set is 0 currently.
[0117] Assume that content 1 is in the "military" classification, and the current submodular gain of the military classification set is 0. If content 1 is added to the military classification set, the submodular gain of the entire military classification set is a1, then the incremental gain of content 1 is a1 - 0, that is, a1. It should be noted that at this time, content 1 is actually not added to the military classification set, so more precisely, a1 is the predicted incremental gain.
[0118] Similarly, it can be calculated that the incremental gain of content 2 is b1, the incremental gain of content 3 is c1, …, and the incremental gain of content 100 is z1.
[0119] Compare the incremental gains of each candidate content from content 1 to content 100, and select the candidate content whose incremental gain meets the set conditions as the selected content in this round. The set conditions here are, for example, to select the candidate content with the largest incremental gain as the selected content in this round. Assume that the candidate content with the largest incremental gain in the first round is content 3, and the classification is "technology". After the first round, as Figure 4 shown, the selected content in the first round is content 3, and at the same time, content 3 is placed in the first position in the final recommended list, as Figure 4As indicated by the dashed arrow for "Content 3" in [description], the selected content chosen in the first round is ranked first in the reordering. Since all category sets are empty in the first round, the candidate content with the largest incremental benefit indicates the greatest contribution to its respective category and, among the candidate contents of all categories, the greatest contribution to the overall benefit of all content to be recommended. Therefore, it can be ranked first in the recommended order of the reordering.
[0120] At the same time, Content 3 is removed from the candidate content in the first round and placed into the category set it belongs to, that is, into the category set of "Technology", as Figure 4 indicated by the solid arrow for "Content 3" in [description]. So after the first round, the category set of "Technology" includes the selected content of Content 3, and Content 3 is no longer included in the remaining candidate content, that is, Content 3 is removed from the original ordering.
[0121] The purpose of removing Content 3 from the candidate content is to not consider it in the next round of selection, that is, it will no longer be a candidate content for the next round of selection because Content 3 has been determined as the selected content in the first round for diversity reordering. In this regard, the candidate content in each round of selection has not been determined as the selected content before that round of selection. The candidate content in each round of selection has not been selected before that round because the candidate content that has been determined as the selected content before this round has been added to the category set it belongs to in the round when it was determined as the selected content, as Figure 4 Content 3 in [description] has been added to its "Technology category" in the first round.
[0122] Second round of selection:
[0123] Only 99 candidate contents participate in the calculation of the submodular gain in the second round of selection because Content 3 has been removed from the candidate content in the previous round and added to the category set of the "Technology category" as the selected content. For the remaining 99 candidate contents, the method introduced above is used to calculate the submodular gain of each candidate content in turn.
[0124] Content 1 is in the "Military" classification, and the category set of the Military classification is still empty, so the submodular gain of Content 1 is still a1.
[0125] Content 2 is in the "Entertainment" classification, and the incremental benefit of Content 2 is b1.
[0126] Content 4 is in the "Sports" classification, and the incremental benefit of Content 4 is d1.
[0127] Content 5 belongs to the "Technology" category. Since the selected content 3 already exists in the category set of the "Technology" category, the current submodular gain of this category set is not 0. Let f represent the submodular function. Then, in the second round, the gain increment e2 of Content 5 is f({Content 5, Content 3}) - f({Content 3}).
[0128] ……
[0129] Content 17 belongs to the "Entertainment" category, and the gain increment of Content 17 is, for example, x1.
[0130] Similarly, the gain increments of all other candidate contents in the second round can be calculated.
[0131] After obtaining the gain increments of each candidate content in the second round, all the gain increments can be compared, and then the candidate contents whose gain increments meet the set conditions are selected as the selected contents in the second round.
[0132] In one possible implementation, the candidate content with the largest gain increment can be used as the selected content in this round. Assume that the gain increment x1 of Content 17 is the largest in the second round. Then, as Figure 5a shown, Content 17 can be used as the selected content in this round, added to the second recommended position in the final recommended list, removed from the candidate contents in this round, and at the same time, Content 17 is added to the category set corresponding to "Sports". In this way, the candidate content with the largest gain increment in this round is selected as the selected content in this round and given priority for reordering. In this way, the overall gain of the final recommended list can be ensured to be maximized as much as possible, thereby improving the effectiveness of the recommendation.
[0133] In another possible implementation, the candidate contents can also be determined in descending order of the predicted recommendation contribution value, starting from the candidate content with the largest predicted recommendation contribution value, and the number of candidate contents is a positive integer greater than or equal to 2, such as 2 or 3 or 4, etc. Further, the selected content in this round is determined from the selected candidate contents. Assume that the number of candidate contents is 2, and the 2 candidate contents with the largest gain increments selected in descending order of the gain increment are Content 17 and Content 5, that is, Content 5 is the candidate content with the second largest gain increment in the second round selection.
[0134] Further, when the gain increments between the candidate contents of the predetermined number are not very different, for example, when the difference between the predicted recommendation contribution values of two adjacent candidate contents among the candidate contents of the predetermined number is less than the difference threshold (such as the third threshold), all the candidate contents of the predetermined number can be determined as the selected contents in this round. For example, if the difference between the submodular increments of Content 17 and Content 5 is less than the aforementioned third threshold, then as Figure 5bAs shown, both Content 17 and Content 5 can be selected as the selected content for this round, such as Figure 5b indicated by the dashed arrows for "Content 5" and "Content 17" in Figure 5b , and they are ranked in the final recommendation list according to the size of the incremental revenue, that is, the larger the incremental revenue, the higher the ranking position in the final recommendation list. By comparing the difference between the predicted incremental recommendation contributions of two adjacent candidate contents with the difference threshold, it can be ensured that the submodular gains among the selected predetermined number of candidate contents are not too different, that is, the contributions of each candidate content in the predetermined number of candidate contents to the total final recommendation revenue are similar. Therefore, they can be rearranged diversely together.
[0135] In another implementation, when the difference between the predicted values of the recommendation contributions of two adjacent candidate contents among the predetermined number of candidate contents is less than the difference threshold (such as the aforementioned third threshold), candidate contents that meet the priority ranking conditions can also be selected from these predetermined number of candidate contents as the selected content for this round. In this way, the candidate contents that meet the priority ranking conditions can be ranked first, that is, they can be arranged in a more forward position, which can maximize the total recommendation revenue as much as possible and improve the effectiveness and accuracy of the recommendation.
[0136] For example, candidate contents among the predetermined number of candidate contents that are of different categories from the selected content chosen in the previous round are used as the selected content for this round. In the previous round, that is, in the first round, the selected content chosen is Content 3, which is in the technology category, while both Content 17 and Content 5 are not in the technology category. So in this round, both Content 17 and Content 5 can be determined as the selected content and ranked first, as Figure 5b shown. Suppose Content 17 is in the technology category while Content 5 is not in the technology category and Content 5 is in the entertainment category. In this method, only Content 5 can be selected as the selected content, as Figure 5c shown. In this way, the contents selected in two adjacent rounds are of different types, and different types of contents can be arranged crosswise as much as possible to enhance diversity.
[0137] Or for example, determine the target category for which the number of selected contents included in at least two categories before this round of selection meets the set quantity limit, and use the candidate contents belonging to the target category among the predetermined number of candidate contents as the selected content for this round. For example, there are four categories in total: "technology", "military", "entertainment", and "sports". Currently, if it is the 10th round of selection, among all the selections in the previous 9 rounds, the content selected as the "technology category" is the least or less than the specified quantity, then the "technology category" can be determined as the target category in the embodiments of this application. Furthermore, then in this round, the content belonging to the "technology category" can be used as the selection result. In this way, diversity can be enhanced as much as possible, and the quantity of each type is ensured not to be too small.
[0138] For another example, the candidate content with the highest original ranking among a predetermined number of candidate contents is used as the selected content for this round. Here, the original ranking of the candidate contents refers to the ranking in the order from high to low according to the matching degree between the candidate contents and the user feature data of the target user. In the second round, when comparing Content 17 and Content 5, the original ranking of Content 5 is more forward. Therefore, as Figure 5c shown, only Content 5 can be determined as the selected content for this round and given priority for ranking. In this way, when the submodular gain difference between two candidate contents is not significant, in order to consider the accuracy of recommendations simultaneously, the one with a more forward original ranking can be given priority for re-ranking. In this way, on the basis of meeting the diversity re-ranking, the requirements of recommendation accuracy can be met as much as possible.
[0139] Referring to the methods of selecting the selected content in the first round and the second round described above, multiple rounds of selection and re-ranking can be carried out until all the content to be recommended is selected as the selected content for re-ranking. For 100 pieces of content to be recommended, assuming that only one candidate content is selected as the selected content for each round, 100 rounds of selection are required. If multiple contents are selected as the selected content for some rounds, fewer than 100 rounds of selection are required until all the candidate contents among the 100 candidate contents are sorted into the final recommended list, thus forming a new re-ranked list.
[0140] Step 204: Perform content recommendation to the target user according to the ranking of each piece of content to be recommended that has completed the recommended order arrangement.
[0141] Furthermore, after obtaining the final recommended list with diversity re-ranking, content recommendation can be carried out to the user in turn according to the recommended order in the final recommended list.
[0142] It should be noted that if at least two categories of content to be recommended in Step 201 are selected from the list of content to be recommended obtained according to accuracy-based recommendation, the obtained final recommended list can be concatenated with the remaining content to be recommended in the aforementioned list of content to be recommended, and then the recommendation can be carried out sequentially according to the concatenated recommended order.
[0143] In the embodiments of the present application, the diminishing marginal benefit is used to achieve re-ranking of multiple categories of content. Please refer to the Figure 6 schematic diagram of the diminishing marginal benefit shown. For the same type of item, as the number of recommended items increases, the total recommended benefit gradually increases. However, as the number of recommended items gradually increases, the increase rate of the total recommended benefit gradually decreases. Therefore, Figure 6 the total benefit curve in becomes gradually flatter as the number of recommended items increases. For this reason, the embodiments of the present application increase the increase rate of the total recommended benefit by increasing the categories of recommended content to try to increase the total recommended benefit as much as possible.
[0144] Therefore, in the embodiment of the present application, since the final recommendation list is re-sorted according to the predicted recommendation contribution value of each content to be recommended relative to the category set to which it belongs, the content to be recommended at the front after re-sorting has a higher probability of belonging to multiple categories, and is recommended to the user in the order after re-sorting. The content to be recommended in multiple categories can be recommended to the user with a higher probability. In this way, the diversity of the recommended content can be improved and the recommendation redundancy can be reduced to a certain extent.
[0145] To facilitate understanding of the diverse re-ranking schemes in the embodiments of this application, the following takes information recommendation as an example. Figure 7 The technical solutions in the embodiments of the present application are further explained.
[0146] Step 701: Obtain information recommendation request from target user.
[0147] Taking the target user as an example, when the target user wants to view information, an information recommendation request can be triggered to the recommendation server through the client. For example, when the user opens the information application, the client can be directly triggered to send an information recommendation request to the recommendation server, or the user can perform specific operations in the application interface of the information application or enter a "latest information" request, which can trigger the client to send an information recommendation request to the recommendation server.
[0148] Step 702: Recall a number of information to be recommended from the information pool based on the user characteristic data of the target user.
[0149] The information pool is a database of massive amounts of information maintained by the recommendation system. All information published by users through the recommendation system can be saved in the information database. In order to improve storage performance, the information pool can be cleaned up regularly. For example, information published six months ago can be deleted from the information pool to store more newly published information.
[0150] Based on the user characteristic data of the target user, some information that is more closely matched with the user characteristic data can be selected from the information pool as recalled information. The degree of match between the information and the user characteristic data can represent the degree to which the user likes the information. For example, the higher the degree of match, the greater the probability that the target user will click on the information after it is recommended to the target user, thus achieving more accurate recommendations.
[0151] Generally speaking, when recalling information to be recommended from the information pool, the recalled amount will generally not be too small, so that a certain amount of information can be continuously recommended to the target user. Moreover, although the recall is matched based on the user feature data of the target user, due to the large number of recalls and the user feature data may also match multiple categories of information, the recalled information to be recommended generally belongs to multiple categories.
[0152] Recall information from the information pool based on user feature data, which can ensure that the information to be recommended recalled is information that the target user is likely to like. In this way, the accuracy of the recommendation can be improved as much as possible.
[0153] Step 703: Score the matching degree of several pieces of content to be recommended recalled according to the level of the matching degree between each piece of information to be recommended and the user feature data.
[0154] For example, a matching degree prediction model can be used to predict the matching degree between each piece of information to be recommended and the user feature data, and score the matching degree of each piece of information to be recommended recalled in such a way that the higher the matching degree, the higher the corresponding score. For example, the scoring range is 50 - 100 points. The higher the score of the information, the greater the degree of the user's preference for this information. In this way, after recommending the information with a higher score to the target user, the probability that the target user clicks on this information is also greater.
[0155] Step 704: Sort in the order from the highest to the lowest matching degree to obtain a sorted list of information to be recommended.
[0156] Suppose the number of pieces of information to be recommended recalled from the information pool is 150. Then, these 150 pieces of information can be sorted according to the level of the matching degree score. The higher the matching degree score of the information, the more forward the sorting. In this way, a sorted list of information to be recommended is obtained. It can be understood that the information ranked first in the sorted list of information to be recommended is the information with the highest matching degree score. From the information ranked first and going backward in turn, the corresponding matching degree scores of the information are lower.
[0157] Step 705: Select the top K pieces of content to be recommended that are ranked forward and belong to multiple categories from the list of information to be recommended, denoted as the topK list.
[0158] In the sorted list of information to be recommended, the information to be re-sorted can be selected starting from the information ranked first. Taking the value of K as 100 as an example, the top 100 pieces of information can be selected from the previously sorted 150 pieces of information as the topK list. It should be noted that the information to be recommended included in the topK list belongs to at least two categories. If the number of categories to which the top 100 pieces of information to be recommended belong is less than two, then information ranked as forward as possible can be selected from the information to be recommended after the 100th ranking to make up at least two categories. In this way, it is ensured that the information to be recommended in the selected topK list is as well-matched with the target user as possible and belongs to multiple categories. In this way, during subsequent re-sorting, diverse re-sorting can be carried out on the basis of ensuring accuracy, which can simultaneously ensure the accuracy and diversity of the recommended content.
[0159] Step 706: Traverse the topK list.
[0160] After obtaining the top-K list, all the to-be-recommended information included in the top-K list can be selected in multiple rounds according to the predicted recommendation contributions of each to-be-recommended information relative to its corresponding category set.
[0161] Step 707: Determine whether the predicted recommendation contribution value of the currently traversed information is the largest.
[0162] For the top-K list, traverse each to-be-recommended information in a certain order, calculate the predicted recommendation contribution value of each to-be-recommended information relative to its corresponding category set in turn, and compare the predicted recommendation contribution value of the currently traversed to-be-recommended information with the predicted recommendation contribution values of the previously traversed to-be-recommended information to determine whether the predicted recommendation contribution value of the currently traversed information is the largest. That is to say, during the traversal process, the largest predicted recommendation contribution value can be determined from all the to-be-recommended information.
[0163] Step 708: When it is determined that the predicted recommendation contribution value of the currently traversed information is the largest, determine whether the top-K list has been traversed completely.
[0164] Step 709: Determine whether the predicted recommendation contribution value of the next information is the largest.
[0165] When it is determined that the predicted recommendation contribution value of the currently traversed information is the largest, the purpose of determining whether the top-K list has been traversed completely is to find the largest predicted recommendation contribution value from all the predicted recommendation contribution values corresponding to all the to-be-recommended information. That is to say, it is necessary to select the to-be-recommended information with the largest predicted recommendation contribution value from the to-be-recommended information of all categories, because the to-be-recommended information with the largest predicted recommendation contribution value indicates that its recommendation benefit for all the to-be-recommended information in the top-K list is the largest. In this way, it can be selected as the first sorting in the diversity sorting.
[0166] When it is determined that the predicted recommendation contribution value of the currently traversed information is not the largest among all the obtained predicted recommendation contribution values, or when it is determined that the predicted recommendation contribution value of the currently traversed information is the largest among all the obtained predicted recommendation contribution values and the top-K list has not been traversed completely, then, in accordance with the traversal order, continue to calculate the predicted recommendation contribution value of the next piece of information, and determine whether the predicted recommendation contribution value of the next piece of information is the largest among all the obtained predicted recommendation contribution values.
[0167] Step 710: Put the information with the largest predicted recommendation contribution value into the diversity list for priority sorting.
[0168] That is to say, in the traversal order, the predicted recommendation contributions of each piece of content to be recommended in the topK list are calculated in sequence, and the piece of information with the largest predicted recommendation contribution is selected from the topK list, and the piece of information with the largest predicted recommendation contribution is placed in the diversity list for priority sorting. For example, the piece of information with the largest predicted recommendation contribution selected in the first round is arranged in the first position in the diversity list.
[0169] Step 711: Remove the piece of information with the largest predicted recommendation contribution from the topK list and add it to the category set to which it belongs.
[0170] After sorting the predicted recommendation contributions in the diversity list, they can be removed from the topK list. After removing them from the topK list, there are actually only K - 1 pieces of content to be recommended in the topK list. That is to say, the topK list is actually updated.
[0171] Step 712: Determine whether the topK list is empty.
[0172] Using the method of the above Step 706 - Step 711, the updated topK list is selected in multiple rounds until all the pieces of content to be recommended in the topK list are removed and added to the corresponding category sets of each piece of content to be recommended. At this time, all the pieces of content to be recommended in the original topK list have new sorting in the diversity list. Finally, the topK list becomes an empty list.
[0173] Step 713: When the topK list is already empty, replace the selected topK list in the list of content to be recommended with the diversity list to obtain the re - sorted list of content to be recommended.
[0174] Step 714: Recommend content to the target user according to the recommendation order in the re - sorted list of content to be recommended.
[0175] After the top-K list becomes an empty list, the diversity list reorders the 100 pieces of information to be recommended included in the original top-K list in a diverse manner. Furthermore, the obtained diversity list can be used to replace all the information to be recommended in the first K (i.e., the first 100) positions in the originally sorted list of information to be recommended. In this way, the order of the top 100 pieces of information after reordering is different from the order of the top 100 pieces of information in the original list of information to be recommended. Compared with the top 100 pieces of information in the original list of information to be recommended, the information of multiple categories in the top 100 pieces of information after reordering is arranged in a scattered manner as much as possible. For example, the information of multiple types is arranged in a disorderly and crosswise manner. In this way, when recommending information to the target user, it can be ensured that as many types of information as possible are preferentially recommended to the target user, thereby improving the diversity of the categories of the recommended information. At the same time, the recommendation order can also be as diverse as possible, and the recommendation redundancy caused by recommending information of a single category can be minimized as much as possible, improving the effectiveness of the recommendation.
[0176] The embodiment of the present application is a technical solution applied to the recommendation scenario to solve the problem of diverse recommendation. The problem of diverse recommendation in the recommendation scenario refers to the problem of how to recommend different types of recommended content in a reasonable order. It can be understood that when recommending different recommended content in different orders, users may generate different browsing behavior data when browsing the recommended content. For example, in the recommendation mechanism that only considers accuracy in the related art, the recommendation system first recommends 8 pieces of entertainment information and 1 piece of sports information to the user, and the user clicks on 3 pieces of entertainment information. Since there are too many entertainment information and the content may be repeated, the overall click-through rate of the 9 pieces of information recommended to the user is not high. Suppose that the recommendation mechanism after diverse reordering provided by the embodiment of the present application is adopted, and the recommendation system also recommends 9 pieces of information to the user, including 5 pieces of entertainment information, 2 pieces of sports information, 1 piece of science and technology information, and 1 piece of military information. The user clicks on 2 pieces of entertainment information, 1 piece of sports information, 1 piece of science and technology information, and 1 piece of military information. It can be seen that the overall click-through rate of the user for these 9 pieces of information is much higher. In this way, it is equivalent to improving the overall benefit of information recommendation. By providing diverse information recommendations for users, not only can the overall recommendation benefit be improved, thereby improving the accuracy and effectiveness of the recommendation, but also, users can be guided to discover some potential interests and hobbies as much as possible. These diverse content recommendations may bring unexpected surprises and delights to users, thereby enhancing the user experience, helping to improve user viscosity, and increasing user retention rate.
[0177] In the embodiment of the present application, the principle of diminishing marginal returns is used to re-sort the recommended content into multiple categories to achieve a diverse recommendation ranking. After the recommended content is re-sorted, the content of multiple categories is scattered as much as possible and then cross-arranged. In the process of recommending in sequence according to the re-sorted recommendation order, multiple categories of content can be recommended, achieving the purpose of increasing the diversity of recommended content. This can provide users with a surprising and surprising recommendation experience, which can to a certain extent explore the user's potential preferences, thereby improving user stickiness and increasing user retention. In addition, the more forward the content is, the greater its recommendation contribution value is for all the content to be recommended. Therefore, recommending in order can also ensure that the overall benefit of all recommended content is maximized, thereby ensuring the accuracy and effectiveness of the recommendation.
[0178] Based on the same inventive concept, the embodiment of the present application provides a device for recommending content, which can be a hardware structure, a software module, or a hardware structure plus a software module. Figure 1 The server 104 itself, or a functional device set in the server 104, can be implemented by a chip system. The chip system can be composed of a chip or include a chip and other discrete devices. Figure 8 As shown, the content recommendation device in the embodiment of the present application includes an acquisition module 801, a sorting module 802 and a recommendation module 803, wherein:
[0179] An acquisition module 801 is configured to acquire at least two categories of content to be recommended for a target user, each category including at least one content to be recommended;
[0180] The ranking module 802 is configured to perform multiple rounds of selection on at least two categories of content to be recommended based on the predicted recommendation contribution values, and determine the recommendation order of the content selected in each round according to the number of selection rounds, until all the content to be recommended is ranked. In each round of selection, the predicted recommendation contribution value of each candidate content in the current round relative to the set of categories to which it belongs is determined, and at least one candidate content whose predicted recommendation contribution value meets a set condition is determined from all candidate content in the current round as the content to be selected in the current round. The selected content is removed from the candidate content and added to the corresponding set of categories to which it belongs. The candidate content in each round of selection has not been selected before the current round of selection.
[0181] The recommendation module 803 is used to recommend content to target users according to the ranking of each content to be recommended in the completed recommendation order.
[0182] In a possible implementation, the acquisition module 801 is configured to:
[0183] Obtain a to-be-recommended list corresponding to the target user. The to-be-recommended content in the to-be-recommended list is recalled based on the user feature data of the target user. The recommendation order of the to-be-recommended content in the to-be-recommended list is related to the matching degree between each to-be-recommended content and the user feature data;
[0184] According to the matching degree of each to-be-recommended content in the to-be-recommended list, start selecting at least two categories of to-be-recommended content from the to-be-recommended content with high matching degree.
[0185] In a possible implementation manner, the sorting module 802 is used to:
[0186] Select the candidate content with the largest predicted recommendation contribution value from all the candidate content in this round as the content selected in this round.
[0187] In a possible implementation manner, the sorting module 802 is used to:
[0188] Determine a predetermined number of candidate content starting from the candidate content with the largest predicted recommendation contribution value in the order of the predicted recommendation contribution value from large to small. The predetermined number is an integer greater than or equal to 2;
[0189] Determine the content selected in this round from the predetermined number of candidate content.
[0190] In a possible implementation manner, the sorting module 802 is used to:
[0191] In the order of the predicted recommendation contribution value from large to small, if the difference between the predicted recommendation contribution values of two adjacent candidate content among the predetermined number of candidate content is less than the difference threshold, then determine the content selected in this round from the predetermined number of candidate content.
[0192] In a possible implementation manner, the sorting module 802 is used to:
[0193] Determine all the predetermined number of candidate content as the content selected in this round; or,
[0194] Select the candidate content that meets the priority sorting condition from the predetermined number of candidate content as the content selected in this round.
[0195] In a possible implementation manner, the sorting module 802 is used to:
[0196] Use the candidate content among the predetermined number of candidate content that is different in category from the content selected in the previous round as the content selected in this round; or,
[0197] Determine the target category for which the number of selected content included in at least two categories before this round of selection meets the set quantity limit, and use the candidate content belonging to the target category among the predetermined number of candidate content as the content selected in this round.
[0198] In a possible implementation manner, the sorting module 802 is configured to:
[0199] Use the candidate content with the highest original sorting among a predetermined number of candidate contents as the content selected in this round, where the original sorting of the candidate contents refers to the sorting in descending order according to the matching degree between the candidate contents and the user feature data of the target user.
[0200] In a possible implementation manner, the sorting module 802 is configured to:
[0201] Calculate the first total recommended contribution prediction value of all the selected contents in the category set to which each candidate content belongs through a submodular function;
[0202] Calculate the second total recommended contribution prediction value of each candidate content and all the selected contents in the category set to which it belongs through a submodular function;
[0203] Determine the difference between the second total recommended contribution prediction value and the first total recommended contribution prediction value of each candidate content as the recommended contribution prediction value of the candidate content.
[0204] In a possible implementation manner, the sorting module 802 is configured to:
[0205] Arrange the recommended order for the content selected in each round in such a way that the earlier the selection round, the higher the recommended order of the selected content.
[0206] All relevant contents of each step involved in the embodiments of the foregoing method for recommending content can be cited in the function descriptions of the corresponding functional modules of the device for recommending content in the embodiments of the present application, and will not be elaborated herein.
[0207] The division of modules in the embodiments of the present application is illustrative, merely a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, the functional modules can be integrated in one processor, or exist separately physically, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0208] Based on the same inventive concept, the embodiments of the present application provide a computing device, which is, for example, the server 104 described above Figure 1 This computing device is capable of executing the method for recommending content provided by the embodiments of the present application, such as Figure 9As shown, the computing device in the embodiment of the present application includes at least one processor 901, as well as a memory 902 and a communication interface 903 connected to the at least one processor 901. In the embodiment of the present application, the specific connection medium between the processor 901 and the memory 902 is not limited. Figure 9 Here, it is taken as an example that the processor 901 and the memory 902 are connected through a bus 900. The bus 900 is represented by a thick line in Figure 9 The connection manners between other components are only for illustrative purposes and are not to be taken as limiting. The bus 900 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 9 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0209] In the embodiment of the present application, the memory 902 stores a computer program executable by the at least one processor 901. By executing the computer program stored in the memory 902, the at least one processor 901 can execute the steps included in the method of the foregoing recommended content.
[0210] Among them, the processor 901 is the control center of the computing device. It can connect various parts of the entire computing device through various interfaces and lines. By running or executing the instructions stored in the memory 902 and calling the data stored in the memory 902, the various functions of the computing device and process data, so as to monitor the computing device as a whole. Optionally, the processor 901 may include one or more processing modules. The processor 901 may integrate an application processor and a modem processor. Among them, the processor 901 mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 901 either. In some embodiments, the processor 901 and the memory 902 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.
[0211] The processor 901 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, which can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0212] The memory 902 serves as a non-volatile computer-readable storage medium and can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 902 may include at least one type of storage medium. For example, it may include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 902 is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 902 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data.
[0213] The communication interface 903 is a transmission interface capable of performing communication. Data can be received or sent through the communication interface 903. For example, data interaction with other devices can be performed through the communication interface 903 to achieve the purpose of communication.
[0214] Furthermore, the computing device further includes a basic input / output system (I / O system) 904 for facilitating information transmission between various components within the computing device, and a mass storage device 908 for storing an operating system 905, application programs 906, and other program modules 907.
[0215] The basic input / output system 904 includes a display 909 for displaying information and input devices 910 such as a mouse and a keyboard for user input of information. Both the display 909 and the input devices 910 are connected to the processor 901 through the basic input / output system 904 connected to the system bus 900. The basic input / output system 904 may further include an input / output controller for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller also provides outputs to a display screen, a printer, or other types of output devices.
[0216] The large-capacity storage device 908 is connected to the processor 901 through a large-capacity storage controller (not shown) connected to the system bus 900. The large-capacity storage device 908 and its associated computer-readable medium provide non-volatile storage for the server package. That is, the large-capacity storage device 908 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM drive.
[0217] According to various embodiments of the present application, the computing device package may also run on a remote computer on the network through a network such as the Internet. That is, the computing device may be connected to the network 911 through the communication interface 903 connected to the system bus 900, or in other words, the communication interface 903 may also be used to connect to other types of networks or remote computer systems (not shown).
[0218] Based on the same inventive concept, an embodiment of the present application also provides a storage medium, which may be a computer-readable storage medium. Computer instructions are stored in the storage medium. When the computer instructions run on a computer, the computer is caused to execute the steps of the method of the recommended content as described above.
[0219] Based on the same inventive concept, an embodiment of the present application also provides a chip system, which includes a processor and may also include a memory for implementing the steps of the method of the recommended content as described above. The chip system may be composed of chips or may include chips and other discrete devices.
[0220] In some possible implementation manners, various aspects of the method for recommended content provided in the embodiments of the present application may also be implemented in the form of a program product, which includes program code. When the program product runs on a computer, the program code is used to cause the computer to execute the steps in the method for recommended content according to various exemplary embodiments of the present application described above.
[0221] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0222] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and variations.
Claims
1. A method for recommended content, characterized in that, The method includes: Obtaining at least two categories of content to be recommended for a target user, where each category includes at least one piece of content to be recommended; Performing multiple rounds of selection on the at least two categories of content to be recommended according to the predicted recommendation contribution values, and determining the recommendation order of the content selected in each round according to the selection round until all the content to be recommended is sorted; wherein, in each round of selection: respectively determining the predicted recommendation contribution value of each candidate content in this round relative to the set of categories to which it belongs; from all the candidate content in this round, determining a predetermined number of candidate content starting from the one with the largest predicted recommendation contribution value in descending order of the predicted recommendation contribution value; determining the target category in which the number of selected content included in the at least two categories before this round of selection meets the set quantity limit, and taking the candidate content belonging to the target category among the predetermined number of candidate content as the content selected in this round; removing the selected content from the candidate content and correspondingly adding it to the set of categories to which it belongs; wherein, the predicted recommendation contribution value of a candidate content is: the incremental predicted recommendation contribution of the one candidate content relative to the selected content in the set of categories to which it belongs; the predetermined number is an integer greater than or equal to 2; the candidate content in each round of selection has not been selected before this round of selection; According to the sorting of each piece of content to be recommended with the recommended order completed, performing content recommendation to the target user.
2. The method according to claim 1, wherein Obtaining at least two categories of content to be recommended for a target user includes: Obtaining the list of content to be recommended corresponding to the target user, where the content to be recommended in the list of content to be recommended is recalled according to the user feature data of the target user, and the recommendation order of the content to be recommended in the list of content to be recommended is related to the matching degree between each piece of content to be recommended and the user feature data; Selecting at least two categories of content to be recommended starting from the content to be recommended with a high matching degree according to the matching degree of each piece of content to be recommended in the list of content to be recommended.
3. The method according to claim 1, wherein Determining the target category in which the number of selected content included in the at least two categories before this round of selection meets the set quantity limit, and taking the candidate content belonging to the target category among the predetermined number of candidate content as the content selected in this round, includes: In descending order of the predicted recommendation contribution value, if the difference between the predicted recommendation contribution values of two adjacent candidate content among the predetermined number of candidate content is less than the difference threshold, then determining the target category in which the number of selected content included in the at least two categories before this round of selection meets the set quantity limit, and taking the candidate content belonging to the target category among the predetermined number of candidate content as the content selected in this round.
4. The method according to any one of claims 1-3, characterized in that, Respectively determining the predicted recommendation contribution value of each candidate content in this round relative to the set of categories to which it belongs, includes: Calculating the first total predicted recommendation contribution value of all the selected content in the set of categories to which each candidate content belongs through a submodular function; Calculating the second total predicted recommendation contribution value of each candidate content and all the selected content in the set of categories to which it belongs through the submodular function; Determine the difference between the second total recommended contribution prediction value and the first total recommended contribution prediction value of each candidate content as the recommended contribution prediction value of the candidate content.
5. The method according to any one of claims 1 to 3, characterized in that, Determine the recommended order of the content selected in each round according to the selection round, including: Arrange the recommended order of the content selected in each round in such a way that the earlier the selection round, the higher the recommended order of the corresponding selected content.
6. A device for recommended content, characterized in that The device includes: An acquisition module, configured to acquire at least two categories of content to be recommended for a target user, where each category includes at least one content to be recommended; A sorting module, configured to perform multiple rounds of selection on the at least two categories of content to be recommended according to the recommended contribution prediction value, and determine the recommended order of the content selected in each round according to the selection round until all the content to be recommended is sorted; where, in each round of selection: respectively determine the recommended contribution prediction value of each candidate content in this round relative to the set of categories to which it belongs; from all the candidate content in this round, determine a predetermined number of candidate content starting from the candidate content with the largest recommended contribution prediction value in descending order of the recommended contribution prediction value; determine the target category in which the number of selected content included in the at least two categories before this round of selection meets the set quantity limit, and use the candidate content belonging to the target category among the predetermined number of candidate content as the content selected in this round; remove the selected content from the candidate content and add it to the set of categories to which it belongs correspondingly, where the recommended contribution prediction value of a candidate content is: the incremental recommended contribution of the candidate content relative to the selected content in the set of categories to which it belongs; the predetermined number is an integer greater than or equal to 2; the candidate content in each round of selection has not been selected before this round of selection; A recommendation module, configured to perform content recommendation to the target user according to the sorting of each content to be recommended for which the recommended order has been completed.
7. The device according to claim 6, characterized in that, The acquisition module is used for: Acquire the list of content to be recommended corresponding to the target user, where the content to be recommended in the list of content to be recommended is recalled according to the user feature data of the target user, and the recommended order of the content to be recommended in the list of content to be recommended is related to the matching degree between each content to be recommended and the user feature data; According to the matching degree of each content to be recommended in the list of content to be recommended, select at least two categories of content to be recommended starting from the content to be recommended with a high matching degree.
8. A computing device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps included in the method according to any one of claims 1-5.
9. A storage medium, characterized in that, The storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the steps included in the method according to any one of claims 1-5.
Citation Information
Patent Citations
Information recommendation method and apparatus
CN109086439A
Content recommendation method and device, storage medium and computer equipment
CN110263244A
Resource recommendation method, device and equipment and storage medium
CN111400615A