A content distribution method and system

By obtaining the relationship network of the intermediate user and combining the relationship network and attribute information of the first publisher, the interest similarity value between the second publisher and the first publisher is calculated, the target publisher is determined and the content to be distributed to the users in the relationship network of the attention network is solved, and the problem of low data usage and low recommendations in the prior art is achieved, and more efficient and personalized content distribution is achieved.

CN116389564BActive Publication Date: 2025-06-17MICRO DREAM TECHTRONIC NETWORK TECH CHINACO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310238638.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2025-06-17
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

The prior art mines the similarity between users or items through user historical behavior data, with few data usage, low recommendation accuracy, poor results, and no personalization.

Method used

By obtaining the relationship network of the intermediate user, combining the relationship network and attribute information of the first publisher, the interest similarity value between the second publisher and the first publisher is calculated, the target publisher is determined and the content to be distributed to the users in the relationship network of the attention network is distributed.

Benefits of technology

The range of available data has been expanded, the recommendation effect has been improved, the recommendation results are closer to the user's own preferences, the recommendation accuracy in social scenarios has been increased, and personalized needs are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116389564B_ABST
    Figure CN116389564B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a content distribution method and system, which relates to the field of content distribution. The method includes: for the content to be distributed by a first publisher, determining all second publishers that each intermediate user follows; for each second publisher, determining a first similarity value between the two based on the follow relationship network of the second publisher and the follow relationship network of the first publisher; determining a second similarity value between the two based on the attribute information of the second publisher and the attribute information of the first publisher; determining an interest similarity value between the two based on the first similarity value and the second similarity value; determining the second publishers with an interest similarity value greater than a preset similarity threshold as target publishers, and using the users corresponding to the follow relationship network of the target publishers as users to be recommended. Through the follow relationship network of the publishers of the intermediate users, the range of users to be recommended becomes wider, improving the recommendation effect; by integrating the attribute information of the publishers, personalized needs are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of content distribution, and particularly to a content distribution method and system. Background Art

[0002] With the rapid development of network technology in recent years, a vast amount of information has come into the view of users. While bringing convenience to users, it also brings the problem of "information overload". It is difficult for users to obtain interesting or valuable information from the vast amount of information. Therefore, recommendation systems emerge to meet the demand. Recommendation systems calculate users' interest preferences based on users' behavioral data, discover users' interest points, thereby guiding users to discover their own interest needs, helping users filter and screen the required information, and improving the efficiency of users in obtaining effective information. The prior art mines the similarity between users or items through users' historical behavioral data, makes decisions and recommendations for users based on the similarity, can use less data, has low recommendation accuracy, poor effects, and the recommendations are not personalized. Summary of the Invention

[0003] Embodiments of the present invention provide a content distribution method and system, which can solve the technical problems in the prior art that the similarity between users or items is mined through users' historical behavioral data, the data that can be used is less, the recommendation accuracy is low, the effects are poor, and the recommendations are not personalized.

[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides a content distribution method, including:

[0005] For the content to be distributed of the first publisher, obtain all intermediate users who have generated user behaviors for the content to be distributed;

[0006] Determine all second publishers followed by each of the intermediate users; wherein, the second publisher refers to other publishers who have published content and are other than the first publisher;

[0007] For each of the second publishers, determine a first similarity value between the second publisher and the first publisher according to the follow-up relationship network of the second publisher and the follow-up relationship network of the first publisher; determine a second similarity value between the second publisher and the first publisher according to the attribute information of the second publisher and the attribute information of the first publisher; and

[0008] Based on the first similarity value and the second similarity value, determine an interest similarity value between the second publisher and the first publisher;

[0009] Determine the second publisher whose interest similarity value is greater than the preset similarity threshold as the target publisher, and use the users corresponding to the attention relationship network of the target publisher as the users to be recommended;

[0010] Distribute the content to be distributed to each of the users to be recommended.

[0011] On the other hand, an embodiment of the present invention provides a content distribution system, including:

[0012] A circumscribing unit, configured to, for the content to be distributed of a first publisher, obtain all intermediate users who have performed user actions on the content to be distributed; determine all second publishers followed by each of the intermediate users; wherein, the second publisher refers to other publishers who have published content and are other than the first publisher;

[0013] The content distribution system further includes: a calculation unit, a determination unit, and a recommendation unit that are each executed for each intermediate user, wherein:

[0014] The calculation unit is configured to, for each of the second publishers, determine a first similarity value between the second publisher and the first publisher according to the attention relationship network of the second publisher and the attention relationship network of the first publisher; determine a second similarity value between the second publisher and the first publisher according to the attribute information of the second publisher and the attribute information of the first publisher; and

[0015] Based on the first similarity value and the second similarity value, determine the interest similarity value between the second publisher and the first publisher;

[0016] The determination unit is configured to determine the second publisher whose interest similarity value is greater than the preset similarity threshold as the target publisher, and use the users corresponding to the attention relationship network of the target publisher as the users to be recommended;

[0017] The recommendation unit is configured to distribute the content to be distributed to each of the users to be recommended.

[0018] The above technical solution has the following beneficial effects: for the content to be distributed by the first publisher, all intermediate users who have generated user behaviors for the content to be distributed are obtained; for each second publisher followed by each intermediate user, according to the attention relationship network of all second publishers of each intermediate user and the attention relationship network of the first publisher, the first similarity value between the second publisher and the first publisher is determined; by using intermediate users to lock more publishers and leveraging the attention relationship networks of the locked publishers, the available data range is expanded, making the range of users to be recommended wider and the quantity larger, thereby improving the recommendation effect. According to the attribute information of the second publisher and the attribute information of the first publisher, the second similarity value between the second publisher and the first publisher is determined; based on the first similarity value and the second similarity value, the interest similarity value between the second publisher and the first publisher is obtained; the second publishers with interest similarity values greater than the preset interest similarity value are determined as target publishers, and the users corresponding to the attention relationship networks of the target publishers are used as users to be recommended; by integrating the attribute information of the publishers, the potential interest characteristics of users can be mined according to the attribute information of the publishers, making the recommendation result closer to the own preferences of the first publisher and increasing the recommendation accuracy in the social scenario; at the same time, integrating the attribute information of the publishers also enables the recommendation to meet personalized needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0020] Figure 1 is a flowchart of a content distribution method according to an embodiment of the present invention;

[0021] Figure 2 is a structural diagram of a content distribution system according to an embodiment of the present invention;

[0022] Figure 3 is a flowchart of another content distribution method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] AsFigure 1 As shown, in combination with the embodiments of the present invention, a content distribution method is provided, including:

[0025] S101: For the content to be distributed of the first publisher, obtain all intermediate users who have performed user actions on the content to be distributed;

[0026] S102: Determine all second publishers followed by each of the intermediate users; wherein, the second publisher refers to other publishers who have published content and are other than the first publisher;

[0027] S103: For each of the second publishers, determine a first similarity value between the second publisher and the first publisher according to the follow-up relationship network of the second publisher and the follow-up relationship network of the first publisher; determine a second similarity value between the second publisher and the first publisher according to the attribute information of the second publisher and the attribute information of the first publisher; and, based on the first similarity value and the second similarity value, determine an interest similarity value between the second publisher and the first publisher;

[0028] S104: Determine the second publishers with the interest similarity value greater than the preset similarity threshold as target publishers, and use the users corresponding to the follow-up relationship network of the target publishers as users to be recommended;

[0029] S105: Distribute the content to be distributed to each of the users to be recommended.

[0030] By using intermediate users to lock in more publishers and leveraging the follow-up relationship networks of the locked-in publishers, the available data range is expanded, making the range of users to be recommended wider and the quantity larger, thereby improving the recommendation effect. By integrating the attribute information of the publishers, the potential interest characteristics of users can be mined according to the attribute information of the publishers, making the recommendation result closer to the own preferences of the first publisher and increasing the recommendation accuracy in the social scenario; at the same time, integrating the attribute information of the publishers also enables the recommendation to meet personalized needs.

[0031] Preferably, the method further includes:

[0032] S106: Take each of the target publishers as the clustering center node of a directed graph, take each user included in the follow-up relationship network of each of the target publishers as an edge node of the directed graph, and take the follow-up relationship of each edge node to the clustering center node as the directed edge of the directed graph, and form the directed graph based on all the clustering center nodes, the edge nodes, and the directed edges; wherein, the target publishers include the first publisher and the second publisher, and the follow-up relationship network of the target publishers is composed of users who follow the target publishers.

[0033] S107: In the directed graph, use the edge nodes of the directed edges pointing to the same cluster center node as the clusters of the cluster center node, and construct a first cluster with the first publisher as the cluster center node and a second cluster with the second publisher as the cluster center node; wherein, the directed edge points from the edge node to the cluster center node. After constructing a directed graph with publishers and their users, extract a first cluster with the first publisher as the cluster center node and a second cluster with the second publisher as the cluster center node from the directed graph for calculating similarity.

[0034] Preferably, S103: For each of the second publishers, determining a first similarity value between the second publisher and the first publisher according to the attention relationship network of the second publisher and the attention relationship network of the first publisher includes:

[0035] S1031: For each of the second publishers, calculate the coincidence degree between the edge nodes of the second cluster corresponding to the second publisher and the edge nodes of the first cluster corresponding to the first publisher to obtain the coincidence degree of the clustering clusters between the second publisher and the first publisher; the coincidence degree of the clustering clusters is the degree of coincidence of the users included in the clusters corresponding to different cluster center nodes, or the count value of the coincidence of other edge nodes except the cluster center nodes in the two clusters. The higher the coincidence degree of the clustering clusters, the higher the similarity between the second publisher and the first publisher.

[0036] S1032: Standardize the coincidence degree of the clustering clusters between the second publisher and the first publisher through the mean and standard deviation of the coincidence degrees of all the second publishers followed by the intermediate users corresponding to the second publisher and the first publisher to obtain the first similarity value between the second publisher and the first publisher. The reason for standardization is that the order of magnitude of the number of cluster center nodes included in each cluster may be different. For example, in Weibo, there are differences in the order of magnitude of the number of fans of each blogger. If no standardization is performed, the order of magnitude of the first similarity value of the center node corresponding to the blogger with a large number of fans will be too large, while weakening other similarity values, such as publisher interaction information and publisher own characteristic information - attribute information, thus causing bias in the recommendation results of content distribution, making it easier for bloggers with a large number of fans and their users to be recommended, resulting in low recommendation accuracy.

[0037] Preferably, S103: Determining a second similarity value between the second publisher and the first publisher according to the attribute information of the second publisher and the attribute information of the first publisher includes:

[0038] S1033: For each of the second publishers, use the user attribute calculation model to calculate the similarity between the attribute information of the second publisher and the attribute information of the first publisher, obtaining a second similarity value between the second publisher and the first publisher; wherein, the user attribute calculation model is obtained through supervised training based on a deep neural network. The attribute information of a publisher includes basic attributes such as gender and age, as well as user interest fields such as entertainment, music, and sports, and behavioral attributes such as recent user interactions and click-through rates in the information stream. The second similarity value of the publisher's attribute information is calculated through machine learning relationships. Use the users corresponding to the attention relationship network of the target publisher as the users to be recommended; by integrating the attribute information of the publisher, the potential interest characteristics of users can be mined based on the attribute information of the publisher, making the recommendation results closer to the own preferences of the first publisher and increasing the recommendation accuracy in the social scenario; at the same time, integrating the attribute information of the publisher also enables the recommendation to meet personalized needs.

[0039] Preferably, S103: The step of determining the interest similarity value between the second publisher and the first publisher based on the first similarity value and the second similarity value includes:

[0040] S1034: Set corresponding weight factors for the first similarity value and the second similarity value respectively;

[0041] S1035: Respectively correct the first similarity value and the second similarity value through their respective weight factors, obtaining a corrected first similarity value and a corrected second similarity value;

[0042] S1036: Take the sum of the corrected first similarity value and the corrected second similarity value as the interest similarity value between the second publisher and the first publisher.

[0043] Set corresponding weight factors for the first similarity value and the second similarity value respectively; enabling the contributions of the attention relationship network corresponding to the first similarity value and the attribute information of the second similarity value publisher to the similarity to be reflected and more rationalized, improving the content distribution recommendation accuracy.

[0044] Preferably, the method further includes:

[0045] S108: Obtain the historical interaction information of the second publisher and the historical interaction information of the first publisher;

[0046] S109: Calculate a third similarity value between the historical interaction information of the second publisher and the historical interaction information of the first publisher based on the collaborative filtering algorithm;

[0047] S103: Determining the interest similarity value between the second publisher and the first publisher based on the first similarity value and the second similarity value includes:

[0048] S1037: Performing a weighted sum of the first similarity value, the second similarity value, and the third similarity value to obtain the interest similarity value between the second publisher and the first publisher.

[0049] When the publisher has historical interaction information, for the historical interaction information of the publisher, based on the traditional collaborative filtering algorithm, for all second publishers of each intermediate user, calculating the third similarity value between the second publisher and the first publisher can also achieve content distribution recommendation; in the present invention, by combining the calculation information of the publisher's attention relationship network and the publisher attribute information, the accuracy of content distribution recommendation is improved.

[0050] As Figure 2 shown, in combination with the embodiments of the present invention, a content distribution system is provided, including:

[0051] A circumscribing unit 21, configured to, for the content to be distributed of the first publisher, obtain all intermediate users who have generated user behaviors for the content to be distributed; determine all second publishers followed by each of the intermediate users; wherein, the second publisher refers to other publishers who have published content and are other than the first publisher;

[0052] The content distribution system further includes: a calculation unit 22, a determination unit 23, and a recommendation unit 24 that are each executed for each intermediate user, wherein:

[0053] The calculation unit 22 is configured to, for each of the second publishers, determine the first similarity value between the second publisher and the first publisher according to the attention relationship network of the second publisher and the attention relationship network of the first publisher; determine the second similarity value between the second publisher and the first publisher according to the attribute information of the second publisher and the attribute information of the first publisher; and

[0054] Based on the first similarity value and the second similarity value, determine the interest similarity value between the second publisher and the first publisher;

[0055] The determination unit 23 is configured to determine the second publishers with the interest similarity value greater than a preset similarity threshold as target publishers, and use the users corresponding to the attention relationship network of the target publishers as users to be recommended;

[0056] The recommendation unit 24 is configured to distribute the content to be distributed to each of the users to be recommended.

[0057] By locking more publishers through intermediate users and leveraging the attention relationship network of the locked publishers, the available data range is expanded, the range of users to be recommended becomes wider and the number increases, improving the recommendation effect. By integrating the attribute information of the publishers, the potential interest characteristics of users can be mined based on the attribute information of the publishers, making the recommendation results more in line with the personal preferences of the first publisher and increasing the recommendation accuracy in the social scenario; at the same time, integrating the attribute information of the publishers also enables the recommendation to meet personalized needs.

[0058] Preferably, the content distribution system further includes:

[0059] A directed graph construction unit, configured to use each target publisher as a clustering center node of the directed graph, use each user included in the attention relationship network of each target publisher as an edge node of the directed graph, and use the attention relationship of each edge node to the clustering center node as a directed edge of the directed graph, and form the directed graph based on all the clustering center nodes, the edge nodes, and the directed edges; wherein, the target publishers include the first publisher and the second publisher, and the attention relationship network of the target publisher is composed of users who follow the target publisher; wherein, the target publishers include the first publisher and the second publisher, and the attention relationship network of the target publisher is composed of users who follow the target publisher.

[0060] A cluster construction unit, configured to, within the directed graph, use the edge nodes of the directed edges pointing to the same clustering center node as the cluster of the clustering center node, and construct a first cluster with the first publisher as the clustering center node and a second cluster with the second publisher as the clustering center node; wherein, the directed edge points from the edge node to the clustering center node. After constructing a directed graph with the publisher and its users, extract from the directed graph a first cluster with the first publisher as the clustering center node and a second cluster with the second publisher as the clustering center node for calculating similarity.

[0061] Preferably, the calculation unit 22 includes a first calculation subunit, configured to, for each second publisher, calculate the coincidence degree between the edge nodes of the second cluster corresponding to the second publisher and the edge nodes of the first cluster corresponding to the first publisher, to obtain the clustering cluster coincidence degree between the second publisher and the first publisher; the clustering cluster coincidence degree is the coincidence degree of the users included in the clusters corresponding to different clustering center nodes, or the count value of the coincidence of other edge nodes except the clustering center nodes in the two clusters. The higher the clustering cluster coincidence degree, the higher the similarity between the second publisher and the first publisher.

[0062] Based on the mean and standard deviation of the coincidence degrees of the clustering clusters of all the second publishers followed by the intermediate users corresponding to the second publisher and the first publisher, the coincidence degree of the clustering clusters of the second publisher and the first publisher is standardized to obtain the first similarity value between the second publisher and the first publisher. The reason for standardization is that the order of magnitude of the clustering center nodes included in each cluster may be different. For example, in Weibo, there are differences in the order of magnitude of the number of followers of each blogger. If no standardization is performed, the order of magnitude of the first similarity value of the central node corresponding to the blogger with a large number of followers will be too large, while weakening other similarity values, such as the publisher interaction information and the publisher's own characteristic information - attribute information. As a result, the recommendation results of content distribution are biased, making it easier for bloggers with a large number of followers and their users to be recommended, resulting in low recommendation accuracy.

[0063] Preferably, the calculation unit 22 includes:

[0064] A second calculation subunit, configured to, for each of the second publishers, calculate the similarity between the attribute information of the second publisher and the attribute information of the first publisher by using a user attribute calculation model to obtain a second similarity value between the second publisher and the first publisher; wherein, the user attribute calculation model is obtained by supervised training based on a deep neural network.

[0065] The attribute information of the publisher includes basic attributes such as gender and age, as well as user interest fields such as entertainment, music, sports, etc., and behavioral attributes such as recent interactions and click-through rates of the user in the information stream. The second similarity value of the publisher's attribute information is calculated through machine learning relationships. The users corresponding to the attention relationship network of the target publisher are used as the users to be recommended; by integrating the attribute information of the publisher, the potential interest characteristics of the users can be mined according to the attribute information of the publisher, making the recommendation results closer to the own preferences of the first publisher and increasing the recommendation accuracy in the social scenario; at the same time, integrating the attribute information of the publisher also enables the recommendation to meet personalized needs.

[0066] Preferably, the calculation unit 22 includes:

[0067] An interest similarity calculation subunit is configured to respectively set corresponding weight factors for the first similarity value and the second similarity value; respectively correct the first similarity value and the second similarity value through their respective weight factors to obtain a corrected first similarity value and a corrected second similarity value; and use the sum of the corrected first similarity value and the corrected second similarity value as the interest similarity value between the second publisher and the first publisher. By respectively setting corresponding weight factors for the first similarity value and the second similarity value, the contributions of the attention relationship network corresponding to the first similarity value and the attribute information of the second similarity value publisher to the similarity can be reflected and made more reasonable, thereby improving the accuracy of content distribution recommendation.

[0068] Preferably, the calculation unit 22 is further configured to obtain the historical interaction information of the second publisher and the historical interaction information of the first publisher; and calculate a third similarity value between the historical interaction information of the second publisher and the historical interaction information of the first publisher based on a collaborative filtering algorithm.

[0069] The calculation unit 22 includes:

[0070] An interest similarity calculation subunit is configured to perform weighted summation on the first similarity value, the second similarity value, and the third similarity value to obtain the interest similarity value between the second publisher and the first publisher.

[0071] The above technical solutions of the embodiments of the present invention will be described in detail below in conjunction with specific application examples. For technical details not introduced during the implementation process, reference can be made to the relevant descriptions above.

[0072] When the publisher has historical interaction information, for the historical interaction information of the publisher, based on the traditional collaborative filtering algorithm, for all second publishers of each intermediate user, calculate the third similarity value between the second publisher and the first publisher, and content distribution recommendation can also be achieved; in the present invention, by combining the calculation information of the publisher's attention relationship network and the publisher's attribute information, the accuracy of content distribution recommendation is improved.

[0073] The embodiments of the present invention provide a content distribution method and system, which are used in various recommendation scenarios, such as the Weibo attention recommendation scenario, and can enable more users who follow key bloggers to participate in interactions. The user historical interaction data, the data related to the user's attention relationship network, and the user's attribute information can be used to guide the learning of the user's interaction preferences. By integrating the calculation information of the attention relationship network and the publisher's attribute information into the interest similarity algorithm, the attention relationship network and the publisher's attribute information are fully utilized.

[0074] To a certain extent, the attention relationship network of the blogger (publisher) is equivalent to clustering the user's attention preferences. The interaction preferences of the fans of the bloggers within similar clusters are also very similar, which is very important prior information; it has a more ideal recommendation effect in content recommendation than the user historical interaction data used by the collaborative filtering algorithm; it can also solve the cold start problem.

[0075] For platforms such as Weibo with a large number of users, there are also quite large characteristic differences between users and users, and between publishers and publishers (such as between bloggers and bloggers). Using the attribute information of users and bloggers can also help the content distribution system (or called content recommendation system) to perform differential distribution.

[0076] A content distribution method according to an embodiment of the present invention, as Figure 3 shown, for the content to be distributed of the first publisher, obtain all intermediate users who have generated user behavior for the content to be distributed; determine all second publishers followed by each of the intermediate users; wherein, the second publisher refers to other publishers who have published content and are outside the first publisher.

[0077] For each intermediate user, based on the social network behavior information of the second publisher - the attention relationship network, based on the historical interaction information of the first publisher and the historical interaction information of the second publisher, the attribute information of the second publisher and the attribute information of the first publisher, set corresponding weight factors for the first similarity value and the second similarity value respectively; correct the first similarity value and the second similarity value respectively through their respective weight factors to obtain the corrected first similarity value and the corrected second similarity value; take the sum of the corrected first similarity value and the corrected second similarity value as the interest similarity value between the second publisher and the first publisher. Based on the interest similarity function sim SN to measure for content distribution recommendation, the interest similarity function sim SN The formula of is:

[0078]

[0079] wherein, i represents the second publisher of the clustering center node; j represents the first publisher of the clustering center node; P i,j represents the first similarity value between the first publisher of the clustering center node j and the second publisher of the clustering center node i; D i,j is the second similarity value; is the third similarity value; k is the weight factor of the first similarity value; s is the weight factor of the second similarity value.

[0080] Determine the second publisher whose interest similarity value is greater than the preset similarity threshold as the target publisher, and use the users corresponding to the attention relationship network of the target publisher as the users to be recommended; distribute the content to be distributed to each of the users to be recommended.

[0081] Interest similarity function The specific definitions and meanings of the similarity values of each part are described as follows:

[0082] I. The first similarity value P based on the attention relationship network i,j

[0083] Take each target publisher as the clustering center node of the directed graph, take each user included in the attention relationship network of each target publisher as the edge node of the directed graph, and take the attention relationship of each edge node to the clustering center node as the directed edge of the directed graph. Form the directed graph G based on all the clustering center nodes, edge nodes, and directed edges; the directed graph G has a large number of nodes, and the set of clustering center nodes is S.

[0084] In the directed graph, take the edge nodes of the directed edges pointing to the same clustering center node as the cluster of the clustering center node, and construct the first cluster with the first publisher as the clustering center node and the second cluster with the second publisher as the clustering center node, that is, the users within the cluster have the same browsing and interaction preferences. Among them, the target publisher includes the first publisher and the second publisher, the attention relationship network of the target publisher is composed of users who follow the target publisher, and the directed edge points from the edge node to the clustering center node.

[0085] For all the second publishers followed by each intermediate user, calculate the coincidence degree between the edge nodes of the second cluster corresponding to the second publisher and the edge nodes of the first cluster corresponding to the first publisher, and obtain the clustering cluster coincidence degree between the second publisher and the first publisher; through the mean and standard deviation of the clustering cluster coincidence degrees between all the second publishers followed by the intermediate user corresponding to the second publisher and the first publisher, standardize the clustering cluster coincidence degree between the second publisher and the first publisher to obtain the first similarity value between the second publisher and the first publisher:

[0086]

[0087] Among them, C i,j is the clustering cluster coincidence degree between the second publisher of the clustering center node i and the first publisher of the clustering center node j, that is, the coincidence degree of the users included in the clusters corresponding to different clustering center nodes, or the count value of the coincidence of other edge nodes except the clustering center nodes in the two clusters; μi The mean degree, σ, of the overlap C of the clustering clusters of all the second publishers followed by the intermediate users corresponding to the second publishers and the first publisher i,k∈s i The mean degree, σ, of the overlap C of the clustering clusters of all the second publishers followed by the intermediate users corresponding to the second publishers and the first publisher i,k∈s The standard deviation, where k represents the publishers in S.

[0088] The first similarity value P i,j In P, the reason for normalizing C i,j is as follows: The order of magnitude of the clustering center nodes contained in each cluster may be different. For example, in Weibo, there are differences in the order of magnitude of the number of fans of each blogger. If no normalization is performed, the order of magnitude of the first similarity value of the central node corresponding to the blogger with a high number of fans will be too large, while weakening other similarity values, such as the publisher interaction information and the publisher's own characteristic information - attribute information, resulting in bias in the recommendation results of content distribution, making it easier for bloggers with a high number of fans and their users to be recommended, leading to low recommendation accuracy. Therefore, the normalized P i,j can accurately characterize the similarity degree of the follow relationship between the clustering center nodes i and j. Furthermore, it can fuse the attribute information of the publisher, such as the user interest preferences implied in the follow information, and the historical interaction information of the traditional collaborative filtering algorithm, to improve the adaptability of content distribution in various scenarios.

[0089] II. The second similarity value D based on the attribute information of the publisher i,j

[0090] The attribute information of the publisher includes basic attributes such as gender and age, as well as user interest fields such as entertainment, music, sports, etc., and behavioral attributes such as the recent interaction and click-through rate of the user in the information stream. The second similarity value of the publisher's attribute information is calculated through machine learning relationships.

[0091] The attribute information of the publisher is transformed into a low-dimensional vector embedding. For each of the second publishers, a user attribute calculation model is used to calculate the similarity between the low-dimensional vector of the attribute information of the second publisher and the low-dimensional vector of the attribute information of the first publisher. The similarity calculation is the cosine similarity, and the second similarity value D between the second publisher and the first publisher is obtained i,j , and the formula is as follows:

[0092]

[0093] where cos represents calculating the second similarity value using cosine, emb represents the low-dimensional vector of the publisher's attribute information, and dim = N represents the dimension of the low-dimensional vector. ​

[0094] The user attribute calculation model is obtained by supervised training of the vector of the publisher's attribute information based on a deep neural network. The dimension of the low-dimensional vector embedding needs to be dynamically adjusted based on different scenarios and features. One of the important reasons for using a deep neural network to depict the recommended item attributes is that the deep network can utilize cross features to a certain extent, improving the accuracy of content distribution. Among them, the cross feature means a high-order feature obtained by combining two or more independent features, which can map linear features to a high-dimensional space and enhance the fitting ability of complex relationships.

[0095] III. Regarding the historical interaction information of users

[0096] Obtain the historical interaction information of the second publisher and the historical interaction information of the first publisher; calculate the third similarity value between the historical interaction information of the second publisher and the historical interaction information of the first publisher based on the traditional collaborative filtering algorithm; further, perform a weighted sum of the first similarity value, the second similarity value, and the third similarity value to obtain the interest similarity value between the second publisher and the first publisher. Specifically:

[0097] Regarding the historical interaction information of the publisher, based on the traditional collaborative filtering algorithm, for all second publishers of each intermediate user, calculate the third similarity value between the second publisher and the first publisher The method cited here is consistent with the traditional collaborative filtering algorithm - the user-based collaborative filtering algorithm (UserCF). The user-based collaborative filtering algorithm recommends content that other users with similar interests to the user itself like. Different calculation methods need to be selected according to different scenarios. Commonly used calculation methods for the third similarity value include:

[0098] 1. Cosine similarity: Measures the angle between the vector of the historical interaction information of the second publisher i and the vector of the historical interaction information of the first publisher j. The smaller the angle, the greater the similarity. The formula is as follows:

[0099]

[0100] Among them, ||i|| represents the norm of the vector of the historical interaction information of the second publisher i, and ||j|| represents the norm of the vector of the historical interaction information of the first publisher j.

[0101] 2. Use the Jaccard similarity coefficient to represent the third similarity value, which is used to measure the similarity between sets. The Jaccard similarity of the attribute set A and the attribute set B is represented by J(A,B)

[0102] It is expressed as:

[0103]

[0104] Among them, A represents the set of attributes of the second publisher i, B represents the set of the first publisher j, |A∩B| represents the number of common attributes of sets A and B, and |A∪B| represents the total number of attributes included in sets A and B.

[0105] The technical means of the content distribution method of the present invention are experimented in specific business scenarios. A total of 41 categories of bloggers, readers, and blog post features are cited. On the premise that the target publisher is a blogger, the publisher is a blogger, the content to be distributed is a blog post, the intermediate user is a reader, and the attention relationship network of the reader is introduced. The blog post features are used to describe the attributes of the blogger. Because the blog posts and readers corresponding to the same blogger have similar attributes, the blog post features can be used to assist in depicting the characteristics of the blogger himself and enhance the representational ability of the features for the blogger's attributes. According to the experiment, the optimal vector dimension is 64 dimensions.

[0106] And in different scenarios, the description method of the attributes (distribution content features) of the publisher can be flexibly changed according to needs: either use a more suitable neural network to deeply describe the recommended items, or use methods such as graph embedding to specifically characterize the attributes of the distributed content recommended in a specific scenario. Finding a method for describing the item attributes suitable for the current business scenario can improve the recommendation effect to a certain extent.

[0107] The embodiment of the present invention makes full use of the attention relationship network by using the second similarity value, and makes full use of the attribute information of the publisher itself by using the third similarity value. It integrates the user's interest, explores the role of cross features in recommendation, improves the degree of fit between the distributed content and the user's preferences, greatly increases the representational ability of the distributed content, thereby improving the recommendation quality and the recommendation effect.

[0108] The beneficial technical effects achieved by the embodiments of the present invention are as follows:

[0109] 1. By calculating the attention relationship network between the second publisher and the first publisher in the attention relationship network through the attention relationship of the social network of the intermediate user, that is, the attention relationship network, integrating the prior information of the user's interest preferences, adapting to the social business scenario, and enhancing the expression ability in different social scenarios. At the same time, the social attention network is transformed into a directed graph topological structure, and for each intermediate user, clustering calculations are performed with the second publisher as the clustering center node and the first publisher as the clustering center node to obtain the first similarity value; the recommendation accuracy in different social scenarios is increased.

[0110] 2. Incorporating the attribute information of the publisher makes the recommendation results closer to the user's own preferences. At the same time, the user attribute calculation model obtained through supervised training based on a deep neural network introduces cross features, mines the latent interest characteristics of users, represents user attributes in high dimensions, and integrates the long-term interests of users, making the recommendation results closer to the user's real needs and realizing personalized recommendation.

[0111] 3. When the user's historical behavior information data is insufficient, cold start recommendation is performed on the distribution content of the publisher only based on the publisher's own attributes or the information of the attention relationship network, so that users lacking historical data can also obtain recommendation results that are more in line with their needs.

[0112] 4. Recommending the distribution content of the publisher only based on the publisher's own attributes or the information of the attention relationship network avoids the problem of inaccurate recommendation results when the data scale is large and the user's behavior matrix is relatively sparse, and can obtain better recommendation effects even when the data is sparse.

[0113] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.

[0114] In the above detailed description, various features are combined in a single embodiment to simplify the present disclosure. This method of disclosure should not be construed as reflecting an intention that the embodiments of the claimed subject matter require more features than are clearly recited in each claim. On the contrary, as reflected by the appended claims, the present invention lies in a state with fewer features than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby clearly incorporated into the detailed description, where each claim stands alone as a separate preferred embodiment of the present invention.

[0115] In order to enable any person skilled in the art to implement or use the present invention, the above disclosed embodiments have been described. For those skilled in the art; various modification methods of these embodiments are obvious, and the general principles defined herein can also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.

[0116] The foregoing description includes examples of one or more embodiments. Of course, it is not possible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but those of ordinary skill in the art should recognize that the various embodiments can be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, this term is inclusive in a manner similar to the term "including" as interpreted when used as a transitional word in a claim. In addition, any use of the term "or" in the specification or claims is to be construed as "non-exclusive or".

[0117] Those skilled in the art will also appreciate that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention can be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the various illustrative components, units, and steps have been generally described in terms of their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the overall system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present invention.

[0118] The various illustrative logical blocks or units described in the embodiments of the present invention can be implemented or operated to perform the described functions by a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other similar configuration.

[0119] In the embodiments of the present invention, the steps of the methods or algorithms described may be directly incorporated into hardware, software modules executed by a processor, or a combination of both. The software modules may be stored in a RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium may be connected to the processor such that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium may also be integrated into the processor. The processor and the storage medium may be provided in an ASIC, and the ASIC may be provided in a user terminal. Optionally, the processor and the storage medium may also be provided in different components of the user terminal.

[0120] In one or more exemplary designs, the above-described functions in the embodiments of the present invention may be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions may be stored on a computer-readable medium or transmitted on a computer-readable medium in the form of one or more instructions or codes. A computer-readable medium includes a computer storage medium and a communication medium that facilitates the transfer of a computer program from one place to another. The storage medium may be any available medium accessible by a general or special computer. For example, such a computer-readable medium may include, but is not limited to, RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and other forms readable by a general or special computer, or a general or special processor. In addition, any connection may be appropriately defined as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote resource via a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wirelessly, such as infrared, wireless, and microwave, it is also included in the defined computer-readable medium. The disks and discs include compact disks, laser disks, optical disks, DVDs, floppy disks, and Blu-ray disks. Disks typically reproduce data magnetically, while discs typically reproduce data optically by laser. The above combinations may also be included in the computer-readable medium.

[0121] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A content distribution method, characterized in that, Including: For the content to be distributed for the first publisher, obtain all intermediate users who have performed user actions on the content to be distributed; Determine all second publishers followed by each of the intermediate users; wherein, the second publisher refers to other publishers who have published content and are other than the first publisher; For each of the second publishers, determine a first similarity value between the second publisher and the first publisher according to the follow-up relationship network of the second publisher and the follow-up relationship network of the first publisher; determine a second similarity value between the second publisher and the first publisher according to the attribute information of the second publisher and the attribute information of the first publisher; and Based on the first similarity value and the second similarity value, determine an interest similarity value between the second publisher and the first publisher; Determine the second publishers with the interest similarity value greater than a preset similarity threshold as target publishers, and use the users corresponding to the follow-up relationship network of the target publishers as users to be recommended; Distribute the content to be distributed to each of the users to be recommended.

2. The content distribution method according to claim 1, characterized in that, The method further includes: Taking each of the target publishers as a clustering center node of a directed graph, taking each user included in the follow-up relationship network of each of the target publishers as an edge node of the directed graph, and taking the follow-up relationship of each of the edge nodes to the clustering center node as a directed edge of the directed graph, and forming the directed graph based on all the clustering center nodes, the edge nodes, and the directed edges; In the directed graph, taking the edge nodes of the directed edges pointing to the same clustering center node as the cluster of the clustering center node, and constructing a first cluster with the first publisher as the clustering center node and a second cluster with the second publisher as the clustering center node; Wherein, the target publishers include the first publisher and the second publisher, the follow-up relationship network of the target publishers is composed of users who follow the target publishers, and the directed edge points from the edge node to the clustering center node.

3. The content distribution method according to claim 2, characterized in that, The step of determining a first similarity value between the second publisher and the first publisher according to the follow-up relationship network of the second publisher and the follow-up relationship network of the first publisher for each of the second publishers includes: For each of the second publishers, calculate the coincidence degree between the edge nodes of the second cluster corresponding to the second publisher and the edge nodes of the first cluster corresponding to the first publisher, to obtain the clustering cluster coincidence degree between the second publisher and the first publisher; Standardize the clustering cluster coincidence degree between the second publisher and the first publisher through the mean and standard deviation of the clustering cluster coincidence degrees between all the second publishers followed by the intermediate users corresponding to the second publisher and the first publisher, to obtain the first similarity value between the second publisher and the first publisher.

4. The content distribution method according to claim 1, characterized in that, The step of determining a second similarity value between the second publisher and the first publisher according to the attribute information of the second publisher and the attribute information of the first publisher includes: For each of the second publishers, a user attribute calculation model is used to calculate the similarity between the attribute information of the second publisher and the attribute information of the first publisher, obtaining a second similarity value between the second publisher and the first publisher; wherein, the user attribute calculation model is obtained through supervised training based on a deep neural network.

5. The content distribution method according to claim 1, characterized in that, Determining the interest similarity value between the second publisher and the first publisher based on the first similarity value and the second similarity value includes: Setting corresponding weight factors for the first similarity value and the second similarity value respectively; Respectively correcting the first similarity value and the second similarity value through their respective weight factors, obtaining a corrected first similarity value and a corrected second similarity value; Taking the sum of the corrected first similarity value and the corrected second similarity value as the interest similarity value between the second publisher and the first publisher.

6. The content distribution method according to claim 1, characterized in that, The method further includes: Obtaining the historical interaction information of the second publisher and the historical interaction information of the first publisher; Calculating a third similarity value between the historical interaction information of the second publisher and the historical interaction information of the first publisher based on a collaborative filtering algorithm; Determining the interest similarity value between the second publisher and the first publisher based on the first similarity value and the second similarity value includes: Performing a weighted sum on the first similarity value, the second similarity value, and the third similarity value to obtain the interest similarity value between the second publisher and the first publisher.

7. A content distribution system, characterized in that, Includes: A circumscribing unit, configured to, for the content to be distributed of the first publisher, obtain all intermediate users who have performed user behaviors on the content to be distributed; determine all second publishers followed by each of the intermediate users; wherein, the second publisher refers to other publishers who have published content and are other than the first publisher; The content distribution system further includes: a calculation unit, a determination unit, and a recommendation unit that are each executed for each intermediate user, wherein: The calculation unit is configured to, for each of the second publishers, determine a first similarity value between the second publisher and the first publisher according to the follow-up relationship network of the second publisher and the follow-up relationship network of the first publisher; determine a second similarity value between the second publisher and the first publisher according to the attribute information of the second publisher and the attribute information of the first publisher; and Based on the first similarity value and the second similarity value, determine the interest similarity value between the second publisher and the first publisher; The determination unit is configured to determine the second publishers with the interest similarity value greater than a preset similarity threshold as target publishers, and use the users corresponding to the follow-up relationship network of the target publishers as users to be recommended; The recommendation unit is configured to distribute the content to be distributed to each of the users to be recommended.

8. The content distribution system according to claim 7, wherein Further includes: A directed graph construction unit is configured to use each target publisher as a clustering center node of the directed graph, use each user included in the attention relationship network of each target publisher as an edge node of the directed graph, and use the attention relationship of each edge node to the clustering center node as a directed edge of the directed graph, and form the directed graph based on all the clustering center nodes, the edge nodes, and the directed edges; wherein, the target publishers include the first publisher and the second publisher, and the attention relationship network of the target publisher is composed of users who follow the target publisher. A cluster construction unit is configured to, within the directed graph, use the edge nodes of the directed edges pointing to the same clustering center node as the cluster of the clustering center node, and construct a first cluster with the first publisher as the clustering center node and a second cluster with the second publisher as the clustering center node; wherein, the directed edge points from the edge node to the clustering center node.

9. The content distribution system according to claim 7, wherein The calculation unit includes: A first calculation subunit is configured to, for each second publisher, calculate the overlap degree between the edge nodes of the second cluster corresponding to the second publisher and the edge nodes of the first cluster corresponding to the first publisher, to obtain the clustering cluster overlap degree between the second publisher and the first publisher; and perform normalization processing on the clustering cluster overlap degree between the second publisher and the first publisher through the mean and standard deviation of the clustering cluster overlap degrees between all the second publishers followed by the intermediate user corresponding to the second publisher and the first publisher, to obtain a first similarity value between the second publisher and the first publisher.

10. The content distribution system according to claim 7, wherein The calculation unit includes: A second calculation subunit is configured to, for each second publisher, calculate the similarity between the attribute information of the second publisher and the attribute information of the first publisher by using a user attribute calculation model, to obtain a second similarity value between the second publisher and the first publisher; wherein, the user attribute calculation model is obtained through supervised training based on a deep neural network.

11. The content distribution system according to claim 7, wherein The calculation unit includes: An interest similarity value calculation subunit is configured to set corresponding weight factors for the first similarity value and the second similarity value respectively; correct the first similarity value and the second similarity value through their respective weight factors to obtain a corrected first similarity value and a corrected second similarity value; and use the sum of the corrected first similarity value and the corrected second similarity value as the interest similarity value between the second publisher and the first publisher.

12. The content distribution system according to claim 7, wherein The calculation unit is further configured to obtain the historical interaction information of the second publisher and the historical interaction information of the first publisher; and calculate a third similarity value between the historical interaction information of the second publisher and the historical interaction information of the first publisher based on a collaborative filtering algorithm. The calculation unit includes: An interest similarity calculation subunit is configured to perform weighted summation on the first similarity value, the second similarity value, and the third similarity value to obtain the interest similarity value between the second publisher and the first publisher.

Citation Information

Patent Citations

  • Label recommendation method based on user attention relationship

    CN110674417A

  • Comparison method and device of publisher, equipment, storage medium and program product

    CN114756709A