A hot search information pushing method based on big data
By classifying and evaluating trending search information and combining it with user behavior data, a personalized push strategy is adopted to solve the problems of interest model bias and content homogenization in existing trending search push technologies. This achieves accurate and diversified trending search information push, improving user experience and system intelligence.
Patent Information
- Application Number
- CN202510291192.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Existing trending search push methods suffer from biased interest models, homogenized content, insufficient user engagement, poor integration of information entropy and user interests, and a lack of personalized strategies. This results in an imbalance between the relevance and diversity of the pushed content, and a lack of differentiation between strategies for new and old users.
Trending topics are manually categorized into politics, economics, technology, entertainment, and sports. User interests are assessed by considering browsing time, interaction depth, and time intervals. Interest weights are calculated, and different push strategies are employed to target both new and returning users. Information entropy is used to control content diversity.
It enables precise delivery of trending information that meets user needs, improves the relevance and diversity of push notifications, avoids information overload, enhances user stickiness, adapts to dynamic changes in user interests, and prevents information cocoons.
Smart Images

Figure CN120196811B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of hot search pushing, in particular to a hot search information pushing method based on big data. BACKGROUND
[0002] With the explosive growth of Internet information, hot search pushing system has become an important way for users to obtain hot information. Traditional pushing methods mainly rely on user historical behavior (such as clicks, browsing time) to model interest, but there are the following problems:
[0003] Invalid browsing data is not effectively filtered, resulting in interest model bias. For example, a user's accidental click but not in-depth reading of the content may be misjudged as an interest point. Over-reliance on historical behavior leads to homogenization of pushed content, and users are exposed to single-domain information for a long time, limiting their vision and exacerbating group polarization. The existing technology does not introduce a time decay factor, which cannot reflect the dynamic changes of user interest. For example, a user's recent interest in technology-related content may be much higher than their historical preference for entertainment. When there is not enough behavior data, the pushing strategy is simplified, usually using a popular list to fill in, which cannot meet the personalized needs. In the existing technology, information entropy is mainly used to evaluate the global content diversity, but it is not combined with user interest. The existing technology uses information entropy to filter high-diversity content, but does not consider user personalized needs when pushing, resulting in an imbalance between content relevance and diversity. In addition, new and old user strategies are not differentiated: new user pushing lacks diversity guidance, and old user pushing may be trapped in an information cocoon due to excessive personalization.
[0004] In summary, there is an urgent need for a hot search information pushing method based on big data to improve the accuracy and diversity of pushing and enhance user stickiness. SUMMARY
[0005] The present application provides a hot search information pushing method based on big data, which solves the problems mentioned in the background.
[0006] The present application provides the following technical solution: a hot search information pushing method based on big data, specifically including:
[0007] Classify hot search information by artificial, divide all hot search information into different types;
[0008] The types include: politics, economy, technology, entertainment and sports;
[0009] Get the user to push the hot search information, recorded as the target user;
[0010] Get all the hot search information browsed by the target user in history, forming a first hot search set;
[0011] According to the data cleaning strategy, it is judged whether the target user's browsing of the hot search information in the first hot search set is valid, specifically including:
[0012] An arbitrary hot search information is selected from the first hot search set, denoted as a target hot search;
[0013] The number of times the target user browses the target hot search is obtained, denoted as the browsing frequency;
[0014] If the browsing frequency = 1, the duration of the target user browsing the target hot search is obtained, denoted as the browsing duration;
[0015] A first duration threshold and a second duration threshold are set;
[0016] The first duration threshold and the second duration threshold are used to judge whether the target user's browsing of the target hot search is valid; wherein the second duration threshold is greater than the first duration threshold;
[0017] If the browsing duration ≤ the first duration threshold;
[0018] It is determined that the target user's browsing of the target hot search is invalid, and the target hot search is removed from the first hot search set;
[0019] If the first duration threshold < the browsing duration < the second duration threshold;
[0020] It is determined that the target user's browsing of the target hot search is valid, and the target hot search is retained in the first hot search set;
[0021] If the browsing duration ≥ the second duration threshold, the interaction depth of the target user on the target hot search is obtained;
[0022] The interaction depth specifically includes 0 and 1;
[0023] If the target user has the behaviors of liking, commenting or sharing during the browsing process of the target hot search, it is determined that the interaction depth of the target user on the target hot search is 1;
[0024] If the target user does not have the behaviors of liking, commenting or sharing during the browsing process of the target hot search, it is determined that the interaction depth of the target user on the target hot search is 0;
[0025] If the interaction depth of the target user on the target hot search is 1, it is determined that the target user's browsing of the target hot search is valid, and the target hot search is retained in the first hot search set;
[0026] If the interaction depth of the target user on the target hot search is 0, it is determined that the target user's browsing of the target hot search is invalid, and the target hot search is removed from the first hot search set.
[0027] Optionally, the method further comprises:
[0028] If the number of times of browsing ≠ 1;
[0029] The time interval between the end time of the first time of browsing the target hot search and the start time of the second time of browsing the target hot search is obtained, denoted as a first time interval, the time interval between the end time of the second time of browsing the target hot search and the start time of the third time of browsing the target hot search is obtained, denoted as a second time interval, and the time interval between the end time of the last but one time of browsing the target hot search and the start time of the last time of browsing the target hot search is obtained, denoted as an A time interval.
[0030] Wherein A = the number of times of browsing - 1;
[0031] The average time interval is calculated as (the first time interval + the second time interval + … + the A time interval) ÷ A.
[0032] A time interval threshold is set, which is used to determine whether the browsing of the target hot search by the target user is effective.
[0033] If the average time interval ≥ the time interval threshold, it is determined that the browsing of the target hot search by the target user is effective, and the target hot search is retained in the first hot search set.
[0034] If the average time interval < the time interval threshold, it is determined that the browsing of the target hot search by the target user is ineffective, and the target hot search is excluded from the first hot search set.
[0035] The above method is used to determine whether the browsing of all hot search information in the first hot search set is effective.
[0036] All hot search information retained in the first hot search set is obtained to form a second hot search set.
[0037] Optionally, the method further comprises:
[0038] The interest weight of the target user for each type of hot search information is calculated, specifically including:
[0039] The hot search information of the political type in the second hot search set forms a political type set, the hot search information of the economic type in the second hot search set forms an economic type set, the hot search information of the science and technology type in the second hot search set forms a science and technology type set, the hot search information of the entertainment type in the second hot search set forms an entertainment type set, and the hot search information of the sports type in the second hot search set forms a sports type set.
[0040] Respectively acquire the time when the target user browses the hot search in each different type set, and calculate the time difference with the current time;
[0041] Set a time decay factor a, which is used to weight the historical browsing data of the target user in the process of calculating the interest weight of the target user for each type hot search information, so as to achieve the effect of paying more attention to recent behavior;
[0042] The interest weight of the target user for each type hot search information is calculated by the following formula:
[0043]
[0044] Wherein:
[0045] w i is the interest weight of the target user for the i-th type hot search information; i=1 represents the interest weight of the political type hot search information, i=2 represents the interest weight of the economic type hot search information, i=3 represents the interest weight of the science and technology type hot search information, i=4 represents the interest weight of the entertainment type hot search information, and i=5 represents the interest weight of the sports type hot search information;
[0046] t ij is the time difference between the j-th browsing behavior of the target user for the i-th type hot search information and the current time;
[0047] a is a time decay factor;
[0048] b ij is the contribution value of the j-th browsing behavior of the target user to the i-th type hot search information; if the target user only browses, the contribution value is 1, if the target user appears like behavior in the browsing process, the contribution value is 2, if the target user appears comment behavior in the browsing process, the contribution value is 3, if the target user appears sharing behavior in the browsing process, the contribution value is 4; if the user appears multiple behaviors in the browsing process, the contribution value corresponding to the behavior with the maximum contribution value is selected as the contribution value;
[0049] The interest weights w1, w2, w3, w4 and w5 are calculated by the above-mentioned manner.
[0050] Optionally, the calculation of the interest weight of the target user for each type hot search information specifically includes:
[0051] The target user interest weight is normalized to make the sum equal to 1, and the target user interest weight is normalized by the following formula:
[0052]
[0053] w1', w2', w3', w4' and w5' are calculated by the above method.
[0054] Optionally, the method further comprises:
[0055] A first quantity threshold is set, which is used to determine whether the target user is a new user or not;
[0056] The quantity of hot search information in the second hot search set is obtained, denoted as a judgment quantity;
[0057] If the judgment quantity is less than or equal to the first quantity threshold, the target user is determined to be a new user;
[0058] If the judgment quantity is greater than the first quantity threshold, the target user is determined to be an old user;
[0059] If the target user is a new user, the target user is pushed with hot search information according to a first strategy;
[0060] If the target user is an old user, the target user is pushed with hot search information according to a second strategy.
[0061] Optionally, if the target user is a new user, the target user is pushed with hot search information according to a first strategy, which specifically comprises:
[0062] All current hot search information is obtained to form a hot search information pool;
[0063] A push quantity, a first information entropy and a second information entropy are set;
[0064] The push quantity, the first information entropy and the second information entropy are used to select hot search information from the hot search information pool for pushing to the target user, wherein the first information entropy is less than the second information entropy;
[0065] A target push set is formed by randomly selecting the push quantity of hot search information from the hot search information pool;
[0066] The hot search information in the target push set is classified according to its type, and the proportion of each type of hot search information to all hot search information in the target push set is obtained, denoted as P(x1), P(x2), P(x3), P(x4) and P(x5) respectively;
[0067] The information entropy of the hot search information in the target push set is calculated by the following formula:
[0068]
[0069] If the first information entropy is less than H(ω) and the second information entropy is greater than H(ω), the hot search information in the target push set is pushed to the target user;
[0070] If H(ω)≤ the first information entropy or H(ω)≥ the second information entropy, then a push quantity of hot search information is randomly selected from the hot search information pool to form a target push set, until the first information entropy < the information entropy of the hot search information in the target push set < the second information entropy.
[0071] Optionally, if the target user is an old user, the target user is pushed with hot search information according to a second strategy, specifically including:
[0072] All current hot search information is acquired to form a hot search information pool;
[0073] A push quantity, a first information entropy and a second information entropy are set;
[0074] The push quantity, the first information entropy and the second information entropy are used to select hot search information from the hot search information pool to push the target user with hot search information, wherein the first information entropy < the second information entropy;
[0075] A push quantity of hot search information is randomly selected from the hot search information pool to form a target push set;
[0076] The hot search information in the target push set is classified according to its type, and the proportion of each type of hot search information to all hot search information in the target push set is acquired, and is respectively denoted as P(x1), P(x2), P(x3), P(x4) and P(x5);
[0077] The weighted information entropy of the hot search information in the target push set is calculated by the following formula:
[0078]
[0079] If the first information entropy < H(ω)' < the second information entropy, the hot search information in the target push set is pushed to the target user;
[0080] If H(ω)'≤ the first information entropy or H(ω)'≥ the second information entropy, then a push quantity of hot search information is randomly selected from the hot search information pool to form a target push set, until the first information entropy < the weighted information entropy of the hot search information in the target push set < the second information entropy.
[0081] The present application has the following advantages:
[0082] 1. The hot search information pushing method based on big data classifies hot search information by artificial means to ensure that different types of hot search information can be accurately matched to the needs of target users. For example, the classification of politics, economy, technology, entertainment, and sports can enable the system to push accurate content according to the interests of different users, improving the relevance of the push. After classification, the information push avoids information overload, and users can receive updated hot search information in their areas of interest, thereby improving the user experience. Personalized push after classification can provide more data support for subsequent algorithm optimization, making the hot search recommendation system more intelligent.
[0083] 2. The hot search information pushing method based on big data determines the attention of users to hot search information by judging the browsing time and interaction depth, which can clearly distinguish whether the user's interest in certain information is real. Invalid information is effectively removed, making the push content more in line with user needs. For hot search information with a browsing frequency of 1, by setting a first time threshold and a second time threshold and comparing them with the target user's browsing time, the interference of junk information can be effectively reduced. If the browsing time is ≤ the first time threshold, it means that the target user's browsing time for this hot search information is very short, which may be due to accidental touch, so the target user's browsing of this target hot search is deemed invalid. If the browsing time is ≥ the second time threshold, the target user's interaction depth with the target hot search is combined to determine whether the target user's browsing of the target hot search is effective. If the interaction depth is 0, it means that the user has been browsing this hot search information for a long time, but has not produced any interaction behavior, which may be due to accidental touch, so the target user's browsing of this target hot search is deemed invalid. In addition, for hot search information with a browsing frequency not equal to 1, the average time interval is obtained by setting a time interval threshold to determine whether the target user's browsing of the target hot search is effective. If the average time interval < the time interval threshold, it means that the target user has browsed the same hot search information multiple times in a short period of time, which is suspicious of brushing hot search, so the target user's browsing of this target hot search is deemed invalid. In this way, valid and invalid behaviors can be accurately distinguished, providing accurate data basis for subsequent hot search information pushing.
[0084] 3. The hot search information pushing method based on big data calculates the interest weight of target users for different types of hot search information, which can realize in-depth mining of user preferences and further push more personalized hot search information, improving the acceptance of users to the pushed content. The calculation of interest weight is not only based on the user's historical behavior, but also combined with a time decay factor, so that the recommended content is more in line with the current user's immediate interest, rather than relying only on historical data. Through the use of the time decay factor, the dynamic changes of user interest can be updated in a timely manner, especially when the user's interest changes, the system can quickly adjust the recommendation strategy.
[0085] 4、The hot search information pushing method based on big data, different pushing strategies are adopted by judging whether the target user is a new user or an old user; if the judgment quantity is less than or equal to the first quantity threshold, it indicates that the browsing record of the target user is less, at this time, the target user is determined to be a new user; if the judgment quantity is greater than the first quantity threshold, it indicates that the browsing record of the target user is more, then the target user is determined to be an old user; for a new user, a pushing quantity of hot search information is selected from the hot search information pool at random, and the information entropy of the selected hot search information is calculated; the greater the information entropy, the more extensive the type of the selected hot search information, the smaller the information entropy, the more narrow the type of the selected hot search information, and the information entropy of the selected hot search information is ensured to be between the first information entropy and the second information entropy, so that the hot search information pushed to the target user is neither too extensive nor too narrow, and the existence of the information cocoon house is avoided; for an old user, a pushing quantity of hot search information is selected from the hot search information pool at random, and the weighted information entropy of the selected hot search information is calculated in combination with the interest weight of the target user for each type of hot search information, so that the weighted information entropy of the selected hot search information is ensured to be between the first information entropy and the second information entropy, and the hot search information pushed to the target user can meet the interest of the target user and avoid the information cocoon house. BRIEF DESCRIPTION OF DRAWINGS
[0086] Figure 1 The flowchart of the present application. DETAILED DESCRIPTION
[0087] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0088] Embodiment, refer to Figure 1 A hot search information pushing method based on big data, specifically comprising:
[0089] Classify the hot search information by artificial, divide all hot search information into different types;
[0090] The types include: politics, economy, technology, entertainment and sports;
[0091] The hot search information pushing method based on big data classifies hot search information by artificial means to ensure that different types of hot search information can be accurately matched to the needs of target users. For example, the classification of politics, economy, technology, entertainment, and sports can enable the system to push accurate content according to the interests of different users, improving the relevance of the push. After classification, the information push avoids information overload, and users can receive updated hot search information in their areas of interest, thereby improving the user experience. Personalized pushing after classification can provide more data support for subsequent algorithm optimization, making the hot search recommendation system more intelligent;
[0092] Obtain a user to which hot search information is pushed, denoted as a target user;
[0093] Obtain all hot search information historically browsed by the target user to form a first hot search set;
[0094] Determine whether the target user's browsing of the hot search information in the first hot search set is valid according to a data cleaning strategy, specifically including:
[0095] Select any one hot search information from the first hot search set, denoted as a target hot search;
[0096] Obtain the number of times the target user browses the target hot search, denoted as a browsing frequency;
[0097] If the browsing frequency = 1, obtain the duration of the target user's browsing of the target hot search, denoted as a browsing duration;
[0098] Set a first duration threshold and a second duration threshold;
[0099] The first duration threshold and the second duration threshold are used to determine whether the target user's browsing of the target hot search is valid; wherein the second duration threshold is greater than the first duration threshold;
[0100] If the browsing duration ≤ the first duration threshold;
[0101] The target user's browsing of the target hot search is deemed invalid, and the target hot search is removed from the first hot search set;
[0102] If the first duration threshold < the browsing duration < the second duration threshold;
[0103] The target user's browsing of the target hot search is deemed valid, and the target hot search is retained in the first hot search set;
[0104] If the browsing duration ≥ the second duration threshold, obtain the interaction depth of the target user with the target hot search;
[0105] The interaction depth specifically includes 0 and 1;
[0106] If the target user likes, comments or shares the target hot search during the browsing process, it is determined that the interaction depth of the target user with the target hot search is 1;
[0107] If the target user does not like, comment or share the target hot search during the browsing process, it is determined that the interaction depth of the target user with the target hot search is 0;
[0108] If the interaction depth of the target user with the target hot search is 1, it is determined that the browsing of the target user with the target hot search is valid, and the target hot search is retained in the first hot search set;
[0109] If the interaction depth of the target user with the target hot search is 0, it is determined that the browsing of the target user with the target hot search is invalid, and the target hot search is excluded from the first hot search set;
[0110] The hot search information pushing method based on big data can determine the attention of the user to the hot search information by judging the browsing time and the interaction depth, and can clearly distinguish whether the interest of the user to certain information is real. Invalid information is effectively removed, so that the pushed content is more in line with the user demand. For the hot search information with a browsing frequency of 1, by setting the first time threshold and the second time threshold and comparing them with the browsing time of the target user, the interference of garbage information can be effectively reduced. If the browsing time is less than or equal to the first time threshold, it means that the target user's browsing time of the hot search information is very short, which may be a short browsing caused by a mistake. Therefore, it is determined that the target user's browsing of the target hot search is invalid. If the browsing time is greater than or equal to the second time threshold, the browsing of the target user with the target hot search is determined to be valid or invalid in combination with the interaction depth of the target user with the target hot search. If the interaction depth is 0, it means that the user has browsed the hot search information for a long time, but has not produced any interaction behavior, which may be a case of staying on the hot search information caused by a mistake. Therefore, it is determined that the target user's browsing of the target hot search is invalid.
[0111] The method further comprises the following steps:
[0112] If the browsing frequency is not equal to 1;
[0113] The time interval between the end time of the first browsing of the target hot search by the target user and the start time of the second browsing of the target hot search is obtained, which is recorded as the first time interval. The time interval between the end time of the second browsing of the target hot search by the target user and the start time of the third browsing of the target hot search is obtained, which is recorded as the second time interval. The time interval between the end time of the last but one browsing of the target hot search by the target user and the start time of the last browsing of the target hot search is obtained, which is recorded as the A time interval.
[0114] Wherein A = browsing frequency - 1;
[0115] Calculate (first time interval + second time interval + … + A time interval) ÷ A, and the result is recorded as the average time interval;
[0116] Set a time interval threshold value for determining whether the target user's browsing of the target hot search is valid;
[0117] If the average time interval ≥ the time interval threshold value, it is determined that the target user's browsing of the target hot search is valid, and the target hot search is retained in the first hot search set;
[0118] If the average time interval < the time interval threshold value, it is determined that the target user's browsing of the target hot search is invalid, and the target hot search is excluded from the first hot search set;
[0119] Determine whether the browsing of all hot search information in the first hot search set is valid by the above method;
[0120] Obtain all hot search information retained in the first hot search set to form a second hot search set;
[0121] In addition, for hot search information with a browsing frequency of 1, by setting a time interval threshold value and obtaining an average time interval, it is determined whether the target user's browsing of the target hot search is valid; if the average time interval < the time interval threshold value, it means that the target user has browsed the same hot search information multiple times in a short period of time, which is suspicious of brushing hot search, so it is determined that the target user's browsing of the target hot search is invalid; in this way, valid and invalid behaviors can be accurately distinguished, providing accurate data basis for subsequent hot search information pushing.
[0122] Further comprising:
[0123] Calculate the interest weight of the target user for each type of hot search information, specifically including:
[0124] The hot search information of the political type in the second hot search set constitutes a political type set, the hot search information of the economic type in the second hot search set constitutes an economic type set, the hot search information of the science and technology type in the second hot search set constitutes a science and technology type set, the hot search information of the entertainment type in the second hot search set constitutes an entertainment type set, and the hot search information of the sports type in the second hot search set constitutes a sports type set;
[0125] Obtain the time when the target user browses the hot search in each different type set, and calculate the time difference from the current time;
[0126] Set a time decay factor α, which is used to weight the target user's historical browsing data in the process of calculating the interest weight of the target user for each type of hot search information, so as to achieve the effect of paying more attention to recent behaviors;
[0127] The interest weight of the target user for each type of hot search information is calculated by the following formula:
[0128]
[0129] Wherein:
[0130] w i is the interest weight of the target user for the i-th type of hot search information; i=1 represents the interest weight of the political type hot search information, i=2 represents the interest weight of the economic type hot search information, i=3 represents the interest weight of the science and technology type hot search information, i=4 represents the interest weight of the entertainment type hot search information, and i=5 represents the interest weight of the sports type hot search information;
[0131] t ij is the time difference between the j-th browsing behavior of the target user for the i-th type of hot search information and the current time;
[0132] α is a time decay factor;
[0133] b ij is the contribution value of the j-th browsing behavior of the target user to the i-th type of hot search information; if the target user only browses, the contribution value is 1, if the target user appears a like behavior in the browsing process, the contribution value is 2, if the target user appears a comment behavior in the browsing process, the contribution value is 3, if the target user appears a sharing behavior in the browsing process, the contribution value is 4; if the user appears multiple behaviors in the browsing process, the contribution value corresponding to the behavior with the maximum contribution value is selected as the contribution value;
[0134] The interest weights q1, q2, w3, w4 and w5 are calculated in the above manner;
[0135] The hot search information pushing method based on big data can realize in-depth mining of user preferences by calculating the interest weights of the target user for different types of hot search information, and further push more personalized hot search information to improve the acceptance of the pushed content by the user. The calculation of the interest weight is not only based on the historical behavior of the user, but also combined with the time decay factor, so that the recommended content is more in line with the immediate interest of the current user, rather than relying only on historical data. Through the use of the time decay factor, the dynamic changes of user interest can be updated in time, especially when the user interest changes, the system can quickly adjust the recommendation strategy.
[0136] The calculation of the interest weight of the target user for each type of hot search information specifically includes:
[0137] The target user interest weight is normalized so that the sum is 1, and the target user interest weight is normalized through the following formula:
[0138]
[0139] The w1', w2', w3', w4' and w5' are calculated and obtained in the above manner.
[0140] Further comprising:
[0141] A first number threshold is set, which is used to determine whether the target user is a new user;
[0142] The number of hot search information in the second hot search set is obtained, denoted as the judgment number;
[0143] If the judgment number is less than or equal to the first number threshold, the target user is determined to be a new user;
[0144] If the judgment number is greater than the first number threshold, the target user is determined to be an old user;
[0145] If the target user is a new user, the target user is pushed with hot search information according to the first strategy;
[0146] If the target user is an old user, the target user is pushed with hot search information according to the second strategy;
[0147] The hot search information pushing method based on big data judges whether the target user is a new user or an old user, and adopts different pushing strategies; if the judgment number is less than or equal to the first number threshold, it means that the target user has less browsing records, and the target user is determined to be a new user; if the judgment number is greater than the first number threshold, it means that the target user has more browsing records, and the target user is determined to be an old user;
[0148] If the target user is a new user, the target user is pushed with hot search information according to the first strategy, specifically including:
[0149] All current hot search information is obtained to form a hot search information pool;
[0150] The push number, the first information entropy and the second information entropy are set;
[0151] The push number, the first information entropy and the second information entropy are used to select hot search information from the hot search information pool for the target user to push hot search information, and the first information entropy is less than the second information entropy;
[0152] Any push number of hot search information is selected from the hot search information pool to form a target push set;
[0153] Classify the hot search information in the target push set according to the types thereof, and obtain the proportion of each type of hot search information to all hot search information in the target push set, denoted as P(x1), P(x2), P(x3), P(x4) and P(x5) respectively;
[0154] The information entropy of the hot search information in the target push set is calculated by the following formula:
[0155]
[0156] If the first information entropy < H(ω) < the second information entropy, the hot search information in the target push set is pushed to the target user;
[0157] If H(ω)≤the first information entropy or H(ω)≥the second information entropy, a number of hot search information is randomly selected from the hot search information pool to form the target push set, until the first information entropy < the information entropy of the hot search information in the target push set < the second information entropy;
[0158] For a new user, a number of hot search information is randomly selected from the hot search information pool, and the information entropy of the selected hot search information is calculated. The greater the information entropy, the more extensive the types of the selected hot search information. The smaller the information entropy, the more narrow the types of the selected hot search information. Ensuring that the information entropy of the selected hot search information is between the first information entropy and the second information entropy can ensure that the hot search information pushed to the target user is neither too extensive nor too narrow, avoiding the existence of an information cocoon house;
[0159] If the target user is an old user, the target user is pushed hot search information according to a second strategy, specifically including:
[0160] Obtain all current hot search information to form a hot search information pool;
[0161] Set the push number, the first information entropy and the second information entropy;
[0162] The push number, the first information entropy and the second information entropy are used to select hot search information from the hot search information pool for pushing hot search information to the target user, wherein the first information entropy < the second information entropy;
[0163] Randomly select a number of hot search information from the hot search information pool to form a target push set;
[0164] Classify the hot search information in the target push set according to the types thereof, and obtain the proportion of each type of hot search information to all hot search information in the target push set, denoted as P(x1), P(x2), P(x3), P(x4) and P(x5) respectively;
[0165] The weighted information entropy of the hot search information in the target push set is calculated by the following formula:
[0166]
[0167] If the first information entropy < H(ω)' < the second information entropy, the hot search information in the target push set is pushed to the target user;
[0168] If H(ω)' ≤ the first information entropy or H(ω)' ≥ the second information entropy, the push quantity of hot search information is randomly selected from the hot search information pool to form a target push set, until the first information entropy < the weighted information entropy of the hot search information in the target push set < the second information entropy.
[0169] For old users, the push quantity of hot search information is randomly selected from the hot search information pool, and the weighted information entropy of the selected hot search information is calculated in combination with the interest weight of the target user for each type of hot search information, so as to ensure that the weighted information entropy of the selected hot search information is between the first information entropy and the second information entropy, which can ensure that the hot search information pushed to the target user meets the interest of the target user and avoids information cocooning.
[0170] It should be noted that in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment.
[0171] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the technical principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.
Claims
1. A method for pushing hot search information based on big data, characterized in that, Specifically comprising: Classifying hot search information by artificial, all hot search information is divided into different types; The type includes: politics, economy, technology, entertainment and sports; Get the user who pushes the hot search information, recorded as the target user; Get all the hot search information browsed by the target user in history, forming a first hot search set; According to the data cleaning strategy, judge whether the target user's browsing of the hot search information in the first hot search set is valid, specifically including: Select a hot search information from the first hot search set at random, recorded as the target hot search; Get the number of times the target user browses the target hot search, recorded as the browsing times; If the browsing times = 1, get the browsing time of the target user browsing the target hot search, recorded as the browsing time; Set the first time threshold and the second time threshold; The first time threshold and the second time threshold are used to judge whether the target user's browsing of the target hot search is valid; wherein the second time threshold is greater than the first time threshold; If the browsing time ≤ the first time threshold; It is determined that the target user's browsing of the target hot search is invalid, and the target hot search is excluded from the first hot search set; If the first time threshold < the browsing time < the second time threshold; It is determined that the target user's browsing of the target hot search is valid, and the target hot search is retained in the first hot search set; If the browsing time ≥ the second time threshold, get the interaction depth of the target user to the target hot search; The interaction depth specifically includes 0 and 1; If the target user appears the behavior of liking, commenting or sharing in the browsing process of the target hot search, it is determined that the interaction depth of the target user to the target hot search is 1; If the target user does not appear the behavior of liking, commenting or sharing in the browsing process of the target hot search, it is determined that the interaction depth of the target user to the target hot search is 0; If the interaction depth of the target user to the target hot search is 1, it is determined that the target user's browsing of the target hot search is valid, and the target hot search is retained in the first hot search set; If the interaction depth of the target user to the target hot search is 0, it is determined that the target user's browsing of the target hot search is invalid, and the target hot search is excluded from the first hot search set. 2.The hot search information pushing method based on big data according to claim 1, characterized in that: The data cleaning strategy is used to judge whether the target user's browsing of the hot search information in the first hot search set is valid, specifically including: If the browsing times ≠ 1; Get the time interval between the end time of the target user's first browsing of the target hot search and the start time of the second browsing of the target hot search, recorded as the first time interval, get the time interval between the end time of the target user's second browsing of the target hot search and the start time of the third browsing of the target hot search, recorded as the second time interval, …, get the time interval between the end time of the target user's last but one browsing of the target hot search and the start time of the last browsing of the target hot search, recorded as the A time interval; Wherein A = browsing times - 1; Calculate (first time interval + second time interval + … + A time interval) ÷ A, the result is recorded as the average time interval; Set the time interval threshold, the time interval threshold is used to judge whether the target user's browsing of the target hot search is valid; If the average time interval is greater than or equal to the time interval threshold, it is determined that the target user's browsing of the target hot search is valid, and the target hot search is retained in the first hot search set; If the average time interval is less than the time interval threshold, it is determined that the target user's browsing of the target hot search is invalid, and the target hot search is excluded from the first hot search set; The above method is used to determine whether the browsing of all hot search information in the first hot search set is valid; All hot search information retained in the first hot search set is obtained to form a second hot search set. 3.The hot search information pushing method based on big data according to claim 2, characterized in that, Further comprising: calculating the interest weight of the target user for each type of hot search information, specifically including: The hot search information of the political type in the second hot search set constitutes a political type set, the hot search information of the economic type in the second hot search set constitutes an economic type set, the hot search information of the science and technology type in the second hot search set constitutes a science and technology type set, the hot search information of the entertainment type in the second hot search set constitutes an entertainment type set, and the hot search information of the sports type in the second hot search set constitutes a sports type set; The time difference between the current time and the time when the target user browses the hot search in each different type set is calculated; A time decay factor a is set, which is used to weight the historical browsing data of the target user in the process of calculating the interest weight of the target user for each type of hot search information, so as to achieve the effect of paying more attention to recent behavior; The interest weight of the target user for each type of hot search information is calculated by the following formula: Wherein: w i is the interest weight of the target user to the i-th type of hot search information; i=1 represents the interest weight of the political type hot search information, i=2 represents the interest weight of the economic type hot search information, i=3 represents the interest weight of the science and technology type hot search information, i=4 represents the interest weight of the entertainment type hot search information, and i=5 represents the interest weight of the sports type hot search information. t ij is the time difference between the jth browsing behavior of the target user to the ith type of hot search information and the current time. a is the time decay factor; b ij is the contribution value of the jth browsing behavior of the target user to the ith type of hot search information; if the target user only browses, the contribution value is recorded as 1; if the target user likes in the browsing process, the contribution value is recorded as 2; if the target user comments in the browsing process, the contribution value is recorded as 3; if the target user shares in the browsing process, the contribution value is recorded as 4; if the user appears multiple behaviors in the browsing process, the contribution value corresponding to the behavior with the maximum contribution value is selected as the contribution value; The interest weights w1, w2, w3, w4 and w5 are calculated in the above manner.
4. The hot search information pushing method based on big data according to claim 3, characterized in that: The calculation of the interest weight of the target user for each type of hot search information specifically includes: The target user interest weight is normalized to make the sum equal to 1, and the target user interest weight is normalized by the following formula; The w1', w2', w3', w4' and w5' are calculated in the above manner.
5. The hot search information pushing method based on big data according to claim 4, characterized in that, Further comprising: A first quantity threshold is set, which is used to determine whether the target user is a new user; The number of hot search information in the second hot search set is obtained, denoted as judgment number; If the judgment number is less than or equal to the first quantity threshold, the target user is determined to be a new user; If the judgment number is greater than the first quantity threshold, the target user is determined to be an old user; If the target user is a new user, the target user is pushed hot search information according to the first strategy; If the target user is an old user, the target user is pushed hot search information according to the second strategy.
6. The hot search information pushing method based on big data according to claim 5, characterized in that: If the target user is a new user, the target user is pushed hot search information according to the first strategy, specifically including: All current hot search information is obtained to form a hot search information pool; A push quantity, a first information entropy and a second information entropy are set; The push quantity, the first information entropy and the second information entropy are used to select hot search information for the target user from the hot search information pool, and the first information entropy is less than the second information entropy; A target push set is formed by randomly selecting push quantity hot search information from the hot search information pool; The hot search information in the target push set is classified according to types, and the proportion of each type of hot search information to all hot search information in the target push set is obtained and denoted as P(x1), P(x2), P(x3), P(x4) and P(x5) respectively; The information entropy of the hot search information in the target push set is calculated by the following formula: If the first information entropy < H(ω) < the second information entropy, the hot search information in the target push set is pushed to the target user; If H(ω) ≤ the first information entropy or H(ω) ≥ the second information entropy, a push quantity of hot search information is randomly selected from the hot search information pool to form a target push set, until the first information entropy < the information entropy of the hot search information in the target push set < the second information entropy.
7. The hot search information pushing method based on big data according to claim 5, characterized in that: If the target user is an old user, the target user is pushed hot search information according to a second strategy, specifically including: All current hot search information is obtained to form a hot search information pool; A push quantity, a first information entropy and a second information entropy are set; The push quantity, the first information entropy and the second information entropy are used to select hot search information from the hot search information pool for pushing hot search information to the target user, wherein the first information entropy < the second information entropy; A push quantity of hot search information is randomly selected from the hot search information pool to form a target push set; The hot search information in the target push set is classified according to types, and the proportion of each type of hot search information to all hot search information in the target push set is obtained and denoted as P(x1), P(x2), P(x3), P(x4) and P(x5) respectively; The weighted information entropy of the hot search information in the target push set is calculated by the following formula: If the first information entropy < H(ω)' < the second information entropy, the hot search information in the target push set is pushed to the target user; If H(ω)' ≤ the first information entropy or H(ω)' ≥ the second information entropy, a push quantity of hot search information is randomly selected from the hot search information pool to form a target push set, until the first information entropy < the weighted information entropy of the hot search information in the target push set < the second information entropy.
Citation Information
Patent Citations
Knowledge management system capable of providing effective information
CN102081635A
Method and device for pushing data
CN102377790A