A sports information recommendation method based on big data
Through the big data-based sports information recommendation method, user portraits are generated and multiple models are recommended, the problem that existing sports information recommendation methods cannot meet personalized needs is solved, personalized, real-time and accurate recommendation effects are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202111518325.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-13
AI Technical Summary
The existing sports information recommendation methods cannot meet users' personalized and accurate scenarios, resulting in users spending a lot of time in filtering content of interest, affecting the user experience.
The sports information recommendation method based on big data is adopted to collect information about user browsing and viewing behaviors, generate user portraits, and use catboost model, information popularity model and embedding model for accurate recommendations.
It realizes personalized, real-time and accurate sports information recommendations, reduces the time for users to screen content they are interested in, and improves the user's reading experience and stickiness.
Smart Images

Figure CN114186130B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data, and in particular to a sports information recommendation method based on big data. Background Art
[0002] At present, the Internet is very popular, and the way people obtain sports information has undergone tremendous changes, from print media to new media on the Internet. More and more users have regarded the Internet as their first choice for obtaining sports information, and the usage rate of related applications and websites has continued to grow.
[0003] At the same time, the huge carrying capacity of the Internet has also resulted in an unprecedented amount of information that users can receive every day on these platforms. Combined with the current vigorous development of the sports culture industry, more and more sports events, athletes, and sports teams have attracted people's attention, and the issue of how to achieve personalized recommendations for sports information has become increasingly prominent. Faced with so many different types of information content, users will spend a lot of time every day looking for content that they are interested in, which will directly affect the user's stickiness to the sports information platform, so it is also very necessary for the platform to solve this demand.
[0004] At present, the sports events and sports news information content of many platforms are mainly pushed based on the popularity of events and the interest tags manually selected by users. The information distribution system mostly recommends relevant sports events or related news information content to users based on this as the basis of user preferences. However, this recommendation method is relatively rough and not intelligent enough, and can no longer meet the increasingly diverse needs of users for personalized and accurate sports information. It cannot completely solve the personalization problem and affects the user experience.
[0005] Many sports information apps have also proposed many solutions for distributing sports information:
[0006] One is to use the latest or most popular news as the sorting basis. After users browse the news, it will be recorded by the system. If many users browse, it will be marked as popular news by the system and recommended to the information recommendation list. Users can see the hot spots and the latest sports news when they open the list. However, the disadvantage of sorting and recommending by the latest or most popular method is that it emphasizes the popularity of the information content too much, and cannot take care of the personalized reading needs of users.
[0007] The second is to search for keywords of interest through search engines. Through users' active search for keywords such as teams, athletes, sports events, etc., relevant sports information content can be obtained. The search method is more suitable for users to actively obtain news of interest. The disadvantage is that users need to actively express their interests. It is too dependent on users' active operations and cannot intelligently and actively recommend content of interest to users.
[0008] The third is to classify news through news category modules, which separates different categories of news and allows users to filter and read according to the categories they are interested in. Classifying news through news category modules can achieve a certain degree of personalized classification, but it is still relatively rough and requires users to take the initiative. It cannot directly help users discover known or unknown potential content of interest.
[0009] Fourth, through personalized recommendation algorithms, such as based on different demographic data such as age, region, occupation, gender, etc. According to the user's reading situation, collaborative recommendations are made with users with similar characteristics, so that users can see news content that they may also be interested in. Although general personalized recommendation algorithms have been applied in many fields such as e-commerce, news, and social networking, such recommendation systems are mostly aimed at comprehensive news and information content. General recommendation methods also face poor results when combined with sports vertical information platforms. For example, the unique sports knowledge of players, teams, and various sports contained in sports information cannot be specialized in the sports vertical field, and the user tags in the sports vertical field carried by such users cannot be very detailed. Summary of the invention
[0010] The purpose of the present invention is to provide a sports information recommendation method based on big data, so as to provide personalized, real-time and accurate recommendations for sports users' usage scenarios. To achieve the above purpose, the present invention adopts the following technical solutions:
[0011] The present invention discloses a sports information recommendation method based on big data, which includes the following process:
[0012] a. Collect information and obtain relevant information about users browsing and watching news or videos.
[0013] b. Store information and integrate the relevant information obtained in process a into the data warehouse.
[0014] c. Generate user portraits
[0015] c1. Big data analysis: Accurately analyze the user click and read behavior logs in the cloud, extract user information and the content of the information the user views, associate the corresponding information tags, and calculate the tag score of each user's click and the user's overall click tag score based on the information content viewed within a certain period of time. c2. Create user portrait tags: Associate user information, and obtain each user's portrait tag based on the decay sorting of the user's overall click tag score, thereby generating a user portrait.
[0016] d. Build training model
[0017] d1. Calculate and generate label word vectors and models: Normalize user portrait labels, normalize each portrait label to between 0 and 1, convert it into the corresponding label word vector, summarize the labels of the information clicked by the user within a period of time, and train the corresponding embedding model.
[0018] d2. Generate information heat model: Generate information heat model by calculating the information exposure click rate, information comment rate and manually set information weight in real time based on big data; calculate information heat based on the information heat model.
[0019] d3. Generate a catboost model: Perform offline training based on the popularity of information, the label word vector of the information, and the embedding model to generate a catboost model.
[0020] e. Information recommendation
[0021] Based on the obtained catboost model, the information that users like in the information candidate pool is predicted and accurately recommended to users.
[0022] Among them, in process a, relevant information about the user browsing and watching news or videos is obtained through user authorization, and the relevant information includes: device identification, news ID, video ID, time, number of page entries, stay time, tags, content category, and playback percentage.
[0023] Preferably, in process b, data is distributed to a collection service cluster based on SLB; the collection service cluster publishes the data to Kafka, and then subscribes to the Kafka data to the data warehouse in real time through data integration.
[0024] Preferably, in process c1, the method for calculating the score of the tag clicked by each user according to the viewed information content is:
[0025] It is calculated based on the frequency decay of information content labels, that is, the total number of labels in the information corresponding to the user's viewing time is t, among which the number of times the label of content A appears is s, then the label score of content A is x=s / t.
[0026] Furthermore, the method for calculating the score of the tag clicked by each user according to the viewed information content in process c1 further includes: calculating according to the decay of the viewing time of the information, that is, defining the initial time as T 0 , the initial time T 0 The score of the label of the information viewed by the user is F 0 , define a time period, the new time T n With the initial time T 0 There are n time periods in the time T n The score of the label of the information viewed by the user F n =(1+n)F 0 .
[0027] Furthermore, the user information in process c includes: the sports of interest selected by the user, the teams followed by the user, and some of the tags blacklisted by the user.
[0028] Preferably, in process d2, important news or videos are assigned information weights. Information popularity R is calculated using the following formula: R=X+Y+Z, where X is the exposure click rate of the information, Y is the information comment rate, and Z is the information weight rate;
[0029]
[0030] Among them, N 1 is the number of times a user browses the information page from the recommendation list, N is the number of times the information is displayed in the recommendation list; n is the number of comments on the information, Q is the manually set information weight, T is the time interval from publication to the present, and G is the time decay factor.
[0031] Preferably, in process e, the catboost model is used to construct a tree model, and the information is recommended to the user after being sorted based on its popularity.
[0032] Due to the adoption of the above method, the present invention has the following beneficial effects:
[0033] 1. The present invention can rely on the big data of a large number of sports users in the platform. According to the users' usage habits in sports information, the big data can collect and analyze the users' behavioral preferences for reading sports information, and form a characteristic portrait of sports users through a large number of general and sports-specific behavioral tags. Through the collaborative catboost model filtering recommendation and embedding model tag recommendation, the user's preference for different information content is calculated and predicted, and combined with the information heat model, personalized reading content that the user is more likely to like is pushed to the user from the information candidate pool. Ultimately, the problem of not being able to meet the personalized needs of users in the process of distributing sports information content is solved. By meeting the personalized sports information needs of different users, the time spent by users to screen and filter the content of interest is reduced. At the same time, the user's reading experience is also improved, which can increase the length and amount of sports information read by users.
[0034] 2. In the present invention, user tags are converted into corresponding word vectors to facilitate subsequent calculation of the similarity of new information and achieve accurate prediction and recommendation.
[0035] 3. Based on big data analysis, the present invention realizes the priority recommendation of instant and popular information, and then intelligently and accurately predicts the information that users are interested in according to their preferences, thereby realizing real-time and accurate recommendations in the vertical field of sports information.
[0036] 4. The present invention uses big data to calculate the exposure click rate of information in real time, and combines big data to calculate the news comment rate in real time, so as to push the hottest information to users in real time and in priority. Since some newly released news is more important, it has no popularity at the beginning and cannot be recommended in priority. To solve this problem, the present invention uses manual setting of corresponding weights. High-weighted news will be added with corresponding popularity, and the final information popularity will be calculated by combining the exposure click rate and the information comment rate, so as to achieve intelligent and accurate real-time recommendation.
[0037] 5. The present invention adopts the highly practical catboost algorithm. Catboost can directly process category features, improve prediction accuracy, have strong generalization and low time complexity, and can take into account both classification and regression problems. However, to achieve accurate recommendations, a single category feature is far from enough. The present invention uses the feature combination function of catboost to add user click label feature word vectors, user overall click label embedding models and other combined features. Catboost will also combine the original category features according to the intrinsic connection of the features in modeling, thereby enriching the feature dimension.
[0038] 6. Utilizing the catboost sorting function, when processing categorical variables and constructing tree models, information samples are processed and calculated based on the popularity sorting of information to obtain unbiased estimates of target variable statistics and model gradient values, effectively avoiding prediction bias and achieving news information to the greatest extent possible, thereby improving the real-time and accuracy of information prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic flow diagram of the recommended method of the present invention. DETAILED DESCRIPTION
[0040] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] The present invention discloses a sports information recommendation method based on big data, which includes the following process:
[0042] a. Collect information
[0043] Obtain relevant information about the user's browsing of news or videos through user authorization. The relevant information includes: device identification, news ID, video ID, time, number of page entries, length of stay, tags, content category, and playback percentage.
[0044] b. Storage Information
[0045] Integrate the relevant information obtained in process a into the data warehouse, specifically: distribute data to the collection service cluster based on SLB; the collection service cluster publishes data to Kafka, and then subscribes to Kafka data in real time to the data warehouse through data integration. SLB is a service load balancing, which distributes and dispatches traffic evenly across multiple cloud servers to eliminate single points of failure and improve the reliability and throughput of the application system. Kafka is a distributed publish / subscribe messaging system with high throughput, low latency, and scalable performance advantages.
[0046] c. Generate user portraits
[0047] c1. Big data analysis: Figure 1 As shown, the user's click and read behavior log is accurately analyzed in the cloud, user information and the content of the information viewed by the user are extracted, and the corresponding information tags are associated. The score of each user's clicked tag and the user's overall click tag score are calculated based on the information content viewed within a certain period of time. User information includes: the user's selected sports of interest, the team the user follows, and the tags the user has blocked.
[0048] The calculation of each user's click score on a tag can be done by using the decay of the tag frequency and the decay of the information viewing time. An example is given below.
[0049] (1) Tag frequency decay
[0050] The total number of tags in the information that a user watches for a certain period of time is t, and the number of times the tag of content A appears is s. Then the tag score of content A is x=s / t.
[0051] For example, if a user reads information a, information b, and information c in a certain day, the tags of information a are: "Lakers", "James". The tags of information b are: "Lakers", "Celtics", "James". The tags of information c are: "James".
[0052] The total number of labels is 6, that is, t=6, and the number of times the "James" label appears is 3 times, that is, s=3, so the "James" label score x=s / t=3 / 6=0.5.
[0053] Similarly, the "Lakers" label score is x=2 / 6=0.33; the "Celtics" label score is x=1 / 6=0.17.
[0054] (2) Information viewing time decay
[0055] Define the initial time as T 0 , the initial time T 0 The score of the label of the information viewed by the user is F 0 , define a time period, the new time T n With the initial time T 0 There are n time periods in the time T n The score of the label of the information viewed by the user F n =(1+n)F 0 . .
[0056] If the initial date is November 1, 2021 (T 0 ), defines the fraction F of labels of information viewed by the user within the initial time 0 =0.1. The user viewed information a on November 1, 2021. Information a has the tag: "Curry". Set the time period to 1 day. On November 2, 2021 (T 1 , n=1) when the user reads information b, and the label of information b is: "James". Then the score of the label "Curry" y= F 0 = 0.1, "James" label score y = F 1 =(1+1)×0.1=0.2.
[0057] Similarly, if the user reads news c on November 6, 2021, and the tag of news c is "Lakers", then n=5, then the score of the "Lakers" tag y=F 5 =(1+5)×0.1=0.6.
[0058] The calculation of the user's overall click tag score is to add up all the scores obtained by each user for the same tag to obtain the total score of the tag.
[0059] For example, the overall user click score for the "James" tag calculated through the above-mentioned tag frequency decay and information viewing time decay is Z=0.5+0.2=0.7.
[0060] c2. Create user portrait tags: associate user information, and get each user's portrait tag according to the decay sorting of the user's overall click tag score, thereby generating a user portrait. User information includes: the user's selected sports of interest, the team the user follows, and several tags that the user has blacklisted.
[0061] For example, the user's chosen sports interest is basketball, and the teams he follows are the Lakers, Warriors, and Nets. The user has previously read information tags, and the overall click score decay of the user's tags is Nets, Lakers, and Guangdong. Then the user's tag profile is Nets, Lakers, Warriors, and Guangdong.
[0062] d. Build training model
[0063] d1. Calculate and generate label word vectors and models
[0064] Normalize the user portrait labels, normalize each portrait label to between 0 and 1, convert them into corresponding label word vectors, summarize the labels of the information clicked by the user within a period of time, and train to obtain the corresponding embedding model.
[0065] The labels obtained in process c are Chinese characters and cannot be directly used in calculations and model training. Therefore, they need to be converted into corresponding word vectors to facilitate the subsequent calculation of the similarity of new information and achieve accurate prediction and recommendation.
[0066] d2. Generate information heat model: Generate information heat model by calculating the exposure click rate, information comment rate and manually set information weight of information in real time based on big data. Calculate information heat based on information heat model.
[0067] The information popularity R is calculated using the following formula: R=X+Y+Z, where X is the exposure click rate of the information, Y is the information comment rate, and Z is the information weight rate;
[0068]
[0069] Among them, N 1is the number of times a user browses the information page from the recommendation list, N is the number of times the information is displayed in the recommendation list; n is the number of comments on the information, Q is the manually set information weight, T is the time interval from publication to the present, and G is the time decay factor.
[0070] In practice, the Q value can be set according to the importance of the information. The more important the information (such as important news or videos), the greater the information weight value is set. The G value can be obtained through ab testing. By testing the recommendation effect under different values (such as which information has better reading time, click-through rate, and interaction rate), a more reasonable G value can be obtained. The G value is used to artificially control the impact of time on the ranking. The value of G determines how fast the ranking decreases over time.
[0071] Sports information has a relatively high real-time requirement, so it is necessary to give priority to pushing instant and popular news to users. The present invention uses big data to calculate the exposure click rate of information in real time, and combines big data to calculate the information comment rate in real time, so as to give priority to pushing popular information to users. Since some newly released news is more important and has no popularity at the beginning, intelligent recommendation cannot give priority to it. To solve this problem, the present invention uses important information to manually set the corresponding information weight. The high-weighted information will be added with the corresponding popularity, and the final information popularity will be calculated in combination with the exposure click rate and the information comment rate, so as to achieve intelligent and accurate real-time recommendation.
[0072] d3. Generate a catboost model: Perform offline training based on the popularity of information, the label word vector of the information, and the embedding model to generate a catboost model.
[0073] The catboost model uses the catboost algorithm and relies on the fast and accurate prediction of the catboost algorithm. Catboost uses the oblivious tree as the basic predictor. This tree is balanced and not prone to overfitting. In the oblivious tree, the index of each leaf node can be encoded as a binary vector with a length equal to the depth of the tree. Catboost first binarizes all floating-point features, statistical information, and one-hot encoded features, and then uses binary features to calculate the model prediction value. This method increases reliability and can greatly accelerate predictions. In terms of performance, it can match advanced machine learning algorithms and provide reliable support for the stability of the system.
[0074] e. Information recommendation
[0075] The obtained catboost model is used to predict the information that users like in the information candidate pool, and the catboost model is used to build a tree model. Based on the popularity of the information, the information samples are processed and calculated and then recommended to the users.
[0076] In summary, the present invention integrates big data behavioral data resources, and forms accurate and vertical user portraits in the sports field based on a rich sports user and content tag system. The personalized sports information content that users are more interested in is accurately and real-time recommended to users through a personalized recommendation system. This can better solve the problem of personalized information recommendation in the sports vertical information reading scenario.
[0077] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A sports information recommendation method based on big data, characterized in that: The process includes: a. Information collected: Obtain information about news or videos that users browse and watch; b. Storage Information: Integrate the relevant information obtained in process a into the data warehouse; c. Generate user portraits: c1. Big data analysis: Accurately analyze the user click and read behavior logs in the cloud, extract user information and the content of the information viewed by the user, associate the corresponding information tags, and calculate the tag score clicked by each user and the user's overall tag click score based on the information content viewed within a certain period of time; the user's overall tag click score is calculated by adding up all the scores obtained by each user for the same tag to obtain the total score of the tag; c2. Create user portrait tags: associate user information, and get each user's portrait tag according to the decay sorting of the user's overall click tag score, thereby generating a user portrait; d. Build a training model: d1. Calculate and generate label word vectors and models: Normalize user portrait labels, normalize each portrait label to between 0 and 1, convert it into the corresponding label word vector, summarize the labels of the information clicked by the user within a period of time, and train the corresponding embedding model; d2. Generate information heat model: Generate information heat model by calculating the information exposure click rate, information comment rate and manually set information weight in real time based on big data; calculate information heat based on the information heat model; d3. Generate CatBoost model: Perform offline training based on information popularity, information label word vectors, and embedding models to generate a CatBoost model; e. Information recommendation: Based on the obtained catboost model, the information that users like in the information candidate pool is predicted and accurately recommended to users.
2. The sports information recommendation method based on big data as claimed in claim 1, characterized in that: In process a, relevant information about the user browsing and watching news or videos is obtained through user authorization, and the relevant information includes: device identification, news ID, video ID, time, number of page entries, stay time, tags, content category, and playback percentage.
3. The sports information recommendation method based on big data as claimed in claim 1, characterized in that: In process b, data is distributed to the collection service cluster based on SLB; the collection service cluster publishes the data to Kafka, and then subscribes to the Kafka data to the data warehouse in real time through data integration.
4. The sports information recommendation method based on big data according to any one of claims 1 to 3, characterized in that: In process c1, the method for calculating the score of the tag clicked by each user based on the viewed information content is: It is calculated based on the frequency decay of information content labels, that is, the total number of labels in the information corresponding to the user's viewing time is t, among which the number of times the label of content A appears is s, then the label score of content A is x=s / t.
5. The sports information recommendation method based on big data as claimed in claim 4, characterized in that: The method for calculating the score of the tag clicked by each user according to the viewed information content in process c1 also includes: The calculation is based on the decay of the viewing time of the information, that is, the initial time is defined as T0, the score of the label of the information viewed by the user within the initial time T0 is F0, and a time period is defined, the new time T n There are n time periods within the initial time T0, then time T n The score of the label of the information viewed by the user F n =(1+n)F0.
6. The sports information recommendation method based on big data as claimed in claim 1, characterized in that: The user information in process c includes: the sports of interest selected by the user, the teams followed by the user, and several tags blacklisted by the user.
7. The sports information recommendation method based on big data as claimed in claim 1, characterized in that: In process d2, the information heat R is calculated using the following formula: R=X+Y+Z Among them, X is the exposure click rate of the information, Y is the information comment rate, and Z is the information weight rate; Among them, N1 is the number of times the user browses the information page from the recommendation list, N is the number of times the information is displayed in the recommendation list; n is the number of comments on the information, Q is the manually set information weight, T is the time interval from publication to the present, and G is the time decay factor.
8. The sports information recommendation method based on big data as claimed in claim 1, characterized in that: In process e, the catboost model is used to build a tree model, and the information is recommended to users after being sorted based on its popularity.
Citation Information
Patent Citations
Social media popularity prediction method and device based on visual semantic relationship
CN113657116A
User label updating method and device based on artificial intelligence, equipment and medium
CN113704620A