A popular science content personalized recommendation system based on big data and a method thereof
By obtaining user registration dates and browsing data from science popularization websites, and combining them with play count rankings and user attention values, high-quality science popularization content that matches user interests is filtered out. This solves the problems of sparse and low-accuracy recommendations in existing technologies, and improves the accuracy of personalized recommendations and user satisfaction.
Patent Information
- Application Number
- CN202510214145.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing personalized science popularization content recommendation systems suffer from data sparsity and homogeneous user behavior, making it difficult to provide comprehensive and diverse recommendations. This leads to a decline in the accuracy of recommendation results and an inability to accurately determine user interests and needs.
By obtaining users' registration dates and current dates on science popularization websites, calculating the number of views and time, key users are identified. Combined with the play count ranking and user attention value, the recommendation value of science popularization content is calculated, and high-quality content that matches user interests is selected.
It enables the adjustment of recommendation strategies based on user interests, accurately pushing high-quality science popularization content that matches user interests, thereby improving user experience and satisfaction.
Smart Images

Figure CN120104875B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of personalized recommendation technology, specifically to a personalized recommendation system and method for popular science content based on big data. Background Technology
[0002] With the rapid development of the internet, major websites have launched personalized science content services, aiming to recommend the most likely-interesting science content to users based on their past browsing history. However, while these personalized recommendation services have improved the user experience to some extent, the recommendation data they provide is often very sparse and limited.
[0003] Due to the limitations of user browsing history, recommender systems may only be able to analyze and make recommendations based on limited data. This means that if a user's past browsing behavior is relatively singular or concentrated in a specific area, the recommender system will struggle to provide comprehensive and diverse science and technology content recommendations. Furthermore, the sparsity of the recommendation data may also lead to a decrease in the accuracy of the recommendations. Because of the limited data volume, the recommender system may struggle to accurately determine the user's true interests and needs, thus providing recommendations that do not meet the user's expectations.
[0004] While personalized science content services have improved user experience to some extent in existing technologies, the content recommended by science websites is often quite limited. They can only characterize and understand a user from a few limited dimensions, making it difficult to accurately determine a user's interests and attributes. Such recommendations are often inaccurate. Furthermore, users' interests may change over time. Therefore, it is necessary to continuously adjust recommendation strategies and methods based on user interests to achieve precise targeting and provide users with a better experience. Summary of the Invention
[0005] The purpose of this invention is to provide a personalized recommendation system and method for popular science content based on big data, thereby solving the above-mentioned technical problems.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] A personalized recommendation method for popular science content based on big data includes the following steps:
[0008] S1: Obtain the user's registration date on the science popularization website. L and the current date D N Calculate the usage date D=D N -D L Get the user's registration date D L Up to the current date D NThe number of views N for popular science content, including popular science articles and videos, is calculated as follows: Daily views R = N / D. If daily views R > R... sta Users are designated as key users, among whom R sta The preset judgment value;
[0009] S2: Obtain the total time t spent by key users browsing popular science content, and calculate the average time spent t. ave =t / N, calculates the browsing intention value of key users for popular science content. Where ts represents the completion time of the science video, A1 and A2 are the preset first and second intention coefficients, A1 < 0, and B is the preset intention threshold.
[0010] Science popularization content belonging to the same section on a science popularization website is categorized as similar content. The section refers to the category directory in the website's navigation bar. The attention value of key users to similar content is calculated. Where λ represents the preset adjustment coefficient and n represents the number of similar content;
[0011] S3: Obtain the view count ranking of science popularization websites, obtain the total number of views S and the number of shares Sp of science popularization content within a preset time period T in the view count ranking, and calculate the quality coefficient of the science popularization content. Where γ1 and γ2 represent the preset first and second quality coefficients, respectively, 0 < γ1 < γ2, P represents the completion rate of the popular science content, and η represents the preset amplification coefficient; the recommendation value of the popular science content is calculated. Among them, G s Based on the attention value of similar content, the popular science content is sorted from high to low according to the recommendation value, and the top 10% of popular science content is selected as recommended content.
[0012] As a further aspect of the present invention: In step S1, if the number of times the user views popular science content N < N min Then stop further operations on this user, where N min This represents the preset minimum number of page views.
[0013] As a further aspect of the present invention: in step S2, the user browsing time t < t min Popular science content was removed and not included in the registration date. L Up to the current date D N The number of views N for popular science content, where t min This represents the preset minimum browsing time.
[0014] As a further aspect of the present invention: In step S2, the popular science content that the user has already viewed is marked as duplicate content, and the recommendation value Rec = 0.5 * Re for the duplicate content is calculated. As a further aspect of the present invention: In step S3, if there are popular science contents with the same recommendation value, the popular science content with the higher attention value is ranked higher.
[0015] As a further aspect of the present invention: in step S1, the daily pageview count R ≤ R sta Users are designated as non-priority users. Popular science content is sorted from high to low according to quality index, and the top 10% of popular science content is selected as recommended content for non-priority users.
[0016] As a further aspect of the present invention: in step S2, the maximum duration t of key users browsing popular science content is obtained. max If the total playback time of the science video is ts > 2*t max If it is not, it will not be recommended.
[0017] A personalized recommendation system for popular science content based on big data includes:
[0018] Filtering module: Retrieves the user's registration date (D) on the science popularization website. L and the current date D N Calculate the usage date D=D N -D L Get the user's registration date D L Up to the current date D N The number of views N for popular science content, including popular science articles and videos, is calculated as follows: Daily views R = N / D. If daily views R > R... sta Users are designated as key users, among whom R sta The preset judgment value;
[0019] Calculation module: Obtain the total time t spent by key users browsing popular science content, and calculate the average time spent t. ave =t / N, calculates the browsing intention value of key users for popular science content. Where ts represents the completion time of the science video, A1 and A2 are the preset first and second intention coefficients, A1 < 0, and B is the preset intention threshold.
[0020] Science popularization content belonging to the same section in a science popularization website is categorized as similar content. The section refers to the category directory in the navigation bar of the science popularization website. The attention value of key users to similar content is calculated as G=λ*Y*n, where λ represents a preset adjustment coefficient and n represents the number of similar content.
[0021] Recommendation module: Retrieves the view count ranking of science popularization websites, obtains the total number of views (S) and the number of shares (Sp) of science popularization content within a preset time period (T) in the view count ranking, and calculates the quality coefficient of the science popularization content. Where γ1 and γ2 represent the preset first and second quality coefficients respectively, 0 < γ1 < γ2, P represents the percentage of complete playback of popular science content, and η represents the preset amplification coefficient.
[0022] Calculate the recommendation value of popular science content. Among them, G s Based on the attention value of similar content, the popular science content is sorted from high to low according to the recommendation value, and the top 10% of popular science content is selected as recommended content.
[0023] The beneficial effects of this invention are as follows: In this invention, it is necessary to first obtain data during the operation and management of science popularization websites and the analysis of user behavior. In order to more accurately understand the user's activity level and consumption of science popularization content, a series of detailed data collection and calculation work is required.
[0024] First, we need to obtain each user's registration date on the science popularization website and label this registration date as DL. This date is the starting point for the user's connection with the science popularization website, and it is of great significance for subsequent analysis of user behavior and engagement. Simultaneously, we also need to obtain the current date in real time and label it as DN. The current date represents the time point for our data analysis; by comparing it with the registration date, we can clearly understand the length of time a user stays on the website.
[0025] Next, we calculate the usage date D, where D reflects the time span from user registration to the current moment. This time span is an important foundation for subsequent analysis of user browsing behavior.
[0026] Next, we count the number of times (N) a user viewed science content from the registration date to the current date. It's important to note that this science content encompasses a wide variety of forms, including science articles and videos. By counting the number of times users viewed various types of science content during this period, we can obtain a comprehensive view count (N), and then calculate the average number of science content views per user per day. This value directly reflects the user's level of attention to and frequency of use of science content.
[0027] Finally, to identify users with high interest and engagement in science popularization content, a preset threshold was set. The calculated daily pageviews R were compared to Rsta; if R was greater than Rsta, the user was designated as a key user. These key users may have greater potential and value in the dissemination and promotion of science knowledge, as well as in their own in-depth learning of science content. Their behavior and needs are of significant reference value for the operation and development of the science popularization website.
[0028] In today's online environment, users' time is becoming increasingly valuable, and the time they can dedicate to popular science content is relatively limited. This characteristic necessitates a more nuanced consideration of the time factor when analyzing and processing popular science content. Specifically, we need to accurately assess the average time users spend on popular science content by deeply analyzing their historical records. This includes not only the time users spend browsing popular science articles and watching popular science videos, but also multiple dimensions such as user participation in popular science interactions and comment feedback. Through this comprehensive and meticulous analysis, we can obtain a more accurate profile of user time spending.
[0029] The reason for emphasizing the limited time users spend on science content is that if such content occupies too much of their time, it may lead to user boredom and reduce their interest in browsing it. This is undoubtedly a problem that science platforms need to avoid. Therefore, we must constantly monitor users' time spending to ensure that science content provides maximum value within a limited time, while maintaining user interest and engagement.
[0030] Furthermore, it's crucial to analyze the types of science content users browse to determine their level of interest in those specific categories. This not only helps in gaining a deeper understanding of user interests and preferences but also provides vital data for subsequent science content recommendations. Through precise recommendation algorithms, we can push science content that users are interested in, thereby improving user satisfaction and loyalty.
[0031] It is important to note that this invention does not impose specific requirements on the categorization of science popularization content. Instead, it primarily relies on the categorization rules of the science popularization platform itself. This means that regardless of how the science popularization platform classifies and manages its content, our method and system can flexibly adapt and provide users with personalized science popularization content recommendations and services. This flexibility not only improves the adaptability of our system but also reduces the difficulty of integrating with different science popularization platforms, providing strong support for the widespread application of our method and system.
[0032] To comprehensively evaluate the influence and user appeal of science popularization website content, it's essential to first obtain the website's view count ranking. This ranking, based on the number of views of science popularization content, clearly demonstrates which content is most popular with users. View count is not only an important indicator of the popularity of science popularization content but also indirectly reflects the quality of the content. Generally, content with high view counts tends to be of higher quality, more engaging, or have a wider range of topic coverage.
[0033] After obtaining the view count rankings, it's necessary to further analyze the specific number of views and shares for each piece of science content within the rankings. View counts reflect users' interest and attention to the content, while share counts reflect users' approval and willingness to share. These two data points together constitute an important basis for evaluating the quality of science content. By comprehensively analyzing view counts and share counts, we can more accurately determine which content is truly welcomed and loved by users.
[0034] Next, the quality of the science popularization content is evaluated based on the calculated quality coefficient. The quality coefficient is a comprehensive indicator that considers multiple factors such as play counts, page views, and share counts, providing a more complete reflection of the content's quality. To more accurately assess content quality, user engagement metrics also need to be considered. Engagement metrics reflect users' attention and interest in a specific topic or field, serving as an important indicator of content-user relevance.
[0035] Finally, the recommendation score (Re) of the science content is calculated using the formula. In this formula, the quality coefficient Z is used as the base, and the attention value G is used as the exponent. This calculation method effectively represents the comprehensive comparison between the quality of the science content and the user's attention value. The recommendation score only increases when both the quality coefficient and the attention value are high. This means that only content that simultaneously meets the conditions of high quality and high user attention can obtain a high recommendation score. In this way, we can filter out science content that is both high-quality and matches user interests, recommending the most suitable content to users.
[0036] In summary, this invention implements a strategy and method for adjusting recommendations based on users' interests, thereby filtering out high-quality science popularization content that aligns with user interests, achieving the goal of precise push notifications, and providing users with a better user experience. Attached Figure Description
[0037] The invention will now be further described with reference to the accompanying drawings.
[0038] Figure 1 This is a flowchart illustrating the personalized recommendation system and method for popular science content based on big data, as described in this invention. Detailed Implementation
[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0040] Please see Figure 1 As shown, this invention is a personalized recommendation method for popular science content based on big data, comprising the following steps:
[0041] S1: Obtain the user's registration date on the science popularization website. L and the current date D N Calculate the usage date D=D N -D L Get the user's registration date D L Up to the current date D N The number of views N for popular science content, including popular science articles and videos, is calculated as follows: Daily views R = N / D. If daily views R > R... sta Users are designated as key users, among whom R sta The preset judgment value;
[0042] S2: Obtain the total time t spent by key users browsing popular science content, and calculate the average time spent t. ave =t / N, calculates the browsing intention value of key users for popular science content. Where ts represents the completion time of the science video, A1 and A2 are the preset first and second intention coefficients, A1 < 0, and B is the preset intention threshold.
[0043] Science popularization content belonging to the same section in a science popularization website is categorized as similar content. The section refers to the category directory in the navigation bar of the science popularization website. The attention value of key users to similar content is calculated as G=λ*Y*n, where λ represents a preset adjustment coefficient and n represents the number of similar content.
[0044] S3: Obtain the view count ranking of science popularization websites, obtain the total number of views S and the number of shares Sp of science popularization content within a preset time period T in the view count ranking, and calculate the quality coefficient of the science popularization content. Where γ1 and γ2 represent the preset first and second quality coefficients respectively, 0 < γ1 < γ2, P represents the percentage of complete playback of popular science content, and η represents the preset amplification coefficient.
[0045] Calculate the recommendation value of popular science content. Among them, G sBased on the attention value of similar content, the popular science content is sorted from high to low according to the recommendation value, and the top 10% of popular science content is selected as recommended content.
[0046] It should be noted that, firstly, in the process of operating and managing science popularization websites and analyzing user behavior, in order to more accurately understand user activity and consumption of science popularization content, a series of detailed data collection and calculation work is required.
[0047] First, we need to obtain each user's registration date on the science popularization website and label this registration date as DL. This date is the starting point for the user's connection with the science popularization website, and it is of great significance for subsequent analysis of user behavior and engagement. Simultaneously, we also need to obtain the current date in real time and label it as DN. The current date represents the time point for our data analysis; by comparing it with the registration date, we can clearly understand the length of time a user stays on the website.
[0048] Next, we calculate the usage date D, where D reflects the time span from user registration to the current moment. This time span is an important foundation for subsequent analysis of user browsing behavior.
[0049] Next, we count the number of times (N) a user viewed science content from the registration date to the current date. It's important to note that this science content encompasses a wide variety of forms, including science articles and videos. By counting the number of times users viewed various types of science content during this period, we can obtain a comprehensive view count (N), and then calculate the average number of science content views per user per day. This value directly reflects the user's level of attention to and frequency of use of science content.
[0050] Finally, to identify users with high interest and engagement in science popularization content, a preset threshold was set. The calculated daily pageviews R were compared to Rsta; if R was greater than Rsta, the user was designated as a key user. These key users may have greater potential and value in the dissemination and promotion of science knowledge, as well as in their own in-depth learning of science content. Their behavior and needs are of significant reference value for the operation and development of the science popularization website.
[0051] In today's online environment, users' time is becoming increasingly valuable, and the time they can dedicate to popular science content is relatively limited. This characteristic necessitates a more nuanced consideration of the time factor when analyzing and processing popular science content. Specifically, we need to accurately assess the average time users spend on popular science content by deeply analyzing their historical records. This includes not only the time users spend browsing popular science articles and watching popular science videos, but also multiple dimensions such as user participation in popular science interactions and comment feedback. Through this comprehensive and meticulous analysis, we can obtain a more accurate profile of user time spending.
[0052] The reason for emphasizing the limited time users spend on science content is that if such content occupies too much of their time, it may lead to user boredom and reduce their interest in browsing it. This is undoubtedly a problem that science platforms need to avoid. Therefore, we must constantly monitor users' time spending to ensure that science content provides maximum value within a limited time, while maintaining user interest and engagement.
[0053] Furthermore, it's crucial to analyze the types of science content users browse to determine their level of interest in those specific categories. This not only helps in gaining a deeper understanding of user interests and preferences but also provides vital data for subsequent science content recommendations. Through precise recommendation algorithms, we can push science content that users are interested in, thereby improving user satisfaction and loyalty.
[0054] It is important to note that this invention does not impose specific requirements on the categorization of science popularization content. Instead, it primarily follows the categorization rules of the science popularization platform itself, with the type of content determined by its placement within a specific navigation bar. This means that regardless of how the science popularization platform categorizes and manages its content, our method and system can flexibly adapt and provide users with personalized science popularization content recommendations and services. This flexibility not only improves the adaptability of our system but also reduces the difficulty of integrating with different science popularization platforms, providing strong support for the widespread application of our method and system.
[0055] To comprehensively evaluate the influence and user appeal of science popularization website content, it's essential to first obtain the website's view count ranking. This ranking, based on the number of views of science popularization content, clearly demonstrates which content is most popular with users. View count is not only an important indicator of the popularity of science popularization content but also indirectly reflects the quality of the content. Generally, content with high view counts tends to be of higher quality, more engaging, or have a wider range of topic coverage.
[0056] After obtaining the view count rankings, it's necessary to further analyze the specific number of views and shares for each piece of science content within the rankings. View counts reflect users' interest and attention to the content, while share counts reflect users' approval and willingness to share. These two data points together constitute an important basis for evaluating the quality of science content. By comprehensively analyzing view counts and share counts, we can more accurately determine which content is truly welcomed and loved by users.
[0057] Next, the quality of the science popularization content is evaluated based on the calculated quality coefficient. The quality coefficient is a comprehensive indicator that considers multiple factors such as play counts, page views, and share counts, providing a more complete reflection of the content's quality. To more accurately assess content quality, user engagement metrics also need to be considered. Engagement metrics reflect users' attention and interest in a specific topic or field, serving as an important indicator of content-user relevance.
[0058] Finally, the recommendation score (Re) of the science content is calculated using the formula. In this formula, the quality coefficient Z is used as the base, and the attention value G is used as the exponent. This calculation method effectively represents the comprehensive comparison between the quality of the science content and the user's attention value. The recommendation score only increases when both the quality coefficient and the attention value are high. This means that only content that simultaneously meets the conditions of high quality and high user attention can obtain a high recommendation score. In this way, we can filter out science content that is both high-quality and matches user interests, recommending the most suitable content to users.
[0059] In another preferred embodiment of the present invention, if the number of times a user views popular science content N < N min Then stop further operations on this user, where N min This represents the preset minimum number of page views.
[0060] It is worth noting that the amount of data is crucial to the reliability and accuracy of the results during data analysis and processing. This is especially true in the statistics of key metrics such as views, shares, and reposts of popular science content. If too little data is available, these metrics may not fully and accurately reflect the actual quality of the popular science content and genuine user feedback.
[0061] Therefore, in this case, we choose not to perform any further operations to avoid misleading conclusions due to insufficient data.
[0062] In another preferred embodiment of the present invention, the user browsing time t < t min Popular science content was removed and not included in the registration date. L Up to the current date D NThe number of views N for popular science content, where t min This represents the preset minimum browsing time.
[0063] Understandably, in practice, we have found that the browsing time for some popular science content is too short. This may not reflect the user's true intention, but rather be due to accidental operation.
[0064] By removing these outliers, we can more accurately assess users' interest in and engagement with science content, providing a more accurate and reliable basis for subsequent analysis and recommendations. This not only helps improve the quality of data analysis but also ensures that the content recommended to users better matches their actual needs and interests, thereby enhancing user satisfaction with science websites.
[0065] In another preferred embodiment of the present invention, popular science content that the user has already viewed is marked as duplicate content, and the recommendation value Rec = 0.5 * Re for duplicate content is calculated.
[0066] It is important to reduce the probability of duplicate content being recommended to users, and to avoid users being repeatedly recommended the same science popularization content, thereby reducing users' goodwill towards the science popularization website.
[0067] In another preferred embodiment of the present invention, if there are popular science content with the same recommendation value, the popular science content with the higher attention value is ranked first.
[0068] It should be noted that popular science content with high attention values is content that users prefer. Therefore, when there are popular science contents with equal recommendation values, the popular science content that users prefer should be recommended.
[0069] In another preferred embodiment of the present invention, the daily pageview count R ≤ R sta Users are designated as non-priority users. Popular science content is sorted from high to low according to quality index, and the top 10% of popular science content is selected as recommended content for non-priority users.
[0070] Understandably, it's difficult to determine a user's preferences based on their browsing history if they have a low number of views. Therefore, recommendations can only be made based on the quality of the science content. High-quality science content is recommended, and once the daily views reach a certain level, the science content that the user likes is then filtered out based on their browsing preferences.
[0071] In another preferred embodiment of the present invention, the maximum duration t of browsing popular science content by key users is obtained. max If the total playback time of the science video is ts > 2*t max If it is not, it will not be recommended.
[0072] It is worth noting that if the total playback time of a video is much longer than what users can accept, in order to avoid invalid recommendations, science videos with long playback times will be removed and not included in the recommended content.
[0073] A personalized recommendation system for popular science content based on big data includes:
[0074] Filtering module: Retrieves the user's registration date (D) on the science popularization website. L and the current date D N Calculate the usage date D=D N -D L Get the user's registration date D L Up to the current date D N The number of views N for popular science content, including popular science articles and videos, is calculated as follows: Daily views R = N / D. If daily views R > R... sta Users are designated as key users, among whom R sta The preset judgment value;
[0075] Calculation module: Obtain the total time t spent by key users browsing popular science content, and calculate the average time spent t. ave =t / N, calculates the browsing intention value of key users for popular science content. Where ts represents the completion time of the science video, A1 and A2 are the preset first and second intention coefficients, A1 < 0, and B is the preset intention threshold.
[0076] Science popularization content belonging to the same section in a science popularization website is categorized as similar content. The section refers to the category directory in the navigation bar of the science popularization website. The attention value of key users to similar content is calculated as G=λ*Y*n, where λ represents a preset adjustment coefficient and n represents the number of similar content.
[0077] Recommendation module: Retrieves the view count ranking of science popularization websites, obtains the total number of views (S) and the number of shares (Sp) of science popularization content within a preset time period (T) in the view count ranking, and calculates the quality coefficient of the science popularization content. Where γ1 and γ2 represent the preset first and second quality coefficients respectively, 0 < γ1 < γ2, P represents the percentage of complete playback of popular science content, and η represents the preset amplification coefficient.
[0078] Calculate the recommendation value of popular science content. Among them, G s Based on the attention value of similar content, the popular science content is sorted from high to low according to the recommendation value, and the top 10% of popular science content is selected as recommended content.
[0079] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A personalized recommendation method for popular science content based on big data, characterized in that, Includes the following steps: S1: Obtain the user's registration date on the science popularization website. L and the current date D N Calculate the usage date D=D N -D L Get the user's registration date D L Up to the current date D N The number of views N for popular science content, including popular science articles and videos, is calculated as follows: Daily views R = N / D. If daily views R > R... sta Users are designated as key users, among whom R sta The preset judgment value; S2: Obtain the total time t spent by key users browsing popular science content, and calculate the average time spent t. ave =t / N, calculates the browsing intention value of key users for popular science content. Where ts represents the completion time of the science video, A1 and A2 are the preset first and second intention coefficients, A1 < 0, and B is the preset intention threshold. Science popularization content belonging to the same section in a science popularization website is categorized as similar content. The section refers to the category directory in the navigation bar of the science popularization website. The attention value of key users to similar content is calculated as G=λ*Y*n, where λ represents a preset adjustment coefficient and n represents the number of similar content. S3: Obtain the view count ranking of science popularization websites, obtain the total number of views S and the number of shares Sp of science popularization content within a preset time period T in the view count ranking, and calculate the quality coefficient of the science popularization content. Where γ1 and γ2 represent the preset first and second quality coefficients, respectively, 0 < γ1 < γ2, P represents the completion rate of the popular science content, and η represents the preset amplification coefficient; the recommendation value of the popular science content is calculated. Among them, G s Based on the attention value of similar content, the popular science content is sorted from high to low according to the recommendation value, and the top 10% of popular science content is selected as recommended content.
2. The personalized recommendation method for popular science content based on big data according to claim 1, characterized in that, In step S1, if the number of times a user views popular science content N < N min Then stop further operations on this user, where N min This represents the preset minimum number of page views.
3. The personalized recommendation method for popular science content based on big data according to claim 1, characterized in that, In step S2, the user browsing time t < t min Popular science content was removed and not included in the registration date. L Up to the current date D N The number of views N for popular science content, where t min This represents the preset minimum browsing time.
4. The personalized recommendation method for popular science content based on big data according to claim 1, characterized in that, In step S2, the popular science content that the user has already viewed is marked as duplicate content, and the recommendation value Rec = 0.5 * Re for duplicate content is calculated.
5. The personalized recommendation method for popular science content based on big data according to claim 1, characterized in that, In step S3, if there are popular science content with the same recommendation value, the popular science content with the higher attention value will be ranked higher.
6. The personalized recommendation method for popular science content based on big data according to claim 1, characterized in that, In step S1, the daily pageview count R ≤ R sta Users are designated as non-priority users. Popular science content is sorted from high to low according to quality index, and the top 10% of popular science content is selected as recommended content for non-priority users.
7. The personalized recommendation method for popular science content based on big data according to claim 1, characterized in that, In step S2, the maximum duration t of key users browsing science popularization content is obtained. max If the total playback time of the science video is ts > 2*t max If it is not, it will not be recommended.
8. A personalized recommendation system for popular science content based on big data, characterized in that, include: Filtering module: Retrieves the user's registration date (D) on the science popularization website. L and the current date D N Calculate the usage date D=D N -D L Get the user's registration date D L Up to the current date D N The number of views N for popular science content, including popular science articles and videos, is calculated as follows: Daily views R = N / D. If daily views R > R... sta Users are designated as key users, among whom R sta The preset judgment value; Calculation module: Obtain the total time t spent by key users browsing popular science content, and calculate the average time spent t. ave =t / N, calculates the browsing intention value of key users for popular science content. Where ts represents the completion time of the science video, A1 and A2 are the preset first and second intention coefficients, A1 < 0, and B is the preset intention threshold. Science popularization content belonging to the same section in a science popularization website is categorized as similar content. The section refers to the category directory in the navigation bar of the science popularization website. The attention value of key users to similar content is calculated as G=λ*Y*n, where λ represents a preset adjustment coefficient and n represents the number of similar content. Recommendation module: Retrieves the view count ranking of science popularization websites, obtains the total number of views (S) and the number of shares (Sp) of science popularization content within a preset time period (T) in the view count ranking, and calculates the quality coefficient of the science popularization content. Where γ1 and γ2 represent the preset first and second quality coefficients, respectively, 0 < γ1 < γ2, P represents the completion rate of the popular science content, and η represents the preset amplification coefficient; the recommendation value of the popular science content is calculated. Where Gs is the attention value of the popular science content in the same type of content. Popular science content is sorted from high to low according to the recommendation value, and the top 10% of popular science content is selected as recommended content.
Citation Information
Patent Citations
Safe and controllable intelligent recommendation system in multimedia content environment
CN107943864A
Content recommendation method, device and system
CN114707074A