A method for analyzing and predicting library popularity based on time-dimensional information
By collecting library user data and activity information, and analyzing library popularity using the time dimension, the problem of difficulty in assessing the popularity of public cultural services has been solved, enabling scientific prediction of library popularity and evaluation of service effectiveness.
Patent Information
- Application Number
- CN202210683506.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-16
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-06-16
AI Technical Summary
Existing technologies lack scientific and accurate methods for analyzing and predicting the popularity of public cultural services, especially in the analysis and prediction of library popularity, making it difficult to effectively assess the service effectiveness of public cultural resources.
By collecting data on library users' entry into the library, their borrowing information, and public cultural activities, and utilizing time-related information, the library calculates the popularity of user borrowing behavior and the popularity of activity impact. Combined with popularity analysis methods, the library's future level of attention is predicted.
It enables scientific, accurate, and reliable analysis and prediction of library popularity, allowing for better evaluation of the service effectiveness of public cultural resources and providing a periodic popularity prediction method to determine the future level of attention a library will receive.
Smart Images

Figure CN114881710B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of public cultural service technology, specifically relating to a method for analyzing and predicting library popularity based on time-dimensional information. Background Technology
[0002] In recent years, the state has successively implemented public digital cultural projects benefiting the people, such as the National Cultural Information Resource Sharing Project, the Digital Library Promotion Project, and the Public Electronic Reading Room Construction Plan. These initiatives have established a nationwide service network, formed a large-scale digital resource database, and developed a technical support service platform represented by the National Public Cultural Cloud, laying the foundation for the digitalization and networking of the public cultural service system. However, existing literature in recent years has primarily focused on the quantitative analysis of online news or public opinion, using content feature extraction and statistical machine learning techniques to predict the popularity or evolution trend of news or public opinion. There has been little analysis and prediction of the popularity of public cultural services. This invention collects and analyzes basic library data. Based on summarizing and organizing public cultural resource and service data, it analyzes and predicts the popularity of public cultural services, achieving a more scientific, realistic, and accurate evaluation of the service effectiveness of public cultural resources. Summary of the Invention
[0003] To overcome the shortcomings of the existing technology, the purpose of this invention is to provide a method for analyzing and predicting library popularity based on time-dimensional information. This method uses user borrowing information and library public cultural activities as a dataset, and calculates a library popularity value through popularity analysis to determine the current level of attention the library receives. Furthermore, it uses the calculated popularity value to predict the future level of attention the library will receive.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0005] A method for analyzing and predicting library popularity based on time-related information includes the following steps:
[0006] Step 1: Collect basic data on user entry information, user borrowing information, public cultural activities held by the library, various seminars and theoretical activities, and exchange activities;
[0007] Step 2: Using data such as book borrowing time, borrowing duration, user type, and number of borrowings, calculate the library user borrowing behavior popularity score. The calculation of the borrowing behavior popularity score is as follows:
[0008] (1) Convert user borrowing date data into timestamps, accurate to the hour, as the x-axis, and the number of books borrowed by the user as the y-axis, and standardize them. The x-axis set is y = (y1,...,y2). n), the set of ordinates z = (z1,...,z2) n The coordinate normalization formula is:
[0009]
[0010]
[0011] (2) Calculate the average distance between all points, denoted as ε. Starting from any point Q, with Q as the center point and ε as the radius, include the data points within the range of that point into the cluster. When no new data points are included in the cluster, start randomly selecting the next center point until all points are traversed.
[0012] (3) Calculate the average height difference of all center points. Randomly select a cluster C1. In the vertical direction, classify the clusters containing sample points whose height difference with any point in cluster C1 is within the range of the average height difference into new clusters. Repeat this process until all old clusters are classified into new clusters, denoted as .
[0013] (4) Assign weights ω1, ..., ω to different clusters respectively. n ;
[0014] (5) The user borrowing behavior popularity value is T v calculate:
[0015]
[0016] Where ω i is the weight coefficient of the i-th class, k is the number of data in the i-th class, and V0 is the initial popularity value of the user's borrowing;
[0017] Step 3: Using data from library-organized public cultural activities, seminars, and exchange events, calculate the activity impact heat value. The activity impact heat value is calculated as follows:
[0018] (1) Utilize data from the library regarding public cultural activities, various seminars, and exchange activities organized by the library, using x1, ..., x n Let each of the following groups represent different types of activities held by the library, from activity 1 to activity n. Keywords are extracted from the activity text information, and the weight of each keyword is calculated:
[0019]
[0020]
[0021] r iLet x be the weight of keyword i, x be the number of times keyword i appears in the activity text information, and f(x) be the normal distribution of the number of occurrences of keyword i in the activity text information. The average frequency of all keywords, σ is the standard deviation of the keyword distribution, and n i N represents the number of activity information texts containing keyword i, and N represents the total number of activity information texts.
[0022] (2) Extract the number of books borrowed related to the event theme and calculate the impact value of the event's popularity:
[0023]
[0024] M represents the influence value of the activity's popularity, r i f represents the weight of each keyword i. number Borrowing volume of books related to keyword type i;
[0025] Step 4: Calculate the total library popularity score, which includes three parts: user borrowing behavior popularity score, activity impact popularity score, and daily activity popularity score. The calculation of the total library popularity score is as follows:
[0026]
[0027] ω1 represents the proportion of users who have borrowed content to the total number of users, ω2 represents the proportion of users who have not borrowed content to the total number of users, and the user borrowing behavior popularity value is T. v The number of daily active users is denoted as Q, representing the daily activity popularity value, and M represents the activity impact popularity value.
[0028] Step 5: Predict future popularity values based on the library popularity scores.
[0029] The aforementioned method for predicting library popularity based on time-dimensional information includes the following steps:
[0030] (1) Calculate the periodicity T of the library's popularity value based on the variation pattern of the library's popularity value;
[0031] (2) Using vector α i =[α i1 ,α i2 ,...,α in The matrix represents the library's popularity at each time step during the i-th period. This indicates the overall popularity score of the library;
[0032] (3) Let the set of n states be S = {s1, s2, ..., sn}. n In the i-th cycle, its state at each position is S. i ={Si1 ,S i2 ,...,S iN};
[0033] (4) Calculate the initial probability distribution of the library's popularity value at each location within the period based on historical data of the library's popularity value, and calculate the state transition probability matrix of the library's popularity value at each moment within the change period, respectively using P1=[P ij ] n×n P2 = [P ij ] n×n ... P n =[P ij ] n×n Let represent the state transition probability of each position within the i-th period;
[0034] (5) Based on the state transition probability at each position within the period, predict the heat value of each position in the (i+1)th period. The prediction result for the first position in the (i+1)th period is only related to the state of the same position in the previous period and is independent of the state of other positions.
[0035]
[0036] Based on the above formula, the predicted result S of the library popularity value at each location in the (i+1)th period is obtained. i+1 =[S (i+1)1 ,S (i+1)2 ,...,S (i+1)n ].
[0037] The beneficial effects of this invention are:
[0038] This invention uses data such as library user entry information, user borrowing information, and information on public cultural activities as basic data, including book borrowing dates, book return dates, user categories, and public cultural activity information. The library popularity analysis is divided into four parts: data collection, calculation of library user borrowing behavior popularity value, calculation of library activity impact popularity value, calculation of library popularity value, analysis of the degree of attention received by the library and user activity, and finally, based on the characteristics of the number of libraries, a periodic library popularity value prediction method is used to judge the degree of attention the library will receive in the future.
[0039] Based on user borrowing information and library public cultural activities as datasets, a popularity score is calculated using a heat analysis method to determine the current level of attention the library receives. Furthermore, the calculated popularity score is used to predict future library popularity, thereby analyzing and predicting the popularity of public cultural services and achieving a more scientific, realistic, and accurate assessment of the service effectiveness of public cultural resources. Attached Figure Description
[0040] Figure 1 This is a flowchart of the museum popularity analysis process of the present invention;
[0041] Figure 2 This is a flowchart of the heat value prediction process of the present invention. Detailed Implementation
[0042] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0043] Since this invention relates to theories and techniques commonly used in the fields of machine learning, probability theory, and mathematical statistics, it is necessary to explain related content, such as: clusters, cluster centers, initial probability distributions, and state transition probability matrices.
[0044] Cluster: A set of samples generated by clustering. Samples within the same cluster are similar to each other, but different from samples in other clusters;
[0045] Cluster centroid: The point that minimizes the average distance from a data point within a cluster to a certain point within the cluster is called the cluster centroid.
[0046] Initial probability distribution: the initial probability of each state;
[0047] State transition probability matrix: All elements of the matrix are non-negative, and the sum of the elements in each row is equal to 1. Each element is represented by a probability, and the states transition between each other under certain conditions, hence the name transition probability matrix. For example, when used in market decision-making, the elements in the matrix represent the probabilities of retaining, acquiring, or losing the market or customers.
[0048] Figure 1 For the library data popularity analysis section of this invention, such as Figure 1 As shown, a method for analyzing library popularity based on time-dimensional information mainly includes the following four parts: data collection, calculation of library user borrowing behavior popularity value, calculation of library activity impact popularity value, and calculation of library popularity value.
[0049] The specific algorithm flow is explained below:
[0050] Step 1: Collect basic data such as user entry information, user borrowing information, public cultural activities held by the library, various seminars and theoretical activities, and exchange activities. Extract information such as user borrowing time, borrowing duration, and user type, as well as text information such as various seminars and theoretical activities, as a dataset.
[0051] Step 2: Calculate the library user borrowing behavior popularity score using data such as book borrowing time, borrowing duration, user type, and number of borrowings; where:
[0052] Step 22: Convert the user's borrowing date data into timestamps, accurate to the hour, as the x-axis, and the number of books borrowed by the user as the y-axis, and standardize the data. The x-axis set is y = (y1,...,y...). n ), the set of ordinates z = (z1,...,z2) n The coordinate normalization formula is:
[0053]
[0054]
[0055] Step 22: The distance between points is calculated using Euclidean distance. The calculation method is as follows:
[0056]
[0057] Calculate the average distance between all points, denoted as ε. Starting from any point Q, with Q as the center point and ε as the radius, include the data points within the range of that point into the cluster. When no new data points are included in the cluster, start randomly selecting the next center point, until all points are traversed.
[0058] Step 23: Calculate the average height difference of all center points. Randomly select a cluster C1. Vertically, reclassify the clusters containing sample points whose height difference from any point in cluster C1 is within the average height difference range as new clusters. Continue this process until all old clusters have been reclassified as new clusters, denoted as […].
[0059] Different clusters are assigned weights ω1, ..., ω1 respectively. n The specific calculations are as follows:
[0060]
[0061] in For clusters The number of sample points contained in C * numbers The total number of sample points;
[0062] Step 24, the user's borrowing behavior popularity value is T. v calculate:
[0063]
[0064] Where ω i is the weight coefficient of the i-th cluster, k is the number of data in the i-th cluster, and V0 is the initial popularity value of the user's borrowing;
[0065] Step 3: Using web scraping technology, send requests to the target website, parse the scraped data after receiving the response, and finally save the required information about various activities held by the library as text. Using data on public cultural activities, seminars, and exchange activities held by the library, calculate the activity impact score; among which:
[0066] Step 31, use x1, ..., x n Let each of the following groups represent different types of activities held by the library, from activity 1 to activity n. Keywords are extracted from the activity text information, and the weight of each keyword is calculated:
[0067]
[0068]
[0069] r i Let x be the weight of keyword i, x be the number of times keyword i appears in the activity text information, and f(x) be the normal distribution of the number of occurrences of keyword i in the activity text information. The average frequency of all keywords, σ is the standard deviation of the keyword distribution, and n i N represents the number of activity information texts containing keyword i, and N represents the total number of activity information texts.
[0070] Step 32: Extract the number of books borrowed related to the event theme and calculate the event's popularity impact value:
[0071]
[0072] M represents the influence value of the activity's popularity, r i f represents the weight of each keyword i. number The number of books borrowed related to keyword type i;
[0073] Step 4: Calculate the library popularity score, which includes the user borrowing behavior popularity score T. v The library's popularity score is composed of three parts: daily activity popularity (Q), activity impact popularity (M). A higher library popularity score indicates greater public attention and higher user activity. The library popularity score is calculated using the following formula, and the specific calculation method is as follows:
[0074]
[0075] ω1 represents the proportion of users who have borrowed content to the total number of users, ω2 represents the proportion of users who have not borrowed content to the total number of users, and the user borrowing behavior popularity value is T. v The number of daily active users is denoted as Q, representing the daily activity popularity value, and M represents the activity impact popularity value.
[0076] Figure 2 This is a method for predicting library popularity based on time-related information, such as... Figure 2 As shown, the specific process is explained below:
[0077] Step 51: Calculate the cycle T of the library's popularity value based on the variation pattern of the library's popularity value;
[0078] Step 52, using vector α i =[α i1 ,α i2 ,...,α in The matrix represents the library's popularity at each time step during the i-th period. This indicates the overall popularity score of the library;
[0079] Step 53, use a set of n state labels as S = {s1, s2, ..., s} n In the i-th cycle, its state at each position is S. i ={S i1 ,S i2 ,...,S iN};
[0080] Step 54: Calculate the initial probability distribution of library popularity values at each location within the period based on historical library popularity data, and calculate the state transition probability matrix of library popularity values at each time point within the change period, using P1 = [P ij ] n×n P2 = [P ij ] n×n ... P n =[P ij ] n×n Let represent the state transition probability of each position within the i-th period;
[0081] Step 5: Based on the state transition probabilities at each position within the period, predict the heat value of each position within the (i+1)th period. The prediction result for the first position within the (i+1)th period is only related to the state of the same position in the previous period and is independent of the states of other positions.
[0082]
[0083] Based on the above formula, the predicted result S of the library popularity value at each location in the (i+1)th period is obtained. i+1 =[S (i+1)1 ,S (i+1)2 ,...,S (i+1)n ].
Claims
1. A method for analyzing and predicting library popularity based on time-dimensional information, characterized in that, Includes the following steps: Step 1: Collect basic data on user entry information, user borrowing information, public cultural activities held by the library, various seminars and theoretical activities, and exchange activities; Step 2: Calculate the library user borrowing behavior popularity score using data on book borrowing time, borrowing duration, user type, and number of borrowings. The calculation of the borrowing behavior popularity score is as follows: (1) Convert user borrowing date data into timestamps, accurate to the hour, as the x-axis, and the number of books borrowed by the user as the y-axis, and standardize them. The x-axis set is y = (y1,...,y2). n ), the set of ordinates z = (z1,...,z2) n The coordinate normalization formula is: (2) Calculate the average distance between all points, denoted as ε. Starting from any point Q, with Q as the center point and ε as the radius, include the data points within the range of that point into the cluster. When no new data points are included in the cluster, start randomly selecting the next center point until all points are traversed. (3) Calculate the average height difference of all center points. Randomly select a cluster C1. In the vertical direction, classify the clusters containing sample points whose height difference with any point in cluster C1 is within the range of the average height difference into new clusters. Repeat this process until all old clusters are classified into new clusters, denoted as . (4) Assign weights ω1, ..., ω to different clusters respectively. n ; (5) The user borrowing behavior popularity value is T v calculate: Where ω i is the weight coefficient of the i-th class, k is the number of data in the i-th class, and V0 is the initial popularity value of the user's borrowing; Step 3: Utilize data from library-organized public cultural activities, various seminars, and exchange events to calculate the activity impact intensity value. The activity impact intensity value is calculated as follows: (1) Utilize data on public cultural activities, various seminars, and exchange activities held by the library, using x1, ..., x n Let each of the following groups represent different types of activities held by the library, from activity 1 to activity n. Keywords are extracted from the activity text information, and the weight of each keyword is calculated: r i Let x be the weight of keyword i, x be the number of times keyword i appears in the activity text information, and f(x) be the normal distribution of the number of occurrences of keyword i in the activity text information. The average frequency of all keywords, σ is the standard deviation of the keyword distribution, and n i N represents the number of activity information texts containing the keyword i, and N represents the total number of activity information texts. (2) Extract the number of books borrowed related to the event theme and calculate the impact value of the event's popularity: M represents the influence value of the activity's popularity, r i f represents the weight of each keyword i. number Borrowing volume of books related to keyword type i; Step 4: Calculate the total library popularity score, which includes three parts: user borrowing behavior popularity score, activity impact popularity score, and daily activity popularity score. The calculation of the total library popularity score is as follows: ω1 represents the proportion of users who have borrowed content to the total number of users, ω2 represents the proportion of users who have not borrowed content to the total number of users, and the user borrowing behavior popularity value is T. v The number of daily active users is denoted as Q, representing the daily activity popularity value, and M represents the activity impact popularity value. Step 5: Predict future popularity values based on the library popularity scores.
2. The method for analyzing and predicting library popularity based on time-dimensional information according to claim 1, characterized in that, The aforementioned method for predicting library popularity based on time-dimensional information includes the following steps: (1) Calculate the periodicity T of the library's popularity value based on the variation pattern of the library's popularity value; (2) Using vector α i =[α i1 ,α i2 ,…,α in The matrix represents the library's popularity at each time step during the i-th period. This indicates the overall popularity score of the library; (3) Let the set of n states be S = {s1, s2, ..., s}. n In the i-th cycle, its state at each position is S. i ={S i1 ,S i2 ,…,S iN }; (4) Calculate the initial probability distribution of the library's popularity value at each location within the period based on historical data of the library's popularity value, and calculate the state transition probability matrix of the library's popularity value at each moment within the change period, respectively using P1=[P ij ] n×N P2 = [P ij ] n×N ... P n =[P ij ] n×N Let represent the state transition probability of each position within the i-th period; (5) Based on the state transition probability at each position within the period, predict the heat value of each position in the (i+1)th period. The prediction result for the first position in the (i+1)th period is only related to the state of the same position in the previous period and is independent of the state of other positions. Based on the above formula, the predicted result S of the library popularity value at each location in the (i+1)th period is obtained. i+1 =[S (i+1)1 ,S (i+1)2 ,...,S (i+1)N ].
Citation Information
Patent Citations
A method for predicting passenger flow change trend of urban rail transit
CN109272168A
A book recommendation method based on data mining
CN109408600A