A book borrowing trend prediction method and system based on big data analysis

By constructing heterogeneous information networks and learning network representations, the interest characteristics of potential user communities are identified, solving the problem of insufficient accuracy in traditional book borrowing prediction methods and realizing precise allocation and personalized recommendations of library resources.

CN121504085BActive Publication Date: 2026-07-21HENAN POLYTECHNIC UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HENAN POLYTECHNIC UNIV
Filing Date
2025-12-11
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional book borrowing prediction methods cannot effectively capture users' potential needs across multiple dimensions, resulting in insufficient accuracy in book purchasing and recommendations, and failing to achieve personalized recommendations and optimized resource allocation.

Method used

By acquiring multiple heterogeneous data sources, constructing a heterogeneous information network, performing network representation learning, extracting the interest features of potential user communities, predicting their borrowing needs for books not yet in the library, and generating accurate book purchasing suggestions.

Benefits of technology

It has optimized the allocation of library resources and made personalized recommendations, improved the accuracy of borrowing demand forecasting, and enhanced the accuracy of book procurement and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504085B_ABST
    Figure CN121504085B_ABST
Patent Text Reader

Abstract

The application discloses a book borrowing trend prediction method and system based on big data analysis, and relates to the technical field of book management. The method comprises the following steps: acquiring a plurality of heterogeneous data sources, and vectorizing all unstructured data to construct a heterogeneous information network; performing network representation learning on the heterogeneous information network to obtain low-dimensional vector representation of all nodes, performing clustering identification based on the low-dimensional vector representation, extracting a potential user community with similar implicit interest characteristics, and generating a corresponding group interest characteristic vector; for any potential user community, predicting its borrowing demand tendency for uncataloged books, and generating a book purchasing suggestion. The application solves the technical problem that the existing book borrowing prediction method cannot fully tap the potential interest demand of users, resulting in insufficient accuracy of book purchasing and recommendation, and achieves the technical effect of accurately identifying the potential demand of users through big data analysis and fusion of heterogeneous data sources, and improving the accuracy of borrowing demand prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of library management technology, specifically to a method and system for predicting book borrowing trends based on big data analysis. Background Technology

[0002] With the rapid development of digitalization and informatization in libraries, users' demand for book resources is becoming increasingly diversified and personalized. However, traditional book borrowing prediction methods mainly rely on users' historical borrowing data and predict borrowing trends through simple statistical analysis. This cannot effectively capture users' potential needs in multiple dimensions, especially in the discovery and recommendation of niche books. It is difficult to accurately predict users' borrowing needs, and library book procurement is somewhat blind, making it difficult to achieve personalized recommendations and optimized resource allocation. Summary of the Invention

[0003] This application provides a method and system for predicting book borrowing trends based on big data analysis, which addresses the technical problem that existing book borrowing prediction methods cannot fully tap into users' potential interests and needs, resulting in insufficient accuracy in book purchasing and recommendations.

[0004] The first aspect of this application provides a method for predicting book borrowing trends based on big data analysis. The method includes: acquiring multiple heterogeneous data sources, including book metadata, historical book borrowing data, user access data to electronic resources on a library's digital portal, and anonymized user profile tags; preprocessing the multiple heterogeneous data sources and vectorizing all unstructured data to construct a heterogeneous information network containing user nodes, interest point nodes, and book nodes; performing network representation learning on the heterogeneous information network to obtain low-dimensional vector representations of all nodes; performing clustering identification based on the low-dimensional vector representations to extract potential user communities with similar latent interest characteristics and generating corresponding group interest feature vectors; for any potential user community, predicting its borrowing demand tendency for books not yet in the library's collection based on its group interest feature vectors and generating corresponding book purchasing suggestions.

[0005] A second aspect of this application provides a book borrowing trend prediction system based on big data analysis. The system includes: a heterogeneous data source acquisition module for acquiring multiple heterogeneous data sources, including book metadata, historical book borrowing data, user access data to electronic resources on the library's digital portal, and anonymized user profile tags; a data vectorization module for preprocessing the multiple heterogeneous data sources and vectorizing all unstructured data to construct a heterogeneous information network containing user nodes, interest point nodes, and book nodes; a network representation learning module for performing network representation learning on the heterogeneous information network to obtain low-dimensional vector representations of all nodes; a clustering identification module for performing clustering identification based on the low-dimensional vector representations, extracting potential user communities with similar latent interest characteristics, and generating corresponding group interest feature vectors; and a borrowing demand prediction module for predicting the borrowing demand tendency of any potential user community for books not yet in the library's collection, based on its group interest feature vector, and generating corresponding book purchasing suggestions.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0007] This application provides a book borrowing trend prediction method and system based on big data analysis, relating to the field of library management technology. By integrating multiple heterogeneous data sources to construct a heterogeneous information network, and utilizing network representation learning to obtain low-dimensional vector representations, it identifies potential user communities through clustering. Based on the community's interest characteristics, it predicts their demand for books not yet in the library's collection, thereby generating accurate book purchasing suggestions, optimizing library resource allocation and personalized recommendations. This solves the technical problem that existing book borrowing prediction methods cannot fully explore users' potential interests and needs, leading to insufficient accuracy in book purchasing and recommendations. It achieves the technical effect of accurately identifying users' potential needs and improving the accuracy of borrowing demand prediction through the integration of big data analysis and heterogeneous data sources. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 A schematic diagram of a book borrowing trend prediction method based on big data analysis provided for an embodiment of this application;

[0010] Figure 2 This is a schematic diagram of a book borrowing trend prediction system based on big data analysis, provided as an embodiment of this application.

[0011] Figure labeling: Heterogeneous data source acquisition module 11, data vectorization module 12, network representation learning module 13, clustering recognition module 14, and borrowing demand prediction module 15. Detailed Implementation

[0012] This application provides a method and system for predicting book borrowing trends based on big data analysis, which addresses the technical problem that existing book borrowing prediction methods cannot fully tap into users' potential interests and needs, resulting in insufficient accuracy in book purchasing and recommendations.

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0014] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0015] Example 1, as Figure 1 As shown, this application provides a method for predicting book borrowing trends based on big data analysis, the method comprising:

[0016] P10: Obtain multiple heterogeneous data sources, including book metadata, historical book borrowing data, user access data of electronic resources on the library digital portal, and anonymized user profile tags.

[0017] The electronic resource access data includes retrieval and browsing records of academic papers, research reports, and multimedia courses. The anonymized user profile tags are sourced from third-party platforms and do not contain directly identifying personal information.

[0018] It should be understood that the first step is to acquire multiple heterogeneous data sources. These data sources come from different channels and platforms, covering book metadata, historical book borrowing data, user access data to electronic resources on the library's digital portal, and anonymized user profile tags. The integration of these data sources is the foundation for achieving accurate prediction of book borrowing trends.

[0019] First, book metadata refers to basic information about a book, including its title, author, publisher, publication date, ISBN number, classification, and keywords. This metadata can be efficiently extracted using library management systems or other book databases and stored in a standardized format. Acquiring this data helps build the library's book catalog database and provides foundational book characteristics for subsequent analysis. These characteristics will play a crucial role when combined with user interest and behavioral data. For example, book classification information and keywords will help analyze which user groups are interested in certain categories of books, thus providing a basis for personalized recommendations and purchasing decisions.

[0020] Secondly, historical book borrowing data is collected. This data records users' borrowing behavior over a period of time, including the borrowing time, duration, frequency, and specific information about the borrowed books, such as title and author. Analyzing this historical data reveals users' preferences for different types of books and their borrowing habits. For example, if a user frequently borrows a certain type of book, it indicates a high interest in that type of book. Analyzing historical borrowing data can reveal users' long-term reading trends and provide a reference for predicting future borrowing needs. Historical borrowing data is usually stored in the library's borrowing management system and can be integrated into the analysis platform using data extraction tools.

[0021] Furthermore, electronic resource access data is also an indispensable part. This type of data includes user activity records on the library's digital portal, especially records of searching, browsing, and downloading electronic resources such as academic papers, research reports, and multimedia courses. This data can reflect users' learning, research, or career development needs. For example, frequent access to papers in a particular academic field may indicate a strong interest in that field, thus providing clues for predicting books or learning materials they might be interested in. At the same time, this data can also reveal the depth of users' knowledge in certain areas, thereby providing a more accurate basis for the library's resource recommendations and acquisition decisions.

[0022] Anonymous user profile tags are data from third-party platforms that do not contain directly identifying personal information. This data is typically provided by users when registering or participating in other platforms, such as social networks, professional networking platforms, or online education platforms. By analyzing these tags, anonymized profiles of users' interests, professional backgrounds, social behaviors, and skill levels can be constructed. This type of data provides libraries with a more comprehensive perspective on user interest analysis. For example, a user's professional tags may reflect their needs in a specific field, while social behavior tags can reveal a user's tendency to participate in community discussions, activities, and resources. By integrating these anonymized tags, user interests and potential needs that are not easily discovered through traditional borrowing data can be identified.

[0023] These data come from different channels and formats, so data preprocessing and standardization are necessary to clean and transform the raw data to ensure that data from different sources can be processed in a unified standard format. This includes removing redundant data, filling in missing values, and standardizing the formats of different data sources, providing rich data support for subsequent user interest analysis, book borrowing demand prediction, and service optimization.

[0024] P20: After preprocessing the multiple heterogeneous data sources, all unstructured data are vectorized to construct a heterogeneous information network containing user nodes, interest point nodes, and book nodes.

[0025] Furthermore, step P20 in this embodiment of the application also includes:

[0026] P21: Extract dynamic points of interest representing the user's short-term focus from the electronic resource access data; P22: Summarize static points of interest representing the user's long-term identity attributes from the anonymized user profile tags; P23: Establish book nodes based on the book metadata; P24: Establish multi-dimensional connection relationships between user nodes and the dynamic points of interest, the static points of interest, and the book nodes to generate the heterogeneous information network.

[0027] Optionally, after preprocessing the data obtained from multiple heterogeneous data sources, the next task is to vectorize all unstructured data. Unstructured data, such as text content and user behavior records, needs to be converted into numerical vector form using specific algorithms for subsequent analysis and processing. For example, natural language processing techniques such as TF-IDF and Word2Vec can be used to convert text data into vectors, or user behavior sequences can be used to generate behavior vectors, thereby transforming complex data into a form that is easy to compute and analyze.

[0028] Next, dynamic points of interest (POIs) representing users' short-term focus are extracted from electronic resource access data. Electronic resource access data typically includes users' search, browsing, and download records on the library's digital portal. These records reflect changes in users' interests and areas of focus over a period of time. By analyzing users' activities within a specific timeframe, their current short-term interests can be determined. For example, if a user has frequently searched for papers and reports related to a particular academic field or technology in the past few weeks, it can be inferred that the user has a strong short-term interest in that field. Extracting dynamic POIs depends not only on the frequency of user behavior but also on its temporality and trends. Therefore, using a time window approach for data segmentation can help capture areas of focused attention for users in the short term.

[0029] Simultaneously, static interest points (POPs) representing users' long-term identity attributes are extracted from anonymized user profile tags. These anonymized user profile tags typically originate from third-party platforms such as social networks, professional networking platforms, and online education platforms. These tags contain information about the user's professional background, interests, and skills. Unlike dynamic POPs, static POPs usually reflect a user's long-term interests, habits, and identity characteristics. For example, a user's professional background may indicate a long-term interest in a particular academic or industry field, or they may have sustained participation in certain social activities. These static POPs provide the system with clues to users' long-term needs and are crucial for constructing a deep user interest model.

[0030] Then, based on book metadata, book nodes are established. Book metadata includes basic information about each book, such as title, author, publication date, book category, and ISBN. The establishment of book nodes is fundamental to linking books with user needs and interests. Each book will be transformed into a book node through its metadata characteristics, becoming an important component of the heterogeneous information network. Through book nodes, user interests can be matched with the content characteristics of books, thereby recommending books that better suit their interests or providing a basis for library acquisition decisions.

[0031] Finally, a multi-dimensional network of connections is established between user nodes and dynamic, static, and book nodes to generate a heterogeneous information network. This involves analyzing user behavioral data and interest tags to establish these connections. For example, user nodes can be connected to their dynamic and static interest points, reflecting changes in user interests and long-term attributes; simultaneously, user nodes can be connected to book nodes, indicating the books the user has borrowed or is interested in. Through these multi-layered and multi-dimensional connections, the resulting heterogeneous information network comprehensively describes the complex interactions between users and library resources.

[0032] For example, when constructing heterogeneous information networks, methods such as graph neural networks (GNNs) can be used to perform information propagation and representation learning within the network. This network can not only help capture changes in users' interests and potential needs, but also provide support for subsequent prediction of book borrowing needs, personalized recommendations, and user behavior analysis.

[0033] Furthermore, in this embodiment, step P24 further includes establishing a multi-way connection relationship between the user node and the dynamic points of interest, the static points of interest, and the book node:

[0034] P24-1: Based on the electronic resource access data, establish a first type of connection edge between the corresponding user node and the dynamic point of interest node according to the user's access behavior to electronic resources; P24-2: Based on the anonymized user profile tags, establish a second type of connection edge between the corresponding user node and the static point of interest node according to the user's profile tags; P24-3: Based on the historical book borrowing data, establish a third type of connection edge between the corresponding user node and the book node according to the user's historical borrowing behavior; P24-4: Establish a fourth type of connection edge between the corresponding book node and the point of interest node according to the content relevance between the book theme and the point of interest.

[0035] In one possible embodiment of this application, the specific method for establishing multi-dimensional connection relationships between user nodes and dynamic points of interest, static points of interest, and book nodes can be further refined to ensure that the heterogeneous information network can more accurately reflect the complexity of user behavior and interests.

[0036] When constructing a heterogeneous information network, the relationship between various data sources and user behavior can be concretized. Different types of connections can be established based on different data and behavioral characteristics to reflect the multi-layered interaction between users and library resources. Firstly, based on electronic resource access data, a first-type connection is established between corresponding user nodes and dynamic interest point nodes according to user access behavior. Electronic resource access data records users' search and browsing behavior on the library's digital portal, reflecting their focus within a specific time period. By analyzing information such as users' search keywords, browsed resource types, and dwell time, the system can identify users' current short-term interests and establish connections between user nodes and corresponding dynamic interest point nodes. For example, if a user frequently searches for academic papers on artificial intelligence and browses related reports for extended periods, the system will identify artificial intelligence as the user's dynamic interest point and establish a first-type connection between the user node and the artificial intelligence interest point node.

[0037] Secondly, based on anonymized user profile tags, a second type of connection edge is established between the corresponding user node and the static interest point node according to the user's profile tags. Static interest points are usually related to the user's long-term attributes and identity characteristics, such as occupation, academic background, and long-term activities. By analyzing user profile tags, the user's deep-seated interests and long-term needs can be revealed, thereby establishing a stable connection relationship between user nodes and static interest points. For example, if the user profile tags show that the user is a computer engineer, the system will identify computer technology as the user's static interest point and establish a second type of connection edge between the user node and the computer technology interest point node.

[0038] Next, based on historical book borrowing data, a third type of connection edge is established between the corresponding user node and book node according to the user's historical borrowing behavior. A user's past borrowing records can reflect their preferences for certain types of books, and establishing connections between users and book nodes can help the system capture users' long-term interest trends. For example, if a user borrows books about literary classics multiple times, the system will establish a third type of connection edge between the user node and these literary classic book nodes.

[0039] Finally, based on the content relevance between the book's theme and the point of interest, a fourth type of connection edge is established between the corresponding book node and the point of interest node. Book metadata provides a description of the book's theme and content; by analyzing this information, the relevance between the book and the point of interest can be determined. For example, if a book's theme is artificial intelligence, the system will establish a fourth type of connection edge between the book node and the artificial intelligence point of interest node. By establishing these connections, the content of the book can be associated with the user's interests and needs. This connection edge helps support the library's book recommendation system by analyzing the degree of fit between book content and user interests.

[0040] P30: Perform network representation learning on the heterogeneous information network to obtain low-dimensional vector representations of all nodes.

[0041] Furthermore, step P30 in this embodiment of the application also includes:

[0042] P31: Traverse the heterogeneous information network according to a predefined semantic path to generate a node sequence rich in contextual information; P32: Train the node sequence using a neural network model to map all nodes in the heterogeneous information network to the same shared vector space to obtain a low-dimensional vector representation of all nodes.

[0043] Specifically, network representation learning is performed on heterogeneous information networks to transform the originally complex network structure into a vector form that can be computed and analyzed through dimensionality reduction. This generates low-dimensional vector representations of all nodes. These low-dimensional vectors can not only capture the feature information of each node, but also reveal the relationships between nodes, providing support for subsequent prediction, recommendation and analysis tasks.

[0044] First, the heterogeneous information network is traversed according to predefined semantic paths to generate node sequences rich in contextual information. A semantic path refers to a path pattern defined in the heterogeneous information network according to specific node types and relationship types. For example, a semantic path could be user node → dynamic point of interest node → book node, representing the path from a user to their dynamic point of interest, and then to the book associated with that point of interest. By traversing the heterogeneous information network according to these predefined semantic paths, a series of node sequences can be generated, each containing rich contextual information. This contextual information reflects the semantic relationships and interaction patterns between nodes, providing an important foundation for subsequent network representation learning.

[0045] Next, a neural network model is used to train the generated node sequence, mapping all nodes in the heterogeneous information network to the same shared vector space to obtain low-dimensional vector representations of all nodes. Neural network models, such as Graph Neural Networks (GNNs) or their variants, can automatically learn low-dimensional vector representations of nodes, allowing these vectors to preserve the semantic and structural information of the nodes within a shared vector space. Specifically, the neural network model learns the contextual information in the node sequence, mapping each node to a low-dimensional vector space. In this space, similar nodes, such as users with similar interests or books related to the same topic, are mapped to nearby locations, allowing the similarity and association between nodes to be measured by the distance or similarity between vectors. For example, if two user nodes are close in the vector space, it indicates a high degree of similarity in their points of interest or borrowing behavior.

[0046] These low-dimensional vector representations can provide rich features for subsequent tasks. For example, in book recommendation, the low-dimensional vector similarity between users and books can be used to predict books that users may be interested in; in cluster analysis, the distance between node vectors can be used to identify similar user groups or book categories.

[0047] Furthermore, in the heterogeneous information network, traversal is performed according to a predefined semantic path. Step P31 in this embodiment further includes:

[0048] P31-1: Define one or more meta-paths that can reflect semantic relationships; P31-2: Based on the guidance of the meta-paths, perform random walks on the heterogeneous information network to capture higher-order association patterns between nodes that go beyond direct adjacency.

[0049] Optionally, the process of traversing heterogeneous information networks according to predefined semantic paths can be further refined by generating node sequences rich in contextual information to capture deep-seated relationships between nodes.

[0050] First, define one or more meta-paths that reflect semantic relationships. A meta-path is a pattern used in heterogeneous information networks to describe sequences of node types and relationship types, explicitly representing the semantic relationships between different nodes. For example, in a heterogeneous information network containing users, books, and points of interest, a meta-path can be defined as "user → dynamic point of interest → book → static point of interest → user." This path represents understanding a user's potential interest in books through the interaction between the user and the point of interest. By defining these meta-paths, the semantic relationships between different node types can be explicitly captured, providing clear path guidance for subsequent network traversal.

[0051] Next, based on the defined meta-paths, random walks are performed on the heterogeneous information network to capture higher-order association patterns between nodes that go beyond direct adjacency. Random walk is a probability-based path traversal method that allows starting from a node, randomly selecting the next node according to certain rules, and continuing traversal along the path. In heterogeneous information networks, the rules of random walks are determined by the guidance of meta-paths, ensuring that the walk process follows paths that reflect semantic connections. Specifically, starting from a starting node, the next node type is randomly selected according to the definition of the meta-path, and a node is randomly selected within that type. This process is repeated until a preset sequence length is reached. For example, starting from a user node, performing a random walk according to the meta-path "user → dynamic interest point → book" might generate a sequence such as "user A → dynamic interest point B → book C". In this way, higher-order association patterns between nodes that go beyond direct adjacency can be captured. For example, user A might be indirectly associated with book C through dynamic interest point B, even though they are not directly connected in the original network.

[0052] By defining meta-paths and performing random walks, we can generate node sequences rich in contextual information. These sequences not only contain direct relationships between nodes but also capture more complex higher-order association patterns. These higher-order association patterns are significant for understanding user behavior, interest propagation, and book recommendations, and can provide rich semantic information for subsequent network representation learning.

[0053] P40: Based on the low-dimensional vector representation, clustering identification is performed to extract potential user communities with similar latent interest features, and corresponding group interest feature vectors are generated.

[0054] Furthermore, step P40 in this embodiment of the application also includes:

[0055] P41: Form a feature matrix from the low-dimensional vector representations of all user nodes; P42: Perform density clustering on the feature matrix to group users whose vectors cluster in the space into the same community; P43: Perform aggregation operation on the low-dimensional vector representations of all user nodes belonging to the same community, and use the calculation result as the group interest feature vector of the community.

[0056] It should be understood that, based on the low-dimensional vector representation of all nodes, clustering is used to extract potential user communities with similar latent interest features and generate corresponding group interest feature vectors, thereby further improving the accuracy of personalized recommendations and resource allocation.

[0057] First, a feature matrix is ​​constructed from the low-dimensional vector representations of all user nodes. Each user's low-dimensional vector representation represents their position in the low-dimensional space, containing feature information about that user across various interest dimensions. In this step, the low-dimensional vectors of all users can be arranged row-wise to form a feature matrix. Each row in the matrix represents a user's vector representation, and each column corresponds to the user's score or weight across different interest dimensions.

[0058] Next, density clustering is performed on the feature matrix to group users whose vectors cluster in the space into the same community. Density clustering is a clustering algorithm based on data point density, which can identify high-density regions in the data and group the data points in these regions into the same cluster. In this application, density clustering algorithms, such as DBSCAN, can be used to cluster user nodes in the feature matrix. Based on the distribution density of user nodes in the low-dimensional vector space, locally high-density regions of user nodes are automatically found, and users in these regions are grouped into the same community. This method can effectively identify user groups with similar latent interest characteristics, even if these groups are irregular in shape or size.

[0059] Finally, after clustering, each community contains a group of user nodes with similar interest characteristics. To generate a group interest feature vector for each community, it is necessary to perform aggregation operations on the low-dimensional vectors of all user nodes within each community. Common aggregation methods include averaging, weighted averaging, or medianing. For example, the mean of the low-dimensional vectors of all user nodes within each community can be calculated to obtain a vector representing the overall interest characteristics of the community, which serves as the feature vector for the entire community. This group interest feature vector can reflect the central interest tendency of the community, providing a basis for subsequent prediction of book borrowing demand and personalized recommendations.

[0060] P50: For any of the aforementioned potential user communities, predict their borrowing demand for books not yet in the collection based on their group interest feature vectors, and generate corresponding book purchase suggestions.

[0061] Furthermore, step P50 in this embodiment of the application also includes:

[0062] P51: Calculate the cosine similarity between the text vectors of the titles and abstracts of the books not included in the collection and the group interest feature vectors; P52: Filter out the books not included in the collection whose cosine similarity exceeds a preset threshold; P53: Generate a book purchase suggestion list with priority ranking based on the number of filtered books not included in the collection and their similarity values.

[0063] Optionally, based on the group interest feature vector of each potential user community, the similarity between the interest features of the user community and the books not yet acquired can be quantified to identify less popular books that users may be interested in, and to provide the library with priority suggestions for book acquisition.

[0064] First, the cosine similarity between the text vectors of the titles and abstracts of uncollected books and the group interest feature vectors is calculated. To assess the interest matching degree between uncollected books and the user community, the titles and abstracts of the books first need to be converted into vector representations, for example, using natural language processing techniques such as TF-IDF, Word2Vec, and BERT to convert them into text vectors. Then, the cosine similarity between the text vector of each uncollected book and the group interest feature vector of that community is calculated to measure the degree of similarity between the two. Cosine similarity is a measure of the similarity between two vectors in a direction, with a value between 0 and 1, where a value closer to 1 indicates a higher similarity. By calculating cosine similarity, the degree of matching between uncollected books and the potential user community's interests can be quantified.

[0065] Next, books not yet added to the collection are filtered based on cosine similarity scores, where the similarity exceeds a preset threshold. This threshold, set using historical data or expert experience, determines which books are sufficiently attractive to the user community. All books exceeding the threshold are considered potential demand books for the community. This filtering process narrows the scope of book recommendations, allowing subsequent recommendations to focus more on books the community is genuinely interested in, thereby improving the accuracy of purchasing decisions.

[0066] Finally, based on the number of books not selected for inclusion in the library and their similarity scores, a priority-based list of recommended book purchases is generated. First, the purchase priority of each book is determined by its similarity score. Books with higher similarity scores indicate a greater alignment with the community's interests and therefore higher priority. Simultaneously, the purchase priority is further optimized by considering the quantity of books and the intensity of community demand. For example, if there is a large demand for books in a particular field, it is recommended that the library conduct centralized purchasing based on this demand. In the final purchase recommendation list, books are ranked according to a comprehensive evaluation of similarity and demand intensity, prioritizing the purchase of those books that best meet user needs.

[0067] By following the steps above, the library can accurately predict the borrowing needs of each potential user community for books not yet in its collection, and generate a priority-based list of recommended book purchases. This not only helps the library optimize its book acquisition strategies and improve resource utilization efficiency, but also better meets the reading needs of potential user communities, enhancing the library's service quality and user experience.

[0068] Furthermore, after making the purchase based on the aforementioned book procurement recommendations, step P50 in this embodiment of the application further includes:

[0069] P54: In response to the information on newly purchased books entering the warehouse, automatically match the target potential user community with a high demand for the target new book from the heterogeneous information network; P55: Proactively push customized book recommendation information related to the target new book to the user terminals of the members of the target potential user community.

[0070] Specifically, after completing the book procurement recommendations and implementing the procurement, the next step is to further optimize the service based on the new books entering the warehouse and ensure that the new books can maximize the satisfaction of the needs of the potential user community.

[0071] First, in response to the acquisition and addition of new books to the library, the system automatically identifies potential user communities with a high demand for the target new book from a heterogeneous information network. Whenever a new book is acquired and added to the library's collection, the system automatically identifies potential user communities with a high degree of demand matching the new book from the existing heterogeneous information network, based on the book's metadata, such as title, category, topic, and abstract, and group interest feature vectors. Specifically, the system recalculates the cosine similarity between the new book's text vector and the group interest feature vectors of each potential user community. If the similarity exceeds a preset threshold, it indicates that the community has a high demand for the new book, and the system marks the community as a target potential user community, ensuring that the new book can be accurately recommended to the user groups most likely to be interested.

[0072] Next, customized book recommendations related to the target new book are proactively pushed to members of the target potential user community's user terminals. This can be done via users' devices (such as mobile phones, tablets, and computers). These recommendations will include detailed information about the new book, such as its title, author, synopsis, reasons for recommendation, highlights, relevant academic background, or application examples, to enhance the book's appeal. Push notifications can be sent via email, in-app notifications, SMS, or other suitable channels to ensure timely delivery and stimulate user interest. By proactively pushing customized book recommendations, libraries can increase user engagement and borrowing willingness, enhancing their ability to meet personalized needs.

[0073] In summary, the embodiments of this application have at least the following technical effects:

[0074] This application constructs a heterogeneous information network and performs network representation learning to accurately identify the latent interest characteristics of potential user communities, thereby accurately predicting their borrowing tendencies for books not yet in the library's collection, significantly improving the accuracy of library services. Based on the prediction results, it generates a scientific list of recommended book purchases and prioritizes them according to demand, helping libraries optimize purchasing decisions, improve resource utilization efficiency, and avoid resource waste. By proactively pushing customized book recommendations to potential user communities, it enhances interaction between users and the library, increases user participation and satisfaction, and promotes knowledge dissemination and cultural exchange. Finally, it provides libraries with data-driven decision support tools, enhancing their capabilities in resource planning, service optimization, and user interaction, and promoting the intelligent and personalized development of library services.

[0075] It achieves the technical effect of accurately identifying potential user needs and improving the accuracy of borrowing demand prediction by integrating big data analysis with heterogeneous data sources.

[0076] Example 2 is based on the same inventive concept as the book borrowing trend prediction method based on big data analysis in the previous examples, such as... Figure 2 As shown, this application provides a book borrowing trend prediction system based on big data analysis. The system and method embodiments in this application are based on the same inventive concept. The system includes:

[0077] The heterogeneous data source acquisition module 11 is used to acquire multiple heterogeneous data sources, including book metadata, historical book borrowing data, user access data of electronic resources on the library digital portal, and anonymized user profile tags.

[0078] The data vectorization module 12 is used to preprocess the multiple heterogeneous data sources and then vectorize all unstructured data to construct a heterogeneous information network containing user nodes, interest point nodes, and book nodes.

[0079] The network representation learning module 13 is used to perform network representation learning on the heterogeneous information network to obtain low-dimensional vector representations of all nodes.

[0080] The clustering identification module 14 is used to perform clustering identification based on the low-dimensional vector representation, extract potential user communities with similar latent interest features, and generate corresponding group interest feature vectors.

[0081] The borrowing demand prediction module 15 is used to predict the borrowing demand tendency of any potential user community for books not yet in the collection based on its group interest feature vector, and generate corresponding book purchase suggestions.

[0082] Furthermore, in the heterogeneous data source acquisition module 11:

[0083] The electronic resource access data includes retrieval and browsing records of academic papers, research reports, and multimedia courses. The anonymized user profile tags are sourced from third-party platforms and do not contain directly identifying personal information.

[0084] Furthermore, the data vectorization module 12 is also used to perform the following steps:

[0085] From the electronic resource access data, dynamic points of interest representing the user's short-term focus are extracted; from the anonymized user profile tags, static points of interest representing the user's long-term identity attributes are summarized; book nodes are established based on the book metadata; and a multi-dimensional connection relationship is established between user nodes and the dynamic points of interest, the static points of interest, and the book nodes to generate the heterogeneous information network.

[0086] Furthermore, the data vectorization module 12 is also used to perform the following steps:

[0087] Based on the electronic resource access data, a first type of connection edge is established between the corresponding user node and the dynamic point of interest node according to the user's access behavior to electronic resources; based on the anonymized user profile tags, a second type of connection edge is established between the corresponding user node and the static point of interest node according to the user's profile tags; based on the historical book borrowing data, a third type of connection edge is established between the corresponding user node and the book node according to the user's historical borrowing behavior; and based on the content relevance between the book theme and the point of interest, a fourth type of connection edge is established between the corresponding book node and the point of interest node.

[0088] Furthermore, the network representation learning module 13 is also used to perform the following steps:

[0089] The heterogeneous information network is traversed according to a predefined semantic path to generate a node sequence rich in contextual information; the node sequence is trained using a neural network model to map all nodes in the heterogeneous information network to the same shared vector space, thereby obtaining a low-dimensional vector representation of all nodes.

[0090] Furthermore, the network representation learning module 13 is also used to perform the following steps:

[0091] Define one or more meta-paths that can reflect semantic relationships; and perform random walks on the heterogeneous information network according to the guidance of the meta-paths to capture higher-order association patterns between nodes that go beyond direct adjacency.

[0092] Furthermore, the clustering identification module 14 is also used to perform the following steps:

[0093] The low-dimensional vector representations of all user nodes are used to form a feature matrix; density clustering is performed on the feature matrix to group users whose vectors are clustered in the space into the same community; aggregation operation is performed on the low-dimensional vector representations of all user nodes belonging to the same community, and the calculation result is used as the group interest feature vector of the community.

[0094] Furthermore, the borrowing demand prediction module 15 is also used to perform the following steps:

[0095] Calculate the cosine similarity between the text vectors of the titles and abstracts of the books not yet included in the collection and the group interest feature vectors; filter out the books not yet included that have a cosine similarity exceeding a preset threshold; and generate a book purchase recommendation list with priority ranking based on the number of the filtered books and their similarity values.

[0096] Furthermore, after purchasing books based on the aforementioned book purchasing recommendations, the borrowing demand prediction module 15 is also used to perform the following steps:

[0097] In response to the information on newly purchased books entering the warehouse, the system automatically matches potential user communities with a high demand for the target new books from the heterogeneous information network; and proactively pushes customized book recommendation information related to the target new books to the user terminals of the members of the target potential user community.

[0098] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0099] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0100] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.

Claims

1. A method for predicting book borrowing trends based on big data analysis, characterized in that, The method includes: Multiple heterogeneous data sources are acquired, including book metadata, historical book borrowing data, user access data of electronic resources on the library digital portal, and anonymized user profile tags; After preprocessing the multiple heterogeneous data sources, all unstructured data are vectorized to construct a heterogeneous information network containing user nodes, interest point nodes, and book nodes. Network representation learning is performed on the heterogeneous information network to obtain low-dimensional vector representations of all nodes; Clustering identification is performed based on the low-dimensional vector representation to extract potential user communities with similar latent interest features and generate corresponding group interest feature vectors. For any of the aforementioned potential user communities, based on their group interest feature vectors, predict their borrowing demand for books not yet in the collection, and generate corresponding book purchasing suggestions. After preprocessing the multiple heterogeneous data sources, all unstructured data are vectorized to construct a heterogeneous information network containing user nodes, point-of-interest nodes, and book nodes, including: From the electronic resource access data, dynamic points of interest that represent the user's short-term focus can be extracted; From the anonymized user profile tags, static interest points that characterize the user's long-term identity attributes are summarized; Book nodes are established based on the aforementioned book metadata; Establish multi-dimensional connection relationships between user nodes and the dynamic points of interest, the static points of interest, and the book nodes to generate the heterogeneous information network; The anonymized user profile tags come from data from third-party platforms, and these tags do not contain directly identifying personal information. This data is provided by users when they register or participate in other platforms. Network representation learning is performed on the heterogeneous information network to obtain low-dimensional vector representations of all nodes, including: The heterogeneous information network is traversed according to a predefined semantic path to generate a node sequence rich in contextual information. The node sequence is trained using a neural network model, and all nodes in the heterogeneous information network are mapped to the same shared vector space to obtain low-dimensional vector representations of all nodes. Traversing the heterogeneous information network according to a predefined semantic path includes: Define one or more meta-paths that reflect semantic relationships; Guided by the meta-path, a random walk is performed on the heterogeneous information network to capture higher-order association patterns between nodes that go beyond direct adjacency. Clustering identification is performed based on the low-dimensional vector representation to extract potential user communities with similar latent interest features, and corresponding group interest feature vectors are generated, including: The low-dimensional vector representations of all user nodes are used to form a feature matrix; Density clustering is performed on the feature matrix to group users whose vectors cluster in the space into the same community; Aggregate the low-dimensional vector representations of all user nodes belonging to the same community, and use the result as the group interest feature vector of the community.

2. The book borrowing trend prediction method based on big data analysis as described in claim 1, characterized in that, After making the purchase based on the aforementioned book purchasing recommendations, the following is included: In response to the new book purchase information, the system automatically matches potential user communities with a high demand for the target new book from the heterogeneous information network. Customized book recommendations related to the target new book are proactively pushed to the user terminals of members of the target potential user community.

3. The book borrowing trend prediction method based on big data analysis as described in claim 1, characterized in that, Establishing a multi-way connection relationship between user nodes and the dynamic points of interest, the static points of interest, and the book nodes includes: Based on the electronic resource access data, according to the user's access behavior to the electronic resources, a first type of connection edge is established between the corresponding user node and the dynamic point of interest node. Based on the anonymized user profile tags, a second type of connection edge is established between the corresponding user node and the static interest point node according to the user's profile tags. Based on the historical book borrowing data, a third type of connection edge is established between the corresponding user node and the book node according to the user's historical borrowing behavior; Based on the content relevance between the book's theme and the point of interest, a fourth type of connection edge is established between the corresponding book node and the point of interest node.

4. The book borrowing trend prediction method based on big data analysis as described in claim 1, characterized in that, For any of the aforementioned potential user communities, based on their group interest feature vectors, predict their borrowing demand for books not yet in the library, and generate corresponding book purchasing suggestions, including: Calculate the cosine similarity between the text vectors of the titles and abstracts of the books not included in the collection and the group interest feature vectors; Uncollected books whose cosine similarity exceeds a preset threshold are filtered out. Based on the number of books not selected for inclusion and their similarity scores, a list of recommended book purchases is generated, including a priority ranking.

5. The book borrowing trend prediction method based on big data analysis as described in claim 1, characterized in that, The electronic resource access data includes retrieval and browsing records of academic papers, research reports, and multimedia courses. The anonymized user profile tags are sourced from third-party platforms and do not contain any directly identifying personal information.

6. A book borrowing trend prediction system based on big data analysis, employing a book borrowing trend prediction method based on big data analysis as described in any one of claims 1-5, characterized in that, The system includes: The heterogeneous data source acquisition module is used to acquire multiple heterogeneous data sources, including book metadata, historical book borrowing data, user access data of electronic resources on the library digital portal, and anonymized user profile tags. The data vectorization module is used to preprocess the multiple heterogeneous data sources and then vectorize all unstructured data to construct a heterogeneous information network containing user nodes, interest point nodes, and book nodes. The network representation learning module is used to perform network representation learning on the heterogeneous information network to obtain low-dimensional vector representations of all nodes; The clustering identification module is used to perform clustering identification based on the low-dimensional vector representation, extract potential user communities with similar latent interest features, and generate corresponding group interest feature vectors. The borrowing demand prediction module is used to predict the borrowing demand tendency of any potential user community for books not yet in the collection, based on its group interest feature vector, and generate corresponding book purchase suggestions.