Intelligent analysis method and system for multi-source book data, terminal and storage medium

By preprocessing and integrating data relationships of multi-source book data, generating knowledge graphs, generating user portraits and predicting borrowing trends, the problem of poor integration of multi-source data in the library system is solved, accurate analysis of user behavior and borrowing trends is achieved, and book resource allocation and user experience are optimized.

CN120471357AActive Publication Date: 2025-08-12COMMUNICATION UNIVERSITY OF CHINA

Patent Information

Application Number
CN202510553168.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

When dealing with the collaborative work of multiple departments, the existing library integrated management system faces the problem of poor integration of multi-source data, and cannot realize user behavior analysis of multi-source book data and book borrowing trend analysis, resulting in difficulty in allocating book resources and unable to meet user needs.

Method used

By obtaining multi-source book data, preprocessing, data relationship integration, target knowledge graphs are generated, user portrait generation and book borrowing trend prediction, and finally visual display is carried out, using machine learning and big data analysis technology to deeply explore reader behavior and resource utilization.

Benefits of technology

It realizes accurate analysis of user behavior and book borrowing trends, optimizes the allocation of book resources, and improves resource utilization and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471357A_ABST
    Figure CN120471357A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent analysis method and system for multi-source book data, a terminal and a storage medium, and the method comprises the steps: obtaining the multi-source book data, and carrying out the preprocessing of the multi-source book data, and obtaining the target multi-source book data; performing data relation integration processing on the target multi-source book data to obtain a target knowledge graph; performing user portrait generation processing and book borrowing trend prediction processing according to the target knowledge graph to obtain a user portrait result and a book borrowing trend result; and visually displaying the target knowledge graph, the user portrait result and the book borrowing trend result. According to the method, the target knowledge graph is obtained by performing preprocessing and data relation integration processing on the multi-source book data, and user portrait generation processing and book borrowing trend prediction processing are performed through the target knowledge graph, so that accurate analysis on user behaviors and book borrowing trends can be realized, and the user experience is improved. And thus, optimized allocation of book resources is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an intelligent analysis method, system, terminal and computer-readable storage medium for multi-source book data. Background Art

[0002] Libraries are institutions that collect, organize and store books and materials for people to read and refer to. With the development of society, people's demand for reading is increasing. Therefore, libraries have become the target of more and more people to borrow books.

[0003] However, when dealing with multi-departmental collaboration, existing library integrated management systems often face the problem of poor integration of multi-source data from various departments. They are unable to perform user behavior analysis of multi-source book data and analysis of book borrowing trends, which leads to difficulties in allocating book resources and cannot meet user needs.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide an intelligent analysis method, system, terminal and computer-readable storage medium for multi-source book data, aiming to solve the problem that the library integrated management system in the existing technology is usually faced with the problem of poor integration of multi-source data from various departments when dealing with the collaborative work of multiple departments, and is unable to realize user behavior analysis of multi-source book data and analysis of book borrowing trends, which leads to difficulties in allocating book resources and cannot meet user needs.

[0006] To achieve the above object, the present invention provides an intelligent analysis method for multi-source book data, which includes the following steps:

[0007] Acquire multi-source book data, and pre-process the multi-source book data to obtain target multi-source book data;

[0008] Performing data relationship integration processing on the target multi-source book data to obtain a target knowledge graph;

[0009] Perform user portrait generation processing and book borrowing trend prediction processing according to the target knowledge graph to obtain user portrait results and book borrowing trend results;

[0010] The target knowledge graph, the user portrait results and the book borrowing trend results are visualized.

[0011] Optionally, the intelligent analysis method for multi-source book data, wherein the step of acquiring multi-source book data and pre-processing the multi-source book data to obtain target multi-source book data, specifically includes:

[0012] Using a distributed log collection tool to acquire multi-source book data, wherein the multi-source book data includes structured data, unstructured data, real-time data streams, and external data sources;

[0013] Performing data standardization processing on the multi-source book data using metadata standards to obtain standardized multi-source book data;

[0014] Using a data mapping tool to perform data mapping processing on the standardized multi-source book data to obtain mapped multi-source book data;

[0015] Data cleaning is performed on the mapped multi-source book data to obtain target multi-source book data.

[0016] Optionally, in the intelligent analysis method for multi-source book data, the data cleaning process includes data deduplication, missing value filling, formatting and outlier processing;

[0017] The data cleaning process of the mapped multi-source book data to obtain target multi-source book data specifically includes:

[0018] Acquiring data identifiers in the mapped multi-source book data, determining duplicate data in the mapped multi-source book data according to the data identifiers, and performing the data deduplication process on the duplicate data to obtain first multi-source book data;

[0019] Obtaining missing values in the first multi-source book data, and performing missing value filling processing on the missing values using a preset filling method to obtain second multi-source book data, wherein the preset filling method includes an interpolation method, a mean filling method, and a missing value deletion method;

[0020] Acquiring a date format corresponding to the second multi-source book data, and performing the formatting process on the second multi-source book data according to the date format to obtain third multi-source book data;

[0021] An outlier detection algorithm is used to identify outliers on the third multi-source book data to obtain target outlier data, and the target outlier data is processed to obtain target multi-source book data.

[0022] Optionally, the intelligent analysis method for multi-source book data, wherein the step of performing data relationship integration processing on the target multi-source book data to obtain a target knowledge graph, specifically includes:

[0023] Constructing a centralized data model, and performing entity relationship structuring processing on the target multi-source book data through the centralized data model to obtain structured multi-source book data;

[0024] Determining a core entity in the structured multi-source book data, and performing identifier assignment processing and fuzzy matching processing on the core entity to obtain associated multi-source book data;

[0025] Entity relationships are extracted from the associated multi-source book data using natural language processing and semantic technology to obtain entity relationship extraction results, and a target knowledge graph is constructed based on the entity relationship extraction results.

[0026] Optionally, the intelligent analysis method for multi-source book data, wherein the user portrait generation processing and book borrowing trend prediction processing are performed based on the target knowledge graph to obtain user portrait results and book borrowing trend results, specifically includes:

[0027] Performing feature extraction and cluster analysis on the target knowledge graph to obtain feature extraction results and user behavior patterns, and performing user portrait generation based on the feature extraction results and the user behavior patterns to obtain user portrait results;

[0028] Obtaining time series data in the target knowledge graph, and performing aggregation processing on the time series data to obtain aggregated time series data;

[0029] A preset deep learning model is determined, and the aggregated time series data is input into the preset deep learning model to output a book borrowing trend result.

[0030] Optionally, the intelligent analysis method for multi-source book data, wherein the step of performing feature extraction and cluster analysis on the target knowledge graph to obtain feature extraction results and user behavior patterns, and performing user portrait generation based on the feature extraction results and the user behavior patterns to obtain user portrait results, specifically includes:

[0031] Performing feature extraction processing on the target knowledge graph to obtain feature extraction results, wherein the feature extraction results include basic features, behavioral features, and semantic features;

[0032] Performing cluster analysis on the feature extraction results using a cluster analysis algorithm to obtain user behavior patterns;

[0033] Determine a pre-trained model, and perform model training and model fine-tuning on the pre-trained model according to the target knowledge graph and the user behavior pattern to obtain an end-to-end portrait model;

[0034] Obtain user reading data corresponding to the target user, input the user reading data into the end-to-end portrait model, and output the user portrait result.

[0035] Optionally, the intelligent analysis method for multi-source book data, wherein the user portrait generation process and the book borrowing trend prediction process are performed based on the target knowledge graph to obtain the user portrait result and the book borrowing trend result, further comprises:

[0036] Determine the corresponding recommended book data and predicted decision direction based on the user portrait result, and push the recommended book data and the predicted decision direction to the corresponding target user;

[0037] A current book resource allocation plan is determined, and the current book resource allocation plan is optimized according to the book borrowing trend result to obtain a target book resource allocation result.

[0038] In addition, to achieve the above-mentioned purpose, the present invention further provides an intelligent analysis system for multi-source book data, wherein the intelligent analysis system for multi-source book data comprises:

[0039] A data preprocessing module, configured to obtain multi-source book data and preprocess the multi-source book data to obtain target multi-source book data;

[0040] A data relationship integration module is used to perform data relationship integration processing on the target multi-source book data to obtain a target knowledge graph;

[0041] A portrait generation and trend prediction module is used to generate user portraits and predict book borrowing trends based on the target knowledge graph to obtain user portrait results and book borrowing trend results;

[0042] A visualization display module is used to visualize the target knowledge graph, the user portrait results and the book borrowing trend results.

[0043] In the present invention, multi-source book data is obtained, and the multi-source book data is pre-processed to obtain target multi-source book data; data relationship integration processing is performed on the target multi-source book data to obtain a target knowledge graph; user portrait generation processing and book borrowing trend prediction processing are performed based on the target knowledge graph to obtain user portrait results and book borrowing trend results; the target knowledge graph, the user portrait results and the book borrowing trend results are visualized. The present invention obtains a target knowledge graph by pre-processing and data relationship integration processing on multi-source book data, and performs user portrait generation processing and book borrowing trend prediction processing through the target knowledge graph. It can achieve accurate analysis of user behavior and book borrowing trends, and then according to the generated user portrait results and book borrowing trend results, it can achieve optimal allocation of book resources, which not only effectively improves the availability of book resources, but also optimizes the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flow chart of a preferred embodiment of the intelligent analysis method of multi-source book data of the present invention;

[0045] Figure 2 Schematic diagram of the overall structure of a preferred embodiment of the intelligent analysis method for multi-source book data of the present invention;

[0046] Figure 3 Schematic diagram of the model architecture of a preferred embodiment of the intelligent analysis method for multi-source book data of the present invention;

[0047] Figure 4 It is a structural diagram of a preferred embodiment of the intelligent analysis system for multi-source book data of the present invention;

[0048] Figure 5 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0050] Currently, the existing technologies used in traditional libraries, such as the Integrated Library Systems (ILS), RFID (Radio-frequency identification) technology, and digital resource management systems, have the following major shortcomings and challenges: 1. Data silo problem: There is a lack of effective data interaction mechanisms between the Integrated Library Management System (ILS) and the Digital Resource Management System (DRMS), resulting in data silos. 2. Poor system compatibility: Different systems may adopt different technical architectures (such as monolithic architecture and microservice architecture) and databases (such as MySQL and MongoDB), resulting in poor compatibility between systems and difficulties in data exchange. 3. Lack of data analysis and visualization tools: There is a lack of powerful data analysis tools (such as Apache Spark and Tableau) to conduct in-depth mining and multi-dimensional analysis of integrated data. 4. Insufficient application of visualization tools (such as Power BI and D3.js) makes it impossible to generate intuitive analysis reports or visualize knowledge graphs.

[0051] To address these issues, the present invention proposes an intelligent analysis method for multi-source book data. This method utilizes advanced database technology, supporting multiple databases such as MySQL, PostgreSQL, and MongoDB, ensuring efficient and scalable data storage. The data acquisition module utilizes a flexible interface to ensure real-time data import and accurate recording. The system's statistical analysis capabilities leverage big data technology to conduct in-depth analysis of borrowing, reader behavior, and resource usage. Graphical presentation utilizes visualization tools such as ECharts and D3.js, providing a rich set of reports and dynamic charts.

[0052] This invention builds an intelligent library data analysis system, integrates multi-source data, and uses machine learning, big data analysis and artificial intelligence technologies to deeply explore multi-dimensional data such as reader behavior, borrowing trends and resource utilization, providing accurate analysis results and personalized services. At the same time, through real-time monitoring and visual display, it supports data-driven decision optimization, which can effectively improve the library's management efficiency, service quality and resource utilization, meet readers' needs and support scientific research and teaching development.

[0053] The intelligent analysis method of multi-source book data described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the intelligent analysis method of multi-source book data includes the following steps:

[0054] Step S10: Acquire multi-source book data, and pre-process the multi-source book data to obtain target multi-source book data.

[0055] The multi-source book data in the present invention includes structured data, unstructured data, real-time data streams and external data sources, which come from different departments and different systems. The present invention collects these data in different formats through data integration. The multi-source book data obtained after the integration can also be called heterogeneous data sources or multimodal data sources.

[0056] The following problems exist with existing library system data: 1. Data integration is difficult: Existing integrated library management systems (ILS) often face problems with data integration when handling the collaborative work of multiple departments such as the Resource Development Department, Loan Services Department, Subject Services Department, and Special Collections Department. This leads to data silos and makes it impossible to form a complete knowledge graph or generate high-quality analytical reports. Specific technical challenges and problems include: a. Inconsistent data formats across heterogeneous systems: The Resource Development Department, Loan Services Department, Subject Services Department, and Special Collections Department may use different data standards and formats (such as MARC (Machine-Readable Cataloging), Dublin Core, and MODS (Metadata Object Description Schema), etc.), making direct data integration difficult. b. The lack of unified metadata standards for data exchange between the ILS and the Digital Resource Management System (DRMS) makes data mapping and conversion difficult. c. Incompatible or limited API interfaces: The API interfaces of different systems (such as RESTful API, SOAP) may use different protocols or data formats (such as JSON, XML), and lack unified interface specifications. The API functions of some systems are limited and cannot support complex data queries or batch operations, resulting in inefficient data synchronization. d. Insufficient data synchronization and real-time performance: Data synchronization between the integrated library management system and the digital resource management system usually relies on scheduled batch updates rather than real-time synchronization, resulting in data delays and inconsistencies. The lack of a real-time data stream processing mechanism based on message queues (such as Kafka, RabbitMQ) makes it impossible to achieve real-time data updates across systems. 2. Insufficient data analysis accuracy: Existing analysis methods often lack depth and detail, making it difficult to accurately analyze multi-dimensional data such as reader behavior and borrowing trends, limiting decision support functions. Specific technical challenges and problems include: a. Insufficient data collection and cleaning: Data collection tools (such as Flume, Logstash) fail to fully cover reader behavior data (such as borrowing records, search history, length of stay, etc.), resulting in incomplete data sources. Data cleaning technologies (such as Python-based Pandas and OpenRefine) are insufficiently applied, failing to effectively address noise, missing values, and outliers in the data, impacting the accuracy of subsequent analysis. b. Limited multi-dimensional data analysis capabilities: The lack of correlation analysis capabilities for multi-dimensional data (such as time, space, user attributes, and resource type) makes it difficult to reveal complex borrowing trends and reader behavior patterns. Existing analytical tools (such as Excel and SPSS) lack performance when processing large-scale, high-dimensional data and cannot support complex analytical tasks.c. Insufficient Application of Machine Learning Models: The lack of in-depth application of machine learning algorithms (such as cluster analysis, classification algorithms, and regression analysis) makes it impossible to discover underlying patterns in massive amounts of data. The inadequate application of deep learning techniques (such as LSTM (Long Short-Term Memory) networks, Transformers, and self-attention neural networks) in time series analysis (such as borrowing trend prediction) results in low prediction accuracy. 3. Poor Visualization: While existing visualization tools can provide basic statistical displays, they lack interactivity and dynamism, making it difficult to meet users' personalized needs. Specific technical challenges and issues include: a. Limited Functionality of Visualization Tools: Existing visualization tools (such as Tableau and Power BI) primarily support static charts (such as bar charts, pie charts, and line charts) and lack support for complex data (such as network diagrams, heat maps, and geospatial data). They also lack the ability to visualize dynamic data (such as real-time borrowing data and user behavior flow data), making it impossible to reflect data changes in real time. b. Insufficient Interactivity: Existing visualization tools have limited interactive features, preventing users from deeply exploring data through dragging, zooming, filtering, and other operations. There is a lack of web-based interactive visualization frameworks (such as D3.js and ECharts), which makes it impossible to provide rich user interaction experiences (such as dynamic filtering and linkage analysis).

[0057] The present invention integrates and standardizes the data of various departments (such as the Resource Development Department, the Borrowing Service Department, the Subject Service Department, the Special Collection Development Department) and systems (such as the Book Integrated Management System and the Digital Resource Management System) through a unified platform, which can eliminate data silos and realize information sharing and interaction. Specific technical implementations include: 1. Unified data standards and metadata management: Use internationally accepted metadata standards (such as MARC, Dublin Core, MODS) to standardize the data of various departments to ensure the consistency of data formats. Use metadata management tools (such as OpenRefine, CatMDEdit) to clean, convert and map metadata to solve the problem of data heterogeneity. 2. Data warehouse and ETL technology: Build a unified data warehouse (such as Hadoop, Snowflake, Amazon Redshift) to centrally store data from various departments and systems. Use ETL (Extract, Transform, Load) tools (such as Apache NiFi, Talend, Informatica) to extract, transform and load data to ensure efficient integration and synchronization of data. 3. API Interface and Microservice Architecture: Provide a unified RESTful API or GraphQL interface to support data exchange between departments and systems. Use a microservice architecture (such as Spring Cloud and Kubernetes) to achieve modularity and loose coupling of the system, facilitating data sharing and functional expansion.

[0058] The system processing flow of the present invention is as follows: 1. Automatic cleaning of data from multiple data sources: data source docking includes crawlers, text table file import, multi-database interface extraction, and three-party system interface docking. 2. Data hierarchical processing relationship integration: hierarchical classification is performed according to data attributes, and data relationships are generated according to upper and lower inclusion relationships and left and right correlation relationships. 3. Intelligent generation of digital model portraits: portraits of different dimensions such as books, readers, and classrooms are automatically generated. The portraits can clearly show their respective reading and management directions. The portraits can be further intelligently analyzed according to the content of the portraits, and appropriate decision-making directions can be given to different portraits. 4. Progressive digital cockpit presentation: progressive presentation from coarse to fine management granularity. Each granularity has manual scheduling and processing capabilities, which is convenient for administrators to arrange tasks and manage books. The main page is mainly composed of analysis reports and data analysis statistical results of each dimension, and the progressive page is drilled down in the form of a relationship map (wherein, drill-down display means that in data analysis, the next layer of data is expanded from the current data to perform more detailed analysis).

[0059] Specifically, a distributed log collection tool is used to obtain multi-source book data, wherein the multi-source book data includes structured data, unstructured data, real-time data streams and external data sources; metadata standards are used to perform data standardization processing on the multi-source book data to obtain standardized multi-source book data; and a data mapping tool is used to perform data mapping processing on the standardized multi-source book data to obtain mapped multi-source book data.

[0060] The first step is to deploy the environment (database deployment). This includes: 1. Selecting a database suitable for the project, such as MySQL, Radis, PostgreSQL, or MongoDB. 2. Installing and configuring the database: Optimizing parameters to ensure high performance and security. 3. Implementing data backup and driver processing: Enhance the system's fault tolerance through up-down maintenance or clustered drivers.

[0061] like Figure 2 As shown, set up your development environment: 1. Install development tools: IDE (Integrated Development Environment) such as VS Code or WebStorm, and configure relevant plugins. 2. Implement version control and collaboration: Use Git and resource management platforms such as GitHub or GitLab to configure a comprehensive development and testing environment to ensure thorough application testing and automated deployment. By employing machine learning and big data analytics, the accuracy and depth of analysis of borrowing and reader behavior can be significantly improved, providing more valuable support for decision-making. Specific technical implementations include: a. Data collection and preprocessing: Use distributed log collection tools (such as Flume and Logstash) to comprehensively collect reader behavior data (such as borrowing records, search history, and page clickstreams). b. Data source access functions include: Data source identification and classification: Structured data (accessing relational databases (such as MySQL and PostgreSQL) through JDBC / ODBC interfaces to read structured data such as borrowing records and reader information). This includes borrowing records (MySQL / Oracle), reader information (LDAP / identity system), and resource metadata (MARC format). Semi-structured data: log files (access logs, retrieval logs, JSON / CSV format), open API data (such as third-party database interfaces); unstructured data: use APIs or crawler technology to access unstructured data (such as reader comments, document abstracts, web page content, and audio and video resources); real-time data streams: access real-time data (such as borrowing operations, retrieval logs) through message queues (such as Kafka and RabbitMQ); external data sources: access external data through open APIs (such as third-party databases and scientific research platforms) to expand analysis dimensions.

[0062] like Figure 2 As shown, data collection technology options include batch data acquisition, real-time data stream acquisition, and API integration. 1. Batch data acquisition: ETL tools (Apache NiFi, Talend): Periodically extract borrowing records from relational databases (such as MySQL). File import: Process CSV / Excel files through scripts (Python Pandas) or tools (Sqoop). 2. Real-time data stream acquisition: Message queues (Apache Kafka, RabbitMQ): Capture real-time events (such as borrowing operations and retrieval requests). Stream processing frameworks (Apache Flink, Spark Streaming): Real-time parsing of log stream data (such as Nginx access logs). 3. API integration: RESTful API / GraphQL: Connect to external systems (such as e-book platforms and scientific research databases) to obtain metadata. Crawler technology (Scrapy, Selenium): Capture public data (such as academic paper citation information).

[0063] Data access technology options include ETL tools, APIs, and file import. ETL tools (such as Apache NiFi and Talend) are used to extract, transform, and load data, supporting the integration of multi-source data. APIs (such as RESTful APIs or GraphQL) are provided to support data exchange between external systems and the library system. File imports (such as CSV, Excel, and JSON) are supported for batch import, facilitating access to historical data.

[0064] Data standardization and mapping include metadata management and data mapping. Metadata management: Standardize multi-source data using metadata standards (such as MARC and Dublin Core). Data mapping: Use data mapping tools (such as OpenRefine) to uniformly map fields from different data sources to the target model.

[0065] Real-time data access includes stream processing frameworks, webhooks, and message queues. Stream processing frameworks use stream processing technologies (such as Apache Flink and Spark Streaming) to access and process data streams in real time. Webhooks and message queues capture system events (such as borrowing and returning) in real time through webhooks or message queues (such as Kafka).

[0066] Obtain data identifiers in the mapped multi-source book data, determine duplicate data in the mapped multi-source book data according to the data identifiers, and perform the data deduplication processing on the duplicate data to obtain first multi-source book data; obtain missing values in the first multi-source book data, and use a preset filling method to perform the missing value filling processing on the missing values to obtain second multi-source book data, wherein the preset filling method includes interpolation method, mean filling method and missing value deletion method; obtain the date format corresponding to the second multi-source book data, and perform the formatting processing on the second multi-source book data according to the date format to obtain third multi-source book data; perform outlier identification on the third multi-source book data through an outlier detection algorithm to obtain target outlier data, and perform the outlier processing on the target outlier data to obtain target multi-source book data.

[0067] The present invention uses data cleaning tools (such as Pandas, OpenRefine) to deduplicate data, fill missing values and process outliers to ensure data quality. Among them, the data cleaning functions are as follows: 1. The data cleaning process includes: Data exploration: Use data exploration tools (such as Pandas Profiling) to analyze data distribution, missing values, outliers and other issues. Data deduplication: Remove duplicate records through unique identifiers (such as ISBN, reader ID). Missing value processing: Use interpolation, mean filling or deletion of missing values to process incomplete data. Outlier processing: Identify and process outliers through statistical methods (such as Z-score, IQR) or machine learning algorithms (such as isolation forest). Format standardization: Unify the formats of fields such as date, time, and text (such as YYYY-MM-DD, UTF-8 encoding). Among them, examples of deduplication and completion in data cleaning are as follows: use Pandas or OpenRefine to process duplicate borrowing records and fill in missing fields (such as reader department information); format unification: standardize different date formats (such as "2023-10-01" and "01 / 10 / 2023") to ISO 8601. Outlier processing: detect abnormal borrowing durations (such as borrowing records exceeding 1 year) through Z-score or IQR methods. 2. Data cleaning technologies include: Rule engine: automatically cleans data based on predefined rules (such as the borrowing date cannot be later than the return date). Machine learning: uses clustering and classification algorithms to identify noise and anomalies in the data. Natural language processing (NLP): performs word segmentation, stop word removal, stemming, and other processing on text data (such as reader comments and literature abstracts). The application and effects of model algorithm cleaning technology in library intelligent analysis systems are as follows: 1. Purpose: Clean and preprocess raw data (such as borrowing records, reader behavior data) to remove noise, missing values, and outliers. Standardize data formats, ensure data quality, and provide reliable input for subsequent analysis. 2. Effect: Improve the accuracy and reliability of data analysis, optimize the training effect of machine learning models, improve prediction accuracy (such as borrowing trends and reader preferences), reduce the risk of incorrect decisions, and provide higher-quality data support for library management and service optimization.

[0068] In summary, the present invention discloses a cross-modal model algorithm cleaning technology, the core uses of which include:

[0069] 1. Multi-source data fusion and cleaning: This is used to uniformly process heterogeneous data such as text (such as user reviews), images (such as book images), and time series data (such as logs), eliminating noise and conflicts (such as inconsistencies between book descriptions and images). It is also used to establish semantic associations through cross-modal alignment (such as through the CLIP model) to ensure data consistency (such as matching the sentiment of user reviews with the content of images).

[0070] 2. Intelligent quality repair: This is used for automatic error correction, such as using BERT / Transformer to fix typos (improving text accuracy) and GANs to repair blurry images (improving PSNR). It is also used for anomaly detection: identifying fake reviews or stolen images through multimodal joint modeling (e.g., text + image). c. Feature enhancement: This uses cross-modal representation learning (e.g., ViLBERT) to generate unified feature vectors, improving the performance of downstream tasks.

[0071] The actual effects include:

[0072] 1. Data quality has been improved. After cleaning, data noise has been significantly reduced, and the completeness of key fields has been significantly improved. 2. The accuracy of identifying multimodal conflicts (such as image and text mismatches) has been improved. 3. Model performance has been optimized. Cross-modal pre-training has effectively improved the average F1 score of NLP / CV tasks. In small sample scenarios, data augmentation is more effective than single-modality data augmentation. 4. Image-text matching has been improved after cleaning; the accuracy of user profile prediction that integrates text, images, and behavior has been improved.

[0073] Step S20: Perform data relationship integration processing on the target multi-source book data to obtain a target knowledge graph.

[0074] Specifically, a centralized data model is constructed, and entity relationship structuring processing is performed on the target multi-source book data through the centralized data model to obtain structured multi-source book data; core entities in the structured multi-source book data are determined, and identifier allocation processing and fuzzy matching processing are performed on the core entities to obtain associated multi-source book data; entity relationship extraction is performed on the associated multi-source book data through natural language processing and semantic technology methods to obtain entity relationship extraction results, and a target knowledge graph is constructed based on the entity relationship extraction results.

[0075] Data Relationship Integration: 1. Data Standardization and Unified Modeling, including: a. Metadata Management: Adopting international standards (such as MARC and Dublin Core) to unify metadata definitions across different data sources, resolving issues with inconsistent field naming and formatting. b. Unified Data Model: Designing a centralized data model (such as a star schema or graph schema) to structure the relationships between entities such as readers, resources, loan records, and subject classifications.

[0076] 2. Entity resolution and association include: a. Unique identifiers: Assign unique IDs (e.g., reader ID, ISBN) to core entities such as readers and resources to enable entity matching across data sources. b. Fuzzy matching techniques: Use string similarity algorithms (e.g., Levenshtein distance) or machine learning models (e.g., Siamese networks) to associate similar but not identical records (e.g., different spellings of author names).

[0077] 3. Graph Databases and Knowledge Graphs: a. Graph Databases (e.g., Neo4j, TigerGraph): Store complex relationships between entities (e.g., "Reader A borrows resource B → associated subject C"), supporting multi-hop queries and semantic reasoning. b. Knowledge Graph Construction: Extract entity relationships through natural language processing (NLP) and semantic technologies (RDF, OWL) to form a unified cross-departmental knowledge network.

[0078] 4. Relationship integration includes real-time relationship updates and stream processing frameworks (such as Apache Kafka and Flink): capturing borrowing and return events in real time, dynamically updating entity relationships (such as automatically expanding nodes after a new borrowing record is added). Incremental ETL: synchronizing only newly added or changed data, reducing the overhead of full integration.

[0079] This invention adopts the multi-database joint technology MySQL, Radis, No4J, ClickHouse based on distributed query, and its core uses include:

[0080] 1. Heterogeneous data integration: Federated query technologies (such as Presto and the ClickHouse table engine) enable seamless joint queries across MySQL (transactional data), Redis (cached / real-time data), Neo4j (relational networks), and ClickHouse (analytic data), breaking down data silos. This also supports cross-database joins, subqueries, and aggregation operations, such as real-time correlation of user behavior (ClickHouse).

[0081] 2. Performance optimizations include: a. Hot data acceleration: Redis caches frequently accessed data, reducing query latency from milliseconds to microseconds. b. Complex analysis acceleration: ClickHouse columnar storage enables aggregation of billions of data points in seconds, faster than traditional MySQL analysis. c. Relational query efficiency: Neo4j handles multi-hop relational queries (such as fraud detection) faster than SQL databases.

[0082] 3. Scenario-specific support includes: a. Real-time business: Redis + MySQL ensures high-concurrency transactions. b. Intelligent analysis: ClickHouse + Neo4j supports real-time user profiling and path analysis (such as recommendation systems). c. Consistency assurance: Data consistency is ensured through distributed transactions (XA protocol) and CDC synchronization (such as Debezium).

[0083] The actual effects include:

[0084] 1. Performance improvement: The response time for complex queries is optimized from minutes to seconds, and the system is highly stable in high-concurrency scenarios.

[0085] 2. Cost optimization: By using the database on demand, resource utilization is improved, redundant processes are reduced, and operation and maintenance costs are effectively reduced.

[0086] 3. Business value includes: a. Accurate recommendations: Neo4j relationship networks + ClickHouse behavioral analysis improve recommendation conversion rates. b. Real-time risk control: Multi-database joint query identifies complex fraud patterns. c. Decision-making efficiency: Dynamic federated query supports real-time cross-source analysis, improving decision-making speed.

[0087] The present invention also discloses a graph relationship integration technology based on relational reasoning, the core uses of which include:

[0088] 1. Multi-source knowledge fusion: Integrate structured (MySQL), semi-structured (JSON logs), and unstructured (PDF papers) data to build a unified knowledge graph, solving the data silo problem. Ontology alignment (such as BERT-OWL) enables cross-domain term mapping, such as linking medical and biological terms.

[0089] 2. Deep relationship discovery: Path-based reasoning (such as PathCon) is used to discover hidden associations (such as drug side effects and gene pathways) and supports complex reasoning. It also includes temporal graph analysis (such as Temporal KG) to track relationship evolution.

[0090] 3. Intelligent Decision Support: Dynamic conflict resolution (such as KG-BERT) resolves conflicts in multi-source knowledge and improves decision consistency. Real-time inference engines (such as Drools + Neo4j) support millisecond-level business response.

[0091] The actual effects include:

[0092] 1. Knowledge completeness: The coverage of graph entities is improved, the relationship density is increased, and the efficiency of discovering interdisciplinary knowledge associations is improved.

[0093] 2. System efficiency: The response time of federated graph queries is reduced, and the automated knowledge update cycle is shortened from days to minutes.

[0094] Step S30: Perform user portrait generation processing and book borrowing trend prediction processing based on the target knowledge graph to obtain user portrait results and book borrowing trend results.

[0095] The machine learning algorithms and large-scale modeling algorithms used in this invention include: 1. Clustering algorithms (such as K-Means and DBSCAN) are used to classify readers and identify behavioral patterns of different groups. 2. Classification algorithms (such as decision trees, random forests, and XGBoost) are used to predict readers' borrowing preferences and churn risk. 3. Regression analysis (such as linear regression and LSTM) is used to predict borrowing trends and resource demand.

[0096] The deep learning technologies employed in this paper include: 1. Using deep learning models (such as LSTM and GRU) to accurately predict time series data (such as borrowing volume and resource utilization). 2. Applying natural language processing technologies (such as BERT and Transformer) to analyze reader comments and search keywords to uncover potential demand.

[0097] The present invention combines the big model with the target knowledge graph, and can accurately construct user portrait results and book borrowing trend results based on the user's reading information and the library's borrowing information, so as to make book recommendations and decision suggestions to users and reasonably allocate library resources.

[0098] Specifically, feature extraction processing is performed on the target knowledge graph to obtain feature extraction results, wherein the feature extraction results include basic features, behavioral features and semantic features; cluster analysis processing is performed on the feature extraction results using a cluster analysis algorithm to obtain user behavior patterns; a pre-trained model is determined, and model training and model fine-tuning are performed on the pre-trained model according to the target knowledge graph and the user behavior patterns to obtain an end-to-end portrait model; user reading data corresponding to the target user is obtained, and the user reading data is input into the end-to-end portrait model to output a user portrait result.

[0099] like Figure 2 As shown, the portrait generation function in the present invention includes: 1. Feature engineering includes basic features: extracting static information (such as reader identity, department, grade) and statistical features (such as average monthly borrowing volume, high-frequency access time period). User behavior interest features: using the LSTM network to model the user's recent book borrowing records and extract the user's short-term interest features; using the multi-head attention mechanism to analyze the user's long-term book borrowing records and construct the user's long-term interest features. Semantic features: using NLP technology (such as BERT) to analyze reader comments and search terms, and extract topic distribution and sentiment tendencies. 2. Portrait modeling technology includes cluster analysis: using algorithms such as K-Means and DBSCAN to divide reader groups (such as "scientific research readers" or "leisure reading readers"). Label system construction: generating labels (such as "high-frequency visiting users" or "interdisciplinary enthusiasts") based on rule engines (such as Drools) and machine learning. Deep learning model: using neural networks (such as Transformer) to build an end-to-end portrait model and associate multi-dimensional behavioral features. 3. Real-time updates and dynamic adjustments, including incremental learning: Dynamically updating reader profiles to adapt to behavioral changes through online machine learning (such as FTRL and online random forests). Stream processing framework: Using Apache Flink to process behavioral data in real time and update profile tags (such as real-time tagging of "recently active users").

[0100] Reader behavior analysis includes: Analysis methods: Cluster analysis (K-Means) to segment reader groups, and association rule mining (Apriori) to discover borrowing patterns. Tools: Spark MLlib, Weka. Resource utilization analysis includes: Analysis methods: Statistics on resource borrowing rates and retention time, and identification of low-utilization resources. Tools: SQL, Tableau. Subject service support includes: Analysis methods: Constructing subject literature citation networks, analyzing research hotspots and trends (such as the graph database Neo4j). Tools: Gephi, Cytoscape.

[0101] Obtain the time series data in the target knowledge graph, and aggregate the time series data to obtain aggregated time series data; determine a preset deep learning model, and input the aggregated time series data into the preset deep learning model to output the book borrowing trend results.

[0102] Furthermore, the corresponding recommended book data and predicted decision direction are determined based on the user portrait results, and the recommended book data and the predicted decision direction are pushed to the corresponding target user; the current book resource allocation plan is determined, and the current book resource allocation plan is optimized according to the book borrowing trend results to obtain the target book resource allocation result.

[0103] In its implementation, the present invention utilizes a multi-head attention mechanism and a long short-term memory (LSTM) network to construct a deep learning architecture to analyze and predict users' book borrowing behavior. The LSTM network is used to model a user's recent book borrowing history and extract short-term interest features. The multi-head attention mechanism is then used to analyze a user's long-term book borrowing history and construct long-term interest features. Finally, the user's long-term and short-term interest features are combined to predict and analyze their next book borrowing. The predicted borrowing results are then recommended to the user, thereby improving the user's borrowing experience.

[0104] like Figure 3As shown, the present invention adopts a multi-head attention mechanism and a long short-term memory network (LSTM) to construct a deep learning architecture to analyze and predict the user's book borrowing behavior. The specific technical implementation process includes: 1. The most recent n book borrowing records of each user are regarded as the short-term borrowing distance, and the name of each book is modeled as a feature vector using a word embedding method; 2. The book feature vector is modeled into a user short-term borrowing interest feature vector using a multi-layer long short-term memory network (LSTM); 3. The remaining book borrowing records of each user are regarded as the user's long-term borrowing records, and the name of each book is modeled as a feature vector using a word embedding method; 4. The feature sequence is modeled into a user long-term borrowing interest feature vector using a multi-layer multi-head attention mechanism network (MUti-atten), and the user's long-term and short-term borrowing interest feature vectors are spliced; 5. A fully connected layer network (FC) is used to compress features, and the compressed features are expanded and combined with the feature vectors of each book; 6. The probability of the user borrowing the book is predicted.

[0105] Build corporate and personal portraits to accurately match corporate needs with student capabilities. For example, by analyzing corporate recruitment data and student practice data, we provide customized talent training programs for colleges and universities, and recommend outstanding talents that meet the needs of enterprises, thus forming an employment ecology of "school-enterprise linkage" and "precise matching", enhancing students' employment competitiveness and meeting the talent needs of enterprises.

[0106] The scenario definition and needs analysis process includes: 1. Clarifying analysis objectives: e.g., optimizing resource procurement, improving reader satisfaction, increasing resource utilization, supporting disciplinary services, etc. 2. Identifying key indicators: e.g., loan volume, search frequency, reader activity, resource utilization, etc. 3. Scenario modeling and disciplinary service support.

[0107] Resource procurement optimization includes: Analytical methods: Predicting future demand based on historical borrowing data and subject hotspots (such as time series analysis and regression models). Tools: Python (Pandas, Scikit-learn), R language.

[0108] The graph relationship integration technology provided in the present invention: Purpose: Mainly used to integrate and associate various resources and data to build a comprehensive knowledge graph. This technology reveals the potential connections between data by analyzing the relationships between different data sources (such as books, reader behavior, borrowing records, etc.), thereby providing more intelligent services. Effect: Data integration eliminates data silos, enabling data from different departments and systems to be seamlessly connected and integrated. Intelligent recommendation: Through graph analysis, the system can provide personalized book recommendations based on readers' historical behaviors and preferences. Decision support: Help library managers understand resource usage, reader needs, etc., to achieve more accurate resource allocation and service optimization.

[0109] By integrating the library's multi-source data resources, accurate teaching support tools and data analysis services are provided to teachers and students. Teachers can use the system's borrowing data, resource usage data, and reader behavior data to optimize course design and teaching content. For example, by analyzing popular borrowed books and subject trends, they can adjust teaching focus and recommend reading materials. Students can use the system's personalized recommendation function to obtain high-quality resources related to the course and improve learning efficiency. At the same time, the system supports real-time data monitoring and dynamic adjustment, helping teachers to keep abreast of students' learning progress and resource usage, thereby optimizing teaching strategies. In addition, the knowledge graph and semantic analysis functions provided by the system can assist teachers in conducting interdisciplinary teaching and research, further improving teaching quality and academic influence. Through data-driven teaching optimization and precise resource matching, the system has significantly improved teaching effectiveness and efficiency, creating a more efficient learning and research environment for teachers and students.

[0110] The library data intelligent analysis system collects statistics on library traffic from various dimensions, exploring reader needs and usage habits. This system then correlates and matches information about library space resources to optimize library resource allocation, providing data support for improving service quality and enhancing the reader experience. Data sources for library traffic analysis include access control systems, loan systems, self-service terminals, Wi-Fi networks, and intelligent surveillance systems. This data records real-time information such as reader arrival time, length of stay, and activity patterns. Data from these different sources can be cross-checked to enhance data value. For example, access control data can be compared with loan records to identify peak hours and hotspots of reader activity.

[0111] Library traffic analysis focuses on identifying reader behavior, such as common visit times, most popular reading areas, and frequently borrowed book types. The results can reveal differences in peak visit times between weekdays and weekends, or indicate the density of traffic in certain areas (such as study rooms or multimedia rooms) during specific hours. This allows libraries to implement appropriate measures, such as extending opening hours, adding seating, and optimizing resource allocation. Traffic analysis helps libraries target resource allocation, such as adjusting the shelving times and quantity of books and periodicals. For high-traffic areas, additional shelves, seating, and electrical outlets can be added, or self-service check-in and return systems can be installed to reduce waiting times. Analyzing seminar room reservations and usage can identify peaks and troughs in user traffic. Analyzing seminar room reservation data can help understand usage and demand for seminar rooms during different time periods, allowing for tailored management strategies. During periods of low usage, consider introducing flexible reservation policies or special events to attract more users. Finally, library traffic analysis can also add predictive functions; by modeling and analyzing historical data, the number of visitors in a certain period of time in the future can be predicted, and coordination and collaboration can be carried out in advance to avoid management loss caused by excessive traffic or waste of resources caused by insufficient traffic; for example, when a large-scale examination is approaching, the library estimates a surge in the number of visitors and increases the opening hours and study room seats accordingly.

[0112] Borrowing data is one of the most important data in libraries. Systematic analysis of borrowing data can help us understand readers' reading preferences and the utilization of collection resources. Borrowing data analysis helps libraries optimize resource scheduling and provide support for collection updates and improvements to spatial layout. Borrowing data includes the number of times a book is borrowed, the borrowing time, the identity of the borrower (such as students, teachers, and the general public), etc. These data are stored in chronological order or by category to form a complete borrowing data set. By analyzing the borrowing volume of students in different grades, we can show the differences in book demand in different academic years and provide more targeted reading resources for students in different grades; for example, lower-grade students may need basic textbooks, while higher-grade students prefer professional literature or research materials. Borrowing statistics for each college help libraries identify which disciplines and fields have high demand for books, so as to rationally allocate book resources. In addition, collection ratio statistics reflect the ratio between library collections and actual borrowed books, evaluate the use of collections, empower book procurement strategies, eliminate books with low borrowing rates, and optimize collection structure.

[0113] By counting borrowing times, we can identify the most popular books. These books may be particularly popular during certain time periods, reflecting a common interest among readers at that time. Secondly, we can analyze borrower behavior. By analyzing the borrowing habits of different groups, such as students, faculty, or other personnel, libraries can better understand their needs and provide more targeted thematic services. Furthermore, we can analyze borrowing frequency and loan cycles, such as whether certain types of books are frequently borrowed during specific seasons or semesters, to help libraries dynamically adjust resources for future planning. Analysis of collection resource utilization is a key component of borrowing data analysis. By calculating the borrowing rates of collection books, we can identify books with high and low borrowing frequencies, enabling more informed decisions when updating the collection. Finally, borrowing data analysis can be combined with other data (such as library traffic and electronic resource access) to generate even more valuable insights. For example, we can analyze the correlation between borrowing volume and library traffic within a specific time period to optimize library hours and service arrangements.

[0114] The present invention discloses a user long-term and short-term behavior interest feature analysis technology based on the LSTM+ multi-head attention mechanism, wherein the uses include: modeling and analyzing the user's historical book borrowing records, constructing the reader's book borrowing interest preference features, and thus constructing the user interest portrait, providing a basis for personalized book recommendations.

[0115] The effects include: 1. Accuracy: To prevent feature analysis errors caused by changes in user borrowing interests, the present invention divides user historical borrowing records into short-term and long-term borrowing, and uses LSTM and multi-head attention mechanisms to model users' short-term or long-term borrowing records, respectively, reducing errors caused by changes in user interests. 2. Lightweight: The multi-head attention mechanism is effective for sequence time series modeling but has high computational complexity and high resource requirements; the LSTM framework has low computational complexity but is affected by gradient explosion and has difficulty modeling long time series. The present invention divides user borrowing records into long-term and short-term records, and uses the LSTM framework to model short-term records, reducing computational complexity while ensuring the effect.

[0116] Step S40: Visually display the target knowledge graph, the user portrait results, and the book borrowing trend results.

[0117] By introducing advanced visualization tools (such as ECharts and D3.js), the interactivity and dynamic effects of charts can be significantly improved to meet the personalized needs of different users. Specific technical implementations include: 1. Interactive Chart Design: Using JavaScript libraries such as ECharts and D3.js, interactive charts (such as bar charts, line charts, and pie charts) are created, allowing users to explore data through clicks, dragging, and zooming. Dynamic filtering and linked analysis functions are implemented, allowing users to select specific data dimensions (such as time ranges and reader groups) and update the charts in real time. 2. Dynamic Data Display: Using ECharts' timeline function to display time series data (such as borrowing trends and resource utilization), animating the data's evolution. Using D3.js's transition and interpolation functions to achieve smooth transitions and dynamic updates of data. 3. Multidimensional Data Visualization: Using parallel coordinates and radar charts to display multidimensional data (such as reader attributes, resource types, and borrowing frequency), users can discover complex relationships. Use heatmaps and treemaps to display the distribution and structure of high-dimensional data. 4. Report model statistical functions include: a. Data aggregation and modeling, including data warehouse technology: Use Hadoop, Snowflake, etc. to build data warehouses to integrate multi-source data such as borrowing records, resource usage, and reader behavior. OLAP analysis: Use OLAP engines such as ClickHouse and Apache Druid to achieve multi-dimensional data aggregation (such as grouping statistics by time, subject, and reader type). Statistical model: Build statistical models based on SQL or Python (such as borrowing volume prediction, resource utilization analysis, and user activity calculation). b. Report template design includes: BI tool integration: Use tools such as Tableau and Power BI to design visual report templates, supporting charts (bar charts, heat maps), tables, and dashboards. Dynamic parameter configuration: Allow users to customize filtering conditions (such as time periods, subject classifications) and generate dynamic reports. c. Automated generation and update include: Scheduled task scheduling: Use tools such as Airflow and Cron to regularly trigger ETL processes and report generation to ensure real-time data. Incremental calculation: Only new or changed data is updated, reducing the cost of full calculations (e.g., daily incremental borrowing statistics). d. Multi-dimensional statistical support includes: Cross-analysis: Supports multi-dimensional combined statistics (e.g., "Student borrowing trends in computer science books in a certain college over the past six months"). Comparative analysis: Provides year-on-year and month-on-month comparisons (e.g., borrowing volume this month vs. the same period last year).

[0118] Relationship graph display features include: 1. Data modeling and storage, graph databases (such as Neo4j and TigerGraph): Store entities (such as readers, books, disciplines, and authors) and relationships (such as "borrowed," "related disciplines," and "co-authors"). Support efficient multi-hop queries (such as "find all books borrowed by a reader and their related disciplines"). 2. Knowledge graph construction: Use semantic technologies (RDF and OWL) to define entity relationships and combine natural language processing (NLP) to extract associations from unstructured data (such as literature and reviews).

[0119] Visualization technologies include: Front-end frameworks (such as D3.js, ECharts, and Vis.js): Dynamic rendering of nodes (entities) and edges (relationships), supporting interactive operations such as dragging, zooming, and highlighting. Hierarchical display by entity type (reader, book) or relationship weight (borrowing frequency). 3D visualization (such as Three.js and WebGL): Display of complex three-dimensional relationship networks (such as interdisciplinary networks and resource citation chains).

[0120] Dynamic query and real-time updates include: Cypher query language (Neo4j): Dynamically generating subgraphs by writing queries (e.g., MATCH (reader) - [borrower] -> (book)). Real-time data streams (e.g., Kafka, Flink): Capturing events such as borrowing and commenting, and updating graph relationships in real time (e.g., automatically expanding nodes after a new borrowing record).

[0121] Interactive features include: Node filtering and focusing: Users can filter nodes by attributes (such as subject, borrowing time), or focus on key entities (such as popular books); Path exploration: Supports finding the shortest path between two entities (such as "the associated path from reader A to subject D"); Semantic search: Generates dynamic graphs based on natural language queries (such as "show books and borrowers related to artificial intelligence").

[0122] Relationship visualization and interaction, including visualization tools (such as D3.js, ECharts, and Vis.js): Dynamic rendering of nodes (entities) and edges (relationships), supporting interactive operations such as dragging, zooming, and highlighting. Semantic search: Generate dynamic graphs based on natural language queries (such as "show books and borrowers related to artificial intelligence").

[0123] Output results include: 1. Visual Charts: Use bar charts, line charts, pie charts, heat maps, and more to display analysis results (such as borrowing trends and resource utilization). 2. Interactive Dashboards: Provide dynamic filtering and linkage analysis capabilities, allowing users to customize data viewing (such as filtering by time or subject). 3. Reports and Documents: Generate reports in PDF, Word, or HTML format, including analysis conclusions and recommendations (such as a monthly borrowing analysis report). 4. API Interface: Provides a RESTful API or GraphQL interface to support other systems calling analysis results (such as a recommendation system).

[0124] Report dynamic analysis technology includes: Purpose: Mainly used to generate real-time, interactive statistical reports to help users dynamically adjust analysis dimensions and filtering conditions, and deeply explore the trends and patterns behind the data. Effect: Real-time update: Users can instantly view the latest borrowing data, reader behavior and resource usage to ensure the timeliness of information. Flexibility: Users can customize report content and display methods according to their needs, and flexibly filter and compare data of different dimensions. Accurate decision support: Through dynamic analysis, managers can quickly identify problems and take corresponding measures to improve library resource allocation and operational efficiency.

[0125] Enhancing visualization capabilities is key to improving user experience and decision-making efficiency. By introducing advanced visualization tools (such as ECharts, D3.js, and Tableau), we can enhance the interactivity and dynamic effects of charts to meet the personalized needs of different users. We design interactive dashboards that support dynamic filtering, linked analysis, and real-time updates. By combining 3D visualization (such as Three.js) and knowledge graphs to display complex relationship networks, we transform complex data into intuitive and easy-to-understand visualizations, lowering the barrier to data comprehension. Furthermore, we support multi-terminal access and personalized customization, significantly improving user experience and decision-making efficiency.

[0126] The present invention discloses a report dynamic analysis technology based on interactive analysis:

[0127] Core uses include:

[0128] 1. Real-time Data Insights: Supports low-latency queries on massive data volumes (e.g., ClickHouse + Presto federated analytics), dynamically generating real-time dashboards showing reading trends and user behavior. Stream processing (e.g., Flink) enables data refresh within seconds, making it suitable for scenarios such as promotion monitoring and IoT device alerts.

[0129] 2. Self-service exploratory analysis: Drag-and-drop interactions (such as Tableau / Power BI) allow business personnel to independently create reports, reducing IT reliance (improving development efficiency). Natural language query (NLQ) supports "voice question-to-chart generation."

[0130] 3. Intelligent Decision-Making: Automatic anomaly detection (such as Prophet time series forecasting) flags data fluctuations and pushes early warning notifications. Enhanced analysis explains data trends, with an explainability score of 80%.

[0131] Actual effect: Efficiency is improved, and chart generation time is shortened from hours to minutes.

[0132] In addition, intelligent data analysis based on other structural scenarios is mainly reflected in the scalability of this technology. By introducing large models (such as BERT, Gork, etc.) and advanced intelligent analysis technologies, the depth, breadth and flexibility of data analysis can be significantly improved to meet the needs of multiple scenarios and fields.

[0133] Across diverse application scenarios, intelligent data analysis technology demonstrates remarkable adaptability and expansion potential. By integrating cutting-edge AI technologies, including deep semantic understanding models in natural language processing (such as DeepSeek and the GPT series), knowledge-enhanced language models (such as BERT), and large models with complex reasoning capabilities (such as Gork), data analysis systems have achieved a qualitative leap. The integrated application of these advanced technologies has enabled data analysis to move beyond traditional structured processing and delve deeper into the semantic connections and underlying patterns within the data.

[0134] The introduction of knowledge graph technology further enhances the ability to analyze cross-domain data associations, while the continuous learning mechanism ensures that the system can adapt to new business scenarios and data characteristics over time. This invention significantly improves capabilities in three dimensions: 1. In terms of analytical depth, it achieves a leap from descriptive analysis to predictive and prescriptive analysis; 2. In terms of application breadth, it covers the comprehensive processing of structured data to unstructured multimodal data; 3. In terms of response flexibility, it supports multiple modes from preset analysis to instant exploration, which enables libraries to respond quickly to changes and maintain a competitive advantage in digital transformation.

[0135] Furthermore, if Figure 4 As shown, based on the above-mentioned intelligent analysis method of multi-source book data, the present invention also provides an intelligent analysis system for multi-source book data, wherein the intelligent analysis system for multi-source book data includes:

[0136] The data preprocessing module 51 is used to obtain multi-source book data and preprocess the multi-source book data to obtain target multi-source book data;

[0137] A data relationship integration module 52 is used to perform data relationship integration processing on the target multi-source book data to obtain a target knowledge graph;

[0138] A portrait generation and trend prediction module 53 is used to generate a user portrait and predict a book borrowing trend based on the target knowledge graph to obtain a user portrait result and a book borrowing trend result;

[0139] The visualization display module 54 is used to visualize the target knowledge graph, the user portrait results and the book borrowing trend results.

[0140] Furthermore, if Figure 5 As shown, based on the above-mentioned intelligent analysis method and system for multi-source book data, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 5 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0141] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal, etc. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, the memory 20 stores an intelligent analysis program 40 for multi-source book data, and the intelligent analysis program 40 for multi-source book data can be executed by the processor 10, thereby realizing the intelligent analysis method for multi-source book data in the present application.

[0142] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program codes or process data stored in the memory 20, such as executing the intelligent analysis method for the multi-source book data.

[0143] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch screen, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0144] In one embodiment, when the processor 10 executes the intelligent analysis program 40 for multi-source book data in the memory 20 , the steps of the intelligent analysis method for multi-source book data are implemented.

[0145] In summary, the present invention provides an intelligent analysis method, system and terminal for multi-source book data, the method comprising: obtaining multi-source book data, and pre-processing the multi-source book data to obtain target multi-source book data; performing data relationship integration processing on the target multi-source book data to obtain a target knowledge graph; performing user portrait generation processing and book borrowing trend prediction processing based on the target knowledge graph to obtain user portrait results and book borrowing trend results; and visually displaying the target knowledge graph, the user portrait results and the book borrowing trend results. The present invention obtains a target knowledge graph by pre-processing and data relationship integration processing on multi-source book data, and performs user portrait generation processing and book borrowing trend prediction processing through the target knowledge graph, which can achieve accurate analysis of user behavior and book borrowing trends, and then achieves optimal allocation of book resources based on the generated user portrait results and book borrowing trend results, which not only effectively improves the availability of book resources, but also optimizes user experience.

[0146] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0147] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0148] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. An intelligent analysis method for multi-source book data, characterized in that: The intelligent analysis method of multi-source book data includes: Acquire multi-source book data, and pre-process the multi-source book data to obtain target multi-source book data; Performing data relationship integration processing on the target multi-source book data to obtain a target knowledge graph; Perform user portrait generation processing and book borrowing trend prediction processing according to the target knowledge graph to obtain user portrait results and book borrowing trend results; The target knowledge graph, the user portrait results and the book borrowing trend results are visualized.

2. The intelligent analysis method of multi-source book data according to claim 1, characterized in that: The step of acquiring multi-source book data and pre-processing the multi-source book data to obtain target multi-source book data specifically includes: Using a distributed log collection tool to acquire multi-source book data, wherein the multi-source book data includes structured data, unstructured data, real-time data streams, and external data sources; Performing data standardization processing on the multi-source book data using metadata standards to obtain standardized multi-source book data; Using a data mapping tool to perform data mapping processing on the standardized multi-source book data to obtain mapped multi-source book data; The mapped multi-source book data is cleaned to obtain target multi-source book data.

3. The intelligent analysis method of multi-source book data according to claim 2, characterized in that: The data cleaning process includes data deduplication, missing value filling, formatting and outlier processing; The data cleaning process of the mapped multi-source book data to obtain target multi-source book data specifically includes: Acquiring data identifiers in the mapped multi-source book data, determining duplicate data in the mapped multi-source book data according to the data identifiers, and performing the data deduplication process on the duplicate data to obtain first multi-source book data; Obtaining missing values in the first multi-source book data, and performing missing value filling processing on the missing values using a preset filling method to obtain second multi-source book data, wherein the preset filling method includes an interpolation method, a mean filling method, and a missing value deletion method; Acquiring a date format corresponding to the second multi-source book data, and performing the formatting process on the second multi-source book data according to the date format to obtain third multi-source book data; An outlier detection algorithm is used to identify outliers on the third multi-source book data to obtain target outlier data, and the target outlier data is processed to obtain target multi-source book data.

4. The intelligent analysis method of multi-source book data according to claim 1, characterized in that: The data relationship integration processing of the target multi-source book data to obtain the target knowledge graph specifically includes: Constructing a centralized data model, and performing entity relationship structuring processing on the target multi-source book data through the centralized data model to obtain structured multi-source book data; Determining a core entity in the structured multi-source book data, and performing identifier assignment processing and fuzzy matching processing on the core entity to obtain associated multi-source book data; Entity relationships are extracted from the associated multi-source book data using natural language processing and semantic technology to obtain entity relationship extraction results, and a target knowledge graph is constructed based on the entity relationship extraction results.

5. The intelligent analysis method of multi-source book data according to claim 1, characterized in that: The user portrait generation process and book borrowing trend prediction process are performed based on the target knowledge graph to obtain user portrait results and book borrowing trend results, specifically including: Performing feature extraction and cluster analysis on the target knowledge graph to obtain feature extraction results and user behavior patterns, and performing user portrait generation based on the feature extraction results and the user behavior patterns to obtain user portrait results; Obtaining time series data in the target knowledge graph, and performing aggregation processing on the time series data to obtain aggregated time series data; A preset deep learning model is determined, and the aggregated time series data is input into the preset deep learning model to output a book borrowing trend result.

6. The intelligent analysis method of multi-source book data according to claim 5, characterized in that: The target knowledge graph is subjected to feature extraction and cluster analysis to obtain feature extraction results and user behavior patterns, and user portrait generation is performed based on the feature extraction results and the user behavior patterns to obtain user portrait results, specifically including: Performing feature extraction processing on the target knowledge graph to obtain feature extraction results, wherein the feature extraction results include basic features, behavioral features, and semantic features; Performing cluster analysis on the feature extraction results using a cluster analysis algorithm to obtain user behavior patterns; Determine a pre-trained model, and perform model training and model fine-tuning on the pre-trained model according to the target knowledge graph and the user behavior pattern to obtain an end-to-end portrait model; Obtain user reading data corresponding to the target user, input the user reading data into the end-to-end portrait model, and output the user portrait result.

7. The intelligent analysis method of multi-source book data according to claim 6, characterized in that: The user portrait generation process and the book borrowing trend prediction process are performed according to the target knowledge graph to obtain the user portrait result and the book borrowing trend result, and then further includes: Determine the corresponding recommended book data and predicted decision direction based on the user portrait result, and push the recommended book data and the predicted decision direction to the corresponding target user; A current book resource allocation plan is determined, and the current book resource allocation plan is optimized according to the book borrowing trend result to obtain a target book resource allocation result.

8. An intelligent analysis system for multi-source book data, characterized in that: The intelligent analysis system for multi-source book data includes: A data preprocessing module, configured to obtain multi-source book data and preprocess the multi-source book data to obtain target multi-source book data; A data relationship integration module is used to perform data relationship integration processing on the target multi-source book data to obtain a target knowledge graph; A portrait generation and trend prediction module is used to generate user portraits and predict book borrowing trends based on the target knowledge graph to obtain user portrait results and book borrowing trend results; A visualization display module is used to visualize the target knowledge graph, the user portrait results and the book borrowing trend results.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and an intelligent analysis program for multi-source book data stored in the memory and runnable on the processor. When the intelligent analysis program for multi-source book data is executed by the processor, the steps of the intelligent analysis method for multi-source book data as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an intelligent analysis program for multi-source book data, and when the intelligent analysis program for multi-source book data is executed by a processor, the steps of the intelligent analysis method for multi-source book data according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Terahertz knowledge graph construction method and system

    CN111813874A

  • Book recommendation retrieval system based on knowledge graph and reader portrait

    CN116595246A

  • Smart library reading promotion management method based on knowledge graph

    CN118552262A

  • User portrait generation query method based on knowledge graph

    CN119149755A

  • Library inventory management method and system

    CN119228271A

Cited By

  • Quick retrieval method and system for multi-hop relationship of data warehouse based on graph embedded index

    CN120723757A

  • Intelligent Chinese text error correction method integrating spelling and semantics

    CN120930635A

  • Book information archive data processing method and system based on big data

    CN121071022A

  • Apartment financial management intelligent auxiliary method and system based on artificial intelligence

    CN121504642A

  • Big data-oriented multi-model fusion analysis processing method, apparatus and device, and medium

    CN121658465A