Online real-time duplicate checking analysis method and system suitable for library

Through multi-source heterogeneous data processing and intelligent analysis, the problem of inefficient library data management and severity checking is solved, and efficient and accurate library resource management and decision-making support are achieved.

CN120386856AActive Publication Date: 2025-07-29SHANDONG CHINESE EDUCATION IND DEVELOPMENT CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510391927.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-29
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing library data management system lacks unified and efficient data maintenance and plagiarism checking tools, resulting in data update lag, entry errors, low efficiency in plagiarism checking, and inability to meet the needs of large-scale data processing. The analysis of plagiarism checking results is not detailed enough, making it difficult to support resource integration and procurement decisions.

Method used

The methods of multi-source heterogeneous data set acquisition and preliminary screening, data deep cleaning and feature extraction, real-time plagiarism check, association rule mining and clustering analysis are used, and book resource classification and evaluation are carried out, and visual output is carried out.

Benefits of technology

It improves the completeness and accuracy of data, improves the speed and accuracy of plagiarism checking, accurately identify duplicate records, discover potential relationships, optimize resource allocation, and provide scientific management decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386856A_ABST
    Figure CN120386856A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of book duplicate checking, in particular to an online real-time duplicate checking analysis method and system suitable for a library, and the method comprises the steps: obtaining a multi-source heterogeneous data set of a book management system, and carrying out the preliminary screening of the obtained data set; performing data deep cleaning and feature extraction based on the screened data set, performing real-time duplicate checking according to the extracted feature vector, performing association rule mining on the duplicate checking result, and performing book resource classification and evaluation by using the association rule, including classifying library resources by using a clustering algorithm; and performing visual output according to association rule mining and clustering analysis results. According to the method, the optimal k value is determined by using the contour coefficient, the clustering result is evaluated in combination with the association rule, the repetition condition and the utilization rate of the book resources in each cluster can be comprehensively analyzed, reasonable resource cleaning, optimization or expansion suggestions are provided for different conditions, and the configuration efficiency and the management level of the book resources are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of book duplicate checking, and in particular to an online real-time duplicate checking and analysis method and system suitable for libraries. Background Art

[0002] With the rapid development of information technology, the scale of library collection resources has grown exponentially, and its types have become increasingly diverse, covering traditional paper books, electronic documents, audio-visual materials, etc. At the same time, business transactions between libraries and external institutions and readers are becoming increasingly frequent, and cooperation in resource procurement, shared lending, academic research support, etc. is constantly deepening. Against this background, efficient management, accurate duplication detection and in-depth analysis of library data have become crucial.

[0003] In terms of library management, existing technologies have many shortcomings. On the one hand, data maintenance in many libraries relies on manual recording and decentralized table management, lacking unified and efficient system support. This approach not only leads to delayed data updates, affecting the timely shelving and service development of new resources, but is also prone to input errors, making it difficult to ensure data consistency and integrity. On the other hand, in terms of duplicate checking functions, existing duplicate checking tools have obvious limitations. Most tools only support duplicate checking of single data and cannot meet the needs of large-scale data processing. When a library needs to compare a large number of collection resources, or authorizes customers to process duplicate checking tasks for a large number of data at one time, the operation is cumbersome and time-consuming, and work efficiency is greatly reduced. At the same time, the records during the duplicate checking process are not detailed enough, making it difficult to trace the specific operation details. The result display and export functions are also extremely limited, which is not conducive to the subsequent in-depth analysis and utilization of the duplicate checking results, and cannot provide strong support for the library's resource integration, procurement decision adjustments, etc. At this stage, a method and system for online real-time duplicate checking analysis suitable for libraries is needed. Summary of the Invention

[0004] In order to solve the problems of low efficiency of book duplicate checking and imperfect analysis of duplicate checking results, the present invention provides an online real-time duplicate checking and analysis method and system suitable for libraries.

[0005] In a first aspect, the present invention provides an online real-time duplicate checking and analysis method for libraries, which adopts the following technical solutions: A real-time online duplicate analysis method for libraries, comprising: Obtain multi-source heterogeneous data sets from the library management system and perform preliminary screening on the obtained data sets; Perform deep data cleaning and feature extraction based on the filtered dataset, including processing missing values using a Bayesian network-based missing value filling method and using the TF-IDF algorithm for feature vectorization; Perform real-time duplicate checking based on the extracted feature vectors, including constructing a hash index, calculating the similarity between records, and finally marking the duplicate checking results; Mine association rules for the duplicate checking results, including mining frequent itemsets and association rules by calculating the support and confidence of item sets; Use association rules for book resource classification and evaluation, including classifying library resources using the K-Means clustering algorithm; Perform visual output according to the results of association rule mining and cluster analysis.

[0006] Furthermore, the preliminary screening of the obtained dataset includes preliminarily screening the collected data using the information entropy theory. For each record, calculate the information entropy of each of its fields, and consider the fields with information entropy lower than the set threshold as low-value information for preliminary elimination. The information entropy calculation formula is: , where, represents the field, represents the probability of the field value of.

[0007] Furthermore, the method for filling missing values based on the Bayesian network processes missing values, including constructing a Bayesian network to describe the dependency relationships between fields in the library dataset, where nodes represent fields in the dataset, and the directed edges between nodes represent the dependency relationships between fields. Use the maximum likelihood estimation method to learn the structure and parameters of the Bayesian network. For the fields with missing values in the dataset, calculate their conditional probabilities according to Bayes' theorem. Introduce a dynamic Bayesian network and divide it into time slices. A Bayesian network structure is constructed for each time slice, and the nodes between adjacent time slices are connected to each other through transition probabilities.

[0008] Furthermore, the real-time duplicate checking based on the extracted feature vectors includes using a Bloom filter to construct a hash index for the extracted features, mapping records into a binary vector through multiple hash functions. For different records, use a Siamese network model based on deep learning to calculate the similarity between records. Among them, the Siamese network consists of two sub-networks with shared weights. Input two records into the two sub-networks respectively, output feature vectors, calculate the similarity between the feature vectors through the Euclidean distance, and mark the records with similarity higher than the set threshold as duplicate records.

[0009] Further, mining frequent item sets and association rules by calculating the support and confidence of item sets includes adopting a dynamic threshold adjustment strategy, performing cluster analysis on historical data, dividing it into multiple clusters according to different characteristics of the data, calculating the standard deviation and coefficient of variation of the data in each cluster to evaluate the data dispersion degree, determining the minimum support threshold required for mining frequent item sets according to the data dispersion degree, applying the weighted Apriori algorithm, assigning weights to each item, calculating the support according to the assigned weights, and screening out frequent item sets through the support and the minimum support threshold.

[0010] Further, mining frequent item sets and association rules by calculating the support and confidence of item sets further includes calculating the confidence based on the frequent item sets, generating association rules when the confidence is greater than the set confidence threshold, and adopting an incremental association rule update method to update the association rules. The confidence calculation formula is: , where represents the association rule from item set X to item set Y, represents the weighted support of the union of item set X and item set Y, represents the weighted support of item set X.

[0011] Further, classifying library resources by using the K-Means clustering algorithm includes determining the optimal k value by using the silhouette coefficient, selecting the feature vectors of k book resources as the initial clustering centers, calculating the distance between the feature vector of each book and each clustering center according to the Euclidean distance, assigning the feature vector to the cluster where the nearest clustering center is located, calculating the mean value of all book feature vectors in the cluster, taking it as the new clustering center, and evaluating the library resources according to the clustering results, and analyzing the duplication situation and utilization rate of the library resources in each cluster.

[0012] In a second aspect, a library online real-time duplicate checking and analysis system is provided, including: A data acquisition module, configured to: acquire a multi-source heterogeneous data set of a library management system and perform a preliminary screening on the acquired data set; A feature extraction module, configured to: perform in-depth data cleaning and feature extraction based on the screened data set, including processing missing values by using a missing value filling method based on a Bayesian network and performing feature vectorization by using the TF-IDF algorithm; A duplicate checking module, configured to: perform real-time duplicate checking according to the extracted feature vectors, including constructing a hash index and calculating the similarity between records, and finally marking the duplicate checking results; An association module, configured to: perform association rule mining on the duplicate checking results, including mining frequent item sets and association rules by calculating the support and confidence of item sets; A resource classification module, configured to: classify and evaluate book resources by using association rules, including classifying library resources by using the K-Means clustering algorithm; An output module, configured to: perform visual output according to the results of association rule mining and clustering analysis.

[0013] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor of a terminal device for a method for online real-time duplicate checking and analysis applicable to a library.

[0014] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor for a method for online real-time duplicate checking and analysis applicable to a library.

[0015] In summary, the present invention has the following beneficial technical effects: 1. By using the information entropy theory for preliminary screening, the present invention can quickly eliminate low-value information, reduce data redundancy, and improve the subsequent processing efficiency. The method for filling missing values based on the Bayesian network, especially the introduction of the dynamic Bayesian network, can fully consider the dependence relationship and dynamic changes between data fields, fill missing values more accurately, and greatly improve the integrity and accuracy of data compared with traditional methods, providing a reliable data basis for subsequent analysis.

[0016] 2. By using the Bloom filter to construct a hash index, the present invention can quickly determine whether a record is likely to be repeated in a large amount of book data, greatly improving the duplicate checking speed. The Siamese network model based on deep learning is used to calculate the similarity. Compared with traditional similarity calculation methods, it can more accurately identify the similarity between book records, accurately mark records with a similarity higher than the set threshold as duplicate records, reduce the misjudgment rate, and effectively avoid duplicate resource procurement and management chaos.

[0017] 3. When using the Apriori algorithm to mine association rules, the present invention adopts a dynamic threshold adjustment strategy and a weighted Apriori algorithm, which can more accurately mine potential and valuable association relationships between book resources according to data characteristics and business requirements, such as the close connection between specific authors and publishers, and Chinese Library Classification numbers, providing in-depth data support for library management decisions.

[0018] 4. When the K-Means clustering algorithm is used in the present invention to classify book resources, the optimal k value is determined by the silhouette coefficient, making the clustering result more in line with the actual distribution of book resources. The clustering result is evaluated in combination with association rules, and the duplication and utilization rate of book resources within each cluster can be comprehensively analyzed. Reasonable suggestions for resource cleaning, optimization or expansion are put forward for different situations, improving the allocation efficiency and management level of book resources.

[0019] 5. The present invention performs visual output according to the results of association rule mining and clustering analysis, and displays information such as the association relationship, classification situation and utilization rate of book resources in an intuitive chart form, facilitating library managers to quickly understand the information behind the data, providing a powerful visual decision-making auxiliary tool for procurement decisions, bookshelf layout adjustment, resource promotion, etc., and enhancing the scientificity and timeliness of library management decisions. Brief Description of the Drawings

[0020] Figure 1 is the overall process schematic diagram of a method for online real-time duplicate checking and analysis applicable to libraries in an embodiment of the present invention. Detailed Embodiment

[0021] The present invention will be further described in detail below with reference to the accompanying drawings.

[0022] Embodiment 1 Refer to Figure 1 , a method for online real-time duplicate checking and analysis applicable to libraries in this embodiment includes: Obtain the multi-source heterogeneous data set of the library management system and conduct a preliminary screening on the obtained data set; Based on the screened data set, conduct in-depth data cleaning and feature extraction, including processing missing values using the missing value filling method based on the Bayesian network, and performing feature vectorization using the TF-IDF algorithm; Conduct real-time duplicate checking based on the extracted feature vectors, including constructing a hash index and calculating the similarity between records, and finally marking the duplicate checking results; Conduct association rule mining on the duplicate checking results, including mining frequent item sets and association rules by calculating the support and confidence of item sets; Use association rules for book resource classification and evaluation, including using the K-Means clustering algorithm to classify library resources; Perform visual output according to the results of association rule mining and clustering analysis.

[0023] Specifically, a method for online real-time duplicate checking and analysis applicable to libraries includes the following steps: As Figure 1 shown, S1. Obtain the multi-source heterogeneous data set of the library management system and conduct a preliminary screening on the obtained data set; Data is obtained from multiple sources such as library management systems, collection databases, purchase order records, and electronic resource platforms, covering various resource information such as paper books, e-journals, and audio-visual materials. The data types collected include structured data (such as database tables) and unstructured data (such as text descriptions and picture captions). After that, preliminary data screening is carried out. The information entropy theory is used to preliminarily screen the collected data. For each record, the information entropy of each field is calculated. Fields with information entropy lower than a set threshold (such as 0.2) can be regarded as low-value information and are preliminarily excluded. The information entropy calculation formula is: , where X is the field, is the probability of the field value .

[0024] S2. Based on the screened data set, data deep cleaning and feature extraction are carried out, including handling missing values using the missing value filling method based on the Bayesian network, and using the TF-IDF algorithm for feature vectorization; After the preliminary screening of the data, there may still be some problems that affect the accuracy and efficiency of subsequent analysis. Missing values in the data are a common problem and may affect the results of subsequent data analysis and mining. Traditional missing value handling methods, such as simple mean filling and median filling, do not consider the dependence relationship between fields in the data set, while the missing value filling method based on the Bayesian network can effectively solve this problem.

[0025] The Bayesian network is essentially a probabilistic graphical model, which represents the probabilistic dependence relationship between variables through a directed acyclic graph. In the traditional Bayesian network, the relationship between variables is modeled based on static data without considering the influence of time factors on the data. However, in the book management scenario, the data has obvious dynamic characteristics. For example, the procurement strategies of the library are different in different years, and the borrowing preferences of readers also change over time, which leads to different characteristics of book-related data (such as borrowing volume and book category distribution) at different times. In this application, a dynamic Bayesian network (Dynamic Bayesian Network, DBN) is used to incorporate the time dimension into the traditional Bayesian network. In the processing of book management data, time slices are divided in chronological order. Assuming the time slice is , in each time slice , a Bayesian network structure is constructed , where is the node set of time slice , corresponding to each field of the book data, such as ; is the set of directed edges between nodes within time slice t, representing the dependence relationship between fields within this time slice.

[0026] For the Bayesian network of each time slice t, the conditional probability distribution between nodes needs to be learned. Let the node have a conditional probability distribution of , where is the set of parent nodes of node in . These conditional probability distributions are determined by the maximum likelihood estimation method. Nodes between adjacent time slices and are connected by transition probabilities. Taking the node X as an example, the transition probabilities of its values and at different time slices are expressed as . Assume that X is "the number of borrowings" and its values are in the discrete value set . The transition probability matrix is determined according to the discrete values and transition probabilities. The elements in this matrix represent the probability of X taking a certain value in time slice transferring to X taking another value in time slice . The adjacent time slices are connected by the transition probability matrix through the transition probabilities between nodes. The specific connection method is as follows: assume that in time slice , the probability that node X is in state is , and the element in the transition probability matrix A represents the probability of transferring from state to state , that is . Then, at time slice , the probability that node X is in state is calculated by the following formula: , The meaning of this formula is that for each possible state of node X in time slice , its probability is the sum of the probabilities of all possible states of node X in time slice transferring to . Through this formula, according to the state probability distribution of the node in time slice and the transition probability matrix, the state probability distribution of the node in time slice can be calculated, thus realizing the connection of nodes between adjacent time slices through transition probabilities.

[0027] After connection, collect all the observed data in time slice except for the missing value fields , denoted as . Collect the data of the past k time slices ( to The observed data of , First, calculate the numerator part of the conditional probability distribution. For , use the Bayesian network structure within the time slice , the conditional probability distribution between nodes and the transition probability of adjacent time slices. According to the dependency relationship of each node in the Bayesian network, combined with the known observed data , calculate the joint probability of the occurrence of these observed data under different values of .

[0028] Next, calculate the denominator part of the conditional probability distribution. The denominator is a normalization constant used to ensure that the sum of the conditional probability distributions of the finally calculated is 1. It is obtained by summing over all possible values of in the numerator, where is the prior probability of . Based on the obtained numerator and denominator parts, according to the Bayesian theorem formula, calculate the probability distribution of different values of under the known and . According to the calculated conditional probability distribution of , select the value with the largest probability in the conditional probability distribution as the filling value.

[0029] Extract key features from the filled data, such as book title, author, ISBN number, publication year, Chinese Library Classification number, etc. For text features (such as book title, author), use the TF-IDF (term frequency-inverse document frequency) algorithm for feature vectorization to convert the text into a numerical feature vector for subsequent duplicate checking and analysis operations.

[0030] S3. Perform real-time duplicate checking based on the extracted feature vectors, including constructing a hash index and calculating the similarity between records, and finally marking the duplicate checking results; Build a hash index for the extracted features using a Bloom filter. A Bloom filter is a probabilistic data structure with high space efficiency. Its principle is to map records into a binary vector through multiple hash functions. For book data, the feature vector of each book (e.g., a vector obtained by vectorizing features such as book title, author, publisher, and publication year) will be processed by multiple hash functions respectively. Each hash function will correspond to a position in the binary vector, and set the value of that position to 1 (initially all bits in the binary vector are 0). In this way, a large number of book records can be efficiently compressed and stored in the Bloom filter, and subsequently, it can be quickly determined whether a record exists in the existing record set (although there is a certain false positive rate, it can greatly improve the search efficiency in large-scale data processing).

[0031] Use a Siamese network model based on deep learning to calculate the similarity between different records. The Siamese network has a unique structure and consists of two sub-networks with shared weights. When processing book records, the two book records to be compared are respectively input into these two sub-networks. The sub-networks will perform a series of complex feature extraction and transformation operations on the input records, and finally output feature vectors. Since the two sub-networks share weights, they can extract features from different input records in the same way, making the output feature vectors comparable. Then, the Euclidean distance is used to calculate the distance between these two feature vectors. The Euclidean distance is a commonly used method to measure the distance between two vector spaces, which can intuitively reflect the proximity of two vectors in space. In the calculation of book record similarity, the smaller the Euclidean distance between feature vectors, the more similar the two book records are.

[0032] According to the calculated similarity, mark the records with similarity higher than the set threshold as duplicate records. This threshold is set in advance according to the actual business requirements and data analysis. For example, if the set threshold is 0.8 (assuming the similarity value range is 0 - 1, and the larger the value, the more similar), when the similarity of two book records calculated by the Siamese network and Euclidean distance is 0.85, these two records will be marked as duplicate records.

[0033] S4. Conduct association rule mining on the duplicate check results, including mining frequent itemsets and association rules by calculating the support and confidence of item sets; The traditional Apriori algorithm uses fixed minimum support and minimum confidence thresholds, which is difficult to adapt to the complex and changeable characteristics of library data. This solution introduces a dynamic threshold adjustment strategy to dynamically adjust the thresholds according to the data distribution and business requirements.

[0034] By performing clustering analysis on historical data, the data is divided into multiple clusters according to different characteristics (such as time, resource type, etc.). For each cluster, the standard deviation and coefficient of variation of the data are calculated respectively to evaluate the degree of dispersion of the data. For clusters with a high degree of dispersion, the minimum support threshold is adjusted towards the lower limit of the initial range. For example, if the coefficient of variation CV of a cluster is greater than a certain set threshold , then the minimum support threshold of this cluster is set to . For clusters with a low degree of dispersion, the minimum support threshold is adjusted towards the upper limit of the initial range. If CV is less than , then . At the same time, in combination with the business objectives of the library, the minimum confidence threshold is determined. Suppose there are m key indicators related to the business objectives of the library. For each key indicator , a weight ( ) is defined, and , indicating the importance of this indicator in the business objectives. For an association rule , an association degree index is defined to measure the degree of association between this association rule R and the key indicator . The larger the value of , the closer the association rule R is associated with the key indicator . Suppose the initial minimum confidence threshold is (0.6 is adopted in this embodiment), and the calculation formula of the final minimum confidence threshold minConf is as follows: , where is an adjustment amplitude parameter , indicating the increase amplitude of the minimum confidence threshold when the association rule is closely associated with the key indicator.

[0035] In library data, the importance of different resource attributes for association rules is not the same. When evaluating book associations, the "Chinese Library Classification Number" may be more able to reflect the internal connection between books than the "location of the publisher". Therefore, this embodiment proposes a weighted Apriori algorithm to assign a weight to each item.

[0036] The support formula is adjusted to: where is the weight of item i. Calculate the support of the item set, filter out the item sets greater than the frequent item sets. When using the Apriori algorithm to mine frequent item sets, if the subset of a certain frequent item set does not meet the minimum support threshold, this subset is removed from the frequent item set search space.

[0037] When calculating the confidence, the weight factor is also considered. Based on the mined frequent item sets, the adjusted confidence formula is as follows: , where represents the association rule from item set X to item set Y, represents the weighted support of the union of item set X and item set Y, represents the weighted support of item set X. When the confidence is greater than the set confidence threshold, an association rule is generated. If new data causes the confidence of some association rules to be lower than the set minimum confidence threshold, these association rules are deleted from the rule base.

[0038] Finally, an incremental association rule update method is used to update the association rules. As the book data continues to increase and change, the existing association rules may no longer accurately reflect the relationship between the data. The incremental update method can efficiently update the association rules when new data is added. When new data is added, first judge the impact of the new data on the existing frequent item sets. For the newly added transactions, calculate the intersection of the item sets contained therein and the existing frequent item sets to determine which frequent item sets may be affected. Then recalculate the weighted support and weighted confidence for the potentially affected frequent item sets. If new data causes the confidence of some association rules to be lower than the set minimum confidence threshold, these association rules are deleted from the rule base; at the same time, new frequent item sets and association rules may be generated according to the new data, and after calculation and screening, they are added to the rule base to ensure that the association rules can always accurately reflect the relationship between the book data.

[0039] S5. Using association rules for book resource classification and evaluation, including classifying library resources using the K-Means clustering algorithm; The classification of library resources using the K-Means clustering algorithm includes determining the optimal k value using the silhouette coefficient. Among them, the silhouette coefficient is used to measure the tightness and separation degree of clustering. For each sample point in the dataset, the calculation of the silhouette coefficient is based on two key distances: one is the average distance between the sample point and other sample points in the same cluster (denoted as a), and the smaller this distance, the higher the tightness of the sample point in its cluster; the other is the average distance between the sample point and the sample points in the nearest other cluster (denoted as b). The calculation formula for the silhouette coefficient s is , and its value range is between [-1, 1]. When s is close to 1, it indicates that the sample point is closely connected to other points within the cluster and has a high separation from other clusters, and the clustering effect is good; when s is close to -1, it means that the sample point may be wrongly assigned to an inappropriate cluster. Select the k value with the largest average silhouette coefficient as the optimal number of clusters, take the feature vectors of k book resources as the initial clustering centers, calculate the distance between the feature vector of each book and each clustering center according to the Euclidean distance, assign the feature vector to the cluster where the nearest clustering center is located, calculate the mean value of all book feature vectors within the cluster, and use it as the new clustering center. Evaluate the book resources according to the clustering results, and analyze the duplication situation and utilization rate of book resources within each cluster.

[0040] S6. Perform visual output according to the results of association rule mining and clustering analysis.

[0041] Comprehensively display the results of association rules and clustering analysis. Based on the clustering graph, mark the important association rules related to certain clusters to more comprehensively understand the relationship and clustering situation between book resources. Finally, make a pie chart or bar chart to display the duplicate checking ratio of different types of books. Divide the books into different categories according to the Chinese Library Classification number, and count the proportion of the number of books marked as duplicates in each category to the total number of books in that category, so as to intuitively see which categories of books have more serious duplication situations, provide reference for the library in resource procurement and management, reasonably adjust the procurement strategy, and reduce unnecessary duplicate purchases.

[0042] Embodiment 2 The difference between this embodiment and Embodiment 1 is that this embodiment provides an online real-time duplicate checking and analysis system applicable to libraries, including: A data acquisition module, configured to: acquire a multi-source heterogeneous data set of a library management system and perform a preliminary screening on the acquired data set; A feature extraction module, configured to: perform in-depth data cleaning and feature extraction based on the screened data set, including processing missing values using a missing value filling method based on a Bayesian network and performing feature vectorization using the TF-IDF algorithm; A duplicate checking module, configured to: perform real-time duplicate checking according to the extracted feature vectors, including constructing a hash index and calculating the similarity between records, and finally marking the duplicate checking results; An association module, configured to: perform association rule mining on the duplicate checking results, including mining frequent item sets and association rules by calculating the support and confidence of item sets; A resource classification module, configured to: classify and evaluate book resources using association rules, including classifying library resources using the K-Means clustering algorithm; An output module, configured to: perform visual output according to the results of association rule mining and clustering analysis.

[0043] A computer-readable storage medium storing multiple instructions, the instructions being adapted to be loaded and executed by a processor of a terminal device for a method for online real-time duplicate checking and analysis applicable to a library.

[0044] A terminal device, comprising a processor and a computer-readable storage medium, the processor being used to implement each instruction; the computer-readable storage medium being used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor for a method for online real-time duplicate checking and analysis applicable to a library.

[0045] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. An online real-time duplicate checking and analysis method applicable to libraries, characterized in that, Including: Obtain the multi-source heterogeneous data set of the library management system, and conduct a preliminary screening on the obtained data set; Based on the screened data set, conduct in-depth data cleaning and feature extraction, including using the missing value filling method based on the Bayesian network to handle missing values, and using the TF-IDF algorithm for feature vectorization; Conduct real-time duplicate checking according to the extracted feature vectors, including constructing a hash index and calculating the similarity between records, and finally marking the duplicate checking results; Mine association rules for the duplicate checking results, including mining frequent item sets and association rules by calculating the support and confidence of item sets; Use association rules for book resource classification and evaluation, including using the K-Means clustering algorithm to classify library resources; Perform visual output according to the results of association rule mining and clustering analysis.

2. The method for online real-time duplicate checking and analysis applicable to a library according to claim 1, wherein The preliminary screening of the obtained data set includes using the information entropy theory to conduct a preliminary screening on the collected data. For each record, calculate the information entropy of each of its fields, and regard the fields with information entropy lower than the set threshold as low-value information for preliminary elimination. The information entropy calculation formula is: , Among them, is represented as a field, is represented as the value of the field probability.

3. A method for online real-time duplicate checking and analysis applicable to libraries according to claim 1, characterized in that The method of handling missing values based on the Bayesian network includes constructing a Bayesian network to describe the dependence relationship between each field in the library data set, where the nodes represent the fields in the data set, and the directed edges between the nodes represent the dependence relationship between the fields. Use the maximum likelihood estimation method to learn the structure and parameters of the Bayesian network. For the fields with missing values in the data set, calculate their conditional probabilities according to Bayes' theorem, introduce a dynamic Bayesian network, and divide the time slices. A Bayesian network structure is constructed for each time slice, and the nodes between adjacent time slices are connected by transition probabilities.

4. The on-line real-time duplicate checking and analysis method applicable to a library according to claim 1, characterized in that, The real-time duplicate checking according to the extracted feature vectors includes using a Bloom filter to construct a hash index for the extracted features, mapping records into a binary vector through multiple hash functions. For different records, use a Siamese network model based on deep learning to calculate the similarity between records. The Siamese network consists of two sub-networks with shared weights. Input two records into the two sub-networks respectively, output feature vectors, calculate the similarity between the feature vectors through the Euclidean distance, and mark the records with similarity higher than the set threshold as duplicate records.

5. A method for online real-time duplicate checking and analysis applicable to libraries according to claim 1, characterized in that The mining of frequent item sets and association rules by calculating the support and confidence of item sets includes adopting a dynamic threshold adjustment strategy, conducting clustering analysis on historical data, dividing it into multiple clusters according to different characteristics of the data, calculating the data standard deviation and coefficient of variation of each cluster to evaluate the data dispersion degree, determining the minimum support threshold required for mining frequent item sets according to the data dispersion degree, using the weighted Apriori algorithm, assigning weights to each item, calculating the support according to the assigned weights, and screening out frequent item sets through the support and the minimum support threshold.

6. The method for online real-time duplicate checking and analysis applicable to libraries according to claim 5, wherein, Mining frequent item sets and association rules by calculating the support and confidence of item sets further includes calculating the confidence based on the frequent item sets, generating association rules when the confidence is greater than the set confidence threshold, and using an incremental association rule update method to update the association rules. The confidence calculation formula is: , Among them, represents the association rule from item set X to item set Y, represents the weighted support of the union of item set X and item set Y, represents the weighted support of item set X.

7. The method for online real-time duplicate checking and analysis applicable to libraries according to claim 1, wherein Using the K-Means clustering algorithm to classify library resources includes determining the optimal k value using the silhouette coefficient, selecting the feature vectors of k book resources as the initial clustering centers, calculating the distance between the feature vector of each book and each clustering center according to the Euclidean distance, assigning the feature vector to the cluster where the nearest clustering center is located, calculating the mean value of all book feature vectors within the cluster, and using it as the new clustering center, and evaluating the book resources according to the clustering results, and analyzing the duplication and utilization rate of book resources within each cluster.

8. An online real-time duplicate checking and analysis system applicable to libraries, which executes the method described in any one of claims 1-7, characterized in that, Including: A data acquisition module configured to: acquire a multi-source heterogeneous data set of a library management system and perform a preliminary screening on the acquired data set; A feature extraction module configured to: perform in-depth data cleaning and feature extraction based on the screened data set, including processing missing values using a missing value filling method based on a Bayesian network and performing feature vectorization using the TF-IDF algorithm; A duplicate checking module configured to: perform real-time duplicate checking based on the extracted feature vectors, including constructing a hash index and calculating the similarity between records, and finally marking the duplicate checking results; An association module configured to: perform association rule mining on the duplicate checking results, including mining frequent item sets and association rules by calculating the support and confidence of item sets; A resource classification module configured to: classify and evaluate library resources using association rules, including classifying library resources using the K-Means clustering algorithm; An output module configured to: perform visual output according to the results of association rule mining and clustering analysis.

9. A computer-readable storage medium storing a plurality of instructions, characterized in that, The instructions are suitable for being loaded and executed by a processor of a terminal device for a method for online real-time duplicate checking and analysis applicable to a library as described in any one of claims 1-7.

10. A terminal device, comprising a processor and a computer-readable storage medium, the processor being configured to implement each instruction; the computer-readable storage medium being configured to store a plurality of instructions, characterized in that, The instructions are suitable for being loaded and executed by a processor for a method for online real-time duplicate checking and analysis applicable to a library as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Standard element duplicate checking method based on text mining

    CN116629228A

  • Application architecture duplicate checking method and system, terminal and storage medium

    CN117851539A

  • Case investigation auxiliary system based on multi-source data association analysis

    CN118643465A

  • Healthcare claims fraud, waste and abuse detection system using non-parametric statistics and probability based scores

    US20170017760A1

  • Analysis method and apparatus for operating data of network function virtualization device

    WO2022083576A1