Online intelligent archive management system based on big data analysis

By designing an online intelligent archive management system based on big data analysis and using natural language processing and machine learning algorithms for document analysis and classification, the problems of inefficiency and classification errors of traditional archive management systems are solved, and efficient and intelligent archive management and storage are achieved.

CN119988709AInactive Publication Date: 2025-05-13JIANGSU VOCATIONAL COLLEGE OF BUSINESS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510079666.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-18
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional archive management systems rely on manual operations, are inefficient and prone to problems such as classification errors, missed archives or lag in archives, and cannot deeply analyze document content and lack the application of big data analysis and machine learning technology.

Method used

Design an online intelligent archive management system based on big data analysis, including document data acquisition and preprocessing module, text content analysis and feature extraction module, document classification and clustering module, automatic archiving and storage module, intelligent monitoring and early warning module, and analysis report and feedback optimization module, and use natural language processing and machine learning algorithms for document analysis and classification.

Benefits of technology

It realizes efficient processing of multiple document formats, automatically analyzing and classifying documents, optimizes storage paths, improves archive efficiency and accuracy, enhances document retrieval capabilities, timely identify and warns of potential storage and access abnormalities, and improves the intelligence and automation level of archive management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988709A_ABST
    Figure CN119988709A_ABST
Patent Text Reader

Abstract

The invention discloses an online intelligent archive management system based on big data analysis, relates to the technical field of archive management, and can efficiently process various formats of documents including PDF and Word and image documents including scanning documents and pictures through the design of a document data acquisition and preprocessing module. And the document format recognition and conversion unit converts unstructured image data into processable text data by utilizing an optical character recognition technology, so that the system can be compatible with different types of input files. The document preprocessing and standardizing unit further cleans and standardizes the extracted texts, and unifies the texts into structured high-quality data. A document management system capable of adaptively changing is established by combining a dynamically generated storage path and monitored access data. No matter a new document is added or an access mode is changed, the system can make a quick response through feature matching, rule adjustment and abnormal early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of archive management, and in particular to an online intelligent archive management system based on big data analysis. Background Art

[0002] In the field of information management, archive management is a key branch, and its main task is to scientifically classify, store, retrieve and maintain various documents, files and data resources in order to support the business operations of enterprises, governments or institutions. In archive management, with the exponential growth of information volume, traditional manual archiving methods have been unable to meet management needs, especially for specific document management scenarios such as contracts, financial documents, and employee information that involve large amounts of text analysis and sensitive information processing, which require more intelligent solutions.

[0003] In the current archive management work, classification and filing still rely on a lot of manual operations, which is not only inefficient but also prone to human errors, such as classification errors, missing documents or delayed filing. Especially when dealing with complex archives such as contracts and financial documents, manual work requires a lot of time to analyze document content, extract key information and manually classify them, which greatly restricts work efficiency.

[0004] Traditional archive management systems are unable to conduct in-depth analysis of document content, lack the ability to utilize big data analysis and machine learning technologies, and rely solely on static rule matching and manual judgment. In addition, manual archiving operations are easily affected by personal subjective factors, resulting in inconsistent classification or misjudgment, further increasing the difficulty of subsequent document retrieval and management. These problems can lead to a series of abnormal effects: for example, important archives cannot be located and called in time in actual needs, resulting in reduced business decision-making efficiency; or important information is lost due to archiving errors, causing data security risks. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention provides an online intelligent archive management system based on big data analysis, which solves the problems mentioned in the background technology.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an online intelligent archive management system based on big data analysis, including a document data collection and preprocessing module, a text content analysis and feature extraction module, a document classification and clustering module, an automatic archiving and storage module, an intelligent monitoring and early warning module, and an analysis report and feedback optimization module;

[0007] The document data collection and preprocessing module collects documents uploaded by users, and converts image documents into text content through optical character recognition technology, performs preprocessing, and obtains processed documents;

[0008] The text content analysis and feature extraction module analyzes the content of the processed document by using a natural language processing algorithm, and extracts features, including keyword K, entity B and time information T features, to form a feature set FV, and converts it into a feature vector FVe;

[0009] The document classification and clustering module classifies the processed documents according to the acquired feature vector FVe and in combination with the K-means clustering algorithm;

[0010] The automatic filing and storage module generates classification rules based on the classified documents. The input new documents will automatically specify the storage path according to the classification rules, and store and file them according to the document keyword K, entity B and time information T factors;

[0011] The intelligent monitoring and early warning module is responsible for monitoring the storage and archiving of archives, and checking the storage status, access frequency Fw and abnormal behavior of new documents in real time;

[0012] The analysis report and feedback optimization module generates an analysis report based on the document storage status and access frequency Fw, providing managers with comprehensive information on archive management.

[0013] Preferably, the document data acquisition and preprocessing module includes a document format recognition and conversion unit and a document preprocessing and standardization unit;

[0014] The document format recognition and conversion unit is responsible for recognizing the format of the document uploaded by the user and processing it according to the document type;

[0015] Specifically, for text-formatted documents, including PDF and Word, the text content is directly extracted; for image-based documents, including scanned copies and picture formats, the image is converted into processable text by using optical character recognition technology;

[0016] The document preprocessing and standardization unit cleans and standardizes the extracted text to obtain a processed document;

[0017] Cleaning includes denoising, removing extra blank characters, special symbols, page numbers, and document headers and tails;

[0018] Standardization processing includes case conversion and text normalization; case conversion includes converting all texts into a unified format, and text normalization includes spelling correction of words, removal of stop words, and synonym replacement.

[0019] Preferably, the text content analysis and feature extraction module includes a feature recognition and extraction unit and a feature quantization and structuring unit;

[0020] The feature recognition and extraction unit analyzes the processed document content by using a natural language processing algorithm, and recognizes the keyword K, entity B and time information T in the processed document through keyword extraction, named entity recognition and time extraction, and fits them into a feature set FV;

[0021] Keyword extraction includes extracting core words from documents, including income, expenditure, and profit in financial documents, and clauses, responsibilities, and deadlines in contracts;

[0022] Named entity recognition involves identifying and labeling entities in documents, including names of people, companies, amounts, transaction information, signatories, and terms;

[0023] Time information T extraction includes identifying time and date information appearing in documents, including the timeliness of contracts and the time nodes of financial statements.

[0024] Preferably, the feature vectorization and structuring unit vectorizes the acquired feature set FV by using the TF-IDF calculation method to generate a structured feature vector FVe; by vectorizing the features, the system can further analyze, classify or cluster the documents.

[0025] The TF-IDF calculation method includes calculating the TF-IDF weights of the extracted features and assigning weight values;

[0026] The internal features of the feature set FV are marked and expressed as FV = {f1, f2, ..., fi}; and the fi-th feature of the feature set FV is weightedly calculated by the TF-IDF calculation method, and the calculation results are fitted to obtain the feature vector FVe.

[0027] Preferably, the document classification and clustering module includes a feature vector clustering unit and a classification and label assignment unit;

[0028] The feature vector clustering unit inputs the feature vector FVe into the K-means clustering algorithm, and uses K-means clustering to classify the processed documents to obtain the clustering result Clusterj, and divides them into different categories, including financial documents, contracts and employee information; the goal of the K-means algorithm is to divide the documents into k clusters, with the document similarity within each cluster being high, and the document similarity between different clusters being low.

[0029] The processed documents are divided into k clusters by using the K-means algorithm. The center of each cluster is Cj. The distance D from the document to the cluster center is calculated, and the feature vector FVe of the document will be assigned to the nearest cluster center Cj. The goal of K-means clustering is to minimize the distance between each document and its cluster center. Through multiple iterations, the cluster centers gradually converge and finally a stable clustering result is obtained.

[0030] The distance D from the document to the cluster center is obtained by the following formula:

[0031]

[0032] Where D(FVe, Cj) represents the distance from the feature vector FVe to the cluster center Cj, Cj,i = {cj,1, cj,2, ..., cj,i} represents the dimension value of the cluster center j in the i-th feature, fi represents the i-th feature in the feature vector FVe, and n represents the total number of clusters;

[0033] Assign documents to the nearest cluster center based on the minimum distance principle:

[0034] FVe∈Clusterj if D(FVe,Cj)≤D(FVe,Ci)for i≠j;

[0035] In the formula, FVe represents the feature vector of the document, Cj and Ci represent the centers of cluster j and cluster i respectively, and D(FVe, Ci) represents the distance from the feature vector FVe to the cluster center Ci;

[0036] The classification and label assignment unit assigns a corresponding label L to each cluster j;

[0037] The label L is obtained by the following formula:

[0038] L(FVe)=Lj if FVe∈Clusterj;

[0039] Where L(FVe) represents the assigned label of the feature vector FVe, and Lj represents the label of cluster j.

[0040] Preferably, the automatic filing and storage module includes a storage path generation and rule application unit and a document filing and storage unit;

[0041] The storage path generation and rule application unit analyzes the classified documents, extracts the keyword K, entity B and time information T features, generates classification rules, and generates a storage path Path according to the classification rules;

[0042] When a new document is input, features are extracted from the new document, including keyword K, entity B and time information T features, and matched with the features of the classified document. Classification rules are selected and the document is stored according to the storage path Path. Predefined path generation rules including "document category / keyword K / time / document name" are applied to determine the path for document storage. Path rules can also be dynamically adjusted according to actual needs, for example, paths can be divided by different dimensions such as year, month, department, etc.

[0043] The storage path Path is obtained by the following formula:

[0044] Path = h(K, B, T);

[0045] In the formula, K represents keywords, including the contract signatory and amount, B represents entities, including financial documents, contracts, and employee information, T represents time information, including the contract signing date and the financial report release date, and h represents the function that generates the storage path. It can be divided into the following two forms:

[0046] Path=Category / Keywords / Year / DocumentName;

[0047] Path = category / keyword / year / document name;

[0048] Path=Category / Entity1 / Entity2 / Time;

[0049] Path = category / entity1 / entity2 / time.

[0050] Preferably, the document archiving and storage unit is responsible for applying the generated storage path Path to the database, archiving and storing the new document according to the storage path Path, and interacting with the database to archive the content, classification information and path of the document;

[0051] Specifically, the document is stored in a specified location according to the generated storage path Path, and the features of the new document are stored as document retrieval information, including the keyword K, the entity B, and the time information T;

[0052] The storage formula of the feature is:

[0053] FV 文档 ={L 文档 ,K 文档 ,B 文档 ,T 文档 ,Path 文档};

[0054] Where FV 文档represents the features of the feature set FV in the document, L 文档 represents the classification label of the document, K 文档 Represents the keyword set of the document, B 文档 Represents the entity information in the document, T 文档 Indicates the time information of the document, Path 文档 Indicates the document storage path.

[0055] Preferably, the intelligent monitoring and early warning module includes a data acquisition and analysis unit and an anomaly detection and early warning generation unit;

[0056] The data collection and analysis unit collects archive storage and access data and performs real-time analysis to identify potential abnormal behaviors by analyzing the access frequency Fw, visitor behavior, and storage status of the document;

[0057] Based on the collected access data, the access frequency Fw of each document is calculated and compared with the access threshold TFw under the normal access mode to identify abnormal access behavior status;

[0058] The access frequency Fw is obtained by the ratio of the number of accesses to the document to the time period to which the number of accesses belongs;

[0059] The access threshold TFw is obtained by the following formula:

[0060] TFw=[μFw-k*σFw,μFw+k*σFw];

[0061] In the formula, μFw represents the mean of the access frequency, σFw represents the standard deviation of the access frequency, and k represents the adjustment factor;

[0062] The abnormal behavior status of the access is obtained by matching in the following ways:

[0063] When the access frequency Fw is within the range of the access threshold TFw, it means that the access frequency is normal;

[0064] When the access frequency Fw is not within the range of the access threshold TFw, it indicates that the access frequency is abnormal.

[0065] Preferably, the anomaly detection and warning generation unit analyzes the storage status, access frequency Fw and abnormal behavior status of access through an anomaly detection algorithm, identifies potential problems, and generates an alarm to notify the administrator based on the abnormal behavior status;

[0066] The abnormal behavior states include abnormal access frequency, abnormal storage capacity, and damaged document integrity;

[0067] When the access frequency is abnormal, the access abnormality score SF is calculated;

[0068] The access anomaly score SF is obtained by the ratio of the absolute value of the difference between the access frequency Fw and the mean μFw of the access frequency to the standard deviation σFw of the access frequency;

[0069] When the storage capacity is abnormal, the storage occupancy ratio RS is calculated and compared with the preset capacity threshold TRS to determine the storage capacity status;

[0070] The storage occupancy ratio RS is obtained by the following formula:

[0071]

[0072] In the formula, CCt represents the total storage capacity, and CCa represents the remaining storage capacity;

[0073] When the storage occupancy ratio RS ≤ the capacity threshold TRS, no alarm is issued;

[0074] When the storage occupancy ratio RS>the capacity threshold TRS, an alarm of insufficient storage capacity is issued.

[0075] Preferably, the analysis report and feedback optimization module includes a report generation unit and a feedback optimization unit;

[0076] The report generation unit calculates the index based on the document storage status and the access frequency Fw, and generates a visual analysis report in the form of charts, text and summaries, which is provided to the manager for review;

[0077] Among them, the indicators include classification accuracy Ac, storage utilization Us, and access trend change rate Rtr;

[0078] The storage utilization rate Us is obtained by the ratio of the difference between the total storage capacity CCt and the remaining storage capacity CCa to the total storage capacity CCt;

[0079] The classification accuracy Ac is obtained by the ratio of the total number of classified documents Nto to the number of correctly classified documents Nco;

[0080] The change rate Rtr of the access trend is obtained by the ratio of the difference between the access frequency Fw at time t and the access frequency Fw at time t+M to the time interval (t+M)-t;

[0081] The feedback optimization unit retrains the K-means clustering algorithm according to the classification accuracy Ac, adjusts the storage path Path according to the storage utilization Us and the access trend change rate Rtr, calculates and obtains the new storage path NPath, and stores the document;

[0082] The new storage path NPath is obtained by the following formula:

[0083] NPath=h(Path,Ac,Us,Rtr);

[0084] Where h represents the function of generating the storage path.

[0085] The present invention provides an online intelligent archive management system based on big data analysis, which has the following beneficial effects:

[0086] (1) When the system is running, through the design of the document data acquisition and preprocessing module, this embodiment can efficiently process documents in various formats, including PDF, Word, and image documents including scans and pictures. The document format recognition and conversion unit uses optical character recognition technology to convert unstructured image data into processable text data, ensuring that the system is compatible with different types of input files. The document preprocessing and standardization unit further cleans and normalizes the extracted text, unifying the text into structured high-quality data. This improvement solves the problem of insufficient processing capabilities of traditional systems for diversified documents, laying a good foundation for subsequent intelligent analysis and archiving.

[0087] The text content analysis and feature extraction module uses the natural language processing algorithm to conduct in-depth analysis of the processed text content through the collaboration of the feature recognition and extraction unit and the feature vectorization and structuring unit, and extracts key features such as keyword K, entity B and time information T. Keyword extraction is targeted at the core content of the document, including income, expenditure and profit in financial documents, and terms, responsibilities and deadlines in contracts, for accurate extraction.

[0088] (2) The feature vector clustering unit processes the feature vector FVe through the Z-score normalization method to unify the dimension of the feature and avoid the influence of dimension difference between features on the clustering effect. Then, the K-means clustering algorithm is used to divide the documents into k clusters. The center Cj of each cluster represents the feature core of the document. The document is assigned to the nearest cluster center Cj according to the minimum distance principle to complete the classification process.

[0089] The classification and label assignment unit assigns a corresponding label Lj to each cluster j, including financial documents, contracts, and employee information. The label assignment formula is based on the characteristic characteristics of the cluster center and dynamically generates labels that best match the cluster characteristics. This approach not only ensures the accuracy of the classification results, but also enables the system to dynamically optimize the classification rules according to the changes in document data, thereby improving the system's adaptability to new documents. This improvement solves the drawbacks of fixed and difficult-to-adapt classification rules in traditional systems, and realizes the dynamic and flexible allocation of classification labels.

[0090] (3) Through the document archiving and storage unit, the system archives the content, classification information, key features and storage path Path of the document into the database. This method not only optimizes the physical storage of the document, but also stores the feature information as the retrieval basis, forming an efficient document retrieval mechanism. Compared with the traditional method of relying only on file name or fixed tag retrieval, this embodiment realizes accurate document retrieval under multi-dimensional and multi-conditions through the storage of key features, effectively improving the efficiency of information acquisition. Through the data acquisition and analysis unit in the intelligent monitoring and early warning module, the storage and access data of the document are collected in real time, the access frequency Fw is calculated, and compared with the threshold TFw under the normal access mode to identify potential abnormal behavior.

[0091] (4) Through the anomaly detection and warning generation unit, key indicators such as storage status, access frequency Fw and document integrity are monitored and analyzed in real time, and the abnormal status is quantified using the access anomaly score SF and storage occupancy ratio RS algorithm. For problems such as abnormal access frequency, insufficient storage capacity or damaged document integrity, the system can generate targeted warnings and notify the administrator. Compared with the traditional static rule detection method, the anomaly detection of this embodiment is based on dynamic data analysis and score calculation, which improves the accuracy of detection and the timeliness of response, and effectively reduces the security risks in archive management.

[0092] By monitoring the storage utilization rate Us, access trend change rate Rt and other indicators, the storage resource usage status and document access rules can be identified. Based on these data, the feedback optimization unit dynamically adjusts the storage path Path and classification rules, and calculates and generates a new storage path NPath in combination with the adjusted K-means clustering algorithm. This real-time optimization capability solves the problem of fixed storage paths and uneven resource allocation in traditional systems, enabling the system to automatically adjust the storage priority of documents when the storage capacity approaches the threshold, thereby improving the utilization efficiency of storage resources and the responsiveness of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 This is a flowchart diagram of an online intelligent archive management system based on big data analysis in the present invention. DETAILED DESCRIPTION

[0094] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0095] Example 1

[0096] The present invention provides an online intelligent archive management system based on big data analysis. Figure 1 , including document data collection and preprocessing module, text content analysis and feature extraction module, document classification and clustering module, automatic archiving and storage module, intelligent monitoring and early warning module and analysis report and feedback optimization module;

[0097] The document data collection and preprocessing module collects documents uploaded by users, and converts image documents into text content through optical character recognition technology, performs preprocessing, and obtains processed documents;

[0098] The text content analysis and feature extraction module analyzes the content of the processed document by using a natural language processing algorithm, and extracts features, including keyword K, entity B and time information T features, to form a feature set FV, and converts it into a feature vector FVe;

[0099] The document classification and clustering module classifies the processed documents according to the acquired feature vector FVe and in combination with the K-means clustering algorithm;

[0100] The automatic filing and storage module generates classification rules based on the classified documents. The input new documents will automatically specify the storage path according to the classification rules, and store and file them according to the document keyword K, entity B and time information T factors;

[0101] The intelligent monitoring and early warning module is responsible for monitoring the storage and archiving of archives, and checking the storage status, access frequency Fw and abnormal behavior of new documents in real time;

[0102] The analysis report and feedback optimization module generates an analysis report based on the document storage status and access frequency Fw, providing managers with comprehensive information on archive management.

[0103] In this embodiment, through the combination of the document data collection and preprocessing module and the text content analysis and feature extraction module, the system uses optical character recognition technology and natural language processing algorithms to digitally transform, clean and extract features from the document content uploaded by the user. The extracted keywords K, entity B and time information T form a feature set FV, which is further converted into a feature vector FVe. This process realizes the automatic analysis and feature extraction of document content, avoids classification errors and inefficiencies caused by human errors in the traditional manual filing mode, and thus significantly improves the filing efficiency and accuracy.

[0104] The document classification and clustering module combines the K-means clustering algorithm to automatically assign documents to the most appropriate categories, including contracts, financial documents, and employee information, by analyzing the feature vector FVe. This process uses the automatic learning ability of machine learning to efficiently classify new documents without manual intervention. This improvement solves the shortcomings of traditional archiving systems that rely on static rules and realizes the dynamic optimization and self-adaptation of classification algorithms.

[0105] The automatic archiving and storage module automatically generates storage paths for new documents based on document classification results and feature information. The generation rules of the storage path are dynamically adjusted according to document attributes and classifications, thereby ensuring that the archives are organized and easy to retrieve in physical storage. This improvement solves the problem of static storage paths in traditional systems that are difficult to adapt to dynamic needs, while optimizing the utilization efficiency of storage resources.

[0106] The intelligent monitoring and early warning module can generate early warnings and notify managers in a timely manner when the access frequency increases abnormally or the storage status is abnormal, through real-time monitoring of the document storage status, access frequency Fw and abnormal behavior. This function significantly improves the system's ability to respond to abnormal file conditions and avoids performance degradation and security risks caused by excessive storage pressure or frequent access.

[0107] The analysis report and feedback optimization module generates a comprehensive management analysis report based on storage utilization, classification accuracy and access frequency Fw, helping managers understand the current status of archive management and potential optimization space. In addition, the module can also use feedback data to optimize classification algorithms and archiving strategies, thereby improving the intelligence level and adaptive capabilities of the system.

[0108] Example 2

[0109] This embodiment is explained in Example 1, please refer to Figure 1 ,Specifically: the document data acquisition and preprocessing module includes a document format recognition and conversion unit and a document preprocessing and standardization unit;

[0110] The document format recognition and conversion unit is responsible for recognizing the format of the document uploaded by the user and processing it according to the document type;

[0111] Specifically, for text-formatted documents, including PDF and Word, the text content is directly extracted; for image-based documents, including scanned copies and picture formats, the image is converted into processable text by using optical character recognition technology;

[0112] The document preprocessing and standardization unit cleans and standardizes the extracted text to obtain a processed document;

[0113] Cleaning includes denoising, removing extra blank characters, special symbols, page numbers, and document headers and tails;

[0114] Standardization processing includes case conversion and text normalization; case conversion includes converting all texts into a unified format, and text normalization includes spelling correction of words, removal of stop words, and synonym replacement.

[0115] The text content analysis and feature extraction module includes a feature recognition and extraction unit and a feature vectorization and structuring unit;

[0116] The feature recognition and extraction unit analyzes the processed document content by using a natural language processing algorithm, and recognizes the keyword K, entity B and time information T in the processed document through keyword extraction, named entity recognition and time extraction, and fits them into a feature set FV;

[0117] Keyword extraction includes extracting core words from documents, including income, expenditure, and profit in financial documents, and clauses, responsibilities, and deadlines in contracts;

[0118] Named entity recognition involves identifying and labeling entities in documents, including names of people, companies, amounts, transaction information, signatories, and terms;

[0119] Time information T extraction includes identifying time and date information appearing in documents, including the timeliness of contracts and the time nodes of financial statements.

[0120] In this embodiment, through the design of the document data acquisition and preprocessing module, this embodiment can efficiently process documents in various formats, including PDF, Word, and image documents including scans and pictures. The document format recognition and conversion unit uses optical character recognition technology to convert unstructured image data into processable text data, ensuring that the system is compatible with different types of input files. The document preprocessing and standardization unit further cleans and normalizes the extracted text, unifying the text into structured high-quality data. This improvement solves the problem of insufficient processing capabilities of traditional systems for diversified documents, and lays a good foundation for subsequent intelligent analysis and archiving.

[0121] The text content analysis and feature extraction module uses the natural language processing algorithm to conduct in-depth analysis of the processed text content through the collaboration of the feature recognition and extraction unit and the feature vectorization and structuring unit, and extracts key features such as keyword K, entity B and time information T. Keyword extraction targets the core content of the document, including income, expenditure and profit in financial documents, and accurately extracts the terms, responsibilities and deadlines in the contract; named entity recognition effectively extracts important entity information in the document, including names, company names, amounts and signatories; time information extraction focuses on important time nodes in the document, including the effective date of the contract and the time of the financial statements. These features are fitted into the feature set FV and converted into feature vectors FVe, providing efficient support for document classification and archiving.

[0122] By cleaning, normalizing and extracting features from document content, this embodiment significantly improves the intelligence of archiving. The processed document features are uniformly structured to meet the needs of automatic classification and storage rules. Compared with the traditional method that relies on manual analysis and manual archiving, this embodiment eliminates errors and redundancy in manual operations by accurately extracting key content, greatly improving archiving efficiency and quality.

[0123] Example 3

[0124] This embodiment is explained in Example 2. Please refer to Figure 1 Specifically: the feature vectorization and structuring unit performs vector transformation on the acquired feature set FV by using the TF-IDF calculation method to generate a structured feature vector FVe;

[0125] The TF-IDF calculation method includes calculating the TF-IDF weights of the extracted features and assigning weight values;

[0126] The internal features of the feature set FV are marked and expressed as FV = {f1, f2, ..., fi}; and the fi-th feature of the feature set FV is weightedly calculated by the TF-IDF calculation method, and the calculation results are fitted to obtain the feature vector FVe.

[0127] The document classification and clustering module includes a feature vector clustering unit and a classification and label assignment unit;

[0128] The feature vector clustering unit inputs the feature vector FVe into the K-means clustering algorithm and uses K-means clustering to classify the processed documents, obtains the clustering result Clusterj, and divides it into different categories, including financial documents, contracts and employee information;

[0129] The processed documents are divided into k clusters by using the K-means algorithm. The center of each cluster is Cj. The distance D from the document to the cluster center is calculated. The feature vector FVe of the document will be assigned to the cluster center Cj with the closest distance.

[0130] The distance D from the document to the cluster center is obtained by the following formula:

[0131]

[0132] Where D(FVe, Cj) represents the distance from the feature vector FVe to the cluster center Cj, Cj,i = {cj,1, cj,2, ..., cj,i} represents the dimension value of the cluster center j in the i-th feature, fi represents the i-th feature in the feature vector FVe, and n represents the total number of clusters;

[0133] Assign documents to the nearest cluster center based on the minimum distance principle:

[0134] FVe∈Clusterj if D(FVe,Cj)≤D(FVe,Ci)for i≠j;

[0135] In the formula, FVe represents the feature vector of the document, Cj and Ci represent the centers of cluster j and cluster i respectively, and D(FVe, Ci) represents the distance from the feature vector FVe to the cluster center Ci;

[0136] The classification and label assignment unit assigns a corresponding label L to each cluster j;

[0137] The label L is obtained by the following formula:

[0138] L(FVe)=Lj if FVe∈Clusterj;

[0139] Where L(FVe) represents the assigned label of the feature vector FVe, and Lj represents the label of cluster j.

[0140] In this embodiment, the TF-IDF algorithm is used to calculate the weight of the feature set FV = {f1, f2, ..., fi} to generate a structured feature vector FVe. The weight value TF-IDF (fi) of each feature fi reflects the importance of the feature in the document. Through the vectorization process, the system can convert the unstructured document content into a numerical feature representation, providing standardized input for subsequent classification and clustering. This improvement solves the problem in traditional archive management that documents cannot be directly classified and analyzed due to their unstructured nature, laying the foundation for intelligent archive management.

[0141] The feature vector clustering unit processes the feature vector FVe through the Z-score normalization method to unify the dimension of the feature and avoid the influence of dimension difference between features on the clustering effect. Then, the K-means clustering algorithm is used to divide the documents into k clusters, and the center Cj of each cluster represents the feature core of the document. The document is assigned to the nearest cluster center Cj according to the minimum distance principle to complete the classification process.

[0142] The classification and label assignment unit assigns a corresponding label Lj to each cluster j, including financial documents, contracts, and employee information. The label assignment formula is based on the characteristic characteristics of the cluster center and dynamically generates labels that best match the cluster characteristics. This approach not only ensures the accuracy of the classification results, but also enables the system to dynamically optimize the classification rules according to the changes in document data, thereby improving the system's adaptability to new documents. This improvement solves the drawbacks of fixed and difficult-to-adapt classification rules in traditional systems, and realizes the dynamic and flexible allocation of classification labels.

[0143] By calculating the distance D(FVe,Cj) from the document to the cluster center, the system can intelligently assign the document to the closest category. This distance-based classification method not only improves the accuracy of classification, but also continuously optimizes the classification rules and results by dynamically adjusting the position of the cluster center Cj. By combining TF-IDF feature weight calculation, Z-score standardization processing and K-means clustering algorithm, a complete intelligent classification process is constructed. The system can dynamically update the cluster center and label assignment rules according to changes in document features, and supports extended processing of different types of documents. This improvement improves the scalability and intelligence of the system, enabling it to maintain efficient classification performance when the number of documents increases or the complexity of classification increases.

[0144] Example 4

[0145] This embodiment is explained in Example 3, please refer to Figure 1 ,Specifically: the automatic archiving and storage module includes a storage path generation and rule application unit and a document archiving and storage unit;

[0146] The storage path generation and rule application unit analyzes the classified documents, extracts the keyword K, entity B and time information T features, generates classification rules, and generates a storage path Path according to the classification rules;

[0147] When a new document is input, the features of the new document are extracted, including the keyword K, entity B and time information T features, and matched with the features of the classified document, and the classification rules are selected and stored according to the storage path Path;

[0148] The storage path Path is obtained by the following formula:

[0149] Path = h(K, B, T);

[0150] Where h represents the function of generating the storage path.

[0151] The document archiving and storage unit is responsible for applying the generated storage path Path to the database, archiving and storing the new document according to the storage path Path, and interacting with the database to archive the content, classification information and path of the document;

[0152] Specifically, the document is stored in a specified location according to the generated storage path Path, and the features of the new document are stored as document retrieval information, including the keyword K, the entity B, and the time information T;

[0153] The storage formula of the feature is:

[0154] FV 文档 ={L 文档 ,K 文档 ,B 文档 ,T 文档 ,Path 文档};

[0155] Where FV 文档 represents the features of the feature set FV in the document, L 文档 represents the classification label of the document, K 文档 Represents the keyword set of the document, B 文档 Represents the entity information in the document, T 文档 Indicates the time information of the document, Path 文档 Indicates the document storage path.

[0156] The intelligent monitoring and early warning module includes a data acquisition and analysis unit and an anomaly detection and early warning generation unit;

[0157] The data collection and analysis unit collects archive storage and access data and performs real-time analysis to identify potential abnormal behaviors by analyzing the access frequency Fw, visitor behavior, and storage status of the document;

[0158] Based on the collected access data, the access frequency Fw of each document is calculated and compared with the access threshold TFw under the normal access mode to identify abnormal access behavior status;

[0159] The access frequency Fw is obtained by the ratio of the number of accesses to the document to the time period to which the number of accesses belongs;

[0160] The access threshold TFw is obtained by the following formula:

[0161] TFw=[μFw-k*σFw,μFw+k*σFw];

[0162] In the formula, μFw represents the mean of the access frequency, σFw represents the standard deviation of the access frequency, and k represents the adjustment factor;

[0163] The abnormal behavior status of the access is obtained by matching in the following ways:

[0164] When the access frequency Fw is within the range of the access threshold TFw, it means that the access frequency is normal;

[0165] When the access frequency Fw is not within the range of the access threshold TFw, it indicates that the access frequency is abnormal.

[0166] In this embodiment, the storage path generation and rule application unit dynamically generates a storage path Path=h(K, B, T) based on the classified document features, including keyword K, entity B, and time information T. For new documents, the system can automatically match the classification rules and specify the storage path based on the extracted features, thus realizing the dynamic and intelligent storage path. This improvement solves the problem of static storage paths in traditional systems, allowing storage rules to be automatically adjusted as document content and classification results change, significantly improving the flexibility and efficiency of storage management.

[0167] Through the document archiving and storage unit, the system archives the document's content, classification information, key features, and storage path Path into the database. This method not only optimizes the physical storage of documents, but also stores feature information as a retrieval basis, forming an efficient document retrieval mechanism. Compared with the traditional method of relying solely on file names or fixed tags for retrieval, this embodiment achieves accurate document retrieval under multiple dimensions and conditions through the storage of key features, effectively improving the efficiency of information acquisition. Through the data acquisition and analysis unit in the intelligent monitoring and early warning module, the storage and access data of the document are collected in real time, the access frequency Fw is calculated, and compared with the threshold TFw under the normal access mode to identify potential abnormal behaviors.

[0168] By calculating the access frequency Fw and the access threshold TFw, the system can generate an early warning in time when the document access frequency increases abnormally, prompting the administrator to pay attention to the frequently accessed documents and their storage status. This mechanism not only effectively avoids system load problems caused by frequent access, but also can discover potential risk documents. In addition, real-time monitoring can also help managers optimize storage resource allocation and improve the overall storage utilization of the system.

[0169] By combining dynamically generated storage paths and real-time monitored access data, a document management system that can adapt to changes has been established. Whether it is a new document or a change in access mode, the system can respond quickly through feature matching, rule adjustment, and abnormal warning. This dynamic adaptability enables the system to maintain efficient operation in complex and changing archive management scenarios, while providing a good technical foundation for expansion and upgrading.

[0170] Example 5

[0171] This embodiment is explained in Example 4. Please refer to Figure 1 ,Specifically: the anomaly detection and warning generation unit analyzes the storage status, access frequency Fw and abnormal behavior status of access through an anomaly detection algorithm, identifies potential problems, and generates an early warning to notify the administrator according to the abnormal behavior status;

[0172] The abnormal behavior states include abnormal access frequency, abnormal storage capacity, and damaged document integrity;

[0173] When the access frequency is abnormal, the access abnormality score SF is calculated;

[0174] The access anomaly score SF is obtained by the ratio of the absolute value of the difference between the access frequency Fw and the mean μFw of the access frequency to the standard deviation σFw of the access frequency;

[0175] When the storage capacity is abnormal, the storage occupancy ratio RS is calculated and compared with the preset capacity threshold TRS to determine the storage capacity status;

[0176] The storage occupancy ratio RS is obtained by the following formula:

[0177]

[0178] In the formula, CCt represents the total storage capacity, and CCa represents the remaining storage capacity;

[0179] When the storage occupancy ratio RS ≤ the capacity threshold TRS, no alarm is issued;

[0180] When the storage occupancy ratio RS>the capacity threshold TRS, an alarm of insufficient storage capacity is issued.

[0181] The analysis report and feedback optimization module includes a report generation unit and a feedback optimization unit;

[0182] The report generation unit calculates the index based on the document storage status and the access frequency Fw, and generates a visual analysis report in the form of charts, text and summaries, which is provided to the manager for review;

[0183] Among them, the indicators include classification accuracy Ac, storage utilization Us, and access trend change rate Rtr;

[0184] The storage utilization rate Us is obtained by the ratio of the difference between the total storage capacity CCt and the remaining storage capacity CCa to the total storage capacity CCt;

[0185] The classification accuracy Ac is obtained by the ratio of the total number of classified documents Nto to the number of correctly classified documents Nco;

[0186] The change rate Rtr of the access trend is obtained by the ratio of the difference between the access frequency Fw at time t and the access frequency Fw at time t+M to the time interval (t+M)-t;

[0187] The feedback optimization unit retrains the K-means clustering algorithm according to the classification accuracy Ac, adjusts the storage path Path according to the storage utilization Us and the access trend change rate Rtr, calculates and obtains the new storage path NPath, and stores the document;

[0188] The new storage path NPath is obtained by the following formula:

[0189] NPath=h(Path,Ac,Us,Rtr);

[0190] Where h represents the function of generating the storage path.

[0191] In this embodiment, the storage status, access frequency Fw, and document integrity and other key indicators are monitored and analyzed in real time through the anomaly detection and warning generation unit, and the abnormal status is quantified using the access anomaly score SF and storage occupancy ratio RS algorithm. For problems such as abnormal access frequency, insufficient storage capacity, or damaged document integrity, the system can generate targeted warnings and notify the administrator. Compared with the traditional static rule detection method, the anomaly detection of this embodiment is based on dynamic data analysis and score calculation, which improves the accuracy of detection and the timeliness of response, and effectively reduces the security risks in archive management.

[0192] By monitoring the storage utilization rate Us, access trend change rate Rt and other indicators, the storage resource usage status and document access rules can be identified. Based on these data, the feedback optimization unit dynamically adjusts the storage path Path and classification rules, and calculates and generates a new storage path NPath in combination with the adjusted K-means clustering algorithm. This real-time optimization capability solves the problem of fixed storage paths and uneven resource allocation in traditional systems, enabling the system to automatically adjust the storage priority of documents when the storage capacity approaches the threshold, thereby improving the utilization efficiency of storage resources and the responsiveness of the system.

[0193] The analysis report and feedback optimization module generates a visual analysis report based on indicators such as document storage status, classification accuracy Ac, storage utilization Us, and access trend change rate Rtr through the report generation unit. The report content includes key information such as classification effect, storage status, and access behavior trends, providing managers with comprehensive data support to help them discover potential system problems and optimize archive management strategies. Compared with the limitations of traditional archive management systems that lack comprehensive analysis, the report generation capability of this embodiment provides strong support for data-driven scientific decision-making.

[0194] Through anomaly detection, resource optimization, and feedback learning, this embodiment significantly improves the operating efficiency, data storage reliability, and management security of the archive management system. Timely warning of access frequency anomalies reduces security risks caused by overload or unauthorized access; dynamic adjustment of storage paths effectively reduces the occurrence of storage bottlenecks; and continuous optimization of classification algorithms ensures the accuracy and consistency of classification results. Together, these functions ensure that the system can operate efficiently and stably in dynamically changing archive management scenarios.

[0195] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An online intelligent archive management system based on big data analysis, characterized by: It includes document data collection and preprocessing module, text content analysis and feature extraction module, document classification and clustering module, automatic archiving and storage module, intelligent monitoring and early warning module and analysis report and feedback optimization module; The document data collection and preprocessing module collects documents uploaded by users, and converts image documents into text content through optical character recognition technology, performs preprocessing, and obtains processed documents; The text content analysis and feature extraction module analyzes the content of the processed document by using a natural language processing algorithm, and extracts features, including keyword K, entity B and time information T features, to form a feature set FV, and converts it into a feature vector FVe; The document classification and clustering module classifies the processed documents according to the acquired feature vector FVe and in combination with the K-means clustering algorithm; The automatic filing and storage module generates classification rules based on the classified documents. The input new documents will automatically specify the storage path according to the classification rules, and store and file them according to the document keyword K, entity B and time information T factors; The intelligent monitoring and early warning module is responsible for monitoring the storage and archiving of archives, and checking the storage status, access frequency Fw and abnormal behavior of new documents in real time; The analysis report and feedback optimization module generates an analysis report based on the document storage status and access frequency Fw, providing managers with comprehensive information on archive management.

2. The online intelligent archive management system based on big data analysis according to claim 1 is characterized by: The document data acquisition and preprocessing module includes a document format recognition and conversion unit and a document preprocessing and standardization unit; The document format recognition and conversion unit is responsible for recognizing the format of the document uploaded by the user and processing it according to the document type; Specifically, for text-formatted documents, including PDF and Word, the text content is directly extracted; for image-based documents, including scanned copies and picture formats, the image is converted into processable text by using optical character recognition technology; The document preprocessing and standardization unit cleans and standardizes the extracted text to obtain a processed document; Cleaning includes denoising, removing extra blank characters, special symbols, page numbers, and document headers and tails; Standardization processing includes case conversion and text normalization; case conversion includes converting all texts into a unified format, and text normalization includes spelling correction, removal of stop words, and synonym replacement.

3. The online intelligent archive management system based on big data analysis according to claim 2 is characterized by: The text content analysis and feature extraction module includes a feature recognition and extraction unit and a feature vectorization and structuring unit; The feature recognition and extraction unit analyzes the processed document content by using a natural language processing algorithm, and recognizes the keyword K, entity B and time information T in the processed document through keyword extraction, named entity recognition and time extraction, and fits them into a feature set FV; Keyword extraction includes extracting core words from documents, including income, expenditure, and profit in financial documents, and clauses, responsibilities, and deadlines in contracts; Named entity recognition involves identifying and labeling entities in documents, including names of people, companies, amounts, transaction information, signatories, and terms; Time information T extraction includes identifying the time and date information appearing in the document, the timeliness of the contract, and the time nodes of the financial statements.

4. The online intelligent archive management system based on big data analysis according to claim 3 is characterized by: The feature vectorization and structuring unit performs vector transformation on the acquired feature set FV by using the TF-IDF calculation method to generate a structured feature vector FVe; The TF-IDF calculation method includes calculating the TF-IDF weights of the extracted features and assigning weight values; The features inside the feature set FV are marked, expressed as FV = {f1, f2, ..., fi}; the fi-th feature of the feature set FV is weightedly calculated by the TF-IDF calculation method, and the calculation results are fitted to obtain the feature vector FVe; 5. The online intelligent archive management system based on big data analysis according to claim 2 is characterized by: The document classification and clustering module includes a feature vector clustering unit and a classification and label assignment unit; The feature vector clustering unit inputs the feature vector FVe into the K-means clustering algorithm and uses K-means clustering to classify the processed documents, obtains the clustering result Clusterj, and divides it into different categories, including financial documents, contracts and employee information; The processed documents are divided into k clusters by using the K-means algorithm. The center of each cluster is Cj. The distance D from the document to the cluster center is calculated. The feature vector FVe of the document will be assigned to the cluster center Cj with the closest distance. The distance D from the document to the cluster center is obtained by the following formula: Where D(FVe, Cj) represents the distance from the feature vector FVe to the cluster center Cj, Cj,i = {cj,1, cj,2, ..., cj,i} represents the dimension value of the cluster center j in the i-th feature, fi represents the i-th feature in the feature vector FVe, and n represents the total number of clusters; Assign documents to the nearest cluster center based on the minimum distance principle: FVe∈Clusterj if D(FVe,Cj)≤D(FVe,Ci)for i≠j; In the formula, FVe represents the feature vector of the document, Cj and Ci represent the centers of cluster j and cluster i respectively, and D(FVe, Ci) represents the distance from the feature vector FVe to the cluster center Ci; The classification and label assignment unit assigns a corresponding label L to each cluster j; The label L is obtained by the following formula: L(FVe)=Lj if FVe∈Clusterj; Where L(FVe) represents the assigned label of the feature vector FVe, and Lj represents the label of cluster j.

6. The online intelligent archive management system based on big data analysis according to claim 5 is characterized by: The automatic filing and storage module includes a storage path generation and rule application unit and a document filing and storage unit; The storage path generation and rule application unit analyzes the classified documents, extracts the keyword K, entity B and time information T features, generates classification rules, and generates a storage path Path according to the classification rules; When a new document is input, the features of the new document are extracted, including the keyword K, entity B and time information T features, and matched with the features of the classified document, and the classification rules are selected and stored according to the storage path Path; The storage path Path is obtained by the following formula: Path = h(K, B, T); Where h represents the function of generating the storage path.

7. The online intelligent archive management system based on big data analysis according to claim 6 is characterized by: The document archiving and storage unit is responsible for applying the generated storage path Path to the database, archiving and storing the new document according to the storage path Path, and interacting with the database to archive the content, classification information and path of the document; Specifically, the document is stored in a specified location according to the generated storage path Path, and the features of the new document are stored as document retrieval information, including the keyword K, the entity B, and the time information T; The storage formula of the feature is: FV 文档 ={L 文档 ,K 文档 ,B 文档 ,T 文档 ,Path 文档 }; Where FV 文档 represents the features of the feature set FV in the document, L 文档 represents the classification label of the document, K 文档 Represents the keyword set of the document, B 文档 Represents the entity information in the document, T 文档 Indicates the time information of the document, Path 文档 Indicates the document storage path.

8. The online intelligent archive management system based on big data analysis according to claim 1 is characterized by: The intelligent monitoring and early warning module includes a data acquisition and analysis unit and an anomaly detection and early warning generation unit; The data collection and analysis unit collects archive storage and access data and performs real-time analysis to identify potential abnormal behaviors by analyzing the access frequency Fw, visitor behavior, and storage status of the document; Based on the collected access data, the access frequency Fw of each document is calculated and compared with the access threshold TFw under the normal access mode to identify abnormal access behavior status; The access frequency Fw is obtained by the ratio of the number of accesses to the document to the time period to which the number of accesses belongs; The access threshold TFw is obtained by the following formula: TFw=[μFw-k*σFw,μFw+k*σFw]; In the formula, μFw represents the mean of the access frequency, σFw represents the standard deviation of the access frequency, and k represents the adjustment factor; The abnormal behavior status of the access is obtained by matching in the following ways: When the access frequency Fw is within the range of the access threshold TFw, it means that the access frequency is normal; When the access frequency Fw is not within the range of the access threshold TFw, it indicates that the access frequency is abnormal.

9. The online intelligent archive management system based on big data analysis according to claim 8 is characterized by: The anomaly detection and warning generation unit analyzes the storage status, access frequency Fw and abnormal behavior status of access through an anomaly detection algorithm, identifies potential problems, and generates an early warning to notify the administrator based on the abnormal behavior status; The abnormal behavior states include abnormal access frequency, abnormal storage capacity, and damaged document integrity; When the access frequency is abnormal, the access abnormality score SF is calculated; The access anomaly score SF is obtained by the ratio of the absolute value of the difference between the access frequency Fw and the mean μFw of the access frequency to the standard deviation σFw of the access frequency; When the storage capacity is abnormal, the storage occupancy ratio RS is calculated and compared with the preset capacity threshold TRS to determine the storage capacity status; The storage occupancy ratio RS is obtained by the following formula: In the formula, CCt represents the total storage capacity, and CCa represents the remaining storage capacity; When the storage occupancy ratio RS ≤ the capacity threshold TRS, no alarm is issued; When the storage occupancy ratio RS>the capacity threshold TRS, an alarm of insufficient storage capacity is issued.

10. The online intelligent archive management system based on big data analysis according to claim 9 is characterized by: The analysis report and feedback optimization module includes a report generation unit and a feedback optimization unit; The report generation unit calculates the index based on the document storage status and the access frequency Fw, and generates a visual analysis report in the form of charts, text and summaries, which is provided to the manager for review; Among them, the indicators include classification accuracy Ac, storage utilization Us, and access trend change rate Rtr; The storage utilization rate Us is obtained by the ratio of the difference between the total storage capacity CCt and the remaining storage capacity CCa to the total storage capacity CCt; The classification accuracy Ac is obtained by the ratio of the total number of classified documents Nto to the number of correctly classified documents Nco; The change rate Rtr of the access trend is obtained by the ratio of the difference between the access frequency Fw at time t and the access frequency Fw at time t+M to the time interval (t+M)-t; The feedback optimization unit retrains the K-means clustering algorithm according to the classification accuracy Ac, adjusts the storage path Path according to the storage utilization Us and the access trend change rate Rtr, calculates and obtains the new storage path NPath, and stores the document; The new storage path NPath is obtained by the following formula: NPath=h(Path,Ac,Us,Rtr); Where h represents the function of generating the storage path.