Intelligent document management method and system
By generating feature sets and indexing tag sets to classify files, identify sensitive files and assign permissions, and record the operation traceability chain, the problems of intelligence and efficiency of file management systems in existing technologies are solved, and the automation and security of intelligent document management are achieved.
Patent Information
- Application Number
- CN202510751730.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, problems such as file classification, permission allocation, and lifecycle management make it difficult for file management systems to achieve intelligence and efficiency, especially in processing massive unstructured files and adapting to flexible business processes.
By generating feature sets and indexing tag sets, files are classified, sensitive files are identified and permissions are assigned, the operation traceability chain is recorded, file lifecycle management is implemented, and disaster recovery backup is performed.
It realizes intelligent classification of files, protection of sensitive information and traceable management throughout the entire process, and improves the automation level and security of document management.
Smart Images

Figure CN120610936A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an intelligent document management method and system. Background Art
[0002] In the modern office environment, document storage, information collection, and storage management are core tasks for comprehensive administrative departments, and their importance is self-evident. With the surge in information volumes and the diversification of business needs, efficient document and information management has become crucial for improving work efficiency and ensuring information security. Whether accessing documents in daily office work or protecting sensitive information under compliance requirements, the quality of document management systems directly impacts the speed of decision-making and operational stability of administrative departments. However, existing technologies present significant implementation challenges in document classification, permission allocation, and lifecycle management.
[0003] First, the classification and cataloging of massive unstructured files is extremely difficult due to the lack of efficient content recognition and automatic metadata indexing capabilities, which results in time-consuming and low accuracy in file organization.
[0004] Secondly, the system is not sufficiently integrated with the organizational structure and authority system, making it difficult to flexibly adapt to changing business processes and access control requirements, resulting in fragmented management.
[0005] Finally, frequent file access and collaborative operations place extremely high demands on system performance, while existing solutions' shortcomings in concurrent processing capabilities further exacerbate efficiency bottlenecks. These technical challenges prevent file management from achieving both intelligent and efficient integration.
[0006] From the above problems, it can be seen that it is urgent to find an efficient and secure document management method among the increasingly numerous files. Summary of the Invention
[0007] To solve the above problems, the present invention proposes an intelligent document management method, which is characterized by comprising the following steps:
[0008] Step 1: Determine the document repository location: Obtain a collection of uploaded files, parse each file, generate a first feature set, and store it in a temporary database. Use the key features obtained from the first feature set to generate a first indexing tag set. Perform multi-dimensional classification on the files within the first indexing tag set, generate a first classification result set in the metadata table, and match it with the system's preset organizational structure mapping table to determine the archive repository location.
[0009] Step 2: Assigning document permissions: Extract sensitive file identifiers from the first classification result set and make a judgment based on the system's preset sensitivity threshold T. If the sensitivity value S is greater than T, the permission control module is triggered to assign read and write permissions based on the user role and generate a first permission profile, where T is the preset sensitivity threshold and S is the file sensitivity value. Obtain the allocation record from the first permission profile, record the operation process, and write it to the audit database through the system's log audit module to generate a first operation traceability chain.
[0010] Step 3, File Lifecycle Management: Based on the first operation traceability chain, the file is segmented and uploaded to cloud storage, a first storage index is generated and associated with the archive location, and the archive is regularly scanned and the file lifecycle status is detected using the first storage index. If the status is marked as expired, the archiving process is executed and updated to the second storage index.
[0011] Preferably, in step 1,
[0012] Data is obtained from the first feature set, and noise is removed through data preprocessing to obtain the first data set. For the first data set, an information extraction algorithm is used to extract key entities, timestamps, and subject fields to obtain the first entity set. Through the first entity set, the first index is generated according to the system's built-in metadata rules to obtain the first label set. If there are duplicate items in the first label set, a second label set is generated through deduplication processing. Otherwise, it is directly determined as the second label set. The second label set is grouped through a clustering algorithm to achieve classification preparation.
[0013] Preferably, in step 1, the second tag base is processed to obtain a classification result set, and the classification result set is matched with a system preset mapping table to determine the storage partition in the archive; if the partition matches the organizational structure, the final location is determined.
[0014] Preferably, in step 2, the file identifier is obtained from the first classification result set, the corresponding sensitivity value S is extracted through parsing, and a threshold judgment is performed on the sensitivity value S according to the preset sensitivity threshold T to determine whether it is greater than T. If the sensitivity value S is greater than T, the trigger module is activated, the user role information is obtained, the allocation rules are used to generate read and write permissions, the first permission configuration file is generated from the read and write permission data, the permission allocation result is stored, and the final permission configuration is obtained.
[0015] Preferably, in step 3,
[0016] The target file is segmented by obtaining the call execution instruction through the first operation tracing chain. If the segmentation is completed, the segmentation data is transmitted to the cloud through the cloud storage upload interface to obtain upload confirmation information. According to the upload confirmation information, a first storage index is generated to associate the first storage index with the corresponding field in the metadata table. The consistency of the file segmentation and the metadata table is confirmed through the first storage index association result to obtain the final storage status.
[0017] Preferably, in step 3,
[0018] Perform regular scans of the archive using the first storage index, use a pre-set operation and maintenance script to obtain the file lifecycle status and status mark data, extract expiration judgment conditions from the status mark data, process the file through the archiving and preservation process if the status mark indicates expiration, and determine the archiving completion status. For files in the archiving completion status, update the second storage index to obtain updated index data;
[0019] Based on the updated index data, the system operation and maintenance log is checked to determine whether the file detection is performed normally and obtain the detection results. The life cycle status change trend is analyzed through the detection results to predict the files that will expire in the next cycle and obtain the predicted data. The regular scanning frequency is adjusted according to the predicted data to obtain new status mark data.
[0020] Preferably, after step 3, step 4 is further included, including
[0021] Step 4: Disaster recovery backup of the file: Based on the second storage index, the disaster recovery backup algorithm is called to replicate the file in multiple regions and generate a first backup confirmation signal, which is then recorded in the operation and maintenance log.
[0022] Preferably, step 4 includes:
[0023] Obtain file data through the second storage index, determine the target region for replication, and generate an initial replication task. Obtain multi-regional distribution information from the initial replication task, allocate files to nodes in each region to obtain distribution results, perform file replication operations based on the distribution results, distribute files to multiple regions through the preset parallel processing tool, generate a replication completion mark, extract regional distribution and log record data, and store them in the second storage index through the data archiving tool to complete the disaster recovery backup process.
[0024] In another technical solution, the present invention proposes an intelligent document management system, which implements any of the intelligent document management methods described above.
[0025] The beneficial effects of using the present invention are:
[0026] The present invention discloses an intelligent document management system. The system processes uploaded files using multiple algorithms, generates and stores feature sets. It then extracts key information, generates indexing tag sets, and updates metadata tables. Files are classified based on the tag sets, and their respective archive locations are determined. The system can also identify sensitive files, assign permissions based on sensitivity, and record the operation traceability chain. The present invention also has a file lifecycle management function and can execute archiving and preservation processes. This invention implements intelligent file classification, sensitive information protection, and fully traceable management, improving the automation and security of document management and providing a comprehensive solution for document management. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flow chart of an intelligent document management method of the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of this technical solution more clear, the following technical solution is further described in detail in conjunction with specific implementation methods. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of this technical solution.
[0029] like Figure 1 As shown, the present invention proposes an intelligent document management method, comprising the following steps:
[0030] Step 1: Determine the document repository location: Obtain a collection of uploaded files, parse each file, generate a first feature set, and store it in a temporary database. Use the key features obtained from the first feature set to generate a first indexing tag set. Perform multi-dimensional classification on the files within the first indexing tag set, generate a first classification result set in the metadata table, and match it with the system's preset organizational structure mapping table to determine the archive repository location.
[0031] Step 2: Assigning document permissions: Extract sensitive file identifiers from the first classification result set and make a judgment based on the system's preset sensitivity threshold T. If the sensitivity value S is greater than T, the permission control module is triggered to assign read and write permissions based on the user role and generate a first permission profile, where T is the preset sensitivity threshold and S is the file sensitivity value. Obtain the allocation record from the first permission profile, record the operation process, and write it to the audit database through the system's log audit module to generate a first operation traceability chain.
[0032] Step 3, File Lifecycle Management: Based on the first operation traceability chain, the file is segmented and uploaded to cloud storage, a first storage index is generated and associated with the archive location, and the archive is regularly scanned and the file lifecycle status is detected using the first storage index. If the status is marked as expired, the archiving process is executed and updated to the second storage index.
[0033] Specifically, in the above steps, in step 1, a collection of uploaded files is obtained, and each file is processed by a text parsing algorithm, an image parsing algorithm, and a structured data parsing algorithm to generate a first feature set and store it in a temporary database.
[0034] Obtain a collection of uploaded files and use a classifier to determine the file type. If it is a text file, use a text parsing algorithm to generate the first text feature. If it is an image file, use an image parsing algorithm to generate the first image feature. If it is structured data, use a structured data parsing algorithm to generate the first data feature to obtain a first feature set. Use a preset threshold to determine the integrity of the first feature set, obtain a complete first feature set, and store it in a temporary database. Obtain the first feature set from the temporary database, use a clustering algorithm to group the first feature set, and obtain a grouped feature set. Based on the grouped feature set, use a statistical tool to calculate the distribution characteristics of each group to obtain a distribution feature set. Obtain the distribution feature set and use a mapping tool to convert it into a standard format to obtain a standard feature set. Use a verification tool to determine the consistency of the standard feature set, obtain a consistent standard feature set, and store it in a temporary database. Obtain a consistent standard feature set from the temporary database, use a merging tool to integrate all features, and obtain a final feature set.
[0035] Then, data is obtained from the first feature set, and noise is removed through data preprocessing to obtain the first data set. For the first data set, an information extraction algorithm is used to extract key entities, timestamps, and subject fields to obtain the first entity set. Through the first entity set, a first index is generated according to the metadata rules built into the system to obtain the first label set. If there are duplicates in the first label set, a second label set is generated through deduplication processing, otherwise it is directly determined as the second label set. According to the second label set, the metadata table is updated to obtain an updated metadata table. After obtaining the updated metadata table, the label set is grouped by a clustering algorithm to obtain a grouped label set. For the grouped label set, statistical analysis is used to determine the distribution characteristics of each group to obtain a distribution feature set.
[0036] After obtaining the distribution feature set, the indexing tags are obtained through the metadata table, and the K-means clustering algorithm is used to perform multi-dimensional classification on the files to obtain the first classification result set. The classification features are extracted from the first classification result set, and the system preset mapping table is used for matching to determine the preliminary attribution location. For the preliminary attribution location, the hierarchical information in the organizational structure is obtained, and the classification results are adjusted according to the hierarchical information to obtain the optimized classification set. According to the optimized classification set, the storage partition in the archive is determined. If the partition matches the organizational structure, the final attribution location is determined. Through the final attribution location, the index structure of the archive is obtained, and the hash algorithm is used to generate the file storage identifier. According to the file storage identifier, it is determined whether the storage partition exists in duplicate. If there is a duplicate, it is distinguished by the timestamp to obtain a unique storage path. Through the unique storage path, the file classification result is bound to the archive location to complete the multi-dimensional classification archiving.
[0037] When extracting the first set of indexing tags from the metadata table, the TF-IDF algorithm is used to calculate the weight of each tag. For example, for a document set containing tags such as "financial statement," "project plan," and "contract approval," a threshold of 75 is set to filter out low-frequency tags, retaining high-weight tags such as "financial statement (weight 82)" and "contract approval (weight 79)" as feature vectors. The documents are then classified using the K-means clustering algorithm, with an initial number of cluster centers set to 5 and a similarity metric measured using Euclidean distance. After 10 iterations, the silhouette coefficient reaches 68, resulting in a first classification result set containing documents in the three categories of "finance," "legal," and "administration." This result set is then matched against the pre-configured organizational structure mapping table using the cosine similarity algorithm to calculate the correlation between classification tags and departmental responsibilities. For example, a "finance" document has a match of 91 with the Finance Department, exceeding the preset threshold of 85. The system automatically assigns it to the "annual audit" subdirectory of the Finance archive. Based on the document's timestamp, 2023Q4, the storage path is set to " / finance / annual-audit / 2023 / 4th-quarter." During this process, secondary clustering is initiated for documents that have not been successfully matched, and the attribution relationship is recalculated after adjusting the feature dimensions to ensure that the classification accuracy reaches more than 98%.
[0038] In step 2, the file identifier is obtained from the first classification result set, and the corresponding sensitivity value S is extracted through parsing. According to the preset sensitivity threshold T, a threshold judgment is performed on the sensitivity value S to determine whether it is greater than T. If the sensitivity value S is greater than T, the trigger module is activated to obtain the user role information. The trigger module generates read and write permissions based on the user role using allocation rules. The first permission configuration file is generated from the read and write permission data, and the permission allocation result is stored. The integrity of the configuration file content is judged by the verification tool to obtain the final permission configuration. The final permission configuration is obtained and loaded into the system to perform permission control.
[0039] Obtain allocation records from the permission configuration file, separate the operation timestamp, user ID, and action type through record extraction, and obtain the original operation data set. For the original operation data set, use time sorting to arrange the operation timestamps in ascending order to generate a time series data set. By identifying and associating the user IDs in the time series data set, determine the action type distribution of each user and obtain a user action mapping table. If the action type in the user action mapping table exceeds the preset threshold, the log audit module writes the abnormal record into the audit database to obtain an abnormal audit log. Based on the abnormal audit log, use the traceability chain generation algorithm to construct the first operation traceability chain and output the traceability chain data set. Obtain the traceability chain data set, synchronize the traceability chain data set with the audit database through module writing, and obtain an updated audit database. For the updated audit database, use the hash algorithm to verify the traceability chain data set, determine the data integrity, and output the verification result.
[0040] For example, first, the system traverses the classification result set and extracts the sensitive identifiers of each file, such as the file's keywords, metadata information, etc. Assuming that the file currently being processed contains the keywords "financial report" and "customer information", the system calculates its sensitivity value S through the preset sensitive word library. The sensitivity calculation adopts a weighted algorithm, where the weight of "financial report" is 6, and the weight of "customer information" is 8, and the final S value is 6+8=4. The system's preset sensitivity threshold T is 0. Since S is greater than T, the system determines that the file is a sensitive file. Next, the permission control module assigns permissions based on user roles. For example, the administrator role has read and write permissions, and ordinary users only have read permissions. The system generates a first permission configuration file, records the file identifier and corresponding permission information, and ensures that sensitive files are only accessible to authorized users. The entire process is implemented through an automated algorithm without the need for human intervention, ensuring data security and operational efficiency.
[0041] In the first rights configuration file, specific information of the allocation record is extracted by parsing the configuration file in XML or JSON format.
[0042] For example, a configuration file might contain the following data: {"timestamp":"2023-10-01T12:34:56","user_id":"user123","action_type":"grant"}. After parsing, the system converts the timestamp "2023-10-01T12:34:56" to the Unix timestamp 1696162496 for subsequent processing. The log audit module then hashes the user ID "user123" and the action type "grant" using the SHA-256 algorithm to generate a unique action identifier, "e99a18c428cb38d5f260853678922e03." This module writes this information, along with a timestamp, into the audit database. Each record in the database is formatted as {"timestamp":1696162496, "user_id_hash":"e99a18c428cb38d5f260853678922e03", "action_type":"grant"}. To generate the first operation traceability chain, the system uses a Merkle tree structure, aggregating the hash values of all relevant operation records layer by layer to form a root hash of "b94d27b9934d3e08a52e52d7da7dabfac484efe37a5380ee9088f7ace2efcde9". This root hash, along with index information for the relevant operation records, is stored in the blockchain, ensuring data immutability and traceability.
[0043] In step 3, file lifecycle management: according to the first operation traceability chain, the file is segmented and uploaded to the cloud storage, a first storage index is generated and associated with the archive location, the archive is scanned regularly through the first storage index and the file lifecycle status is detected. If the status is marked as expired, the archiving process is executed and updated to the second storage index.
[0044] The call execution instruction is obtained through the operation traceability chain, triggering the storage management module to process the target file. For file sharding processing, a sharding algorithm is used to generate multiple sharding data to determine the sharding completion status. If the sharding generation is completed, the sharding data is transmitted to the cloud through the cloud storage upload interface to obtain the upload confirmation information. Based on the upload confirmation information, a first storage index is generated, and the index is associated with the corresponding field in the metadata table. Through the index association result, the archive location is obtained to determine whether the storage location is consistent with the preset location. If the storage location is consistent, the storage metadata is extracted from the archive location to determine the data integrity. Based on the integrity judgment result, a verification algorithm is used to confirm the consistency of the file sharding and the metadata table to obtain the final storage status.
[0045] Perform regular scans of the archive through the first storage index, use the preset operation and maintenance script to obtain the file life cycle status, and obtain status mark data. Extract the expiration judgment condition from the status mark data. If the status is marked as expired, process the file through the archiving and preservation process to determine the archiving completion status. For files in the archiving completion status, update to the second storage index to obtain the updated index data. Based on the updated index data, check the system operation and maintenance log to determine whether the file detection is performed normally and obtain the detection results. Analyze the trend of life cycle status changes based on the detection results, use the random forest algorithm to predict the files that will expire in the next cycle, and obtain the predicted data. Adjust the regular scanning frequency based on the predicted data, optimize the scanning task through the preset script, and determine the optimized task configuration. Use the optimized task configuration to perform the next scan and obtain new status mark data.
[0046] For example, the client request is parsed to extract the file to be processed. The SHA-256 algorithm is used to calculate the file's fingerprint: 3a7bd3e2360a3d29eea436fcfb7e44c735d117c7d1b6e5b. This triggers the sharding strategy module to split the 2GB video file into 300 shards according to the preset 4MB fixed shard size, generating 300 shard blocks and appending a CRC32 checksum (for example, the checksum for shard 0025 is 0xEDB88320). The storage management module then calls the AWS S3 SDK's PutObject interface to concurrently upload the shards. Each shard is assigned a unique object ID (for example, s3: / / bucket / prefix / 3a7bd3e2 / 002bin). A consistent hashing algorithm is then used to map the shards to the storage node cluster (for example, selecting node 171103 based on the partition of keyspace 0x58A3FC). The metadata service generates a MongoDB document with the primary key FILE_3a7bd3e2, records the shard topology [{"shard0025":{"size":4194304,"offset":104857600,"node":"171103"}}], and uses Elasticsearch to create an inverted index associated with the archive coordinates {"archive_id":"ARC-2023-12-001","storage_path":" / cold / zone4"}. Finally, in the MySQL transaction, the file status is updated to UPLOAD_COMPLETED and written to the audit log {"operation":"slice_upload","duration_ms":1248].
[0047] The archive is scanned regularly using the primary storage index. The system automatically initiates a scan task at 2:00 AM each day, using a pre-configured operation and maintenance script to check the file lifecycle status. The script calculates the lifecycle status of each file based on its creation time and the preset retention policy.
[0048] For example, for files that have been created for more than 365 days, the system will mark their status as "expired". During the detection process, the script will traverse all files in the archive and record each file's unique identifier, creation time, current status and other information. For files with a status marked as "expired", the system will trigger the archiving and preservation process to migrate these files from the first storage index to the second storage index. During the migration process, the system will use the SHA-256 algorithm to verify the files to ensure data integrity. After the migration is completed, the system will update the second storage index, record the file's archiving time, storage location and other information, and update the file status in the first storage index to "archived". The entire process is implemented through automated scripts without the need for manual intervention, ensuring efficient management of the archive and data security.
[0049] Step 4, disaster recovery backup of the file: According to the second storage index, the disaster recovery backup algorithm is called to replicate the file in multiple regions and generate a first backup confirmation signal, and the first backup confirmation signal is recorded in the operation and maintenance log.
[0050] Step 4 includes: obtaining file data through the second storage index and determining the target region for replication and generating an initial replication task; obtaining multi-region distribution information from the initial replication task; allocating files to nodes in each region to obtain distribution results; performing file replication operations based on the distribution results; distributing files to multiple regions through a preset parallel processing tool and generating a replication completion flag; extracting regional distribution and log record data; storing them in the second storage index through a data archiving tool to complete the disaster recovery backup process.
[0051] For example, based on the hash value 123456789 of the second storage index, the system first calls the multi-region replication module in the disaster recovery backup algorithm and uses the CRC32 checksum algorithm to divide the file into blocks, each block size is 4MB, to ensure the integrity of data transmission. By analyzing the block hash values of the file, the system automatically selects the three closest and lowest-loaded data centers for replication: Beijing, Shanghai, and Guangzhou, to ensure high data availability and low latency. During the replication process, the system monitors network bandwidth and latency in real time and dynamically adjusts the transmission rate to ensure that the replication time for each data center does not exceed 30 seconds. After the copy is complete, the system generates a primary backup confirmation signal. This signal contains information such as the file ID, backup timestamp, and storage paths in each data center. This signal is encrypted using the SHA256 algorithm and recorded in the operation and maintenance log. The log entry format is "File ID: 987654321, Backup Time: 2023-10-01 12:00:00, Beijing Path: / backup / beijing / 987654321, Shanghai Path: / backup / shanghai / 987654321, Guangzhou Path: / backup / guangzhou / 987654321." At the same time, the system automatically updates the backup status field of the secondary storage index, marking the backup status as "Completed," and triggers the scheduling analysis for the next backup task.
[0052] When extracting the first backup confirmation signal from the operation and maintenance log, a regular expression is first used to match the backup completion marker in the log, for example, matching the "backup_status:success" field. The timestamp and backup size (e.g., backup file size 3TB) are also recorded. The system collects real-time performance monitoring data, including storage node IOPS (e.g., current average IOPS is 4500), latency (e.g., average latency 12ms), and CPU utilization (e.g., CPU utilization of node A is 78%). Based on this data, a weighted round-robin load balancing algorithm is implemented, with the weight calculation formula set as (1 / latency) × (1 / CPU utilization) × IOPS. For example, the weight of node A is (1 / 12) × (1 / 78) × 4500 ≈ 487, and the weight of node B is (1 / 8) × (1 / 65) × 5200 ≈ 1000. By comparing the weights, storage access requests are dynamically allocated, for example, 60% of requests are allocated to node B and 40% to node A. The generated access policy includes concurrent connection limits (e.g., a maximum of 500 concurrent connections per node) and priority rules (e.g., prioritizing high IOPS tasks). The policy is then deployed to the storage management system through the API. After the policy takes effect, real-time monitoring shows that the node load is becoming more balanced (e.g., the CPU utilization of node A drops to 65%, and the latency of node B drops to 9ms).
[0053] The present invention discloses an intelligent document management system. The system processes uploaded files using multiple algorithms, generates and stores feature sets. It then extracts key information, generates indexing tag sets, and updates metadata tables. Files are classified based on the tag sets, and their respective archive locations are determined. The system can also identify sensitive files, assign permissions based on sensitivity, and record the operation traceability chain. The present invention also has a file lifecycle management function and can execute archiving and preservation processes. This invention implements intelligent file classification, sensitive information protection, and fully traceable management, improving the automation and security of document management and providing a comprehensive solution for document management.
[0054] The above content is only a preferred embodiment of the present invention. For ordinary technicians in this field, many changes can be made in the specific implementation methods and application scope based on the ideas of the present technical content. As long as these changes do not deviate from the concept of the present invention, they all fall within the scope of protection of this patent.
Claims
1. An intelligent document management method, characterized by: The following steps are included: Step 1: Determine the document repository location: Obtain a collection of uploaded files, parse each file, generate a first feature set, and store it in a temporary database. Use the key features obtained from the first feature set to generate a first indexing tag set. Perform multi-dimensional classification on the files within the first indexing tag set, generate a first classification result set in the metadata table, and match it with the system's preset organizational structure mapping table to determine the archive repository location. Step 2: Assigning document permissions: Extract sensitive file identifiers from the first classification result set and make a judgment based on the system's preset sensitivity threshold T. If the sensitivity value S is greater than T, the permission control module is triggered to assign read and write permissions based on the user role and generate a first permission profile, where T is the preset sensitivity threshold and S is the file sensitivity value. Obtain the allocation record from the first permission profile, record the operation process, and write it to the audit database through the system's log audit module to generate a first operation traceability chain. Step 3, File Lifecycle Management: Based on the first operation traceability chain, the file is segmented and uploaded to cloud storage, a first storage index is generated and associated with the archive location, and the archive is regularly scanned and the file lifecycle status is detected using the first storage index. If the status is marked as expired, the archiving process is executed and updated to the second storage index.
2. The intelligent document management method according to claim 1, wherein: In step 1, Data is obtained from the first feature set, and noise is removed through data preprocessing to obtain the first data set. For the first data set, an information extraction algorithm is used to extract key entities, timestamps, and subject fields to obtain the first entity set. Through the first entity set, the first index is generated according to the system's built-in metadata rules to obtain the first label set. If there are duplicate items in the first label set, a second label set is generated through deduplication processing. Otherwise, it is directly determined as the second label set. The second label set is grouped through a clustering algorithm to achieve classification preparation.
3. The intelligent document management method according to claim 2, wherein: In step 1, the second tag base is processed to obtain a classification result set, and the classification result set is matched with the system preset mapping table to determine the storage partition in the archive. If the partition matches the organizational structure, the final location is determined.
4. The intelligent document management method according to claim 1, wherein: In step 2, the file identifier is obtained from the first classification result set, and the corresponding sensitivity value S is extracted through parsing. According to the preset sensitivity threshold T, a threshold judgment is performed on the sensitivity value S to determine whether it is greater than T. If the sensitivity value S is greater than T, the trigger module is activated, the user role information is obtained, and the allocation rules are used to generate read and write permissions. The first permission configuration file is generated from the read and write permission data, the permission allocation result is stored, and the final permission configuration is obtained.
5. The intelligent document management method according to claim 1, wherein: In step 3, The target file is segmented by obtaining the call execution instruction through the first operation tracing chain. If the segmentation is completed, the segmentation data is transmitted to the cloud through the cloud storage upload interface to obtain upload confirmation information. According to the upload confirmation information, a first storage index is generated to associate the first storage index with the corresponding field in the metadata table. The consistency of the file segmentation and the metadata table is confirmed through the first storage index association result to obtain the final storage status.
6. The intelligent document management method according to claim 5, characterized in that: In step 3, Perform regular scans of the archive using the first storage index, use a pre-set operation and maintenance script to obtain the file lifecycle status and status mark data, extract expiration judgment conditions from the status mark data, process the file through the archiving and preservation process if the status mark indicates expiration, and determine the archiving completion status. For files in the archiving completion status, update the second storage index to obtain updated index data; Based on the updated index data, the system operation and maintenance log is checked to determine whether the file detection is performed normally and obtain the detection results. The life cycle status change trend is analyzed through the detection results to predict the files that will expire in the next cycle and obtain the predicted data. The regular scanning frequency is adjusted according to the predicted data to obtain new status mark data.
7. The intelligent document management method according to claim 1, wherein: After step 3, step 4 is also included, including Step 4, disaster recovery backup of the file: Based on the second storage index, the disaster recovery backup algorithm is called to perform multi-region replication of the file and generate a first backup confirmation signal, and the first backup confirmation signal is recorded in the operation and maintenance log.
8. The intelligent document management method according to claim 7, wherein: Step 4 includes: Obtain file data through the second storage index, determine the target region for replication, and generate an initial replication task. Obtain multi-regional distribution information from the initial replication task, allocate files to nodes in each region to obtain distribution results, perform file replication operations based on the distribution results, distribute files to multiple regions through the preset parallel processing tool, generate a replication completion mark, extract regional distribution and log record data, and store them in the second storage index through the data archiving tool to complete the disaster recovery backup process.
9. An intelligent document management system, characterized in that: The system implements the intelligent document management method according to any one of claims 1 to 8.
Citation Information
Cited By
Embedded device log management method and system
CN120950474A
File global sensing control intelligent protection method and system
CN121093387A
Steel rail flaw detection audio and video automatic synchronization merging and filing method, equipment and medium
CN121985165A