File full life cycle management system and method based on cloud computing
Through the full life cycle management system and methods of archives based on cloud computing, the problem of inefficient electronic archive management in the existing technology is solved, efficient, intelligent and standardized management of archive data is achieved, and the efficiency and security of archive management are improved throughout the process.
Patent Information
- Application Number
- CN202510518355.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing technology is difficult to achieve efficient, intelligent and standardized management of the entire life cycle of electronic files, especially in terms of centralized data management, secure access and remote sharing.
The full life cycle management system and method of archives based on cloud computing are adopted, and the archive data is collected and reconstructed into a unified standard format, and metadata extraction is performed by combining text vector representations fused by TF-IDF and BERT, dynamically divide the document life cycle state and bind access permissions and storage policies to realize dynamic encryption and intelligent storage scheduling to ensure data security and compliance.
It significantly improves the efficiency and security of the entire process of archive management, realizes the precise classification and standardized processing of multi-source heterogeneous archives, dynamic security and compliance guarantees, storage resource optimization and cost control, and ensures the confidentiality, integrity and tamper-proof of data.
Smart Images

Figure CN120045520A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file management, and particularly to a file full-life cycle management system and method based on cloud computing. Background Art
[0002] With the continuous deepening of digital office work and information management, file management work is gradually transforming from traditional paper files to electronic files. The types, quantities, and complexities of file data are increasing explosively. How to achieve efficient, intelligent, and standardized management of the full-life cycle process of electronic files from generation, classification, storage, retrieval, approval, backup to destruction has become an important issue faced by various enterprises, institutions, government agencies, and social organizations.
[0003] A method and system for constructing a smart file management throughout the life cycle are disclosed in the publication number CN117671714A; the method includes: collecting file materials and creating a primary file library; sorting files based on the primary file library, generating an ultimate file library and storing it; performing file retrieval and sharing by establishing retrieval rules; recording the operation history and access history of files and performing marking and traceability; establishing filing requirements and performing filing and backup; setting file destruction rules and regularly performing destruction and update; by collecting, identifying, sorting, storing, retrieving and counting, and destroying file information resources, establishing a full-cycle file management from file collection to file destruction, so as to realize the integrated utilization and management of file resources; thereby improving the efficiency and quality of file management and promoting information sharing and collaboration.
[0004] At present, cloud computing technology, with its elastic resource scheduling, distributed storage, and service capabilities, has been widely applied in the construction of various information systems, providing strong support for the centralized management, unified scheduling, secure access, and remote sharing of file data. At the same time, the integrated application of emerging technologies such as artificial intelligence and blockchain also provides new technical paths for intelligent identification, intelligent transfer, and compliance auditing in the file life cycle. Summary of the Invention
[0005] The object of the present invention is to propose a file full-life cycle management system and method based on cloud computing for the problems existing in the background art.
[0006] The technical solution of the present invention: A file full-life cycle management method based on cloud computing includes the following specific implementation steps: S1. Collect file data, evaluate the matching degree with standard files based on the format adaptability function, combine optical character recognition and semantic density analysis to determine the filing value of unstructured documents, define a policy function to decide the processing action according to the adaptability threshold and semantic density threshold, and map the data reconstruction into a unified standard format; S2. Adopt a text vector representation that combines TF-IDF and BERT, extract metadata, perform multi-label semantic classification after weighted fusion, and generate reconstructed labels by combining label hierarchy completion and cross-inference; S3. Dynamically divide the initial state into draft, review, archive, and deletion phases based on document metadata tags and identifiers, and automatically bind access permissions and storage policies according to the state; S4. Dynamically adjust the encryption method and storage policy according to the document lifecycle state. Use AES encoding in the draft state, RSA encoding in the review state, and an archive static encoding mechanism based on matrix decomposition in the archive state. Combine the access frequency to intelligently schedule resources and optimize the storage cost and secure access efficiency; S5. Dynamically allocate user permissions, set fine-grained access policies according to the document lifecycle state, record user behaviors, operation types, and results in real time to generate audit logs, and detect abnormal access to trigger an alarm mechanism; S6. Automatically migrate the archive state according to user behaviors, time rules, and the policy engine. Configure migration rules through a conditional trigger mechanism and verify the consistency of permissions and states. Combine automated migration and event-driven trigger policies to adjust and synchronously update the storage location and audit records; S7. Real-time evaluate the operation compliance through a periodic audit mechanism, generate a warning report by combining anomaly detection and risk assessment, and automatically lock abnormal archives; S8. Automatically trigger a compliance review at the end of the archive retention period, adaptively select a destruction policy according to the type and execute it, and generate a traceable audit log after verifying the destruction result.
[0007] Preferably, the implementation process of the policy function for decision-making on processing actions is defined according to the adaptation threshold and semantic density threshold as follows: S21. Based on a unified data access model, connect to several data sources, automatically identify the input data format, define a format adaptation function, and evaluate its adaptation to the standard archive format: ; where F adapt represents the overall adaptation of the archive data to the standard format; f i represents the i-th field format in the data source of the archive data; F std represents a predefined standard archive metadata template; sim(f i ,F std ) represents the field format similarity matching function; n represents the total number of input fields; α i represents the field importance weight, which is set according to business rules; S22. Perform OCR recognition and semantic density analysis on unstructured documents, define a semantic density function, and determine whether it has archiving value: ; wherein, D sem represents the semantic density score; w j represents the j-th semantic keyword; represents the semantic weight of the keyword in the document; β j represents the preset semantic keyword weight factor; L represents the number of words in the document; S23. Define the policy function , and determine the access action according to the format adaptability and semantic density, specifically: ; wherein, θ 1 represents the format adaptability threshold; θ 2 represents the semantic density threshold; sourceType represents the data source type; criticalList represents the non-discardable source list.
[0008] Preferably, the classification process of multi-label semantic classification is as follows: S31. Use the TF-IDF + BERT word embedding model for fusion representation: ; wherein, represents the statistical word frequency weight vector; represents the statistical word frequency weight vector; represents the context embedding vector; α represents the weighted fusion coefficient; S32. Adopt a three-channel fusion of the structural rule channel + semantic understanding channel + named entity recognition channel to extract metadata, and fuse the results of each channel; S34. Combine the extracted metadata M for multi-label semantic classification: ; wherein, W represents the classification weight matrix; b represents the bias term; σ represents the Sigmoid activation function; C represents the archived label set.
[0009] Preferably, the extraction process of extracting metadata by the three-channel fusion of the structural rule channel + semantic understanding channel + named entity recognition channel is as follows: Structural rule channel: Parse the data based on the existing template rules, extract metadata for documents with known formats, and identify the keyword fields in the document according to the predefined rules, that is, the structured metadata M struct ; Semantic extraction channel: Input the feature vector into the pre-trained model BERT for context semantic analysis, use the text classification method to identify the meaningful metadata items, and output the semantic metadata M sem ; Named entity recognition channel: Based on the pre-trained bidirectional long short-term memory network BiLSTM-Conditional Random Field CRF model, combined with the context information in the feature vector, perform named entity recognition and output named entity metadata M NER ; Fuse the results of each channel: ; Among them, M represents the fused metadata; λ 1 , λ 2 and λ 3 represent the weighting coefficients of each channel.
[0010] Preferably, the implementation process based on the matrix factorization archive static encoding mechanism is as follows: S51. Initialization settings: Select two positive integers n and m, satisfying m > n, and select a finite field F q ; q = 251; Define two special matrix sets: A matrix set S with elements being n×n invertible matrices and satisfying ; A matrix set M with elements being n×m matrices and satisfying rank(M)=n; Among them, I is the identity matrix, T is the matrix transpose operation; rank(M) represents the rank of matrix M; Archive static encoding and archive static parsing generation: Randomly select an orthogonal matrix A ∈ S and a full-rank matrix B ∈ M; and calculate the archive static parsing code Pd: ; S52. Archive document encoding operation: Extract the document to be archived, perform binary encoding doc on it, and perform primary encoding on it to obtain the encoded data data = H(doc); H is a predefined hash function; Calculate the encoding transition parameter Et: , where u is randomly selected vector, satisfying , solve the equation to get ; Among them, B + is the pseudo-inverse of matrix B; Among them, is the m-dimensional vector space over the finite field F q ; Output the archive static encoding u, and archive {u, data}; S53. Archive static parsing operation: Obtain {u, data}, extract the archive data data and the archive static encoding u, and calculate the parsing transition parameter ; If =data, the archived data data has been confirmed to be extracted, parsed, and the archival data is restored.
[0011] Preferably, the destruction strategy formulation process is as follows: S61. Obtain the creation time D created and the corresponding retention period P time , and calculate the retention period T of the file according to the current time T curr to determine whether it has ended: exp ; ; If the condition is met, it is determined to enter the destruction process; S62. Before destroying the file, perform a compliance check according to the type of the file and the corresponding compliance rule R rule , and generate a destruction authorization C for files that meet the destruction conditions compliance =1, otherwise mark it as "further review required"; S63. Automatically select an appropriate destruction strategy according to the type, sensitivity, and compliance requirements of the file: ; Among them, S strategy represents the destruction strategy; Type doc represents the file type; A audit represents the audit record requirement, whether a destruction audit report needs to be generated; S() represents the destruction strategy formulation mechanism.
[0012] Preferably, the destruction strategy formulation mechanism is as follows: Select the destruction method according to the file type: For ordinary files, select the logical destruction strategy; for sensitive files, select a more secure destruction strategy; Adjust the compliance requirements: If the compliance check shows that the file cannot be destroyed, the destruction strategy returns "delayed destruction" and re-evaluates the retention period.
[0013] Preferably, the evaluation process of anomaly detection and risk assessment is as follows: S81. Dynamic compliance detection: Automatically determine whether the current operation meets the compliance requirements by monitoring the real-time data stream of document operations: ; Among them, C com represents the legality evaluation result, with a value of 1 indicating that the operation meets the compliance requirements and a value of 0 indicating that it does not meet the compliance requirements; C rule represents the applied compliance rule; C() represents the core function of the compliance evaluation; S82. Build an intelligent compliance evaluation model to evaluate document operations in real time, learn and predict potential compliance risks by analyzing historical operation data, and provide an intelligent evaluation report: ; Among them, C context represents the context information of the current operation, including but not limited to the time, user, device, and geographical location of the operation; M() represents a risk assessment function used to evaluate the compliance risk of the current document based on the historical operations of the document, the current operation context, and the characteristics of the document itself.
[0014] Technical solution of the present invention: A cloud computing-based full life cycle management system for archives, which is used to execute the above-mentioned cloud computing-based full life cycle management method for archives, includes: An archive collection module, which is used to support the automatic collection of multi-source format files and provide optical character recognition OCR and content structured parsing; A cloud archiving and classification module, which is used to automatically classify and archive based on intelligent tags and knowledge graphs, support user-defined metadata models, and realize the standardization of archive structures; A life cycle management module, which introduces a four-stage management mechanism: creation, review, archiving, and destruction; access permissions, processing policies, and automatic audit rules are bound to each stage; A cloud permission scheduling and auditing module, which is based on dynamic permission access control and blockchain evidence storage, records operation logs in real time, and has the functions of backtracking and error correction; An intelligent scheduling and storage optimization module, which combines file popularity, permission levels, and life cycles to automatically optimize storage resources and perform dynamic resource allocation using the elastic capabilities of cloud computing; A destruction and archiving retention judgment module, which automatically executes an archiving retention destruction judgment program.
[0015] Compared with the prior art, the above technical solution of the present invention has the following beneficial technical effects: The present invention designs a cloud computing-based full life cycle management system and method for archives. Through modular integration and technological innovation, the overall process efficiency and security of archive management are significantly improved: (1) Improvement of overall process automation and classification accuracy: Based on an intelligent access mechanism combining a format adaptation function and semantic density analysis, and a metadata extraction method integrating multiple models such as TF-IDF and BERT, accurate classification and standardized processing of multi-source heterogeneous archives are realized, effectively solving the problems of low efficiency and classification errors caused by the dependence on manual intervention in traditional systems; (2) Dynamic security and compliance guarantee: Through the integration mechanism of blockchain evidence storage and RBAC / ABAC permission control, the immutability of operation logs and fine-grained permission management are ensured. Combined with an AI-driven compliance destruction judgment module (based on regulations and semantic analysis), the risk of illegal operations is significantly reduced, meeting the security and legal requirements of enterprise-level document management; (3) Storage resource optimization and cost control: The intelligent scheduling module dynamically allocates high-performance storage, encrypted storage, or low-cost cold storage resources according to the document lifecycle status (draft, review, archive) and access frequency, and realizes the efficient and secure storage of archived data through the matrix factorization and stationary encoding mechanism, taking into account both storage cost optimization and data recovery reliability; (4) Stationary encoding and parsing: Utilize the linear algebraic properties of the orthogonal matrix (A) and the full-rank matrix (B) to combine the binary encoded data (data) of the document to be archived with the random vector (u) to generate an irreversible archived stationary encoding, and ensure that the data can only be parsed by legitimate authorized parties through pseudo-inverse operation and the irreversibility of the orthogonal matrix; This mechanism ensures data confidentiality (preventing unauthorized recovery) through the dual processing of matrix factorization and hash function during data storage, and also ensures the integrity and anti-tampering of archived data by verifying the consistency between the transition parameters and the original data; At the same time, its mathematical architecture based on finite fields optimizes the storage efficiency. Brief Description of the Drawings
[0016] Figure 1 It is the system architecture diagram of an archive full-life cycle management system based on cloud computing proposed by the present invention; Figure 2 It is the method flow chart of an archive full-life cycle management method based on cloud computing proposed by the present invention. Detailed Embodiments
[0017] Embodiment 1, as Figure 1 shown, an archive full-life cycle management system based on cloud computing proposed by the present invention includes: an archive collection module, a cloud archiving and classification module, a life cycle management module, a cloud permission scheduling and auditing module, an intelligent scheduling and storage optimization module, and a destruction and archive retention judgment module.
[0018] The archive collection module supports the automatic collection of multi-source format files (including but not limited to PDF, images, audio and video), and provides OCR (Optical Character Recognition) recognition and content structured parsing; The cloud archiving and classification module automatically classifies and archives based on intelligent tags and knowledge graphs, supports user-defined metadata models, and realizes the standardization of archive structures; The life cycle management module introduces a five-stage management mechanism: creation, use, storage, maintenance, and destruction; Each stage is bound with access permissions, processing policies, and automatic auditing rules; The cloud permission scheduling and auditing module is based on role-based access control (RBAC) and blockchain evidence storage, records operation logs in real time, and has the functions of backtracking and error correction; Intelligent scheduling and storage optimization module, which combines file popularity (access frequency), permission level and life cycle to automatically optimize storage resources and perform dynamic resource allocation using the elastic capabilities of cloud computing; Destruction and archiving retention judgment module, which introduces an AI algorithm to judge whether the archives can be destroyed or should be retained for a long time (in combination with regulations / unit policies), and provides automatic backup and compliance reminders.
[0019] Embodiment 2, as Figure 2 shown, a cloud computing-based full life cycle management method for archives proposed by the present invention is applied to the cloud computing-based full life cycle management system proposed in Embodiment 1, and its specific implementation steps are as follows: S1. The archive collection module automatically identifies, standardizes and structurally accesses multi-source heterogeneous archive data, providing a unified entry for subsequent archive classification, management and analysis. The specific implementation process is as follows: S11. Based on a unified data access model, connect to several data sources {including but not limited to: local document directory (file system); third-party business systems (OA, ERP, CRM); mail servers (Exchange, SMTP); scanner or mobile terminal input (image / scanned copy)}, automatically identify the input data format, and evaluate its compatibility with the standard archive format. Specifically: Define the format compatibility function: ; Among them, F adapt represents the overall compatibility of the archive data with the standard format; f i represents the format of the i-th field in the data source of the archive data; F std represents a predefined standard archive metadata template; sim(f i ,F std ) represents the field format similarity matching function, which calculates the similarity of the type, length and format rules of the input field and the standard field; n represents the total number of input fields; α i represents the field importance weight, which is set according to business rules (for example: the weight of "archive number" is higher than the weight of "remark"); S12. Perform OCR recognition and semantic density analysis on unstructured documents (including but not limited to PDF, Word, scanned copies), and judge whether they have archiving value. Define the semantic density function: ; Among them, D sem represents the semantic density score, which reflects the effective information density per unit length of the document; w j represents the j-th semantic keyword (including but not limited to "contract", "project number", "approver"); Represents the semantic weight of keywords in the document, which is obtained based on the TF-IDF semantic vector in this embodiment; β j Represents a preset semantic keyword weight factor; L represents the number of words in the document; S13. Define the policy function , and judge the access action according to the format adaptability and semantic density. Specifically: ; Among them, θ 1 Represents the format adaptability threshold; θ 2 Represents the semantic density threshold; sourceType represents the data source type (including but not limited to: approval system, official document system); criticalList represents the non-discardable source list (including but not limited to: official documents, financial budget documents); S14. Map the received or reconstructed data into a unified standard file format.
[0020] S2. The cloud archiving and classification module performs semantic parsing on the accessed data, extracts key metadata, and automatically classifies it into a multi-dimensional file label system according to the content characteristics and context, providing a data basis for subsequent life cycle management and retrieval. The specific implementation process is as follows: S21. Use the TF-IDF + word embedding model (BERT) for fusion representation: ; Among them, Represents the statistical word frequency weight vector; Represents the statistical word frequency weight vector; Represents the context embedding vector (output by the pre-trained model BERT in this embodiment); α represents the weighted fusion coefficient; S22. Adopt a three-channel fusion of "structural rule channel + semantic understanding channel + named entity recognition channel" to extract metadata: Structural rule channel: Parse the data based on the existing template rules, extract metadata for documents with known formats, and identify the keyword fields in the document according to the predefined rules, that is, the structured metadata M struct : ; Among them, f struct () represents the structured data extraction function, which depends on the structure pattern of the document; Semantic extraction channel: Input the feature vector into the pre-trained model BERT for context semantic analysis, and use the text classification method to identify meaningful metadata items (including but not limited to dates, project numbers), and output the semantic metadata M sem : ; Among them, f NLP () represents a semantic extraction function based on the pre-trained BERT model; Named entity recognition channel: Based on the pre-trained BiLSTM (Bidirectional Long Short-Term Memory Network)-CRF (Conditional Random Field) model, combined with the context information in the feature vector, named entity recognition is performed, and the named entity metadata M is output NER : ; After the metadata processed independently by each channel (structured metadata M struct , semantic metadata M sem , named entity metadata M NER ), different weights are assigned based on the extraction importance of each channel, and the results of each channel are fused: ; Among them, M represents the fused metadata; λ 1 , λ 2 and λ 3 represent the weighting coefficients of each channel; S23. Combine the extracted metadata M to perform multi-label semantic classification: ; Among them, W represents the classification weight matrix; b represents the bias term; σ represents the Sigmoid activation function; C represents the archived label set; S24. Since the labels have a hierarchical and cross-dimensional context, set the label fusion strategy: ; Among them, F hier () represents the label tree automatic completion function, which is completed according to the upper concept; F cross () represents the label cross-inference function; represents the reconstructed label, that is, the completed label; S25. Output the reconstructed label .
[0021] S3. The lifecycle management module initializes the lifecycle status of the document and binds the corresponding lifecycle management policies to ensure the efficient utilization and protection of the document throughout its lifecycle. The specific implementation process is as follows: S31. Build the file status initialization division rules and dynamically allocate the lifecycle status: Creation stage: If the document is a newly created document (the label contains the "creation date" and has not been reviewed), the initial status is "Draft"; Review stage: If the document is marked as needing review (the label contains "review"), its lifecycle status is initialized to "In Review"; Archiving Phase: If the document has completed the approval process and meets the archiving conditions (such as the expiration of the archiving period), it is initialized as "Archived". Deletion Phase: If the document reaches the preset end date of its life cycle or is marked as useless, it is initialized as "Deleted". S32. Once the life cycle status of the document is determined, the corresponding life cycle management policies are automatically bound to the document, including but not limited to access rights and storage requirements, which are dynamically set according to the life cycle stage of the document. Specifically: Draft Status: The document is only visible to the creator and is stored in a temporary area without backup. Under Review Status: The document is visible to specific reviewers, stored in the review area, and a review period is set. Archived Status: The document is stored in a long-term archiving system, with data encryption and backup policies enabled, and can only be modified by specific personnel. Deleted Status: The document is marked for deletion and is automatically cleared or completely destroyed after a specified time. Accordingly: Through the metadata tag M and the unique identifier of the document, the initial life cycle status of the document is judged, and according to the initial life cycle status and the complemented tag M, the corresponding life cycle management policy P is automatically bound, and the life cycle status S of the document is output. init and the corresponding management policy P.
[0022] S4. The intelligent scheduling and storage optimization module stores the document in the cloud according to the life cycle status and management policy of the document, and ensures the security, privacy, and high availability of the document. At the same time, considering the different access requirements of the document and different stages of the life cycle, elastic scheduling is performed, so that the document can automatically adjust its storage and processing resources according to its access frequency, storage requirements, etc. The specific implementation process is as follows: S41. Dynamically adjust the data encoding method according to different life cycle stages of the document to improve the security of the document and reduce the complexity of encryption according to requirements: Documents in Draft Status: Use a symmetric encryption algorithm (including but not limited to AES) to encode the document, and the key is stored in a temporary storage area, and the encryption process is completed before the document is uploaded. Documents Under Review Status: Encapsulate and encode the data, and use the RSA public-private key encryption algorithm to ensure higher security during the transmission of the document. Documents in Archived Status: When archiving and extracting / viewing archived documents, encoding and parsing are both based on the matrix decomposition archiving static encoding mechanism. S42. Automatically select different cloud storage resources for storage according to the lifecycle status and management strategy of the document, and determine the storage location of the document by analyzing the access frequency, storage period, and encryption requirements of the document: (1) Draft status documents: Storage location: Stored in a high-performance, fast-response cloud storage area, including but not limited to SSD storage pools; Storage configuration: Enable multi-copy storage to ensure high availability and automatically expand capacity when access is frequent; Storage cost: Higher, but can ensure fast read and write; (2) Reviewed status documents: Storage location: Stored in a high-security, encrypted storage pool, using block storage or encrypted object storage services; Storage configuration: Enable data access control (DAC) and encryption policies; Storage cost: Higher, ensuring confidentiality during transmission and storage; (3) Archived status documents: Storage location: Stored in a cold storage area or archival storage; Storage configuration: Use long-term storage services, including but not limited to low-frequency access storage (LRS) or object storage; Storage cost: Low, suitable for documents with low access frequency but long-term preservation; S43. Automatically adjust the allocation of storage and computing resources according to the lifecycle stage of the document, optimizing storage costs and resource utilization: ; Among them, R sched represents the resource scheduling scheme; f() represents a predefined scheduling function; F access represents the access frequency of the document (including but not limited to read frequency, modification frequency); R compute represents the computing resource requirements; S init represents the lifecycle status of the document (draft, reviewed, archived, deleted); Specifically: Draft status documents: Due to high access frequency, faster response time and higher computing resources are required to handle modifications and reviews, and storage resources need to have a high response speed; Archived status documents: Low access frequency, reduce computing and storage resources, and use low-cost, long-term storage solutions; Reviewed status documents: Higher guarantees are required for storage and computing resources to support fast document review and modification; Accordingly: The intelligent scheduling and storage optimization module is based on the lifecycle status S of the document init, the management policy P and the document content D are used to perform corresponding encryption processing on the document according to the life cycle state and the management policy, generating the encrypted document D enc , according to the life cycle state and the policy, determine the storage location and storage type of the document, and intelligently schedule the storage and computing resources according to the access frequency and life cycle state of the document to ensure the efficient storage and secure access of the document, and then output the encrypted storage document and the corresponding resource scheduling scheme.
[0023] S5. When the user initiates operation requests such as file access, editing, downloading, forwarding, etc., the cloud access permission scheduling and auditing module is triggered in real time; combined with RBAC (role-based), ABAC (attribute-based) and behavior auditing mechanisms, it is judged whether the user has the corresponding operation permission; at the same time, all access operations are recorded in the blockchain audit chain to ensure that the operation behavior is traceable and tamper-proof. The specific implementation process is as follows: S51. Pre-define roles and perform permission allocation. The permission allocation table is shown in Table 1: Table 1 Permission Allocation Table ; In addition, the access permission is dynamically determined based on the attributes of the user, resource, and environment: ; Among them, A() represents the attribute-based access control function; U role represents the user role; P history represents the user behavior history; S status represents the life cycle state of the document; F doc represents the sensitivity of the document; E env represents the environmental factors, including but not limited to the access time (working hours / non-working hours), device (company intranet / extranet); Resource attributes: including but not limited to the sensitivity level of the document, document tags (including but not limited to financial documents, personal information); Environmental attributes: including but not limited to access time, access location, access device; User attributes: including but not limited to the user's access history, account status, affiliated department; S52. According to the life cycle state and management policy of the document, design fine-grained role and permission management. The requirements for user permissions for documents in different stages are different: ; Among them, C access represents the access control policy; Pe() represents the access control function, which generates the policy according to the life cycle state, user role and access request; S init represents the life cycle state of the document; U roleIndicates user roles (including but not limited to administrators, editors, viewers); R access () Indicates the type of access request (view, edit, delete); Specifically: (1) Draft status documents: Edit permission: Granted to the document owner and authorized editors, allowing modification of the document content; View permission: Allows some high-privilege users (administrators) to view the document, but does not allow external personnel to access; (2) Reviewed status documents: View permission: Only reviewers and administrators can view; Modification permission: Limited to administrators and some specific roles to prevent ordinary users from modifying the document; (3) Archived status documents: View permission: Archived documents can be opened to limited viewers, but editing is not allowed; Permission lock: Once archived, the modification permission will be frozen to prevent tampering; S53. During the document life cycle, build a monitoring mechanism to track the access situation of the document in real time and audit all operations of users, especially the monitoring of sensitive operations (including but not limited to editing, deleting), specifically: (1) Content of the audit record: User information: Record who performed the operation (including but not limited to user ID, role); Operation type: Record the operation type (including but not limited to view, edit, delete, modify status); Document information: ID, name, and type of the document being operated on; Operation time: Timestamp of when the operation was executed; Operation result: Whether the operation was successful, and if it failed, record the reason for failure; An example of the audit log is shown in Table 2: Table 2 Example table of audit log ; (2) When detecting abnormal behaviors (such as frequent failed operations or the behavior of a user accessing sensitive documents), the system should automatically trigger an alarm mechanism to notify the administrator. The audit alarm formula is: ; Among them, R alert Indicates the alarm trigger mechanism; R() indicates the audit alarm mechanism; L audit () Indicates the audit log; T threshold Indicates the set threshold (the number of consecutive failed operations exceeds 5 times); F anomalyIndicates abnormal behavior identification (including but not limited to high - profile access within a short period).
[0024] S6. The lifecycle management module automatically updates the file status according to user behaviors (including but not limited to editing, archiving, and signing), time rules (including but not limited to reaching the retention period), and the judgment results of the policy rule engine. The file status migrates according to the state transition function, and the state transition operation will trigger the adjustment of the encrypted storage location and the audit chain record again. The specific implementation process is as follows: S61. Build an automatic migration mechanism triggered by conditions, that is, when the document meets specific conditions, automatically promote the document to migrate from one state to another. For each state transition, the following mechanism must be used to control and change the lifecycle state: (1) State transition rule configuration: The change of the lifecycle state of each document depends on specific rules; for example: (a) From the draft state to the review state: The document completes preliminary editing and is submitted for review; (b) From the review state to the archive state: The document passes the review, is approved, and archived; (2) State transition condition definition: Each state transition is attached with specific conditions, for example: (a) Draft to review: All required fields of the document are filled in and submitted for review; (b) Review to archive: The document has passed the reviewer's review and has no modification opinions; S62. To ensure that the lifecycle migration of the document strictly complies with the regulations, perform verification during each state change to confirm whether the document meets the migration conditions. Specifically: (1) Permission verification: During each state transition, verify whether the current user has sufficient permissions to perform the operation; (2) State consistency verification: When performing a state transition, check whether the current state of the document allows the target state transition; For example, the document can only be migrated to the "approved" state after passing the review. If the states are inconsistent, automatically reject the migration request; S63. To improve the efficiency of document lifecycle management, introduce intelligent automated migration rules. Specifically: (1) Automated migration: Automatically migrate the document from one state to another according to the lifecycle state of the document and internal and external conditions; For example: After the document passes the reviewer's review, automatically change the document status from "review" to "archive"; (2) Event - triggered migration: Automatically execute state transitions by setting specific event - driven mechanisms (including but not limited to timed triggering, business process completion); For example, when a certain document exceeds the storage time limit, it is automatically migrated to the "archived" state.
[0025] S7. The lifecycle management module periodically triggers the audit mechanism to conduct a compliance assessment on all file operations and the process of status evolution: compare with the unit's policies and national regulations (including but not limited to the Data Protection Law and the Archives Law), generate a compliance report or a risk warning. If any abnormal behavior is found (including but not limited to files being retained beyond the due date and unauthorized access to permissions), the administrator will be automatically notified and the file will be locked. The specific implementation process is as follows: S71. Record each key operation in the document lifecycle in real time and conduct regular audits. These operations include document creation, editing, review, archiving, and deletion. By recording these operations in detail, it can ensure that any sensitive operation can be traced and audited, preventing any unauthorized behavior or incorrect operation. (1) Record each event in the document lifecycle in detail, including but not limited to the time when the event occurred, the user who performed the operation, the type of operation, and the object of the operation. That is, every change in the document status must be recorded as an audit log, and the following content is automatically recorded: ; Among them, T time represents the timestamp when the operation occurred; R role represents the user role who performed the operation; A action represents the operation performed (including but not limited to creation, editing, review); D doc represents the document involved in the operation; P para represents the parameters related to the operation (including but not limited to approval opinions and modified content); AutoLog() represents the audit log. (2) Use the hash algorithm to encrypt and store the audit log to ensure that the log content will not be modified during the storage process. S72. Combine the dynamic compliance detection and real-time compliance assessment mechanisms to conduct a compliance assessment on the document management information: (1) Dynamic compliance detection: By monitoring the real-time data flow of document operations, automatically determine whether the current operation meets the compliance requirements: ; Among them, C com represents the result of the legality assessment. A value of 1 indicates that the operation meets the compliance requirements, and a value of 0 indicates that it does not meet the compliance requirements; C rule represents the applied compliance rules, that is, the predefined regulations, policies, and compliance requirements; C() represents the core function of the compliance assessment, which is used to determine whether the current document operation meets the predefined compliance rules and standards. It should be noted that the core function C() of the compliance assessment calculates the compliance score based on the operation type, user role, document characteristics and applicable compliance rules, specifically: Operation type matching: For operation type A action , according to the predefined compliance rule set C rule Determine whether the operation is allowed; For example, if A action = "edit", and C rule The rule in the table states that editing operations are only allowed to senior users. If the user role R role If you are not an "Administrator", your compliance score is 0, which means you do not meet the compliance requirements. Document feature matching: Check the document type D doc Whether it complies with the operating rules; for example, if D doc If the file is confidential and the action is deletion, the compliance score will be adjusted based on whether the action approval process is met; Role check: determine the user role R role Whether the operation is allowed. If the role does not meet the requirements, the operation is not compliant. Final Compliance Score: If the action meets all rule conditions, then C com =1, otherwise 0; (2) Combining big data analysis and machine learning algorithms, we build an intelligent compliance assessment model to evaluate document operations in real time. By analyzing historical operation data, we learn and predict potential compliance risks and provide intelligent assessment reports: ; Among them, C context Represents the context information of the current operation, including but not limited to the time, user, device, and geographic location of the operation; M() represents a risk assessment function, which is used to assess the compliance risk of the current document based on the document's historical operations, the current operation context, and the characteristics of the document itself; S73. Introduce abnormal behavior detection and risk assessment mechanisms to monitor and analyze potential risks in document operations in real time, detect behaviors that do not meet compliance requirements in advance, and prevent violations. Specifically: (1) Abnormal behavior detection: Based on user behavior analysis and pattern recognition, it automatically identifies abnormal behaviors in document operations, including but not limited to sudden additions of deletion operations and unconventional file modifications: ; Among them, E detect () indicates the abnormal behavior detection result, which is 1 if an abnormality is detected, otherwise 0; oper Indicates the behavior data of the current operation (including but not limited to file modification and deletion); thr() represents the set abnormal threshold, which is used to identify whether the behavior is abnormal; represents the abnormal behavior detection function, which is used to identify and judge the abnormal behavior of the user during the document operation process. Based on the historical behavior pattern and the setting of the threshold, it detects whether there is a suspicious or non-compliant operation. By comparing the deviation between the current operation and the historical record, combined with the threshold to judge whether an abnormal behavior has occurred. If an abnormality is detected, the function output result is 1; otherwise, it is 0. (2) Based on big data analysis, the system evaluates the compliance risk of the document by calculating the risk score of all operations in the document life cycle: ; Among them, R com () represents the compliance risk score; T risk represents the set risk threshold; P policy () represents the compliance policy, which is defined according to the laws and regulations of document management; Risk() represents the final compliance risk assessment function, which is used to comprehensively evaluate according to the abnormal behavior detection result, operation threshold and compliance policy to obtain the final compliance risk score.
[0026] S8. At the end of the archive retention period, the destruction and archive retention judgment module automatically and intelligently decides whether to destroy the archive according to factors such as the type, importance, and legal requirements of the archive, and performs a safe destruction operation after confirmation. The specific implementation process is as follows: S81. Obtain the creation time D of the archive created and the corresponding retention period P time , and according to the current time T curr , calculate whether the retention period T of the archive exp has ended: ; If the condition is met, it is judged to enter the destruction process; S82. Before destroying the archive, conduct a review according to the compliance requirements to ensure that the destruction behavior complies with relevant regulations and enterprise policies. The compliance check includes, but is not limited to, confirming the sensitivity of the archive, legal or regulatory requirements (including but not limited to financial and tax documents), and whether there are other specific prohibited destruction clauses, that is: according to the type of the archive and the corresponding compliance rule R rule , perform a compliance check. If the archive belongs to a sensitive archive (including but not limited to financial and personal privacy archives), confirm again whether there are unexpired compliance requirements or whether the retention period needs to be extended. For archives that meet the destruction conditions, generate a destruction authorization C compliance =1, otherwise mark it as "needs further review"; S83. Automatically select an appropriate destruction strategy according to the type, sensitivity, and compliance requirements of the archives. The destruction strategies include, but are not limited to, physical destruction (including but not limited to hard disk destruction), and logical destruction (including but not limited to file erasure, encrypted destruction): ; Among them, S strategy represents the destruction strategy; Type doc represents the archive type; A audit represents the audit record requirement, whether an audit report on destruction needs to be generated; S() represents the mechanism for formulating the destruction strategy; It should be noted that the mechanism for formulating the destruction strategy is as follows: (1) Select the destruction method according to the archive type: For ordinary archives, select the logical destruction strategy (including but not limited to file erasure, encrypted destruction); For sensitive archives (including but not limited to confidential documents), select a more secure destruction strategy, including but not limited to physical destruction (hard disk destruction, printed document destruction) or encrypted destruction; (2) Adjustment of compliance requirements: If the compliance check shows that the archive cannot be destroyed (including but not limited to extending the retention period), the destruction strategy returns "delayed destruction" and re-evaluates the retention period; (3) Audit record requirement: If it is required to generate an audit record, add an audit function when selecting the destruction strategy to ensure that each destruction operation can be traced; S84. After the destruction strategy is determined, automatically execute the actual destruction operation O destruction : ; Among them, D doc represents the archive for the operation, that is, the archive to be destroyed; T destroy represents the timestamp of the destruction operation; O() represents the destruction mechanism; It should be noted that the destruction mechanism O() is specifically: According to the selected destruction strategy, execute the archive destruction operation (physical destruction: including but not limited to destroying hard disks, paper documents; logical destruction: execute file erasure, encrypted destruction to ensure that the files cannot be restored); Each destruction operation generates a timestamp T destroy to ensure that the operation is traceable; S85. After the destruction operation is completed, verify the destruction result and generate an audit record, that is: Verify the destruction result: Verify whether the destruction operation is complete through technical means (for physical destruction, verify whether the hard disk or storage medium is shredded or destroyed; for logical destruction, use data recovery tools to check whether the files can be restored); Audit generation: Record the destruction operation and generate an audit report, recording the destroyed archive, destruction time, and executor information; S86. Archive the logs and audit records generated during the destruction process, and generate a destruction report L according to requirements. log , the destruction log contains detailed information for each operation: L log ={O destruction ,D doc ,A audit}.
[0027] Embodiment 3. A method for managing the entire life cycle of archives based on cloud computing proposed by the present invention further includes a matrix decomposition-based archiving static encoding mechanism, and its specific implementation steps are as follows: Initialization settings: Select two positive integers n and m, where m > n, and select a finite field F q (q = 251); Define two special matrix sets: S (n×n invertible matrices that satisfy , I is the identity matrix, and T is the matrix transpose operation), M (n×m matrices that satisfy rank(M) = n, where rank(M) represents the rank of matrix M).
[0028] Archiving static encoding and archiving static parsing generation: Randomly select an orthogonal matrix A ∈ S and a full-rank matrix B ∈ M; and calculate the archiving static parsing code Pd: .
[0029] Archiving document encoding operation: Extract the document to be archived, perform binary encoding doc on it, and perform first-level encoding on it to obtain the encoded data data = H(doc); H is a predefined hash function; Calculate the encoding transition parameter Et: , where u is a randomly selected vector that satisfies , solve the equation to obtain ; where B + is the pseudoinverse of matrix B; Accordingly, obtain the archiving static encoding u, and archive {u, data}.
[0030] Archiving static parsing operation: Obtain {u, data}, extract the archived data data and the archiving static encoding u, and calculate the parsing transition parameter ; If = data, it is confirmed that the archived data data has been extracted, and it is parsed to restore the archive data.
[0031] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the relevant art.
Claims
1. A cloud computing-based archive life cycle management method, characterized in that: The specific implementation steps include the following: S1. Collect archive data, evaluate the compatibility with standard archives based on the format fitness function, determine the archiving value of unstructured documents by combining optical character recognition and semantic density analysis, define the policy function decision processing action based on the fitness threshold and semantic density threshold, and reconstruct and map the data into a unified standard format; S2, using the text vector representation fused with TF-IDF and BERT, extracting metadata, performing multi-label semantic classification after weighted fusion, and generating reconstructed labels by combining label level completion and cross-reasoning; S3, dynamically divide the initial status into draft, review, archive, and delete stages based on document metadata tags and identifiers, and automatically bind access rights and storage policies based on the status; S4. Dynamically adjust the encryption method and storage strategy according to the document life cycle status. The draft status uses AES encoding, the review status uses RSA encoding, and the archive status is based on the matrix decomposition archive static encoding mechanism. In addition, resources are intelligently scheduled based on the access frequency to optimize storage costs and security access efficiency. S5. Dynamically allocate user permissions, set fine-grained access policies based on the document lifecycle status, record user behavior, operation types and results in real time to generate audit logs, and detect abnormal access to trigger alarm mechanisms; S6. Automatically migrate archive status based on user behavior, time rules and policy engine, configure migration rules and verify permissions and status consistency through conditional trigger mechanism, combine automated migration with event-driven trigger policy adjustment, and synchronously update storage location and audit records; S7. Real-time evaluation of operational compliance through periodic audit mechanism, combined with anomaly detection and risk assessment to generate early warning reports and automatically lock abnormal files; S8. Automatically trigger compliance review at the end of the archive retention period, adaptively select and execute the destruction strategy based on the type, and generate a traceable audit log after verifying the destruction results.
2. According to the cloud computing-based archive life cycle management method of claim 1, it is characterized in that: The implementation process of defining the strategy function decision processing action based on the fitness threshold and the semantic density threshold is as follows: S21. Based on a unified data access model, connect to several data sources, automatically identify the input data format, define the format adaptability function, and evaluate its adaptability with the standard archive format: ; Among them, F adapt Indicates the overall adaptability of archival data to the standard format; f i Indicates the format of the i-th field in the data source of the archive data; F std Represents a predefined standard archive metadata template; sim(f i ,F std ) represents the field format similarity matching function; n represents the total number of input fields; α i Indicates the importance weight of the field, which is set according to business rules; S22. Perform OCR recognition and semantic density analysis on unstructured documents, define semantic density function, and determine whether they have archiving value: ; Among them, D sem represents the semantic density score; w j represents the jth semantic keyword; Indicates the semantic weight of the keyword in the document; β j represents the preset semantic keyword weight factor; L represents the number of words in the document; S23. Define the strategy function , the access action is determined according to the format adaptability and semantic density, specifically: ; Among them, θ1 represents the format adaptation threshold; θ2 represents the semantic density threshold; sourceType represents the data source type; criticalList represents the list of sources that cannot be discarded.
3. According to the cloud computing-based archive life cycle management method of claim 1, it is characterized in that: The classification process of multi-label semantic classification is as follows: S31. Use TF-IDF+ BERT word embedding model to combine representation: ; in, Represents the statistical word frequency weight vector; Represents the statistical word frequency weight vector; represents the context embedding vector; α represents the weighted fusion coefficient; S32, extract metadata by fusing three channels: structural rule channel, semantic understanding channel, and named entity recognition channel, and fuse the results of each channel; S34. Combine the extracted metadata M to perform multi-label semantic classification: ; Among them, W represents the classification weight matrix; b represents the bias term; σ represents the Sigmoid activation function; C represents the archive label set.
4. The method for managing archives throughout their life cycle based on cloud computing according to claim 3 is characterized in that: The extraction process of metadata by fusing three channels, namely structural rule channel + semantic understanding channel + named entity recognition channel, is as follows: Structural rule channel: Parse data based on existing template rules, extract metadata for documents of known formats, match according to predefined rules, and identify key fields in documents, namely structured metadata M struct ; Semantic extraction channel: feature vector Input into the pre-trained model BERT, perform contextual semantic analysis, use text classification methods to identify meaningful metadata items, and output semantic metadata M sem ; Named entity recognition channel: Based on the pre-trained bidirectional long short-term memory network BiLSTM-conditional random field CRF model, combined with the context information in the feature vector, named entity recognition is performed and named entity metadata M is output. NER ; Fusion of the results of each channel: ; Wherein, M represents the fused metadata; λ1, λ2 and λ3 represent the weighting coefficients of each channel.
5. According to the cloud computing-based archive life cycle management method of claim 1, it is characterized in that: The implementation process of the static coding mechanism based on matrix decomposition archiving is as follows: S51, initialization settings: select two positive integers n and m, satisfying m>n, and select a finite field F q ;q=251; Define two special matrix sets: one element is an n×n reversible matrix that satisfies The matrix set S; A matrix set rank M whose elements are n×m matrices and satisfy rank(M)=n; Where I is the identity matrix, T is the matrix transpose operation; rank(M) represents the rank of the matrix M; Archive static encoding and archive static analysis generation: randomly select an orthogonal matrix A∈S and a full-rank matrix B∈M; and calculate the archive static analysis code Pd: ; S52, archive document encoding operation: extract the document to be archived, perform binary encoding doc on it, and perform primary encoding on it to obtain encoded data data=H(doc); H is a predefined hash function; Calculate the encoding transition parameter Et: , where u is a randomly selected Vector, satisfying , solving the equation gives Among them, B + is the pseudo-inverse of matrix B; in, is a finite field F q m-dimensional vector space on ; The output is the archive static code u, and {u, data} is archived; S53, archive static analysis operation: obtain {u, data}, extract archive data data and archive static code u, calculate analysis transition parameters ; like =data, it is confirmed that the archive data data has been extracted, parsed, and restored.
6. The method for managing archives throughout their life cycle based on cloud computing according to claim 1, characterized in that: The destruction strategy formulation process is as follows: S61. Get the creation time of the archive D created and the corresponding retention period P time , and according to the current time T curr , calculate the archive retention period T exp Has it ended: ; If the conditions are met, the destruction process will be entered; S62. Before destroying the files, according to the type of files and the corresponding compliance rules, rule , perform compliance checks, and generate destruction authorization C for files that meet the destruction conditions compliance =1, otherwise marked as "needs further review"; S63. Automatically select the appropriate destruction strategy based on the file type, sensitivity and compliance requirements: ; Among them, S strategy Indicates the destruction strategy; Type doc Indicates the file type; A audit Indicates the audit record requirements and whether a destruction audit report needs to be generated; S() indicates the destruction policy formulation mechanism.
7. The method for managing archives throughout their life cycle based on cloud computing according to claim 6, characterized in that: The destruction strategy mechanism is: Select the destruction method based on the file type: for common files, select the logical destruction strategy; for sensitive files, select a safer destruction strategy; Compliance requirement adjustment: If the compliance check shows that the archive cannot be destroyed, the destruction policy returns to "delayed destruction" and the retention period is re-evaluated.
8. The method for managing archives throughout their life cycle based on cloud computing according to claim 1, characterized in that: The evaluation process of anomaly detection and risk assessment is as follows: S81. Dynamic compliance detection: By monitoring the real-time data flow of document operations, it automatically determines whether the current operation meets the compliance requirements: ; Among them, C com Indicates the result of the legality assessment. A value of 1 indicates that the operation complies with the compliance requirements, and a value of 0 indicates that the operation does not comply with the compliance requirements. rule Represents the compliance rules of the application; C() represents the core function of compliance assessment; S82. Build an intelligent compliance assessment model to evaluate document operations in real time, learn and predict potential compliance risks by analyzing historical operation data, and provide intelligent assessment reports: ; Among them, C context Represents the context information of the current operation, including but not limited to the time, user, device, and geographic location of the operation; M() represents a risk assessment function, which is used to assess the compliance risk of the current document based on the document's historical operations, the current operation context, and the characteristics of the document itself.
9. A cloud computing-based archive life cycle management system, which is used to execute a cloud computing-based archive life cycle management method according to any one of claims 1 to 8, characterized in that: include: Archive collection module, used to support automatic collection of files in multiple source formats, and provide optical character recognition (OCR) and content structured analysis; Cloud archiving and classification module, which is used for automatic classification and archiving based on intelligent tags and knowledge graphs, supports user-defined metadata models, and realizes standardization of archive structures; The lifecycle management module introduces a four-stage management mechanism: creation, review, archiving, and destruction; each stage is bound to access rights, processing strategies, and automatic audit rules; The cloud permission scheduling and auditing module records operation logs in real time based on dynamic permission access control and blockchain evidence storage, and has backtracking and error correction functions; Intelligent scheduling and storage optimization module automatically optimizes storage resources based on file popularity, permission level, and life cycle, and uses the elasticity of cloud computing to dynamically allocate resources; The destruction and archive retention judgment module automatically executes the archive retention and destruction judgment procedure.
Citation Information
Patent Citations
Intelligent archive management method and system for constructing full life cycle
CN117671714A
Personnel archive digital automatic classification method and system based on vocabulary statistics
CN117851869A
Intelligent archive management method and system based on artificial intelligence
CN118673528A
Remodification, identification and alarm system and method for sensitive archives
CN119046933A
Intelligent file cabinet management method based on RFID technology
CN119477206A
Cited By
Retrieval optimization method and device in retrieval system, equipment, medium and product
CN120407516A
Search optimization method, device, equipment, medium and product in search system
CN120407516B
Big data archive management full-life-cycle credible traceability data processing method
CN120408574A
Digital archive platform based on paper document digital conversion
CN120580703A
Cooperative office safety management method and system for preventing information leakage
CN120833122A