Data security encryption method and system for archive management
Through natural language processing technology, the file sensitivity is identified and matched with the encryption policy library, combined with data compression and distributed storage, the problem of time-consuming and labor-consuming development of manual classification and encryption policy is solved, and efficient and secure file management is achieved.
Patent Information
- Application Number
- CN202510196290.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The process of formulating manual classification and encryption strategies in the prior art is time-consuming and labor-intensive, and it is difficult to quickly and accurately evaluate the sensitivity level of massive archives, resulting in security risks and waste of storage resources.
Natural language processing technology to identify the sensitivity of archives and match them with the preset encryption policy library to achieve accurate encryption. When archives are stored, a data compression algorithm is used to remove redundant content, extract core information, and optimize storage space. According to the frequency and importance of archive access, a distributed storage system is used for load balancing to ensure data security and access efficiency.
It realizes accurate file encryption and optimized storage solutions, improves data security and storage efficiency, reduces the time and cost of manual operations, and reduces security risks and waste of storage resources.
Smart Images

Figure CN120124084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of file management and data security, and particularly to a data security encryption method and system for file management. Background Art
[0002] Data security and storage efficiency in file management are a pair of technical contradictions. To ensure the security of file data, it is usually necessary to encrypt and distribute the storage of file data, but this will increase the occupation of storage space and the complexity of accessing data. File administrators need to formulate differentiated encryption strategies and storage plans according to the sensitivity level and importance of files, which requires a scientific and perfect file classification system. However, in the scenario of a large amount of file data, it is a very difficult task to comprehensively evaluate the sensitivity level of each file and implement personalized encryption and storage strategies. How to automatically identify the sensitivity level of files at the stage of file storage and quickly match them with suitable security strategies is a technical problem to be solved urgently. In addition, on the premise of not destroying the original appearance of the files, how to accurately extract key information from redundant file content to compress the storage space is also a test for algorithms. The problems of data security and storage efficiency in file management need to balance security and efficiency on the basis of in-depth business scenarios and comprehensively apply technologies such as classification, encrypted storage, and data compression to achieve the unity of data security and storage optimization.
[0003] In an existing technology, the specific implementation method includes: adopting a method combining classification and encrypted storage to ensure data security and storage efficiency. In the implementation process, file administrators first classify the files preliminarily, formulate corresponding encryption strategies according to their sensitivity levels and importance, and then distribute the storage of file data to multiple nodes. Finally, the administrator will also apply data compression technology to process redundant information, so as to output the files.
[0004] However, the process of manual classification and formulation of encryption strategies in the existing technology is often time-consuming and laborious, and in the environment of a large amount of data, it is difficult to quickly and accurately evaluate the sensitivity level of each file, resulting in potential safety hazards and waste of storage resources. Summary of the Invention
[0005] The present invention provides a data security encryption method and system for file management to solve the problems that the process of manual classification and formulation of encryption strategies in the existing technology is often time-consuming and laborious, and in the environment of a large amount of data, it is difficult to quickly and accurately evaluate the sensitivity level of each file, resulting in potential safety hazards and waste of storage resources.
[0006] In a first aspect, to solve the above technical problems, the present invention provides a data security encryption method for file management, including:
[0007] Obtain the original file data;
[0008] Identify the file content according to the original file data to obtain the sensitivity level;
[0009] Perform algorithm matching on the sensitivity level based on a pre-stored encryption policy library to obtain an encryption policy;
[0010] Encrypt the file according to the encryption policy to obtain an encrypted file;
[0011] Extract core information from the original file data to obtain standardized data;
[0012] Classify and store the standardized data according to pre-stored storage rules to obtain storage nodes;
[0013] Allocate the storage nodes based on the access frequency of the file to obtain a storage plan;
[0014] Dynamically adjust the encryption policy and the storage plan based on the application scenario to obtain an optimized encryption policy and an optimized storage plan;
[0015] Decrypt the encrypted file using a distributed storage index mechanism to obtain a decryption plan;
[0016] Output according to the optimized encryption policy, the optimized storage plan and the decryption plan to obtain a file management plan.
[0017] In one implementable manner of the first aspect, the identifying the file content according to the original file data to obtain the sensitivity level includes:
[0018] Read the content according to the original file data to obtain text data;
[0019] Perform word segmentation on the text data based on natural language processing technology to obtain key information;
[0020] Judge according to the key information using a pre-stored sensitive word library. If the file content in the original file data contains sensitive words, mark the file corresponding to the original file data as a sensitive file;
[0021] Match sensitive words and sensitive sentence patterns according to the sensitive file to obtain sensitive information; wherein, the sensitive information includes the frequency of occurrence of sensitive words and the number of occurrences of sensitive sentence patterns;
[0022] Calculate a score using a pre-stored sensitivity level evaluation model according to the sensitive file to obtain a sensitivity level score;
[0023] Classify according to the pre - stored file classification system based on the sensitive level score to obtain the sensitive level.
[0024] In an implementable manner of the first aspect, the calculating the sensitive level score by using the pre - stored sensitive level evaluation model according to the sensitive file includes:
[0025] The sensitive level score is calculated by the following formula:
[0026]
[0027] where S represents the sensitive level score, n represents the number of different types of sensitive words in the pre - stored sensitive word library, w i is the weight coefficient of the i - th type of sensitive word, f i represents the occurrence frequency of the sensitive word of the i - th type, m represents the number of different types of sensitive sentence patterns in the pre - stored sensitive word library, v j is the weight coefficient of the j - th sensitive sentence pattern, g j represents the occurrence times of the j - th sensitive sentence pattern.
[0028] In an implementable manner of the first aspect, the obtaining the encryption policy by performing algorithm matching on the sensitive level based on the pre - stored encryption policy library includes:
[0029] Extract the encryption algorithm level from the pre - stored encryption policy library according to the sensitive level to obtain the encryption algorithm level;
[0030] Perform specific algorithm matching on the encryption algorithm level based on the pre - stored encryption policy library to obtain the encryption policy.
[0031] In an implementable manner of the first aspect, the obtaining the storage scheme by allocating the storage nodes based on the access frequency of the file includes:
[0032] Classify the files based on the pre - stored access frequency to obtain the file access level; wherein, the file access level includes high access heat, medium access heat and low access heat;
[0033] Match the file access level with the storage nodes to obtain the file storage location;
[0034] Integrate the file storage locations matched according to each level in the file access level to obtain the storage scheme.
[0035] In an implementable manner of the first aspect, the obtaining the optimized encryption policy and the optimized storage scheme by dynamically adjusting the encryption policy and the storage scheme based on the application scenario includes:
[0036] Read the file in real time based on the application scenario to obtain the real-time access frequency and real-time sensitivity level;
[0037] Adjust the encryption policy and the storage scheme in real time according to the real-time access frequency and the real-time sensitivity level to obtain an optimized encryption policy and an optimized storage scheme.
[0038] In one implementable manner of the first aspect, the decrypting the encrypted file using the distributed storage index mechanism to obtain a decryption scheme includes:
[0039] Read the location using the distributed storage index mechanism according to the encrypted file to obtain a physical address;
[0040] Extract the file according to the physical address to obtain a file identifier;
[0041] Query the file identifier based on the pre-stored encryption policy library to obtain key information;
[0042] Load according to the key information in the decryption algorithm library to obtain a decryption scheme.
[0043] In a second aspect, the present invention provides a data security encryption system for file management, including:
[0044] A data acquisition module for acquiring original file data;
[0045] A content recognition module for recognizing the file content according to the original file data to obtain a sensitivity level;
[0046] A policy matching module for performing algorithm matching on the sensitivity level based on a pre-stored encryption policy library to obtain an encryption policy;
[0047] A file encryption module for encrypting the file according to the encryption policy to obtain an encrypted file;
[0048] An information extraction module for extracting core information from the original file data to obtain standardized data;
[0049] A data storage module for classifying and storing according to the standardized data using a pre-stored storage rule to obtain a storage node;
[0050] A storage allocation module for allocating the storage nodes based on the access frequency of the file to obtain a storage scheme;
[0051] A scheme optimization module for dynamically adjusting the encryption policy and the storage scheme based on the application scenario to obtain an optimized encryption policy and an optimized storage scheme;
[0052] An archive decryption module, configured to decrypt according to the encrypted archive using a distributed storage index mechanism to obtain a decryption solution;
[0053] An output module, configured to output according to the optimized encryption policy, the optimized storage solution, and the decryption solution to obtain an archive management solution.
[0054] In an implementable manner of the second aspect, the identifying the sensitive level for the archive content according to the original archive data includes:
[0055] Reading the content according to the original archive data to obtain text data;
[0056] Performing word segmentation on the text data based on natural language processing technology to obtain key information;
[0057] Judging according to the key information using a pre-stored sensitive word library. If the archive content in the original archive data contains sensitive words, marking the corresponding archive of the original archive data as a sensitive archive;
[0058] Performing matching of sensitive words and sensitive sentence patterns according to the sensitive archive to obtain sensitive information; wherein, the sensitive information includes the occurrence frequency of sensitive words and the occurrence times of sensitive sentence patterns;
[0059] Calculating a score using a pre-stored sensitive level evaluation model according to the sensitive archive to obtain a sensitive level score;
[0060] Classifying according to the sensitive level score using a pre-stored archive classification and grading system to obtain the sensitive level.
[0061] In an implementable manner of the second aspect, the calculating a score using a pre-stored sensitive level evaluation model according to the sensitive archive to obtain a sensitive level score includes:
[0062] Calculating to obtain the sensitive level score through the following formula:
[0063]
[0064] wherein, S represents the sensitive level score, n represents the number of different types of sensitive words in the pre-stored sensitive word library, w i is the weight coefficient of the i-th type of sensitive word, f i represents the occurrence frequency of the i-th type of sensitive word, m represents the number of different types of sensitive sentence patterns in the pre-stored sensitive word library, v j is the weight coefficient of the j-th sensitive sentence pattern, g j represents the occurrence times of the j-th sensitive sentence pattern.
[0065] In an implementable manner of the second aspect, the algorithm matching of the sensitivity level based on the pre-stored encryption policy library to obtain an encryption policy includes:
[0066] Extract the algorithm from the pre-stored encryption policy library according to the sensitivity level to obtain the encryption algorithm level;
[0067] Perform specific algorithm matching on the encryption algorithm level based on the pre-stored encryption policy library to obtain an encryption policy.
[0068] In an implementable manner of the second aspect, the allocation of the storage nodes based on the access frequency of the archives to obtain a storage plan includes:
[0069] Classify the archives based on the pre-stored access frequency to obtain the archive access level; wherein, the archive access level includes high access heat, medium access heat, and low access heat;
[0070] Match the archive access level with the storage nodes to obtain the archive storage location;
[0071] Integrate the archive storage locations obtained by matching each level in the archive access level to obtain a storage plan.
[0072] In an implementable manner of the second aspect, the dynamic adjustment of the encryption policy and the storage plan based on the application scenario to obtain an optimized encryption policy and an optimized storage plan includes:
[0073] Read the archives in real time based on the application scenario to obtain the real-time access frequency and the real-time sensitivity level;
[0074] Perform real-time adjustment on the encryption policy and the storage plan according to the real-time access frequency and the real-time sensitivity level to obtain an optimized encryption policy and an optimized storage plan.
[0075] In an implementable manner of the second aspect, the decryption of the encrypted archive using the distributed storage index mechanism to obtain a decryption plan includes:
[0076] Read the location of the encrypted archive using the distributed storage index mechanism to obtain the physical address;
[0077] Extract the archive according to the physical address to obtain the archive identifier;
[0078] Query the archive identifier based on the pre-stored encryption policy library to obtain the key information;
[0079] Load according to the key information in the decryption algorithm library to obtain a decryption plan.
[0080] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the data security encryption method for file management described in any one of the above is implemented.
[0081] In a fourth aspect, the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, the device where the computer-readable storage medium is located is controlled to execute the data security encryption method for file management described in any one of the above.
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] The present invention discloses a data security encryption method for file management, including obtaining original file data;
[0084] Identifying the file content according to the original file data to obtain a sensitivity level; performing algorithm matching on the sensitivity level based on a pre-stored encryption policy library to obtain an encryption policy; encrypting the file according to the encryption policy to obtain an encrypted file; extracting core information according to the original file data to obtain standardized data; classifying and storing the standardized data according to a pre-stored storage rule to obtain a storage node; allocating the storage node based on the access frequency of the file to obtain a storage scheme; dynamically adjusting the encryption policy and the storage scheme based on the application scenario to obtain an optimized encryption policy and an optimized storage scheme; decrypting the encrypted file using a distributed storage index mechanism to obtain a decryption scheme; outputting according to the optimized encryption policy, the optimized storage scheme, and the decryption scheme to obtain a file management scheme. The method of the present invention identifies the sensitivity of the file through natural language processing technology and matches it with a preset encryption policy library to achieve precise encryption. When the file is stored in the library, a data compression algorithm is used to remove redundant content, extract core information, and optimize the storage space. According to the access frequency and importance of the file, a distributed storage system is used for load balancing to ensure data security and access efficiency. The present invention also uses a machine learning algorithm to continuously optimize the sensitivity recognition model to improve the accuracy rate. When accessing the file, the data is quickly located and decrypted through a distributed index. In addition, the present invention can dynamically adjust the encryption policy and the storage scheme according to business requirements to adapt to the actual application scenario. The present invention not only improves data security but also optimizes storage efficiency, providing a comprehensive solution for file management. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1It is a schematic flowchart of a data security encryption method for file management provided by the first embodiment of the present invention;
[0086] Figure 2 It is a schematic structural diagram of a data security encryption system for file management provided by the second embodiment of the present invention. Specific embodiments
[0087] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0088] Referring to Figure 1 , the first embodiment of the present invention provides a data security encryption method for file management, including the following steps:
[0089] S1. Obtain the original file data;
[0090] S2. Identify the file content according to the original file data to obtain the sensitivity level;
[0091] S3. Perform algorithm matching on the sensitivity level based on the pre-stored encryption policy library to obtain the encryption policy;
[0092] S4. Encrypt the file according to the encryption policy to obtain the encrypted file;
[0093] S5. Extract the core information according to the original file data to obtain the standardized data;
[0094] S6. Classify and store according to the standardized data using the pre-stored storage rules to obtain the storage node;
[0095] S7. Allocate the storage node based on the access frequency of the file to obtain the storage scheme;
[0096] S8. Dynamically adjust the encryption policy and the storage scheme based on the application scenario to obtain the optimized encryption policy and the optimized storage scheme;
[0097] S9. Decrypt the encrypted file using the distributed storage index mechanism to obtain the decryption scheme;
[0098] S10. Output according to the optimized encryption policy, the optimized storage scheme and the decryption scheme to obtain the file management scheme.
[0099] In step S1, the original file data is obtained.
[0100] In a specific embodiment, a preset file classification system is adopted to obtain the text data in the file content, and the text data is segmented through natural language processing technology, and keywords and key information are extracted as the original file data.
[0101] In step S2, the file content is identified according to the original file data to obtain the sensitivity level.
[0102] In the above step S2, the process of identifying the file content according to the original file data to obtain the sensitivity level specifically further includes the following steps:
[0103] S21, read the content according to the original file data to obtain the text data;
[0104] S22, segment the text data based on natural language processing technology to obtain the key information;
[0105] S23, judge using the pre-stored sensitive word library according to the key information. If the file content in the original file data contains sensitive words, mark the file corresponding to the original file data as a sensitive file;
[0106] S24, match the sensitive words and sensitive sentence patterns according to the sensitive file to obtain the sensitive information; wherein, the sensitive information includes the frequency of occurrence of sensitive words and the number of occurrences of sensitive sentence patterns;
[0107] S25, calculate the score using the pre-stored sensitivity level evaluation model according to the sensitive file to obtain the sensitivity level score;
[0108] S26, classify according to the sensitivity level score using the pre-stored file classification system to obtain the sensitivity level.
[0109] In this embodiment, the process of calculating the sensitivity level score by using the pre-stored sensitivity level evaluation model according to the sensitive file specifically obtains the sensitivity level score through the following formula:
[0110]
[0111] Wherein, S represents the sensitivity level score, n represents the number of different types of sensitive words in the pre-stored sensitive word library, w i is the weight coefficient of the i-th type of sensitive word, f i represents the frequency of occurrence of the i-th type of sensitive word, m represents the number of different types of sensitive sentence patterns in the pre-stored sensitive word library, v j is the weight coefficient of the j-th sensitive sentence pattern, g j represents the number of occurrences of the j-th sensitive sentence pattern.
[0112] It should be noted that in the above steps S21 to S26, the specific implementation process includes: according to the extracted key information, combined with a preset sensitive word library, determine whether the file content contains sensitive words. If it contains sensitive words, mark the file as a sensitive file. For the files marked as sensitive, use a preset sensitivity assessment model to calculate the sensitivity score of the file. The higher the score, the higher the sensitivity. According to the sensitivity score, combined with a preset file grading and classification system, determine the sensitivity level of the file. If the score exceeds the preset threshold, classify the file as a high-sensitivity level. For the files with a high-sensitivity level, use encryption storage technology to encrypt the file content to ensure the security of the file content. Through natural language processing technology, perform secondary analysis on the encrypted file content, extract the key content in the file, and generate a file summary. According to the file summary and the sensitivity level, store the file information in a preset file management system to complete the identification, grading, and classification processing of the file sensitivity.
[0113] Exemplarily, the file grading and classification system is a classification standard formulated based on the importance and sensitivity of file content. For example, files can be classified into three levels: ordinary, confidential, and top-secret. In practical applications, a certain government department classifies documents related to national security as top-secret, financial statements as confidential, and daily work reports as ordinary. Natural language processing technology enables computers to process language information like humans by analyzing the structure, semantics, and context of language, and can be used for word segmentation and key information extraction of text data. Taking a report on a new energy project as an example, keywords such as "solar energy", "wind energy", and "investment amount" can be identified through word segmentation, and then key information such as project type and investment scale can be extracted. The establishment of a sensitive word library needs to consider multiple dimensions, such as politics, economy, military, etc. For example, words such as "nuclear weapons", "state secrets", and "internal price" are listed as sensitive words. In practical applications, if a phrase like "nuclear weapon R & D progress" appears in a file, the system will immediately mark it as a sensitive file. The sensitivity assessment model can be scored based on multiple factors, such as the frequency, location, and context of sensitive words. Suppose a model has a full score of 100 points, with 10 points for each occurrence of a sensitive word, and an additional 5 points if it appears in the title. If a file mentions "state secrets" multiple times and appears in the title, its score will exceed 50 points and be determined by the system as a high-sensitivity level. Encryption storage technology is crucial for protecting files at a high-sensitivity level. Common encryption methods include symmetric encryption and asymmetric encryption. For example, the AES-256 algorithm can be used to encrypt the file content to ensure that even if the data is stolen, the content cannot be read without the key. File abstract generation is to extract the core content of the document to form a concise and clear overview. For example, for a 50-page research report, the abstract will include key points such as research purpose, methods, main findings, and conclusions, usually controlled within about 300 words. Such an abstract can help managers quickly understand the file content without having to refer to the full text. The design of the file management system needs to consider security, retrievability, and scalability. For example, the system can adopt a multi-layer authentication mechanism, and only those with corresponding permissions can access files at a specific level. At the same time, the system should support multi-dimensional retrieval, such as querying by keywords, time, sensitivity level, etc. Such a design not only ensures information security but also improves work efficiency. Through this series of steps, the identification and grading and classification processing of file sensitivity can not only effectively protect sensitive information but also improve the efficiency and accuracy of file management. This method is particularly suitable for processing a large number of complex files.
[0114] In step S3, an algorithm matching is performed on the sensitivity level based on the pre-stored encryption policy library to obtain an encryption policy.
[0115] In step S3 above, the algorithm matching for the sensitivity level based on the pre-stored encryption policy library to obtain an encryption policy specifically further includes the following steps:
[0116] S31. Extract an encryption algorithm level from the pre-stored encryption policy library according to the sensitivity level.
[0117] In a specific embodiment, the encryption policy library is a pre-established database containing various levels of encryption algorithms. Its form usually includes the following parts: Encryption algorithm classification, that is, according to different sensitivities, the encryption algorithms are divided into multiple levels, such as low, medium, and high levels, and each level corresponds to different encryption strengths; Specific encryption algorithms: Each level contains specific encryption algorithms. For example, the low level may use AES-128, the medium level uses AES-192, and the high level uses AES-256 or elliptic curve encryption (ECC), etc.; Key management information: The rules for key generation, storage, and management corresponding to each encryption algorithm to ensure the security and accessibility of the keys; Encryption parameter configuration: Includes parameter configurations such as encryption mode, padding method, and initialization vector (IV) to ensure the security and consistency of the encryption process; Dynamic adjustment mechanism: A mechanism for dynamically adjusting the encryption policy according to actual requirements and security assessment results to ensure that the encryption policy can adapt to different security requirements; Through this structured encryption policy library, the system can quickly match and apply appropriate encryption algorithms according to the sensitivity of the file to ensure the security and access efficiency of the data.
[0118] S32. Perform specific algorithm matching for the encryption algorithm level based on the pre-stored encryption policy library to obtain an encryption policy.
[0119] In a specific embodiment, the processes of steps S31 to S32 above specifically include: According to the sensitivity level, extract the corresponding encryption algorithm level division from the pre-established encryption policy library. For the extracted encryption algorithm level division, match the specific encryption algorithms in the encryption policy library. If the match is successful, use the corresponding encryption algorithm to encrypt the data. If the match fails, determine whether dynamic adjustment is required according to a preset threshold. Use the dynamic adjustment mechanism to obtain a new encryption algorithm level division and re-match the encryption policy library. According to the result of the re-match, determine the final decryption policy.
[0120] Exemplarily, sensitivity level identification is a key step in data classification and protection. By analyzing the data content, its sensitivity level can be determined. For example, a document containing a company's financial statements will be identified as highly sensitive, while a publicly available press release will be determined to have a low sensitivity level. Based on the identified sensitivity level, the system will extract the corresponding encryption algorithm level from a pre-established encryption policy library. This process is similar to selecting the appropriate level of a multi-layer safe. For highly sensitive data, a more complex and secure encryption algorithm will be selected, such as asymmetric encryption with a long key; while for low-sensitivity data, a relatively simple but faster symmetric encryption algorithm will be chosen. When matching a specific encryption algorithm, the system will look for the corresponding algorithm in the encryption policy library according to the previously determined encryption algorithm level. For example, if it is determined to use high-level encryption, the system will select a high-strength algorithm such as Elliptic Curve Cryptography (ECC). If the match is successful, the system will use the selected algorithm to encrypt the data to ensure its security. However, in some cases, a match failure may occur. At this time, the system will judge whether dynamic adjustment is needed according to a preset threshold. This threshold can be understood as a tolerance level. If the match result is not very different from the expected value, the system will accept the sub-optimal choice; but if the gap is too large, the dynamic adjustment mechanism needs to be activated. The role of the dynamic adjustment mechanism is to find alternative solutions when the original encryption policy cannot meet the requirements. This may involve reducing or increasing the encryption strength, or trying different types of encryption algorithms. For example, if asymmetric encryption was originally planned but cannot be implemented due to certain reasons, the system will switch to a symmetric encryption algorithm with a comparable strength. Through this dynamic adjustment, the system can maintain a certain degree of flexibility while ensuring security. Finally, the system will determine the most suitable encryption algorithm based on the adjusted result and complete the data encryption process.
[0121] In step S4, the file is encrypted according to the encryption policy to obtain an encrypted file.
[0122] Exemplarily, a corresponding key or password information is set for the file according to the encryption policy, so that the file can be safely stored.
[0123] In step S5, core information is extracted from the original file data to obtain standardized data.
[0124] In a specific embodiment, the original archive data is processed using a preset compression algorithm to remove redundant content, obtaining preliminary compressed data. Through key information extraction technology, the core content is identified and extracted from the preliminary compressed data to generate a key information data set. The key information data set is combined with the compression algorithm for secondary compression processing to generate the final compressed archive data. According to the requirements for archive storage, format standardization processing is performed on the final compressed archive data to obtain standardized compressed data.
[0125] Exemplarily, archive data compression and storage are crucial links in an information management system. First, the original archive data is processed using a preset compression algorithm. For example, using common ZIP or 7z algorithms can effectively reduce the file size. For instance, a text file containing a large amount of repetitive content will reduce the storage space by more than 50% after compression. Key information extraction technology is the core of archive management. Through natural language processing and machine learning methods, important content can be identified from the preliminary compressed data. For example, for an enterprise annual report, the system will extract key information such as financial data, major events, and future plans to form a streamlined data set. This process can not only further reduce the data volume but also improve the efficiency of subsequent retrieval. Secondary compression processing combines the key information data set and the compression algorithm to further optimize the storage efficiency. For example, if the key information data set contains a large amount of numerical data, a compression algorithm specifically for numerical values, such as differential encoding, can be used to store the differences between adjacent numerical values instead of storing the original numerical values, thus significantly reducing the storage space. Format standardization processing ensures the consistency and interoperability of archive data.
[0126] In step S6, according to the standardized data, classification storage is performed using pre-stored storage rules to obtain storage nodes.
[0127] In a specific embodiment, using preset storage rules, the standardized compressed data is classified and stored in the archive to complete the archive storage operation. If the amount of archive data exceeds a preset threshold, a distributed storage mechanism is activated to shard and store the data in multiple storage nodes.
[0128] Exemplarily, storage rules can store data on different storage media according to factors such as the type, importance level, and usage frequency of the files. For example, important and frequently used files will be stored on high-speed solid-state drives, while infrequently used historical files will be stored in a large-capacity but relatively slow tape library. This hierarchical storage strategy not only ensures fast access to important data but also reduces the overall storage cost. When the file data volume exceeds a preset threshold, starting a distributed storage mechanism can effectively solve the capacity and performance bottlenecks of a single storage node. For example, the annual financial statements of a large enterprise can reach a scale of hundreds of terabytes. At this time, the data can be sliced by year or department and stored on multiple physical servers. This not only increases the storage capacity but also improves the data processing speed through parallel reading and writing. At the same time, distributed storage can also improve the reliability of the data because even if one node fails, the data on other nodes is still available. The design of the entire file compression and storage process aims to balance storage efficiency, data integrity, and access performance. Through multi-level compression and intelligent storage strategies, it not only saves storage space but also ensures the fast extraction and long-term preservation of important information. This method is not only applicable to traditional document management but also can be applied to the processing of massive information in the big data era, providing an efficient and reliable data management solution for organizations.
[0129] In step S7, the storage nodes are allocated based on the access frequency of the files to obtain a storage plan.
[0130] In the above step S7, the storage nodes are allocated based on the access frequency of the files to obtain a storage plan, which specifically further includes the following steps:
[0131] S71, classify the files based on the pre-stored access frequency to obtain file access levels; wherein, the file access levels include high access popularity, medium access popularity, and low access popularity;
[0132] Exemplarily, the system obtains the access frequency data of each file. For example, a certain medical file has been retrieved 500 times in the past month. Combining with a preset importance scoring model, which considers factors such as the content type and creation time of the file, a comprehensive weight value is calculated for each file. For example, the above-mentioned medical file obtains a high weight value of 0.9 due to its importance and high access frequency. According to the calculated comprehensive weight value, the system classifies the files into three access popularity levels: high, medium, and low. For example, a weight value of 0.7 - 1.0 is high access popularity, 0.4 - 0.7 is medium access popularity, and 0 - 0.4 is low access popularity. At the same time, the classification of the access frequency levels is also pre-stored and will not be described in the embodiments of the present invention.
[0133] S72, match the file access levels with the storage nodes to obtain the file storage locations;
[0134] S73. Integrate the file storage locations obtained by matching each level in the file access level to obtain a storage plan.
[0135] It should be noted that in the above steps S71 to S73, the specific implementation process includes: obtaining the access frequency data of the files, and calculating the comprehensive weight value of each file in combination with a preset importance scoring model. According to the comprehensive weight value, the files are divided into three levels: high access heat, medium access heat, and low access heat, and a file access heat classification result is generated. For the storage nodes in the distributed storage system, obtain the storage capacity, load status, and network bandwidth information of the nodes, and generate a storage node resource distribution map. If the file access heat classification result is high access heat, it is preferentially allocated to the storage nodes with lower load and higher network bandwidth to generate a storage location allocation plan for high access heat files. If the file access heat classification result is medium access heat, a storage location allocation plan for medium access heat files is generated according to the remaining capacity of the storage nodes and the load balancing strategy. If the file access heat classification result is low access heat, the compression storage technology is used to centrally store the low access heat files to specific nodes to generate a storage location allocation plan for low access heat files.
[0136] Exemplarily, the classification of file access heat is an important basis for optimizing the allocation of storage resources. This classification method can effectively identify important files that need to be accessed quickly, providing a basis for formulating subsequent storage strategies. In a distributed storage system, it is crucial to understand the resource status of each storage node. The system collects information on the storage capacity, current load status, and network bandwidth of each node to generate a detailed distribution map of storage node resources. For example, Node A has 2TB of storage space, a current load of 60%, and a network bandwidth of 10Gbps; while Node B has 5TB of space, a load of only 30%, and a bandwidth of 20Gbps. For files with high access heat, the system will preferentially allocate them to nodes with lower load and higher network bandwidth. For instance, the aforementioned medical files will be allocated to Node B because it has more available resources and higher bandwidth, ensuring fast access. This strategy ensures that important and frequently used files can obtain the best storage and access performance. For files with medium access heat, the system will consider the remaining capacity of the storage node and the load balancing strategy. For example, a financial report of medium importance will be allocated to a node with a moderate load, which will neither occupy high-performance resources nor ensure a reasonable access speed. This balancing strategy helps to effectively utilize the resources of the entire system. Files with low access heat adopt compression storage technology and are centrally stored on specific nodes. For example, old files from many years ago are compressed and stored on a high-capacity, low-performance node dedicated to long-term archiving. This method not only saves storage space but also avoids low-usage files from occupying high-performance resources. Through this intelligent storage strategy based on access heat, the system can achieve the optimal allocation of storage resources, improve the overall storage efficiency and access performance. At the same time, this dynamic adjustment mechanism also reserves flexibility for future changes in file usage patterns, ensuring that the storage system can continue to operate efficiently.
[0137] In step S8, based on the application scenario, dynamically adjust the encryption policy and the storage scheme to obtain an optimized encryption policy and an optimized storage scheme.
[0138] In the above step S8, the dynamically adjusting the encryption policy and the storage scheme based on the application scenario to obtain an optimized encryption policy and an optimized storage scheme specifically further includes the following steps:
[0139] S81, read the files in real time based on the application scenario to obtain the real-time access frequency and the real-time sensitivity level;
[0140] S82, perform real-time adjustment on the encryption policy and the storage scheme according to the real-time access frequency and the real-time sensitivity level to obtain an optimized encryption policy and an optimized storage scheme.
[0141] In a specific embodiment, the file type and data classification are determined by combining the analysis results of the business scenario. According to the file type and data classification, a mapping relationship between the encryption strength and the storage location is established in advance to generate an initial encryption policy and a storage plan. If the sensitivity level of the file is higher than the preset threshold, a high-encryption-strength algorithm is used to encrypt the file, and a high-security-level storage location is selected. If the storage requirement of the file includes the characteristic of high-frequency access, the storage plan is dynamically adjusted, and the file is stored in a low-latency storage location while keeping the encryption strength unchanged. According to the changes in the business scenario, the change data of the access frequency and sensitivity level of the file are obtained in real time to determine whether the encryption policy or the storage plan needs to be updated. Exemplarily, if the sensitivity level of the file decreases or the storage requirement changes, the encryption policy and the storage plan are dynamically adjusted to reduce the encryption strength or migrate the storage location.
[0142] In step S9, the encrypted file is decrypted using the distributed storage index mechanism to obtain a decryption scheme.
[0143] In the above step S9, the step of decrypting the encrypted file using the distributed storage index mechanism to obtain a decryption scheme specifically further includes the following steps:
[0144] S91, Use the distributed storage index mechanism to read the location according to the encrypted file to obtain a physical address;
[0145] S92, Extract the file according to the physical address to obtain a file identifier;
[0146] S93, Query the file identifier based on the pre-stored encryption policy library to obtain key information;
[0147] S94, Load according to the key information through the decryption algorithm library to obtain a decryption scheme.
[0148] It should be noted that in the above steps S91 to S94, the specific implementation process includes: obtaining the storage location information of the archival data through the indexing mechanism of the distributed storage system to determine the physical address of the archival data. Extracting the archival data from the distributed storage system according to the physical address of the archival data to obtain the unique identifier of the archive. In the distributed storage system, the technology of quickly locating the data storage location through the indexing mechanism; distributed storage disperses the data storage on multiple nodes, and the indexing mechanism is used to manage and find these dispersed data. Using the unique identifier of the archive to query the corresponding encryption policy in the encryption policy library to determine the encryption algorithm type and key information of the archive. Loading the corresponding decryption algorithm module from the decryption algorithm library according to the encryption algorithm type and key information to initialize the decryption parameters. Using the initialized decryption parameters to decrypt the extracted archival data to generate the decrypted archival data. If the integrity check of the decrypted archival data passes, the decrypted archival data is stored in the temporary buffer for subsequent access.
[0149] S10. Output according to the optimized encryption policy, the optimized storage scheme and the decryption scheme to obtain an archive management scheme.
[0150] Exemplarily, a data readability detection mechanism is established to verify the data integrity during the decompression process, and if data corruption is found, a data repair process is triggered. Through the storage space monitoring module, the usage of the storage space is obtained in real time, and combined with the access frequency prediction model, the compression policy is dynamically adjusted. According to the optimized encryption policy, the optimized storage scheme and the decryption scheme, an archive management scheme is generated to continuously optimize the utilization rate of the storage space.
[0151] In a realizable manner of the first aspect, the decrypting according to the encrypted archive using the distributed storage indexing mechanism to obtain a decryption scheme includes:
[0152] In summary, the content of the file is identified according to the original file data to obtain the sensitivity level; an encryption policy is obtained by performing algorithm matching on the sensitivity level based on a pre-stored encryption policy library; the file is encrypted according to the encryption policy to obtain an encrypted file; core information is extracted from the original file data to obtain standardized data; the standardized data is classified and stored according to pre-stored storage rules to obtain storage nodes; storage nodes are allocated based on the access frequency of the file to obtain a storage plan; the encryption policy and the storage plan are dynamically adjusted based on the application scenario to obtain an optimized encryption policy and an optimized storage plan; the encrypted file is decrypted using a distributed storage index mechanism to obtain a decryption plan; and a file management plan is obtained according to the optimized encryption policy, the optimized storage plan, and the decryption plan. Through the method of the present invention, the sensitivity of the file is identified through natural language processing technology and matched with a preset encryption policy library to achieve precise encryption. When the file is stored in the library, a data compression algorithm is used to remove redundant content, extract core information, and optimize the storage space. According to the access frequency and importance of the file, a distributed storage system is used for load balancing to ensure data security and access efficiency. The present invention also uses a machine learning algorithm to continuously optimize the sensitivity recognition model to improve the accuracy rate. When the file is accessed, the data is quickly located and decrypted through a distributed index. In addition, the present invention can dynamically adjust the encryption policy and the storage plan according to business requirements to adapt to the actual application scenario. The present invention not only improves data security but also optimizes storage efficiency, providing a comprehensive solution for file management.
[0153] Referring to Figure 2 , the second embodiment of the present invention provides a data security encryption system for file management, including:
[0154] A data acquisition module 101, configured to acquire original file data;
[0155] A content recognition module 102, configured to identify the content of the file according to the original file data to obtain a sensitivity level;
[0156] A policy matching module 103, configured to perform algorithm matching on the sensitivity level based on a pre-stored encryption policy library to obtain an encryption policy;
[0157] A file encryption module 104, configured to encrypt the file according to the encryption policy to obtain an encrypted file;
[0158] An information extraction module 105, configured to extract core information from the original file data to obtain standardized data;
[0159] A data storage module 106, configured to classify and store the standardized data according to pre-stored storage rules to obtain storage nodes;
[0160] A storage allocation module 107, configured to allocate the storage nodes based on the access frequency of the files to obtain a storage plan;
[0161] A scheme optimization module 108, configured to dynamically adjust the encryption policy and the storage plan based on the application scenario to obtain an optimized encryption policy and an optimized storage plan;
[0162] An archive decryption module 109, configured to decrypt the encrypted archive using a distributed storage index mechanism to obtain a decryption plan;
[0163] An output module 110, configured to output according to the optimized encryption policy, the optimized storage plan, and the decryption plan to obtain an archive management plan.
[0164] It should be noted that a data security encryption system for archive management provided by an embodiment of the present invention is used to execute all process steps of a data security encryption method for archive management in the above embodiment. The working principles and beneficial effects of the two correspond one by one, and thus will not be elaborated here.
[0165] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a data security encryption program for archive management. When the processor executes the computer program, the steps in the above embodiments of various data security encryption methods for archive management are implemented, such as Figure 1 the step S1 shown. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above device embodiments are implemented, such as the data acquisition module.
[0166] Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0167] The electronic device may be a computing device such as a desktop computer, a notebook, a palm computer, and a smart tablet. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above components are only examples of the electronic device and do not constitute a limitation on the electronic device. It may include more or fewer components than the above, or combine some components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0168] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects various parts of the entire electronic device through various interfaces and circuits.
[0169] The memory can be used to store the computer program and / or module. The processor realizes various functions of the electronic device by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.), etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0170] Among them, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0171] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative work.
[0172] The above-described specific embodiments have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A data security encryption method for archive management, characterized in that: Executed by a computer, including: Access to original archival data; Identify the archive content according to the original archive data to obtain a sensitivity level; Performing algorithm matching on the sensitivity level based on a pre-stored encryption strategy library to obtain an encryption strategy; Encrypting the file according to the encryption strategy to obtain an encrypted file; Extracting core information from the original archival data to obtain standardized data; Classify and store the standardized data using a pre-stored storage rule to obtain storage nodes; Allocating the storage nodes based on the access frequency of the archive to obtain a storage solution; Dynamically adjusting the encryption strategy and the storage solution based on the application scenario to obtain an optimized encryption strategy and an optimized storage solution; Decrypting the encrypted archive using a distributed storage index mechanism to obtain a decryption solution; Output is performed based on the optimized encryption strategy, the optimized storage solution and the decryption solution to obtain a file management solution.
2. The data security encryption method for archive management according to claim 1, characterized in that: The identifying of the archive content according to the original archive data to obtain a sensitivity level includes: Read the content according to the original archive data to obtain text data; Perform word segmentation on the text data based on natural language processing technology to obtain key information; Using a pre-stored sensitive word library to judge based on the key information, if the archive content in the original archive data contains sensitive words, then marking the archive corresponding to the original archive data as a sensitive archive; Matching sensitive words and sensitive sentences according to the sensitive files to obtain sensitive information; wherein the sensitive information includes the frequency of occurrence of sensitive words and the number of occurrences of sensitive sentences; Calculating a score based on the sensitive file using a pre-stored sensitivity level assessment model to obtain a sensitivity level score; The sensitivity level is graded using a pre-stored archive grading classification system according to the sensitivity level score to obtain a sensitivity level.
3. The data security encryption method for file management according to claim 2 is characterized in that: The step of calculating the score based on the sensitive file using a pre-stored sensitivity level assessment model to obtain a sensitivity level score includes: The sensitivity level score is calculated using the following formula: Among them, S represents the sensitivity level score, n represents the number of different types of sensitive words in the pre-stored sensitive word library, and w i is the weight coefficient of the i-th sensitive word, f i represents the frequency of occurrence of sensitive words in the i-th category, m represents the number of different types of sensitive sentences in the pre-stored sensitive word library, v j is the weight coefficient of the jth sensitive sentence, g j Represents the number of occurrences of the j-th sensitive sentence pattern.
4. The data security encryption method for archive management according to claim 1, characterized in that: The algorithm matching of the sensitivity level based on the pre-stored encryption strategy library to obtain the encryption strategy includes: Extracting an algorithm from a pre-stored encryption strategy library according to the sensitivity level to obtain an encryption algorithm level; The encryption algorithm level is matched with a specific algorithm based on a pre-stored encryption strategy library to obtain an encryption strategy.
5. The data security encryption method for archive management according to claim 1, characterized in that: The storage nodes are allocated based on the access frequency of the archive to obtain a storage solution, including: Classifying the archives based on the pre-stored access frequencies to obtain archive access levels; wherein the archive access levels include high access heat, medium access heat and low access heat; Matching the archive access level with the storage node to obtain the archive storage location; The archive storage locations obtained by matching the archive access levels are integrated to obtain a storage solution.
6. The data security encryption method for file management according to claim 1, characterized in that: The dynamically adjusting the encryption strategy and the storage solution based on the application scenario to obtain an optimized encryption strategy and an optimized storage solution includes: Read archives in real time based on application scenarios to obtain real-time access frequency and real-time sensitivity; The encryption strategy and the storage scheme are adjusted in real time according to the real-time access frequency and the real-time sensitivity to obtain an optimized encryption strategy and an optimized storage scheme.
7. The data security encryption method for file management according to claim 1, characterized in that: Decrypting the encrypted archive using a distributed storage index mechanism to obtain a decryption solution includes: Using a distributed storage index mechanism to read the encrypted archive, a physical address is obtained; Extracting the file according to the physical address to obtain a file identifier; querying the archive identifier based on a pre-stored encryption policy library to obtain key information; The key information is loaded through a decryption algorithm library to obtain a decryption solution.
8. A data security encryption system for archive management, characterized in that: include: A data acquisition module, used to acquire original archive data; A content identification module, used to identify the archive content according to the original archive data to obtain a sensitivity level; A policy matching module, used to perform algorithm matching on the sensitivity level based on a pre-stored encryption policy library to obtain an encryption policy; The file encryption module is used to encrypt the file according to the encryption strategy to obtain an encrypted file; An information extraction module, used to extract core information from the original archive data to obtain standardized data; A data storage module, used to classify and store the standardized data using pre-stored storage rules to obtain storage nodes; A storage allocation module, used to allocate the storage nodes based on the access frequency of the archives to obtain a storage solution; A scheme optimization module, used to dynamically adjust the encryption strategy and the storage scheme based on the application scenario to obtain an optimized encryption strategy and an optimized storage scheme; An archive decryption module, used to decrypt the encrypted archive using a distributed storage index mechanism to obtain a decryption solution; The output module is used to output according to the optimized encryption strategy, the optimized storage scheme and the decryption scheme to obtain the archive management scheme.
9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the data security encryption method for archive management as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the data security encryption method for archive management as described in any one of claims 1 to 7.
Citation Information
Cited By
Human resource archive filing management system
CN120743852A
Automatic processing system for archiving and destroying employee data
CN121071922A