Document encryption method and system, program product and storage medium
Through technical means such as dynamically generating data keys, decentralized storage and time-efficient session keys, the problems of unreasonable key leakage and resource consumption in existing document encryption technology are solved, and the security protection and management of documents under different secret levels are realized.
Patent Information
- Application Number
- CN202510521900.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-22
AI Technical Summary
The existing document encryption and secure storage technology lacks flexibility and cannot dynamically adjust the encryption method according to the document encryption level. It relies on single storage and static permission management, resulting in a reduction in overall security when key leaks and unreasonable resource consumption.
The data key is dynamically generated according to the document cryptographic level, the data key is encrypted using the inaccessible master key, and a mapping relationship storage is established. The encrypted document is divided into multiple data blocks and stored separately in multiple storage nodes. Combined with time-sensitive session keys and permission verification, operation information is recorded to build a traceability map.
It realizes proper protection of documents in different secret scenarios, enhances key confidentiality, avoids security threats caused by single key leakage, reasonably controls resource consumption, and improves security management level.
Smart Images

Figure CN120354453A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data security, and in particular, to a document encryption method, system, program product, and storage medium. Background Art
[0002] In the current era of rapid digital development, a vast amount of important information is carried in documents, covering corporate trade secrets, personal privacy data, government classified documents, and so on. The security of these document data is of crucial importance. Once security issues occur, such as leakage, tampering, etc., it will bring serious consequences to the relevant parties.
[0003] Most of the current mainstream document encryption and secure storage technologies adopt relatively fixed modes. In terms of encryption, basically, fixed encryption algorithms are used to process documents, that is, regardless of the classification level of the documents, the same encryption method is uniformly adopted. For example, both ordinary notice documents within an enterprise and confidential documents involving major business decisions are encrypted according to the established encryption algorithm. In terms of storage, it often relies on a single storage server or a centralized cloud storage platform. For example, many enterprises store documents uniformly in a self-built single storage server or choose the centralized cloud storage service of a certain cloud service provider. In the permission management link, most are based on relatively simple role divisions, such as roles like employees, department heads, enterprise executives, etc. to assign corresponding permissions. The permission labels are relatively static and lack the ability to be dynamically adjusted according to the actual situation.
[0004] However, with the increasingly complex actual application scenarios and the continuous improvement of document security requirements, such related technical solutions have exposed many problems. If the key is accidentally leaked or cracked, the security of all documents will be greatly reduced. If the complexity of the key algorithm is increased, the resource consumption will also be unbearable. Summary of the Invention
[0005] The present application provides a document encryption method, system, program product, and storage medium, which are used to reasonably control resource consumption according to the actual situation, so that documents can be properly protected in different classification level scenarios.
[0006] First aspect, the present application provides a document encryption method, including: receiving an uploaded document, dynamically generating a data key corresponding to the document according to the classification level of the document, where the data key is used to encrypt or decrypt the document; encrypting the document with the data key to obtain an encrypted document; encrypting the data key with a built-in master key to obtain an encrypted data key, where the master key is an inaccessible key; establishing a mapping relationship between the encrypted data key and the encrypted document and storing it; when receiving an access request for the document from a user, verifying the access permission of the user; after the verification passes, obtaining the encrypted data key and the encrypted document; decrypting the encrypted data key with the master key to obtain the data key; decrypting the encrypted document with the data key to obtain the document; generating a time-limited session key and sending it to the user, where the validity period of the session key is a preset duration; during the user's access to the document, verifying whether the user holds the session key and whether the session key is within the validity period; if the user holds the session key and the session key is within the validity period, allowing the user to access the document; if the user does not hold the session key or the session key has expired, prohibiting the user from accessing the document.
[0007] By adopting the above technical solution, a corresponding data key is dynamically generated according to the classification level of the document, so that documents with different classification levels have encryption methods adapted to their own security requirements. The data key is encrypted with an inaccessible master key, enhancing the confidentiality of the key itself. Even if the data key is accidentally leaked, it cannot be cracked without the master key. Then, a mapping relationship between the encrypted data key and the encrypted document is established and stored, facilitating subsequent calls and management. In the user access link, the permission is verified and a session key with a preset duration is generated. The session key needs to be within the validity period and held by the user to access the document, effectively preventing the situation where the security of all documents is threatened due to the leakage of a single key. It not only ensures the security of the document but also reasonably controls resource consumption according to the actual situation, enabling the document to be properly protected in different classification level scenarios.
[0008] Combined with some embodiments of the first aspect, in some embodiments, after the step of establishing a mapping relationship between the encrypted data key and the encrypted document and storing it, the method further includes: when receiving a classification level adjustment instruction for the document, obtaining the encrypted data key and the encrypted document; decrypting the encrypted data key with the master key to obtain the data key; decrypting the encrypted document with the data key to obtain the document; regenerating a new data key corresponding to the document according to the adjusted classification level; encrypting the document with the new data key to obtain a new encrypted document; encrypting the new data key with the built-in master key to obtain a new encrypted data key; establishing a mapping relationship between the new encrypted document and the new encrypted data key and storing it, and replacing the encrypted data key and the encrypted document.
[0009] By adopting the above technical solution, when a document classification level adjustment instruction is received, the original encrypted data key is decrypted using the master key to obtain the original data key, and then the original encrypted document is decrypted. Then, a new data key corresponding to the adjusted classification level is regenerated. The new data key is a key element adapted to the new classification level and can encrypt the document according to the new security requirements. The document is re-encrypted using the new data key to obtain a newly encrypted document, and the new data key is encrypted again using the master key to obtain a newly encrypted data key. Finally, a mapping relationship is established between the newly encrypted document and the newly encrypted data key and stored to replace the original one. The function of flexibly adjusting the access rights and classification levels of documents according to business requirements is realized, ensuring that the document security always conforms to the changes in the actual business scenario.
[0010] Combined with some embodiments of the first aspect, in some embodiments, the step of establishing a mapping relationship between the encrypted data key and the encrypted document and storing it specifically includes: splitting the encrypted document into multiple data blocks; where the size of the encrypted document is positively correlated with the number of data blocks; the classification level of the encrypted document is positively correlated with the number of data blocks; among multiple storage nodes in a file repository, at least two target storage nodes are selected for each data block; the data blocks are respectively stored in their corresponding target storage nodes; a data check is performed on the data blocks to obtain a check value, and the check value is stored together with the data blocks; the metadata information of the encrypted document is stored in a relational database, where the metadata information includes: document identifier, document name, classification level, encryption algorithm type, the encrypted data key, and the storage location information of the data blocks in the target storage nodes.
[0011] By adopting the above technical solution, the encrypted document is split into multiple data blocks, and the size and classification level of the document are positively correlated with the number of data blocks. This means that documents with higher classification levels and larger volumes will be split more finely, facilitating subsequent refined management and storage. Then, at least two target storage nodes are selected for each data block among multiple storage nodes, so that the document data is stored dispersedly, avoiding the risk of data loss caused by a single point of failure, and at the same time serving as the basis for redundant backup. Next, a data check is performed on the data blocks to obtain a check value and store it together. If data inconsistency is found during the check, the data corruption situation can be detected in time and remedial measures can be taken. By using the redundant backup and data check mechanisms, the integrity and reliability of the document are ensured, enabling the document to be safely and completely stored in a complex storage environment.
[0012] In some embodiments in combination with some embodiments of the first aspect, after the step of allowing a user to access a document if the user holds a session key and the session key is within the valid period, the method further includes: recording operation information of the document during its life cycle, where the operation information includes: operation timestamp, operating user identifier, operation type, operation object, operation result status code, operation device information, operation network address; where the operation type includes: document creation, document access, document modification, document deletion, classification adjustment, permission change; encrypting the operation information with the master key and storing it; where, when the classification level of the document is higher than a preset level, starting a screen recording program to record the operation process of the user on the document to obtain screen recording content; encrypting and storing the screen recording content with the master key.
[0013] By adopting the above technical solution, first, the operation information of the document during its life cycle is recorded, comprehensively and meticulously capturing every operation situation of the document. The operation information is encrypted with the master key and stored. Relying on the confidentiality of the master key, the operation records are protected from being illegally obtained, ensuring the security of these key information. Especially when the classification level of the document is higher than the preset level, the screen recording program is started to record the operation process of the user on the document and also encrypted and stored with the master key, further strengthening the monitoring of the operation behavior of high-classification documents. When an abnormal situation occurs in the document, such as suspected information leakage, by viewing the encrypted stored operation information and the operation process record, it is possible to clearly trace which specific link and which user had problems, thereby improving the security management level of the document and ensuring the security and compliance of the document during use.
[0014] In some embodiments in combination with some embodiments of the first aspect, after the step of encrypting and storing the operation information with the master key, the method further includes: taking the operation object as the first associated node; taking the operating user identifier as the second associated node; taking the operation type and the operation timestamp as the associated edge between the first associated node and the second associated node; obtaining a traceability graph; receiving a query request including a query time interval, the operating user identifier, and the operation type; retrieving the associated edge in the traceability graph according to the query request to obtain the associated path of the operation information.
[0015] By adopting the above technical solution, the operation object is used as the first associated node, the operator user identifier is used as the second associated node, and the operation type and operation timestamp are used as the associated edge between the two. In this way, a traceability graph is constructed, which clearly presents the sequence of document operations and the association relationship between each operation in a structured form. After that, when a query request containing key elements such as the query time range, operator user identifier, and operation type is received, the associated edge can be accurately retrieved in the traceability graph according to these conditions, and then the association path of the operation information can be quickly sorted out. By providing such field-level fine-grained audit logs and visual traceability graphs, in the event of a security incident such as document leakage, security management personnel can quickly and accurately trace back to the source along the association path presented in the graph, such as determining which user performed what improper operation at what specific time, and then can take targeted measures in a timely manner to deal with it, improving the security management level of the document and reducing the losses caused by security risks.
[0016] Combined with some embodiments of the first aspect, in some embodiments, after the step of using the operation type and the operation timestamp as the associated edge between the first associated node and the second associated node; obtaining the traceability graph, the method further includes: judging whether the traceability graph exceeds a preset storage period according to the operation timestamp; if it exceeds the preset storage period, migrating the traceability graph to a backup database.
[0017] By adopting the above technical solution, it will judge whether the traceability graph exceeds the preset storage period according to the operation timestamp. If the traceability graph exceeds the preset storage period, it will be migrated to the backup database. Doing so can, on the one hand, avoid occupying too much precious storage space by storing a large number of traceability graphs in the main database for a long time, ensure the efficient operation of the main database, and maintain the overall fluency of the system; on the other hand, it can properly save these historical traceability information, and when it is necessary to conduct a detailed audit of the past document operation situation or trace some long-term remaining problems in the future, the corresponding traceability graph can be queried from the backup database at any time.
[0018] Combined with some embodiments of the first aspect, in some embodiments, after the step of receiving the uploaded document, the method further includes: extracting the text content of the document; identifying the keywords in the text content; predicting the classification level of the document based on the keywords using a language model; generating a classification level suggestion value according to the prediction result, the classification level suggestion value including: public level, internal level, confidential level, secret level, top secret level; outputting the classification level suggestion value; receiving a confirmation instruction or a modification instruction for the classification level suggestion value; determining the classification level of the document according to the confirmation instruction or the modification instruction.
[0019] By adopting the above technical solution, after receiving the uploaded document, the text content of the document is first extracted, which is the basic material for subsequent analysis. Then, the keywords in the text content are identified. Next, based on these keywords, the language model is used to predict the confidentiality level of the document, and a relatively scientific and accurate confidentiality level recommendation value is given according to factors such as keyword association, such as different levels like the public level and the top-secret level. Then, the confidentiality level recommendation value is output for the user to refer to. After that, the confirmation instruction or modification instruction from the user is received, and finally, the confidentiality level of the document is determined based on these instructions. This improves the efficiency and accuracy of confidentiality level determination, making the document confidentiality level more in line with its actual security requirements.
[0020] In a second aspect, the present application provides a document encryption system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the document encryption system to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0021] In a third aspect, the present application provides a computer program product containing instructions, which, when running on a document encryption system, enables the document encryption system to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium, including instructions, which, when running on a document encryption system, enables the document encryption system to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0023] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Dynamically generate corresponding data keys according to the confidentiality level of the document, so that documents with different confidentiality levels have encryption methods adapted to their own security requirements. Encrypt the data keys with an inaccessible master key to enhance the confidentiality of the keys themselves. Even if the data keys are accidentally leaked, they cannot be cracked without the master key. Then, establish a mapping relationship between the encrypted data keys and the encrypted documents for storage, which is convenient for subsequent invocation and management. In the user access link, verify the permissions and generate a session key with a preset duration. The session key can only be used to access the document within the validity period and when held by the user, effectively preventing the situation where the security of all documents is threatened due to the leakage of a single key. This not only ensures the security of the documents but also reasonably controls resource consumption according to the actual situation, enabling the documents to be properly protected in different confidentiality level scenarios.
[0024] 2. When receiving an instruction to adjust the classification level of a document, use the master key to decrypt the original encrypted data key to obtain the original data key, and then decrypt the original encrypted document. Then, regenerate a new data key corresponding to the adjusted classification level. The new data key is a key element adapted to the new classification level and can encrypt the document according to the new security requirements. Encrypt the document again with the new data key to obtain a newly encrypted document, and encrypt the new data key with the master key again to obtain a newly encrypted data key. Finally, establish a mapping relationship between the newly encrypted document and the newly encrypted data key, store it, and replace the original one. This realizes the function of flexibly adjusting the access rights and classification levels of documents according to business requirements, ensuring that the document security always conforms to the changes in the actual business scenario.
[0025] 3. Split the encrypted document into multiple data blocks. The size and classification level of the document are positively correlated with the number of data blocks. This means that documents with higher classification levels and larger volumes will be split more finely, facilitating subsequent refined management and storage. Then, select at least two target storage nodes for each data block among multiple storage nodes, so that the document data is stored dispersedly, avoiding the risk of data loss caused by a single point of failure, and at the same time serving as the basis for redundant backup. Next, perform data verification on the data blocks to obtain verification values and store them together. If data inconsistency is found during verification, the data corruption situation can be detected in a timely manner and remedial measures can be taken. By using the redundant backup and data verification mechanisms, the integrity and reliability of the document are ensured, enabling the document to be safely and completely stored in a complex storage environment. Brief Description of the Drawings
[0026] Figure 1 is a flowchart of a document encryption method in an embodiment of the present application; Figure 2 is another flowchart of a document encryption method in an embodiment of the present application; Figure 3 is another flowchart of a document encryption method in an embodiment of the present application; Figure 4 is another flowchart of a document encryption method in an embodiment of the present application; Figure 5 is an exemplary hardware structure diagram of a document encryption system in an embodiment of the present application. Detailed Implementation Manner
[0027] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items.
[0028] Hereinafter, the terms "first" and "second" are only for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0029] Please refer to Figure 1 , Figure 1 which is a schematic flow diagram of the document encryption method in the embodiments of the present application; S101. Receive the uploaded document, and dynamically generate a data key corresponding to the document according to the confidentiality level of the document, where the data key is used to encrypt or decrypt the document; It should be noted that the document here refers to the files of various types of information uploaded by the user.
[0030] Specifically, identify the received document, and then generate a unique data key corresponding to it according to the confidentiality level of the document; this data key will be used for subsequent encryption and decryption operations of the document.
[0031] The confidentiality level can be pre-entered, determined by the user, or judged according to the pre-set confidentiality evaluation rules to determine which confidentiality level the document belongs to. For example, it can be determined according to factors such as the business scope and content nature involved in the document, and no limitation is made here.
[0032] It is worth noting that there are two implications here: First, different confidentiality levels correspond to different encryption algorithms, and the more complex the confidentiality level, the higher the corresponding security factor; Second, each file has its unique data key, and files with the same confidentiality level do not share a data key.
[0033] In some embodiments, a five-layer confidentiality model is adopted, namely public, internal, secret, confidential, and top secret.
[0034] S102. Encrypt the document with the data key to obtain an encrypted document; encrypt the data key with the built-in master key to obtain an encrypted data key, where the master key is an inaccessible key; Using the data key generated in the previous step, encrypt the original document through a specific encryption algorithm to convert the document content into an encrypted document, keeping its content confidential. Then, use the built-in and inaccessible master key and apply the supporting encryption algorithm to encrypt the generated data key, thereby obtaining an encrypted data key, which further enhances the confidentiality of the data key, prevents it from being easily leaked or cracked, and ensures the security of the entire document encryption system.
[0035] S103. Establish a mapping relationship between the encrypted data key and the encrypted document and store it. Associate the encrypted data key with the encrypted document, for example, establish a mapping relationship by assigning the same identification number to them or recording them in the same index table. Then, according to the established storage strategy, select a storage location to save the encrypted data key and the encrypted document with this mapping relationship, so that when the document needs to be accessed or related operations are performed on the document later, the corresponding encrypted data key can be quickly and accurately found based on this mapping relationship, and then subsequent decryption and other operations can be carried out.
[0036] In some embodiments, the encrypted data key and the encrypted document can also be stored at one address, which is not limited here.
[0037] It should be emphasized that steps S102 and S103 are actually solutions to the new technical problems caused by S101. In S101, since each document is equipped with its own unique data key and the generation algorithms of the data keys of different documents are also different, a new and difficult problem will arise, that is, the difficulty of storing the data key increases significantly. If only the method described in step S103 is used to try to solve this problem, it is very likely that the data key can be easily obtained while the encrypted document is obtained. This obviously violates the original intention of encryption and greatly reduces the security of the entire encryption system.
[0038] Step S103 specifically includes: S1031. Divide the encrypted document into multiple data blocks; where the size of the encrypted document is positively correlated with the number of data blocks; the classification level of the encrypted document is positively correlated with the number of data blocks. First, relevant attribute information of the encrypted document will be obtained, including its size, classification level, and other content. Then, according to the pre-set splitting rules, the basic splitting granularity will be determined based on the size of the encrypted document. The larger the document, the more data blocks will be split for easier management and subsequent storage, transmission, and other operations. At the same time, in combination with the classification level of the encrypted document, for documents with a high classification level, considering higher security and reliability, they will also be split more finely, that is, more data blocks will be divided to ensure a more refined management mode for important documents during storage and reduce the risk of affecting the entire document due to local problems. For example, a large project document with a top-secret classification will be split into more data blocks, while a relatively small daily notice document with a lower classification level will have a relatively smaller number of split data blocks.
[0039] S1032. Among multiple storage nodes in the file repository, select at least two target storage nodes for each data block; First, the status of all storage nodes in the file repository will be detected to understand information such as the available storage space, network connection status, and current load of each storage node to ensure they are in a normal usable state. Then, for each data block, according to the pre-set storage policies and algorithms, considering various performance indicators of the storage nodes, the importance of the data block (related to the classification of the original encrypted document), and other factors, at least two suitable target storage nodes will be selected from numerous storage nodes. For example, for data blocks with a higher classification level, storage nodes with higher security and better stability may be preferentially selected, and through a decentralized selection method, all backups of the same data block are avoided from being stored on adjacent or strongly related storage nodes, thereby achieving decentralized storage of data and reducing the risk of data loss due to the failure of a certain storage node.
[0040] In some embodiments, the classification level of the encrypted document is positively correlated with the number of target storage nodes, that is, the higher the importance of the data block, the lower the risk of data loss.
[0041] S1033. Store the data blocks into their corresponding target storage nodes respectively; S1034. Perform data verification on the data blocks to obtain verification values and store the verification values together with the data blocks; S1035. Store the metadata information of the encrypted document into a relational database, where the metadata information includes: document identifier, document name, classification level, encryption algorithm type, encrypted data key, and storage location information of the data blocks in the target storage nodes.
[0042] For document identification, a unique identification string will be generated according to the established generation rules, combining elements such as the document creation time, the affiliated project, and the creator, through corresponding algorithms, and it will be recorded. The document name is given by the user at the beginning of document creation or during upload. If there are naming non-conformities, appropriate adjustments may be made to make its meaning clear and comply with the overall naming rules. The classification level is determined by professional security personnel based on a comprehensive evaluation of the nature of the document content, the confidential information involved, etc., and the corresponding records will be updated in a timely manner as the importance of the document changes. The encryption algorithm type is determined when the document is encrypted, and the specific encryption algorithm used at that time will be recorded. For example, the symmetric encryption algorithm AES is used to balance encryption efficiency and security. The encrypted data key is generated in relation to factors such as the classification level of the document. After generation, its relevant information will be saved, and its accuracy and integrity will be ensured during recording to prevent errors or losses in the key information. For the storage location information of data blocks in the target storage nodes, when the data blocks are stored in the target storage nodes before, the storage management will record in real time the detailed paths and other location situations where each data block is specifically stored. At this time, these accurate location information will be collected, and then the metadata including the document identification, document name, classification level, encryption algorithm type, encrypted data key, and data block storage location information will be sorted out in a certain format and order, so as to be stored in the corresponding table structure in the relational database for subsequent query and management through various conditions.
[0043] It can be seen that when an encrypted document is divided into multiple data blocks, the size and classification level of the document are positively correlated with the number of data blocks. This means that documents with higher classification levels and larger volumes will be divided more finely, facilitating subsequent refined management and storage. Then, at least two target storage nodes are selected for each data block among multiple storage nodes, so that the document data is stored dispersedly, avoiding the risk of data loss caused by single-point failures, and at the same time serving as the basis for redundant backup. Next, data verification is performed on the data blocks to obtain the verification values and store them together. If data inconsistency is found during verification, the data corruption situation can be detected in a timely manner and remedial measures can be taken. By using the redundant backup and data verification mechanisms, the integrity and reliability of the document are ensured, enabling the document to be safely and completely stored in a complex storage environment.
[0044] S104. When receiving a user's access request for a document, verify the user's access rights; First, the identity of the user initiating the access request will be identified, and it will be confirmed which user is initiating the request by verifying the login account, password, or other authentication methods. Then, according to the user's role setting in [relevant system], combined with the classification level of the document itself, and referring to the pre-set permission rule table, it will be judged whether the user has the access right to this document.
[0045] S105. After verification, the encrypted data key and encrypted document are obtained; When it is determined that the user has permission to access the document, the encryption data key associated with the document and the encrypted document itself will be obtained by searching the corresponding storage location or index information based on the previously established mapping relationship between the encryption data key and the encrypted document, preparing for subsequent decryption operations and ensuring that the encrypted document can be successfully restored to the original document content that can be read and operated normally.
[0046] S106, using the master key to decrypt the encrypted data key to obtain the data key; using the data key to decrypt the encrypted document to obtain the document; First, use the built-in master key and the matching decryption algorithm to decrypt the encrypted data key and restore the originally generated data key. This step is because the encrypted data key is encrypted by the master key, so only the master key can decrypt it. Then, use the restored data key and the corresponding decryption algorithm to decrypt the encrypted document and restore the ciphertext document content to the original document state that can be read and operated normally, thereby achieving decryption access to the document.
[0047] S107, generating a time-effective session key and sending it to the user, wherein the validity period of the session key is a preset period; The operation is performed before the document is provided to the user for access. Specifically, a specific key generation algorithm is used to combine the current time, user information, and some document-related parameters to generate a time-limited session key. This session key is temporary and only works for the user's access to the document. Then, the generated session key is sent to the user who initiated the access request. The user needs to use this session key to prove the legitimacy of his access when accessing the document later, and the access operation must be completed within the preset time, otherwise the access right will be lost.
[0048] S108. During the period when the user accesses the document, verify whether the user holds the session key and whether the session key is within the validity period; The user's access status will be monitored in real time, and by interacting with the user end, such as sending verification requests regularly or verifying when the user performs key operations, it will check whether the user holds the session key sent to him before, and also check whether the session key is still within the preset validity period. Only when the user meets both conditions of holding the session key and the key being within the validity period, the user will be allowed to continue to access the document normally, otherwise corresponding access restriction measures will be taken.
[0049] S109. If the user holds the session key and the session key is within the validity period, the user is allowed to access the document; When it is verified that the user does hold the session key and the session key is still within the preset validity period, this indicates that the user's behavior of accessing the document this time meets the security requirements and has legitimate access rights. Therefore, the user will be allowed to continue to normally view, edit, and perform other corresponding operations on the document, ensuring that the user can smoothly utilize the document to carry out work or obtain information under compliance.
[0050] S110. If the user does not hold the session key or the session key has expired, the user is prohibited from accessing the document.
[0051] When it is found that the user does not hold the session key, or although the user holds it but the session key has exceeded the specified validity period, it means that the user's behavior of accessing the document this time does not meet the set security requirements and cannot guarantee the legality and security of the access. Therefore, the user will be prohibited from continuing to perform any operations on the document to avoid security risks such as possible leakage of document information, thereby maintaining the confidentiality of the document and the overall security protection mechanism.
[0052] It can be seen that by dynamically generating corresponding data keys according to the classification levels of the documents, different classification-level documents have encryption methods that adapt to their own security requirements. Using the inaccessible master key to encrypt the data keys enhances the confidentiality of the keys themselves. Even if the data keys are accidentally leaked, they cannot be cracked without the master key. Then, a mapping relationship is established and stored between the encrypted data keys and the encrypted documents for convenient subsequent invocation and management. In the user access link, the permissions are verified and a session key with a preset duration is generated. The session key needs to be within the validity period and held by the user to access the document, effectively preventing the situation where the security of all documents is threatened due to the leakage of a single key. This not only ensures the security of the documents but also reasonably controls resource consumption according to the actual situation, enabling the documents to be properly protected in different classification-level scenarios.
[0053] Please refer to Figure 2 , Figure 2 which is another process schematic diagram of the document encryption method in the embodiments of the present application; In some embodiments, after step S103, it further includes: S201. When receiving an instruction to adjust the classification level of the document, obtain the encrypted data key and the encrypted document; Among them, the classification level adjustment instruction refers to a command message issued by a person or module with corresponding permissions, used to indicate the change of the current classification level of a certain document.
[0054] S202. Use the master key to decrypt the encrypted data key to obtain the data key; use the data key to decrypt the encrypted document to obtain the document; It should be noted that the principle and process of this application are similar to those of step S106. The relevant principle and process can be referred to step S106 and will not be limited here.
[0055] S203. Regenerate a new data key corresponding to the document according to the adjusted confidentiality level; It should be noted that the principle and process of this application are similar to those of step S201. The relevant principle and process can be referred to step S201 and will not be limited here.
[0056] S204. Encrypt the document with the new data key to obtain a newly encrypted document; encrypt the new data key with the built-in master key to obtain a newly encrypted data key; It should be noted that the principle and process of this application are similar to those of step S202. The relevant principle and process can be referred to step S202 and will not be limited here.
[0057] S205. Establish a mapping relationship between the newly encrypted document and the newly encrypted data key and store them, and replace the encrypted data key and the encrypted document.
[0058] It should be noted that the principle and process of this application are similar to those of step S203. The relevant principle and process can be referred to step S203 and will not be limited here.
[0059] It can be seen that when a document confidentiality level adjustment instruction is received, the original encrypted data key is decrypted with the master key to obtain the original data key, and then the original encrypted document is decrypted. Then, a new data key corresponding to it is regenerated according to the adjusted confidentiality level. The new data key is a key element adapted to the new confidentiality level and can encrypt the document according to the new security requirements. The document is re-encrypted with the new data key to obtain a newly encrypted document, and the new data key is encrypted again with the master key to obtain a newly encrypted data key. Finally, a mapping relationship between the newly encrypted document and the newly encrypted data key is established and stored to replace the original ones. The function of flexibly adjusting the access permission and confidentiality level of the document according to business requirements is realized, and it is ensured that the document security always conforms to the changes in the actual business scenario.
[0060] Please refer to Figure 3 , Figure 3 which is another process schematic diagram of the document encryption method in the embodiment of this application; In some embodiments, after step S109, it further includes: S301. Record the operation information of the document during its life cycle. The operation information includes: operation timestamp, operating user identifier, operation type, operation object, operation result status code, operation device information, operation network address; where the operation type includes: document creation, document access, document modification, document deletion, confidentiality level adjustment, permission change; Whenever an operation on a document occurs, various types of information related to that operation are automatically collected. Whether the operation is initiated manually or executed automatically, the recording mechanism will be triggered. The operation dynamics of the document will be monitored in real time. Once an operation is detected, the corresponding operation information will be collected according to the predefined rules and processes, in preparation for subsequent encrypted storage and security management. When a document operation is detected, the operation information will be obtained from different data sources according to the preset rules. The operation timestamp can be obtained from the clock module to ensure the accuracy of time; the operator identification is obtained through user authentication to ensure the uniqueness and reliability of the identification; the operation type is judged and classified according to the specific behavior of the operation; the operation object is determined through the index information of document management; the operation result status code is returned by the operation execution module to reflect the actual result of the operation; the operation device information is obtained through device identification technology, such as the hardware information and operation information of the device; the operation network address is obtained through the network monitoring module to record the network location when the operation occurs.
[0061] S302. Encrypt the operation information with the master key and store it. S303. Among them, when the confidentiality level of the document is higher than the preset level, start the screen recording program, record the operation process of the user on the document to obtain the screen recording content, and encrypt and store the screen recording content with the master key.
[0062] The confidentiality level of the document is monitored in real time. When it is found that the confidentiality level of a certain document is higher than the preset level, the screen recording program will be automatically triggered. The screen recording program will start recording all the operation processes on the screen when the user operates on the document. After the recording is completed, the master key will be obtained by calling the key management, and the master key will be used to encrypt the screen recording content. The encrypted screen recording content will be stored in a dedicated storage area, such as a high-security database or an encrypted storage device. At the same time, corresponding identifiers and metadata will be added to it for subsequent query and management.
[0063] It can be seen that first, the operation information of the document during its life cycle is recorded, comprehensively and meticulously capturing every operation of the document. The operation information is encrypted with the master key and stored. Relying on the confidentiality of the master key, the operation records are protected from being illegally obtained, ensuring the security of these key information. Especially when the confidentiality level of the document is higher than the preset level, the screen recording program is started, the operation process of the user on the document is recorded and also encrypted and stored with the master key, further strengthening the monitoring of the operation behavior of high-confidentiality documents. When an abnormal situation occurs in the document, such as suspected information leakage, by viewing the encrypted operation information and the recorded operation process, it is possible to clearly trace which specific link and which user had problems, thereby improving the security management level of the document and ensuring the security and compliance of the document during use.
[0064] After step S302, the method further includes: S304, taking the operation object as the first associated node; Extract the content related to the operation object from the previously recorded and stored operation information, and set it as the first associated node according to the established rules. For example, for each document operation record, the corresponding document identification or name and other information will be extracted separately and standardized to meet the format requirements of the nodes in the graph construction, so that it can be accurately connected to other nodes through associated edges in the future, thereby gradually building a traceability graph that can reflect the association of document operations.
[0065] S305, taking the operating user identifier as a second associated node; Extract the relevant content of the operation user ID from the stored operation information, and perform normalization processing according to the set rules so that it can adapt to the requirements of the nodes in the traceability graph. For example, no matter what format the operation user ID is originally stored in, it will be uniformly converted into a standard string format, and necessary verification will be performed to ensure its uniqueness and accuracy, and then it will be set as the second associated node to prepare for the subsequent connection with the operation object (first associated node) through the associated edge to build a complete traceability graph relationship network.
[0066] S306: Use the operation type and the operation timestamp as an association edge between the first association node and the second association node to obtain a traceability graph; The operation type and operation timestamp content corresponding to each operation are extracted from the stored operation information in turn, and then integrated according to certain rules. For example, the operation type and timestamp are concatenated into a string in a specific format, or constructed into a structured data object containing these two attributes, and then used as an association edge to connect the corresponding first association node (operation object) and the second association node (operation user identifier). In this way, each operation record is converted into the association relationship between nodes and edges in the graph, and a complete traceability graph that can clearly reflect the overall picture of the document operation is gradually constructed.
[0067] S307, receiving a query request including a query time interval, the operation user identifier, and the operation type; The user will enter the query time interval, operation user ID, operation type and other related information in the specified format through the provided query interface or specific query interface, and then submit the query request. After receiving this request, the legality and format of the request will be verified first, such as checking whether the format of the time interval is correct and whether the operation user ID exists. Only after the verification is passed, will the search operation be prepared in the constructed traceability map based on the key elements in the request to find the associated path of the operation information that meets the requirements.
[0068] S308. Retrieve the associated edges in the traceability graph according to the query request to obtain the associated path of the operation information.
[0069] Based on conditions such as the query time range, operator user identifier, operation type, etc. included in the query request, traversal search is performed starting from each node in the traceability graph. For example, first locate the corresponding second associated node (user node) based on the operator user identifier, then combine the operation type, and filter out the edges that meet the operation type requirements among the associated edges connected to this user node. Further, according to the query time range, determine the edges whose time is also within the range among these associated edges that meet the operation type. Through such layer-by-layer filtering and traversal methods, a series of interconnected nodes and edges are gradually found. The path formed by these nodes and edges is the associated path of the operation information that meets the query request, which can accurately reflect the document operation situation to be queried.
[0070] It can be seen that taking the operation object as the first associated node, the operator user identifier as the second associated node, and the operation type and operation timestamp as the associated edges between the two, the traceability graph is constructed in this way. This graph clearly presents the sequence of document operations and the associated relationships between operations in a structured form. After that, when a query request containing key elements such as the query time range, operator user identifier, operation type, etc. is received, the associated edges can be accurately retrieved in the traceability graph according to these conditions, and then the associated path of the operation information can be quickly sorted out. By providing this field-level fine-grained audit log and visual traceability graph, when security incidents such as document leakage occur, security managers can quickly and accurately trace back to the source along the associated path presented in the graph, such as determining which user performed what improper operation at what specific time, and then can take targeted measures in a timely manner to deal with it, improving the security management level of the document and reducing the losses caused by security risks.
[0071] After step S306, it further includes: S309. Judge whether the traceability graph exceeds the preset storage period according to the operation timestamp; Extract the operation timestamp information from the associated edges of the traceability graph, then compare the current time with the operation timestamp, and calculate the time elapsed from the occurrence of the operation to the current moment for the traceability graph. Then, compare this elapsed time with the preset storage period to judge whether the traceability graph exceeds the storage period.
[0072] S310. If it exceeds the preset storage period, migrate the traceability graph to the backup database.
[0073] First, confirm whether the connection to the backup database is normal. Then, completely copy the traceability graphs that have exceeded the storage period from the main database and insert them into the corresponding tables or storage areas of the backup database according to the storage format and structure requirements of the backup database. After the migration is completed, the migrated traceability graph data will be deleted from the main database to ensure the effective release of the storage space of the main database.
[0074] It can be seen that it will judge whether the traceability graph exceeds the preset storage period based on the operation timestamp. If the traceability graph exceeds the preset storage period, it will be migrated to the backup database. Doing so can, on the one hand, avoid occupying too much precious storage space by storing a large number of traceability graphs in the main database for a long time, ensure the efficient operation of the main database, and maintain the overall fluency; on the other hand, it can properly save these historical traceability information, and when it is necessary to conduct a detailed audit of past document operation situations or trace some long-standing problems in the future, the corresponding traceability graphs can be queried from the backup database at any time.
[0075] Please refer to Figure 4 , Figure 4 which is another process schematic diagram of the document encryption method in the embodiments of the present application; In some embodiments, after step S101, it further includes: S401. Extract the text content of the document; It will select a suitable parsing method according to the format type of the document. For common text formats, such as TXT files, the text content can be directly read; for Word documents, a special document parsing library (such as the python-docx library in python) will be used to parse its text; for PDF files, a PDF parsing tool (such as pdfplumber, etc.) will be used to extract the text. During the extraction process, some unnecessary format information and special characters will be removed, and only the pure text content will be retained to prepare for identifying keywords later.
[0076] S402. Identify the keywords in the text content; It will preprocess the extracted text content, such as word segmentation and removing stop words (such as meaningless words like "of", "is", "in", etc.). Then, keyword extraction algorithms, such as the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, the TextRank algorithm, etc., will be used to calculate the scores of each word according to factors such as the frequency of occurrence and importance of the word in the text. The words with higher scores are the keywords. Finally, the identified keywords will be screened and sorted, and some repeated or irrelevant words will be removed to obtain the final keyword list.
[0077] S403. Based on the keywords, use a language model to predict the confidentiality level of the document; It will combine keywords into a format suitable for input to the language model. For example, the keywords are concatenated into a string using specific delimiters. Then this input data is sent to the language model, which will analyze and process the input keywords according to its internal neural network structure and the weight parameters obtained through training, and calculate the probabilities of the document belonging to different classification levels. Finally, the most likely classification level of the document is determined based on the probability values.
[0078] S404. Generate a classification level recommendation value based on the prediction result. The classification level recommendation value includes: public level, internal level, confidential level, secret level, top secret level; Based on the probabilities of the document belonging to different classification levels in the prediction result, it will select the classification level with the highest probability as the preliminary recommendation value. Then, this recommendation value is verified and adjusted, taking into account some special situations and rules. For example, certain keywords may imply that the document requires a higher classification level, etc. Finally, the adjusted classification level is used as the final classification level recommendation value.
[0079] S405. Output the classification level recommendation value; It will select an appropriate way to output the classification level recommendation value according to different application scenarios and user interface designs. If it is a web application, the classification level recommendation value will be prominently displayed on the page, such as using large fonts, different colors, etc.; if it is a desktop application, the recommendation value will be displayed in a pop-up window or a specific prompt area; if it provides services through an API interface, the classification level recommendation value will be returned to the caller in formats such as JSON or XML.
[0080] S406. Receive a confirmation instruction or a modification instruction for the classification level recommendation value; It will provide corresponding operation buttons or input boxes on the user interface, such as a "confirm" button and a "modify" button. When the user clicks the "confirm" button, the confirmation instruction will be captured; when the user clicks the "modify" button and selects or enters a new classification level, the modification instruction will be captured. It will monitor the user's operations in real time. Once it detects that the user issues an instruction, it will receive and parse the instruction.
[0081] S407. Determine the classification level of the document according to the confirmation instruction or the modification instruction.
[0082] It will parse the received instruction. If it is a confirmation instruction, it will set the classification level of the document to the previously generated classification level recommendation value; if it is a modification instruction, it will update the classification level of the document to the value modified by the user. Then, the determined classification level information will be stored in the database, and at the same time, the metadata information of the document will be updated so that the corresponding classification level rules can be accurately applied during the subsequent management and use of the document.
[0083] It can be seen that after receiving the uploaded document, the text content of the document is first extracted, which is the basic material for subsequent analysis. Then, the keywords in the text content are identified. Next, based on these keywords, a language model is used to predict the classification level of the document, and a relatively scientific and accurate classification level recommendation value is given according to factors such as keyword association, such as different levels like the public level and the top-secret level. Then, the classification level recommendation value is output for the user to refer to. After that, the confirmation instruction or modification instruction of the user is received, and finally, the classification level of the document is determined according to these instructions. This improves the efficiency and accuracy of classification level determination, making the classification level of the document more in line with its actual security requirements.
[0084] The following introduces the exemplary document encryption system 500 provided by the embodiments of the present application. Figure 5 It is an exemplary hardware structure diagram of the document encryption system 500 provided by the embodiments of the present application.
[0085] In some embodiments, the document encryption system 500 is a computer device or the document encryption system 500 includes a computer device. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with other external terminals or servers through a network connection. In some embodiments, the network interface can be a wired network interface, and in some embodiments, the network interface can also be a wireless network interface. The computer program, when executed by the processor, implements the method in the embodiments of the present application.
[0086] Those skilled in the art can understand that Figure 5 the structure shown in [figure reference] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0087] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0088] As used in the foregoing embodiments, depending on the context, the term "when" may be construed to mean "if", "after", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "when determining" or "if (the stated condition or event) is detected" may be construed to mean "if determined", "in response to determining", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0089] In the foregoing embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0090] Those of ordinary skill in the art can understand all or part of the processes in the methods of the foregoing embodiments. These processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the foregoing method embodiments. The foregoing storage media include various media that can store program codes, such as ROM or random access memory RAM, magnetic disks, or optical discs.
Claims
1. A document encryption method, characterized in that, Including: Receiving the uploaded document, dynamically generating a data key corresponding to the document according to the classification level of the document, where the data key is used to encrypt or decrypt the document; Encrypting the document using the data key to obtain an encrypted document; encrypting the data key using a built-in master key to obtain an encrypted data key, where the master key is an inaccessible key; Establishing a mapping relationship between the encrypted data key and the encrypted document and storing it; When receiving a user's access request for the document, verifying the user's access permission; After passing the verification, obtaining the encrypted data key and the encrypted document; Decrypting the encrypted data key using the master key to obtain the data key; decrypting the encrypted document using the data key to obtain the document; Generating a time-limited session key and sending it to the user, where the validity period of the session key is a preset duration; During the user's access to the document, verifying whether the user holds the session key and whether the session key is within the validity period; If the user holds the session key and the session key is within the validity period, allowing the user to access the document; If the user does not hold the session key or the session key has expired, prohibiting the user from accessing the document.
2. The method according to claim 1, characterized in that, After the step of establishing a mapping relationship between the encrypted data key and the encrypted document and storing it, the method further includes: When receiving an instruction to adjust the classification level of the document, obtaining the encrypted data key and the encrypted document; Decrypting the encrypted data key using the master key to obtain the data key; decrypting the encrypted document using the data key to obtain the document; Regenerating a new data key corresponding to the document according to the adjusted classification level; Encrypting the document using the new data key to obtain a new encrypted document; encrypting the new data key using the built-in master key to obtain a new encrypted data key; Establishing a mapping relationship between the new encrypted document and the new encrypted data key and storing it, and replacing the encrypted data key and the encrypted document.
3. The method according to claim 1, characterized in that The step of establishing a mapping relationship between the encrypted data key and the encrypted document and storing it specifically includes: Dividing the encrypted document into multiple data blocks; where the size of the encrypted document is positively correlated with the number of data blocks; the classification level of the encrypted document is positively correlated with the number of data blocks; Selecting at least two target storage nodes for each data block among multiple storage nodes in the file repository; Storing the data blocks into their corresponding target storage nodes respectively; Performing data verification on the data blocks to obtain a verification value, and storing the verification value together with the data blocks; Storing the metadata information of the encrypted document into a relational database, where the metadata information includes: Document identifier, document name, classification level, encryption algorithm type, the encrypted data key, and the storage location information of the data blocks in the target storage nodes.
4. The method according to any one of claims 1 to 3, characterized in that, After the step of allowing the user to access the document if the user holds the session key and the session key is within the validity period, the method further includes: Recording operation information of the document during its life cycle, where the operation information includes: Operation timestamp, operator identifier, operation type, operation object, operation result status code, operation device information, operation network address; where the operation type includes: document creation, document access, document modification, document deletion, classification adjustment, permission change; Encrypting the operation information with the master key and storing it; Wherein, when the classification level of the document is higher than a preset level, starting a screen recording program to record the operation process of the user on the document to obtain screen recording content; encrypting and storing the screen recording content with the master key.
5. The method according to claim 4, wherein After the step of encrypting the operation information with the master key and storing it, the method further includes: Regarding the operation object as the first associated node; Regarding the operator identifier as the second associated node; Regarding the operation type and the operation timestamp as the associated edge between the first associated node and the second associated node; obtaining a traceability graph; Receiving a query request including a query time range, the operator identifier, and the operation type; Retrieving the associated edge in the traceability graph according to the query request to obtain the associated path of the operation information.
6. The method according to claim 5, wherein Regarding the operation type and the operation timestamp as the associated edge between the first associated node and the second associated node; After the step of obtaining the traceability graph, the method further includes: Judging whether the traceability graph exceeds a preset storage period according to the operation timestamp; If it exceeds the preset storage period, migrating the traceability graph to a backup database.
7. The method according to claim 1, characterized in that, After the step of receiving the uploaded document, the method further includes: Extracting the text content of the document; Identifying keywords in the text content; Performing classification prediction on the document based on the keywords using a language model; Generating a classification recommendation value according to the prediction result, where the classification recommendation value includes: public level, internal level, confidential level, secret level, top secret level; Outputting the classification recommendation value; Receiving a confirmation instruction or a modification instruction for the classification recommendation value; Determining the classification level of the document according to the confirmation instruction or the modification instruction.
8. A document encryption system, characterized in that, The document encryption system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the document encryption system to execute the method according to any one of claims 1-7.
9. A computer program product comprising instructions, characterized in that, When the computer program product runs on the document encryption system, enabling the document encryption system to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the document encryption system, enabling the document encryption system to execute the method according to any one of claims 1-7.
Citation Information
Cited By
Urban business strategy data processing method
CN121327869A