An implementation system for document encryption
Through the exception perception and key optimization module, the encryption parameter adaptive adjustment module, the resource allocation and task scheduling module and the encryption risk assessment module, the problems of insecure key management and unreasonable resource allocation are solved, and the security and efficiency of document encryption are improved.
Patent Information
- Application Number
- CN202510622412.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In the existing document encryption technology, key management is unsafe, keys are prone to cracking or leaking, resource allocation is unreasonable, and risk assessment mechanisms are lacking, resulting in poor encryption effect and reduced security.
The exception perception and key optimization module, the encryption parameter adaptive adjustment module, the resource allocation and task scheduling module, and the encryption risk assessment and analysis module are used to optimize the keys using deep neural networks and genetic algorithms, and the resource task information matrix is established, resource allocation and task scheduling is performed, and early warning signals are generated through encryption risk assessment.
It improves the security and encryption efficiency of keys, rationally allocates resources, promptly discovers abnormal access behaviors and potential security threats, and enhances encryption effect and security.
Smart Images

Figure CN120145426B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an implementation system for document encryption. Background Art
[0002] In current document encryption processing, the following problems still exist: key management problem. The security of the key in document encryption is crucial, but currently, there may be risks of the key being cracked or leaked. At the same time, the update and optimization of the key may not be timely, resulting in reduced encryption security; fixed encryption parameters. Traditional document encryption often uses fixed encryption parameters and cannot be adaptively adjusted according to document types, importance, or environmental changes, which may lead to poor encryption effects or resource waste; unreasonable resource allocation. During the document encryption process, resource allocation and task scheduling may not be reasonable enough, resulting in too long encryption time for some documents or insufficient encryption resources, affecting the overall encryption efficiency and performance; lack of risk assessment mechanism. The encrypted documents may face various security risks. Existing document encryption methods may lack an effective risk assessment mechanism and cannot detect potential security threats in a timely manner, thus increasing the risk of document leakage or tampering. For this reason, the present invention proposes an implementation system for document encryption. Summary of the Invention
[0003] The purpose of the present invention is to solve the problems in the background art and propose an implementation system for document encryption.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] An implementation system for document encryption includes: an anomaly perception and key optimization module, an encryption parameter adaptive adjustment module, a resource allocation and task scheduling module, and an encryption risk assessment and analysis module;
[0006] Anomaly perception and key optimization module: Collect document access data and perform formatting processing. Use a deep neural network model to identify normal or abnormal access behaviors for the processed document access data. When an abnormal access behavior is identified, a key optimization event is triggered. At the same time, use a genetic algorithm to start the key optimization process;
[0007] Encryption parameter adaptive adjustment module: Initialize document encryption parameters, create a document encryption agent, and let the document encryption agent learn and adjust the encryption parameters to find the optimal encryption parameters;
[0008] Resource allocation and task scheduling module: Establish a resource task information matrix and use a fuzzy logic algorithm for resource allocation and task scheduling;
[0009] Encryption Risk Assessment and Analysis Module: After completing the scheduling of encryption tasks and performing encryption, it extracts features from the encrypted document and quantitatively assesses the encryption risk, obtains the encryption risk assessment result, and compares and analyzes the encryption risk assessment result with the preset risk threshold to generate a warning signal.
[0010] Furthermore, the process of the Anomaly Perception and Key Optimization Module collecting document access data, formatting it, and using a deep neural network model to identify normal or abnormal access behaviors for the processed document access data includes:
[0011] Collect document access data during the document access process from various access interfaces and components, including access time, access source IP, and user operation sequence;
[0012] Format all the original document access data and convert it into the format required for input to the deep neural network. Specifically: convert the time data into a timestamp value, convert the IP address into a digital code, and convert the user operation sequence into a digital code sequence in chronological order of timestamps, obtaining three data types: access time, IP address encoding, and operation sequence encoding.
[0013] Use a pre-trained deep neural network model with the processed document access data as the input; the deep neural network model includes an input layer, multiple hidden layers, and an output layer for identifying normal or abnormal access behaviors; among them, the number of nodes in the input layer is determined according to the dimension of the processed document access data, including three data types: access time, IP address encoding, and operation sequence encoding, and there are a total of h data features after encoding, so the number of nodes in the input layer is h; the number of hidden layers is set to f layers, where the number of nodes in each layer is set with a decreasing number of nodes; the number of nodes in the output layer is 2, representing normal and abnormal access behaviors respectively.
[0014] Furthermore, the process of the Anomaly Perception and Key Optimization Module optimizing the key using the genetic algorithm includes:
[0015] For the key optimization event, initialize a group of key populations using the genetic algorithm; start simulating the biological evolution process according to the set fitness function; among them, the fitness function combines two indicators: the complexity and randomness of the key.
[0016] According to the fitness value of each key individual, use the roulette wheel selection method to calculate the proportion of each key individual in the total fitness value of the population as its probability of being selected.
[0017] Through crossover and mutation operations on the selected key individuals, make the key population evolve in a better direction and generate key individuals with higher fitness.
[0018] Further, the process of the encryption parameter adaptive adjustment module initializing the document encryption parameters includes:
[0019] Obtain the initial attributes of the document; among them, the initial attributes of the document include file type and the initial estimated size of the document.
[0020] According to the initial attributes of the document, set a set of initial encryption parameters, including the number of encryption rounds and the key length; among them, the number of encryption rounds is set according to the file type, and the key length is set according to the initial estimated size of the document.
[0021] Further, the process of the encryption parameter adaptive adjustment module creating a document encryption agent and having the document encryption agent learn and adjust the encryption parameters to find the optimal encryption parameters includes:
[0022] Create a document encryption agent. The document encryption agent obtains the CPU resources and memory resources for document encryption by perceiving the environmental information of the document encryption process.
[0023] The document encryption agent encrypts the document according to the currently set encryption parameters, and obtains corresponding feedback information based on the evaluation results of the encryption time and the ciphertext security; among them, the set encryption parameters include the number of encryption rounds and the key length; during the encryption process, record the time spent from the start to the end of the encryption operation to obtain the encryption time; adopt a ciphertext security evaluation method, and the evaluation result is represented by a security level, which is divided into three levels: high, medium, and low. If the ciphertext security reaches the high or medium level, it is considered that the ciphertext security meets the standard; if it is at the low level, it is considered that the ciphertext security does not meet the standard; the feedback information specifically includes positive feedback, mild feedback, and negative feedback.
[0024] Based on the feedback information, the document encryption agent learns in the parameter space and tries different parameter combinations: if the encryption time exceeds the encryption time threshold but the ciphertext security meets the standard, the agent tries to reduce the number of encryption rounds; if the ciphertext security does not meet the standard, increase the key length; through continuous iteration to find the optimal encryption parameter combination, determine the optimal encryption parameters for the current document.
[0025] Further, the process of the resource allocation and task scheduling module establishing a resource task information matrix includes:
[0026] Collect the priority information of all documents to be encrypted, the remaining encryption time, and the current environment information obtained by the document encryption agent; among them, for the priority information of the documents to be encrypted, it is obtained by weighted calculation of the data volume of the documents to be encrypted, the nature of the documents to be encrypted, and the reliability of the source of the documents to be encrypted; for the remaining encryption time, it is calculated based on the encryption progress and the estimated total encryption time of the documents to be encrypted; for the environmental status structure, the current environment information obtained by the document encryption agent, including CPU resources and memory resources, is integrated into the environmental status structure;
[0027] Obtain the priority information list, the remaining encryption time list, and the environmental status structure, and summarize them into a resource task information matrix.
[0028] Furthermore, the process of the resource allocation and task scheduling module using the fuzzy logic algorithm for resource allocation and task scheduling includes:
[0029] Obtain the resource task information matrix as the input of the fuzzy logic algorithm;
[0030] Define the priority fuzzy set, the remaining encryption time fuzzy set, and the available CPU resource fuzzy set based on the resource task information matrix; allocate resources to the encryption tasks according to the defined fuzzy sets;
[0031] Calculate the comprehensive evaluation value of the encryption tasks; sort the encryption tasks in descending order according to the calculated comprehensive evaluation value of the encryption tasks to determine the execution order of the encryption tasks.
[0032] Furthermore, the process of the encryption risk assessment and analysis module extracting features and quantitatively evaluating the encryption risk of the encrypted documents, obtaining the encryption risk assessment result, and comparing and analyzing the encryption risk assessment result with the preset risk threshold to generate a warning signal includes:
[0033] Obtain the encrypted file size feature by getting the number of bytes of the encrypted document;
[0034] Analyze the encryption parameter feature in the file header, that is, the key length identifier;
[0035] Adopt the block encryption method during the encryption process, calculate the statistical information of the encryption block size, and obtain the distribution feature of the encryption blocks;
[0036] Select the decision tree algorithm to construct the encryption risk assessment model, and classify the encrypted documents according to the obtained features: for the normally encrypted documents, mark them as 0, indicating no risk; for the encrypted documents with risks, mark them as 1, indicating the existence of risks;
[0037] During decision tree training, starting from the root node, select an optimal feature as the splitting node;
[0038] Divide the encrypted documents into different subsets according to the selected feature, and continue to repeat the process of selecting the splitting node for each subset until the stopping condition is met, that is, the encrypted documents in the subset all belong to the same category or reach the preset depth of the tree;
[0039] Extract features from the new encrypted document, including the feature of the encrypted file size, the feature of the encryption parameters in the file header, and the distribution feature of the encrypted blocks;
[0040] Input the features extracted from the new encrypted document into the trained encryption risk assessment model; the input feature vector is judged according to the nodes of the decision tree and finally reaches a leaf node; among them, the category corresponding to the leaf node is the preliminary encryption risk assessment result, 0 indicates no risk, and 1 indicates there is risk;
[0041] Set a risk threshold. If the risk quantitative assessment result is higher than the risk threshold, mark the encrypted document as a suspicious document and generate a warning signal.
[0042] Compared with the existing technologies, the advantages of the implementation system for document encryption provided by the present invention are as follows:
[0043] 1. By collecting document access data and performing formatting processing, the present invention uses a deep neural network model to identify normal or abnormal access behaviors of the processed document access data. When an abnormal access behavior is identified, a key optimization event is triggered. At the same time, a genetic algorithm is used to start the key optimization process, which is conducive to the timely discovery of abnormal behaviors and the timely optimization of keys;
[0044] 2. By initializing document encryption parameters, creating a document encryption agent, and using the document encryption agent to learn and adjust the encryption parameters to find the optimal encryption parameters, the present invention can, while ensuring encryption security, improve the encryption efficiency as much as possible, enhance the encryption effect, and reduce the encryption time; by establishing a resource task information matrix and using a fuzzy logic algorithm for resource allocation and task scheduling, it can more accurately evaluate the encryption requirements and resource status of different documents, and thus make more reasonable scheduling decisions to improve resource utilization;
[0045] 3. By extracting features from the encrypted document and performing encryption risk quantitative assessment to obtain the encryption risk assessment result, and comparing and analyzing the encryption risk assessment result with the preset risk threshold to generate a warning signal, it can remind the administrator to take measures in time to deal with potential security threats. Description of the Drawings
[0046] Figure 1A module diagram of an implementation system for document encryption proposed by the present invention. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] Refer to Figure 1 , an implementation system for document encryption, the system includes an anomaly perception and key optimization module, an encryption parameter adaptive adjustment module, a resource allocation and task scheduling module, and an encryption risk assessment and analysis module;
[0049] Anomaly perception and key optimization module: Collect document access data and perform formatting processing. Use a deep neural network model to identify normal or abnormal access behaviors for the processed document access data. When an abnormal access behavior is identified, trigger a key optimization event. At the same time, start the key optimization process using a genetic algorithm;
[0050] Encryption parameter adaptive adjustment module: Initialize document encryption parameters, create a document encryption agent, and let the document encryption agent learn and adjust the encryption parameters to find the best encryption parameters;
[0051] Resource allocation and task scheduling module: Establish a resource task information matrix and use a fuzzy logic algorithm for resource allocation and task scheduling;
[0052] Encryption risk assessment and analysis module: After completing the scheduling of the encryption task and performing the encryption, extract features from the encrypted document and quantitatively evaluate the encryption risk, obtain the encryption risk assessment result, compare and analyze the encryption risk assessment result with a preset risk threshold, and generate a warning signal.
[0053] The steps for the anomaly perception and key optimization module to collect document access data, perform formatting processing, use a deep neural network model to identify normal or abnormal access behaviors for the processed document access data, trigger a key optimization event when an abnormal access behavior is identified, and start the key optimization process using a genetic algorithm include:
[0054] Step 101: For the network access interface in document encryption, deploy a data collection agent to capture network data packets when accessing the document, extract information related to the access time from the network data packets, accurate to the millisecond level, to ensure the accuracy of the time data;
[0055] For the collected access time data, convert it into a timestamp value calculated from a fixed starting time point (such as the encryption startup time); adopt a time encoding algorithm to convert the time difference into an integer value in seconds, which is convenient for deep neural network processing;
[0056] Set up a data collection end at the authentication component for document encryption to obtain the access source IP address. At the same time, monitor the operation logs and collect the user operation sequence, such as the timestamps and operation types corresponding to operations like opening, editing, and saving;
[0057] For the access source IP address, convert it into a digital code according to the division rules of the network number and host number of the IP address: regard the four parts of the IP address as four-digit decimal numbers respectively; combine them into a long decimal number, or adopt a hash coding method to ensure that each IP address has a unique digital representation;
[0058] For the user operation sequence, establish a mapping table from the operation type to the digital code: map the "open" operation to 1, "edit" to 2, and "save" to 3; convert the user operation sequence into a digital code sequence in timestamp order to form an ordered digital vector;
[0059] Step 102: Construct a deep neural network model with an input layer, multiple hidden layers, and an output layer; among them, the number of input layer nodes is determined according to the dimension of the processed document access data, including three data types: access time, IP address encoding, and operation sequence encoding. After encoding, there are a total of h data features, so the number of input layer nodes is h; the number of hidden layers is set to f layers (f ∈ [3, 5]), where the number of nodes in each layer is set with a decreasing number of nodes, that is, the number of nodes in the first hidden layer is m (m > h), the second layer is m / 2, the third layer is m / 4...; the number of output layer nodes is 2, representing two behaviors: normal access and abnormal access respectively;
[0060] Divide the processed document access data into a training set, a validation set, and a test set, and use the training set to train the deep neural network; among them, during the training process, adopt the backpropagation algorithm to calculate the error between the predicted result of the output layer and the actual label (the true label of normal or abnormal access); according to the error, start from the output layer and gradually adjust the weights of the deep neural network backward to the input layer; in each iteration, update the weights according to the set learning rate (such as 0.001) to make the error gradually decrease; through multiple iterations (such as 1000 - 5000 times, determined according to the data volume and model complexity), evaluate the performance of the deep neural network model on the validation set to avoid overfitting; when the accuracy on the validation set reaches the preset validation threshold (such as above 90%), stop training to obtain the trained deep neural network model;
[0061] During the document encryption process, the newly collected and processed document access data is input into the trained deep neural network model in real time. Among them, the model analyzes according to the learned features and outputs the normal or abnormal probability corresponding to each access behavior. If the abnormal probability of the output access behavior exceeds the preset abnormal threshold (such as 0.8), it is determined as an abnormal access behavior and a key optimization event is triggered;
[0062] Step 103: After triggering the key optimization event, determine the format and length range of the key: The key length is set to 128 bits and represented in hexadecimal; Randomly generate a group of initial key populations, and the population size is set to 100 - 500 key individuals; Each key individual randomly generates a hexadecimal digit combination within the set length range;
[0063] Define a fitness function to evaluate the quality of each key individual: The fitness function combines two indicators of key complexity and randomness; For complexity, calculate the number of different characters (hexadecimal digits) in the key and their distribution uniformity; For randomness, use a randomness test algorithm (such as some test methods in the NIST randomness test suite) to evaluate the key; Combine these two indicators according to the preset weights (such as complexity accounting for 0.6 and randomness accounting for 0.4) to form a fitness value;
[0064] According to the fitness value of each key individual, use the roulette wheel selection method to calculate the proportion of each key individual in the total fitness value of the population as its probability of being selected; The higher the fitness value of the key individual, the greater the probability of being selected to participate in subsequent genetic operations;
[0065] Pair the selected key individuals in pairs and perform a crossover operation at a random position of each key individual (for example, between the 32nd and 64th bits in hexadecimal representation) to exchange part of the key content to generate new key individuals; Among them, the crossover probability is set to 0.8 to ensure sufficient diversity;
[0066] For each hexadecimal digit in the newly generated key individuals, perform a mutation operation with a low probability (such as 0.01); Among them, the mutation operation is to randomly replace the digit with other hexadecimal digits, thereby introducing new gene mutations and increasing the diversity of the key population; Through multiple (such as 100 - 500 times) selection, crossover, and mutation operation cycles, the key population continuously evolves to generate key individuals with higher fitness.
[0067] The encryption parameter adaptive adjustment module initializes the document encryption parameters, creates a document encryption agent, and the steps for the document encryption agent to learn and adjust the encryption parameters to find the best encryption parameters include:
[0068] Step 201. Set the encryption rounds based on the file type: Obtain the file type and classify the file types: classify the files into office documents (such as Word, Excel, etc.), image files (JPEG, PNG, etc.), audio files (MP3, WAV, etc.), and video files (MP4, AVI, etc.);
[0069] It should be further noted that for office documents, since their content is text and simple format information, the initial encryption rounds are set to 3 rounds. Among them, the structure of office documents is relatively regular, and fewer encryption rounds can ensure a certain level of security without excessive resource consumption; for image files, their data structure is relatively complex and contains a large amount of information, so the initial encryption rounds are set to 5 rounds. Among them, pixel information in image files requires more rounds of encryption to ensure security; for audio and video files, their data volume is large and continuous, and the initial encryption rounds are set to 4 rounds. Among them, this is a setting after comprehensively considering security and encryption efficiency;
[0070] Set the key length based on the initial estimated size of the document: If the initial estimated size of the document is less than 100KB, considering that the file is small, to avoid excessive resource consumption while ensuring basic security, the key length is initially set to 64 bits; if the initial estimated size of the document is between 100KB and 1MB, the key length is set to 128 bits. Among them, for files in this range, a 128-bit key can provide better security protection; if the initial estimated size of the document exceeds 1MB, since the file is large and may contain more important information, the key length is set to 256 bits to ensure high security;
[0071] Step 202. Create a document encryption agent for interacting and communicating with the document encryption process and sensing the environmental information during the document encryption process; during the document encryption process, the agent obtains CPU resources and memory resources;
[0072] The document encryption agent encrypts the document according to the currently set encryption parameters (encryption rounds and key length); during the encryption process, record the time spent from the start to the end of the encryption operation to obtain the encryption time;
[0073] Determine the encryption time threshold, for example, set it to 30 seconds according to user requirements. If the encryption time exceeds the encryption time threshold, adjust the encryption parameters;
[0074] Adopt a ciphertext security evaluation method, and the evaluation results are represented by security levels, divided into three levels: high, medium, and low. If the ciphertext security reaches a high or medium level, it is considered that the ciphertext security meets the standard; if it is at a low level, it is considered that the ciphertext security does not meet the standard, and adjust the encryption parameters to improve security;
[0075] Give the corresponding feedback to the document encryption agent according to the encryption time and the evaluation result of ciphertext security: If the encryption time does not exceed the encryption time threshold and the ciphertext security meets the standard, give positive feedback indicating that the current encryption parameters are appropriate; if the encryption time exceeds the encryption time threshold but the ciphertext security meets the standard, give mild feedback instructing the document encryption agent to reduce the number of encryption rounds; if the ciphertext security does not meet the standard, give negative feedback instructing the document encryption agent to increase the key length;
[0076] When receiving mild feedback, the document encryption agent attempts to reduce the number of encryption rounds. The number of encryption rounds reduced each time is set to 1 round, and the encryption operation is performed again with the new number of encryption rounds. Then, re-evaluate the encryption time and ciphertext security; continue the loop process until the encryption time does not exceed the encryption time threshold or reaches the minimum allowed number of encryption rounds (the minimum is set to 1 round);
[0077] If receiving negative feedback, the document encryption agent attempts to increase the key length. The amplitude of the key length increased each time is set to 32 bits, and the encryption operation is performed again with the new key length. Then, re-evaluate the encryption time and ciphertext security; repeat the loop process until the ciphertext security meets the standard or reaches the maximum key length (the maximum supports a 512-bit key);
[0078] By continuously adjusting the number of encryption rounds and the key length according to the feedback, the document encryption agent repeatedly tries and optimizes in the parameter space; after iteration, finally determine the optimal combination of encryption parameters (the number of encryption rounds and the key length) with the encryption time within the encryption time threshold and the ciphertext security meeting the standard.
[0079] The steps for the resource allocation and task scheduling module to establish a resource task information matrix and perform resource allocation and task scheduling using the fuzzy logic algorithm include:
[0080] Step 301, set the priority of each document to be encrypted input by the user ; where , 1 is the highest priority and 10 is the lowest priority, is the index of the document to be encrypted;
[0081] After setting the user priority, associate the set user priority with the unique identifier of the document and store it; where, use the document name or document number as the unique identifier of the document;
[0082] It should be further noted that the automatic judgment process of the priority of the document to be encrypted is: Define the data volume of the document to be encrypted , where the data volume of the document to be encrypted is measured in bytes; by analyzing the nature of the document to be encrypted Make a judgment (including financial statements, contracts, etc.), such as checking the extension of the document to be encrypted; set the reliability degree of the source of the document to be encrypted as , where the source of the document to be encrypted is an internal trusted network or an external untrusted network; calculate the priority of all documents to be encrypted using the formula :
[0083] ,
[0084] In the formula, are preset different weight coefficients, and ;
[0085] Sort the priority information of all documents to be encrypted in the order of document identifiers into a priority information list, which is part of the construction of the resource task information matrix;
[0086] Based on the document access data, obtain the document content data corresponding to each document to be encrypted, and split the document content data into fixed-size data blocks for encryption: define the number of encrypted data blocks as , the total number of data blocks as , then the encryption progress of each document to be encrypted is: ;
[0087] Set the encryption speed as , and obtain the estimated total encryption time :
[0088] ,
[0089] In the formula, is the total data volume of the document;
[0090] Then the remaining encryption time is:
[0091] ;
[0092] Associate the remaining encryption time of each document to be encrypted with the document identifier, and sort it into a remaining encryption time list, which is part of the construction of the resource task information matrix;
[0093] Obtain the current environmental information through the document encryption agent, that is, CPU resources and memory resources; sort the obtained CPU resources and memory resources into an environmental status structure, which is part of the construction of the resource task information matrix; among them, for CPU resources, set the available amount of CPU resources as , and the total amount of CPU resources is ; for memory resources, set the physical memory as , and the virtual memory is , the currently used physical memory is , the used virtual memory is , then the total available amount of memory is , and at the same time the memory utilization rate ;
[0094] Integrate the obtained priority information list, remaining encryption time list, and environmental status structure into a resource task information matrix, and use the resource task information matrix as the input of the fuzzy logic algorithm;
[0095] Step 302, for the priority, define high priority , medium priority and low priority three priority fuzzy sets: If , it is determined as high priority ; if , it is determined as medium priority ; if , it is determined as low priority ;
[0096] For the remaining encryption time, define short time , medium time and long time three remaining encryption time fuzzy sets: If , it is determined as short time ; if , it is determined as medium time ; if , it is determined as long time ; where represents that the remaining encryption time is between 0 - 10 seconds, represents that the remaining encryption time is between 10 - 60 seconds, represents that the remaining encryption time exceeds 60 seconds;
[0097] For the available amount of CPU resources, define high - volume CPU resources , medium CPU resources and low - volume CPU resources three available - amount - of - CPU - resources fuzzy sets: If , it is determined as high - volume CPU resources ; if , it is determined as medium CPU resources ; if , it is determined as low - volume CPU resources ;
[0098] When the document to be encrypted is of high priority and the remaining encryption time is and it is high-volume CPU resources when, then the CPU resource allocation is , and at the same time according to the memory usage rate allocate memory , where is the first resource allocation adjustment factor, and ; is the allocated CPU resources; is the allocated memory resources; is the total memory;
[0099] When the document to be encrypted is of low priority and the remaining encryption time is and it is low-volume CPU resources when, then the CPU resource allocation is , and at the same time according to the memory usage rate allocate memory , where is the second resource allocation adjustment factor, and ;
[0100] For other intermediate situations, then allocate CPU cores to this encryption task, and the memory allocation ;
[0101] Use the formula to calculate the comprehensive evaluation value of the encryption task : ,
[0102] In the formula, are preset different weight coefficients;
[0103] Arrange the encryption tasks in descending order according to the calculated comprehensive evaluation value of the encryption tasks to determine the execution order of the encryption tasks;
[0104] It should be further noted that during the execution of the encryption task, the resource usage and the progress of the encryption task are monitored in real time, and the resource allocation and execution order of the task are dynamically adjusted according to the new resource task information matrix. For example, if a low-priority task being executed has more available resources due to some reason (such as other high-priority tasks finishing ahead of schedule and releasing resources), and its remaining encryption time is relatively long, then appropriately increase the resources allocated to it to speed up its encryption speed.
[0105] After the encryption risk assessment and analysis module completes the scheduling of the encryption task and executes the encryption, it extracts the features of the encrypted document and quantitatively evaluates the encryption risk to obtain the encryption risk assessment result. The steps of comparing the encryption risk assessment result with the preset risk threshold and generating a warning signal include:
[0106] Step 401: Obtain the encrypted file size feature by getting the number of bytes of the encrypted document;
[0107] It should be further noted that different types of documents may have a relatively stable size range after encryption. For example, after a plain text file is encrypted, its size may increase by a certain proportion according to the characteristics of the encryption algorithm. If the file size deviates from the expected range after normal encryption, it may imply that something is wrong with the encryption process, such as the encryption algorithm being tampered with or the file being maliciously modified before encryption;
[0108] Step 402: Analyze the encryption parameter feature in the file header, that is, the key length identifier;
[0109] It should be further noted that for the analysis of the key length identifier, determine the expected key length range, find the position in the file header where the key length identifier is stored, read the identifier data, and convert it into a recognizable form representing the key length. Compare the read key length identifier with the expected key length. If the two are equal, it means that the key length identifier matches the actual key length and the encryption process is normal. If they are not equal, it may imply that the encryption process has been interfered with or the file has been tampered with, thus affecting the security of the document;
[0110] Step 403: In the encryption process, use the block encryption method to calculate the statistical information of the encryption block size (i.e., calculate the average and standard deviation of the block size) to obtain the distribution characteristics of the encryption blocks: When analyzing the average value of the encryption block size, combine the standard deviation to obtain the overall distribution characteristics of the encryption blocks. The average value provides the central tendency of the encryption block size, while the standard deviation reveals the degree of dispersion of the encryption blocks. If a low standard deviation is combined with an encryption block size distribution close to the average value, it indicates that the encryption blocks are evenly divided. On the contrary, if the combination of the average value and the standard deviation shows a large deviation, it may indicate uneven division or potential tampering during the encryption process;
[0111] Step 404: Select the decision tree algorithm to build an encryption risk assessment model, and classify the encrypted documents according to the obtained features: For normally encrypted documents, mark them as 0, indicating no risk; For encrypted documents with risks, mark them as 1 according to the specific type of risk (such as key leakage, encryption algorithm being cracked, file being maliciously tampered with, etc.), indicating that there is a risk;
[0112] Step 405: When training the decision tree, starting from the root node, select an optimal feature as the splitting node;
[0113] It should be further noted that the optimal feature can distinguish different categories of encrypted documents to the greatest extent. For example, it is selected according to the information gain or Gini index of the feature for the encrypted documents; the information gain and Gini index are metrics to measure the influence degree of a feature on the classification result. When constructing the decision tree, starting from the root node, select the feature with the maximum information gain or the minimum Gini index as the splitting node, indicating that this feature can most effectively separate different categories of encrypted documents, enabling the decision tree to grow faster in the direction of correct classification. For example, if a certain feature can distinguish most of the normal encrypted documents from the risky encrypted documents, then select this feature as the splitting node;
[0114] Step 406: Divide the encrypted documents into different subsets according to the selected feature, and continue to repeat the process of selecting the splitting node for each subset until the stopping condition is met, that is, the encrypted documents in the subset all belong to the same category or reach the preset depth of the tree;
[0115] It should be further noted that once the splitting node is selected, the encrypted documents are divided into different subsets according to different values of this feature, and then continue to find the optimal splitting node in each subset, repeating this process; the stopping condition is to prevent the decision tree from overgrowing. If the encrypted documents in the subset all belong to the same category, it means that accurate classification can already be achieved and there is no need to continue splitting. The preset tree depth is an artificially set limit to avoid overfitting caused by the decision tree being too complex, that is, the situation where the model performs well on the training data but poorly on the new encrypted document data;
[0116] Step 407: Extract features from the new encrypted documents, including the feature of the file size after encryption, the encryption parameter feature in the file header, and the distribution feature of the encrypted blocks;
[0117] Step 408: Input the features extracted from the new encrypted documents into the already trained encryption risk assessment model; the input feature vector is judged according to the nodes of the decision tree and finally reaches a leaf node; among them, the category (0 or 1) corresponding to the leaf node is the preliminary encryption risk assessment result, where 0 indicates no risk and 1 indicates there is risk;
[0118] It should be further noted that for the encryption risk assessment model constructed by the decision tree, according to the input feature vector, it goes down along the structure of the tree, and each node is a judgment condition about the feature;
[0119] Step 409: Set a risk threshold. If the result of the risk quantification assessment is higher than the risk threshold, mark the encrypted document as a suspicious document and generate a warning signal. At the same time, take corresponding security measures, specifically including re-encrypting the suspicious encrypted document with a new security key or checking the encryption key to see if there is a possibility of key leakage.
[0120] It should be further noted that the setting of the risk threshold is determined according to specific security requirements and the degree of acceptance of risks. In some scenarios that are very sensitive to security, even a small risk possibility cannot be accepted, so a lower risk threshold will be set. When the risk assessment result is higher than the set risk threshold, it indicates that there may be a security risk in the encrypted document, and relevant personnel need to be notified in a timely manner for further inspection and handling.
[0121] In the embodiments of the present invention, by collecting and formatting document access data and using a deep neural network model to identify normal and abnormal access behaviors, potential malicious access or illegal operations can be discovered in a timely manner, effectively preventing security risks. When an abnormal access behavior is identified, a key optimization event is triggered, and a genetic algorithm is used to optimize the key, enhancing the complexity and randomness of the key, greatly increasing the difficulty of cracking the key, and improving the security of encryption. Initial encryption parameters are set according to the initial attributes of the document (such as file type, initial estimated size), and the document encryption agent learns and adjusts the encryption parameters to find the best encryption scheme, improving the flexibility and adaptability of encryption. According to the priority information of the document to be encrypted, the remaining encryption time, and the environmental status structure, a resource task information matrix is established, and resource allocation and task scheduling are performed to ensure that high-priority documents can be processed first, improving the timeliness of encryption. Using a fuzzy logic algorithm for resource allocation and task scheduling can more accurately evaluate the encryption requirements and resource status of different documents, and thus make more reasonable scheduling decisions. By extracting features from the encrypted document and performing a quantitative assessment of the encryption risk, the encryption risk assessment result is obtained, which helps to discover potential risks and problems in the encryption process in a timely manner. By comparing and analyzing the encryption risk assessment result with a preset risk threshold and generating a warning signal, it can remind the administrator to take measures in a timely manner to deal with potential security threats. By selecting a decision tree algorithm to construct an encryption risk assessment model and classifying the encrypted documents according to the obtained features, decision-making support is provided for the administrator to help them better understand the security status and risk level of the encrypted documents. In summary, the embodiments of the present invention solve the problem of poor current document encryption effect. In actual situations, more data and context information may be required to make specific decisions and optimization plans.
[0122] In addition, the formulas involved above are all calculated by removing the dimension and taking their numerical values. They are obtained by collecting a large amount of data for software simulation to get a formula that is closest to the actual situation. The proportionality coefficient in the formula and each preset threshold value in the analysis process are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data. The magnitude of the proportionality coefficient is a specific numerical value obtained by quantifying each parameter, which is convenient for subsequent comparison. Regarding the magnitude of the proportionality coefficient, it depends on the amount of sample data and the processing coefficients initially set by those skilled in the art for each group of sample data, as long as the proportional relationship between the parameters and the quantified values is not affected.
[0123] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically based on the method embodiments, they are described relatively simply, and reference can be made to the relevant parts of the method embodiments for the relevant content.
[0124] For the convenience of description, when describing the above device, it is divided into various units according to functions for separate description. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.
[0125] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more flows or multiple flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the one or more of the processes and / or blocks.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the one or more of the processes and / or blocks.
[0129] Second: In the drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments of the present disclosure are involved. For other structures, reference may be made to the general design. Without conflict, the same and different embodiments of the present invention may be combined with each other;
[0130] Finally: The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. An implementation system for document encryption, characterized in that: It includes an anomaly perception and key optimization module, an encryption parameter adaptive adjustment module, a resource allocation and task scheduling module, and an encryption risk assessment and analysis module; Anomaly perception and key optimization module: Collect document access data and perform formatting processing. Use a deep neural network model to identify normal or abnormal access behaviors for the processed document access data. When an abnormal access behavior is identified, a key optimization event is triggered. At the same time, use the genetic algorithm to start the key optimization process; Encryption parameter adaptive adjustment module: Initialize the document encryption parameters, create a document encryption agent, and let the document encryption agent learn and adjust the encryption parameters to find the optimal encryption parameters; Resource allocation and task scheduling module: Establish a resource task information matrix and use the fuzzy logic algorithm for resource allocation and task scheduling; Encryption risk assessment and analysis module: After completing the scheduling of the encryption task and executing the encryption, perform feature extraction and encryption risk quantitative assessment on the encrypted document, obtain the encryption risk assessment result, compare and analyze the encryption risk assessment result with the preset risk threshold, and generate a warning signal; Among them, the process of the encryption parameter adaptive adjustment module creating a document encryption agent and letting the document encryption agent learn and adjust the encryption parameters to find the optimal encryption parameters includes: Create a document encryption agent. The document encryption agent obtains the CPU resources and memory resources for document encryption by perceiving the environmental information of the document encryption process; The document encryption agent encrypts the document according to the currently set encryption parameters and obtains corresponding feedback information based on the evaluation results of the encryption time and the security of the ciphertext. Among them, the set encryption parameters include the number of encryption rounds and the key length; during the encryption process, record the time spent from the start to the end of the encryption operation to obtain the encryption time; use the ciphertext security assessment method, and the evaluation result is represented by a security level, which is divided into three levels: high, medium, and low. If the ciphertext security reaches a high or medium level, it is considered that the ciphertext security meets the standard; if it is a low level, it is considered that the ciphertext security does not meet the standard; the feedback information specifically includes positive feedback, mild feedback, and negative feedback; Based on the feedback information, the document encryption agent learns in the parameter space and tries different parameter combinations: if the encryption time exceeds the encryption time threshold but the ciphertext security meets the standard, the agent tries to reduce the number of encryption rounds; if the ciphertext security does not meet the standard, increase the key length; through the process of continuously iterating to find the optimal encryption parameter combination, determine the optimal encryption parameters for the current document.
2. The implementation system for document encryption according to claim 1, wherein: The process of the anomaly perception and key optimization module collecting document access data and performing formatting processing, and using a deep neural network model to identify normal or abnormal access behaviors for the processed document access data includes: Collect document access data during the document access process from various access interfaces and components, including access time, access source IP, and user operation sequence; Format all the original document access data and convert it into the format required for input to a deep neural network. Specifically: convert time data into timestamp values, convert IP addresses into digital encodings, and convert the user operation sequence into a sequence of digital codes in timestamp order, obtaining three data types: access time, IP address encoding, and operation sequence encoding respectively. Use a pre-trained deep neural network model with the processed document access data as the input. The deep neural network model includes an input layer, multiple hidden layers, and an output layer for identifying normal or abnormal access behaviors. Among them, the number of input layer nodes is determined according to the dimension of the processed document access data, including three data types: access time, IP address encoding, and operation sequence encoding. After encoding, there are a total of h data features, so the number of input layer nodes is h. The number of hidden layers is set to f layers, and the number of nodes in each layer is set to decrease. The number of output layer nodes is 2, representing normal and abnormal access behaviors respectively.
3. The implementation system for document encryption according to claim 1, characterized in that: The process of the anomaly perception and key optimization module using the genetic algorithm to optimize the key includes: For the key optimization event, use the genetic algorithm to initialize a group of key populations and start simulating the biological evolution process according to the set fitness function. Among them, the fitness function combines two indicators: the complexity and randomness of the key. According to the fitness value of each key individual, use the roulette wheel selection method to calculate the proportion of each key individual in the total fitness value of the population as the probability of its being selected. Perform crossover and mutation operations on the selected key individuals to make the key population evolve in a more optimal direction and generate key individuals with higher fitness.
4. The implementation system for document encryption according to claim 1, wherein: The process of the encryption parameter adaptive adjustment module initializing the document encryption parameters includes: Obtain the initial attributes of the document. Among them, the initial attributes of the document include file type and the initial estimated size of the document. According to the initial attributes of the document, set a set of initial encryption parameters, including the number of encryption rounds and the key length. Among them, the number of encryption rounds is set according to the file type, and the key length is set according to the initial estimated size of the document.
5. The implementation system for document encryption according to claim 1, wherein: The process of the resource allocation and task scheduling module establishing the resource task information matrix includes: Collect the priority information, remaining encryption time of all documents to be encrypted, and the current environmental information obtained by the document encryption agent. Among them, for the priority information of the document to be encrypted, it is obtained by weighted calculation of the data volume, nature, and source reliability of the document to be encrypted. For the remaining encryption time, it is calculated according to the encryption progress and the estimated total encryption time of the document to be encrypted. For the environmental status structure, the current environmental information obtained by the document encryption agent, including CPU resources and memory resources, is integrated into the environmental status structure. Obtain the priority information list, remaining encryption time list, and environmental status structure and summarize them into a resource task information matrix.
6. The implementation system for document encryption according to claim 5, wherein: The process of the resource allocation and task scheduling module using the fuzzy logic algorithm for resource allocation and task scheduling includes: Obtain the resource task information matrix as the input of the fuzzy logic algorithm. Define the priority fuzzy set, the remaining encryption time fuzzy set, and the CPU resource availability fuzzy set based on the resource task information matrix; allocate resources to the encryption tasks according to the defined fuzzy sets; Calculate the comprehensive evaluation value of the encryption tasks; sort the encryption tasks in descending order according to the calculated comprehensive evaluation value of the encryption tasks to determine the execution order of the encryption tasks.
7. The implementation system for document encryption according to claim 1, wherein: The process of the encryption risk assessment and analysis module extracting features and quantitatively evaluating the encryption risk of the encrypted document to obtain the encryption risk assessment result, and comparing and analyzing the encryption risk assessment result with a preset risk threshold to generate a warning signal includes: Obtain the encrypted file size feature by getting the number of bytes of the encrypted document; Analyze the encryption parameter feature in the file header, that is, the key length identifier; Adopt a block encryption method during the encryption process, calculate the statistical information of the encryption block size, and obtain the distribution feature of the encryption blocks; Select the decision tree algorithm to construct an encryption risk assessment model, and classify the encrypted documents according to the obtained features: for the normally encrypted documents, mark them as 0, indicating no risk; for the encrypted documents with risks, mark them as 1, indicating there is a risk; During the decision tree training, start from the root node and select an optimal feature as the splitting node; Divide the encrypted documents into different subsets according to the selected feature, and continue to repeat the process of selecting the splitting node for each subset until the stopping condition is met, that is, the encrypted documents in the subset all belong to the same category or reach the preset depth of the tree; Extract features from the new encrypted document, including the encrypted file size feature, the encryption parameter feature in the file header, and the distribution feature of the encryption blocks; Input the features extracted from the newly encrypted document into the already trained encryption risk assessment model; the input feature vector is judged according to the nodes of the decision tree and finally reaches a leaf node; among them, the category corresponding to the leaf node is the preliminary encryption risk assessment result, 0 indicates no risk, and 1 indicates there is a risk; Set a risk threshold. If the risk quantitative assessment result is higher than the risk threshold, mark the encrypted document as a suspicious document and generate a warning signal.
Citation Information
Patent Citations
Electronic document authority management method and system based on localized OFD
CN117786715A
Archive system and method based on artificial intelligence
CN119066201A