An electronic information security system for archives
By designing an electronic archive information security system, using time series data analysis and user operation mode prediction, potential risk behavior is identified, and through encryption, multi-factor authentication and dynamic resource scheduling, the problem of inability to effectively identify potential risk behaviors in the existing technology is solved, achieving more efficient information security protection.
Patent Information
- Application Number
- CN202411975222.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-12-31
AI Technical Summary
When existing technologies face changing security threats, they fail to provide sufficient forward-looking protection and fail to effectively identify potential risky behaviors, resulting in the exposure of security vulnerabilities.
An archive electronic information security system is designed, including an archive operation behavior prediction module, an archive data encryption and secure storage module, an archive security protection policy execution module, and an archive resource scheduling and security monitoring module. The system identifies potential risk behaviors through time series data analysis and prediction of user operation modes, and improves information security through encryption, multi-factor authentication and dynamic resource scheduling.
By identifying potential risky behaviors in advance, preventing illegal access and abnormal operations, improving information security, reducing losses caused by human error or malicious behavior, and ensuring the integrity and confidentiality of data. The system can flexibly adjust storage space and bandwidth resources according to the risk level of actual operation, and improve the system's response speed and resource allocation capabilities.
Smart Images

Figure CN119377998B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and in particular to an information security system for electronic archives. Background Art
[0002] The field of information security technology involves protecting the confidentiality, integrity, and availability of information, preventing information from being accessed, tampered with, damaged, or leaked without authorization. This field covers various technical methods such as encryption technology, identity authentication, access control, data integrity verification, network security protection, and security auditing, and is widely applied in links such as data storage, transmission, and processing to ensure that information systems have high security and reliability in a complex and changing threat environment.
[0003] The existing technology mainly relies on encryption technology, access control, and security auditing to ensure the security of archives. However, when facing continuously changing security threats, it fails to provide sufficient forward-looking protection. The current methods lack real-time prediction and risk assessment of user operation behaviors and cannot effectively identify potential risk behaviors. Even through hierarchical access control and data encryption, there are still missed detections of operation deviations or abnormal behaviors, resulting in the exposure of security vulnerabilities. For example, file management cannot promptly identify and respond to minor deviations in illegal access or operations, thus generating unnecessary security risks. The existing technology also has deficiencies in dynamic resource scheduling and real-time monitoring, and cannot flexibly adjust storage space and bandwidth resources according to the risk level of actual operations, resulting in limitations in the system's response speed and resource allocation ability in emergency situations. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose an information security system for electronic archives.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: An information security system for electronic archives includes:
[0006] An archive operation behavior prediction module extracts time series features of archive operations based on the operation data of archive users, predicts the next operation of the user according to the original operation mode of the user, obtains the operation behavior prediction result, identifies potential risk behaviors, and generates an archive access security risk assessment result;
[0007] An archive data encryption and secure storage module divides the data into blocks for storage according to the archive access security risk assessment result and in combination with the archive data content, divides it according to archive sensitivity and access frequency to generate a set of block data, encrypts the archive data block by block, and stores the encrypted data blocks through a B-tree index structure to generate an encrypted archive data block set;
[0008] The file security protection policy execution module identifies risk operation behaviors based on the encrypted file data block set, generates file security protection requirements, executes multi-factor authentication security measures by adjusting access permissions, monitors the deviation between operation behaviors and predictions, and generates the execution result of the security protection policy;
[0009] Based on the execution result of the security protection policy, the file resource scheduling and security monitoring module dynamically schedules file resources in real time, preferentially allocates resources to risk operations, optimizes the resource allocation of access permissions, storage space, and bandwidth by monitoring user operations in real time, and obtains resource usage and security protection records.
[0010] As a further solution of the present invention, the specific steps for obtaining the operation behavior prediction result are as follows:
[0011] Based on the operation data of file users, extract the operation data characteristics of file users, including access time, operation type, and access object, analyze the time series feature vectors of each data characteristic, and generate a time series matrix through the sequence trend of operation time;
[0012] Analyze the feature vectors in the time series matrix, extract the change trend and correlation of operation behavior characteristics, and perform operations on the time interval distribution and feature change rate. Using the formula:
[0013] ;
[0014] Calculate the amplitude of the feature change trend;
[0015] Among them, represents the amplitude of the feature change trend, is the eigenvalue of the th feature in the time series, is the mean value of the eigenvalues, is the feature weight parameter, represents the total number of features;
[0016] Call the amplitude of the feature change trend, compare it with the user's original operation mode, predict the user's next operation behavior through the feature change trend, perform prediction probability analysis on each operation type, and perform normalization processing to obtain the operation behavior prediction result.
[0017] As a further solution of the present invention, the specific steps for obtaining the file access security risk assessment result are as follows:
[0018] Based on the operation behavior prediction result, extract the original data of each user's access to files, perform access frequency statistics and classify access types, analyze the distribution of access time periods, and perform screening according to permission configuration and access behavior characteristics to generate access behavior statistics results;
[0019] Based on the statistical results of the access behavior, combined with the frequency, access time, and user permission scope of the access behavior, identify abnormal behavior patterns, mark the access behaviors with high frequency and inconsistent permissions, and compare them with the normal mode to obtain the characteristics of abnormal access behaviors;
[0020] Based on the characteristics of the abnormal access behaviors, combined with the access path information, operation permissions, and behavior patterns, conduct an evaluation to determine the risk level of each type of abnormal behavior, divide the different risk levels according to the risk criteria, and generate the evaluation results of the security risks of the file access.
[0021] As a further solution of the present invention, the steps for obtaining the block data set are specifically as follows:
[0022] Analyze the evaluation results of the security risks of the file access, extract the sensitivity level and access frequency characteristics of each file in the file access records, convert the sensitivity data into numerical weights, and generate a file sensitivity and access frequency characteristic matrix through normalizing the number of accesses within a time period;
[0023] Call the file sensitivity and access frequency characteristic matrix, segment the data according to the differences in sensitivity and access frequency, adjust the weight parameters in combination with the block size limit and data continuity requirements, and use the formula:
[0024] ;
[0025] Calculate the segmentation standard value of the data block to obtain the characteristics of the file data block;
[0026] Wherein, represents the segmentation standard value of the data block, is the sensitivity weight, is the access frequency characteristic value, is the average sensitivity of the data block, is the current file sensitivity value, is the content continuity parameter of the data block;
[0027] Call the characteristics of the file data block, determine the range of the data block to which each file belongs by comparing the segmentation standard with the sensitivity and frequency values of the files in the characteristic matrix, and generate a block data set.
[0028] As a further solution of the present invention, the steps for obtaining the encrypted file data block set are specifically as follows:
[0029] Based on the block data set, split the data according to the preset data block size, encrypt each data block, and encrypt the content of each data block using the corresponding encryption technology to generate encrypted data blocks;
[0030] Based on the encrypted data blocks, sort them according to the unique identifier of each data block, and sequentially insert the encrypted data blocks into the B-tree structure. During the insertion process, sort and index each node to generate a B-tree index structure;
[0031] Based on the B-tree index structure, associate the encrypted data blocks with the index entries, determine the storage location of each data block according to the location information of the data block, and store the location of each encrypted data block in the B-tree node through an index linked list to generate an encrypted archive data block set.
[0032] As a further solution of the present invention, the step of obtaining the file security protection requirements is specifically as follows:
[0033] Extract the access behavior logs from the encrypted archive data block set, use the access timestamps, user identifiers, and operation type information in the logs to screen for illegal access and abnormal operations in chronological order, and generate an initial set of risk operations;
[0034] Analyze the access behavior characteristics of each record in the initial set of risk operations, including operation frequency, access path length, and operation type weight, and use the formula:
[0035] ;
[0036] Calculate the access risk characteristic value of each record to obtain an updated risk operation evaluation set;
[0037] Among them, represents the access risk characteristic value, represents the operation frequency, represents the average value of the operation frequency, represents the standard deviation of the operation frequency, represents the access path length, represents the maximum access path length, represents the operation type weight, are the adjustment coefficients for frequency, path, and weight respectively;
[0038] Perform behavior pattern clustering on the records in the updated risk operation evaluation set, group the records with the same type of risk characteristic values into one group, and judge whether there are illegal access and abnormal operations according to the distribution of the intra-group risk values in the clustering result to generate file security protection requirements.
[0039] As a further solution of the present invention, the step of obtaining the execution result of the security protection strategy is specifically as follows:
[0040] Based on the above file security protection requirements, analyze the adjustment conditions of access permissions, call the access permission parameter set and the user hierarchical permission parameter set, determine the permissions to be adjusted, and generate a set of users whose permissions need to be adjusted;
[0041] Based on the set of users whose permissions need to be adjusted, combine the user multi-factor authentication parameter set and the access permission parameter set, and through the formula:
[0042] ;
[0043] Calculate the corrected user access permission value and generate an adjusted access permission set;
[0044] Wherein, is the corrected user access permission value, is the current permission level, is the preset permission reference value, is the multi-factor authentication weight parameter;
[0045] Call the adjusted access permission set, combine the user behavior monitoring parameter set and the deviation prediction model parameter set, compare the user operation deviation with the predicted deviation, and generate the execution result of the security protection strategy.
[0046] As a further solution of the present invention, the step of obtaining the resource usage and security protection record is specifically as follows:
[0047] Based on the execution result of the security protection strategy, according to the user operation logs and resource access data recorded in real-time monitoring, screen the user behavior parameters associated with risk operations, classify the user behaviors and mark the operation types of sensitive file access, and combine the timestamps, user identity identifiers and resource call types marked in the logs to obtain a set of user risk behavior parameters;
[0048] Based on the behavior frequency and sensitivity level in the set of user risk behavior parameters, use the formula:
[0049] ;
[0050] Calculate the priority score of the user operation, evaluate the impact of the user risk operation and quantify the priority to obtain the user risk operation score result;
[0051] Wherein, represents the priority score of the user operation, represents the behavior frequency, represents the sensitivity level, represents the resource call type, represents the operation time interval, is the weight parameter;
[0052] Combined with the user risk operation scoring results, dynamically optimize the resource allocation strategy, and preferentially allocate resources such as bandwidth, storage space, and access permissions to high-scoring risk operations, and obtain records of resource usage and security protection.
[0053] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0054] In the present invention, through the analysis of time-series data and the prediction of user operation patterns, potential risk behaviors can be identified in advance, preventing the occurrence of illegal access and abnormal operations, thereby enhancing the security of information. This ability to identify potential threats in advance can effectively reduce the losses caused by human errors or malicious behaviors, ensure the integrity and confidentiality of data, store the archival data in blocks, and encrypt them according to sensitivity and access frequency, which not only improves the security of storage but also more precisely meets the security requirements at different levels. By dynamically adjusting access permissions and implementing multi-factor authentication, the flexibility and accuracy of protection measures are enhanced, effectively reducing the occurrence of security vulnerabilities. Real-time monitoring and resource scheduling optimization ensure that during the operation process, the access to sensitive archives is preferentially protected, avoiding waste of security resources and unnecessary risk exposure. The measures work together synergistically to overall enhance the security protection ability of archival management in a complex threat environment. Description of the Drawings
[0055] Figure 1 is the system flow chart of the present invention;
[0056] Figure 2 is the flow chart of the operation behavior prediction result in the present invention;
[0057] Figure 3 is the flow chart of the archival access security risk assessment result in the present invention;
[0058] Figure 4 is the flow chart of the block data set in the present invention;
[0059] Figure 5 is the flow chart of the encrypted archival data block set in the present invention;
[0060] Figure 6 is the flow chart of the archival security protection requirements in the present invention;
[0061] Figure 7 is the flow chart of the security protection strategy execution result in the present invention;
[0062] Figure 8 is the flow chart of the resource usage and security protection record in the present invention. Detailed Embodiments
[0063] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0064] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.
[0065] Please refer to Figure 1 , an electronic information security system for archives includes:
[0066] The file operation behavior prediction module performs time series data analysis based on the operation data of file users, including access time, operation type, and access object, extracts the time series characteristics of file operations, predicts the next operation of the user according to the original operation mode of the user, obtains the operation behavior prediction result, identifies potential risk behaviors, and generates a file access security risk assessment result;
[0067] The file data encryption and secure storage module divides the data into blocks for storage according to the file access security risk assessment result and in combination with the file data content. Each data block is divided according to file sensitivity and access frequency to generate a set of block data. The file data is encrypted block by block, and the encrypted data blocks are stored through a B-tree index structure to generate an encrypted file data block set;
[0068] The file security protection policy execution module identifies risk operation behaviors, including illegal access and abnormal operations, based on the encrypted file data block set, generates file security protection requirements, executes multi-factor identity authentication security measures by adjusting access permissions, triggers corresponding protection policies, monitors the deviation between the operation behavior and the prediction, and generates a security protection policy execution result;
[0069] The file resource scheduling and security monitoring module dynamically schedules file resources in real time based on the security protection policy execution result, preferentially allocates resources to risk operations, including access to sensitive files, and optimizes the resource allocation of access permissions, storage space, and bandwidth by monitoring user operations in real time to obtain resource usage and security protection records.
[0070] The operation behavior prediction results include the predicted operation type, predicted access time, and predicted access object. The file access security risk assessment results include risk level assessment, potential risk behavior identification, and user behavior deviation assessment. The chunked data set includes sensitivity category chunks, frequency category chunks, and data chunk indexes. The encrypted file data block set includes encrypted data blocks, index structures, and metadata information. The file security protection requirements include anti-illegal access requirements, abnormal operation monitoring requirements, and identity authentication enhancement requirements. The security protection policy execution results include permission adjustment records, verification measure implementation details, and monitoring deviation reports. The resource usage and security protection records include resource allocation details, protection activity records, and operation monitoring logs.
[0071] Please refer to Figure 2 , and the specific steps for obtaining the operation behavior prediction results are as follows:
[0072] Based on the operation data of the file user, extract the operation data characteristics of the file user, including access time, operation type, and access object, analyze the time series feature vectors of each data characteristic, and generate a time series matrix through the sequence trend of the operation time;
[0073] Calculate the interval between each time point and the subsequent time point as a time difference array, use the time difference as the basic index for evaluating the operation density, establish a time characteristic description by statistically calculating the mean and standard deviation of the time intervals, analyze the operation type, extract the type classification tags of each user operation, count the occurrence frequencies of various types of operations, use relative frequency and absolute frequency as the characteristic values of the operation type, extract the access object information, including key attributes such as the ID, category, and size of the access object, form an access object feature matrix by quantifying the attributes, integrate the above three characteristics into a time series matrix, and through normalization processing, make the data characteristics of each dimension tend to be consistent, reduce the calculation deviation caused by uneven dimensions, and form a data basis that can accurately reflect the user operation characteristics to provide support for subsequent analysis and prediction.
[0074] Analyze the feature vectors in the time series matrix, extract the change trends and correlations of the operation behavior characteristics, and through operations on the time interval distribution and feature change rate, use the formula:
[0075] ;
[0076] Calculate the feature change trend amplitude;
[0077] Among them, represents the feature change trend amplitude, is the eigenvalue of the th feature in the time series, is the mean value of the eigenvalues, is the feature weight parameter, Represents the total number of features;
[0078] The advantage of the formula is that by introducing a weight parameter , it is possible to adjust the weight according to the variation law of operation characteristics within a specific time period, improve the capture accuracy and calculation fineness of the change trend, and thus better meet the analysis requirements of various operation characteristic patterns;
[0079] Extract the time node feature values from the time series matrix , calculate the mean value of each feature value through statistical calculation , calculate the weight parameter according to the importance distribution rule of time series nodes , for example, assume that the time series contains 8 nodes, and the feature values are respectively , , , , , , , , the mean value , the weight parameter , substitute the values into the formula:
[0080] ;
[0081] ;
[0082] ;
[0083] This value indicates that the change amplitude of the operation behavior is small, providing a basis for subsequent prediction. The specific value comes from the construction process of the time series matrix, where the setting of the weight parameter is based on the operation frequency weight function of the time period, and the higher the frequency, the greater the weight, which is obtained through the analysis of the frequency distribution.
[0084] Call the amplitude of the feature change trend, compare it with the user's original operation mode, predict the user's next operation behavior through the feature change trend, conduct a prediction probability analysis for each operation type, and perform normalization processing to obtain the operation behavior prediction result;
[0085] First, filter the corresponding eigenvalue sequence in the time series matrix from the user's historical operation records. Group the eigenvalues by time period and statistically analyze the main feature distribution rules in each group. Extract the change eigenvalues of the operation type, and generate a benchmark template according to the fluctuation pattern of the time series. In the benchmark template, the distribution range and operation frequency of the eigenvalues become important references. By comparing the change amplitude of the current operation behavior characteristics with the change trend of the corresponding characteristics in the benchmark template, determine whether there is a significant deviation. Calculate the probability weights for each operation type respectively, combine the operation frequency, the change amplitude of the characteristics, and the characteristic weights, and obtain the predicted probability distribution of each operation type through normalization processing. Select the operation type with the highest probability as the prediction result of the operation behavior, and further apply this result to the generation of specific operation suggestions to provide the final result of user operation behavior prediction.
[0086] Please refer to Figure 3 , the steps to obtain the assessment result of the file access security risk are specifically as follows:
[0087] Based on the prediction result of the operation behavior, extract the original data of each user's file access, conduct access frequency statistics and classify the access types, analyze the distribution of the access time period, and perform screening according to the permission configuration and access behavior characteristics to generate the access behavior statistics result;
[0088] First, conduct a detailed record analysis of the user's access behavior. Based on the user's login time, operation path, and access purpose, map each operation of the user into different behavior sequences. By gradually decomposing these sequences, distinguish various access types such as browsing, editing, and downloading, and record the operation frequencies corresponding to each type. After the statistics are completed, combined with the time dimension of the data, further analyze the access distribution of the user at different times of the day. By constructing time intervals and accumulating the access frequencies within each interval, form a description of the distribution rule of the high-frequency access time periods. Perform screening according to the permission configuration and access behavior characteristics. The screening process is based on the judgment of the rationality of the access behavior. For example, verify whether the access type matches the user's permissions and whether the access duration is abnormal. Filter out the records of accessing multiple sensitive files frequently in a short period of time, and at the same time retain the behaviors within the expected permission range to generate the access behavior statistics result.
[0089] Based on the access behavior statistics result, combine the frequency, access time, and user permission range of the access behavior to identify abnormal behavior patterns. Mark the access behaviors with high frequency and inconsistent permissions, and compare them with the normal pattern to obtain the characteristics of abnormal access behaviors;
[0090] By gradually dividing the distribution interval of access frequency, the user's behavior data is refined into multiple feature sets. By comparing the statistical patterns of normal access behaviors, cross-verification is performed on high-frequency access users and their permission scopes. For behaviors where the operation frequency within the permission scope is significantly higher than the historical statistical mean, segmented classification is performed in combination with the time characteristics of the operation behaviors to identify potential abnormal behaviors. At the same time, each access record with inconsistent permissions is inspected item by item, and the access behaviors that do not conform to the user role permissions are compared with the key features of the normal mode from multiple angles, including the similarity of access paths and the deviation of access time periods, to obtain the characteristics of abnormal access behaviors, and continuous tracking analysis is performed based on the marked data.
[0091] Based on the characteristics of abnormal access behaviors, combined with access path information, operation permissions, and behavior patterns, an assessment is made to determine the risk level of each type of abnormal behavior. According to the risk criteria, different risk levels are classified to generate the assessment results of the security risk of file access;
[0092] The abnormal behaviors are classified into levels from low to high according to the risk criteria. By gradually analyzing the key nodes in the access path, situations of repeated or abnormal jumps in the path are identified, and a comparative assessment is made on the permission level and actual operation of the relevant behaviors. Combining the distribution characteristics of abnormal behaviors in the time dimension, the degree of deviation of the behavior pattern is further analyzed, and behaviors that are significantly inconsistent with the historical access pattern are marked as high risk. At the same time, the different risk levels are refined, and abnormal behaviors of different degrees are included in the corresponding levels by setting risk thresholds. Finally, the assessment results of the security risk of file access are generated, enabling clear hierarchical response bases for security management at the implementation level.
[0093] Please refer to Figure 4 , and the specific steps for obtaining the block data set are as follows:
[0094] Parse the assessment results of the security risk of file access, extract the sensitivity level and access frequency characteristics of each file in the file access records, convert the sensitivity data into numerical weights, and generate a file sensitivity and access frequency feature matrix by normalizing the number of accesses within a time period;
[0095] First, extract the data characteristics of each file, including the sensitivity level, access frequency, and time distribution. Digitally encode the sensitivity data and convert it into a quantified sensitivity level value. At the same time, based on the timestamp data in the access records, count the access frequency of each file, normalize the access frequency to eliminate the influence of time span or access cycle, and construct a file sensitivity and access frequency feature matrix by calculating the mean and variance of the sensitivity level value and the normalized access frequency. Each row in the matrix represents a file, and each column represents a certain eigenvalue of the file. This matrix accurately represents the distribution state of the file in terms of sensitivity and access characteristics.
[0096] Call the file sensitivity and access frequency feature matrix, segment the data according to the differences in sensitivity and access frequency, adjust the weight parameters in combination with the block size limit and data continuity requirements, and use the formula:
[0097] ;
[0098] Calculate the segmentation standard value of the data block to obtain the file data block segmentation characteristics;
[0099] Among them, represents the segmentation standard value of the data block, is the sensitivity weight, is the access frequency eigenvalue, is the average sensitivity of the data block, is the current file sensitivity value, is the content continuity parameter of the data block;
[0100] The advantage of the formula is that by combining the sensitivity weight, access frequency, and sensitivity difference in the operation, it can effectively balance the sensitivity and access characteristic distribution of the file data block, thus achieving the optimization of data segmentation;
[0101] Extract the sensitivity weight and access frequency values from the feature matrix. The sensitivity weight is allocated based on the sensitivity level of the file. For example, a file with a sensitivity level of 1 corresponds to a weight value of 0.5, and a file with a sensitivity level of 2 corresponds to a weight value of 1.0; the access frequency is obtained through normalization;
[0102] Extract the average sensitivity of the segmented data block, which is obtained by calculating the average value of the sensitivity values of the files within the block. The sensitivity of the current file is directly obtained from the feature matrix;
[0103] Combine the content continuity parameter to control the logical consistency of the segmentation, and set calculated based on the data logical distribution;
[0104] Assume that the sensitivity weight of a certain file in the feature matrix is , the access frequency is , the average sensitivity of the data block is , the current file sensitivity is , and the content continuity parameter is . Substitute the values into the formula:
[0105] ;
[0106] The result shows that the segmentation standard value of the data block is 0.624, which is used to subsequently determine whether the current file needs to be block-integrated with a certain data block or independently block-divided, providing a basis for data block division.
[0107] Call the file data block division characteristics, and determine the range of the attributed data block for each file by comparing the segmentation standard with the sensitivity and frequency values of the files in the feature matrix, generating a block-divided data set;
[0108] Call the file data block division characteristics. First, conduct a comparison of the segmentation standard value across all files, calculate the randomness of the attributed data block for each file based on its sensitivity weight and access frequency value, determine whether block division is required by setting a segmentation standard threshold, directly classify the files that meet the segmentation standard into the current data block, record the files that do not meet the standard as independently block-divided, and at the same time, during the block division process, adjust the continuity parameter of the block division content in real time based on the difference in file sensitivity to ensure the integrity and logic of the block division. Finally, integrate the list of files after block division, generate a block-divided data set, and complete the archiving with a structure that is logically clear and has a balanced distribution of sensitivity and access frequency, ensuring that the generated result meets the requirements of data storage and security management.
[0109] Please refer to Figure 5 , and the specific steps for obtaining the encrypted file data block set are as follows:
[0110] Based on the block-divided data set, split the data according to the preset data block size, perform encryption processing on each data block, and use the corresponding encryption technology to encrypt the content of each data block to generate encrypted data blocks;
[0111] Read the original data set one by one and perform a splitting operation with a fixed byte size. Each split data block is appended with a serial number as a unique identifier. After completing the data splitting, perform encryption processing on each data block separately. Use the encryption protocol to convert the data content byte by byte to ensure that the encryption result meets the confidentiality requirements. During the encryption process, perform an effectiveness check on the encryption result of each data block to verify whether the encrypted data is completely stored and matches the input data. At the same time, record the encryption time and ciphertext size of each encrypted data block. All encrypted data blocks are archived and saved in a structured manner to generate encrypted data blocks.
[0112] Based on the encrypted data blocks, sort them according to the unique identifier of each data block, and insert the encrypted data blocks into the B-tree structure in sequence. During the insertion process, sort and index each node to generate a B-tree index structure;
[0113] Perform a full scan of all data blocks using the identifiers attached to the encrypted data blocks, sort them in ascending order. After sorting, insert the encrypted data blocks into the nodes of the B-tree structure one by one. During the insertion process, preferentially allocate leaf nodes to store the data blocks, and at the same time check the storage status of each node to ensure that the data blocks within the node remain in an ordered state. Record the maximum and minimum values of each node in the parent node to form an index relationship. During the sorting and insertion process, for data blocks that exceed the storage capacity of the node, allocate them to new child nodes by splitting the node to maintain the balance of the B-tree, and finally generate the B-tree index structure.
[0114] Based on the B-tree index structure, associate the encrypted data blocks with the index entries, determine the storage location of each data block according to the location information of the data block, and store the location of each encrypted data block in the B-tree node through an index linked list to generate an encrypted archive data block set;
[0115] Bind the unique identifier of each encrypted data block to the corresponding index entry through a two-way mapping mechanism to ensure that the location of the data block can be quickly located. According to the storage path of the data block in the B-tree, trace back the node information level by level to determine the physical storage address of each data block, and add a linked list structure in the B-tree node to record the location relationship of the encrypted data blocks. Recheck all the linked list information to ensure that the storage location pointed to by the linked list is consistent with the actual data block location. Merge and store the location index table of the encrypted data blocks with the B-tree index structure to generate an encrypted archive data block set, providing an accurate index basis for subsequent retrieval and decryption.
[0116] Please refer to Figure 6 , and the steps for obtaining the requirements for archive security protection are specifically as follows:
[0117] Extract the access behavior logs from the encrypted archive data block set, and use the access timestamps, user identifiers, and operation type information in the logs to filter out illegal accesses and abnormal operations in chronological order to generate an initial set of risk operations;
[0118] First, it is necessary to perform a structural analysis on the block data, extract the timestamps of each record and sort them in chronological order. The user identification is matched with the independent identification code one by one through the system log recording function. The operation type needs to analyze the operation characteristics of reading, writing, and deleting through the behavior fields in the record. At the same time, duplicate removal and consistency verification are performed on the access records. The consistency verification includes the matching verification of user permissions and operation types and the rationality evaluation of duplicate operations within the time series. After completing the verification, the logs are further reorganized in chronological order to form a complete set of operation sequences, screening access behavior records with suspicious characteristics, and identifying whether there are unauthorized or abnormal operations by comparing characteristics such as operation frequency, path length, and permission type. Finally, an initial set of risk operations is formed, which will serve as the basis for subsequent analysis.
[0119] Analyze the access behavior characteristics based on each record in the initial set of risk operations, including operation frequency, access path length, and operation type weight, using the formula:
[0120] ;
[0121] Calculate the access risk characteristic value of each record to obtain an updated set of risk operation evaluations;
[0122] Among them, represents the access risk characteristic value, represents the operation frequency, represents the average value of the operation frequency, represents the standard deviation of the operation frequency, represents the access path length, represents the maximum access path length, represents the operation type weight, are the adjustment coefficients for frequency, path, and weight respectively;
[0123] The benefit of the formula is to balance the importance of each characteristic value by adjusting the weight parameters, making it more in line with the actual scenario of risk operations, and improving the accuracy of the evaluation by combining the standardized numerical range.
[0124] The operation frequency is obtained by counting the number of operations within a specific time period;
[0125] The average value of the operation frequency is obtained by averaging the frequencies of all records;
[0126] The standard deviation of the operation frequency is calculated using the formula calculate, is the total number of records;
[0127] The access path length is obtained by parsing the hierarchical depth of the user access path in the log;
[0128] The maximum path length is the maximum value of the path hierarchy levels in the access log;
[0129] The operation type weight maps the importance of a specific operation type through a given weight value, and the weight is determined based on the potential impact of the operation on data integrity and security;
[0130] The adjustment coefficients representing frequency, path, and weight respectively are set according to the risk characteristic distribution and are related to the fluctuations in the actual scenario;
[0131] Set , , , , , , , , , substitute into the formula:
[0132] ;
[0133] ;
[0134] ;
[0135] ;
[0136] The result shows that the risk characteristic value of the access behavior is 1.143, and further classification and analysis can be carried out based on this value.
[0137] Cluster the records in the updated risk operation evaluation set according to the behavior patterns, group the records with the same risk characteristic values into one group, and judge whether there are illegal accesses and abnormal operations based on the distribution of the risk values within the group in the clustering result, and generate the requirements for file security protection;
[0138] First, the risk characteristic value of each record needs to be analyzed, and the similarity between records is calculated using the Euclidean distance formula. Records with similar characteristic values are grouped into the same group based on the similarity. For each group of records, the internal maximum risk value needs to be evaluated. The calculation of the maximum risk value is based on the highest characteristic value of the records in the group. By setting a dynamic risk threshold, the rationality of the distribution of risk characteristic values of each group is verified. In the verification step, the risk characteristic value is compared with the dynamic threshold. When the characteristic values in the group are all lower than the threshold, the group is marked as low risk. If the characteristic value of a group is higher than the threshold, it is marked as high risk. The grouping results and risk tags are combined to further determine whether it constitutes illegal access or abnormal operation. Finally, the archive security protection requirements are generated based on the analysis results of all grouped records to guide the subsequent implementation of security policies.
[0139] See also Figure 7 , the specific steps for obtaining the security protection strategy execution results are:
[0140] Based on the requirements of archive security protection, analyze the adjustment conditions of access rights, call the access rights parameter set and user hierarchical permission parameter set, determine the permissions to be adjusted, and generate the user set whose permissions need to be adjusted;
[0141] Analyze the adjustment conditions of access rights. First, obtain access rights parameter set A and user graded access rights parameter set B. Compare the user access rights in parameter set A with the user permission levels in parameter set B item by item, establish a user permission difference set, and screen users with large permission differences by analyzing the distribution trend of the set and the boundary values of the permission level difference intervals. Then, by comparing the historical data of user permission changes in the permission difference set, confirm the specific direction and scope of the permission adjustment. By matching the analysis results with the permission change standards in the preset adjustment rules item by item, obtain the user set whose permissions need to be adjusted.
[0142] Based on the user set that needs to adjust permissions, combined with the user multi-factor authentication parameter set and the access permission parameter set, the formula ;
[0143] Calculate the modified user access permission value and generate an adjusted access permission set;
[0144] in, is the corrected user access permission value. is the current permission level, is the preset authority base value, Weight parameter for multi-factor authentication;
[0145] The formula is beneficial by introducing the weight parameter of multi-factor authentication , which can effectively balance the current permissions Compared with the preset authority base value The impact on the correction result makes the corrected permissions more accurate, and at the same time enhances the system's sensitivity to multi-factor authentication information;
[0146] Suppose the current permission level of a certain user , the preset permission benchmark value , the multi-factor authentication weight parameter , the steps to calculate the corrected user access permission value are as follows:
[0147] Calculate the numerator part: ;
[0148] Calculate the denominator part: ;
[0149] Calculate the corrected permission: ;
[0150] This result indicates that the corrected user access permission value is 5.67, which is improved compared to the current permission level, conforms to the directional adjustment of the preset permission benchmark value, and provides a data basis for the further adjustment of the permission set P.
[0151] Call the adjusted access permission set, combine the user behavior monitoring parameter set and the deviation prediction model parameter set, compare the user operation deviation with the predicted deviation, and generate the execution result of the security protection strategy;
[0152] Call the adjusted access permission set, obtain the deviation value between the user operation behavior and the predicted behavior through the user behavior monitoring parameter set and the deviation prediction model parameter set. The calculation process includes: extracting features from the historical operation behavior records of the user, obtaining the operation mode and behavior frequency of each user, performing deviation quantization analysis on the current operation data of the user and the historical operation mode, classifying using a difference threshold, generating the user behavior prediction value through the preset prediction formula in the deviation prediction model parameter set E, then comparing the actual operation behavior with the prediction value one by one to calculate the deviation value, matching the deviation value distribution with the security protection strategy library to obtain the specific difference level between the operation behavior deviation and the prediction, and combining the security protection strategy execution conditions to compare the difference level with the trigger conditions of the protection strategy to generate the execution result of the security protection strategy.
[0153] Please refer to Figure 8 , the specific steps for obtaining the resource usage and security protection records are as follows:
[0154] Based on the execution results of security protection policies, filter the user behavior parameters associated with risky operations according to the user operation logs and resource access data recorded in real-time monitoring. By classifying user behaviors and marking the operation types of sensitive file accesses, and combining the timestamps, user identity identifiers, and resource call types marked in the logs, obtain the user risk behavior parameter set;
[0155] Extract the core parameters in the operation logs according to the user operation logs and resource access data recorded in real-time monitoring, including operation types, timestamps, user identity identifiers, resource call identifiers, etc. Filter the operation records involving sensitive file accesses. Further analyze based on the content types in the operation logs, classify different operation behaviors and classify and mark sensitive operations. By calling the resource call record database, screen each operation's associated resource call history item by item, identify the resources associated with risky operations, extract the sensitivity levels and access frequencies of the resources, match the user identity identifiers with the operation behaviors, generate the basic records of user risk behaviors. Through parameter analysis and behavior pattern modeling of the basic records of risk behaviors, extract the behavior characteristic parameters such as the operation frequency, resource call type, and operation time interval of the user, integrate to obtain the user behavior characteristic set, sort the extracted user operation characteristics according to the sensitivity level, and store them as the user risk behavior parameter set for subsequent priority calculation and adjustment of resource allocation strategies.
[0156] Based on the behavior frequency and sensitivity level in the user risk behavior parameter set, use the formula:
[0157] ;
[0158] Calculate the priority score of the user operation, evaluate the impact of the user risk operation and quantify the priority to obtain the user risk operation score result;
[0159] Among them, represents the priority score of the user operation, represents the behavior frequency, represents the sensitivity level, represents the resource call type, represents the operation time interval, is the weight parameter;
[0160] The benefit of the formula is that by introducing multiple parameters such as the user operation frequency, sensitivity level, resource call type, and operation time interval, and combining different weights, it quantifies the user behavior priority, improving the resource allocation efficiency and security;
[0161] Extract the behavior frequency 、sensitivity level (quantified from 1 to 5, 5 being the highest), resource call type (Normal resource call is 1, priority resource call is 2), operation time interval Unit time;
[0162] Set weight parameter , substitute into the formula:
[0163] ;
[0164] The result shows that the priority score of the user operation is 23.25. The value is used in subsequent dynamic resource allocation to guide the execution of the priority resource allocation strategy and ensure the security requirements of sensitive operations and critical resources.
[0165] Combined with the user risk operation score result, dynamically optimize the resource allocation strategy, preferentially allocate resources such as bandwidth, storage space, and access rights to high-score risk operations, and obtain resource usage and security protection records;
[0166] Extract the resource sensitivity level, call frequency, and bandwidth requirements corresponding to high-priority user operations from the score results, dynamically sort the user operation priorities, preferentially allocate resources to user operations with higher scores, dynamically adjust resources such as sensitive file access rights, storage space, and bandwidth by real-time retrieving the resource allocation strategy database, combine the sensitivity level coefficients of the score data, optimize the resource allocation strategy, record the user operations, timestamps, resource types, and allocation results of each resource allocation, synchronously update the allocation records in the resource usage log, and periodically monitor the usage of the allocated resources to ensure the security and efficiency of resource usage, and generate resource usage and security protection records.
[0167] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical content of the technical solution of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. A file electronic information security system, characterized in that: The system comprises: The archive operation behavior prediction module extracts the time series characteristics of archive operations based on the operation data of archive users, predicts the user's next operation according to the user's original operation mode, obtains the operation behavior prediction results, identifies potential risk behaviors, and generates archive access security risk assessment results; The archive data encryption and security storage module stores the data in blocks according to the archive access security risk assessment results and the archive data content, divides the data according to the archive sensitivity and access frequency, generates a block data set, encrypts the archive data block by block, stores the encrypted data blocks through the B-tree index structure, and generates an encrypted archive data block set; The archive security protection strategy execution module identifies risky operation behaviors based on the encrypted archive data block set, generates archive security protection requirements, adjusts access rights, executes multi-factor identity authentication security measures, monitors deviations between operation behaviors and predictions, and generates security protection strategy execution results; The steps for obtaining the archive security protection requirements are specifically as follows: Extracting access behavior logs from the encrypted archive data block set, using access timestamps, user identifiers and operation type information in the logs, filtering illegal access and abnormal operations in chronological order, and generating an initial set of risky operations; The access behavior characteristics are analyzed according to each record in the initial set of risk operations, including operation frequency, access path length and operation type weight, using the formula: ; Calculate the access risk feature value of each record to obtain an updated risk operation assessment set; in, Indicates the access risk characteristic value, Indicates the operating frequency, represents the average value of the operating frequency, represents the standard deviation of the operating frequency, Indicates the access path length, Indicates the maximum access path length, Indicates the weight of the operation type, are the adjustment coefficients for frequency, path, and weight, respectively; Performing behavior pattern clustering on the records in the updated risk operation assessment set, grouping records with similar risk characteristic values together, judging whether there is illegal access and abnormal operation based on the distribution of risk values within the group in the clustering result, and generating archive security protection requirements; The archive resource scheduling and security monitoring module dynamically schedules archive resources in real time based on the execution results of the security protection strategy, allocates resources to risk operations first, optimizes access rights, storage space, and bandwidth resource allocation by real-time monitoring of user operations, and obtains resource usage and security protection records.
2. The electronic archive information security system according to claim 1, characterized in that: The steps for obtaining the operation behavior prediction result are specifically as follows: Based on the operation data of archive users, extract the operation data features of archive users, including access time, operation type and access object, analyze the time series feature vector of each data feature, and generate a time series matrix through the sequence trend of operation time; The characteristic vectors in the time series matrix are analyzed to extract the changing trend and correlation of the operation behavior characteristics. The time interval distribution and the characteristic change rate are calculated using the formula: ; Calculate the characteristic change trend amplitude; in, Indicates the magnitude of the characteristic change trend. For the time series The eigenvalues of the features, is the mean of the eigenvalues, is the feature weight parameter, Indicates the total number of features; The feature change trend amplitude is called, compared with the user's original operation mode, the user's next operation behavior is predicted through the feature change trend, the prediction probability analysis is performed on each operation type, and normalization processing is performed to obtain the operation behavior prediction result.
3. The electronic archive information security system according to claim 2, characterized in that: The steps for obtaining the archive access security risk assessment results are specifically as follows: Based on the operation behavior prediction results, extract the original data of each user's access profile, perform access frequency statistics and classify access types, analyze the distribution of access time periods, perform screening based on permission configuration and access behavior characteristics, and generate access behavior statistics results; Based on the access behavior statistics, combined with the access behavior frequency, access time and user authority range, abnormal behavior pattern recognition is performed, high-frequency access and access behaviors that do not match authority are marked, and compared with normal patterns to obtain abnormal access behavior characteristics; Based on the abnormal access behavior characteristics, an assessment is performed in combination with access path information, operation permissions and behavior patterns to determine the risk level of each type of abnormal behavior, and differentiated risk levels are divided according to risk standards to generate archive access security risk assessment results.
4. The electronic archive information security system according to claim 3 is characterized in that: The steps of obtaining the block data set are specifically as follows: Analyze the archive access security risk assessment results, extract the sensitivity level and access frequency characteristics of each archive in the archive access record, convert the sensitivity data into a numerical weight, and generate an archive sensitivity and access frequency characteristic matrix by normalizing the number of accesses within a time period; The sensitivity and access frequency characteristic matrix of the archive is called, and the data is segmented according to the differences in sensitivity and access frequency. The weight parameters are adjusted in combination with the block size limit and data continuity requirements, and the formula is used: ; Calculate the segmentation standard value of the data block and obtain the segmentation characteristics of the archive data; in, Indicates the segmentation standard value of the data block, is the sensitivity weight, To access the frequency eigenvalues, is the average sensitivity of the data block, is the current file sensitivity value, is the content continuity parameter of the data block; The file data block feature is called, and the data block range to which each file belongs is determined by comparing the segmentation standard with the sensitivity and frequency values of the files in the feature matrix, so as to generate a block data set.
5. The electronic archive information security system according to claim 4 is characterized in that: The steps of obtaining the encrypted archive data block set are specifically as follows: Based on the block data set, the data is split according to a preset data block size, each data block is encrypted, and the content of each data block is encrypted using a corresponding encryption technology to generate an encrypted data block; Based on the encrypted data blocks, sorting is performed according to the unique identifier of each data block, and the encrypted data blocks are sequentially inserted into the B-tree structure, and each node is sorted and indexed during the insertion process to generate a B-tree index structure; Based on the B-tree index structure, the encrypted data blocks are associated with index entries, the storage location of each data block is determined according to the location information of the data block, and the location of each encrypted data block is stored in the B-tree node through an index linked list to generate an encrypted archive data block set.
6. The electronic archive information security system according to claim 1, characterized in that: The steps for obtaining the security protection strategy execution result are specifically as follows: Based on the archive security protection requirements, the access permission adjustment conditions are analyzed, the access permission parameter set and the user hierarchical permission parameter set are called, the permissions to be adjusted are determined, and a user set whose permissions need to be adjusted is generated; Based on the user set whose permissions need to be adjusted, combined with the user multi-factor authentication parameter set and the access permission parameter set, the formula is: ; Calculate the modified user access permission value and generate an adjusted access permission set; in, is the corrected user access permission value. is the current permission level, is the preset authority base value, Weight parameter for multi-factor authentication; The adjusted access permission set is called, combined with the user behavior monitoring parameter set and the deviation prediction model parameter set, the user operation deviation is compared with the prediction deviation, and the security protection strategy execution result is generated.
7. The electronic archive information security system according to claim 6, characterized in that: The steps for obtaining the resource usage and security protection records are as follows: Based on the execution results of the security protection strategy, according to the user operation logs and resource access data recorded by real-time monitoring, the user behavior parameters associated with the risky operation are screened, and the user risk behavior parameter set is obtained by classifying the user behavior and marking the operation type of sensitive file access, combined with the timestamp, user identity and resource call type marked in the log; Based on the behavior frequency and sensitivity level in the user risk behavior parameter set, the formula is used: ; Calculate the priority score of user operations, evaluate the impact of user risk operations and quantify the priority, and obtain the user risk operation score result; in, Represents the priority score of the user operation. Represents the frequency of behavior, Represents the sensitivity level, Represents the resource call type. Represents the operation time interval, is the weight parameter; Combined with the user risk operation scoring results, the resource allocation strategy is dynamically optimized to prioritize bandwidth, storage space, and access rights to risk operations with high scores, and resource usage and security protection records are obtained.
Citation Information
Patent Citations
Intelligent file borrowing and utilizing platform based on zero-trust authentication
CN116228167A
Archive management method and system based on cloud archive library
CN116595556A