Big data electronic archive management method
By building a risk factor identification model and collecting and analyzing electronic archive data in real time, it solves the problem that traditional methods are difficult to monitor and evaluate electronic archive management risks in real time, and realizes intelligent risk assessment and report generation, which improves risk prevention and control capabilities and intelligence levels.
Patent Information
- Application Number
- CN202510253205.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional electronic archive management methods are difficult to discover potential security risks in real time, cannot accurately assess the severity of risks, and cannot intelligently generate risk assessment reports, providing file managers with decision-making support for risk prevention and control.
By collecting data related to electronic archives in real time, a risk factor identification model is built, including deep Q network, convolutional LSTM and BiLSTM-Attention models, real-time monitoring and analysis of data leakage risks, storage device failure risks, and regulatory compliance risks, and intelligently generate risk assessment reports.
Real-time monitoring and early warning of various risks is achieved, the severity of risks is accurately assessed, and intelligent risk assessment reports are provided, which has improved the risk prevention and control capabilities and intelligence level of electronic file management.
Smart Images

Figure CN120179610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic file management, and specifically provides a big data electronic file management method. Background Art
[0002] With the rapid development of information technology, electronic files have become an indispensable information carrier in modern society. A large number of electronic documents are generated by various institutions and organizations in their daily operations, and these documents need to be effectively managed, stored, and utilized. The rise of big data technology has provided new possibilities and challenges for electronic file management.
[0003] There are deficiencies in traditional electronic file management methods. On the one hand, traditional technologies cannot help file managers timely discover potential security risks and accurately evaluate the severity of risks by quantifying risk impact factors. On the other hand, this method cannot intelligently generate risk assessment reports to provide decision-making support for file managers in risk prevention and control.
[0004] In summary, traditional technologies are significantly insufficient and difficult to meet the requirements of electronic file management in the big data environment. Therefore, it is particularly important to develop a big data electronic file management method. Summary of the Invention
[0005] The purpose of the present invention is to make up for the deficiencies of the existing technology and provide a big data electronic file management method. It can collect relevant data of electronic files in real time, construct a risk factor identification model, deeply mine and analyze the data, realize real-time monitoring and early warning of various risks such as data leakage risk, storage device failure risk, and regulatory compliance risk. At the same time, this method should also be able to intelligently generate risk assessment reports to provide decision-making support for file managers in risk prevention and control, and further improve the risk prevention and control ability and intelligent level of electronic file management.
[0006] To solve the above technical problems, the present invention provides the following technical solution: A big data electronic file management method, and the specific steps of this method are as follows:
[0007] S1. Data collection and integration
[0008] Deploy data collectors at each key node to collect relevant data of electronic files in real time, including but not limited to access records of files, modification records, operating parameters of storage devices, system log files, updated information of laws and regulations, and industry standard documents. For access records, record the timestamp t of each access a , the unique identifier u of the accessing user id , the access operation type identifier o t , the modification record includes the modification time t m , the hash value h of the modified contentc , the identity identifier p of the modifier id , the operating parameters of the storage device include temperature T, humidity H, and disk space occupancy rate S o , read / write speed V r , error rate E r ;
[0009] Clean the collected multi-source heterogeneous data to remove duplicate, incorrect, and incomplete data. Use a similarity-based cleaning algorithm. For two data items d i and d j in the dataset, calculate their similarity Sim(d i ,d j ):
[0010]
[0011] where n is the number of attributes of the data item, w k is the weight of the k-th attribute. First, construct an attribute importance judgment matrix A, calculate its maximum eigenvalue λ max and the corresponding eigenvector W, perform normalization on W to obtain w k , and then integrate the cleaned data into a unified big data platform. Use a graph model-based data integration method. Represent the data from different data sources as nodes and the association relationships between the data as edges to construct a data integration graph G=(V, E), where V is the set of nodes and E is the set of edges. Implement data integration through a graph traversal algorithm;
[0012] S2. Construction of risk factor identification model
[0013] For the risk of data leakage, construct a risk identification model. Define the state space S, which includes the privilege level r l of the current access user, h historical access frequency f s , and the sensitivity s action of the current access as features, the action space A n includes judging normal access a a and abnormal access a
[0014] , and the reward function R. Give a positive reward when the judgment is correct and a negative reward when the judgment is incorrect. Use the Deep Q-Network (DQN) algorithm for model training. The network structure includes an input layer, multiple hidden layers, and an output layer. The input layer receives the features of the state space S. After non-linear transformation by the hidden layers, the Q value of each action is obtained at the output layer. During the training process, continuously interact with the environment and update the network parameters θ according to the reward feedback. The update formula is: t where, θDenote the network parameters of the model at time step t, α is the learning rate, r is the reward value given by the reward function R, and γ is the discount factor. Denote the maximum Q value predicted by the model for all possible actions a' in state s'. Q(s, a; θ t ) denotes the Q value predicted by the model when taking action a in the current state s. is the gradient of the Q function with respect to the network parameters θ t , and its optimal value is determined through cross-validation.
[0015] For the risk of storage device failure, a prediction model is established. The operating parameters of the storage device are organized into multi-dimensional data in a time series as the input of the model. The model consists of a convolutional layer, a ConvLSTM layer, and a fully connected layer. The convolutional layer is used to extract local features of the data, and the calculation formula for the convolution operation is:
[0016]
[0017] where is the feature map at the i-th row and j-th column of the l-th layer, is the convolutional kernel weight, M represents the size of the convolutional kernel in the vertical direction, m represents the index of the convolutional kernel in the vertical direction (row direction), N represents the size of the convolutional kernel in the horizontal direction (column direction), n represents the index of the convolutional kernel in the horizontal direction (column direction), and b l is the bias.
[0018] The ConvLSTM layer is used to capture long-term dependencies in time series data, and its calculation formula is:
[0019] i t = σ(W xi x t + W hi h t-1 + W ci c t-1 + b i )
[0020] f t = σ(W xf x t + W hf h t-1 + W cf c t-1 + b f )
[0021] c t = f t × c t-1 + i t × tanh(W xc x t + Whc h t-1 +b c )
[0022] o t =σ(W xo x t +W ho h t-1 +W co c t-1 +b o )
[0023] h t =o t ×tanh(c t )
[0024] Among them, i t 、f t 、o t are the input gate, forget gate and output gate respectively, c t is the cell state, h t is the hidden state, σ is the sigmoid function, W xi 、W hi 、W ci are the weight matrices connecting the input gate with the current input, the hidden state at the previous moment, and the cell state at the previous moment respectively, x t is the input data at the current moment, h t-1 is the hidden state at the previous moment, c t-1 is the cell state at the previous moment, W xc 、W hc are the weight matrices connecting with the current input and the hidden state at the previous moment respectively, b i 、b f are the biases of the input gate and the forget gate respectively, W xf 、W hf 、W cf are the weight matrices connecting the forget gate with the current input, the hidden state at the previous moment, and the cell state at the previous moment respectively, tanh is the hyperbolic tangent function, b c is the bias used to calculate the current cell state, W xo 、W ho 、W co are the weight matrices connecting the output gate with the current input, the hidden state at the previous moment, and the cell state at the previous moment respectively, b o is the bias of the output gate. The fully connected layer maps the output of the ConvLSTM layer to the prediction result, and the model is trained by minimizing the mean square error loss function;
[0025] For regulatory compliance risks, the BiLSTM-Attention model based on the attention mechanism is used to preprocess the laws and regulations documents and the operation process documents in the electronic file management system, converting them into word vector sequences. The BiLSTM layer processes the word vector sequences from both the forward and backward directions to obtain the forward hidden state and the backward hidden state Concatenate to obtain the hidden state The attention mechanism calculates the attention weight α of each hidden state t , and performs a weighted sum of the hidden states to obtain the feature representation r of the document:
[0026]
[0027] Among them, this formula calculates the correlation between the hidden states h i and h j , and e ij is used to calculate the attention weight α j , and T represents the transpose operation of the matrix;
[0028]
[0029] Among them, W iearnable is a learnable weight matrix, L seq is the sequence length, exp is the mathematical representation of the exponential function, and e ik represents the correlation score between the hidden state h i and the hidden state h k , k is an index variable, and finally the classifier classifies the feature representation of the document to determine whether it meets the regulatory requirements;
[0030] S3. Real-time risk monitoring and analysis
[0031] Input the integrated data into the above-built risk factor identification model in real time to monitor the data leakage risk, storage device failure risk, and regulatory compliance risk in the electronic file management process. For the identified potential risks, start the risk analysis module. For the data leakage risk, analyze the sensitivity of the leaked data, the number of users involved, and the business area. For the storage device failure risk, evaluate the number of files and the importance that may be affected by the failure. For the regulatory compliance risk, determine the specific clauses violated by the regulations and the possible degree of punishment. Quantify the risk degree by calculating the risk impact factor RI. For the data leakage risk:
[0032]
[0033] Among them, is the influence degree of the i-th influencing factor on the data leakage risk, is its weight;
[0034] For the risk of storage device failure:
[0035] For the risk of regulatory compliance:
[0036] S4, Intelligent early warning and risk assessment report generation
[0037] According to the risk impact factor RI obtained from the risk analysis, set the thresholds for different risk levels. When RI reaches the corresponding threshold, the system automatically sends out early warning information, which is sent to relevant management personnel through multiple channels such as text messages, emails, and instant messaging tools, including detailed information on the risk type, risk level, time, and location of the risk occurrence. At the same time, the system generates a risk assessment report, and the report content includes a detailed description of the risk, an analysis of the cause of the risk occurrence, an assessment of the impact of the risk on the electronic file management system and business, and a prediction of the risk development trend. The report is presented in a visual form, and various charts such as bar charts, line charts, and radar charts are used to display risk-related information;
[0038] S5, Generation and push of countermeasure suggestions
[0039] For different types and levels of risks, the system generates corresponding countermeasure suggestions based on the pre-established expert knowledge base and policy library. For the risk of data leakage, the suggestions include immediately freezing relevant user accounts, strengthening network security protection measures, and encrypting the leaked data. For the risk of storage device failure, the suggestions include starting an emergency backup program, arranging equipment repair or replacement plans. For the risk of regulatory compliance, the suggestions include conducting regulatory training, adjusting management processes, and performing internal audits. The generated countermeasure suggestions are pushed to relevant management personnel, and at the same time, the implementation steps and precautions of the suggestions are provided.
[0040] Furthermore, in the data collection and integration step, when cleaning data from different sources, for the text data contained in the dataset, a cleaning method based on word vector similarity is adopted. For two sentences S1 and S2 in the text data, they are converted into word vector representations and The similarity of the sentences is measured by calculating the cosine similarity:
[0041]
[0042] where is the dot product of the vectors, and are the norms of the vectors respectively. For text data with a similarity higher than the set threshold θ t it is considered duplicate data for cleaning, and the threshold θ tBy performing clustering analysis on historical text data, in practical applications, K-means clustering is performed on a large amount of historical text data to analyze the similarity distribution between different clusters, and an appropriate threshold is selected to ensure that duplicate text data can be effectively removed without accidentally deleting valuable information.
[0043] Furthermore, when constructing the data leakage risk identification model, in order to improve the generalization ability and accuracy of the model, a transfer learning method is adopted. In the pre-training stage, a large-scale public network access data is used to pre-train the model to learn general network access behavior patterns. Then, in the fine-tuning stage, the electronic file access record data within the enterprise is used to fine-tune the pre-trained model. During the fine-tuning process, the parameters of some pre-trained layers are fixed, and only the parameters of some layers are updated. Let the pre-trained model be M p , whose parameters are θ p , and the set of parameters updated during fine-tuning is θ f , for the input data x, the output y of the fine-tuned model is:
[0044] y = M p (x; θ p , θ f )
[0045] where, θ f is updated by minimizing the loss function L on the internal data of the enterprise f :
[0046]
[0047] where, is the true label, is the model prediction result, N is the number of samples of the internal data of the enterprise. Through transfer learning, the general features learned from the public data can converge to the model parameters suitable for the internal data of the enterprise faster, improving the performance of the model in the enterprise-specific environment.
[0048] Furthermore, when constructing the storage device failure risk prediction model, considering the differences in the operating characteristics of different types of storage devices, the model is trained individually. For different types of storage devices, historical operation parameter data and failure records are collected respectively. For each type of storage device, an independent ConvLSTM model is constructed. During the model training process, hyperparameters such as the convolutional kernel size and the number of hidden units in the ConvLSTM layer are adjusted to adapt to the characteristics of different devices. For hard disk devices, since their read and write operations have strong periodicity, the convolutional kernel size can be set to a larger value to better capture the periodic characteristics of the data. For solid-state drives, since their read and write speeds are fast and the data changes frequently, the number of hidden units can be appropriately increased to improve the model's ability to capture complex data patterns. Through this individualized training, the prediction accuracy of the model for the failure risks of different types of storage devices can be improved.
[0049] Furthermore, when constructing the regulatory compliance risk identification model, to better understand the semantic information in legal and regulatory documents, an external knowledge graph is introduced. The knowledge graph contains entity, relationship, and attribute information related to laws and regulations. The entities in the operation process documents and legal and regulatory documents in the electronic file management system are matched with the entities in the knowledge graph. Through the relationship reasoning of the knowledge graph, the semantic representation of the documents is enriched. For entity e in the document, the set of its neighbor entities in the knowledge graph is N(e). The representation of entity e is updated by aggregating the information of neighbor entities:
[0050]
[0051] where w n is the weight of neighbor entity n, which is determined by calculating the relationship strength between entity e and neighbor entity n. By introducing the external knowledge graph, the model's understanding and judgment ability of regulatory compliance risks can be enhanced.
[0052] Furthermore, in the real-time risk monitoring and analysis step, to timely detect the dynamic changes of risks, the sliding window technique is adopted. For each risk type, a time window T with a fixed size is set w , and data is collected and analyzed in real time within the window. As time goes by, the window slides continuously, each time sliding by a time step Δt. Within each window, the risk impact factor RI is recalculated. For the data leakage risk, access records are collected within window T w , and the frequency and scope of abnormal access in different time periods are calculated. Based on this information, RI is recalculated d . Through the sliding window technique, the dynamic change trend of risks can be timely captured, and the timeliness and accuracy of risk monitoring can be improved.
[0053] Furthermore, in the step of generating the intelligent early warning and risk assessment report, for the visual display of the risk assessment report, interactive visualization technology is adopted. Users can, through interactive operations, view the charts in the report in a personalized manner. For the line chart of the risk development trend, users can view the details of risk changes within a specific time period through zooming operations. For the radar chart showing different risk factors, users can filter out the risk factors of concern for key viewing. By providing such an interactive visualization interface, users can more conveniently obtain the required risk information, deeply analyze the risk situation, and provide more powerful support for decision-making.
[0054] Furthermore, in the step of generating and pushing response suggestions, to improve the effectiveness and operability of response suggestions, case-based reasoning technology is introduced. When generating response suggestions, the system first searches the case base for historical cases similar to the current risk situation. Each case in the case base contains risk descriptions, response measures, and implementation effect information. By calculating the similarity between the current risk and historical cases, the case with the highest similarity is selected as a reference. For the calculation of similarity, a method based on multi-attribute decision-making is adopted, comprehensively considering multiple attributes such as risk type, risk level, and influence scope. Let the attribute vector of the current risk be The attribute vector of the historical case be Similarity Is:
[0055]
[0056] Where n represents the number of attributes considered when calculating similarity, w i Is the weight of the i-th attribute, Similarity i (a i , b i ) is used to calculate the similarity between the i-th attribute in the current risk attribute vector a and the i-th attribute in the historical case attribute vector b. a i The value of the i-th attribute in the current risk attribute vector a, b i The value of the i-th attribute in the historical case attribute vector b. The response suggestions are adjusted based on the response measures of the reference case and the current actual situation to generate the final response suggestions. Through case-based reasoning technology, historical experience can be drawn on to improve the pertinence and effectiveness of response suggestions.
[0057] Furthermore, to ensure the stability and reliability of the entire big data electronic file management system, a distributed architecture is adopted. The data collection, storage, analysis, and early warning function modules are distributed on multiple nodes, and communication and collaboration are carried out through the network. In terms of data storage, a distributed file system is used to disperse the electronic file data on multiple storage nodes, improving the storage capacity and fault tolerance of the data. In terms of data processing, a distributed computing framework is adopted to decompose large-scale data processing tasks into multiple subtasks and process them in parallel on multiple computing nodes, improving the processing efficiency. At the same time, by setting up a master node and slave nodes, load balancing and fault recovery of the system are achieved. When a certain node fails, the master node can automatically reassign tasks to other normal nodes to ensure the normal operation of the system. Through the distributed architecture, the high concurrency and large-scale data processing requirements in the big data environment can be effectively addressed, improving the overall performance and reliability of the system.
[0058] Compared with the prior art, the big data electronic file management method has the following beneficial effects:
[0059] First, by collecting relevant data of electronic files in real time and constructing a risk factor identification model, this method can monitor and analyze data leakage risks, storage device failure risks, and regulatory compliance risks in real time. This can not only help file managers timely discover potential security hazards but also accurately evaluate the severity of risks by quantifying risk impact factors. Thus, file managers can quickly take corresponding countermeasures to avoid or reduce losses caused by risks. In addition, this method can also intelligently generate risk assessment reports, providing decision-making support for risk prevention and control for file managers and further enhancing the risk prevention and control ability of electronic file management.
[0060] Second, by constructing a risk factor identification model, a real-time risk monitoring and analysis system, an intelligent early warning and risk assessment report generation system, and a response suggestion generation and push system, this method can automatically complete risk identification, analysis, early warning, and response tasks in the process of electronic file management. This not only reduces the workload of file managers but also improves the efficiency and accuracy of file management. At the same time, this method adopts a distributed architecture and has high concurrency and large-scale data processing capabilities, which can meet the requirements of electronic file management in the big data environment and provide strong support for the intelligent and automated development of electronic file management.
[0061] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0063] Figure 1 It is a flowchart operation diagram of a big data electronic file management method. Detailed implementation manners
[0064] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific implementation manners, structures, features and their effects of the present invention as follows.
[0065] Embodiment 1
[0066] This embodiment describes that a large enterprise has a vast amount of electronic files, covering customer information, financial statements, and project documents, stored in multiple servers in the enterprise data center. These files are frequently accessed and modified by employees of different departments every day and need to strictly comply with industry regulations.
[0067] Deploy data collectors at key nodes of servers, network devices and application systems in the data center to comprehensively collect data related to electronic files. For example, collect file access records (including access timestamp t a , unique access user identifier u id , access operation type identifier o t ) at the server end, modification records (modification time t m , modified content hash value h c , modified personnel identity identifier p id ), and storage device operation parameters (temperature T, humidity H, disk space occupancy rate S o , read / write speed V r , error rate E r ). Collect traffic data at network devices to assist in analyzing access situations, obtain operation logs and relevant process documents from application systems, and use a similarity measure cleaning algorithm to clean the data. For example, for two data items d i and d j in the dataset, calculate their similarity Construct an attribute importance judgment matrix A, calculate the maximum eigenvalue λ max and the corresponding eigenvector W, and perform normalization processing to obtain w k, Assume that the data item has three attributes: access time, user identifier, and operation type. By analyzing historical data, the attribute weights are determined to be 0.3, 0.4, and 0.3 respectively. Taking the cleaned archive access records as an example, if there are two records d1(t a1 ,u id1 ,o t1 ) and d2(t a2 ,u id2 ,o t2 ), Similarity1(t a1 ,t a2 ) calculates the similarity score within a certain range of time differences (for example, getting 0.8 within 1 hour of difference, otherwise getting 0.2), Similarity2(u id1 ,u id2 ) determines whether the user identifiers are the same (getting 1 if the same, 0 if different), Similarity3(o t1 ,o t2 ) determines whether the operation types are the same (getting 1 if the same, 0 if different). Assume d1(10:00, user001, read) and d2(10:30, user001, read), then Sim(d1,d2) = 0.3×0.8 + 0.4×1 + 0.3×1 = 0.94. If the similarity threshold is set to 0.9, these two records may be duplicate data for cleaning.
[0068] Adopt the data integration method of the graph model, represent the data from different data sources as nodes, and the association relationships as edges, to construct the data integration graph G=(V,E), and integrate the data through the graph traversal algorithm.
[0069] Construct the state space S, which includes the current access user privilege level r ι , historical access frequency f h , and the sensitivity s of this access s features. The action space A includes judging as normal access a n and abnormal access a a . The reward function R gives a positive reward when the judgment is correct and a negative reward when it is wrong. Adopt the Deep Q-Network (DQN) algorithm to train the model. The input layer of the network receives the features of the state space S, and after non-linear transformation through the hidden layer, the output layer obtains the Q value of each action. During the training process, update the network parameters θ according to the formula , where the learning rate α = 0.001 and the discount factor γ = 0.9, which are determined through cross-validation.
[0070] Organize the storage device operation parameters into multi-dimensional data in time series and input them into the model. The model consists of a convolutional layer, a ConwLSTM layer, and a fully connected layer. The convolutional layer extracts local features, and the convolution operation formula Assume a 3×3 convolutional kernel, and the weights W of the convolutional kernel are initialized to random values, and the bias b is set to 0.1.
[0071] The ConVISIM layer extracts long-term dependencies, and the calculation formula is as follows:
[0072] i t = σ(W xi x t + W hi h t-1 + W ci c t-1 + b i )
[0073] f t = σ(W xf x t + W hf h t-1 + W cf c t-1 + b f )
[0074] c t = f t × c t-1 + i t × tanh(W xc x t + W hc h t-1 + b c )
[0075] o t = σ(W xo x t + W ho h t-1 + W co c t-1 + b o )
[0076] h t = o t × tanh(c t )
[0077] Among them, i t , f t , o t are the input gate, forget gate, and output gate respectively, c t is the cell state, h t is the hidden state, σ is the sigmoid function, and the fully connected layer maps the output of the ConwLSTM layer to the prediction result, and the model is trained by minimizing the mean square error loss function.
[0078] Using the BiLSTM-Attention model, preprocess the laws and regulations documents and operation process documents into word vector sequences. The BiLSTM layer processes forward and backward to obtain the forward hidden state and the backward hidden state and the backward hidden state The attention mechanism calculates the attention weight α t , and performs a weighted sum of the hidden states to obtain the document feature representation r. The formula is where W is the learnable weight matrix, T is the sequence length, and finally, a classifier is used to determine whether it meets the regulatory requirements.
[0079] Set a threshold according to the risk impact factor RI. For example, when the data leakage risk RI reaches 0.7, the system automatically sends a warning to the management personnel via text message, email, and instant messaging tool, including risk type, level, time and location information, generates a risk assessment report, and presents the detailed risk description, cause analysis, impact assessment, and development trend prediction in a visual form (such as a bar chart showing the proportion of different risk types, a line chart presenting the risk change trend over time, and a radar chart comparing the influence degrees of different risk factors). Adopt interactive visualization technology, and users can view the charts personalized.
[0080] Based on the expert knowledge base and policy library, the system generates countermeasures for different risks. For the data leakage risk, it is recommended to immediately freeze the relevant user accounts, strengthen network security protection measures, and encrypt the leaked data. For the storage device failure risk, start an emergency backup program, arrange equipment repair or replacement plans. For the regulatory compliance risk, conduct regulatory training, adjust management processes, and conduct internal audits. Introduce case-based reasoning technology, calculate the similarity between the current risk and historical cases, select similar cases to adjust countermeasures to generate the final suggestions. For example, after calculating the similarity between the current data leakage risk (high level, involving financial data, affecting multiple departments) and the cases in the case library, refer to the countermeasures of the reference cases (such as strengthening data encryption intensity, adding access permission review links) and adjust according to the actual situation, and push relevant suggestions to the management personnel and provide implementation steps and precautions.
[0081] Example 2
[0082] This example describes a government department responsible for managing various government affairs archives, including policy documents, resident information, and administrative approval records. The archives involve public interests and need to ensure data security and regulatory compliance. At the same time, the storage devices are aging and face the risk of failure, and effective monitoring of archive access and operations is required.
[0083] Deploy collectors at key nodes of the internal office network of the department to collect file access records (e.g., staff member C accessed resident information at 14:00:00 on June 11, 2023, the access type was query, and the identifier was 003), modification records (e.g., staff member D modified the policy document at 10:15:00 on June 5, 2023, the hash value was yyy, and the identity identifier was 004), storage device parameters (temperature 32°C, humidity 35%, disk space occupancy rate 55%, read / write speed 80MB / s, error rate 0.08%), system logs, and regulatory update information. When cleaning the data, determine the word vector similarity threshold to be 0.75 for text data through clustering analysis, and integrate it into the big data platform after cleaning.
[0084] Build a data leakage risk model. The state space considers the access user permissions (determined according to the position), historical access frequency (counted quarterly), and the sensitivity of the current access (classified according to the degree of public interest involved in the file). Use transfer learning, pre-train and then fine-tune to improve the accuracy of the model.
[0085] For the risk of storage device failure, build models for different storage devices (such as the hard disks of old servers and newly purchased storage devices) respectively, and adjust the hyperparameters of the ConvLSTM model. For example, set the convolution kernel size to 2x2 and the number of hidden units to 32 for old hard disks.
[0086] Build a regulatory compliance model, introduce a knowledge graph, and associate the government affairs operation process with the regulatory knowledge graph. For example, match the "administrative approval process" with relevant regulatory entities to enhance semantic understanding.
[0087] Monitor risks with real-time input data. If it is found that a large amount of resident information is frequently accessed within a certain period, analyze that the number of users involved is large and it affects public trust. For storage devices, if the error rate is detected to increase, evaluate that it may affect important approval record files, and calculate the risk factor through a sliding window (window 2 hours, step size 15 minutes).
[0088] When the risk impact factor reaches the set threshold (such as the regulatory compliance risk factor 0.7), send warnings to department leaders and relevant management personnel via text messages and emails, inform them of the details of the risk, generate a visual report, and display the comprehensive situation of different risk types with a radar chart for leaders to view interactively.
[0089] For the data leakage risk, it is recommended to suspend the relevant access permissions, check for network security vulnerabilities, encrypt the storage of resident information, use case-based reasoning, refer to similar government affairs data leakage cases, and push the adjusted suggestions, such as strengthening personnel security training measures.
[0090] The above are only the preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments with equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A big data electronic archive management method, characterized in that: The specific steps of this method are: S1. Data collection and integration Data collectors are deployed at each key node to collect data related to electronic archives in real time, including but not limited to archive access records, modification records, operating parameters of storage devices, system log files, legal and regulatory update information, and industry standard documents. For access records, the timestamp of each access is recorded. a , the unique identifier of the accessing user u id , access operation type identifier o t , the modification record contains the modification time t m , modify the hash value h of the content c , the identity of the person who modified it id , storage device operating parameters include temperature T, humidity H, disk space occupancy S o , read and write speed V r , error rate E r ; The collected multi-source heterogeneous data is cleaned to remove duplicate, erroneous and incomplete data. A cleaning algorithm based on similarity measurement is used to calculate the similarity between two data items d in the data set. i and d j , calculate its similarity Sim(d i ,d j ): Where n is the number of attributes of the data item, w k is the weight of the kth attribute. First, construct the attribute importance judgment matrix A and calculate its maximum eigenvalue λ max And the corresponding eigenvector W, normalize W to get w k , and then integrate the cleaned data into a unified big data platform. Using the data integration method of the graph model, the data from different data sources are represented as nodes, and the associations between the data are represented as edges. The data integration graph G = (V, E) is constructed, where V is the node set and E is the edge set. The data integration is achieved through the graph traversal algorithm. S2. Construction of risk factor identification model In view of the risk of data leakage, a risk identification model is constructed to define the state space S, which includes the permission level r of the current access user. l , historical access frequency f h 、The sensitivity of this visit s Features, action space A action Contains the judgment of normal access a n and abnormal access to a a , reward function R, when the judgment is correct, a positive reward is given, and when the judgment is wrong, a negative reward is given. The deep Q network (DQN) algorithm is used for model training. The network structure includes an input layer, multiple hidden layers and an output layer. The input layer receives the characteristics of the state space S. After the nonlinear transformation of the hidden layer, the Q value of each action is obtained in the output layer. During the training process, the network parameters θ are updated according to the reward feedback through continuous interaction with the environment. The update formula is: Among them, θ t represents the network parameters of the model at time step t, α is the learning rate, r is the reward value given by the reward function R, γ is the discount factor, Indicates the maximum Q value predicted by the model for all possible actions a′ in state s′, Q(s,a;θ t ) represents the Q value predicted by the model when taking action a in the current state s. is the Q function with respect to the network parameters θ t The gradient of , and its optimal value is determined by cross validation; For the risk of storage device failure, a prediction model is established. The operating parameters of the storage device are organized into multi-dimensional data in time series as the input of the model. The model consists of a convolutional layer, a ConvLSTM layer, and a fully connected layer. The convolutional layer is used to extract local features of the data. The calculation formula of the convolution operation is: in, is the feature map of the i-th row and j-th column of the l-th layer, is the convolution kernel weight, M represents the size of the convolution kernel in the vertical direction, m represents the index of the convolution kernel in the vertical direction, N represents the size of the convolution kernel in the horizontal direction, n represents the index of the convolution kernel in the horizontal direction, and b l is bias; The ConvLSTM layer is used to capture long-term dependencies in time series data. Its calculation formula is: i t =σ(W xi x t +W hi h t-1 +W ci c t-1 +b i ) f t =σ(W xf x t +W hf h t-1 +W cf c t-1 +b f ) c t =f t ×c t-1 +i t ×tanh(W xc x t +W hc h t-1 +b c ) o t =σ(W xo x t +W ho h t-1 +W co c t-1 +b o ) h t =o t ×tanh(c t ) Among them, i t 、f t , o t They are input gate, forget gate and output gate, c t is the cell state, h t is the hidden state, σ is the sigmoid function, W xi , W hi , W ci They are the weight matrices connecting the input gate with the current input, the hidden state at the previous moment, and the cell state at the previous moment, respectively. t The input data at the current moment, h t-1 is the hidden state at the previous moment, c t-1 is the cell state at the last moment, W xc , W hc are the weight matrices connected to the current input and the previous hidden state, respectively, i , b f , are the bias of the input gate and the bias of the forget gate, respectively, W xf , W hf , W cf are the weight matrices connecting the forget gate with the current input, the hidden state at the previous moment, and the cell state at the previous moment, tanh is the hyperbolic tangent function, and b c The bias used to calculate the current cell state, W xo , W ho , W co are the weight matrices connecting the output gate with the current input, the hidden state at the previous moment, and the cell state at the previous moment, respectively. o The fully connected layer is the bias of the output gate, and the output of the ConvLSTM layer is mapped to the prediction result. The model is trained by minimizing the mean square error loss function; In response to regulatory compliance risks, the BiLSTM-Attention model based on the attention mechanism is used to pre-process legal and regulatory documents and operation process documents in the electronic archive management system and convert them into word vector sequences. The BiLSTM layer processes the word vector sequence from the forward and reverse directions to obtain the forward hidden state. and the backward hidden state Concatenate to get hidden state The attention mechanism calculates the attention weight α for each hidden state t , perform weighted summation on the hidden states to obtain the feature representation r of the document: Among them, the formula calculates the hidden state h i and h j The correlation between ij Used to calculate the attention weight α j , T represents the transpose operation of the matrix; Among them, W iearnable is a learnable weight matrix, L seq is the length of the sequence, exp is the mathematical representation of the exponential function, e ik Represents the hidden state h i With the hidden state h k The correlation score between them, k is an index variable, and finally the feature representation of the document is classified by the classifier to determine whether it meets the regulatory requirements; S3. Real-time risk monitoring and analysis The integrated data is input into the risk factor identification model constructed above in real time to monitor the data leakage risk, storage device failure risk and regulatory compliance risk in the electronic archive management process in real time. For the identified potential risks, the risk analysis module is started. For the data leakage risk, the sensitivity of the leaked data, the number of users involved and the business field are analyzed. For the storage device failure risk, the number and importance of archives that may be affected by the failure are evaluated. For the regulatory compliance risk, the specific provisions of the violation and the degree of possible penalties are determined. The risk level is quantified by calculating the risk impact factor RI. For the data leakage risk: Among them, RI d represents the risk impact factor of data leakage risk, n is the number of factors affecting data leakage risk, is the impact of the i-th influencing factor on the risk of data leakage, is its weight; For storage device failure risks: Among them, RI s represents the risk influencing factor of storage device failure risk, m is the number of factors that affect the storage device failure risk, is the weight of the jth influencing factor on the risk of storage device failure, is the influence degree of the jth influencing factor on the risk of storage device failure; For regulatory compliance risks: Among them, RI r represents the risk impact factor of regulatory compliance risk, p is the number of factors affecting regulatory compliance risk, is the weight of the kth influencing factor on regulatory compliance risk, is the degree of influence of the kth influencing factor on regulatory compliance risk; S4. Intelligent early warning and risk assessment report generation According to the risk impact factor RI obtained from risk analysis, thresholds of different risk levels are set. When RI reaches the corresponding threshold, the system automatically issues warning information, which is sent to relevant managers through SMS, email, and instant messaging tools. The warning information includes detailed information on risk type, risk level, time and location of risk occurrence. At the same time, the system generates a risk assessment report, which includes a detailed description of the risk, analysis of the cause of the risk, assessment of the impact of the risk on the electronic archive management system and business, and prediction of the risk development trend. The report is presented in a visual form, using bar charts, line charts, radar charts and other charts to display risk-related information; S5. Generation and delivery of response suggestions For different types and levels of risks, the system generates corresponding response suggestions based on the pre-established expert knowledge base and policy library. For data leakage risks, suggestions include immediately freezing relevant user accounts, strengthening network security protection measures, and encrypting leaked data. For storage device failure risks, suggestions include starting emergency backup procedures and arranging equipment repair or replacement plans. For regulatory compliance risks, suggestions include conducting regulatory training, adjusting management processes, and conducting internal audits. The generated response suggestions will be pushed to relevant managers, and the implementation steps and precautions of the suggestions will be provided.
2. A big data electronic archive management method according to claim 1, characterized in that: In the data collection and integration step, when cleaning the data from different sources, the word vector similarity cleaning method is used for the text data contained in the data set. For the two sentences S1 and S2 in the text data, they are converted into word vector representations. and The similarity of sentences is measured by calculating cosine similarity: in, is the dot product of the vectors, and are the moduli of the vectors, and for similarities above the set threshold θ t The text data is considered to be repeated data and cleaned, and the threshold θ t By performing cluster analysis on historical text data, in practical applications, K-means clustering is performed on a large amount of historical text data, the similarity distribution between different clusters is analyzed, and the appropriate threshold is selected to ensure that duplicate text data can be effectively removed without accidentally deleting valuable information.
3. A big data electronic archive management method according to claim 1, characterized in that: When constructing the data leakage risk identification model, in order to improve the generalization ability and accuracy of the model, the transfer learning method is adopted. In the pre-training stage, the model is pre-trained using large-scale public network access data to learn the general network access behavior pattern. Then, in the fine-tuning stage, the pre-trained model is fine-tuned using the electronic archive access record data within the enterprise. In the fine-tuning process, the parameters of some pre-training layers are fixed, and only the parameters of some layers are updated. Suppose the pre-trained model is M p , whose parameter is θ p , the parameter set updated during fine-tuning is θ f , for input data x, the fine-tuned model output y is: y=M p (x;θ p ,i f ) Among them, θ f By minimizing the loss function L on the enterprise's internal data f To update: in, is the true label, is the prediction result of the model, N is the number of samples of internal enterprise data, and the common features learned from public data are learned through transfer learning.
4. A big data electronic archive management method according to claim 1, characterized in that: When constructing the storage device failure risk prediction model, the differences in operating characteristics of different types of storage devices are taken into consideration, and the model is trained individually. For different types of storage devices, their historical operating parameter data and fault records are collected respectively. For each type of storage device, an independent ConvLSTM model is constructed. During the model training process, the convolution kernel size and the number of hidden units in the ConvLSTM layer are adjusted to adapt to the characteristics of different devices.
5. A big data electronic archive management method according to claim 1, characterized in that: When constructing the regulatory compliance risk identification model, in order to better understand the semantic information in the legal and regulatory documents, an external knowledge graph is introduced. The knowledge graph contains entities, relationships and attribute information related to laws and regulations. The entities in the operation process documents and legal and regulatory documents in the electronic archive management system are matched with the entities in the knowledge graph. Through the relational reasoning of the knowledge graph, the semantic representation of the document is enriched. For entity e in the document, its neighbor entity set in the knowledge graph is N(e). The representation of entity e is updated by aggregating the information of neighbor entities: Among them, w n is the weight of neighbor entity n, which is determined by calculating the strength of the relationship between entity e and neighbor entity n. By introducing an external knowledge graph, the model's ability to understand and judge regulatory compliance risks can be enhanced.
6. A big data electronic archive management method according to claim 1, characterized in that: In the real-time risk monitoring and analysis step, in order to timely discover the dynamic changes of risks, a sliding window technology is used to set a fixed-size time window T for each risk type. w , collect and analyze data in real time within the window. As time goes by, the window keeps sliding, each time sliding a time step Δt. In each window, the risk impact factor RI is recalculated. For data leakage risk, in window T w Collect access records within a certain period of time, calculate the frequency and impact of abnormal access in different time periods, and recalculate RI based on this information d , through the sliding window technology, it is possible to capture the dynamic changing trend of risks in a timely manner.
7. A big data electronic archive management method according to claim 1, characterized in that: In the step of generating the intelligent early warning and risk assessment report, interactive visualization technology is used for the visual display of the risk assessment report, and the user can personalize the charts in the report through interactive operations.
8. A big data electronic archive management method according to claim 1, characterized in that: In the generation and push of response suggestions, in order to improve the effectiveness and operability of response suggestions, case reasoning technology is introduced. When generating response suggestions, the system first searches the case library for historical cases similar to the current risk situation. Each case in the case library contains risk description, response measures and implementation effect information. By calculating the similarity between the current risk and the historical case, the case with the highest similarity is selected as a reference. For the calculation of similarity, a method based on multi-attribute decision-making is adopted, which comprehensively considers multiple attributes such as risk type, risk level, and impact range. Suppose the attribute vector of the current risk is The attribute vector of the historical case is Similarity for: Where n represents the number of attributes considered when calculating similarity, and w i is the weight of the i-th attribute, Similarity i (a i ,b i ) is used to calculate the similarity between the i-th attribute in the current risk attribute vector a and the i-th attribute in the historical case attribute vector b, a i The i-th attribute value in the current risk attribute vector a, b i The i-th attribute value in the historical case attribute vector b is adjusted according to the response measures of the reference case and the current actual situation to generate the final response suggestion. Through case reasoning technology, historical experience can be learned.
9. A big data electronic archive management method according to claim 1, characterized in that: In order to ensure the stability and reliability of the entire big data electronic archive management system, a distributed architecture is adopted to distribute the data collection, storage, analysis and early warning function modules on multiple nodes, and communicate and collaborate through the network. In terms of data storage, a distributed file system is adopted to store the electronic archive data in multiple storage nodes. In terms of data processing, a distributed computing framework is adopted to decompose large-scale data processing tasks into multiple subtasks, which are distributed on multiple computing nodes for parallel processing. At the same time, by setting up master nodes and slave nodes, the system load balancing and fault recovery are achieved. When a node fails, the master node can automatically reallocate tasks to other normal nodes. The distributed architecture can effectively cope with the high concurrency and large-scale data processing requirements in the big data environment.
Citation Information
Patent Citations
Power-supply-enterprise electronic file safety risk evaluation system
CN105205581A
Compliance inspection system and method for enterprise business process
CN119378993A
Cited By
Cultural industry digital monitoring system and method based on machine learning
CN121543114A