Archive data security integration management system

Through the methods of archival data collection, intelligent knowledge analysis and dynamic construction of knowledge graphs, the existing system's insufficient security management and situation assessment are solved, and the efficient, security management and real-time situation assessment of archival data are achieved, which enhances the security and adaptability of the system.

CN120257327AInactive Publication Date: 2025-07-04CHINA SHENHUA ENERGY CO LTD SHENDONG COAL BRANCH

Patent Information

Application Number
CN202510398469.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing archive data security integration management system has shortcomings in security management and security situation assessment, and it is difficult to detect and respond to advanced persistent threats and zero-day vulnerabilities in a timely manner, and the security situation assessment is not comprehensive and real-time enough.

Method used

The archive data acquisition module, knowledge intelligent analysis module, knowledge graph dynamic construction module and security situation evaluation module are adopted to collect metadata and operation behavior data in a structured manner, and the knowledge graph with a confidential label is dynamically built. The improved PageRank algorithm is used to quantify security influence, dynamically divide security domains, and the identity authentication and bandwidth allocation of cross-domain communication is realized through the SM4 algorithm.

Benefits of technology

It realizes comprehensive and accurate collection and management of archival data, dynamically manages data quality, evaluates security situations in real time, promptly detects potential threats, and enhances the security of cross-domain communications and the flexibility and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257327A_ABST
    Figure CN120257327A_ABST
Patent Text Reader

Abstract

The invention discloses an archive data security integration management system, which relates to the technical field of knowledge maps and comprises an archive data acquisition module, a knowledge intelligent analysis module, a knowledge map dynamic construction module, a security situation evaluation module and a security communication center module. The data acquisition module structurally acquires metadata and operation behaviors, the intelligent analysis module dynamically governs data quality, the map construction module constructs a three-dimensional security model, the situation evaluation module quantifies security influence and generates a strategy, and the communication center module realizes dynamic identity authentication and bandwidth allocation. Through the knowledge graph technology, comprehensive integration and safety management of the archive data are achieved, the data quality and utilization efficiency are improved, dynamic evaluation and coping with safety threats are achieved, and the safety and integrity of the archive data are ensured. Meanwhile, through dynamic identity authentication and bandwidth allocation, the safety and efficiency of cross-domain communication are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graphs, and particularly to an archival data security integration management system. Background Art

[0002] An archival data security integration management system is a comprehensive information management system designed to achieve the full life cycle management of archives and ensure the security and integrity of archival data. This system uses modern information technology means such as computer technology, database technology, and multimedia technology to centrally manage, scientifically classify, orderly store, and efficiently utilize various types of archival information. It covers all aspects of archival entry, sorting, storage, retrieval, utilization, archiving, and destruction. Through digital and information-based methods, it improves the security, integrity, availability, and reliability of archival information. In an archival data security integration management system, the security and integration of data are highly emphasized. The system ensures that only authorized personnel can access and operate relevant archival information by adopting means such as multiple identity authentication technologies, data backup and recovery mechanisms, and permission management mechanisms. At the same time, the system also has data integration capabilities, capable of integrating and uniformly managing archival data from different sources and in different formats, facilitating cross-library retrieval and utilization by administrators. Generally speaking, an archival data security integration management system is an efficient, secure, and convenient archival management tool, which is of great significance for improving the archival management efficiency and security of enterprises, organizations, or individuals.

[0003] To address the deficiencies in the security management and security posture assessment of archival data security integration management systems, the existing technologies mainly adopt traditional security control measures, such as encrypted storage, access control lists, and regular security audits for processing. These measures can, to a certain extent, ensure the basic security of archival data and prevent unauthorized access and data leakage. However, in practice, there are still some new security threats and vulnerabilities that are difficult to be discovered and addressed in a timely manner. For example, some advanced persistent threats and zero-day vulnerabilities may bypass traditional security protection measures, leading to the risk of illegal acquisition or tampering of archival data. In addition, the existing systems also have limitations in security posture assessment, making it difficult to comprehensively and real-time grasp the security status of the system, and thus unable to take effective countermeasures in a timely manner. Therefore, in order to improve the security and management efficiency of archival data security integration management systems, an archival data security integration management system is proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide an archival data security integration management system to solve the problems raised in the above background art.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is: an archival data security integration management system, including an archival data collection module, a knowledge intelligent parsing module, a knowledge graph dynamic construction module, a security situation assessment module, and a security communication central module; The archival data collection module structurally collects the metadata and operation behavior data of the archival system; The knowledge intelligent parsing module calculates the text complexity level based on the metadata of the archival system, and preprocesses and extracts features from the metadata of the archival system; The knowledge graph dynamic construction module constructs a knowledge graph with confidentiality level tags, embeds archival security attributes into the nodes of the knowledge graph, and optimizes the semantic representation and security relevance of the knowledge graph based on the text complexity level; The security situation assessment module quantifies the security influence of the nodes of the knowledge graph by improving the PageRank algorithm, identifies high-sensitivity archives, dynamically divides the security domain and generates a least privilege policy based on the knowledge graph and the text complexity level; The security communication central module binds the confidentiality level tags in the knowledge graph through the SM4 algorithm, performs dynamic identity authentication for cross-domain communication, calculates the comprehensive sensitivity based on the text complexity level, and dynamically constructs a bandwidth allocation decision tree.

[0006] A further improvement of the technical solution of the present invention is that in the archival data collection module, the process of structurally collecting the metadata and operation behavior data of the archival system includes: Adopt a dual-machine hot standby deployment mode, deploy a database protocol converter and a log collection terminal, connect to the archival management system database through the JDBC / ODBC interface, and capture structured metadata and operation logs in real time; Through the database interface polling mechanism, periodically scan the system tablespace to obtain the metadata of the archival system, generate structured records according to the preset field template containing the file identifier and storage location, and for unstructured metadata, use regular expressions to match the key attribute fields and convert them into the attribute name: attribute value format; Deploy a log parsing engine on the database server side of the archival management system, extract operation behavior data through online log scanning and archived log backtracking technologies, the parsing engine is built with a transaction recognition algorithm, reorganize the original log stream according to the timestamp - operation subject - behavior type structure, deploy a national secret SM4 encryption gateway between the collection terminal and the security communication central module, perform shard encryption on the transmission fields, use a distributed message queue to implement multi-source data buffering, and set up a primary and backup transmission channel; Analyze transaction records in the online logs of the database in real time, identify behavior categories through the operation type coding mapping table, adopt an incremental capture strategy, locate uncollected records based on the log sequence number, and retain the previous image data for deletion operations; The archive system metadata and operation behavior data are aligned in time and space, with the file identifier as the primary key, and the operation behavior record data is added to the corresponding metadata entry. Normalization is performed, and heterogeneous time formats are uniformly converted into system standard time codes. The structured data packets are graded according to field sensitivity and encrypted in pieces using the SM4 algorithm. Highly sensitive fields containing permission attributes and operation subjects are encrypted using independent keys, and low-sensitivity fields containing timestamps and operation types are encrypted using group keys. The encrypted data packets are asynchronously pushed to the secure communication hub module through the publish-subscribe mode of the message queue, and the breakpoint resume mechanism is automatically triggered when the transmission fails.

[0007] A further improvement of the technical solution of the present invention is that: in the knowledge intelligent parsing module, the text complexity level is calculated based on the archive system metadata, and the process of preprocessing the archive system metadata includes: The total number of characters in the archive description text is divided by the number of text paragraphs, and the ratio of the number of entities identified in the text to the total number of characters in the text is calculated, and then multiplied by 1000 to convert it into the number of thousand-word entities. The number of entities includes names of people, institutions and places. The text complexity level L is divided into 1, 2 and 3 through linear weighted comprehensive text length and semantic density. When L=1, it is judged as a simple text, and when L≥2, it is judged as a medium-to-high complexity text; When L=1, the sliding window mean is used to fill in the local missing values, the window width is set to 5 adjacent sample points in the time series, the standard deviation threshold detection is performed, and abnormal records that deviate from the mean by more than 3 times the standard deviation are eliminated. The bidirectional long short-term memory network and conditional random field joint model are used to capture the basic entity relationship through the time series features, and the keyword weight is calculated based on the improved TF-IDF algorithm; When L ≥ 2, time series prediction is used to fill missing values, modeling is based on the periodic characteristics of historical data, dynamic cutting of isolation forests is enabled, outliers in complex distributions are adaptively identified through tree depth, pre-trained language models and conditional random field joint models are loaded, and deep semantic understanding capabilities are used to parse professional terms and nested structures in archives. Graph embedding vectors and text semantic vectors are fused, cross-modal features are extracted through convolution operations, and distributed parallel computing is enabled.

[0008] A further improvement of the technical solution of the present invention is that: in the knowledge intelligent parsing module, the process of performing feature extraction according to the text complexity level includes: The feature levels are divided into basic features, semantic features, and topological features. The basic features include text length, entity density, and keyword frequency. The semantic features include relationships between entities and context-dependent structures. The topological features include connection attributes of entity nodes in the knowledge graph; When L = 1, an improved term frequency-inverse document frequency algorithm is used to calculate keyword weights. By suppressing the weights of high-frequency general words, the discrimination of professional entity words is improved. Based on the bidirectional long short-term memory network architecture, text sequence features are extracted. Its forward network encodes context information in character order, and the reverse network captures suffix dependency relationships in reverse order, outputting forward and backward hidden state vectors of character positions. A transition probability matrix is introduced to calculate the optimal entity label sequence and generate entity relationship triples; When L ≥ 2, text semantic vectors are generated through a multi-layer self-attention mechanism. The input text is segmented into words and converted into embedding vectors. After passing through 12 layers of Transformer encoders to extract global semantics, context-related vector representations of words are output. The graph embedding vectors in the knowledge graph are concatenated with the text semantic vectors, and feature crossing is performed through convolutional kernels to form fused features. Multi-scale convolutional kernels are used to perform sliding window calculations on the fused features. Each convolutional kernel generates a feature map, and significant features are extracted through max pooling. The fully connected layer outputs the final feature vectors. Mutual information analysis is performed on the extracted feature vectors to calculate the mutual information values between each feature dimension and the target entity label. Feature dimensions with mutual information below 0.05 are removed. A piecewise normalization strategy is adopted, with min-max normalization for statistical features and L2 norm unit normalization for semantic vectors; The feature vectors are encapsulated into JSON objects according to the entity-relationship-attribute structure. The feature transmission priority is set according to the L value. When L = 1, it is marked as a normal priority with a weight of 0.5. When L ≥ 2, it is marked as a high priority with a weight of 0.8 and enters the graph construction queue first.

[0009] A further improvement of the technical solution of the present invention lies in: in the knowledge graph dynamic construction module, a knowledge graph with a confidentiality level label is constructed. The process of embedding the file security attribute into the knowledge graph nodes includes: Define the basic attributes of the file entity. The basic attributes of the file entity include file identifier, storage location, confidentiality level label, and responsible department. Among them, the confidentiality level label is divided into top secret, confidential, and internal, describing the interaction relationships between entities. The interaction relationships between entities include access relationships, file derivation links, and revision histories. Bind the basic permission attributes and the confidentiality level label to the corresponding nodes in the knowledge graph to form a node-confidentiality level mapping table; Taking the collected archival system metadata as the basic attributes of nodes, and converting the operation behavior data in the operation logs into relational edges. The types of relational edges include access, copy, and authorization. Based on the entity-relationship triples output by the knowledge intelligent parsing module, a cross-archival semantic association network is constructed, the security levels of the relationship types are labeled, and the classification labels are stored as node attribute fields. Through embedding space constraints, high-classification nodes are connected to edges of specific relationship types, and the relationship paths across security domains meet the minimum classification transition conditions; Check the completeness of the security attributes of nodes. If there are missing attribute nodes, they are prohibited from participating in relationship reasoning. Verify the classification compliance of the relationship paths. If the classification of the head node is confidential, the classification of the tail node must not be lower than confidential, otherwise a security alert is triggered. When a permission change operation is detected, use the graph traversal algorithm to locate the nodes affected by the permission change operation, recalculate their embedding vectors and security distances, and synchronize the real-time security status of the knowledge graph.

[0010] A further improvement of the technical solution of the present invention lies in: in the knowledge graph dynamic construction module, the process of optimizing the semantic representation and security relevance of the knowledge graph based on the text complexity level includes: According to the text complexity level L value, evaluate the semantic relationship complexity of the knowledge graph. If L = 1, it is considered that the relationship between entities is mainly linear one-to-one. If L ≥ 2, it is considered that there are multi-hop asymmetric relationships between entities; When L = 1, adopt the translational embedding model, establish a simple security association of head entity + relationship ≈ tail entity through linear transformation, and calculate the loss function. When L ≥ 2, adopt the rotational embedding model, model complex paths through the Hadamard product formula in the complex space, support asymmetric and reverse relationship reasoning, and capture multi-hop security dependencies; In the process of vectorizing the knowledge graph, define a security distance threshold, construct a classification constraint space, use the classification label as an embedding space constraint condition, allow internal-level nodes to be distributed in a larger common subspace, limit confidential-level nodes to an independent area, keep the minimum distance from the common subspace, and completely isolate top-secret-level nodes, only allowing them to be connected through specific relational edges. When training the embedding model, superimpose a security constraint loss term, and the security constraint loss term includes semantic loss and security violation penalty; When an abnormal cross-classification access frequency is detected, automatically increase the security coefficient and strengthen isolation, and restore the basic value during low-risk periods. Before the change of the L value, pre-load the parameters of the target embedding model into the memory, adopt dual-model parallel reasoning, gradually switch 10% of the traffic to the new model, and perform a full-scale switch after verifying the stability. If the inference error rate of the new model exceeds 5%, automatically roll back to the previous stable version; For the generated entity-relationship triples, verify whether the confidentiality levels of the head and tail nodes meet the security policy of the relationship type. If the relationship type is replication, the confidentiality level of the tail node must not be higher than that of the head node. If the relationship type is authorization, the confidentiality level of the head node must be at least one level higher than that of the tail node. Scan the embedding space regularly to detect abnormal neighboring nodes. For cross-confidentiality node pairs with a distance less than the security threshold, trigger manual review and automatically correct the vector coordinates of the illegal nodes to return them to the compliant area.

[0011] A further improvement of the technical solution of the present invention is that in the security situation assessment module, by improving the PageRank algorithm, the security influence of the knowledge graph nodes is quantified, and the process of identifying highly sensitive files includes: Quantify the confidentiality labels of knowledge graph nodes into security weight coefficients, set the weight of top-secret nodes to 3, the weight of confidential nodes to 2, and the weight of internal nodes to 1, define the basic influence value of nodes, set the propagation attenuation rate according to the edge type, set the attenuation rate of cross-confidentiality replication to 0.3, the attenuation rate of same-confidentiality access to 0.7, and the attenuation rate of internal authorization to 0.9; Introducing security parameters into the PageRank algorithm, the security parameters include confidentiality difference, access frequency and complexity attenuation factor; Real-time monitoring of cross-level access frequency. When a single node has >50 cross-level accesses per hour, the attenuation rate is updated and the security weight is updated every 6 hours. Calculate the global node influence mean, standard deviation and high sensitivity threshold, mark nodes whose influence value exceeds the high sensitivity threshold in real time, generate a high-risk list, perform reverse graph traversal on the marked nodes, identify the key nodes in all their incoming edge paths, and if the path contains ≥2 top-secret nodes, mark it as a core risk source. If the path crosses more than 3 confidentiality-level areas, mark it as a cross-domain transmission hotspot. Automatically add temporary access restrictions for highly sensitive nodes, prohibit cross-level read operations, force the enablement of two-factor authentication writes, re-shard and store node-related data by security domain, and physically isolate high-risk data.

[0012] A further improvement of the technical solution of the present invention is that in the security situation assessment module, the process of dynamically dividing the security domain and generating the minimum privilege policy based on the knowledge graph and the text complexity level includes: Calculate the number of inbound and outbound edges of each node and the average confidentiality difference of adjacent nodes, identify high connection density areas and isolated nodes, set path weights according to the relationship edge type, cross-confidentiality path weight = basic value × confidentiality difference, same-confidentiality path weight = basic value, if L = 1, use spectral clustering algorithm to divide security domains, number of eigenvectors = 5, Laplace matrix regularization parameter is fixed to 0.1, if L ≥ 2, dynamically adjust regularization parameters, enhance the cutting accuracy of complex relationship networks, generate security domains including core domains, transition domains and public domains, where the core domain is a connected subgraph containing ≥ 2 top-secret nodes, set independent physical storage partitions, transition domains are areas containing confidential nodes and have access relations with the core domain, enable logical isolation, and public domains are low-risk areas containing internal nodes, allowing cross-domain reading but restricting writing; Statistics are collected on the access frequency, operation type and cross-domain ratio of nodes in the past 72 hours, behavioral baselines are generated, risk scores are calculated and risk thresholds are set, and basic permissions are predefined according to the node classification level. If the risk score is greater than the risk threshold, it is identified as a high-risk node, cross-domain writing is prohibited, and copy operations are restricted. If it is a core domain node, two-factor authentication is mandatory, and operation logs are fully audited. During the operation, only the minimum permissions required to complete the current operation are granted; The security domain division is recalculated every 6 hours. If the average risk score of regional nodes increases by 15%, domain splitting is triggered, and the region is divided into sub-domains according to the difference in security levels. If the interaction frequency between adjacent domains is lower than the risk score threshold, they are merged into a single domain to reduce management overhead. The frequency of cross-domain requests of a single node is detected in real time. If the cross-domain request is >20 times within 1 minute, the circuit breaker mechanism is triggered to suspend the cross-domain permission of the node for 10 minutes. After the circuit breaker is restored, the permission is reactivated through administrator approval. The generated permission policy is simulated and deduced, and a virtual operation chain is constructed to verify whether there is a path to bypass the permission constraint through a legitimate operation combination. If a vulnerability is found, an operation chain length limit policy is added.

[0013] A further improvement of the technical solution of the present invention is that in the secure communication hub module, the process of dynamically authenticating the identity of cross-domain communication by binding the confidentiality level labels in the knowledge graph through the SM4 algorithm includes: The confidentiality labels of nodes in the knowledge graph are used as encryption factors, and a dynamic session key is generated through the key derivation function of the SM4 algorithm. The confidentiality levels are quantified according to top secret = 3, confidential = 2, and internal = 1. When a cross-domain communication request is initiated, the security attributes of the requesting node are extracted, including the security domain to which it belongs and the access permission list. The requesting party identifier, target domain identifier, and operation type are encrypted using the session key to generate an authentication token. Intercept the header metadata of the data packet during the communication process, extract the encrypted token and the target domain identifier, query the security policy of the target domain through the knowledge graph, the security policy includes the scope of confidentiality allowed to be accessed and the list of authorized departments, use the metadata of the requesting node to recalculate the session key, decrypt the token to obtain the original request information, verify whether the decrypted operation code is in the set of allowed operations of the target domain, and satisfy the confidentiality level of the requesting party is higher than the minimum confidentiality level of the target domain, and the requesting party department belongs to the list of authorized departments of the target domain. If the verification is successful, use the new session key to encrypt the communication content, set the key validity period to 15 minutes, if the permissions do not match and the token decryption fails, then record the security event and generate an alarm, and freeze the cross-domain communication permissions of the requesting node for 30 minutes; When the confidentiality level and permissions of a node in the knowledge graph change, an update event is automatically pushed to the communication hub, immediately invalidating all tokens issued by the node and forcing re-authentication. If the same node initiates more than 10 cross-domain requests within 1 minute, its confidentiality weight is automatically reduced. The updated authentication rules are synchronized to all communication nodes within 50 milliseconds.

[0014] A further improvement of the technical solution of the present invention is that: in the secure communication hub module, the process of calculating the comprehensive sensitivity based on the text complexity level and dynamically constructing the bandwidth allocation decision tree includes: The basic sensitivity value is set according to the confidentiality label of the node in the knowledge graph, where top secret node = 3, confidential node = 2, and internal node = 1. The dynamic risk score is superimposed to calculate the comprehensive sensitivity. When L = 1, it is marked as a regular transmission task, and when L ≥ 2, it is marked as a critical transmission task. Construct a decision tree and divide the main branches of the decision tree according to the L value. If L=1, it is divided into the left branch and enters the conventional transmission strategy to determine whether the comprehensive sensitivity of the node is ≥2. If the comprehensive sensitivity is ≥2, 1.2 times of the basic bandwidth is allocated and redundancy check is enabled. If the comprehensive sensitivity is <2, 0.8 times of the basic bandwidth is allocated and the number of retransmissions is limited to ≤3. If L≥2, it is divided into the right branch and enters the key transmission strategy to determine whether the node is located in the core security domain. If the node is located in the core security domain, the allocated bandwidth = total bandwidth × comprehensive sensitivity / global sensitivity. If the node is not located in the core security domain, the allocated bandwidth = remaining bandwidth × comprehensive sensitivity / non-core sensitivity sum; Monitor the current network throughput and dynamically correct bandwidth allocation. When a node fails to transmit for three consecutive times, downgrade its sensitivity weight. If the weight is lower than the threshold, suspend the node transmission and trigger an alarm. Regenerate the decision tree every 2 hours, count the transmission success rate of each branch, prune branches with a success rate lower than 85%, merge branches with similar strategies, and reduce the depth of the decision tree. If L≥2 and the sensitivity≥2, it monopolizes 40% of the bandwidth, allows preemption of low-priority resources, the transmission interval≤100ms, and enables end-to-end encryption. If L≥2 and the sensitivity≥2 are not satisfied, it shares the remaining bandwidth, adopts round-robin scheduling, extends the transmission interval to 500ms, implements fragmentation transmission for highly sensitive data streams, and if the top-secret data fragment≤64KB, it attaches a hash checksum. If the confidential data fragment≤128KB, it enables breakpoint resumption.

[0015] Due to the adoption of the above technical solutions, the technical progress achieved by the present invention compared with the prior art is as follows: 1. The present invention provides an archival data security integration management system. Through the efficient acquisition technology of the archival data acquisition module, it can comprehensively and accurately obtain the metadata and operation behavior information of the archival system, providing a solid foundation for subsequent security management and data governance, ensuring the integrity and accuracy of the data, and effectively avoiding security management loopholes caused by data loss or errors.

[0016] 2. The present invention provides an archival data security integration management system. The knowledge intelligent parsing module can dynamically select preprocessing methods and entity recognition models according to the complexity levels of different texts, realizing the dynamic governance of data quality. It not only improves the efficiency and accuracy of data processing, but also can perform customized processing for different types of archival data, further enhancing the flexibility and adaptability of the system.

[0017] 3. The present invention provides an archival data security integration management system. The security situation assessment module can quantify the security influence of the nodes in the knowledge graph and dynamically divide the security domain by improving the PageRank algorithm, generating a least privilege policy, enabling the system to perceive and evaluate the security situation of archival data in real time, timely discover potential security threats, and take effective countermeasures, thus greatly improving the security of archival data. At the same time, the secure communication central module realizes the dynamic authentication of identities for cross-domain communication by binding the SM4 algorithm with the knowledge graph metadata, further enhancing the communication security of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. Example, such as Figure 1 As shown, the present invention provides an archival data security integration management system, including an archival data collection module, a knowledge intelligent parsing module, a knowledge graph dynamic construction module, a security situation assessment module, and a security communication central module; The archival data collection module structurally collects the metadata and operation behavior data of the archival system, adopts a dual-machine hot standby deployment mode, deploys a database protocol converter and a log collection terminal, connects to the archival management system database through the JDBC / ODBC interface, captures structured metadata and operation logs in real time, periodically scans the system tablespace through the database interface polling mechanism to obtain the archival system metadata, generates structured records according to a preset field template containing file identifiers and storage locations, for unstructured metadata, uses regular expressions to match key attribute fields and converts them into the attribute name: attribute value format, deploys a log parsing engine on the archival management system database server side, extracts operation behavior data through online log scanning and archived log backtracking technologies, the parsing engine is built with a transaction recognition algorithm, reorganizes the original log stream according to the timestamp - operation subject - behavior type structure, deploys a national secret SM4 encryption gateway between the collection terminal and the security communication central module, shards and encrypts the transmission fields, uses a distributed message queue to implement multi-source data buffering, and sets up a primary and backup transmission channel, parses the transaction records in the database online log in real time, identifies the behavior category through the operation type coding mapping table, adopts an incremental capture strategy, locates the uncollected records based on the log sequence number, for deletion operations, retains the pre-image data, aligns the archival system metadata and operation behavior data in time and space, uses the file identifier as the primary key, records the operation behavior data to the corresponding metadata entry, performs normalization processing, uniformly converts the heterogeneous time formats into the system standard time code, classifies the structured data packets according to field sensitivity, shards and encrypts them using the SM4 algorithm, encrypts the highly sensitive fields containing permission attributes and operation subjects with independent keys, and encrypts the low-sensitive fields containing timestamps and operation types with group keys, and asynchronously pushes the encrypted data packets to the security communication central module through the publish-subscribe mode of the message queue, and automatically triggers the breakpoint resumption mechanism when the transmission fails; The knowledge intelligent parsing module calculates the text complexity level based on the archive system metadata, preprocesses and extracts features from the archive system metadata, counts the total number of characters in the archive description text divided by the number of text paragraphs, calculates the ratio of the number of entities identified in the text to the total number of characters in the text, and then multiplies it by 1000 to convert it into the number of thousand-word entities. The number of entities includes names of people, institutions and places. The text complexity level L is divided into 1, 2 and 3 through linear weighted comprehensive text length and semantic density. When L=1, it is judged as a simple text. When L≥2, it is judged as a medium-to-high complexity text. When L=1, the sliding window mean is used to fill in the local missing values. The window width is set to 5 adjacent sample points in time sequence. The labeling is performed. Quasi-error threshold detection, remove abnormal records that deviate from the mean by more than 3 standard deviations, use a bidirectional long short-term memory network and a conditional random field joint model, capture basic entity relationships through time series features, calculate keyword weights based on the improved TF-IDF algorithm, and use time series prediction to fill in missing values ​​when L ≥ 2. Model based on the periodic characteristics of historical data, enable dynamic cutting of isolated forests, identify outliers in complex distributions through tree depth adaptation, load pre-trained language models and conditional random field joint models, use deep semantic understanding capabilities to parse professional terms and nested structures in archives, fuse graph embedding vectors and text semantic vectors, extract cross-modal features through convolution operations, enable distributed parallel computing, and divide features. The hierarchy is basic features, semantic features and topological features. The basic features include text length, entity density and keyword frequency. The semantic features include the relationship between entities and the context dependency structure. The topological features include the connection attributes of entity nodes in the knowledge graph. When L=1, the improved word frequency-inverse document frequency algorithm is used to calculate the keyword weight. By suppressing the weight of high-frequency general words, the discrimination of professional entity words is improved. The text sequence features are extracted based on the bidirectional long short-term memory network architecture. The forward network encodes the context information in character order, and the reverse network captures the suffix dependency in reverse order. The forward and reverse hidden state vectors of the character position are output. The transition probability matrix is ​​introduced to calculate the optimal entity label sequence and generate entity relationship ternary. When L≥2, a text semantic vector is generated through a multi-layer self-attention mechanism. The input text is segmented by words and converted into an embedding vector. The global semantics is extracted through a 12-layer Transformer encoder. The context-related vector representation of the word is output. The graph embedding vector in the knowledge graph is concatenated with the text semantic vector. The convolution kernel is used to perform feature crossover to form a fusion feature. A multi-scale convolution kernel is used to perform sliding window calculation on the fusion feature. Each convolution kernel generates a feature map. The significant features are extracted through maximum pooling. The fully connected layer outputs the final feature vector. The extracted feature vector is subjected to mutual information analysis. The mutual information value of each feature dimension and the target entity label is calculated. The vectors with a mutual information value lower than 0 are removed.For the feature dimension of 05, a piecewise normalization strategy is adopted. For statistical features, min-max normalization is used, and for semantic vectors, L2 norm unit normalization is used. The feature vectors are encapsulated into a JSON object according to the entity-relationship-attribute structure. The feature transmission priority is set according to the L value. When L = 1, it is marked as the normal priority with a weight of 0.5. When L ≥ 2, it is marked as the high priority with a weight of 0.8 and enters the graph construction queue first;. Knowledge graph dynamic construction module, which constructs a knowledge graph with confidentiality level tags, embeds the file security attributes into the knowledge graph nodes, optimizes the semantic representation and security relevance of the knowledge graph based on the text complexity level, defines the basic attributes of the file entities, and the basic attributes of the file entities include file identifiers, storage locations, confidentiality level tags, and responsible departments. Among them, the confidentiality level tags are divided into top secret, confidential, and internal, describes the interaction relationships between entities, and the interaction relationships between entities include access relationships, file derivation links, and revision histories. Bind the basic permission attributes and confidentiality level tags to the corresponding nodes of the knowledge graph to form a node-confidentiality level mapping table. Use the collected archive system metadata as the node basic attributes, and convert the operation behavior data in the operation logs into relationship edges. The relationship edge types include access, copy, and authorization. Based on the entity relationship triples output by the knowledge intelligence parsing module, construct a cross-archive semantic association network, label the security levels of the relationship types, store the confidentiality level tags as node attribute fields, and through embedding space constraints, connect high-confidentiality level nodes with edges of specific relationship types, and make the relationship paths across security domains meet the minimum confidentiality level transition conditions. Check the completeness of the node security attributes. If there are nodes missing attributes, they are prohibited from participating in relationship reasoning. Verify the confidentiality compliance of the relationship paths. If the head node confidentiality level is confidential, the tail node confidentiality level must not be lower than confidential, otherwise a security alarm is triggered. When a permission change operation is detected, use the graph traversal algorithm to locate the nodes affected by the permission change operation, recalculate their embedding vectors and security distances, and synchronize the real-time security status of the knowledge graph. According to the text complexity level L value, evaluate the semantic relationship complexity of the knowledge graph. If L = 1, it is considered that the relationship between entities is mainly linear one-to-one. If L ≥ 2, it is considered that there are multi-hop asymmetric relationships between entities. When L = 1, use the translational embedding model to establish a simple security association of head entity + relationship ≈ tail entity through linear transformation and calculate the loss function. When L ≥ 2, use the rotational embedding model to model complex paths through the Hadamard product formula in the complex space, support asymmetric and reverse relationship reasoning, and capture multi-hop security dependencies. During the vectorization process of the knowledge graph, define a security distance threshold, construct a confidentiality constraint space, use the confidentiality level tags as embedding space constraint conditions, allow internal-level nodes to be distributed in a larger-radius common subspace, limit confidential-level nodes to independent regions, maintain the minimum distance from the common subspace, and completely isolate top-secret-level nodes, only allowing them to be connected through specific relationship edges. During the training of the embedding model, superimpose security constraint loss terms, and the security constraint loss terms include semantic loss and security violation penalties. When the cross-confidentiality level access frequency anomaly is detected, automatically increase the security coefficient and strengthen isolation, and restore the basic value during low-risk periods. Before the L value changes, pre-load the parameters of the target embedding model into the memory, use dual-model parallel reasoning, gradually switch 10% of the traffic to the new model, and after verifying the stability, switch all. If the inference error rate of the new model exceeds 5%, automatically roll back to the previous stable version. For the generated entity relationship triples,Verify whether the confidentiality level of the head and tail nodes meets the security policy of the relationship type. If the relationship type is replication, the confidentiality level of the tail node must not be higher than that of the head node. If the relationship type is authorization, the confidentiality level of the head node must be at least one level higher than that of the tail node. Scan the embedding space regularly to detect abnormal neighboring nodes. For cross-confidentiality node pairs with a distance less than the security threshold, trigger manual review and automatically correct the vector coordinates of the illegal nodes to return them to the compliance area. Security situation assessment module, by improving the PageRank algorithm, quantifies the security influence of knowledge graph nodes, identifies highly sensitive files, dynamically divides security domains and generates minimum privilege policies based on the knowledge graph and text complexity levels, quantifies the classification labels of knowledge graph nodes into security weight coefficients, sets the weight of top-secret nodes to 3, the weight of confidential nodes to 2, and the weight of internal nodes to 1, defines the basic influence value of nodes, sets the propagation attenuation rate according to the edge type, sets the attenuation rate of cross-classification replication to 0.3, the attenuation rate of same-classification access to 0.7, and the attenuation rate of internal authorization to 0.9, introduces security parameters into the PageRank algorithm, the security parameters include classification difference, access frequency and complexity attenuation factor, monitors the cross-classification access frequency in real time, when the cross-level access of a single node per hour > 50 times, updates the attenuation rate, performs a security weight update every 6 hours, calculates the global node influence mean, standard deviation and high-sensitivity threshold, marks the nodes whose influence values exceed the high-sensitivity threshold in real time, generates a high-risk list, performs a reverse graph traversal on the marked nodes, identifies the key nodes in all incoming edge paths, if the path contains ≥ 2 top-secret nodes, marks it as a core risk source, if the path crosses more than 3 classification regions, marks it as a cross-domain propagation hotspot, automatically adds temporary access restrictions to highly sensitive nodes, prohibits cross-classification read operations, forces the use of two-factor authentication for writing, re-shards the node-associated data according to security domains, physically isolates high-risk data, calculates the number of incoming and outgoing edges of each node and the average classification difference of adjacent nodes, identifies high-connection density regions and isolated nodes, sets the path weight according to the relationship edge type, cross-classification path weight = base value × classification difference, same-classification path weight = base value, if L = 1, uses the spectral clustering algorithm to divide the security domain, the number of eigenvectors = 5, and the Laplacian matrix regularization parameter is fixed at 0.1. If L≥2, dynamically adjust the regularization parameter to enhance the cutting accuracy of the complex relationship network, and generate a security domain that includes a core domain, a transition domain, and a common domain. Among them, the core domain is a connected subgraph containing ≥2 top-secret nodes, and an independent physical storage partition is set. The transition domain is an area containing confidential nodes and having an access relationship with the core domain, and logical isolation is enabled. The common domain is a low-risk area containing internal nodes, allowing cross-domain reading but restricting writing. Statistically analyze the access frequency, operation type, and cross-domain ratio of nodes in the past 72 hours to generate a behavior baseline, calculate the risk score, and set the risk threshold. Pre-define the basic permissions according to the node classification level. If the risk score is greater than the risk threshold, it is identified as a high-risk node, and cross-domain writing is prohibited, and copy operations are restricted. If it is a core domain node, two-factor authentication is mandatory, and all operation logs are audited. During the operation process, only the minimum permissions necessary to complete the current operation are granted. Recalculate the security domain division every 6 hours. If the average risk score of the nodes in a region increases by 15%, trigger domain splitting, and split the region into sub-domains according to the classification level difference. If the interaction frequency between adjacent domains is lower than the risk score threshold, merge them into a single domain to reduce management overhead. Real-time detect the cross-domain request frequency of a single node. If the number of cross-domain requests within 1 minute > 20 times, trigger the fuse mechanism, and suspend the cross-domain permissions of this node for 10 minutes. After the fuse recovery, reactivate the permissions through the approval of the administrator. Perform a simulation deduction on the generated permission policy, construct a virtual operation chain, and verify whether there is a path to bypass the permission constraints through legal operation combinations. If a vulnerability is found, append a policy to limit the length of the operation chain;. Secure Communication Central Module. It binds the classification tags in the knowledge graph through the SM4 algorithm, performs dynamic authentication of identities for cross-domain communication, calculates the comprehensive sensitivity based on the text complexity level, and dynamically constructs a bandwidth allocation decision tree. It uses the classification tags of the nodes in the knowledge graph as encryption factors, generates a dynamic session key through the key derivation function of the SM4 algorithm, and quantifies the classification encoding as top secret = 3, secret = 2, and internal = 1. When a cross-domain communication request is initiated, it extracts the security attributes of the requesting node, and the security attributes include the affiliated security domain and the access permission list. It encrypts the requester identifier, the target domain identifier, and the operation type using the session key to generate an authentication token. It intercepts the header metadata of the data packet during the communication process, extracts the encrypted token and the target domain identifier, queries the security policy of the target domain through the knowledge graph, and the security policy includes the allowed classification range and the list of authorized departments. It recalculates the session key using the metadata of the requesting node, decrypts the token to obtain the original request information, and verifies whether the decrypted operation code is in the allowed operation set of the target domain and satisfies that the classification of the requester is higher than the lowest classification of the target domain, and the department of the requester belongs to the list of authorized departments of the target domain. If the verification passes, it encrypts the communication content using the new session key and sets the key validity period to 15 minutes. If the permissions do not match and the token decryption fails, it records the security event and generates an alarm, freezing the cross-domain communication permission of the requesting node for 30 minutes. When the classification and permissions of the nodes in the knowledge graph change, it automatically pushes an update event to the communication central, immediately invalidates all tokens issued by this node, and forces re-authentication. If a node initiates more than 10 cross-domain requests within 1 minute, its classification weight is automatically reduced, and the updated authentication rules are synchronized to all communication nodes within 50 milliseconds. It sets a basic sensitivity value according to the classification tags of the nodes in the knowledge graph, where top secret nodes = 3, secret nodes = 2, and internal nodes = 1, and superimposes the dynamic risk score to calculate the comprehensive sensitivity. When L = 1, it is marked as a regular transmission task. When L ≥ 2, it is marked as a critical transmission task. It constructs a decision tree, divides the main branch of the decision tree according to the L value. If L = 1, it is divided into the left branch and enters the regular transmission strategy. It judges whether the comprehensive sensitivity of the node is ≥ 2. If the comprehensive sensitivity ≥ 2, it allocates 1.2 times the basic bandwidth and enables redundant verification. If the comprehensive sensitivity < 2, it allocates 0.8 times, limit the retransmission times ≤ 3. If L ≥ 2, it is divided into the right branch and enters the key transmission strategy. Determine whether the node is located in the core security domain. If the node is located in the core security domain, the allocated bandwidth = total bandwidth × comprehensive sensitivity / global sensitivity. If the node is not located in the core security domain, the allocated bandwidth = remaining bandwidth × comprehensive sensitivity / non-core sensitivity sum. Monitor the current network throughput and dynamically correct the bandwidth allocation. When the node fails to transmit 3 times continuously, degrade its sensitivity weight. If the weight is lower than the threshold, suspend the node transmission and trigger an alarm. Regenerate the decision tree every 2 hours, count the transmission success rate of each branch, prune the branches with a success rate lower than 85%, merge the branches with similar strategies, and reduce the depth of the decision tree. If L ≥ 2 and sensitivity ≥ 2, monopolize 40% of the bandwidth, allow preemption of low-priority resources, the transmission interval ≤ 100 ms, and enable end-to-end encryption. If it does not meet L ≥ 2 and sensitivity ≥ 2, share the remaining bandwidth, adopt round-robin scheduling, extend the transmission interval to 500 ms, and perform fragmentation transmission on highly sensitive data streams. If the top-secret data fragment ≤ 64 KB, attach a hash checksum. If the confidential data fragment ≤ 128 KB, enable resume from breakpoint.

[0021] First, the system needs to be initialized and configured, including setting database interface parameters, log capture rules, etc., to ensure that the archive data collection module can correctly collect metadata and operation behavior data from the archive system. At the same time, according to actual requirements, relevant parameters of the knowledge intelligent parsing module, knowledge graph dynamic construction module, security situation assessment module, and security communication central module are configured. Start the archive data collection module, which will automatically perform structured collection of metadata and operation behavior data in the archive system through the preset database interface and log capture technology. These data will be used for subsequent knowledge intelligent parsing and knowledge graph construction. Next, the knowledge intelligent parsing module will calculate the text complexity level L of the collected archive data, and dynamically select preprocessing methods and entity recognition models according to the L value to perform dynamic governance of data quality and feature extraction. Then, the knowledge graph dynamic construction module will use these data to construct a knowledge graph with confidentiality level tags, embed the archive security attributes into the graph nodes, and form a three-dimensional model of data-relationship-security. The security situation assessment module will, based on the constructed knowledge graph, quantify the security influence of nodes by improving the PageRank algorithm and identify highly sensitive archives. At the same time, this module will also dynamically divide the security domain and generate the least privilege policy according to the topological structure of the graph and the L value parameter to ensure the security of archive data. During cross-domain communication and data transmission, the security communication central module will bind with the knowledge graph metadata through the SM4 algorithm to achieve dynamic authentication of identities. At the same time, this module will also dynamically construct a bandwidth allocation decision tree according to the L value and the sensitivity of the knowledge graph nodes to ensure the security and efficiency of data transmission. Finally, the system will continuously monitor the security situation of archive data and adjust and optimize the parameters of each module according to the actual situation to ensure the stability and security of the system. At the same time, the administrator can also regularly export security reports and data analysis results as needed for decision-making reference. Through the above steps, the administrator can efficiently utilize this archive data security integration management system to achieve functions such as comprehensive collection, intelligent parsing, graph construction, security assessment, and security communication of archive data.

[0022] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

Claims

1. An archival data security integration management system, characterized in that: It includes an archive data collection module, a knowledge intelligent parsing module, a knowledge graph dynamic construction module, a security situation assessment module, and a security communication central module; The archive data collection module structurally collects metadata and operation behavior data of the archive system; The knowledge intelligent parsing module calculates the text complexity level based on the metadata of the archive system, and preprocesses and extracts features from the metadata of the archive system; The knowledge graph dynamic construction module constructs a knowledge graph with confidentiality level tags, embeds archive security attributes into the nodes of the knowledge graph, and optimizes the semantic representation and security relevance of the knowledge graph based on the text complexity level; The security situation assessment module quantifies the security influence of the nodes of the knowledge graph by improving the PageRank algorithm, identifies highly sensitive archives, dynamically divides the security domain and generates the least privilege policy based on the knowledge graph and the text complexity level; The security communication central module binds the confidentiality level tags in the knowledge graph through the SM4 algorithm, performs dynamic authentication of identities for cross-domain communication, calculates the comprehensive sensitivity based on the text complexity level, and dynamically constructs a bandwidth allocation decision tree.

2. The archival data security integration management system according to claim 1, characterized in that: In the archive data collection module, the process of structurally collecting metadata and operation behavior data of the archive system includes: Adopt a dual-machine hot standby deployment mode, deploy a database protocol converter and a log collection terminal, connect to the archive management system database through the JDBC / ODBC interface, and capture structured metadata and operation logs in real time; Through the database interface polling mechanism, periodically scan the system tablespace to obtain the metadata of the archive system, and generate structured records according to the preset field template containing the file identifier and storage location; Deploy a log parsing engine on the database server side of the archive management system, extract operation behavior data through online log scanning and archived log backtracking technology, the parsing engine has a built-in transaction recognition algorithm, reorganize the original log stream according to the timestamp - operation subject - behavior type structure, deploy a national secret SM4 encryption gateway between the collection terminal and the security communication central module, perform sharding encryption on the transmission fields, use a distributed message queue to implement multi-source data buffering, and set up a primary and backup transmission channel; Parse the transaction records in the database online log in real time, identify the behavior category through the operation type coding mapping table, adopt an incremental capture strategy, locate the uncollected records based on the log sequence number, and for delete operations, retain the pre-image data; Align the metadata and operation behavior data of the archive system in time and space, use the file identifier as the primary key, record the operation behavior data to the corresponding metadata entry, perform normalization processing, uniformly convert the heterogeneous time format to the system standard time code, classify the structured data packet by field sensitivity, perform sharding encryption using the SM4 algorithm, encrypt the highly sensitive fields containing permission attributes and operation subjects with independent keys, and encrypt the low-sensitive fields containing timestamps and operation types with group keys. Asynchronously push the encrypted data packet to the security communication central module through the publish-subscribe mode of the message queue, and automatically trigger the breakpoint resumption mechanism in case of transmission failure.

3. The archival data security integration management system according to claim 2, characterized in that: In the knowledge intelligent analysis module, the process of calculating the text complexity level based on the archive system metadata and preprocessing the archive system metadata includes: The total number of characters in the archive description text is divided by the number of text paragraphs, and the ratio of the number of entities identified in the text to the total number of characters in the text is calculated, and then multiplied by 1000 to convert it into the number of thousand-word entities. The number of entities includes names of people, institutions and places. The text complexity level L is divided into 1, 2 and 3 through linear weighted comprehensive text length and semantic density. When L=1, it is judged as a simple text, and when L≥2, it is judged as a medium-to-high complexity text; When L=1, the sliding window mean is used to fill in the local missing values, the window width is set to 5 adjacent sample points in the time series, the standard deviation threshold detection is performed, and abnormal records that deviate from the mean by more than 3 times the standard deviation are eliminated. The bidirectional long short-term memory network and conditional random field joint model are used to capture the basic entity relationship through the time series features, and the keyword weight is calculated based on the improved TF-IDF algorithm; When L ≥ 2, time series prediction is used to fill missing values, modeling is based on the periodic characteristics of historical data, dynamic cutting of isolation forests is enabled, outliers in complex distributions are adaptively identified through tree depth, pre-trained language models and conditional random field joint models are loaded, and deep semantic understanding capabilities are used to parse professional terms and nested structures in archives. Graph embedding vectors and text semantic vectors are fused, cross-modal features are extracted through convolution operations, and distributed parallel computing is enabled.

4. The archival data security integration management system according to claim 3, characterized in that: In the knowledge intelligent parsing module, the process of performing feature extraction according to the text complexity level includes: Dividing the feature hierarchy into basic features, semantic features and topological features, the basic features include text length, entity density and keyword frequency, the semantic features include the relationship between entities and context dependency structure, and the topological features include the connection attributes of entity nodes in the knowledge graph; When L=1, the improved word frequency-inverse document frequency algorithm is used to calculate the keyword weight. By suppressing the weight of high-frequency general words, the discrimination of professional entity words is improved. The text sequence features are extracted based on the bidirectional long short-term memory network architecture. The forward network encodes the context information in character order, and the reverse network captures the suffix dependency in reverse order. The forward and reverse hidden state vectors of the character position are output, and the transition probability matrix is ​​introduced to calculate the optimal entity label sequence and generate entity relationship triples. When L ≥ 2, a text semantic vector is generated through a multi-layer self-attention mechanism. The input text is segmented into words and then converted into embedding vectors. After passing through 12 layers of Transformer encoders to extract global semantics, a context-related vector representation of the words is output. The graph embedding vector in the knowledge graph is concatenated with the text semantic vector, and feature cross is performed through a convolutional kernel to form a fused feature. A multi-scale convolutional kernel is used to perform a sliding window calculation on the fused feature. Each convolutional kernel generates a feature map, and significant features are extracted through max pooling. The fully connected layer outputs the final feature vector. Mutual information analysis is performed on the extracted feature vector, the mutual information value between each feature dimension and the target entity label is calculated, and the feature dimensions with a mutual information lower than 0.05 are removed. A piecewise normalization strategy is adopted, with min-max normalization for statistical features and L2 norm unit normalization for semantic vectors; The feature vector is encapsulated into a JSON object in the entity-relationship-attribute structure. The feature transmission priority is set according to the L value. When L = 1, it is marked as a normal priority with a weight = 0.

5. When L ≥ 2, it is marked as a high priority with a weight = 0.8 and enters the graph construction queue first.

5. A file data security integration management system according to claim 4, characterized in that: In the knowledge graph dynamic construction module, the process of constructing a knowledge graph with a confidentiality label attached and embedding the file security attributes into the knowledge graph nodes includes: Define the basic attributes of the archive entity. The basic attributes of the archive entity include file identifier, storage location, confidentiality label, and responsible department. Among them, the confidentiality label is divided into top secret, secret, and internal. Describe the interaction relationships between entities. The interaction relationships between entities include access relationships, file derivation links, and revision histories. Bind the basic permission attributes and the confidentiality label to the corresponding nodes in the knowledge graph to form a node-confidentiality mapping table; Use the collected archive system metadata as the basic attributes of the nodes, and convert the operation behavior data in the operation logs into relationship edges. The types of relationship edges include access, copy, and authorization. Based on the entity relationship triples output by the knowledge intelligent parsing module, construct a cross-archive semantic association network, mark the security level of the relationship type, store the confidentiality label as a node attribute field, and through the embedding space constraint, connect high-confidentiality nodes with edges of specific relationship types, and make the relationship paths across security domains meet the minimum confidentiality transition condition; Check the completeness of the node security attributes. If there are nodes missing attributes, they are prohibited from participating in relationship reasoning. Verify the confidentiality compliance of the relationship paths. If the head node confidentiality level is secret, the tail node confidentiality level must not be lower than secret, otherwise a security warning is triggered. When a permission change operation is detected, use the graph traversal algorithm to locate the nodes affected by the permission change operation, recalculate their embedding vectors and security distances, and synchronize the real-time security status of the knowledge graph.

6. The archival data security integration management system according to claim 5, characterized in that: In the knowledge graph dynamic construction module, the process of optimizing the semantic representation and security relevance of the knowledge graph based on the text complexity level includes: According to the L value of the text complexity level, evaluate the complexity of the semantic relationships in the knowledge graph. If L = 1, it is considered that the relationships between entities are mainly linear one-to-one. If L ≥ 2, it is considered that there are multi-hop asymmetric relationships between entities; When L=1, a translation embedding model is used to establish a simple security association of head entity + relationship ≈ tail entity through linear transformation and calculate the loss function. When L≥2, a rotation embedding model is used to model complex paths in complex space through the Hadamard product formula, support asymmetric and reverse relationship reasoning, and capture multi-hop security dependencies. In the process of knowledge graph vectorization, a security distance threshold is defined, a confidentiality constraint space is constructed, and confidentiality labels are used as embedding space constraints. Internal nodes are allowed to be distributed in a public subspace with a larger radius, confidential nodes are restricted to independent areas, and the minimum distance from the public subspace is maintained. Top secret nodes are completely isolated and only allowed to be connected through specific relationship edges. When training the embedding model, a security constraint loss term is superimposed, and the security constraint loss term includes semantic loss and security violation penalty. When abnormal cross-classification access frequency is detected, the safety factor is automatically increased, isolation is strengthened, and the basic value is restored during the low-risk period. Before the L value changes, the parameters of the target embedding model are pre-loaded into the memory, and dual-model parallel reasoning is used. 10% of the traffic is gradually switched to the new model, and the full switch is performed after the stability is verified. If the reasoning error rate of the new model exceeds 5%, it will automatically roll back to the previous stable version; For the generated entity-relationship triples, verify whether the confidentiality levels of the head and tail nodes meet the security policy of the relationship type. If the relationship type is replication, the confidentiality level of the tail node must not be higher than that of the head node. If the relationship type is authorization, the confidentiality level of the head node must be at least one level higher than that of the tail node. Scan the embedding space regularly to detect abnormal neighbor nodes. For cross-confidentiality node pairs with a distance less than the security threshold, trigger manual review and automatically correct the vector coordinates of the illegal nodes to return them to the compliant area.

7. A file data security integration management system according to claim 6, characterized in that: In the security situation assessment module, the process of quantifying the security influence of knowledge graph nodes and identifying highly sensitive files by improving the PageRank algorithm includes: Quantify the confidentiality labels of knowledge graph nodes into security weight coefficients, set the weight of top-secret nodes to 3, the weight of confidential nodes to 2, and the weight of internal nodes to 1, define the basic influence value of nodes, set the propagation attenuation rate according to the edge type, set the attenuation rate of cross-confidentiality replication to 0.3, the attenuation rate of same-confidentiality access to 0.7, and the attenuation rate of internal authorization to 0.9; Introducing security parameters into the PageRank algorithm, the security parameters include confidentiality difference, access frequency and complexity attenuation factor; Real-time monitoring of cross-level access frequency. When a single node has >50 cross-level accesses per hour, the attenuation rate is updated and the security weight is updated every 6 hours. Calculate the global node influence mean, standard deviation and high sensitivity threshold, mark nodes whose influence value exceeds the high sensitivity threshold in real time, generate a high-risk list, perform reverse graph traversal on the marked nodes, identify the key nodes in all their incoming edge paths, and if the path contains ≥2 top-secret nodes, mark it as a core risk source. If the path crosses more than 3 confidentiality-level areas, mark it as a cross-domain transmission hotspot. Automatically add temporary access restrictions for highly sensitive nodes, prohibit cross-level read operations, force the enablement of two-factor authentication writes, re-shard and store node-related data by security domain, and physically isolate high-risk data.

8. A file data security integration management system according to claim 7, characterized in that: In the security situation assessment module, the process of dynamically dividing security domains and generating a minimum privilege policy based on the knowledge graph and text complexity level includes: Calculate the number of inbound and outbound edges of each node and the average confidentiality difference of adjacent nodes, identify high connection density areas and isolated nodes, set path weights according to the relationship edge type, cross-confidentiality path weight = basic value × confidentiality difference, same-confidentiality path weight = basic value, if L = 1, use spectral clustering algorithm to divide security domains, number of eigenvectors = 5, Laplace matrix regularization parameter is fixed to 0.1, if L ≥ 2, dynamically adjust regularization parameters, enhance the cutting accuracy of complex relationship networks, generate security domains including core domains, transition domains and public domains, where the core domain is a connected subgraph containing ≥ 2 top-secret nodes, set independent physical storage partitions, transition domains are areas containing confidential nodes and have access relations with the core domain, enable logical isolation, and public domains are low-risk areas containing internal nodes, allowing cross-domain reading but restricting writing; Statistics are collected on the access frequency, operation type and cross-domain ratio of nodes in the past 72 hours, behavioral baselines are generated, risk scores are calculated and risk thresholds are set, and basic permissions are predefined according to the node classification level. If the risk score is greater than the risk threshold, it is identified as a high-risk node, cross-domain writing is prohibited, and copy operations are restricted. If it is a core domain node, two-factor authentication is mandatory, and operation logs are fully audited. During the operation, only the minimum permissions required to complete the current operation are granted; The security domain division is recalculated every 6 hours. If the average risk score of regional nodes increases by 15%, domain splitting is triggered, and the region is divided into sub-domains according to the difference in security levels. If the interaction frequency between adjacent domains is lower than the risk score threshold, they are merged into a single domain to reduce management overhead. The frequency of cross-domain requests of a single node is detected in real time. If the cross-domain request is >20 times within 1 minute, the circuit breaker mechanism is triggered to suspend the cross-domain permission of the node for 10 minutes. After the circuit breaker is restored, the permission is reactivated through administrator approval. The generated permission policy is simulated and deduced, and a virtual operation chain is constructed to verify whether there is a path to bypass the permission constraint through a legitimate operation combination. If a vulnerability is found, an operation chain length limit policy is added.

9. The archival data security integration management system according to claim 8, wherein: In the secure communication hub module, the process of dynamically authenticating the identity of cross-domain communication by binding the confidentiality level labels in the knowledge graph through the SM4 algorithm includes: The confidentiality labels of nodes in the knowledge graph are used as encryption factors, and a dynamic session key is generated through the key derivation function of the SM4 algorithm. The confidentiality levels are quantified according to top secret = 3, confidential = 2, and internal = 1. When a cross-domain communication request is initiated, the security attributes of the requesting node are extracted, including the security domain to which it belongs and the access permission list. The requesting party identifier, target domain identifier, and operation type are encrypted using the session key to generate an authentication token. Intercept the header metadata of the data packet during the communication process, extract the encrypted token and the target domain identifier, query the security policy of the target domain through the knowledge graph, the security policy includes the scope of confidentiality allowed to be accessed and the list of authorized departments, use the metadata of the requesting node to recalculate the session key, decrypt the token to obtain the original request information, verify whether the decrypted operation code is in the set of allowed operations of the target domain, and satisfy the confidentiality level of the requesting party is higher than the minimum confidentiality level of the target domain, and the requesting party department belongs to the list of authorized departments of the target domain. If the verification is successful, use the new session key to encrypt the communication content, set the key validity period to 15 minutes, if the permissions do not match and the token decryption fails, then record the security event and generate an alarm, and freeze the cross-domain communication permissions of the requesting node for 30 minutes; When the confidentiality level and permissions of a node in the knowledge graph change, an update event is automatically pushed to the communication hub, immediately invalidating all tokens issued by the node and forcing re-authentication. If the same node initiates more than 10 cross-domain requests within 1 minute, its confidentiality weight is automatically reduced. The updated authentication rules are synchronized to all communication nodes within 50 milliseconds.

10. A file data security integration management system according to claim 9, characterized in that: In the secure communication hub module, the process of calculating the comprehensive sensitivity based on the text complexity level and dynamically constructing the bandwidth allocation decision tree includes: The basic sensitivity value is set according to the confidentiality label of the node in the knowledge graph, where top secret node = 3, confidential node = 2, and internal node = 1. The dynamic risk score is superimposed to calculate the comprehensive sensitivity. When L = 1, it is marked as a regular transmission task, and when L ≥ 2, it is marked as a critical transmission task. Construct a decision tree and divide the main branches of the decision tree according to the L value. If L=1, it is divided into the left branch and enters the conventional transmission strategy to determine whether the comprehensive sensitivity of the node is ≥2. If the comprehensive sensitivity is ≥2, 1.2 times of the basic bandwidth is allocated and redundancy check is enabled. If the comprehensive sensitivity is <2, 0.8 times of the basic bandwidth is allocated and the number of retransmissions is limited to ≤3. If L≥2, it is divided into the right branch and enters the key transmission strategy to determine whether the node is located in the core security domain. If the node is located in the core security domain, the allocated bandwidth = total bandwidth × comprehensive sensitivity / global sensitivity. If the node is not located in the core security domain, the allocated bandwidth = remaining bandwidth × comprehensive sensitivity / non-core sensitivity sum; Monitor the current network throughput and dynamically correct bandwidth allocation. When a node fails to transmit for three consecutive times, downgrade its sensitivity weight. If the weight is lower than the threshold, suspend the node transmission and trigger an alarm. Regenerate the decision tree every 2 hours, count the transmission success rate of each branch, prune branches with a success rate lower than 85%, merge branches with similar strategies, and reduce the depth of the decision tree. If L≥2 and sensitivity≥2, 40% of the bandwidth will be exclusively used, low-priority resources can be preempted, the transmission interval is ≤100ms, and end-to-end encryption is enabled. If L≥2 and sensitivity≥2 are not met, the remaining bandwidth will be shared, polling scheduling will be used, the transmission interval will be extended to 500ms, and highly sensitive data streams will be transmitted in fragments. If the top-secret data fragment is ≤64KB, a hash check will be added. If the confidential data fragment is ≤128KB, breakpoint resumption will be enabled.

Citation Information

Patent Citations

  • Policy information management data processing method and device, equipment and storage medium

    CN116595173A

  • Electronic archive management system based on knowledge graph

    CN117171105A

  • Multivariate data management method and system based on knowledge graph technology

    CN117235281A

  • Coal mine potential safety hazard characterization method and system based on knowledge graph

    CN119272868A

Cited By

  • AI authority intelligent distribution system and method based on behavior prediction

    CN120449209A

  • Intelligent archive management method and system

    CN120524969A

  • Cross-department government affair data security sharing and analysis method based on deep learning

    CN120632954A

  • File management all-in-one machine cross-region management system and method based on hybrid architecture

    CN120705855A

  • Archive management all-in-one machine cross-region management system and method based on hybrid architecture

    CN120705855B