Information security management method and system based on sensitive data

By identifying data sensitivity levels and dynamically updating permissions, combining quantum key distribution and dynamic chaotic key stream encryption, zero-trust data tunneling and multi-dimensional detection are established, which solves the problem of insufficient protection in the face of internal abuse by traditional information security management methods, and achieves efficient and reliable security management of sensitive data.

CN120449206APending Publication Date: 2025-08-08HANGZHOU LUXIANG TECH CO LTD

Patent Information

Application Number
CN202510594986.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

When facing the complexity and internal abuse of data usage scenarios, the existing information security management methods lack protection capabilities, lack the ability to identify and dynamically manage data sensitivity, and cannot achieve refined, dynamic, and visual management. In addition, traditional mechanisms lack protection capabilities when facing the illegal use of internal legal identities.

Method used

By collecting data to identify sensitive information, generating tags and dividing data sensitivity levels, using time series prediction models for real-time evaluation, dynamically updating permission matrix, and combining distributed redundant disaster recovery systems with quantum key distribution and dynamic chaotic key stream encryption, a zero-trust data tunnel is established for unified identity authentication, and an isolated forest model and LSTM time series prediction model are used for abnormal detection to build a multi-dimensional security protection mechanism.

Benefits of technology

Real-time and multi-dimensional security monitoring of sensitive data is realized, and the anti-tampering, anti-forgery and anti-leakage capabilities of the data access link are improved, ensuring the recovery and security of data in extreme disasters, and enhancing the real-time monitoring and abnormal detection capabilities of user behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449206A_ABST
    Figure CN120449206A_ABST
Patent Text Reader

Abstract

The invention discloses an information security management method and system based on sensitive data, relates to the technical field of data management, and comprises the following steps: realizing quantum encryption transmission of remote high-sensitive data by combining a quantum key distribution link and a dynamic chaotic key stream; the partitioning and synchronization priority is dynamically adjusted according to the node load, the data integrity is verified through CRC, the distributed storage disaster recovery system with extremely high safety and usability is achieved, it is ensured that data can still be recovered even under extreme disasters, a zero-trust data channel based on an IPSec tunnel is constructed in the data transmission process, and the reliability of data transmission is improved. The tamper-proof, anti-counterfeiting and anti-leakage capabilities of a data access link are greatly improved, the access frequency abnormity is detected through the isolated forest model, the behavior trend is detected through the LSTM time sequence prediction model, the baseline deviation analysis is performed through user behavior portrait comparison, and the real-time, multi-modal and chained security monitoring of the user data operation behavior is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and in particular to an information security management method and system based on sensitive data. Background Art

[0002] Against the backdrop of the rapid development of informatization and digitalization, data has become a core driver of socioeconomic development. Data contains vast amounts of personal privacy information, commercially sensitive data, and important national data. Consequently, data security and privacy protection issues are becoming increasingly severe. This is especially true for data involving user privacy, core business operations, and nationally sensitive areas. The leakage, tampering, or misuse of data can have serious consequences. Currently, the industry's commonly adopted information security management approaches are based on perimeter protection and unified policy configuration, emphasizing the use of access control, identity authentication, firewalls, and other measures to prevent unauthorized intrusion and data leakage. However, with the increasing complexity of data usage scenarios, data security risks have gradually shifted from "peripheral attacks" to "internal misuse." Traditional perimeter-based security mechanisms are clearly insufficient in protecting against "legitimate identities, illicit use." Furthermore, existing data security products, such as database auditing, data loss prevention (DLP), and data encryption systems, while providing some protection in specific scenarios, still suffer from complex policy configuration, high maintenance costs, and a lack of granular identification and dynamic management capabilities for data sensitivity. Furthermore, they are unable to implement automated data-level control and traceability, making them unable to meet business demands for refined, dynamic, and visual management.

[0003] Although current research has attempted to improve the situation by introducing technical means such as behavioral analysis, dynamic permission management, and deep encryption, limitations still exist. First, with the diversification of data types and structures, the difficulty of automatically identifying and classifying sensitive data continues to increase. Existing solutions mostly rely on static rules or manual labeling, which are difficult to adapt to the dynamically changing data flow environment. Secondly, data access permission settings are still mainly based on roles, and fail to fully consider dynamic factors such as changes in data sensitivity, user behavior patterns, and contextual environments, which can easily lead to security vulnerabilities such as permission abuse or failure of "least privilege". In addition, there is a lack of effective hierarchical encryption and key management mechanisms during data storage, transmission, and use. Once a link is breached, the overall protection system will be destroyed. Therefore, there is an urgent need to build a comprehensive information security management method that integrates data perception, dynamic response, and multi-dimensional prevention and control. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an information security management method and system based on sensitive data to solve the problem of a comprehensive information security management method that integrates data perception, dynamic response and multi-dimensional prevention and control.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides an information security management method based on sensitive data, which comprises:

[0008] Collect data, identify sensitive information, generate labels, and classify data sensitivity levels to build a data access permission mapping matrix. Use a time series prediction model to evaluate changes in data sensitivity in real time, and dynamically update labels and corresponding permission matrices.

[0009] Execute a hierarchical encryption strategy on the data based on the update results, and store the encrypted data in a distributed redundant disaster recovery system that integrates quantum key distribution and dynamic chaotic key stream encryption;

[0010] During the process of user access to data, a zero-trust data tunnel is established and unified identity authentication access control is performed. Through the isolation forest model, LSTM time series prediction model and user behavior profile comparison, data security is continuously analyzed and detected, and security protection measures are generated.

[0011] As a preferred solution of the information security management method based on sensitive data described in the present invention, the continuous analysis and detection of data security and generation of security protection measures through the isolation forest model, LSTM time series prediction model and user behavior portrait comparison refers to the use of blockchain technology to chain log records for each data access, modification and deletion action, and the chain storage of each log in sequence, and the triple superposition anomaly detection through the isolation forest model, LSTM time series prediction model and user behavior portrait comparison.

[0012] As a preferred solution of the information security management method based on sensitive data described in the present invention, wherein: the establishment of a zero-trust data tunnel and unified identity authentication access control during the user access to data refers to the construction of a zero-trust data tunnel when the user accesses the data, the use of IPSec combined with RSA to encrypt the authentication information, binding the X.509 device certificate, and using SHA-256 to generate the data hash value. The data flow passes through the relay node hop-by-hop triple verification mechanism in the path, combined with OAuth2.0, the national secret SM series algorithm and two-factor authentication, access to the unified authentication platform, dynamic authorization based on the RBAC and ABAC fusion model, combining static roles and dynamic attributes to build a joint permission matrix, when the matrix judgment result is 1, access is authorized, otherwise it is denied.

[0013] As a preferred solution of the information security management method based on sensitive data described in the present invention, wherein: the encrypted data is stored in a distributed redundant disaster recovery system that integrates quantum key distribution and dynamic chaotic key stream encryption, which means that the encrypted data is encapsulated using CMS, pushed to the target distributed storage node using the Kafka message queue, and a one-time key encryption link is established through quantum key distribution technology to protect the data, and all data blocks are further dynamically encrypted using dynamic chaotic key stream encryption to generate a continuous key stream value. ;

[0014] from A part of the key is sampled for block-level encryption, the encrypted data is logically divided into blocks through hyperbolic cyclic mapping, the bilinear differential scheduling function is used to dynamically adjust the load balancing, and the synchronization priority index is calculated through the priority mechanism. , prioritize synchronization of important data blocks;

[0015] During synchronization, the optimal delay path is dynamically selected based on network delay , real-time monitoring system synchronization efficiency and node robustness, calculate the synchronization efficiency of monitoring node u , set the threshold , when the synchronization efficiency Less than threshold When the node is in the state of failure, node reselection and disaster recovery switching are triggered, otherwise normal synchronization is performed. After the synchronization is completed, a cyclic redundancy check (CRC) is used to verify the integrity of the data.

[0016] As a preferred solution of the information security management method based on sensitive data of the present invention, wherein: the performing of a hierarchical encryption strategy on the data according to the update result refers to encrypting the highly sensitive data, the medium sensitive data and the low sensitive data respectively according to the updated adjustment;

[0017] For highly sensitive data, the national secret SM4 symmetric encryption is used. The SM4 encryption algorithm is used to encrypt in the form of a block cipher, and the individual encrypted blocks are finally connected to form a ciphertext;

[0018] For medium-sensitive data, it is encrypted in the form of a block cipher using AES-256, the ciphertext encrypted by AES is hashed, and the private key is used to perform ECDSA signature to generate a signature;

[0019] For low-sensitivity data, ChaCha20 is used to encrypt with a 256-bit key and a 64-bit random number, and discrete cosine transform is used to embed watermarks in the low-frequency part of the encrypted data.

[0020] As a preferred solution of the information security management method based on sensitive data described in the present invention, wherein: the use of the time series prediction model to evaluate the change of data sensitivity in real time, dynamically updating the label and the corresponding authority matrix means monitoring the indicator in real time, obtaining the historical sensitivity score sequence, normalizing the sensitivity score by Z-Score to generate training samples, inputting them into the double-layer stacked LSTM network, extracting the local time series change characteristics and the global sensitivity evolution trend, obtaining the sensitivity prediction sequence, and calculating the mean sensitivity change during the prediction period. , set thresholds b and e, and b is greater than e, when When it is greater than the threshold b, the sensitivity is on an upward trend, the data label T is automatically upgraded, the sensitivity comprehensive score is recalculated, and the authority matrix P is tightened. When it is less than e, the sensitivity is decreasing and the access rights are relaxed. When it is greater than e and less than b, the sensitivity is stable, the status quo is maintained, and only logs are recorded.

[0021] As a preferred solution of the information security management method based on sensitive data described in the present invention, wherein: the data collection and identification of sensitive information to generate labels, and the classification of data sensitivity levels to construct a data access permission mapping matrix refers to collecting internal, external and terminal data, uniformly streaming them into the Apache Flink cluster, identifying sensitive information based on preset rules and the NER model, extracting context features to train a decision tree to generate label triples T, and constructing a permission matrix based on roles and sensitivity scores. , score each data according to the T label and calculate the comprehensive sensitivity score S;

[0022] Set permission control thresholds A and B, and A is greater than B. When S is greater than A, it is highly sensitive data and only system administrators can fully access it. When S is greater than B and less than or equal to A, it is medium-sensitive data and internal employees can read but not modify it. When S is less than or equal to B, it is low-sensitive data and is open to external collaborators for readable access.

[0023] In a second aspect, the present invention provides an information security management system based on sensitive data, comprising:

[0024] Data collection and sensitive identification module, used to uniformly collect data and identify sensitive information, generate labels and sensitivity levels;

[0025] Dynamic sensitivity prediction module, used for LSTM to predict sensitivity changes and dynamically adjust labels and permissions;

[0026] Hierarchical encryption mechanism module, used for encryption based on sensitivity classification, using SM4, AES, and ChaCha20 algorithms;

[0027] A secure storage disaster recovery module, used to integrate quantum key and chaotic encryption to build a highly secure and disaster-tolerant distributed system;

[0028] Zero Trust Access Control module, used to establish a zero trust tunnel and use multi-factor and certificate chain unified authentication;

[0029] The anomaly detection and protection module is used to combine the forest model with behavioral profiling to detect anomalies and provide protection in multiple dimensions.

[0030] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the information security management method based on sensitive data as described in the first aspect of the present invention is implemented.

[0031] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the information security management method based on sensitive data as described in the first aspect of the present invention.

[0032] The beneficial effects of the present invention are as follows: the present invention combines quantum key distribution links and dynamic chaotic key streams to realize quantum encrypted transmission of highly sensitive data in different locations, dynamically adjusts the block and synchronization priorities according to the node load, and adopts CRC to verify the data integrity, thereby realizing a distributed storage disaster recovery system with extremely high security and availability, ensuring that data can be recovered even in extreme disasters, and constructing a zero-trust data channel based on IPSec tunnel during data transmission, which greatly improves the anti-tampering, anti-forgery and anti-leakage capabilities of the data access link, and detects access frequency anomalies through the isolation forest model, detects behavioral trends through the LSTM time series prediction model, and performs baseline deviation analysis by comparing user behavior portraits, thereby realizing real-time, multi-modal and chain-type security monitoring of user data operation behavior. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0034] Figure 1 This is a flowchart of an information security management method based on sensitive data in Example 1.

[0035] Figure 2 This is a structural diagram of an information security management system based on sensitive data in Example 1.

[0036] Figure 3 This is a flowchart of hierarchical encryption and storage disaster recovery in Example 1. DETAILED DESCRIPTION

[0037] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0038] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0039] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0040] Example 1, with reference to Figures 1 to 3 , which is the first embodiment of the present invention, provides an information security management method based on sensitive data, comprising the following steps:

[0041] S1. Collect data and identify sensitive information to generate labels. Then, classify data sensitivity levels and build a data access permission mapping matrix. Use a time series prediction model to evaluate changes in data sensitivity in real time and dynamically update labels and the corresponding permission matrix.

[0042] Specifically, collecting data and identifying sensitive information to generate labels, and dividing data sensitivity levels to build a data access permission mapping matrix refers to deploying lightweight edge gateways near the interfaces of ERP, CRM and SCADA systems (internal data), accessing IoT devices and mobile terminal data streams through the MQTT protocol (terminal data), and accessing external APIs and databases through RESTful API adapters (external data). All accessed data will flow uniformly into the Apache Flink cluster, and sensitive information in the collected data will be preliminarily identified and processed through preset sensitive dictionaries (ID cards, bank account numbers), regular expressions and NER deep models. The identified sensitive fragments and their contextual features (word frequency, field name, historical access behavior) are used as input to train the decision tree classification algorithm to generate label triples T (confidentiality level, industry type and field category), and a hybrid model based on role-based access control and sensitivity weight scoring is used to establish a permission matrix. ,in is a collection of roles (administrators, internal employees, external collaborators), It is a set of data sensitivity levels. The matrix value represents the permission level of each role for data of different sensitivity levels. Each piece of data is scored according to the T label. The sensitivity scoring formula is:

[0043]

[0044] Among them, S is the comprehensive sensitivity score, is the score corresponding to the confidentiality level (based on the rules system and expert consultation to define the confidentiality level standards), is the industry sensitivity score (using data analysis and machine learning models to define the industry sensitivity score table), is the field category sensitivity score (using regular expression matching and NER model to define the field category sensitivity score table), , , is the weight coefficient, set through experiments;

[0045] By setting permission control thresholds A and B based on the role-based access control model, and A is greater than B, when S is greater than A, it is highly sensitive data and only system administrators can fully access it; when S is greater than B and less than or equal to A, it is medium-sensitive data and internal employees are allowed to read but not modify it; when S is less than or equal to B, it is low-sensitive data and is open to external collaborators for readable access.

[0046] By constructing the label triple T, a scalable and portable data labeling system is formed, which is suitable for different industries and business scenarios. It effectively improves the accuracy and stability of sensitive information classification and enhances the model's adaptability to new fields or context changes. By building a role-based access permission matrix and combining it with a sensitivity scoring model, flexible and dynamic data access control is achieved, avoiding the rigidity of static configuration permissions, so that the system can still ensure data security under dynamic conditions such as user changes and business expansion. By setting access permission thresholds A and B and combining the scoring results with the permission matrix, a fine-grained and hierarchical access control strategy is implemented, reducing the risk of misjudgment, improving system execution efficiency and security credibility, and supporting automated permission configuration.

[0047] Furthermore, the time series prediction model is used to evaluate the changes in data sensitivity in real time, and the labels and corresponding permission matrices are dynamically updated. A unified risk monitoring agent and abnormal behavior detection engine are used to monitor indicators (access frequency, changes in access location, interface call anomalies, and changes in user behavior patterns) in real time to obtain the historical sensitivity score sequence of each data. ,in, is the sensitivity score at time t. The sensitivity score is normalized by Z-Score. A fixed sliding window length a is used to generate training samples. The training samples are input into a two-layer stacked LSTM network. The first layer extracts the local time series change features and updates the hidden state:

[0048]

[0049] in, is the hidden state, is the output gate, is the unit state;

[0050] Unit Status renew:

[0051]

[0052] in, , They are the forget gate and input gate control variables, is a candidate memory unit;

[0053] Process the training samples in the sliding window in sequence according to the time step to obtain the hidden state output sequence of the entire window. Update the hidden state sequence as the second layer input to extract the global sensitivity evolution trend, obtain the sensitivity prediction sequence for the next k steps, and calculate the mean sensitivity change during the prediction period. ;

[0054] Thresholds b and e are set by analyzing historical sensitivity fluctuations, and b is greater than e. When it is greater than the threshold b, the sensitivity is on an upward trend, the data label T is automatically upgraded, the sensitivity comprehensive score is recalculated, and the authority matrix P is tightened. When it is less than e, the sensitivity is decreasing and the access rights are relaxed. When it is greater than e and less than b, the sensitivity is stable, the status quo is maintained, and only logs are recorded.

[0055] By using a sliding window to generate training samples and extracting local and global sensitivity features based on the LSTM model, a comprehensive and in-depth prediction of sensitivity is achieved, which can more accurately analyze the trend of sensitivity changes and improve the system's ability to perceive and prevent potential risks in advance. By setting thresholds and dynamically adjusting data labels and access rights, flexible permission management is achieved. While enhancing data protection, it can also flexibly adjust data accessibility according to system load and security threats, effectively reducing the rigidity of access control and improving the system's responsiveness and flexibility. By monitoring user access behavior in real time based on the abnormal behavior detection engine, it can immediately respond and adjust permissions or trigger alarms when abnormal access behavior occurs, further enhancing the security of data access.

[0056] S2. Execute a hierarchical encryption strategy on the data based on the update results, and store the encrypted data in a distributed redundant disaster recovery system that integrates quantum key distribution and dynamic chaotic key stream encryption;

[0057] Specifically, executing a hierarchical encryption strategy on data according to the update result means encrypting the highly sensitive data, the medium sensitive data and the low sensitive data respectively according to the updated adjustment;

[0058] For highly sensitive data, we use the national secret SM4 symmetric encryption, use a standard random number generator to generate a 16-byte key, and use the SM4 encryption algorithm to encrypt in the form of a block cipher. The plaintext is divided into blocks of fixed length (128 bits), and encrypted block by block. Finally, the encrypted blocks are connected to form the ciphertext.

[0059] For moderately sensitive data, AES-256 encryption is used to generate a key using a standard random number generator, which is then encrypted in the form of a block cipher. The AES-encrypted ciphertext is hashed to generate a hash value, and the private key is used for ECDSA signing to generate a signature (which is stored or transmitted with the encrypted data as metadata and verified using the public key).

[0060] For low-sensitivity data, ChaCha20 is used to encrypt with a 256-bit key and a 64-bit random number, and discrete cosine transform is used to embed a watermark in the low-frequency part of the encrypted data to ensure that it is not easily lost or tampered with during compression and transmission, thereby ensuring the concealment and anti-attack capabilities of the watermark.

[0061] By using SM4 to symmetric encrypt highly sensitive data, combined with national secret standards to improve data security, ensure national security requirements, and enhance the system's compliance and trust in national security reviews, AES-256 encryption and hashing combined with ECDSA signature technology are used to protect medium-sensitive data. Medium-sensitive data is not only protected, but also guaranteed to have integrity and tamper-proof capabilities, which helps to verify the accuracy of data during data transmission or storage, and improves the reliability and trust of the system. ChaCha20 encryption combined with watermark embedding technology is used to encrypt low-sensitivity data, which not only enhances the tamper-proof capabilities of low-sensitivity data, but also improves the tracking and traceability of data, ensuring security and reliability during data transmission.

[0062] Furthermore, storing encrypted data in a distributed redundant disaster recovery system that integrates quantum key distribution and dynamic chaotic key stream encryption means encapsulating the encrypted data using CMS (encrypted data body, signature information, metadata (such as sensitivity label, encryption algorithm identifier)), pushing it to the target distributed storage node using the Kafka message queue, generating N copies of each data, and storing them in different nodes for off-site central backup. By using quantum key distribution technology to establish a one-time key encryption link, sensitive data is synchronized and encrypted to prevent eavesdropping in the middle, and synchronization between nodes ensures ultimate security.

[0063] After the sensitive data is synchronously protected, dynamic chaotic key stream encryption is used to further dynamically encrypt all data blocks to improve the overall system's anti-cracking strength and prevent the leakage of non-sensitive data from becoming an attack breakthrough. The chaotic stream generation formula is:

[0064]

[0065] in, is the dynamic chaotic key stream at time The continuous key stream values generated on is the initial bias, is the amplitude control factor, is the input frequency parameter, is the phase shift, , , is the parameter that controls the complexity of chaos;

[0066] from A portion of the key is sampled for block-level encryption, and the encrypted data is logically divided into blocks through hyperbolic cyclic mapping. The block size is dynamically adjusted according to the current load of each node. The bilinear differential scheduling function is used to dynamically adjust the load balancing. The threshold is set by the upper limit of memory usage. , when the node load is greater than the threshold When the load of a single node is less than the threshold, dynamic redistribution is triggered. , avoid single node overload and ensure that the entire system is always in the optimal operating state;

[0067] After load balancing, the priority mechanism is used to synchronize important data blocks first to ensure the system's disaster recovery capability:

[0068]

[0069] in, is the synchronization priority indicator of node u. The smaller the value, the closer the synchronization state of the current node is to the target state, the higher the priority, and the earlier it should be synchronized. g is the total number of objects to be synchronized on the node. is the weight coefficient of the jth data unit, reflecting the priority of data unit j, and is set through machine learning. is the expected synchronization progress of the jth data unit, is the actual synchronization progress of the jth data unit, and q is an exponential parameter for the adjustment amplitude, which is used to adjust the sensitivity of synchronization deviation. The larger q is, the greater the weight of the data unit with larger deviation in the overall priority;

[0070] The synchronization status is achieved and confirmed through the Raft protocol, according to the synchronization priority index Adjust the order of node synchronization requests. The higher the priority, the earlier the node will participate in the synchronization process. Each node will send synchronization requests to other nodes during the negotiation process, and data synchronization will only be carried out after reaching an agreement.

[0071] During synchronization, the optimal latency path is dynamically selected based on network latency, improving overall system responsiveness and reliability:

[0072]

[0073] in, is the maximum delay of node v, is the set of dependent nodes of node v, is the synchronous arrival time of node u, is the synchronization delay of node u;

[0074] During the synchronization process, the system synchronization efficiency and node robustness are monitored in real time to ensure the system continues to operate efficiently. Synchronization efficiency monitoring:

[0075]

[0076] in, is the synchronization efficiency of node u;

[0077] Setting thresholds through fuzzy logic , when the synchronization efficiency Less than threshold When , node reselection and disaster recovery switching are triggered, otherwise normal synchronization is performed;

[0078] After synchronization is completed, a cyclic redundancy check (CRC) is used to verify data integrity and ensure that the data is not damaged. If an anomaly is found, the disaster recovery mechanism is immediately triggered.

[0079] The cyclic redundancy check CRC is:

[0080]

[0081] in, It is a copy generated for each data, indicating the message that needs to be CRC checked. It is the index used for bit shift operations, which determines the number of bits that the data needs to be shifted left, and is usually related to the length of the data. is the length of the CRC polynomial;

[0082] When the CRC check result is 0, the data is correct. Otherwise, the data is wrong and the backup node is automatically switched to restore the complete data.

[0083] Through encrypted data segmentation and dynamic load balancing adjustment, resource allocation can be optimized under different node load conditions, single node overload can be avoided, and the system stability and processing capacity can be improved. Through the priority mechanism and Raft protocol to optimize data synchronization, high-priority data blocks can be synchronized first, improving the efficiency and consistency of data synchronization and reducing unnecessary delays. At the same time, the Raft protocol provides an efficient consistency mechanism to ensure that the system can reliably achieve synchronization in a distributed environment. Through network delay and synchronization efficiency monitoring, the system can be quickly adjusted in the face of network fluctuations and node failures, improving the system's stability and emergency response capabilities, and ensuring high reliability and fault tolerance of data synchronization. Through the combination of CRC verification and disaster recovery mechanism, the system can automatically repair data when it is damaged or the node fails, ensuring data integrity and reliability and enhancing the system's disaster recovery capabilities.

[0084] S3. During user data access, a zero-trust data tunnel is established and unified identity authentication and access control is performed. Through the isolation forest model, LSTM time series prediction model, and user behavior profile comparison, continuous analysis and detection of data security are carried out and security protection measures are generated.

[0085] Specifically, in the process of user access to data, a zero-trust data tunnel is established and unified identity authentication access control is performed. This means that when users access data, a zero-trust data tunnel is built, and the tunnel authentication information is encrypted using the RSA algorithm through the IPSec tunnel mode to ensure that the identity cannot be forged. After the tunnel is established, the source node identity of each data stream is determined, and the X.509 certificate chain is used to bind the device. The SHA-256 algorithm is applied to the data to generate a hash value to ensure that the data cannot be tampered with. The source ID, access permission level, and processing requirements (encryption storage requirements) fields are embedded in the data packet for overall encryption to prevent intermediate leakage or tampering. As the data flows through the tunnel, the relay node hop-by-hop verification mechanism is used (authentication is performed at each gateway) to perform three real-time verifications at each hop node (source legitimacy verification, tampering detection, and access purpose verification) to prevent tampering and illegal access, combined with OAuth 2.0 and national encryption SM2 / SM3 / SM4 CA certificates and two-factor authentication are used for unified authentication of the access control platform. Based on the permission matrix, a hybrid model of role-based access control and attribute-based access control is adopted for dynamic authorization. Static role-behavior-resource tables (RBAC foundation) are defined, dynamic attribute factors (ABAC extension) are defined, and a joint permission matrix is constructed:

[0086]

[0087] Where r is the behavior, is the behavior, res is the resource, ctx is the context attribute (IP, device type), and are all Boolean values (1 or 0);

[0088] When the product is 1, access is allowed; otherwise, access is denied.

[0089] Through the X.509 certificate chain and SHA-256 hash generation, the integrity of the data transmission process is ensured. Any tampering will cause the hash value to mismatch, ensuring the authenticity and credibility of the data. Through the hop-by-hop verification mechanism and multiple authentication methods, the intermediate nodes are effectively prevented from being attacked or tampered, and the system's resistance to malicious nodes or tampering is enhanced, ensuring the data consistency and security between each node. Access control is carried out through OAuth 2.0 and two-factor authentication, which improves the security of identity authentication and access authorization. It can flexibly adjust permissions according to user identity and access scenarios, ensure fine-grained control of data access, and prevent illegal access and abuse of permissions. Dynamic authorization is carried out through the joint permission matrix, which effectively improves the accuracy and flexibility of access control. It can adjust access rights in real time according to user behavior and environmental factors, ensuring secure access in a dynamic environment.

[0090] Furthermore, through the isolation forest model, LSTM time series prediction model and user behavior profile comparison, continuous analysis and detection of data security and generation of security protection measures means that every data access, modification and deletion action is synchronously recorded in the log (user unique identifier, data object unique identifier, operation type, operation timestamp and source IP address). Using blockchain technology, each log is stored in a sequential chain to ensure that once written, the log cannot be tampered with. Tampering with any log will cause the subsequent chain to break and be detected. Through the isolation forest model, LSTM time series prediction model and user behavior profile comparison, triple superposition anomaly detection is carried out;

[0091] The isolation forest detection is used to detect abnormal access frequency (such as accessing high-frequency or rare resources in a short period of time), synchronize the latest block from the on-chain node, extract the log records in each block in chronological order, add them to the log sequence, and set a sliding window. ,in It is When a new block is generated on the chain, the analysis is automatically triggered, and all logs in the new block are extracted and inserted into the window one by one. Get all log data sets in the window:

[0092]

[0093] in, is the set of all access log transactions within the sliding window. is the latest confirmed block number, is the current block number (block height), ranging from arrive , is the sliding window size (number of blocks). If the number of logs in the window exceeds , then remove the earliest log to ensure that the sliding window always keeps the latest Block log data, It is Blocks The Access logs (raw data);

[0094] The logs in the sliding window are converted into feature vectors through feature engineering. The features include source IP address segment (converted into numerical features through hash mapping), access time (normalized to the interval [0,1] using min-max), accessed resource path (one-hot encoding), response code (categorical variable), and source block number (retaining numerical features) to obtain a feature set. ,in , , use k-means clustering to divide the log into (determined by the elbow rule) different clusters (subsets) to reveal access behavior patterns, and initialize cluster centers using k-means++ , effectively avoiding the convergence problem caused by random initialization, each cluster center , assign the feature vectors in the feature set to the nearest cluster center and update each cluster center:

[0095]

[0096] in, is the updated cluster center, is the cluster center The total number of eigenvectors in , belongs to the cluster center The eigenvector of

[0097] Setting thresholds using the K-means clustering algorithm , when the change of cluster center is less than the threshold When , stop the iteration and get the clustering result , each cluster contains a set of eigenvectors and corresponding logs, and the intra-cluster density is calculated:

[0098]

[0099] in, It is The lower the score, the more concentrated the cluster. is the cluster center The total number of eigenvectors in , belongs to the cluster center The eigenvector of It is a cluster the center, is the Euclidean distance, which is used to measure the distance between the feature vector and the cluster center;

[0100] Calculate the compactness scores of all clusters, sort them in descending order, select the cluster with the lowest compactness (most compact) as the representative cluster, collect all sample feature vectors in the cluster, and construct a sample set ;

[0101] in, It is The feature vectors belonging to the representative cluster, d is the feature dimension, is the number of samples in the cluster;

[0102] For each feature dimension in the feature vector, calculate the statistics (minimum, maximum, mean) to form a statistical vector :

[0103]

[0104] in, It is a dimension The statistical vector of is the feature dimension The minimum value of is the feature dimension The maximum value of is the feature dimension The mean of

[0105] Combine the statistical vectors of all dimensions to obtain the characteristic statistical matrix (used to reflect the numerical distribution characteristics of samples in each dimension in the representative cluster), where d is the feature dimension, 3 is min, max, mean, through the recent Select data from the consensus-completed blocks (extract useful information from the part that has been accepted and confirmed as valid and permanent by the blockchain network) to construct the sample space:

[0106]

[0107] in, is the total sample space of the isolation forest, It is blocks that have completed consensus, is the latest confirmed block number, is the sliding window size;

[0108] In the sample space In the block height, the modular operation sampling is performed to define the subsample set :

[0109]

[0110] in, It is a sample The height of the block, is the sampling step size, which is adjusted according to the access frequency of the block to control the sampling interval;

[0111] Use the Box-Plot algorithm to analyze the subsample set Remove outliers, improve sample quality, and calculate the quartiles of all sample feature dimensions 、 , interquartile range ;

[0112] When the outlier meets:

[0113]

[0114] Then the sample Is an abnormal sample from Otherwise, the samples are retained normally for model training;

[0115] The subsample set after removing outliers With the statistical matrix Perform dimensional consistency check and confirm subsample set Match the feature dimensions and statistical features of the model to ensure the consistency of feature interpretation of subsequent models;

[0116] By sub-sample set Set forest parameters, including the number of trees in the forest Z and the sample subset size of each tree ,Right now =| |, the maximum tree depth is set to:

[0117]

[0118] in, is the maximum depth of the isolation tree, ensuring that the average number of splits does not exceed the number of samples;

[0119] During the construction of each isolated tree, when the internal node is split, a feature dimension is randomly selected from the dimension set and the selected matrix is selected. Extract the statistical set of feature dimensions , in turn as candidate segmentation thresholds, As the preferred split point (candidate threshold), construct the split condition:

[0120]

[0121] in, It is a sample In the feature dimension The value of ;

[0122] If it cannot be split, use Construct the splitting conditions. When all three cannot be split further (for example, all samples are the same or fall on the same side), it degenerates into ordinary random splitting:

[0123]

[0124] in, From the feature dimension Minimum value of to the maximum value In the range, randomly sample a value according to the uniform distribution. It is a continuous uniform distribution within this range;

[0125] The current node stops splitting and is marked as a leaf node when any of the following conditions are met:

[0126] The current node depth reaches the maximum depth hour;

[0127] When the number of current node samples is less than or equal to 1;

[0128] When the values of samples in the current node are the same in all dimensions;

[0129] For each isolated tree in all trees Z, record the node partition feature dimension, partition threshold (priority statistic partitioning, alternative random partitioning), left and right subtree pointers and node sample number to form an isolation forest ;

[0130] Exploiting Isolation Forests For each isolated tree in , calculate the log Path length:

[0131]

[0132] in, It's a log The average path length in the isolation forest, Z is the total number of isolated trees, In the i-th tree, the log The path length (the number of steps from the root to the split termination node);

[0133] Through the log Average path length in the isolation forest, calculate the anomaly score:

[0134]

[0135] in, It's a log The anomaly score, is the total number of logs, log The expected value of the path length in multiple isolated trees, is a normalization factor used to calibrate the expected value of paths under different data sizes;

[0136] The calculation formula is:

[0137]

[0138] in, is the Euler–Mascheroni constant;

[0139] Setting thresholds through statistical analysis ,when Greater than threshold , it is abnormal, otherwise it is normal;

[0140] The LSTM time series prediction model is used to detect abnormal access time series patterns (such as abnormal access period and abnormal operation sequence). Serialize into time window data Input into the time series forecasting model to predict the sequence , calculate the difference between the true sequence and the predicted sequence by Euclidean distance:

[0141]

[0142] Adaptive threshold setting through dynamic distribution , when Error is greater than the set threshold If , it is abnormal, otherwise, it is normal;

[0143] The user behavior profile comparison is constructed by historical audit logs to build the user's standard behavior profile G (common operation set, common access time distribution, common source IP address range) and log Compare information:

[0144]

[0145] in, is the similarity between the current log and the portrait, is the total number of features used for comparison, It is the number of expected characteristics in the current behavior and user profile;

[0146] Setting thresholds through fixed ratio screening ,when Less than threshold When , it is marked as abnormal access, otherwise, normal;

[0147] Based on the Isolation Forest Anomaly Score , LSTM prediction error Similarity between the current log and the portrait Comprehensive judgment abnormality:

[0148]

[0149] For abnormal access, when all three judgment conditions are abnormal, the current user session will be immediately interrupted, the user account will be frozen, login will be prohibited, and the investigation work order will be automatically pushed to the security audit system. When two judgment conditions are abnormal, a real-time security alert will be sent to the SOC, and user permissions will be temporarily restricted (such as restricting read and write operations and retaining only read-only permissions). The user will be added to the key observation list and subsequent behavior will be continuously tracked. When only one condition is abnormal, a detailed audit log will be recorded, the abnormal characteristics will be retained, no restrictive measures will be taken for the time being, and continuous detection and observation will be carried out.

[0150] By recording each data access operation log and combining blockchain technology to ensure that the log cannot be tampered with, it provides reliable traceability for data access, enhances the transparency of data operations, and provides credible evidence for subsequent security audits. Through log recording and feature engineering conversion, it ensures a comprehensive understanding of data behavior, provides rich feature input for subsequent anomaly detection, and ensures that data can be input into the isolation forest model in a structured manner, so that the model can extract valuable information from the log data. Through K-means clustering and identification of representative clusters, it can accurately grasp the most representative behavior patterns, further improving the sensitivity and accuracy of the isolation forest model to abnormal behavior. By calculating the statistics of each feature, it can fully grasp the distribution characteristics of the data in each dimension, helping the isolation forest model to better understand the normal behavior of the data, thereby improving the ability to detect anomalies. By sampling from different blocks, it can construct a diverse training sample space, increase the representativeness of the samples, improve the isolation forest's ability to recognize different abnormal patterns, and avoid time mismatch or loss between samples. The ox-Plot algorithm can effectively remove extreme outliers in the data, allowing the isolation forest model to be trained based on high-quality data, avoiding the impact of noise data on model performance, reducing training errors caused by abnormal data, and improving the accuracy of anomaly detection. Through the splitting and path calculation of the isolation tree, it can effectively separate abnormal samples from normal samples, thereby improving detection accuracy. By randomly selecting feature dimensions for splitting, it can effectively capture complex nonlinear relationships in the data, thereby improving the adaptability and detection capabilities of the model. Through real-time calculation of anomaly scores, it can respond to abnormal access immediately, thereby effectively improving the real-time and flexibility of data protection. Through the LSTM time series prediction model, abnormal access time series patterns are detected, enabling the system to predict and identify abnormal patterns in time, improving the ability to prevent potential abnormal behaviors and enhancing the intelligence of data protection. Through user behavior profile comparison, a standard behavior model is constructed, enabling the system to accurately identify abnormal deviations in user behavior, further strengthening the system's ability to identify abnormal access.

[0151] This embodiment also provides an information security management system based on sensitive data, including:

[0152] Data collection and sensitive identification module, used to uniformly collect data and identify sensitive information, generate labels and sensitivity levels;

[0153] Dynamic sensitivity prediction module, used for LSTM to predict sensitivity changes and dynamically adjust labels and permissions;

[0154] Hierarchical encryption mechanism module, used for encryption based on sensitivity classification, using SM4, AES, and ChaCha20 algorithms;

[0155] A secure storage disaster recovery module, used to integrate quantum key and chaotic encryption to build a highly secure and disaster-tolerant distributed system;

[0156] Zero Trust Access Control module, used to establish a zero trust tunnel and use multi-factor and certificate chain unified authentication;

[0157] The anomaly detection and protection module is used to combine the forest model with behavioral profiling to detect anomalies and provide protection in multiple dimensions.

[0158] This embodiment also provides a computer device, which is suitable for an information security management method based on sensitive data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement an information security management method based on sensitive data proposed in the above embodiment.

[0159] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.

[0160] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements an information security management method based on sensitive data as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

Claims

1. A method for information security management based on sensitive data, characterized by: include, Collect data and identify sensitive information to generate labels, classify data sensitivity levels and build a data access permission mapping matrix. Use time series prediction models to conduct real-time assessments of data sensitivity changes and dynamically update labels and corresponding permission matrices. Execute a hierarchical encryption strategy on the data based on the update results, and store the encrypted data in a distributed redundant disaster recovery system that integrates quantum key distribution and dynamic chaotic key stream encryption; During the process of user access to data, a zero-trust data tunnel is established and unified identity authentication access control is performed. Through the isolation forest model, LSTM time series prediction model and user behavior profile comparison, data security is continuously analyzed and detected, and security protection measures are generated.

2. The information security management method based on sensitive data according to claim 1, characterized in that: The continuous analysis and detection of data security and generation of security protection measures through the isolation forest model, LSTM time series prediction model and user behavior profile comparison refers to the use of blockchain technology to chain log records for each data access, modification and deletion action, storing each log in sequence in a chain, and performing triple superposition anomaly detection through the isolation forest model, LSTM time series prediction model and user behavior profile comparison.

3. The sensitive data-based information security management method according to claim 2, characterized in that: The establishment of a zero-trust data tunnel and unified identity authentication access control during the user access to data refers to the construction of a zero-trust data tunnel when the user accesses the data, the use of IPSec combined with RSA to encrypt authentication information, binding X.509 device certificates, and using SHA-256 to generate data hash values. The data flow path passes through a hop-by-hop triple verification mechanism through relay nodes, combined with OAuth 2.0, the national secret SM series algorithm and two-factor authentication, access to a unified authentication platform, dynamic authorization based on the RBAC and ABAC fusion model, and the construction of a joint permission matrix combining static roles and dynamic attributes. When the matrix judgment result is 1, access is authorized, otherwise it is denied.

4. The sensitive data-based information security management method according to claim 3, characterized in that: The encrypted data is stored in a distributed redundant disaster recovery system that integrates quantum key distribution and dynamic chaotic key stream encryption. The encrypted data is encapsulated using CMS, pushed to the target distributed storage node using the Kafka message queue, and a one-time key encryption link is established through quantum key distribution technology for data protection. All data blocks are further dynamically encrypted using dynamic chaotic key stream encryption to generate a continuous key stream value. ; from A part of the key is sampled for block-level encryption, the encrypted data is logically divided into blocks through hyperbolic cyclic mapping, the bilinear differential scheduling function is used to dynamically adjust the load balancing, and the synchronization priority index is calculated through the priority mechanism. , prioritize synchronization of important data blocks; During synchronization, the optimal delay path is dynamically selected based on network delay , real-time monitoring system synchronization efficiency and node robustness, calculate the synchronization efficiency of monitoring node u , set the threshold , when the synchronization efficiency Less than threshold When the node is in the state of failure, node reselection and disaster recovery switching are triggered, otherwise normal synchronization is performed. After the synchronization is completed, a cyclic redundancy check (CRC) is used to verify the integrity of the data.

5. The information security management method based on sensitive data according to claim 4, characterized in that: The execution of a hierarchical encryption strategy for data according to the update result refers to encrypting the highly sensitive data, the medium sensitive data and the low sensitive data respectively according to the updated adjustment; For highly sensitive data, the national secret SM4 symmetric encryption is used. The SM4 encryption algorithm is used to encrypt in the form of a block cipher, and the individual encrypted blocks are finally connected to form a ciphertext; For medium-sensitive data, it is encrypted in the form of a block cipher using AES-256, the ciphertext encrypted by AES is hashed, and the private key is used to perform ECDSA signature to generate a signature; For low-sensitivity data, ChaCha20 is used to encrypt with a 256-bit key and a 64-bit random number, and discrete cosine transform is used to embed watermarks in the low-frequency part of the encrypted data.

6. The information security management method based on sensitive data according to claim 5, characterized in that: The time series prediction model is used to evaluate the change of data sensitivity in real time, dynamically update the label and the corresponding authority matrix to monitor the indicator in real time, obtain the historical sensitivity score sequence, normalize the sensitivity score through Z-Score to generate training samples, input them into the double-layer stacked LSTM network, extract the local time series change characteristics and the global sensitivity evolution trend, obtain the sensitivity prediction sequence, and calculate the mean sensitivity change during the prediction period. , set thresholds b and e, and b is greater than e, when When it is greater than the threshold b, the sensitivity is on an upward trend, the data label T is automatically upgraded, the sensitivity comprehensive score is recalculated, and the authority matrix P is tightened. When it is less than e, the sensitivity is decreasing and the access rights are relaxed. When it is greater than e and less than b, the sensitivity is stable, the status quo is maintained, and only logs are recorded.

7. The information security management method based on sensitive data according to claim 6, characterized in that: The data collection and identification of sensitive information to generate labels, and the classification of data sensitivity levels to build a data access permission mapping matrix refers to collecting internal, external and terminal data, uniformly streaming them into the Apache Flink cluster, identifying sensitive information based on preset rules and the NER model, extracting contextual features to train the decision tree to generate label triples T, and building a permission matrix based on roles and sensitivity scores. , score each data according to the T label and calculate the comprehensive sensitivity score S; Set permission control thresholds A and B, and A is greater than B. When S is greater than A, it is highly sensitive data and only system administrators can fully access it. When S is greater than B and less than or equal to A, it is medium-sensitive data and internal employees can read but not modify it. When S is less than or equal to B, it is low-sensitive data and is open to external collaborators for readable access.

8. An information security management system based on sensitive data, based on the information security management method based on sensitive data according to any one of claims 1 to 7, characterized in that: include, Data collection and sensitive identification module, used to uniformly collect data and identify sensitive information, generate labels and sensitivity levels; Dynamic sensitivity prediction module, used for LSTM to predict sensitivity changes and dynamically adjust labels and permissions; Hierarchical encryption mechanism module, used for encryption based on sensitivity classification, using SM4, AES, and ChaCha20 algorithms; A secure storage disaster recovery module, used to integrate quantum key and chaotic encryption to build a highly secure and disaster-tolerant distributed system; Zero Trust Access Control module, used to establish a zero trust tunnel and use multi-factor and certificate chain unified authentication; The anomaly detection and protection module is used to combine the forest model with behavioral profiling to detect anomalies and provide protection in multiple dimensions.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the information security management method based on sensitive data described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the information security management method based on sensitive data described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Enterprise sensitive data security access management method and system

    CN118656870A

  • Data security management system and method in smart power grid

    CN119598484A

  • A dynamic data encryption method and device integrating quantum state and chaotic system

    CN119788274A

  • Sequential Encryption Method Based On Multi-Key Stream Ciphers

    US20190207745A1

Cited By

  • Dynamic authority intelligent contract generation method and system based on data sensitivity self-adaption

    CN120850349A

  • Data security transmission method based on digital archive multi-protection

    CN120880802A

  • Data access security intelligent management method based on network intrusion detection

    CN121012668A

  • A Smart Management Method for Data Access Security Based on Network Intrusion Detection

    CN121012668B

  • Private network elastic security situation awareness method and device based on time-space bimodal

    CN121463041A