Data encryption method, encryption equipment and storage medium
By analyzing the multidimensional features of data and dynamically generating encryption keys, this method solves the problem of balancing security strength and efficiency in existing encryption methods, and achieves fine-grained differentiated protection and improved anti-attack capabilities, making it suitable for data encryption in complex scenarios.
Patent Information
- Application Number
- CN202511542660.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-12-05
AI Technical Summary
Existing encryption methods are difficult to dynamically adapt to differences in data complexity and content sensitivity, making it difficult to balance security strength and processing efficiency. They cannot meet the refined security protection requirements in complex scenarios, and the static key management is vulnerable to attacks.
By analyzing the multidimensional features of the data, the system can match the best encryption strategy in real time and dynamically generate encryption keys. It can also use data sensitivity features, business risk features, and user behavior features to perform quantitative scoring, combine them with a decision tree model to determine the security level, and select appropriate encryption algorithms and key generation methods according to the level.
It achieves fine-grained differentiated protection based on data characteristics, improves the efficiency of encrypted resource utilization, enhances anti-attack capabilities, and dynamically generated keys improve unpredictability, reduce the risk of brute-force attacks and side-channel attacks, and is suitable for the data encryption needs of economic industries in high-concurrency and multi-scenario applications.
Smart Images

Figure CN121077809A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data encryption technology, and in particular to a data encryption method, encryption device and storage medium. Background Technology
[0002] Mainstream encryption methods generally adopt an architecture that combines fixed encryption algorithms with periodic key management. At the encryption algorithm application level, they mostly rely on fixed strategies (i.e., single or fixed combinations of encryption algorithms) and fail to dynamically adapt to differences in data complexity and content sensitivity. At the key management level, they are usually based on rotating symmetric encryption keys at fixed periods (such as monthly or quarterly), and their key generation patterns are easily captured by brute-force attacks or side-channel attacks.
[0003] Faced with the massive and multi-dimensional security needs of industrial data, the above encryption methods are unable to effectively balance security strength and processing efficiency, resulting in a waste of encryption resources and failing to meet the refined security protection requirements in complex scenarios. Summary of the Invention
[0004] The purpose of this application is to provide a data encryption method, encryption device, and storage medium, which aims to solve the problems that current encryption methods cannot achieve differentiated protection based on the characteristics of the data itself, and the security risks caused by static key management.
[0005] To address the aforementioned technical problems, embodiments of this application provide a data encryption method applied to an encryption device. The method includes: acquiring data to be encrypted uploaded by a terminal; extracting multiple data features from the data to be encrypted, wherein the data features include data sensitivity features, business risk features, and user behavior features; quantifying and scoring the multiple data features to obtain multiple feature scores, and calculating a total data score based on the multiple feature scores; inputting the total data score into a pre-trained decision tree model to determine the security level of the data to be encrypted; determining a target encryption algorithm based on the security level from a pre-defined correspondence between security levels and encryption algorithms; and encrypting the data to be encrypted using the target encryption algorithm and a dynamically generated encryption key.
[0006] Embodiments of this application also provide an encryption device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the data encryption method described above.
[0007] Embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described data encryption method.
[0008] The technical solutions provided in this application have at least the following beneficial effects: The technical solution provided in this application aims to improve the efficiency of encryption resource utilization while ensuring high security by analyzing the multi-dimensional features of data to match the best encryption strategy in real time and dynamically generating encryption keys. In this method, after the encryption device obtains the data to be encrypted uploaded by the terminal, it extracts three core features from the data: data sensitivity features, business risk features, and user behavior features. These three data features correspond to the core influencing factors of data security, namely the sensitivity of the data itself, the risk level of the business it belongs to, and the behavioral risk of the accessing user. Next, each of these three core features is quantitatively scored to obtain a feature score, transforming qualitative security requirements into calculable objective values and ensuring the standardization of security level classification. Through a pre-trained decision tree model, the security level is automatically output based on the total data score, achieving fine-grained differentiated protection and avoiding the imbalance in protection between high and low sensitive data; furthermore, different levels are output according to subtle differences in data features to avoid the imbalance in protection between high-sensitive and low-sensitive data at the same level. Then, according to the predefined correspondence between security levels and encryption algorithms, the encryption strength is matched, ensuring that high-value data receives high-strength encryption protection, while low-value data avoids redundant computational resources, effectively balancing security and efficiency. Unlike fixed-period key management, this scheme dynamically generates encryption keys, improving key unpredictability, significantly reducing the risk of brute-force attacks and side-channel attacks, and enhancing the system's resistance to attacks. This leads to the construction of an encryption method that dynamically adapts encryption strength and key generation based on data characteristics, ensuring the core data's resistance to attacks while achieving intelligent optimization of encryption resource allocation.
[0009] Furthermore, the generation of the encryption key includes: acquiring the load data of the encryption device; adding noise to the load data; and then dynamically generating the encryption key based on the noise-added load data. In this way, through dynamic load correlation and noise-added privacy protection, the predictability problem of traditional fixed keys is solved, the anti-attack capability is strengthened, and a dynamic balance between security and efficiency is achieved. This is particularly suitable for the data encryption needs of high-concurrency, multi-scenario economic industries. Attached Figure Description
[0010] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0011] Figure 1 This is an exemplary flowchart of a data encryption method according to some embodiments of this application; Figure 2 This is an exemplary flowchart of a method for dynamically generating encryption keys according to some embodiments of this application; Figure 3 This is a schematic diagram of the structure of an encryption device according to some embodiments of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.
[0013] It should be understood that the terms "system," "device," "unit," and / or "module" used in this specification are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other words can achieve the same purpose, they may be replaced by other expressions.
[0014] Unless otherwise specified, the technical terms used to describe components, elements, etc. in this specification are not singular but may include plural. Generally speaking, terms such as "comprising" or "including" only indicate that explicitly identified steps, elements, or components are included, and these steps, elements, and components do not constitute an exclusive list, as the described method or apparatus may also include other steps or components.
[0015] This specification uses flowcharts to illustrate the operational steps performed by the apparatus or system of related embodiments. However, unless otherwise specified, the order in which these steps are described should not be construed as a limitation on the order of execution. Those skilled in the art can adjust the order of these steps based on the knowledge and information conveyed by the embodiments in this specification. Adjustments include, but are not limited to, reversing the order of steps, merging multiple steps, and splitting a step.
[0016] Mainstream encryption methods generally employ an architecture combining fixed encryption algorithms with periodic key management. At the algorithm application level, they often rely on fixed strategies (i.e., single or fixed combinations of encryption algorithms), failing to dynamically adapt to differences in data complexity and content sensitivity. At the key management level, they typically rotate symmetric encryption keys on a fixed period (e.g., monthly or quarterly), making their key generation patterns vulnerable to brute-force attacks or side-channel attacks. However, when faced with the security needs of massive, multi-dimensional industrial data, these encryption methods struggle to effectively balance security strength and processing efficiency, leading to wasted encryption resources and failing to meet the refined security protection requirements of complex scenarios.
[0017] To address the aforementioned issues, some embodiments of this application provide a data encryption method aimed at improving the efficiency of encryption resource utilization while ensuring high security by analyzing the multi-dimensional features of data to match the optimal encryption strategy in real time and dynamically generating encryption keys. In this method, an encryption device acquires the data to be encrypted uploaded by the terminal and extracts multiple data features from the data, including data sensitivity features, business risk features, and user behavior features. These features are then quantified and scored to obtain multiple feature scores, and a total data score is calculated based on these scores. The total data score is input into a pre-trained decision tree model to determine the security level of the data to be encrypted. Based on the security level, a target encryption algorithm is determined from a pre-defined correspondence between security levels and encryption algorithms. Finally, the data to be encrypted is encrypted using the target encryption algorithm and the dynamically generated encryption key. This method solves the problems of related encryption methods' difficulty in achieving differentiated protection based on the data's own characteristics and the security risks caused by static key management.
[0018] Figure 1 This is an exemplary flowchart illustrating a data encryption method according to some embodiments of this application, which can be executed by an encryption device. In some embodiments, Figure 1 The process shown may include the following steps.
[0019] Step 110: The encryption device acquires the data to be encrypted uploaded by the terminal and extracts multiple data features from the data to be encrypted.
[0020] Data characteristics include data sensitivity characteristics, business risk characteristics, and user behavior characteristics.
[0021] In some embodiments, data sensitivity features include the level, density, and field relevance factor of sensitive fields in the data to be encrypted. Optionally, in this embodiment, a field identification model based on the BERT (Bidirectional Encoder Representations from Transformers) architecture is deployed on the encryption device. This model, in conjunction with a regular expression matching engine, scans structured data and performs entity alignment by calling the National Sensitive Information Directory Application Programming Interface (API), thereby accurately extracting sensitive fields from the data to be encrypted. This process simultaneously determines the level of each sensitive field (e.g., distinguishing between highly sensitive fields such as financial transaction records, semi-sensitive fields such as company registration addresses, and non-sensitive fields such as publicly available industrial policies) and calculates the sensitive field density. This scheme achieves differentiated identification of data privacy value.
[0022] In some embodiments, business risk characteristics include the regulatory level, influence coefficient, and scope of impact of the data to be encrypted within the industry. Optionally, in this embodiment, the encryption device connects to the real-time database of the national data supervision platform to obtain the regulatory level of the industry to which the data to be encrypted belongs. The device further uses an industry influence assessment model (which integrates GDP contribution and employment elasticity coefficient) to calculate the industry's influence coefficient and employs Fault Tree Analysis (FTA) to assess the scope of impact of the data breach. Based on the above characteristics, the encryption device can achieve differentiated identification of business scenario risks (e.g., determining that bank data has a strong regulatory level and a high influence coefficient). Optionally, the assessment of the scope of impact of the breach can be combined with Monte Carlo simulation methods. This solution can dynamically adjust security strategies according to industry characteristics and potential impact.
[0023] In some embodiments, user behavior characteristics include the frequency of user access to encrypted data, operation type, user permissions, and behavioral credibility. Optionally, in this embodiment, the encryption device deploys a blockchain distributed log system based on the Hyperledger Fabric log framework to record user operation behavior. This system associates with a Lightweight Directory Access Protocol (LDAP) permission directory service to construct a role-based access control (RBAC) matrix and calculates user access frequency using a sliding window algorithm. Optionally, off-chain caching and batch on-chain mechanisms (e.g., supporting TPS≥1000) are employed to ensure the real-time performance of data processing. This scheme achieves differentiated identification of user access risks.
[0024] The extraction of data sensitivity features, business risk features, and user behavior features enables a comprehensive quantification of data value and risk, providing a precise basis for subsequent hierarchical encryption strategies and effectively overcoming the problems of over-protection or under-protection caused by uniform processing in traditional encryption schemes. Furthermore, each feature data is processed using the SHA-256 hash algorithm and stored on the blockchain, with immutable evidence storage achieved by recording block headers containing Merkle root values and timestamps.
[0025] Step 120: Quantify and score multiple data features to obtain multiple feature scores, and calculate the total data score based on the multiple feature scores.
[0026] Traditional encryption methods, lacking a data grading mechanism, typically employ a single algorithm to process all data. Applying strong encryption to low-value data excessively consumes CPU and memory resources; conversely, using weak encryption algorithms for high-value data risks data leakage, resulting in a severe imbalance between security and resource efficiency. Therefore, this embodiment provides a decision-making basis for differentiated encryption through a quantitative scoring mechanism.
[0027] In some embodiments, multiple data features are quantitatively scored to obtain multiple feature scores, and the total data score is calculated based on the multiple feature scores. This can be achieved in the following ways: 1. Quantify the level and density of sensitive fields, as well as the field correlation factor, to calculate the feature score of the data sensitivity feature; 2. Quantify the regulatory level, influence coefficient, and leakage impact scope to calculate the feature score of the business risk feature; 3. Quantify the access frequency, operation type, user permissions, and behavior credibility to calculate the feature score of the user behavior feature; 4. Weight the aforementioned three feature scores according to a preset weighting rule (data sensitivity feature 0.4, business risk feature 0.3, user behavior feature 0.3) and sum them.
[0028] An example of an optional implementation of feature scoring, taking the calculation of feature scores for data-sensitive features as an example. This embodiment quantifies data-sensitive features through the following elements: sensitive field density (i.e., the proportion of sensitive fields in the data); sensitive field level, which is divided according to sensitivity (such as PII or SPI) and assigned a score range (example: sensitive field: 8-10 points; semi-sensitive field: 4-7 points; non-sensitive field: 1-3 points); and field relevance factor, which uses the Pearson correlation coefficient (…). Assess the strength of the association between fields. When When this occurs, a correlation weighting mechanism is triggered (example: add 1-2 points). The data sensitivity feature score can be calculated using the following formula: Data sensitivity feature score = [Σ(Sensitive field level score × Sensitive field density)] × 0.4 + Correlation weighting score.
[0029] An example of an optional implementation of feature scoring, taking the calculation of feature scores for business risk features as an example. This embodiment quantifies the characteristics of business risk features through the following elements: based on the regulatory level obtained from the API of the "Data Classification and Grading Guidelines" and the industry influence coefficient, the impact range parameter after data leakage is introduced to quantify the business risk features. Specifically, the business risk feature score is calculated as follows: Business Risk Score = Regulatory Level × Industry Influence Coefficient × 0.3 + Impact Range Weighted Score. Typical scoring examples are as follows: Data with a strong regulatory level and a high industry influence coefficient (e.g., bank credit data, power grid dispatch data) scores 8 to 10 points; data with a medium regulatory level and a medium industry influence coefficient (e.g., retail sales data) scores 4 to 7 points; data with a weak regulatory level and a low industry influence coefficient (e.g., publicly available survey data) scores 1 to 3 points.
[0030] An example of an optional implementation of feature scoring, taking the calculation of feature scores for user behavior features as an example. This embodiment quantifies the user behavior feature score by considering factors such as the frequency of user access to the data to be encrypted (in minutes), access control permissions (based on a Role-Based Access Control (RBAC) matrix), and operation type (with a combined weight of 0.3), combined with a behavior credibility score derived from Bayesian estimation based on historical audit logs (1 point is added for compliant records, and 1 to 2 points are deducted for violations). The specific calculation method is: User behavior score = (Access frequency score + Permission score) / 2 × 0.3 + Credibility score. Typical scoring examples are as follows: 8 to 10 points are awarded for high-frequency access behavior (e.g., more than 10 accesses per day) and high permissions (e.g., read, write, and delete permissions); 4 to 7 points are awarded for medium-frequency access behavior (e.g., 3 to 10 accesses per day) and medium permissions (e.g., read-only permissions only); and 1 to 3 points are awarded for low-frequency access behavior (e.g., less than 3 accesses per day) and low permissions (e.g., query permissions only).
[0031] Step 130: Input the total data score into a pre-trained decision tree model to determine the security level of the data to be encrypted.
[0032] In step 130, a multi-dimensional data classification decision tree strategy is adopted. Through three-dimensional feature quantization and multi-node reasoning, four security levels are output: public, general confidential, highly confidential, and top-secret. The algorithm optimization strategy in this embodiment is as follows: To address the information loss problem caused by discretizing continuous attributes, a multi-interval partitioning mechanism based on information gain ratio is introduced: the optimal partitioning point set is determined by calculating attribute entropy and splitting information entropy. For example, for user access frequency features, a dynamic threshold generation algorithm is used to divide continuous values into intervals such as [0,2), [2,5), [5,10), and [10,+∞) to maximize information gain ratio and improve feature discriminability.
[0033] Construct a random forest model consisting of 10-15 CART trees: diversify the training set through bootstrap sampling and use the feature bag method to select the optimal feature subset; the model output adopts a weighted voting mechanism to assign higher confidence weights to key features such as business risk weights in order to reduce the overfitting risk of a single decision tree and improve the F1-score by 12%-15%.
[0034] A parameter optimizer based on reinforcement learning is introduced to monitor the model confusion matrix metrics (accuracy, recall, FPR) in real time: when the validation set accuracy is lower than the default threshold of 90%, the parameter adjustment process is automatically triggered, and the optimal hyperparameters (minimum number of split samples 5-20, maximum tree depth 8-16) are searched using the Bayesian optimization algorithm. The model parameters are updated through incremental training to ensure adaptive recognition of new sensitive fields.
[0035] This application, in this embodiment, employs an optimized decision tree model constructed using an integrated improved C4.5 algorithm and a random forest algorithm, mapping the total feature score to security levels of "public, general confidential, highly confidential, and top-secret." This decision tree is a four-layer optimized decision tree, with the decision nodes as follows: First layer: Initial classification based on total score thresholds (total score ≥ 8 points leads to the highly confidential / top-secret branch; total score between 5 and 7 points is classified as general confidential; total score ≤ 4 points is classified as public); Second layer: National-level economic lifeline data identification is performed on data leading to the highly confidential / top-secret branch (achieved through predefined keyword matching and regular expressions); Third layer: Strategic value assessment is performed on data identified as national-level economic lifeline data (using the Analytic Hierarchy Process (AHP) to calculate its irreplaceability index to distinguish between highly confidential and top-secret levels); Fourth layer: Emergency state trigger node (when temporary data involving national security is detected, regardless of the current processing level, it is given the highest priority and directly triggers the top-secret classification).
[0036] For example, the security level assessment process for the data to be encrypted is as follows. The data to be assessed contains the following characteristics and their scores: Sensitive data characteristics include enterprise credit reports (belonging to Personal Identity Information (PII), score 9 points), repayment records (belonging to PII, score 9 points), and enterprise names (belonging to semi-sensitive information, score 5 points), with a sensitivity density of 85% and a field correlation coefficient r=0.85. The calculated sensitive feature score is ((9×0.85)+(5×0.15))×0.4+1=4.38; Business risk characteristics are reflected in the banking regulatory level score of 9 points, industry influence coefficient of 0.9, and cross-institutional influence. The calculated risk weight score is (9×0.9)×0.3+1=3.43; User behavior characteristics are reflected in the administrator privilege score of 6 points, the average daily access of 12 times score of 9 points, and the credibility score of 1 point. The calculated user behavior score is ((6+9) / 2)×0.3+1=3.25. The total feature score is calculated as 4.38 + 3.43 + 3.25 = 11.06. Based on the pre-defined decision tree mapping rule (a total score ≥ 8 points leads to a high / top-secret branch), this data is assessed as "highly confidential".
[0037] Step 140: Based on the security level, determine the target encryption algorithm from the preset correspondence between security levels and encryption algorithms.
[0038] In step 140, the correspondence between security levels and encryption algorithms is as follows: if the security level is public, the corresponding encryption algorithm is SM4 encryption algorithm; if the security level is general confidential, the corresponding encryption algorithm is adaptive AES encryption algorithm, wherein the number of encryption rounds of the adaptive AES encryption algorithm is determined according to the load data of the encryption device; if the security level is high confidential, the corresponding encryption algorithm is RSA-ECC hybrid encryption algorithm; if the security level is top confidential, the corresponding encryption algorithm is CRYSTALS-Kyber quantum-resistant encryption algorithm.
[0039] This application's embodiments are based on federated learning to optimize parameters (such as dynamically adjusting the number of AES rounds), combined with system resource monitoring (CPU / memory load) and algorithm performance matrices, to select the optimal encryption combination in real time (such as using SM4 for public data and RSA-ECC hybrid encryption for confidential data). Specifically, due to the large volume of public-level data (such as publicly available industry survey data), the fixed 32-round structure of SM4 has a fast computation speed and requires no additional resource overhead, avoiding the resource waste of traditional fixed strong algorithms for encrypting low-value data. When the device processes a large amount of ordinary confidential data (such as encrypting sales data during peak periods in retail enterprises), a fixed 14 rounds under high load will lead to excessively long encryption time; while the adaptive number of rounds, optimized through federated learning, is reduced to 10 rounds under high load, which can effectively improve encryption efficiency, while restoring 14 rounds under low load to ensure security, achieving a dynamic balance effect perceived by the actual load. For highly confidential data, such as bank corporate credit data, core transaction records of supply chain finance, and non-national-level data of power grid dispatch, this type of data is highly sensitive and has high business risks, requiring simultaneous protection of data encryption security, key transmission security, and data integrity. Therefore, this hybrid algorithm is adopted to compensate for the shortcomings of a single algorithm. Top-secret data, such as national-level credit risk early warning data, core data of nationwide power grid dispatch, and production capacity data of industries vital to the national economy, are directly related to national security and financial stability and must be protected against the threat of quantum computing.
[0040] The correspondence between the preset encryption levels and algorithms can be shown in the table below.
[0041] Table 1
[0042] Step 150: Encrypt the data to be encrypted according to the target encryption algorithm and the dynamically generated encryption key.
[0043] Through steps 110 to 150, an encryption method based on dynamically adapting encryption strength and key generation according to data characteristics is constructed. This ensures the anti-attack capability of core data while achieving intelligent optimization and configuration of encryption resources. In steps 110 to 150, after the encryption device obtains the data to be encrypted uploaded by the terminal, it extracts three core features from the data: data sensitivity features, business risk features, and user behavior features. These three data features correspond to the core influencing factors of data security, namely the sensitivity of the data itself, the risk level of the business to which it belongs, and the behavioral risk of the accessing user. Next, each of these three core features is quantitatively scored to obtain feature scores, transforming qualitative security requirements into calculable objective values and ensuring the standardization of security level classification. Through a pre-trained decision tree model, the security level is automatically output based on the total data score, achieving fine-grained differentiated protection and avoiding the imbalance problem of protection for high and low sensitive data; furthermore, different levels are output according to subtle differences in data features to avoid the imbalance problem of protection for high and low sensitive data at the same level. Furthermore, based on the predefined correspondence between security levels and encryption algorithms, encryption strength is matched to provide high-value data with strong encryption protection, while avoiding redundant computational resources for low-value data, effectively balancing security and efficiency. Unlike fixed-period key management, the data encryption method provided in this application dynamically generates encryption keys, improving key unpredictability, significantly reducing the risk of brute-force attacks and side-channel attacks, and enhancing the system's resistance to attacks.
[0044] Figure 2 This is an exemplary flowchart of a dynamic encryption key generation method according to some embodiments of this application. In some embodiments, such as Figure 2 As shown, the encryption key generation process used in step 150 includes the following steps.
[0045] Step 210: Obtain the load data of the encrypted device; Step 220: Add noise to the payload data, and then dynamically generate an encryption key based on the noise-added payload data.
[0046] In this way, by dynamically associating the load and adding noise for privacy protection, the predictability problem of traditional fixed keys is solved, the anti-attack capability is strengthened, and a dynamic balance between security and efficiency is achieved. It is especially suitable for the data encryption needs of economic industries with high concurrency and multiple scenarios.
[0047] This application embodiment uses a non-intrusive acquisition system to collect real-time running features and construct a 128-dimensional feature vector to provide dynamic basic data for key generation. Specifically, it includes CPU load, memory page swapping efficiency, disk I / O response time, TCP establishment interval, and the Poisson process strength of the key generation request.
[0048] First, a multi-source heterogeneous data acquisition framework was adopted to extract 128-dimensional time-series feature vectors of 5 categories. ,; among which, at the hardware layer, CPU load statistics (mean) ,variance The 1-minute sliding window analysis meets the needs of monitoring dynamic changes in system resources. Related research shows that this short-time window statistics can effectively capture fluctuations in system operation; memory page swapping frequency. (Unit: Hz) and disk I / O response time distribution It is also an important indicator for evaluating hardware performance and stability. At the network layer, the exponential distribution parameter of the TCP connection establishment interval... Data packet retransmission rate This can reflect the network's load and reliability; similar network feature extraction is widely used in network security situational awareness research. In the application layer, the Poisson process strength of key requests... Process context switching frequency This is of great significance for analyzing the patterns and stability of applications' access to system resources.
[0049] This application's embodiments utilize eBPF kernel-mode probes to achieve non-intrusive data acquisition. This method has been widely used in recent operating system performance monitoring research because it can efficiently acquire data without affecting the normal operation of the system. Sampling period The sampling period can be adjusted according to the real-time data requirements of different systems. Related research indicates that this sampling period can effectively balance data acquisition accuracy and system resource consumption in most scenarios. This can be achieved by using a sliding window (window size...). Feature aggregation, time complexity This ensures efficient data processing; the sliding window technique is a mature and commonly used method in time series data processing. After collecting the load data, it is necessary to perform data standardization. In the data standardization stage, Z-score transform is used. ,in and This is a training set statistic, a classic operation in data preprocessing. Numerous studies have demonstrated its effectiveness in eliminating the influence of data units, making different features comparable. Anomaly detection employs the Isolation Forest algorithm (contamination = 0.01), which boasts high accuracy and robustness in anomaly detection, effectively eliminating outliers and ensuring feature stability (coefficient of variation). This coefficient of variation standard has been verified in multiple experiments to ensure the quality of feature data.
[0050] Furthermore, during the feature acquisition process, referring to relevant side-channel attack research, there is a potential risk that attackers can infer the system's operational status by analyzing feature vectors. In terms of defense mechanisms, the eBPF probe runs in a kernel-mode isolated environment, which, from the perspective of operating system security architecture design, effectively prevents unauthorized external access; the collected data is transmitted after memory encryption (AES-256-XTS), a standard industry-compliant encryption method that prevents memory sniffing; the feature extraction algorithm is proven to have no information leakage vulnerabilities using formal verification tools (such as Coq), and formal verification is authoritative in ensuring algorithm security. Regarding security effectiveness, experiments and related simulation analyses have reduced the success rate of side-channel attacks to [missing information]. It meets the FIPS 140-2 side channel protection requirements.
[0051] In the above encryption key generation process, the payload data is noise-added in the following way: the CPU load, memory page swapping efficiency, disk I / O response time, TCP establishment interval, and the Poisson process strength of the key generation request are vectorized to obtain a multi-dimensional feature vector; the quantum random bits generated by the quantum random number generator (QRNG) are injected into the multi-dimensional feature vector to generate the noise-added multi-dimensional feature vector.
[0052] In the above encryption key generation process, the encryption key is dynamically generated based on the noisy payload data. This can be done in the following ways: perform a hash operation on the noisy multidimensional feature vector to generate key seed data; perform secondary noisy processing on the key seed data based on CPU load as physical noise; use a chaotic mapping algorithm to map the secondary noisy key seed data into binary sequence data; and input the binary sequence data into the key derivation function to generate the encryption key.
[0053] Specifically, noise is added to the extracted time-series features to prevent attackers from using feature vectors to infer system state and predict keys. Details are as follows. When determining the basic configuration, core financial data... General government data ,satisfy This setup references extensive research on the application of differential privacy under varying data sensitivity levels, allocating the privacy budget rationally based on the importance and sensitivity of the data. Dynamic adjustments are made based on system load. piecewise functions ,when Time triggers budget expansion Duration By recording the cumulative privacy loss through an exponential mechanism, this dynamic adjustment mechanism can flexibly balance privacy protection and data availability when dealing with changes in system load, and is widely used in research on the combination of cloud computing resource scheduling and privacy protection.
[0054] Sensitivity calculation is based on feature vectors Sensitivity (After normalization), this calculation method is a commonly used sensitivity measure in differential privacy noise injection research. Laplace noise generation. ,in After adding noise, it satisfies This aligns with the mathematical definition of differential privacy, ensuring strong data privacy protection. Monte Carlo simulations based on IBM DPLibrary verify the accuracy of an attacker reconstructing the original features after 1000 iterations. Information leakage amount This verification method is reliable and reproducible in differential privacy verification research and can effectively evaluate the privacy protection effect.
[0055] The noise generation algorithm may contain biases, leading to insufficient privacy protection. Attackers can launch differential attacks by combining multiple queries. In terms of defense mechanisms, the noise generator uses quantum random numbers (QRNG) as the entropy source. Tested by NIST SP 800-22, the randomness and unpredictability of the noise are guaranteed. A privacy budget monitor is introduced, limiting the number of queries per user (≤10 times per day) through smart contracts to prevent budget exhaustion attacks. Smart contracts offer advantages in automation and immutability in ensuring data security and privacy budget management. Regarding security effectiveness, differential attack simulation tests show that the probability of an attacker successfully distinguishing adjacent datasets is [not specified]. This meets the privacy protection requirements of the GDPR.
[0056] The noisy features are processed by chaotic mapping to generate a highly random binary sequence, providing an entropy source for the key. Specific details are as follows: Set initial values (Non-periodic point), control parameters (Lyapunov exponent is in a chaotic region) Based on multiple studies on the application of chaos theory in cryptography, this parameter setting enables the Logistic chaotic mapping to produce good chaotic properties, remaining in the chaotic region and possessing a suitable Lyapunov exponent, thus guaranteeing the randomness and unpredictability of the sequence. Iterative equations It satisfies the Devaney definition of chaos, and this iterative equation is a classic method for generating chaos in chaotic encryption research.
[0057] Add randomness enhancement process, seed generation ,in To generate a noisy feature vector, a hash function is used to generate a seed, ensuring the seed's security and unpredictability. Mapping initialization. The seed is transformed into an initial value suitable for the chaotic mapping. The sequence generation is completed after 2048 iterations. Quantized into a 256-bit binary sequence This process is a common operation in the research of chaotic sequence generation and application, and can effectively generate high-quality random sequences.
[0058] To ensure data reliability, the random sequences were subjected to the full NISTSP800-22 test suite, including frequency testing of key metrics. Longest run test Linear complexity test The NIST SP800-22 test is an internationally recognized authoritative standard for evaluating the quality of random sequences. Passing this test demonstrates that the generated sequences have good randomness and unpredictability.
[0059] Chaotic mappings may contain periodic vulnerabilities, allowing attackers to predict keys by analyzing sequence patterns; leakage of initial values can also lead to complete sequence reconstruction. To address this, the mapping parameters are cryptographically verified (proving the absence of short periods using the CryptoMiniSat solver) to ensure security. Seed generation introduces physical noise sources (such as CPU thermal noise), updating the initial value every 1000 keys generated to increase randomness and unpredictability. The iteration process incorporates random perturbations (based on quantum random numbers) to break potential periodicity and further enhance the randomness of the chaotic sequence. In terms of security effectiveness, the unpredictability of the key sequence reaches [a certain level]. Its resistance to sequence prediction attacks meets the NIST SP 800-203 requirements for resistance to quantum randomness.
[0060] A standardized algorithm is used to derive a dedicated key from the chaotic sequence, distinguishing between the data encryption key and the key encryption key to ensure storage security. HKDF (HMAC-based Key Derivation Function) is employed to derive the key from the chaotic sequence; the HKDF algorithm is a mature and secure method in key derivation research. Master Key The output is 512 bits, a setting that meets the key length requirements of high-security applications. Key splitting involves using a 256-bit Data Encryption Key (DEK) for symmetric encryption and a 256-bit Key Encryption Key (KEK) stored in the HSM (compliant with FIPS 140-2 Level 3). This splitting and storage of the keys, combined with a hardware security module, effectively ensures key security.
[0061] The key derivation process may lead to key leakage due to algorithm implementation flaws; a breach in the physical security of the HSM can result in KEK leakage. In terms of defense mechanisms, the HKDF implementation is verified by NIST CAVP, ensuring no key recovery vulnerabilities; the HSM employs a tamper-proof design, where physical attacks trigger key self-destruction (based on sensor detection), ensuring security at the hardware design level; the DEK uses a one-time pad mechanism, destroying the key immediately after each encryption to prevent security risks from key reuse. Regarding security effectiveness, the probability of key leakage is significantly reduced. HSM's resistance to physical attacks reaches Common Criteria EAL 5+ level.
[0062] To achieve trusted storage and end-to-end traceability of encrypted data, in addition to data encryption, the encryption results and feature data can also be managed on-chain. Therefore, in some embodiments, after encrypting the data to be encrypted, the encrypted ciphertext data is written to the blockchain, along with multiple data features. This leverages the immutability of the blockchain to ensure the integrity of the ciphertext data and the traceability of the feature data, providing a foundation for subsequent trusted data analysis.
[0063] For the management and distribution of dynamically generated encryption keys, a secure key-sharing mechanism must be considered when the encryption key is a symmetric key. Therefore, in some embodiments, after the encryption key is dynamically generated, the public key in the encryption key is broadcast to the consensus nodes in the blockchain. In this way, the distributed verification of the blockchain consensus nodes ensures the security and legitimacy of the public key transmission, providing a trusted foundation for the subsequent use of the key (such as ciphertext decryption and signature verification).
[0064] The key distribution protocol used by the dynamically generated encryption key in this application embodiment can be implemented in the following way.
[0065] Choose the elliptic curve secp256k1, whose order is... Generator The secp256k1 curve is widely used in the field of cryptography, and its security and performance have been verified in a large number of studies.
[0066] User key pair satisfy Public key hash Write the blockchain identity contract and associate attribute vector This blockchain-based identity contract registration method, combined with the immutability of blockchain, can effectively guarantee the authenticity and security of identity information.
[0067] During the certification process, the certifier is selected. ,calculate ,send ; Verifier receives Generate challenges Send c; the prover calculates the response. Send s; Verifier verification If it passes, it is accepted. This protocol process is a classic Schnorr protocol implementation in zero-knowledge proof research, satisfying the zero-knowledge property of special honest verifiers.
[0068] Introducing a 30-second validity period timestamp (nonce) prevents replay attacks. Authentication logs are uploaded to the blockchain via smart contracts to achieve non-repudiation. Timestamps and smart contracts are effective in preventing replay attacks and ensuring the non-repudiation of authentication records. Related research shows that this approach can effectively improve the security of authentication protocols.
[0069] The protocol may be vulnerable to man-in-the-middle attacks, where attackers could tamper with it. and Identity forgery; the predictable nonce generation mechanism leads to replay attacks. For defense, a two-way authentication model is adopted, where both parties exchange signed public key certificates to enhance authentication reliability; the nonce is generated based on quantum random numbers, achieving high unpredictability. This ensures the security of the nonce; the protocol uses an automated verification tool (ProVerif) to prove its resistance to man-in-the-middle attacks. The ProVerif tool provides a formal verification of the protocol's security. In terms of security effectiveness, the success rate of identity forgery is [not specified]. The replay attack defense rate reaches 100%.
[0070] Specifically, a quantum key distribution (QKD) subsystem is used to distribute quantum keys. The transmitting end employs a quantum random number generator (QRNG, entropy source rate ≥ 1 Gbps) and a LiNbO3 phase modulator (response time < 1 ns). The QRNG provides high-quality random numbers, and the LiNbO3 phase modulator meets the requirements for high-speed quantum state modulation. The transmission link uses G.655 single-mode fiber (attenuation ≤ 0.19 dB / km) and a series EDFA optical amplifier (noise figure < 5 dB). This combination of fiber and optical amplifier is a commonly used configuration in quantum communication link research, effectively reducing transmission loss and noise. The receiving end uses a superconducting single-photon detector (SPD, dark count rate < 100 Hz, detection efficiency > 80%), enabling efficient detection of single photons. During quantum state transmission, Alice transmits polarization states at a rate of 1 Mbit / s. Encoding basis vectors This is the classic quantum state encoding method in the BB84 protocol. Basis vector alignment involves exchanging basis vector information via an SSL / TLS 1.3 channel with a screening rate of approximately 45%. The SSL / TLS 1.3 channel ensures the security of basis vector information exchange. Error correction employs cascaded codes, reducing the QBER from 3% to <0.1%. Cascaded codes have high error correction efficiency in quantum key distribution research. Privacy amplification uses Toeplitz matrix hash compression to generate a 256-bit secure key. Toeplitz matrix hashing is a commonly used method in privacy amplification research. Every 10 4 A 10% sampling rate of qubits is used to check the QBER. If the QBER > 6%, a link self-check is triggered, and a backup key pool is activated. This eavesdropping detection mechanism is commonly used in quantum key distribution research and can effectively detect eavesdropping behavior in the link. Quantum channels are susceptible to intercept-retransmission attacks; classical channel basis vector information can be eavesdropped on, leading to key leakage; and device defects (such as modulator non-ideals) can introduce side-channel attacks. In terms of defense mechanisms, the quantum no-cloning theorem is utilized, as any eavesdropping behavior will significantly increase the QBER (detection rate 100%), which is the core guarantee of quantum communication security. Classical channels employ post-quantum encryption (CRYSTALS-Kyber) to protect basis vector information, which effectively resists attacks from quantum computers. The device compensates for non-ideals through a self-calibration algorithm to eliminate side-channel vulnerabilities. The self-calibration algorithm plays an important role in the performance optimization research of quantum communication devices. In terms of security performance, the key generation rate is 500 kbps@50 km, and the security against quantum attacks reaches an unconditional security level (information-theoretical security).
[0071] In this embodiment of the application, the AES-256-GCM mode is adopted. Encrypt the key Generate ciphertext With a 128-bit authentication tag, the AES-256-GCM mode is a secure and efficient method in data encryption and authentication research.
[0072] The physical layer transmits data via a QKD link. Quantum communication is used to ensure the security of key transmission; the application layer transmits data via TCP / IP protocol. And Schnorr authentication tokens, using congestion control algorithms to ensure transmission reliability, TCP / IP protocol and congestion control.
[0073] After data encryption, a data polynomial is constructed based on Kate commitments, supporting batch non-interactive verification (no data transmission required), improving efficiency by 75%. Anomalies are located using vector clocks to mark operation timing and the A* search algorithm, completing tampering tracing within 3 seconds (traditional log auditing takes hours). This enables integrity verification of the original data copy and tracing of data tampering nodes.
[0074] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this patent.
[0075] Another embodiment of this application relates to an encryption device, such as... Figure 3 As shown, it includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned data encryption method.
[0076] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0077] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0078] Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.
[0079] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0080] Those skilled in the art will understand that the above embodiments are specific implementations of this application, and in practical applications, various changes can be made in form and detail without departing from the spirit and scope of this application.
Claims
1. A data encryption method characterized by, Applied to an encryption device, the method comprises: obtaining terminal uploaded data to be encrypted, and extracting a plurality of data features from the data to be encrypted, wherein the data features include data sensitive features, business risk features and user behavior features; respectively quantifying the plurality of data features to obtain a plurality of feature scores, and calculating a data total score according to the plurality of feature scores; inputting the data total score into a pre-trained decision tree model to determine the security level of the data to be encrypted; determining a target encryption algorithm from a pre-set correspondence between security levels and encryption algorithms according to the security level; encrypting the data to be encrypted according to the target encryption algorithm and a dynamically generated encryption key.
2. The data encryption method of claim 1, wherein, The security level includes public level, ordinary confidential level, high confidential level, and top confidential level; The correspondence between the security level and the encryption algorithm comprises: If the security level is public level, the corresponding encryption algorithm is SM4 encryption algorithm; If the security level is ordinary confidential level, the corresponding encryption algorithm is adaptive AES encryption algorithm, wherein the encryption round number of the adaptive AES encryption algorithm is determined according to the load data of the encryption device; If the security level is high confidential level, the corresponding encryption algorithm is RSA-ECC hybrid encryption algorithm; If the security level is top confidential level, the corresponding encryption algorithm is CRYSTALS-Kyber quantum-resistant encryption algorithm.
3. The data encryption method as claimed in claim 1, wherein, The generation of the encryption key comprises: obtaining load data of the encryption device; performing noise addition processing on the load data, and then dynamically generating the encryption key according to the noise-added load data.
4. The data encryption method of claim 3, wherein, The load data includes CPU load, memory page exchange efficiency, disk I / O response time, TCP establishment interval, and Poisson process intensity of key generation request; The noise addition processing on the load data comprises: vectorizing the CPU load, the memory page exchange efficiency, the disk I / O response time, the TCP establishment interval, and the Poisson process intensity of the key generation request to obtain a multi-dimensional feature vector; injecting quantum random bits generated by a quantum random number generator into the multi-dimensional feature vector to obtain a noise-added multi-dimensional feature vector.
5. The data encryption method of claim 4, wherein, The dynamically generating the encryption key according to the noise-added load data comprises: performing hash operation on the noise-added multi-dimensional feature vector to generate key seed data; performing secondary noise addition processing on the key seed data based on the CPU load as physical noise; mapping the secondary noise-added key seed data to binary sequence data using a chaotic mapping algorithm; inputting the binary sequence data into a key derivation function to generate the encryption key.
6. The data encryption method of claim 1, wherein, The data sensitive feature comprises a level, a density and a field correlation degree factor of a sensitive field in the data to be encrypted, the business risk feature comprises a regulatory level, an influence coefficient and a leakage influence range of an industry to which the data to be encrypted belongs, and the user behavior feature comprises an access frequency, an operation type, a user permission and a behavior credibility of the user to the data to be encrypted. The quantification and scoring of the multiple data features respectively to obtain multiple feature scores, and the calculation of a data total score according to the multiple feature scores, comprises: The level and the density of the sensitive field, and the field correlation degree factor are quantified respectively to obtain a feature score of the data sensitive feature; The regulatory level, the influence coefficient and the leakage influence range are quantified respectively to obtain a feature score of the business risk feature; The access frequency, the operation type, the user permission and the behavior credibility are quantified respectively to obtain a feature score of the user behavior feature; The feature score of the data sensitive feature, the feature score of the business risk feature and the feature score of the user behavior feature are summed according to a preset weight to obtain the data total score.
7. The data encryption method of claim 1, wherein, The method further comprises: After the encryption processing of the data to be encrypted, the encrypted ciphertext data is written into a blockchain, and the multiple data features are written into the blockchain.
8. The data encryption method of claim 1, wherein, The encryption key is a symmetric key; The method further comprises: After the dynamic generation of the encryption key, a public key in the encryption key is broadcast to a consensus node in the blockchain.
9. An encryption device, characterized by Comprise: At least one processor; And A memory in communication connection with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data encryption method as claimed in any one of claims 1 to 8.
10. A computer readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to implement the data encryption method as claimed in any one of claims 1 to 8.
Citation Information
Cited By
Cross-network data security interaction method and system
CN121509120A
Cross-network data security interaction method and system
CN121509120B
Meteorological observation data hybrid encryption method and device based on polling key, and medium
CN121547282A
Data encryption and decryption method, system and equipment based on national cryptographic algorithm and storage medium
CN121711153A