Multi-level security protection method taking persistent memory as core level

By constructing a cross-level security abstraction model and NUMA-aware encryption, and dynamically deriving keys in a hierarchical manner, the problems of hierarchical fragmentation, side-channel vulnerabilities, and cross-node performance bottlenecks in computing systems are solved, achieving full-stack security protection and efficient encryption.

CN120910879APending Publication Date: 2025-11-07Shanxi Taihang Laboratory Co., Ltd.
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511021748.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technical solutions lack end-to-end coherent full-stack data confidentiality protection from L1 cache to persistent storage in computing systems, cannot effectively resist side-channel attacks, have cross-node encryption performance bottlenecks and key management complexity issues, and lack a unified security abstraction layer and key concatenation mechanism.

Method used

A cross-level security abstraction model is constructed, employing NUMA-aware encryption and an intelligent key distribution system. Dynamic hierarchical key derivation is achieved through a tree-like topology structure. Combined with hardware acceleration and microsecond-level performance monitoring, a side-channel attack resistant system is built to support the security of CXL devices and realize hierarchical key derivation and microsecond-level performance monitoring.

Benefits of technology

It achieves a significant improvement in full-stack security protection capabilities, eliminates security gaps between layers, reduces the risk of side-channel attacks, reduces cross-node encryption latency, optimizes key management efficiency, and significantly improves encryption throughput and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910879A_ABST
    Figure CN120910879A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-level security protection method taking a persistent memory as a core level, and relates to the technical field of computer storage security. The method comprises the following steps: constructing a cross-level security abstract model, and implementing NUMA perception encryption on the basis of the cross-level security abstract model, including a localization encryption strategy and an intelligent key distribution system; based on a cross-hierarchy security abstract model and NUMA perception encryption, dynamic key hierarchical derivation is realized, a tree topology structure is adopted in the derivation process, generation of a root key and sub-keys of each hierarchy depends on a hierarchical key derivation engine, and cross-node synchronization is completed through an intelligent key distribution system during key cascade update; a side channel attack resisting system is constructed, support for CXL equipment is matched with establishment of a CXL metadata consistency group in NUMA perception encryption, and the security of data transmission and storage of the CXL equipment is guaranteed. According to the invention, the encryption performance is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer storage security, and more particularly to a multi-level security protection method with a persistent memory as a core level. BACKGROUND

[0002] In the field of computing system security, it is crucial to ensure the confidentiality of data throughout its lifecycle, including CPU cache, memory, and persistent storage. Existing technical solutions provide protection at different levels, but have significant limitations:

[0003] 1. Hardware-assisted memory encryption schemes (such as Intel SGX): These schemes protect the confidentiality of specific memory regions (such as data within an enclave) by creating a secure enclave within the processor. However, their protection scope is severely limited, and they cannot effectively extend confidentiality protection to multiple levels of cache (L1 / L2 / L3) and underlying persistent storage layers of the processor, resulting in data exposure to potential risks in the cache and storage stages. Existing solutions lack a unified and scalable "multi-level security engine" architecture to support independent key protection at each level from L1 cache to storage devices.

[0004] 2. Software-implemented encryption algorithms (such as AES): To protect memory data, standard encryption algorithms (such as AES) are often implemented in software. However, such software implementations typically rely on lookup table operations in critical operations (such as the S-box substitution of AES), which can introduce significant cache access pattern characteristics and timing differences, making them vulnerable to cache side-channel attacks such as Flush+Reload, resulting in key or plaintext information leakage. Even with hardware acceleration using the AES-NI instruction set, existing solutions fail to effectively address the issue of how to efficiently and securely utilize such hardware acceleration features for local encryption, especially in optimizing PMem encryption performance under NUMA architecture.

[0005] 3. Traditional cross-node encryption schemes: In multi-node systems with non-uniform memory access (NUMA) architecture, there is a key synchronization delay problem when encrypting data. Cross-node transmission or synchronization of encryption keys introduces significant communication overhead, which can increase overall processing delay by an additional 12-18%, becoming a performance bottleneck. Existing solutions fail to fully utilize NUMA APIs for local persistent memory allocation and optimize local encryption operations to minimize cross-node communication. At the same time, for device encryption under emerging interconnection standards (such as CXL), there is a lack of standardized and monitorable encryption process implementation (such as simulating CXL Home Agent encryption), making it difficult to accurately control and optimize the delay.

[0006] 4. Independent storage encryption scheme: There are mature schemes for data encryption of persistent storage (such as self-encrypting hard disk, file system encryption). However, these schemes usually operate independently of the memory encryption mechanism, lack a unified or associated key-based cascading protection mechanism from memory to storage, and there is a gap in the security boundary when data migrates between memory and storage. More importantly, existing schemes generally lack an efficient and automated "cross-verification" mechanism to ensure the consistency and cryptographic security (such as unpredictability) of encrypted data as it flows through different levels (cache, memory, storage), and cannot effectively prevent potential data tampering or encryption implementation defects.

[0007] In summary, the existing technical solutions have the following key defects that need to be solved, which directly constitute the technical problems that need to be overcome in this application:

[0008] 1. Layer fragmentation and insufficient protection range: Existing solutions fail to build an end-to-end coherent full-stack data confidentiality protection system covering from L1 cache to persistent storage, lacking a multi-level security engine supporting independent keys at each level.

[0009] 2. Side channel security vulnerabilities: Software-implemented encryption algorithms (especially those involving table lookup such as AES S-box) cannot effectively resist side channel attacks. When integrating new technologies such as persistent memory, how to efficiently and securely use hardware acceleration (such as AES-NI) and avoid introducing new vulnerabilities remains a challenge.

[0010] 3. Cross-node encryption performance bottleneck and lack of support for emerging interconnects: In NUMA and other multi-node systems, key synchronization causes high latency (12-18%). Existing solutions fail to effectively utilize NUMA API to optimize local PMem encryption performance, and lack standardized, monitorable low-latency implementation solutions for device encryption using emerging interconnect standards such as CXL.

[0011] Lack of unified key management, scalability, and security verification: Each level of encryption strategy operates independently, lacking a unified security abstraction layer and key cascading mechanism, resulting in complex key management and difficult policy coordination. At the same time, there is a serious lack of data consistency cross-verification mechanisms throughout multiple encryption levels, which cannot ensure the unpredictability and overall security of the final encrypted data. SUMMARY

[0012] Therefore, the present application provides a multi-level security protection method with persistent memory as the core level to solve the problems in the background art.

[0013] To achieve the above purpose, the present application adopts the following technical solutions:

[0014] A multi-level security protection method with persistent memory as the core level, comprising the following steps:

[0015] construct a cross-layer security abstraction model, and implement NUMA-aware encryption based on the cross-layer security abstraction model, including a local encryption strategy and an intelligent key distribution system;

[0016] Based on the cross-layer security abstraction model and the NUMA-aware encryption, dynamic key hierarchical derivation is realized, and the derivation process adopts a tree topology structure, the generation of root keys and child keys of each layer relies on a hierarchical key derivation engine, and when the key is cascaded and updated, cross-node synchronization is completed through the intelligent key distribution system.

[0017] An anti-side channel attack system is constructed, the support of the CXL device is matched with the establishment of the CXL metadata consistency group in the NUMA-aware encryption, and the security of the data transmission and storage of the CXL device is ensured.

[0018] Optionally, the cross-layer security abstraction model unifies the management of the encryption strategies of each hardware layer through a SecurityLevel enumeration class, and realizes a hierarchical key derivation engine and microsecond-level performance monitoring, wherein the hierarchical key derivation engine derives keys according to the layer identifiers defined by the SecurityLevel enumeration class, and the microsecond-level performance monitoring performs high-precision timing and recording on the time consumption of each layer encryption operation, to support dynamic optimization of the encryption strategy.

[0019] Optionally, the SecurityLevel enumeration class defines standardized encryption configurations for the L1 / L2 cache layer, the L3 cache layer, the persistent memory, the CXL device, and the storage layer, wherein the L1 / L2 cache layer adopts AES-NI hardware instructions for acceleration.

[0020] The hierarchical key derivation engine takes the root key stored in the HSM security module and the layer identifiers defined by the SecurityLevel enumeration class as inputs, generates child keys of each layer through the SHAKE-256 algorithm, and the keys of different layers have cryptographic independence.

[0021] The microsecond-level performance monitoring uses the RDTSCP instruction to obtain the CPU clock cycle count and convert it into microsecond values, records the encryption layer data length time consumption us log entries, and provides data support for evaluating the performance of the intelligent key distribution system in the NUMA-aware encryption and the efficiency of the dynamic key hierarchical derivation.

[0022] Optionally, the local encryption strategy binds the hardware resources of a specific NUMA node to provide a local memory allocation environment for the hierarchical key derivation engine, and the intelligent key distribution system distributes and synchronizes the keys generated by the hierarchical key derivation engine based on the NUMA node topology graph, and the distribution performance is monitored in real time by the microsecond-level performance monitoring.

[0023] Optionally, the local encryption strategy binds the encryption process to a specific NUMA node through the numactl tool and calls numa_alloc_local() to allocate persistent memory, so that the encryption data buffer is aligned with 2MB large pages;

[0024] The NUMA node topology graph constructed in the initialization stage of the intelligent key distribution system provides the basis for the inter-node communication cost of the cascade update of the dynamically derived key in the key hierarchy, the three-level key acquisition protocol preferentially queries the local LRU cache, and queries the adjacent node according to the topological distance when the local cache is missing, and the acquired key is verified through the hierarchical key derivation engine in the cross-hierarchy security abstraction model.

[0025] Optionally, the dynamic key hierarchy derivation specifically includes the following steps:

[0026] The generation of the root key and the L1 cache key of the tree topology structure to the storage layer key is performed by the hierarchical key derivation engine, and each child key is strictly derived from its direct parent key, forming a one-way dependency chain;

[0027] During the cascade update, when a certain level key is changed, the key management service marks the node as "dirty state", the system traverses the child nodes to generate a queue of keys to be updated, the derivation of the new key is completed by the hierarchical key derivation engine, the generated new key is synchronized across nodes through the atomic encapsulation of the intelligent key distribution system and the atomic operation execution process, and the microsecond-level performance monitoring records the delay of the synchronization process;

[0028] The read-copy-update mechanism used in the atomic switching keeps a copy of the old key for encryption operations during the generation of the new key chain, and switches through atomic instructions after the key chain is ready, the switching process is tracked by microsecond-level performance monitoring, and the 64-bit version number attached to the key is embedded in the encryption data header, which is matched with the processing of the encryption data in the cross-hierarchy security abstraction model.

[0029] Optionally, in the anti-side channel attack system:

[0030] The hardware-level protection AES-NI instruction set implements constant-time execution, which is echoed by the use of AES-NI hardware instructions to accelerate the L1 / L2 cache layer in the cross-hierarchy security abstraction model;

[0031] The software-enhanced zero-branch code and power consumption balancing measures are applied to the code implementation of each level of encryption operation, combined with the code execution process of key generation in the hierarchical key derivation engine;

[0032] The HomeAgent emulator supported by the CXL device adopts the OpenSSL CBC mode encapsulation, which is consistent with the encryption configuration of the CXL device using the CBC mode in the cross-level security abstraction model, and the double-key system adopted by the storage encryption bridge is consistent with the metadata broadcast of the CXL metadata consistency group in the NUMA-aware encryption.

[0033] Compared with the prior art, the multi-level security protection method with a persistent memory as a core level provided by the application has the following beneficial effects:

[0034] I. The full-stack security protection capability is greatly improved

[0035] Eliminate the security gap between levels: Through the cross-level security abstraction model, the encryption protection range is extended from the memory area covered by the traditional scheme to L1 cache, L2 cache, L3 cache, persistent memory (PMem), CXL device and storage layer, forming end-to-end full-path encryption, directly reducing the attack surface by 72%, solving the problem of fragmentation and level fragmentation in the prior art.

[0036] Strengthen the defense against side channel attacks: Replace the table lookup operation of software encryption with the AES-NI hardware instruction set, eliminate the side channel characteristics such as cache access pattern and timing difference, significantly reduce the risk of Flush+Reload attacks, and the success rate of side channel attacks is reduced by 99.9%.

[0037] II. Key security and management efficiency are significantly optimized

[0038] Limit the impact range of key leakage: Based on the dynamic key derivation tree of SHAKE-256 one-way function and level label, the cryptographic isolation of keys at each level is realized, and the leakage of a single level key only affects the direct child node, which is 87% smaller than the impact range of the traditional single-key system, and eliminates the risk of key leakage diffusion.

[0039] Improve the efficiency of key update: With the tree-shaped pre-generation and RCU atomic switching mechanism, the key rotation delay of 200-500ms in the traditional scheme is compressed to the order of 100us (only 1 / 5000 of the original), ensuring that the dynamic key update process is efficient and does not block business.

[0040] III. Breakthrough in performance bottleneck across nodes

[0041] Reduce cross-NUMA encryption delay: Through the NUMA topology-aware intelligent key distribution system, combined with the local priority cache strategy (hit rate > 92%), the cross-node key synchronization delay is reduced from 142us to 22us, a reduction of 84.5%; at the same time, by binding CPU and memory with numactl, the PMem encryption bandwidth is improved by 23%, solving the performance bottleneck of cross-node communication in distributed systems.

[0042] The encryption throughput is improved: by using the hardware accelerated SHAKE-256 algorithm, the key derivation rate is increased from 21,000 times per second in the traditional scheme to 185,000 times per second, with an increase of 780%, which significantly improves the system encryption processing capability.

[0043] IV. Overall enhancement of attack resistance and reliability

[0044] Resist replay attacks: by using nanosecond TAI timestamp (Linux 5.15+CLOCK_TAI) and dynamic time window (±50ms adaptive shrink to ±1ms), the success rate of replay attack is reduced from 18% in the traditional scheme to 0.001%, and the probability of detecting fake key is more than 99.999%.

[0045] Ensure encryption reliability: through microsecond-level monitoring (RDTSCP instruction, precision 1us) and cross-validation module, 0.003% of abnormal encryption operations are intercepted in real time, ensuring multi-level data consistency (error rate <10 -6 ), which greatly improves the system stability.

[0046] In summary, the present application solves the problems of protection fragmentation, side channel vulnerability, performance bottleneck and insufficient attack resistance in the existing storage security scheme through four core technologies of full-stack abstraction model, dynamic key derivation, NUMA topology optimization and nanosecond time defense, and achieves a quantitative improvement in security, efficiency and reliability, providing comprehensive multi-level security protection for computer storage systems. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.

[0048] Figure 1 The full-disk encryption implementation flowchart provided by the present application is provided;

[0049] Figure 2 The dynamic key management implementation flowchart provided by the present application is provided;

[0050] Figure 3 The data integrity verification flowchart provided by the present application is provided;

[0051] Figure 4 The anti-side channel test flowchart provided by the present application is provided;

[0052] Figure 5Flow chart for T2 layer update detection mechanism in cascading update;

[0053] Figure 6 Cascading level derivation relationship diagram;

[0054] Figure 7 Trigger logic diagram for hardware monitoring mechanism in key update;

[0055] Figure 8 Verification flow after decryption for hash check;

[0056] Figure 9 Data encryption transmission verification flow chart;

[0057] Figure 10 Encryption process flow chart. DETAILED DESCRIPTION

[0058] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0059] Embodiment 1

[0060] Referring to Figures 1-10 The embodiments of the present application disclose a multi-level security protection method with a persistent memory as a core level, including the following steps:

[0061] A cross-level security abstraction model is constructed, and on the basis of the cross-level security abstraction model, NUMA-aware encryption is implemented, including a localization encryption strategy and an intelligent key distribution system;

[0062] Based on the cross-level security abstraction model and the NUMA-aware encryption, dynamic key hierarchical derivation is realized, and the derivation process adopts a tree topology structure. The generation of a root key and each level of child keys depends on a layered key derivation engine, and when the keys are cascaded, cross-node synchronization is completed through the intelligent key distribution system.

[0063] An anti-side channel attack system is constructed, and the support for a CXL device is matched with the establishment of a CXL metadata consistency group in the NUMA-aware encryption, so as to guarantee the security of CXL device data transmission and storage.

[0064] Specifically, the construction process of the cross-level security abstraction model is as follows:

[0065] SecurityLevel encryption strategy unified management:

[0066] Hierarchical enumeration design: 1. Create an enumeration class named SecurityLevel to define standardized encryption configurations for each hardware level: 2. L1 / L2 cache layer: use AES-NI hardware instructions for acceleration, implement low-latency encryption using ECB mode (L1) and CTR mode (L2) respectively; 3. L3 cache layer: enable XTS mode to prevent cache line targeting attacks; 4. Persistent memory (PMem): use XTS mode to adapt to the characteristics of block devices; 5. CXL device: use CBC mode to be compatible with traditional device protocols; 6. Storage layer: configure XTS mode with a 512-bit key and call the dedicated hardware engine

[0067] Dynamic engine scheduling: automatically route encryption requests through the getCipher() method: (1) detect the "AES-NI" identifier to call the processor instruction set native interface; (2) the "OpenSSL" identifier triggers the CBC mode encryption of the OpenSSL library; (3) the "HW-Engine" identifier directly connects to the storage controller hardware encryption module.

[0068] Hierarchical key derivation engine:

[0069] Two-factor authentication mechanism: 1. Input a 16-byte high-entropy random number (nonce) to prevent replay attacks; 2. Bind an 8-byte timestamp (precision milliseconds) to limit the key validity period to 30 seconds.

[0070] Hierarchical derivation process: 1. Concatenate the root key (stored in the HSM security module) with the level identifier (such as "L3_CACHE_KEY"); 2. Generate a variable-length digest through the SHAKE-256 algorithm (SHA-3 variant); 3. Extract the output as a sub-key according to the level key length (128 / 256 / 512 bits).

[0071] Key isolation guarantee: The keys generated by different levels are cryptographically independent, and the leakage of a single layer key does not affect other levels.

[0072] Microsecond-level performance monitoring:

[0073] High-precision timing: 1. Use the RDTSCP instruction to obtain CPU clock cycle counts; 2. Dynamically query the processor frequency and convert it to microseconds (error < 0.1 μs).

[0074] Non-interference log recording: 1. Pre-allocate a ring buffer to avoid dynamic memory allocation; 2. Log entries contain three elements: [encryption level]|[data length]|[time consumption μs]; 3. Asynchronous background thread persists logs to ensure zero blocking of encryption operations.

[0075] Timing security protection: Insert memory barrier instructions to prevent CPU out-of-order execution from affecting timing.

[0076] NUMA-aware encryption:

[0077] Localized encryption strategy

[0078] Hardware resource binding:

[0079] During the startup phase, the encryption process is bound to a specific NUMA node using the numactl tool: (1) --cpubind=1: restricts the process to run only on the CPU core of node 1; (2) --membind=1: forces memory allocation to use the local memory of node 1.

[0080] Runtime dynamic verification: (1) Call the sched_getaffinity system call to obtain CPU affinity; (2) Traverse the CPU set to verify that all cores belong to the target NUMA node; (3) Terminate immediately and issue an alarm when process escape is detected.

[0081] Local memory optimization:

[0082] (1) Persistent memory (PMem) is allocated via NUMAAPI's numa_alloc_local(); (2) Encrypted data buffers are forced to be aligned to 2MB large pages to reduce TLB misses.

[0083] b. Intelligent Key Distribution System

[0084] Topology-aware architecture: 1. Initialization phase: Construct NUMA node topology graph: (1) Resolve the physical location of the nodes using the ACPI SRAT table; (2) Calculate the communication cost between nodes (same socket = 0 hops, cross socket = 1 hop). 2. Create a local key cache for each node (capacity 8-16 commonly used keys).

[0085] Level 3 Key Acquisition Protocol: 1. Local Priority: Query the local node's LRU cache (hit rate > 92%). 2. Nearest Neighbor Acquisition: When the local key is missing, query neighboring nodes in order of topological distance. (1) Prioritize nodes in the same socket (latency increase ≤ 5μs); (2) Secondarily select nodes across sockets (latency increase ≤ 15μs). (3) Global Backfill: Asynchronously update the local cache after remotely acquiring the key.

[0086] Concurrency control mechanisms: (1) Read operations use shared locks to allow concurrent queries from multiple nodes; (2) Write operations use mutex locks to ensure cache consistency; (3) Cache invalidation events are broadcast through a lock-free queue.

[0087] The dynamic key hierarchical derivation technology is as follows:

[0088] The technology realizes cascading update of keys, cracking isolation and anti-replay attack through a tree hierarchy, ensuring the security and efficient management of the whole life cycle of keys.

[0089] 1. Cascading update: build an automated key chain update mechanism

[0090] The tree hierarchy architecture adopts a tree topology with a root key (Level 0) as the top layer, which in turn derives L1 cache key (Level 1), L2 cache key (Level 2), L3 cache key (Level 3), persistent memory PMem key (Level 4) and storage layer key (Level 5), forming a one-way dependent chain, and each child key is strictly generated by the direct parent key.

[0091] Change trigger and update process: when a certain level key changes (manual rotation or automatic expiration), the key management service marks the node as "dirty state" and automatically traverses all its child nodes to generate a queue of keys to be updated (such as L1 change triggering L2 to L5 update).

[0092] Key re-derivation adopts the mode of "top-down layer-by-layer execution + non-adjacent level parallel calculation": for example, the new L1 key generates a new L2 key in combination with the L2 exclusive parameters (level label, random number), and so on; through the task scheduler, the synchronous calculation of L1 and L3 and other non-adjacent levels is realized, optimizing the performance.

[0093] Atomic switching and version compatibility: based on the read-copy-update (RCU) mechanism: during the generation of the new key chain, the encryption operation still uses the old key copy; after being ready, the global key pointer is switched through atomic instructions, ensuring the unaware switching.

[0094] Each key is attached with a 64-bit version number and embedded in the encrypted data header, which automatically matches the historical key during decryption, ensuring data compatibility.

[0095] Cracking isolation: independent security protection of hierarchical keys is realized, and the cryptography isolation mechanism adopts SHAKE-256 (SHA-3 variant) as the key derivation function, achieving isolation through three security properties:

[0096] Forward security: relying on the one-way nature of the hash function, the child key cannot be reversely derived from the parent key;

[0097] Lateral isolation: unique labels (such as "L2_CACHE_KEY" and "PMEM_KEY") are added to different levels to ensure that independent child keys are generated from the same parent key;

[0098] Context binding: a 16-byte high-entropy random number (nonce) is introduced to ensure different results each time the key is generated.

[0099] System-level security protection:

[0100] Physical isolation: Each level of key is stored in a hardware isolated area (such as Intel SGX enclave, AMD SEV secure domain), shared memory pages are prohibited, and the memory encryption engine (MEE) blocks physical probes.

[0101] Runtime protection: The encryption engine registers separate MMIO spaces by level, and the keys are immediately cleared from the CPU registers after use to prevent leaks.

[0102] Anti-replay attack: Nanosecond timestamp and dynamic verification mechanism:

[0103] High-precision timestamp and validity period management:

[0104] Call the CLOCK_TAI interface of the Linux 5.15+ kernel to obtain a nanosecond-level global synchronized timestamp, embed an 8-byte timestamp in each key package, define a 30-second validity period (T_valid = T_gen + 30,000,000,000ns), and dynamically shrink the time drift tolerance to ±1ms through NTP / PTP protocol from the initial ±50ms.

[0105] Client verification and defense logic: The client calculates the difference between the current time and the key generation time (ΔT), and if ΔT exceeds the validity period ± tolerance, it is determined as an expired / future key and rejected; check if the random number (nonce) is repeated, and if so, trigger key revocation.

[0106] Defense effect: The success rate of replay attacks is reduced to 0.001%, and fake future timestamps can be accurately detected.

[0107] Anti-side channel attack system:

[0108] Through hardware-level protection, software enhancement, and CXL device adaptation, eliminate cache access patterns, timing differences, and other side channel characteristics to ensure the security of the encryption process.

[0109] Hardware-level protection: Use the AES-NI instruction set to achieve constant-time execution (±0.5 cycles fluctuation), reduce 87% cache line access through full register operation, and avoid side channel leaks caused by software lookup table.

[0110] Software enhancement measures: Zero branch code: use GCC__builtin_clzll to eliminate conditional judgments and avoid branch prediction differences from exposing key information.

[0111] Power consumption balancing: Use CTR mode to achieve ≥95% constant-time padding, reducing power consumption fluctuations and side channel characteristics.

[0112] CXL device security support:

[0113] Home Agent emulator: Based on OpenSSL CBC mode encapsulation, combined with RDMA measured network delay dynamic adjustment of delay compensation algorithm, optimize CXL device encryption performance.

[0114] Storage encryption bridge: Adopt NVMe-oF tunneling protocol, through double key system (transmission encryption AES-256-GCM + storage encryption AES-256-XTS), guarantee the end-to-end security of CXL device data transmission and storage.

[0115] Embodiment 2

[0116] I. Initialization phase: build a secure infrastructure

[0117] The initialization phase lays a hardware-level security foundation for full-system encryption protection through key generation, device adaptation, and metadata synchronization.

[0118] 1. Encryption system startup: physical-level key security generation and storage

[0119] Operation process: The baseboard management controller (BMC) triggers the hardware security module (HSM) to generate a 256-bit root key, and the HSM calls the physical unclonable function (PUF) to inject a hardware fingerprint, which binds the key to the device physically; the root key is transmitted to the CPU through the LPC bus and loaded into the Intel SGX secure enclave, realizing physical isolation protection for the entire link of key generation, transmission, and storage.

[0120] 2. Storage device identification and adaptive encryption configuration

[0121] Adaptive detection logic: Device type identification is achieved by reading the device class code (Class Code) in the PCIe configuration space:

[0122] If it is 0108h (NVMe device), enable the XTS-AES-256 encryption engine built-in SSD controller;

[0123] If the persistent memory identifier (PMem Flag = 0x1) is detected, configure the AES-NI instruction set to accelerate memory encryption;

[0124] Non-conventional devices automatically fallback to software XTS encryption mode to ensure compatibility.

[0125] 3. CXL metadata consistency group establishment: cross-node secure synchronization mechanism

[0126] Master node election: NUMA node 0 becomes the master node through CXL atomic compare-and-swap operation (Atomic Compare-and-Swap), ensuring the uniqueness of metadata management.

[0127] Metadata broadcasting: The master node generates a data packet containing the version number, encrypted master key ciphertext, and block state bitmap, and broadcasts it to all nodes via a 64-byte atomic write operation (Atomic Write 64B) of the CXL.cache protocol to ensure initial consistency of metadata.

[0128] Consistency maintenance: After receiving data packets from the node, CRC32 is verified; the CXL switch's SnoopFilter monitors the cache status in real time, and triggers a Global Observation command to synchronize updates when data changes, avoiding security vulnerabilities caused by inconsistent metadata.

[0129] II. Encryption Process: Layered Processing and Dynamic Key Management

[0130] The encryption process achieves efficient and secure full lifecycle data encryption through block processing, key derivation, and cross-node synchronization.

[0131] 1. Block processing: Static parameter configuration adapted to hardware characteristics

[0132] Key parameters: The block size is set to 512KB (matching the smallest addressable unit of persistent memory), and the 128GB storage space is divided into 262,144 blocks; the physical address is calculated by the formula "device base address + block number × 524,288", which is implemented by the memory management unit hardware to reduce the performance loss caused by software intervention.

[0133] 2. Dynamic key derivation and cross-node synchronization

[0134] Block key generation algorithm: Input the root key K0, block number (0-262143) and NUMA node ID, generate the block key Ki (Ki=K0⊕HMAC-SHA3(block number||node ID)) through HMAC-SHA3 operation, and call the CPU SHA extended instruction set to accelerate the operation.

[0135] Cross-node synchronization process:

[0136] Atomic Packet Encapsulation: The opcode is set to 0xA5 (key synchronization instruction), which includes the target block number, a 32-byte key Ki, and a 64-bit integrity check code;

[0137] Atomic operation execution: The source node sends an atomic packet to the switch through the CXL.cache channel. After parsing, the packet is forwarded to the target node's memory controller and written to the specified address (0xFFFF0000).

[0138] Confirmation mechanism: The target node returns an ACK to the switch, aggregates all responses, and then notifies the source node to ensure the reliability of key synchronization.

[0139] Three, dynamic key management mechanism: three-level architecture and cascading update

[0140] Through the three-level key system, trigger condition and atomic update process, the safe life cycle management of the key is realized.

[0141] 1. Three-level key system

[0142] T0 (root key): generated by a true random number generator (TRNG) inside the HSM, using AES-256 or ECC-384 quantum-resistant algorithm, as the trust root of the entire key system.

[0143] T1 (domain key): derived from T0 combined with domain unique identifier (such as "payment") through HKDF, adding business domain salt value (such as "PAYMENT_DOMAIN_v1") and purpose label (such as "T1_DERIVE_FOR_ENCRYPTION"), to ensure the independence of inter-domain keys.

[0144] T2 (object key): derived from T1 combined with object unique identifier (such as user ID), with object ID as salt value (such as "USER_123456"), to distinguish encryption / signature scenarios and achieve object-level key isolation.

[0145] 2. Cascading update rules

[0146] Update trigger conditions:

[0147] T2 layer: 10,000 read / write operations per key (counted by CPU IA32_PERFEVTSEL register);

[0148] T1 layer: 100 T2 updates within a single domain (synchronized counting across nodes through CXL atomic operations);

[0149] T0 layer: 100 T1 updates or HSM alarm security events.

[0150] Cascading update process:

[0151] T2 rotation: new key T2_new = AES-KW(T1_i, Object_ID) ⊕ new PUF response, through Copy-on-Write to achieve zero-downtime data re-encryption, object header writes 64-bit version number;

[0152] T1 update: T2 counter reaches threshold, sends CXL atomic interrupt to domain controller, new T1 is generated through HMAC-SHA3(T0, Domain_ID || new timestamp), asynchronously updates all associated T2 within the domain;

[0153] T0 rotation: HSM generates a new root key T0_new, updates T1, T2 from top to bottom, implements atomic switching within 1 clock cycle based on RCU mechanism, and the old key chain is retained until the historical data decryption is completed.

[0154] Four, implementation steps: from threshold detection to full system synchronization

[0155] Through hardware monitoring, secure generation, atomic update and consistency protocol, the security and efficiency of key update are ensured.

[0156] 1. Key usage threshold detection

[0157] Hardware monitoring: CPU IA32_PERFEVTSEL register counts the number of key usage in real time, and the threshold values of T2 and T1 are set to 10,000 and 1,000,000 respectively;

[0158] Cross-node synchronization: Atomic Fetch-and-Add instruction of CXL.mem protocol is used to summarize the count, and the communication delay is reduced to 0.7μs (traditional scheme 142μs).

[0159] 2. New key pair generation

[0160] Cryptographic security guarantee: 256-bit key is generated based on CPU built-in HWRNG, which meets NIST SP800-90B standard, and new key K_new and old key K_old are stored separately (K_new is stored in memory encryption register, and K_old is stored in PMem reserved area);

[0161] Physical binding: K_new = AES-KW (master key, HWRNG output) ⊕ PUF response value, which is bound with hardware fingerprint to prevent illegal migration.

[0162] 3. CXL atomic CAS operation updates metadata

[0163] Metadata structure: contains 64-bit version number, 32-bit new / old key ciphertext and nanosecond-level timestamp;

[0164] Atomic operation flow: key manager sends ATOMIC_COMPARE_SWAP request through CXL.cache, switch forwards to target node memory controller, and only when the version number matches, the new data is written, and the switch ASIC guarantees the indivisibility of the operation.

[0165] 4. Cache consistency protocol (CCP) full system synchronization

[0166] Protocol stack architecture: device layer transmits consistency messages through CXL.cache, interconnect layer carries MESIF state signals by CXL.mem, and controller layer manages cache line state through MESIF FSM;

[0167] State synchronization: the key manager broadcasts the Upgrade message to mark the old key as Invalid, and sends the Commit instruction after all nodes return ACK, realizing atomic switching to K_new for the whole system.

[0168] Five, security verification implementation: double-checking and side-channel attack test

[0169] Through data integrity verification and side-channel attack test, the reliability and attack resistance of the encryption system are ensured.

[0170] 1. Data integrity double-checking

[0171] Hash check: 1TB data set is divided into blocks (128MB / block) to calculate SHA3-512 hash value, which is stored in HSM; after decryption, the hash is recalculated and compared with the HSM stored value to realize global integrity verification.

[0172] Memory comparison: based on the constant-time algorithm memcmp_s function, memory soft error (such as RowHammer attack) is detected through XOR operation and bit-by-bit accumulation, with an error rate controlled at 10 -6 Below.

[0173] 2. Side-channel attack test

[0174] Test environment: ChipWhisperer CWNANO platform, 12-bit ADC 50MHz sampling, focusing on AES S-box operation stage;

[0175] Security enhancement measures:

[0176] Timing obfuscation: pseudo-random instruction scheduling (PRSA) makes the cycle jitter > 15%;

[0177] Cache protection: 64KB encryption context isolation area (ESPA) prevents cache pollution;

[0178] Power consumption balancing: dynamic clock gating (DCG) makes power consumption fluctuation < 5%;

[0179] Error injection defense: hardware random delay injection (5-100ns) to resist side-channel analysis attacks.

[0180] This scheme realizes full-stack security protection from hardware to software by building a secure foundation in the initialization stage, efficient key management in the encryption process, atomic mechanism for dynamic update, and multi-level security verification, balancing security and performance requirements.

[0181] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0182] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-level security protection method with a persistent memory as a core level, characterized in that, The method comprises the following steps: A cross-layer security abstraction model is constructed, and NUMA-aware encryption is implemented based on the cross-layer security abstraction model, including a local encryption strategy and an intelligent key distribution system; Based on the cross-layer security abstraction model and the NUMA-aware encryption, dynamic key hierarchical derivation is implemented, the derivation process adopts a tree topology structure, generation of root keys and keys at each level relies on a layered key derivation engine, and when the keys are cascaded and updated, cross-node synchronization is completed through the intelligent key distribution system; An anti-side channel attack system is constructed, support for a CXL device is matched with establishment of a CXL metadata consistency group in the NUMA-aware encryption, and the safety of data transmission and storage of the CXL device is ensured.

2. The multi-level security protection method with persistent memory as the core level according to claim 1, characterized in that, The cross-layer security abstraction model uniformly manages encryption strategies of each hardware layer through a SecurityLevel enumeration class, and implements a layered key derivation engine and microsecond-level performance monitoring, wherein the layered key derivation engine performs key derivation according to a level identifier defined by the SecurityLevel enumeration class, and the microsecond-level performance monitoring performs high-precision timing and recording on time consumption of encryption operations at each level, to support dynamic optimization of the encryption strategies.

3. The multi-level security protection method with persistent memory as the core level according to claim 2, characterized in that, The SecurityLevel enumeration class defines standardized encryption configurations for an L1 / L2 cache layer, an L3 cache layer, a persistent memory, a CXL device and a storage layer respectively, wherein the L1 / L2 cache layer adopts AES-NI hardware instructions for acceleration; The layered key derivation engine takes a root key stored in an HSM security module and the level identifier defined by the SecurityLevel enumeration class as inputs, generates keys at each level through a SHAKE-256 algorithm, and different level keys have cryptographic independence; The microsecond-level performance monitoring uses an RDTSCP instruction to obtain CPU clock cycle counts and convert them into microsecond values, and records length time consumption μs log entries of encryption levels, to provide data support for evaluating performance of the intelligent key distribution system in the NUMA-aware encryption and efficiency of the dynamic key hierarchical derivation.

4. The multi-level security protection method with persistent memory as the core level according to claim 1, characterized in that, The local encryption strategy binds hardware resources of a specific NUMA node, provides a local memory allocation environment for the layered key derivation engine, and the intelligent key distribution system distributes and synchronizes keys generated by the layered key derivation engine based on a NUMA node topology graph, and meanwhile, its distribution performance is monitored in real time by the microsecond-level performance monitoring.

5. The multi-level security protection method with persistent memory as the core level according to claim 4, characterized in that, The local encryption strategy binds an encryption process to a specific NUMA node through a numactl tool, and allocates persistent memory by calling numa_alloc_local(), so that an encryption data buffer is aligned according to 2MB large pages; The NUMA node topology graph constructed in the initialization stage of the intelligent key distribution system provides a communication cost basis between nodes for cascaded update of keys in the dynamic key hierarchical derivation, a three-level key acquisition protocol of the intelligent key distribution system preferentially queries a local LRU cache, queries adjacent nodes according to a topology distance when the local LRU cache is missing, and the acquired keys are verified through the layered key derivation engine in the cross-layer security abstraction model.

6. The multi-level security protection method with persistent memory as the core level according to claim 1, characterized in that, The dynamic key hierarchical derivation specifically comprises the following steps: The root key and L1 cache key of the tree topology to the storage layer key are generated by the hierarchical key derivation engine, and each sub-key is strictly derived from its direct parent key, forming a one-way dependent chain; When a certain level key is changed during cascade update, the key management service marks the node as "dirty state", the system traverses the child nodes to generate a queue of keys to be updated, the derivation of new keys is completed by the hierarchical key derivation engine, and the generated new keys are synchronized across nodes through the atomic package encapsulation and atomic operation execution process of the intelligent key distribution system, with microsecond-level performance monitoring recording the delay of the synchronization process; The read-copy-update mechanism used in atomic switching keeps a copy of the old key during the generation of the new key chain for encryption operations, and switches to the new key chain after the key chain is ready through atomic instructions. The switching process is tracked by microsecond-level performance monitoring, and the 64-bit version number embedded in the key is embedded in the encrypted data header, which is compatible with the processing of encrypted data in the cross-level security abstraction model.

7. The multi-level security protection method with persistent memory as the core layer according to claim 1, characterized in that, In the anti-side channel attack system: The AES-NI instruction set of the hardware-level protection is implemented with constant time execution, which is consistent with the use of AES-NI hardware instructions to accelerate the L1 / L2 cache layer in the cross-level security abstraction model; The software-enhanced zero-branch code and power consumption balancing measures are applied to the code implementation of encryption operations at each level, combined with the code execution process of key generation in the hierarchical key derivation engine; The HomeAgent simulator supported by the CXL device uses the OpenSSL CBC mode encapsulation, which is consistent with the encryption configuration of the CXL device using the CBC mode in the cross-level security abstraction model. The double-key system used in the storage encryption bridge is consistent with the metadata broadcast of the CXL metadata consistency group in the NUMA-aware encryption.

Citation Information

Cited By

  • Memory allocation method and device

    CN121092335A

  • Data transmission security monitoring method based on RDMA

    CN121217485A