A data operation request processing method and system
Patent Information
- Application Number
- CN202610654661.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-05-13
AI Technical Summary
然而,这种堆叠模式的安全控制灵活性不足,加密目标驱动作为通用块加密驱动,无法感知缓存目标驱动的调度策略与数据冷热状态,只能执行无差别加密,难以根据数据类型或地址范围动态调整加密规则,导致安全与性能无法灵活平衡
Smart Images

Figure CN122219853B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer data storage and security processing technology, and in particular to a data operation request processing method and system. Background Technology
[0002] In Linux systems, to simultaneously meet the requirements of data security and I / O acceleration, a Device Mapper (DM) framework is commonly used to construct independent encryption and caching layers, forming a serially stacked architecture for encryption and caching. In this architecture, the encryption target driver (e.g., built on dm-crypt) and the caching target driver (e.g., built on dm-cache) are two independent DM modules, each registered with the DM core layer. All I / O requests must be parsed and forwarded through the DM core layer, serially transmitted between the two driver modules. However, this stacked model lacks sufficient flexibility in security control. As a general-purpose block encryption driver, the encryption target driver cannot perceive the scheduling strategy and data hot / cold status of the caching target driver, and can only perform indiscriminate encryption. It is difficult to dynamically adjust encryption rules based on data type or address range, resulting in an inflexible balance between security and performance. Summary of the Invention
[0003] This application provides a data operation request processing method and system that enables fine-grained control of security policies, consistent encryption and decryption operations throughout the data lifecycle, and non-intrusive, low-risk dynamic deployment and update capabilities. It effectively balances data security and system performance, improves the maintainability of the storage system, and reduces upgrade and maintenance costs.
[0004] This application discloses a data manipulation request processing method, including: The cache target driver receives data operation requests transferred from the virtual file system. The first hook function associated with the cache target driver write entry point obtains and parses the data operation request. If the data operation request is a data write operation, the encryption tag is determined based on the write target data obtained by parsing the data operation request. Based on the encrypted tag, modify the private field corresponding to the write target data in the data operation request to obtain an update operation request. The first hook function sends the update operation request to the cache target driver. The cache target driver parses the update operation request to obtain and store the write target data, and generates and stores the metadata of the write target data with the encrypted tag.
[0005] Optionally, determining the encryption token based on the write target data obtained by parsing the data operation request includes: Determine the data attributes of the target data to be written, wherein the data attributes are the data type of the target data to be written or the physical address to which the target data is written; Based on the data attributes, determine whether the target data to be written needs to be encrypted; if so, set the encryption flag to indicate that encryption is required. If not, set the encryption flag to "no encryption required".
[0006] Optionally, modifying the private field corresponding to the write target data in the data operation request based on the encrypted tag to obtain the update operation request includes: If the encryption flag indicates that encryption is required, the private field in the data operation request in the form of an I / O request is set to 1; if the encryption flag indicates that encryption is not required, the private field in the data operation request in the form of an I / O request is set to 0, thus obtaining the update operation request.
[0007] Optionally, the data manipulation request processing method may also include: When the cache target driver initiates data flushing, the cache target driver generates and sends a data flushing request based on the encrypted token; The second hook function associated with the cache target driver flushing exit obtains the data flushing request and parses the dirty data and corresponding encryption tags of the data flushing; If the encryption flag indicates that encryption is required, the data flush request is transmitted to the encryption target driver for encryption and then transmitted to the physical storage device. If the encryption flag indicates that encryption is not required, the data flush request is transmitted to the physical storage device.
[0008] Optionally, transmitting the data flush request to the encrypted target driver for encryption before transmitting it to the physical storage device includes: The data refresh request is transmitted to the DM core layer, and then transmitted to the physical storage device through the DM core layer.
[0009] Optionally, if the encryption flag indicates that encryption is not required, transmitting the data flush request to the physical storage device includes: If the encryption flag indicates that encryption is not required, the update flush request is obtained by modifying the I / O interface in the data flush request to the device interface corresponding to the physical storage device. The update flush request is transmitted to the DM core layer, and then transmitted to the physical storage device through the DM core layer.
[0010] Optionally, the cache target driver forms a data refresh request based on the encrypted token and sends it, including: The cache target driver acquires the dirty data to be flushed; Read the metadata corresponding to the dirty data, the metadata containing the encryption tag; The dirty data and its corresponding encryption token are encapsulated to form a data flush request and sent out through the flush exit.
[0011] Optionally, the data manipulation request processing method may also include: If the data operation request is a data read operation, the cache target driver retrieves the corresponding read target data from the cache based on the data read operation, and if it exists, returns the read target data; If it does not exist, the third hook function associated with the physical storage device output port obtains the data reading result returned by the physical storage device, which carries the target data to be read; Based on the data reading result, determine whether the target data is encrypted. If so, transmit the data reading result through the DM core layer to the encrypted target driver for decryption and then transmit it to the cache target driver. If not, transmit the data reading result through the DM core layer to the cache target driver.
[0012] Optionally, transmitting the data read result to the cache target driver through the DM core layer includes: Modify the I / O interface in the data read result to cache the target driver to obtain the updated read result; The update read result is sent to the DM core layer, and then transmitted to the cache target driver through the DM core layer.
[0013] This application also discloses a data operation request processing system, including a cache target driver, a first hook function, and an encryption target driver; The cache target driver receives data operation requests transferred from the virtual file system. The first hook function associated with the cache target driver write entry point obtains and parses the data operation request. If the data operation request is a data write operation, the encryption tag is determined based on the write target data obtained by parsing the data operation request. Based on the encrypted tag, modify the private field corresponding to the write target data in the data operation request to obtain an update operation request. The first hook function sends the update operation request to the cache target driver. The cache target driver parses the update operation request to obtain and store the write target data, and generates and stores the metadata of the write target data with the encrypted tag.
[0014] As can be seen from the above technical solution, while maintaining the existing DM independent cache target driver and independent encryption target driver architecture, this application, by integrating a first hook function at the cache write source, can parse the attributes of each data write request in real time and dynamically generate differentiated encryption tags. This achieves fine-grained and programmable control of security policies at the granularity of the write target data of a single data operation request, effectively solving the performance waste problem caused by indiscriminate encryption and decryption, and balancing data security and system performance. Simultaneously, the encryption tags are injected into the request's private fields and persisted in the cache metadata, realizing the binding collaboration between encryption decision information and data, ensuring the consistency of encryption and decryption operations throughout the data's lifecycle. Furthermore, the encryption and decryption decision logic is implemented in the form of hook functions, which can be dynamically loaded, unloaded, and updated during system runtime without modifying the cache target driver code, restarting the service, or reconstructing the device stack. This provides non-intrusive, low-risk dynamic deployment and update capabilities, effectively improving the maintainability of the storage system and the response speed to dynamic security needs, while reducing the risks and costs of system upgrades and maintenance. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a data operation request processing method according to an embodiment of this application; Figure 2 This is a flowchart illustrating a data operation request processing method S200 in an embodiment of this application; Figure 3 This is a schematic diagram of the data refresh process of a data operation request processing method in an embodiment of this application; Figure 4 This is a flowchart illustrating a data operation request processing method S430 in an embodiment of this application. Figure 5 This is a flowchart illustrating a data operation request processing method S410 in an embodiment of this application. Figure 6 This is a flowchart illustrating a data read operation of a data operation request processing method according to an embodiment of this application. Figure 7 This is a flowchart illustrating a data operation request processing method S530 in an embodiment of this application. Figure 8 This is a schematic diagram of the structure of a data operation request processing system according to an embodiment of this application. Detailed Implementation
[0017] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without such specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.
[0018] In this application embodiment, to facilitate understanding of the technical solution, a typical Linux storage system architecture and its processing flow in the prior art are first described. In existing Linux storage systems, the DM framework is commonly used to construct independent encryption and caching layers, forming a serially stacked block device stack.
[0019] A specific example of existing technology is as follows: In a storage device configuration project, the backend physical storage device is an independent redundant disk array 0 composed of two mechanical hard drives. The caching medium is a solid-state drive using the non-volatile memory host controller interface specification protocol. An encryption target driver is built based on dm-crypt, and a cache target driver is built based on dm-cache. The cache target driver provides mountable block device nodes. When an application initiates a write operation, the virtual file system converts the write operation into a block device input / output request and sends it to the cache target driver. After completing cache scheduling, the cache target driver forwards the request to the encryption target driver through the DM core layer. After encrypting the data, the encryption target driver forwards the request again to the backend physical storage device through the DM core layer. When performing a cache flush operation, the cache target driver also needs to submit a request through the interface provided by the DM core layer, and this request will re-enter the scheduling queue of the DM core layer.
[0020] In this existing technical architecture, the encryption target driver, acting as a generic block device encryption driver, is unaware of the data scheduling strategy and data status of the cache target driver. It performs indiscriminate encryption and decryption operations on all passing I / O data, and cannot dynamically adjust encryption rules based on data type or write address range. For example, file system metadata typically does not require encryption protection, but it is still forcibly encrypted in this architecture, increasing unnecessary computational overhead and processing latency. When administrators want to adjust the encryption strategy, such as changing data in a logical block address range from plaintext storage to encrypted storage, the DM mapping table must be reconfigured, resulting in a complex operation process that may cause business interruption.
[0021] In view of this, embodiments of this application provide a data operation request processing method and system, and the technical solution of this application will be described in detail below with reference to specific embodiments.
[0022] This application provides a data operation request processing method, such as... Figure 1 As shown, it includes: S100: Caches data operation requests received by the target driver from the virtual file system; S200: The first hook function associated with the cache target driver write entry obtains and parses the data operation request. If the data operation request is a data write operation, the encryption tag is determined based on the write target data obtained by parsing the data operation request. S300: Based on the encrypted tag, modify the private field of the write target data in the data operation request to obtain an update operation request. The first hook function sends the update operation request to the cache target driver. The cache target driver parses the update operation request to obtain and store the write target data, and generates and stores the metadata of the write target data with the encrypted tag.
[0023] It should be noted that, in this embodiment, the cache target driver refers to a cache management module implemented based on the DM framework, which presents itself as a standard block device node upwards and is associated with physical storage devices downwards. This cache target driver is used to maintain the data mapping and scheduling between the high-speed cache medium and the backend slow physical storage devices. It does not implement complete encryption / decryption algorithms itself, but rather completes data security processing in conjunction with an attached independent encryption target driver. The encryption target driver refers to an independent driver module implemented based on the DM framework that provides block device-level encryption and decryption functions. The physical storage device refers to the backend block device that ultimately persists the data, which can be a hard disk drive, a solid-state drive, or an array thereof.
[0024] The virtual file system is a unified interface layer in the Linux kernel that abstracts operations from various specific file system operations. It shields applications from the differences between different file systems and provides a standard file operation interface. Data operation requests in the block device subsystem typically exist as a BIO (Block Input / Output) structure, containing information such as the operation type, logical block address, data length, the memory address corresponding to the data, a callback function to be executed after the operation, and the private data fields corresponding to the request. The cache target driver receives this data operation request through a registered set of block device operation functions.
[0025] In this embodiment, the data operation request processing method is implemented based on the Linux kernel's DM framework, but differs from the existing technology's simple stacking of independent encryption and cache drivers. The cache target driver can receive data operation requests transmitted by the virtual file system, and the first hook function associated with the cache target driver's write entry point obtains and parses the data operation request. The write entry point corresponds to the processing entry point of the cache target driver receiving data operation requests, which is a hook point in the Linux kernel that can mount Extended Berkeley Packet Filter (eBPF) programs. The first hook function is an eBPF hook program, whose loading and mounting can be completed through a user-space loading tool. During the loading phase, the eBPF program undergoes security verification by the kernel verifier to ensure that it will not affect kernel stability.
[0026] The first hook function can parse multiple pieces of information in the request, including the operation type, flag information carried in the request, the logical block address range corresponding to the request, and the private data of the request. If the data operation request is determined to be a data write operation, the write target data to be written can also be parsed from the data operation request, and an encryption flag can be determined based on the write target data. The write target data is the data content and associated attributes to be stored in the storage device by the write operation. The encryption flag is used to indicate whether the data needs to be encrypted by the encryption target driver before being subsequently flushed back to the physical storage device, and its value can be one of two states: encryption required or encryption not required.
[0027] In this embodiment, eBPF is a technical framework provided by the Linux kernel that allows user-defined programs to run in kernel space. After verification by the kernel verifier, the eBPF program runs in a security sandbox and can capture kernel events and data through preset hook points to execute custom processing logic. The eBPF program can be dynamically loaded and unloaded during system operation without modifying the kernel source code or recompiling the kernel or driver modules. An eBPF mapping is a key-value data structure stored in kernel space that can be accessed by both eBPF programs and user-space programs. It supports various formats such as hash mapping, array mapping, and ring buffer mapping, and can be used to transmit security policies and configuration parameters.
[0028] After determining the encryption tag, the first hook function modifies the private field of the write target data in the data operation request based on the encryption tag, thereby obtaining the update operation request. Specifically, the BIO structure contains a private field named bi_private, which is used by the driver module to carry custom data. The first hook function securely accesses this field using an eBPF helper function and writes the value of the encryption tag into it. The first hook function sends the update operation request to the main processing flow of the cache target driver. The cache target driver parses the update operation request to obtain the write target data and stores it in the cache medium. At the same time, it generates metadata for the write target data and stores the encryption tag in the metadata. Metadata is management information describing the state of cached data, including logical block address, cache medium physical address, dirty data flag, access frequency, etc. The encryption tag in this application is persistently stored as an attribute item in the metadata and is bound to the lifecycle of the data block.
[0029] Through the above process, this embodiment of the application completes the recording of control information required for subsequent encryption and decryption processing at the starting node when the data enters the cache storage. This tag also exists in the private field of the I / O request for use in this processing flow, and is also persisted in the metadata for subsequent asynchronous processing flow queries.
[0030] In alternative implementations, such as Figure 2 As shown, step S200, which determines the encryption token based on the write target data obtained by parsing the data operation request, includes: S210: Determine the data attribute of the target data to be written, wherein the data attribute is the data type of the target data to be written or the physical address to which the target data to be written is written.
[0031] S220: Determine whether the target data to be written needs to be encrypted based on the data attributes. If so, set the encryption flag to require encryption.
[0032] S230: If not, set the encryption flag to not require encryption.
[0033] Specifically, determining the encryption tag based on the write target data obtained from parsing the data operation request requires first determining the data attributes of the write target data. The data attributes are either the data type of the write target data or the physical address to which the write target data is written. By pre-setting the association mapping relationship between different data attributes to determine whether data needs encryption, flexible control over encryption and decryption during data storage can be achieved. Data that does not need encryption does not need to undergo encryption / decryption processing driven by the encryption target, thereby avoiding frequent DM core layer parsing mappings and improving the flexibility and efficiency of data operation request processing.
[0034] In specific examples, data types can be identified through flags carried in the request. For instance, the Linux kernel block layer provides flags like REQ_META, indicating that the request operates on file system metadata; it can also be information indicating priority, such as REQ_PRIO. The physical address refers to the range of logical block addresses targeted by the request, i.e., the starting and ending logical block addresses or length. The first hook function obtains the security policy configured by user space through eBPF mapping, i.e., the mapping rules for whether different data attributes need encryption. eBPF mapping can use hash mapping, with the device number or policy identifier as the key and the rule set as the value. Various filtering rules can be defined in the security policy, such as data type rules: file system metadata is not encrypted; address range rules: data within the logical block address range 0x000000 to 0x0FFFFF is not encrypted, other addresses are encrypted, or specific database address ranges need encryption. Priorities can be set between rules, for example, global switches have the highest priority, followed by data type rules, and then address rules.
[0035] The first hook function determines whether the target data needs to be encrypted based on the data attributes and the security policy: if the data attributes match a rule requiring encryption, the encryption flag is set to require encryption; if they match a rule not requiring encryption, the encryption flag is set to not require encryption; if no rule is matched, the default policy applies, such as requiring encryption or not requiring encryption by default. For example, if the security policy specifies that metadata does not need encryption, when the first hook function parses the write request flag containing REQ_META, it determines that the data is metadata and sets the encryption flag to not require encryption; if the security policy specifies that logical block addresses between 0x400000 and 0x500000 require encryption, when the requested logical block address falls within this range, even if the data is ordinary file data and may have previously been in the unencrypted range, it will be determined to require encryption after the policy update.
[0036] It should be noted that the security policy configuration is written to the eBPF mapping by the administrator through userspace tools, and this mapping can be read in real time by the first hook function. When the administrator needs to change the policy, such as dynamically including a database address range that was originally unencrypted into the encrypted range, they only need to update the corresponding rule in the eBPF mapping. Subsequent write requests arriving at that address range will immediately obtain the new encryption tag at the first hook function, without needing to restart any services or drivers, or reconfigure the DM device stack.
[0037] The step S300, which modifies the private field of the target data in the data operation request based on the encryption tag to obtain the update operation request, includes: S310: If the encryption flag indicates that encryption is required, set the private field in the data operation request in the form of an I / O request to 1; if the encryption flag indicates that encryption is not required, set the private field in the data operation request in the form of an I / O request to 0, and obtain the update operation request.
[0038] In one implementation, modifying a private field based on an encryption flag to obtain an update operation request includes: if the encryption flag indicates encryption is required, setting the private field in the data operation request (in the form of an I / O request) to 1, i.e., setting a preset data bit (based on requirements) in the private field to represent the encryption flag to 1; if the encryption flag indicates encryption is not required, setting the private field to 0, i.e., setting a preset data bit in the private field to represent the encryption flag to 0. This assignment operation is securely completed in the kernel context through auxiliary functions such as eBPF's bpf_set_priv or direct pointer access, changing only control information and not actually modifying the write target data itself. Furthermore, the private field can further store information such as the encryption rule index and key identifier, so that more detailed processing parameters can be directly obtained in subsequent stages, but this application does not limit its specific encoding format. The modified update operation request is returned to the write path of the cache target driver by the first hook function.
[0039] After receiving an update request, the cache target driver extracts the write target data and writes it to the cache medium in plaintext. The write cache strategy can employ either write-back or write-through mode, depending on the specific cache target driver configuration. Simultaneously, the cache target driver creates or updates the cache metadata corresponding to the logical block address storing the write target data, writing the encryption token as a field in the metadata. This double-recording method ensures that even after the I / O request is processed and the BIO is released, the encryption intent of the data block can still be queried through the metadata.
[0040] The following example illustrates the write process described above. In this example, the backend physical storage device is a mechanical hard disk array, the cache medium is a solid-state drive, and the user-configured security policy is: file system metadata does not need to be encrypted, but ordinary user data needs to be encrypted. When the virtual file system issues a metadata write operation request with the REQ_META flag, the cache target driver receives the request and triggers the first hook function at the write entry point. The first hook function parses the flag of the request, confirms that it is REQ_META, and then queries the security policy in the eBPF mapping to find that the file system metadata does not need to be encrypted. The first hook function sets the preset data bit of the BIO private field to 0, generates an update operation request, and returns. The cache target driver writes the file system metadata to the cache medium in plaintext and records the encryption flag as 0 in the corresponding metadata. When a user data write operation request without the REQ_META flag is received, the first hook function parses its logical block address and flag, determines that the data needs to be encrypted according to the policy, and sets the preset data bit of the BIO private field to 1. The cache target driver writes the user data in plaintext to the cache medium, marks it as dirty data, and records the encryption flag as 1 in the corresponding metadata. It should be noted that regardless of whether encryption is required, the data is stored in plaintext at the cache layer. This ensures that when a cache hit occurs, the plaintext is returned immediately without decryption.
[0041] In alternative implementations, such as Figure 3 As shown, the data operation request processing method further includes: S410: When the cache target driver initiates data flushing, the cache target driver generates and sends a data flushing request based on the encrypted token; S420: The second hook function associated with the cache target driver flush exit obtains the data flush request and parses the dirty data and corresponding encryption tokens of the data flush; S430: If the encryption flag indicates that encryption is required, the data flush request is transmitted to the encryption target driver for encryption and then transmitted to the physical storage device. If the encryption flag indicates that encryption is not required, the data flush request is transmitted to the physical storage device.
[0042] Furthermore, this application embodiment also covers a data flushing process. When the cache target driver initiates data flushing, for example, due to triggering conditions such as the cache medium's free space falling below a set threshold, scheduled flushing, device removal, or system shutdown, the cache target driver forms a data flushing request based on the stored encrypted markers and sends it. Specifically, the cache target driver obtains the dirty data to be flushed. Dirty data refers to data blocks in the cache that have been modified, are inconsistent with the corresponding data block content on the backend physical storage device, and have not yet been written back. The cache target driver reads the metadata corresponding to the dirty data, which contains the previously written encrypted marker field. Then, the cache target driver encapsulates the dirty data content and the corresponding encrypted marker to form a data flushing request, which is sent through the flushing exit. The flushing request is also represented as a BIO structure, whose private fields again carry the encrypted markers, ensuring that hook functions can directly read from the request.
[0043] The second hook function associated with the cache target driver's flush exit node obtains the data flush request and parses it to obtain the dirty data to be flushed and the corresponding encryption flag. The flush exit node is also a hook point that can mount eBPF programs. The second hook function makes a decision based on the value of the encryption flag: if the encryption flag indicates that encryption is required, the second hook function transmits the data flush request to the encryption target driver for encryption, and the encrypted data is then transmitted to the physical storage device; if the encryption flag indicates that encryption is not required, the second hook function transmits the data flush request directly to the physical storage device without going through the encryption target driver.
[0044] The step S430 involves transmitting the data refresh request to the encrypted target driver for encryption before transmitting it to the physical storage device, including: S431: The data refresh request is transmitted to the DM core layer, and then transmitted to the physical storage device through the DM core layer.
[0045] Specifically, when the encryption flag indicates that encryption is required, the second hook function maintains the target device of the flush request as the mapped device corresponding to the encryption target driver, and transmits the data flush request to the DM core layer. The DM core layer then forwards the request to the encryption target driver according to its own mapping table. Upon receiving the request, the encryption target driver uses the encryption key and algorithm pre-configured for the device corresponding to the encryption target driver to encrypt the data in the request, generating ciphertext data and replacing the original plaintext data in the request. Subsequently, the encryption target driver again transmits the request containing the ciphertext data to the driver interface of the physical storage device through the DM core layer, completing the data persistence to disk. This path fully utilizes the hierarchical forwarding capabilities of the existing DM framework, ensuring that the data requiring encryption undergoes complete encryption processing and guaranteeing data confidentiality.
[0046] In alternative implementations, such as Figure 4As shown, step S430, if the encryption flag indicates that encryption is not required, transmitting the data flush request to the physical storage device includes: S432: If the encryption flag indicates that encryption is not required, modify the I / O interface in the data refresh request to the device interface corresponding to the physical storage device to obtain the update refresh request; S433: Transmit the update flush request to the DM core layer, and then transmit the update flush request to the physical storage device through the DM core layer.
[0047] When encryption is marked as unnecessary, continuing through the encryption target driver via the original path introduces unnecessary encryption calculations and requires frequent mapping and forwarding by the DM core layer, resulting in additional path latency. Therefore, the second hook function modifies the I / O interface in the data flush request to the device interface corresponding to the physical storage device, directly pointing to the driver submission interface of the backend physical storage device, generating an update flush request. Then, the second hook function transmits this update flush request to the DM core layer, which directly forwards the request to the physical storage device, bypassing the encryption target driver's mapping to the device. In this way, data that does not require encryption skips the encryption target driver during flushing, shortening the I / O processing path and reducing CPU encryption calculation overhead and the number of forwarding operations by the DM core layer.
[0048] It should be noted that the second hook function can modify the target interface of the request by modifying the bi_bdev or bi_disk fields in the BIO structure through eBPF helper functions, making them point to the physical storage device, and simultaneously setting the corresponding commit function pointer. These modifications are performed within the kernel eBPF security constraints, ensuring that the integrity of kernel data structures is not compromised.
[0049] In alternative implementations, such as Figure 5 As shown, the S410 cache target driver forms a data refresh request based on the encrypted token and sends it, including: S411: The cache target driver obtains the dirty data to be flushed.
[0050] S412: Read the metadata corresponding to the dirty data, wherein the metadata contains the encryption tag.
[0051] S413: Encapsulate the dirty data and the corresponding encryption tag to form a data refresh request and send it out through the refresh exit.
[0052] Specifically, the cache target driver scans the cache for data blocks marked as dirty according to its internal flushing policy and eviction algorithm. For each dirty data block to be flushed, the cache target driver first obtains the metadata corresponding to the dirty data. This metadata was generated along with the data and stored in the cache medium or memory during the previous cache write, and it fully records the specific value of the encryption flag set by the first hook function. The cache target driver reads this encryption flag, then encapsulates the dirty data content and the encryption flag into a new BIO request, and fills the encryption flag into the private field of this BIO, thus constructing a data flushing request. This request is then submitted to the cache target driver's flushing exit, thereby triggering the second hook function. In this way, the encryption flag in the metadata is passed into the flushing request for direct use by the second hook function.
[0053] In a file system metadata write operation scenario, the file system metadata is marked as dirty data in the cache, and its encryption flag is 0. When the cache target driver triggers a flush, it reads the file system metadata to obtain the encryption flag 0, constructs a data flush request, and sends it through the flush exit. The second hook function, after receiving this flush request, parses it and finds the encryption flag is 0, determining that the data does not need encryption. The second hook function modifies the I / O interface of the request to the physical storage device's driver submission interface, generates an update flush request, and then sends it directly to the mechanical hard drive array through the DM core layer. Ultimately, the file system metadata is written to the physical storage device in plaintext and does not enter the encryption target driver during the entire flush path.
[0054] In a user data write operation scenario, the encryption flag for this data block is set to 1. After the second hook function parses the encryption flag as 1, it does not change the path of the flush request and forwards it to the encryption target driver through the DM core layer using the default path. The encryption target driver uses its configured key to encrypt the data, generating ciphertext, and then writes the ciphertext to the physical storage device through the DM core layer. In this way, the same cache target driver uses different transmission paths for flushing different data blocks, achieving differentiated security control.
[0055] In one extended implementation, when an administrator updates the security policy online, changing a logical block address range from unencrypted to encrypted, newly written data within that range will be marked as requiring encryption at the first hook function. Subsequent flushes will always be encrypted via the encryption target driver before being written to disk. For old, dirty data that was cached before the policy change but has not yet been flushed, its metadata still records the old encryption marker. During a flush, the second hook function will still select the path based on the actual encryption marker stored in the metadata; that is, the old data will still execute the flush path according to the policy at the time of writing. This gradual policy activation mechanism avoids the risk of data inconsistency during policy switching and does not require a forced flush of all dirty data. As the system runs, the old data is naturally discarded or overwritten, and the new policy will take full effect.
[0056] In alternative implementations, such as Figure 6 As shown, the data operation request processing method further includes: S510: If the data operation request is a data read operation, the cache target driver obtains the corresponding read target data in the cache based on the data read operation, and if it exists, returns the read target data.
[0057] S520: If it does not exist, the third hook function associated with the physical storage device output port obtains the data reading result returned by the physical storage device carrying the target data.
[0058] S530: Based on the data reading result, determine whether the target data to be read is encrypted data. If so, transmit the data reading result through the DM core layer to the encrypted target driver for decryption and then transmit it to the cache target driver. If not, transmit the data reading result through the DM core layer to the cache target driver.
[0059] In this embodiment, if the data operation request is a data read operation, the cache target driver first searches for the corresponding read target data in the cache based on the data read operation. The cache target driver maintains a mapping table between logical block addresses and physical addresses of the cache medium. If the requested logical block address has a valid mapping in the table and the data block is not marked as invalid, it indicates a cache hit. At this time, the cache stores plaintext data, and the cache target driver directly returns the read target data to the virtual file system to complete the read operation response without triggering an encryption / decryption process.
[0060] If a cache miss occurs, meaning the requested data is not in the cache, the cache target driver needs to read the data from the backend physical storage device. This read operation constructs a read request and sends it to the backend device. Once the physical storage device completes the read operation, the returned data read result, carrying the target data, is returned along the block device stack. At this point, the third hook function associated with the physical storage device's output port will retrieve this data read result. The physical storage device's output port refers to a node on the data return path after the physical storage device has completed the read operation, where an eBPF hook can be attached.
[0061] After obtaining the data read result, the third hook function determines whether the target data is encrypted based on this result. This can be done by checking if the private fields of the data read result contain an encryption flag. If the encryption flag was persisted to the physical storage device along with the data during writing (e.g., through additional metadata areas or extended attributes), it can be retrieved as is during reading. If the physical storage device does not store the flag, the third hook function can obtain the current address security policy from the eBPF mapping and determine whether the address range belongs to the encryption range based on the requested logical block address. If the current policy specifies encryption for that address, the data is considered encrypted; otherwise, it is unencrypted. Alternatively, it can also be determined by combining this information with any residual metadata information in the cache.
[0062] If the third hook function determines that the target data to be read is encrypted, the data read result is transmitted through the DM core layer to the encryption target driver for decryption. The encryption target driver uses the corresponding device key and algorithm to decrypt the ciphertext data, generating plaintext data, which replaces the data in the read result. The decrypted plaintext is then transmitted by the encryption target driver to the cache target driver through the DM core layer. If the third hook function determines that the target data to be read is unencrypted, the data read result is directly transmitted to the cache target driver through the DM core layer without decryption.
[0063] In alternative implementations, such as Figure 7 As shown, the S530 transmits the data reading result to the cache target driver through the DM core layer, including: S531: Modify the I / O interface in the data read result to the cache target driver to obtain the updated read result; S532: Send the update read result to the DM core layer, and then transmit the update read result to the cache target driver through the DM core layer.
[0064] In one implementation, when the third hook function determines that the data is unencrypted and does not need to be decrypted by the encryption target driver, the third hook function modifies the I / O interface in the data read result to the device interface corresponding to the cache target driver, that is, modifies the target device to the block device of the cache target driver, and obtains an updated read result; then the updated read result is sent to the DM core layer, which transmits it to the commit path of the cache target driver according to the updated target device. In this way, plaintext data can directly reach the cache target driver without entering the encryption target driver, reducing processing steps.
[0065] It should be noted that the way the third hook function modifies the I / O interface can be referenced from the operation of the second hook function during the flush, which is accomplished by modifying the BIO target block device pointer.
[0066] When an upper-layer application initiates a read operation on a previously written and flushed user data block, and the cache target driver finds a cache miss after querying the mapping table, the cache target driver sends a read request to the backend physical storage device. The encrypted data returned by the physical storage device is intercepted by a third hook function along the data read path. The third hook function queries the current address policy in the eBPF mapping based on the requested logical block address, determines that the address range belongs to the encrypted range, and thus confirms that the target data to be read is encrypted data. The third hook function forwards the data read result to the encryption target driver through the DM core layer. The encryption target driver decrypts the ciphertext using the key within its corresponding device and returns the plaintext data to the cache target driver through the DM core layer. The cache target driver can write the plaintext to the cache medium, establish a mapping, and simultaneously return the plaintext data to the upper-layer application.
[0067] When a read request targets metadata, the physical storage device returns plaintext data. The third hook function determines that the data is unencrypted based on the address policy or by checking private fields. It then modifies the target interface in the read result to the cache target driver interface, generates an updated read result, and sends it to the DM core layer, directly transmitting it to the cache target driver. The cache target driver can also write plaintext data to the cache medium and return it to the upper-layer application.
[0068] Those skilled in the art will understand that the third hook function, when handling read operations, may encounter situations where the data returned by the physical storage device lacks an inherent cryptographic marker. In an alternative implementation, a persistent header with a cryptographic marker can be appended to the end of each data entry requiring encryption at a specific offset during data writing, and this header can be written by the second hook function before or after encryption. This way, during reading, regardless of whether cached metadata exists, the cryptographic marker can be recovered from the data itself, avoiding inconsistencies before and after policy changes that might result from relying entirely on external policies.
[0069] It should be noted that the first, second, and third hook functions mentioned above can be developed, loaded, unloaded, and updated independently. They share security policies and some tagging information through eBPF mapping, and can also directly pass tags through private fields in data requests, thus ensuring the consistency of encryption and decryption operations throughout the entire data processing lifecycle. The loading order of the hook functions is not limited; system administrators can mount some or all hook programs according to actual needs. For example, if the business scenario does not require on-demand decryption of read paths, only the first and second hook functions can be mounted, and the read path can continue to use the original method.
[0070] In this embodiment, the differentiated encryption / decryption control is implemented using eBPF hook functions without modifying the existing code of the cache target driver and encryption target driver. The cache target driver only needs to expose standard kernel hook points, such as the block device commit entry (write entry), flush exit, and appropriate mount points on the read return path. These hook points can be provided through mechanisms such as Linux kernel tracepoints or kprobe, or by explicitly adding eBPF program calls to the driver code. This application does not limit the specific hook mounting mechanism. Since it does not require modification of the mapping logic of the DM core layer, nor does it require merging the encryption and caching modules, this scheme has good compatibility with the existing Linux storage stack ecosystem.
[0071] For example, the complete lifecycle of the configuration example above will be used for illustration. During system initialization, the administrator writes the security policy into the eBPF mapping using the eBPF userspace tool. The policy content is that metadata is not encrypted and user data is encrypted. At the same time, the encryption target driver is created and configured with the encryption key, and the cache target driver is created and bound to the encryption target driver as the backend. Each hook program is compiled, verified, and loaded into the corresponding hook point.
[0072] When a metadata write request arrives, the first hook function sets the encryption flag to "no encryption required" according to the policy, sets the private field to 0, and stores the data in plaintext in the cache target driver, recording flag 0. During subsequent flushes, the second hook function reads flag 0, modifies the target interface to the physical storage device, and writes the data to disk in plaintext. When a read request arrives, if the cache is hit, it returns the plaintext directly; if the cache is not hit, the third hook function reads the data from the backend and determines it to be plaintext based on the policy or flag, directly routing the data to the cache target driver and caching it.
[0073] When a user data write request comes in, the first hook function sets the encryption flag to indicate that encryption is required, sets the private field to 1, caches the plaintext, and records flag 1. During a flush, the second hook function reads flag 1, sends the request to the encryption target driver for encryption, and then writes it to disk. If a read miss occurs, the third hook function determines that the data is ciphertext, forwards the read result to the encryption target driver for decryption, returns it to the cache, and caches it again.
[0074] When an administrator needs to change the database address range 0x400000 to 0x500000 from unencrypted to encrypted, the address policy in the eBPF mapping is updated in user space. Subsequent write requests falling within this range will be marked as requiring encryption in the first hook function, and their flush and read paths will be automatically processed as encrypted data. Dirty data already existing in this range will be processed with the old mark during its own flush, until it is naturally discarded. The entire process does not require uninstalling the device or restarting the driver or service, achieving online dynamic application of the policy.
[0075] For read paths, if the policy for a certain address changes from unencrypted to encrypted, but the physical storage device still contains plaintext data written before the policy change, the third hook function will mistakenly identify it as encrypted data based on the current policy and send it to the encryption target driver for decryption. Decryption of plaintext data will lead to data corruption. To prevent this, a relevant cache invalidation operation can be triggered when the policy changes, or persistent encrypted markers can be used for differentiation as mentioned earlier. In actual deployment, it can be ensured that dirty cache data within the relevant address range has been flushed before the policy change, or the read-side judgment logic can be updated simultaneously to add additional verification.
[0076] In addition, when the cache target driver needs to clear the cache due to system shutdown or device removal, the hook program can traverse the encryption tags in the cache metadata and perform zeroing operations on sensitive information in the relevant memory areas. For example, it can overwrite the memory area storing encryption tags and key indexes with all zeros to ensure that no sensitive policy information remains.
[0077] This application also provides a data manipulation request processing system, such as... Figure 8 As shown, the system includes a cache target driver 11, a first hook function 12, and an encryption target driver 13. The cache target driver 11 receives data operation requests transmitted by the virtual file system. The first hook function 12, associated with the write entry point of the cache target driver 11, obtains and parses the data operation request. If it is a data write operation, it determines an encryption flag based on the parsed write target data. The first hook function 12 is also used to modify the private fields of the write target data in the data operation request based on the encryption flag to obtain an update operation request, and sends the update operation request to the cache target driver 11. The cache target driver 11 is also used to parse the update operation request to obtain and store the write target data, and generate and store metadata of the write target data with encryption flags.
[0078] Optionally, the data operation request processing system also includes a second hook function 14, which is mounted on the refresh exit of the cache target driver. It is used to obtain the data refresh request and parse the encryption tag during refresh: if encryption is required, the data refresh request is transmitted to the encryption target driver 13 for encryption and then transmitted to the physical storage device 15; if encryption is not required, the data refresh request is transmitted to the physical storage device 15.
[0079] Optionally, the data operation request processing system also includes a third hook function 16, which is mounted on the output port of the physical storage device 15 or the read data return path. It is used to obtain the data read result returned by the physical storage device 15 when a read miss occurs. If the target data to be read is encrypted data, the data read result is transmitted through the DM core layer to the encrypted target driver 13 for decryption and then transmitted to the cache target driver 11. If it is unencrypted data, it is transmitted directly to the cache target driver 11 through the DM core layer.
[0080] For example, the hook functions in the above system can be centrally managed through a unified eBPF userspace management program, and security policies can be uniformly configured through eBPF mappings. The hook functions can share policy mappings in read-only mode, or pass encrypted tags for specific data through BIO private fields. The cache target driver integrates logic for cache address mapping table management, hot / cold data grading, dirty data flushing, and cache data eviction, and is compatible with existing dm-cache or similar cache target drivers, requiring no additional encryption / decryption code. The encryption target driver retains its original design, only handling flushing and read requests routed by the hook functions.
[0081] Those skilled in the art will understand that the mounting locations of the first hook function, the second hook function, and the third hook function in the embodiments of this application are not limited to nodes with specific names. Any mountable point on the cache target driver data operation request entry, the flush exit, and the physical storage device read completion return path can be used as the mounting point for each hook function. Those skilled in the art can set the specific code of the first hook function, the second hook function, and the third hook function according to actual needs based on the specific functions of the first hook function, the second hook function, and the third hook function and implement the mounting. This is a conventional technical means in the art and will not be elaborated here.
[0082] Furthermore, in this embodiment, the use of eBPF mappings is not limited to a single mapping. Multiple mappings can be created to store global switches, data type rules, address range rules, key indexes, etc. User-space programs access these mappings through file descriptors, while kernel-space hook functions retrieve data through mapping IDs. This separation facilitates the management of various policies and supports individual updates to single rules.
[0083] When hook functions modify BIO private fields, care should be taken to avoid conflicts with the cache target driver's own use of that field. In implementation, it can be agreed that the high-order bits or specific bits of the private field are used to store the encryption tag, while the low-order bits are reserved for the original purpose of the cache target driver; alternatively, the cache target driver can allocate an additional pointer in its private data structure to store the encryption tag in the memory area pointed to by the pointer. This application does not limit the specific usage of private fields, but it must ensure that the tag can be correctly read in subsequent processing.
[0084] It should be noted that, in addition to being based on data type and address range, security policies can also incorporate information such as the requesting process context and user ID. This information can be obtained through eBPF helper functions in some versions of the Linux kernel, such as `bpf_get_current_uid_gid`. Therefore, the technical solution of this application can be further extended to differentiated encryption based on processes or users. For example, data from specific users could be encrypted, while data from other users could be stored in plaintext. This extension still falls within the scope of this application and should be protected within its scope.
[0085] Other features of the data operation request processing system provided in this application embodiment are similar to the data operation request processing method described above, and can be found in the description of the method above, and will not be repeated here.
[0086] In summary, the embodiments of this application insert eBPF hook functions into the critical path of the existing DM independent cache and encryption driver architecture, and use encryption tags to serialize the three stages of writing, flushing, and reading. This enables online hot updates of differentiated encryption control and strategies for different data operation requests without modifying the original driver code or reconstructing the device stack, thus solving the problem of flexibly balancing security and performance.
[0087] The above description is only a partial embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A data operation request processing method, characterized in that, include: The cache target driver receives data operation requests transferred from the virtual file system. The first hook function associated with the cache target driver write entry point obtains and parses the data operation request. If the data operation request is a data write operation, the encryption tag is determined based on the write target data obtained by parsing the data operation request. Based on the encrypted tag, modify the private field corresponding to the write target data in the data operation request to obtain an update operation request. The first hook function sends the update operation request to the cache target driver. The cache target driver parses the update operation request to obtain and store the write target data, and generates and stores the metadata of the write target data with the encrypted tag. The data operation request processing methods also include: When the cache target driver initiates data flushing, the cache target driver generates and sends a data flushing request based on the encrypted token; The second hook function associated with the cache target driver flushing exit obtains the data flushing request and parses the dirty data and corresponding encryption tags of the data flushing; If the encryption flag indicates that encryption is required, the data flush request is transmitted to the encryption target driver for encryption and then transmitted to the physical storage device. If the encryption flag indicates that encryption is not required, the data flush request is transmitted to the physical storage device.
2. The data operation request processing method according to claim 1, characterized in that, The encryption tokens determined based on the write target data obtained by parsing the data operation request include: Determine the data attributes of the target data to be written, wherein the data attributes are the data type of the target data to be written or the physical address to which the target data is written; Based on the data attributes, determine whether the target data to be written needs to be encrypted; if so, set the encryption flag to indicate that encryption is required. If not, set the encryption flag to "no encryption required".
3. The data operation request processing method according to claim 2, characterized in that, The update operation request, which modifies the private field corresponding to the write target data in the data operation request based on the encrypted tag, includes: If the encryption flag indicates that encryption is required, the private field in the data operation request in the form of an input / output request is set to 1; if the encryption flag indicates that encryption is not required, the private field in the data operation request in the form of an input / output request is set to 0, thus obtaining the update operation request.
4. The data operation request processing method according to claim 1, characterized in that, The process of transmitting the data flush request to the encrypted target driver for encryption before transmitting it to the physical storage device includes: The data refresh request is transmitted to the DM core layer, and then transmitted to the physical storage device through the DM core layer.
5. The data operation request processing method according to claim 1, characterized in that, If the encryption flag indicates that encryption is not required, transmitting the data flush request to the physical storage device includes: If the encryption flag indicates that encryption is not required, the update flush request is obtained by modifying the I / O interface in the data flush request to the device interface corresponding to the physical storage device. The update flush request is transmitted to the DM core layer, and then transmitted to the physical storage device through the DM core layer.
6. The data operation request processing method according to claim 1, characterized in that, The cache target driver generates a data refresh request based on the encrypted token and sends it, including: The cache target driver acquires the dirty data to be flushed; Read the metadata corresponding to the dirty data, the metadata containing the encryption tag; The dirty data and its corresponding encryption token are encapsulated to form a data flush request and sent out through the flush exit.
7. The data operation request processing method according to claim 1, characterized in that, Also includes: If the data operation request is a data read operation, the cache target driver retrieves the corresponding read target data from the cache based on the data read operation, and if it exists, returns the read target data; If it does not exist, the third hook function associated with the physical storage device output port obtains the data reading result returned by the physical storage device, which carries the target data to be read; Based on the data reading result, determine whether the target data is encrypted. If so, transmit the data reading result through the DM core layer to the encrypted target driver for decryption and then transmit it to the cache target driver. If not, transmit the data reading result through the DM core layer to the cache target driver.
8. The data operation request processing method according to claim 7, characterized in that, Transmitting the data reading result to the cache target driver through the DM core layer includes: Modify the I / O interface in the data read result to cache the target driver to obtain the updated read result; The update read result is sent to the DM core layer, and then transmitted to the cache target driver through the DM core layer.
9. A data operation request processing system, characterized in that, This includes the cache target driver, the first hook function, and the encryption target driver; The cache target driver receives data operation requests transferred from the virtual file system. The first hook function associated with the cache target driver write entry point obtains and parses the data operation request. If the data operation request is a data write operation, the encryption tag is determined based on the write target data obtained by parsing the data operation request. Based on the encrypted tag, the private field corresponding to the write target data in the data operation request is modified to obtain an update operation request. The first hook function sends the update operation request to the cache target driver. The cache target driver parses the update operation request to obtain and store the write target data, and generates and stores metadata of the write target data with the encrypted tag; wherein, When the cache target driver initiates data flushing, the cache target driver forms a data flushing request based on the encryption flag and sends it; the second hook function associated with the cache target driver's flushing exit obtains the data flushing request, parses it to obtain the dirty data of the data flushing and the corresponding encryption flag; if the encryption flag indicates that encryption is required, the data flushing request is transmitted to the encryption target driver for encryption and then transmitted to the physical storage device; if the encryption flag indicates that encryption is not required, the data flushing request is transmitted to the physical storage device.
Citation Information
Patent Citations
Cache security defense method and system based on RISC-V
CN119720298A
Transparent encryption and decryption method and system based on section programming and interceptor
CN121744363A