Container data storage system, method and device and electronic equipment
By introducing a three-tiered data storage strategy of dedicated memory space, dedicated NVMe hard disk, and shared hard disk in confidential containers, the problem of low storage efficiency of confidential containers in big data processing is solved, and flexible and efficient data storage and performance improvement are achieved.
Patent Information
- Application Number
- CN202511468376.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-11
AI Technical Summary
In confidential container application scenarios, data storage efficiency is low, especially when processing big data, there are performance degradation issues caused by memory capacity limitations and frequent physical address translation.
A container data storage system is provided, including dedicated memory space, dedicated NVMe hard disk and shared hard disk. Data is stored through hardware encryption and device pass-through. A three-level data storage strategy is adopted: dedicated memory space, dedicated NVMe hard disk and shared hard disk, respectively storing data with different frequency.
It improves the flexibility and rationality of data storage, enhances data storage efficiency, ensures data security, and improves the overall performance of confidential containers in big data processing scenarios.
Smart Images

Figure CN120930192A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of privacy computing technology and multi-party secure computing technology, and in particular to a container data storage system, method, apparatus and electronic device. Background Technology
[0002] Currently, in the practical application of privacy computing technology and multi-party secure computing technology, data computing can be processed based on confidential container technology in a TEE (Trusted Execution Environment), thereby further improving the security of data use.
[0003] In scenarios where confidential containers are used for big data processing, such as big data analysis, model training and inference based on big data, the confidential container technology is based on hardware isolation through memory encryption, which has the problem of memory capacity limitations. Therefore, in practical applications, it is usually accompanied by the storage of a large amount of data on disk.
[0004] For example, a large amount of intermediate result data is generated when performing big data analysis. Confidential containers have two main security features: kernel isolation and memory encryption. However, these two security features limit the performance of confidential containers when processing large amounts of data: (1) The confidentiality of confidential containers is runtime memory encryption. Its specific implementation requires setting a fixed amount of memory space (such as 128GB) for the container when it starts up, and encrypting the memory space with a specific hardware key. This makes it impossible for confidential containers to share the memory of the host machine and the memory space of other containers. The limitation of the memory access space of confidential containers means that data exceeding the memory space can only be written to disk for processing, which determines that there will be more read and write operations on external storage space (hard disk) during its operation. (2) The kernel isolation of confidential containers is manifested in that the underlying layer of confidential containers is a lightweight virtual machine. Confidential containers have a separate Linux kernel, thereby achieving kernel-to-kernel isolation security with the host machine's operating system. Malicious attacks on the host machine's kernel will not affect confidential containers. However, virtual machines need multiple layers of address translation to access external hard disk I / O: from the virtual machine logical address (Guest Virtual Address, GVA) to the virtual machine physical address (Guest Physical Address, GPA), and then to the host physical address (HPA). In big data analytics scenarios, the large number of I / O operations on external storage devices and the frequent physical address translation will cause the performance of confidential containers to drop significantly in this scenario. Summary of the Invention
[0005] This application provides a container data storage system, method, apparatus, and electronic device to solve the problem of low data storage efficiency in confidential container application scenarios in the prior art.
[0006] This application provides a container data storage system, including: multiple confidential containers, a dedicated memory space corresponding to each confidential container, a dedicated NVMe hard disk corresponding to each confidential container, and a shared hard disk of the physical host to which the multiple confidential containers belong; The dedicated memory space is used to store data from the corresponding confidential container using hardware encryption. The dedicated NVMe hard drive is directly connected to the corresponding confidential container via a device pass-through method, and is used to store data from the corresponding confidential container in an encrypted manner; The shared hard drive is used to store encrypted data from each of the confidential containers. The encrypted data generated by each confidential container is obtained by encrypting the data using the key generated by the confidential container.
[0007] Furthermore, the memory key used to store data in the dedicated memory space is a derived key generated by obfuscating the hardware root of trust with the container ID of the corresponding confidential container.
[0008] Furthermore, the dedicated NVMe hard drive employs hardware-assisted virtualization technology and a device pass-through method to directly access the corresponding confidential container.
[0009] Furthermore, the confidential container contains an application; The application is used to store the first type of data generated in the dedicated memory space corresponding to the confidential container using hardware encryption, according to a preset data storage strategy; to store the second type of data generated in the dedicated NVMe hard drive corresponding to the confidential container using encryption; and to store the third type of data generated in the shared hard drive after encryption, wherein the data heat of the first type of data, the second type of data and the third type of data decreases in that order.
[0010] Furthermore, the system includes multiple physical hosts and a container management module connected to the multiple confidential containers on the multiple physical hosts. The container management module is used to control the operation of the multiple confidential containers.
[0011] This application also provides a container data storage method based on any of the container data storage systems described above, applied to an application installed in a confidential container, including: During the execution of data processing tasks, data to be stored is generated; When the data to be stored is the first type of data, hardware encryption is used to store the first type of data in the dedicated memory space corresponding to the confidential container. When the data to be stored is the second type of data, an encryption method is used to store the second type of data in the dedicated NVMe hard drive corresponding to the confidential container. When the data to be stored is a third type of data, the third type of data is encrypted and then stored in the shared hard drive; The popularity of the first type of data, the second type of data, and the third type of data decreases in that order.
[0012] Furthermore, the application installed in the confidential container is the Executor module in the distributed computing engine Apache Spark; During the execution of data processing tasks, data to be stored is generated, including: During the execution of big data analysis tasks, data is generated that needs to be stored; When the data to be stored is a first type of data, hardware encryption is used to store the first type of data in the dedicated memory space corresponding to the confidential container, including: When the data to be stored is non-disk data, hardware encryption is used to store the non-disk data in the dedicated memory space corresponding to the confidential container. When the data to be stored is the second type of data, an encryption method is used to store the second type of data in the dedicated NVMe hard drive corresponding to the confidential container, including: When the data to be stored is intermediate result data generated by performing Spark Shuffle operation, the intermediate result data is stored in the dedicated NVMe hard disk corresponding to the confidential container using an encrypted method. When the data to be stored is a third type of data, encrypting the third type of data and storing it in the shared hard drive includes: When the data to be stored is the calculation result data generated after the task calculation is completed, the calculation result data is encrypted and then stored in the shared hard disk.
[0013] This application also provides a container data storage device based on any of the container data storage systems described above, applied to an application installed in a confidential container, including: The data generation module is used to generate data to be stored during the execution of data processing tasks; The data storage module is used to store the first type of data in the dedicated memory space corresponding to the confidential container when the data to be stored is the first type of data using hardware encryption. When the data to be stored is the second type of data, an encryption method is used to store the second type of data in the dedicated NVMe hard drive corresponding to the confidential container. When the data to be stored is a third type of data, the third type of data is encrypted and then stored in the shared hard drive; The popularity of the first type of data, the second type of data, and the third type of data decreases in that order.
[0014] Furthermore, the application installed in the confidential container is the Executor module in the distributed computing engine Apache Spark; The data generation module is specifically used to generate data to be stored during the execution of big data analysis tasks; The data storage module is specifically configured to, when the data to be stored is non-disk data, use hardware encryption to store the non-disk data in the dedicated memory space corresponding to the confidential container. When the data to be stored is intermediate result data generated by performing Spark Shuffle operation, the intermediate result data is stored in the dedicated NVMe hard disk corresponding to the confidential container using an encrypted method. When the data to be stored is the calculation result data generated after the task calculation is completed, the calculation result data is encrypted and then stored in the shared hard disk.
[0015] This application also provides an electronic device, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor is prompted by the machine-executable instructions to implement any of the container data storage methods described above.
[0016] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the container data storage methods described above.
[0017] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the container data storage methods described above.
[0018] The beneficial effects of this application include: The container data storage system provided in this application includes multiple confidential containers, each confidential container having its own dedicated memory space, a dedicated NVMe hard drive, and a shared hard drive on the physical host to which the multiple confidential containers belong. The dedicated memory space is used to store data from the corresponding confidential container using hardware encryption. The dedicated NVMe hard drive is directly connected to the corresponding confidential container using a device pass-through method and is used to store data from the corresponding confidential container using encryption. The shared hard drive is used to store ciphertext data from each confidential container. The ciphertext data generated by each confidential container is obtained by encrypting the data using the key generated by the confidential container. As can be seen, this container data storage system provides three levels of data storage methods. Level 1 is a dedicated memory space data storage method, which has the highest access performance and the smallest storage space. Level 2 is a dedicated NVMe hard drive data storage method, which has relatively high access performance and moderate storage space. Level 3 is a shared hard drive data storage method on a shared physical host, which has the worst access performance and the largest storage space. This allows for the selection of an appropriate data storage method from the three levels based on the actual data storage needs for data generated by confidential containers. While ensuring data security, this improves the flexibility and rationality of data storage, and further enhances data storage efficiency.
[0019] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0020] The accompanying drawings are provided to further illustrate the present application and form part of the specification. They are used together with the embodiments of the present application to explain the application and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the structure of a container data storage system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a container data storage system provided in another embodiment of this application; Figure 3 A flowchart illustrating the container data storage method provided in this application embodiment; Figure 4 A flowchart illustrating a container data storage method provided in another embodiment of this application; Figure 5 This is a schematic diagram of the structure of the container data storage device provided in the embodiments of this application; Figure 6This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0021] To provide a solution for improving data storage efficiency in confidential container application scenarios, this application provides a container data storage system, method, apparatus, and electronic device. The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features described in this application can be combined with each other unless otherwise specified.
[0022] This application provides a container data storage system, such as... Figure 1 As shown, it includes: multiple confidential containers, dedicated memory space for each confidential container, dedicated NVMe hard disk for each confidential container, and shared hard disk of the physical host to which multiple confidential containers belong; Dedicated memory space is used to store data from corresponding confidential containers using hardware encryption. The dedicated NVMe drive uses a device pass-through method to directly access the corresponding confidential container, which is used to store data from the corresponding confidential container in an encrypted manner; The shared hard drive is used to store encrypted data from various confidential containers. The encrypted data generated by each confidential container is obtained by encrypting the data using the key generated by that confidential container.
[0023] The container data storage system provided in this application embodiment offers three levels of data storage methods. Level 1 is a dedicated memory space data storage method, characterized by the highest access performance and the smallest storage space. Level 2 is a dedicated NVMe hard drive data storage method, characterized by relatively high access performance and moderate storage space. Level 3 is a shared hard drive data storage method on a shared physical host, characterized by the worst access performance and the largest storage space. This allows for the selection of an appropriate data storage method from the three levels based on the actual data storage needs for data to be stored generated by confidential containers. While ensuring data security, this improves the flexibility and rationality of data storage, further enhancing data storage efficiency.
[0024] In one embodiment of this application, such as Figure 2 As shown, a container data storage system may include multiple physical hosts, each physical host including multiple confidential containers. The data storage system also includes a container management module, which is connected to the multiple confidential containers on the multiple physical hosts and is used to control the operation of the multiple confidential containers.
[0025] from Figure 2As shown in the system architecture, the container data storage system includes multiple physical hosts, and each physical host includes multiple confidential containers. It can be understood as a confidential container cluster, that is, confidential containers are deployed in a cluster on multiple physical hosts. Correspondingly, the container management module can be a K8S container cluster management module, which is used to manage the multiple confidential containers in a cluster and control the operation of the multiple confidential containers.
[0026] In this embodiment of the application, the container data storage system has three levels of data storage methods. Level one is a data storage method with dedicated memory space. Figure 1 and Figure 2 As shown in the system architecture, each confidential container corresponds to a dedicated memory space, which is a block of physical address space in the physical host memory. This physical address space is allocated by the CPU security processor when the container starts. It features runtime memory encryption, whereby the container uses a hardware key (also known as a memory key) to encrypt the memory access range of the confidential container during runtime.
[0027] The memory key used to store data in dedicated memory space is a derived key generated by obfuscating the hardware root of trust with the container ID of the corresponding confidential container.
[0028] For dedicated memory space, a fixed amount of memory space (such as 128GB) needs to be set for the confidential container when it starts. Confidential containers cannot share the host machine's memory or the memory space of other containers. Confidential containers have space limitations in memory access.
[0029] Therefore, the data storage method with dedicated memory space represented by level one has the characteristics of the highest access performance and the smallest storage space.
[0030] Tier 2 is a dedicated NVMe hard drive data storage method, which exclusively uses the storage space of one or more NVMe hard drives on the physical host. The dedicated NVMe hard drive can use hardware-assisted virtualization technology and device pass-through to directly access the corresponding confidential container, which is exclusively used by the confidential container. Other containers and the host machine do not have access rights.
[0031] Dedicated NVMe SSDs use encrypted storage, which can be implemented using either hardware or software encryption. Hardware-assisted virtualization technology can accelerate CPU access to memory and NVMe SSD (a PCIe device) access. Specifically, it accelerates the two-level address translation between the CPU and physical memory (from the virtual machine's logical address to its physical address GVA->GPA, and from the virtual machine's physical address to the host's physical address GPA->HPA) using hardware. This acceleration is achieved through the IOMMU module, which performs hardware acceleration for the two-level address translation between NVMe SSD (a PCIe device) and memory access.
[0032] Therefore, the data storage method of dedicated NVMe hard drives represented by tier 2 has the characteristics of high access performance and moderate storage space.
[0033] The capacity of the dedicated memory space and dedicated NVMe hard drive storage space corresponding to the confidential container can be flexibly set based on the amount of data in the specific application scenario. No further examples or detailed descriptions will be provided here.
[0034] Level 3 is the data storage method of shared hard drives on shared physical hosts. This is actually to meet the need for persistent disk storage of confidential container data and to encrypt the disk data.
[0035] In this embodiment, the shared hard drive represented by level three can be other hard drive storage space on the physical host, such as a traditional mechanical hard drive or a high-speed solid-state drive. When the data generated by the confidential container needs to be stored on the shared hard drive, the confidential container first generates its own key. The key can be stored in the trusted execution environment in the confidential container, which cannot be obtained from the outside. The key is used to encrypt the data to obtain ciphertext data, and the ciphertext data is stored in the shared hard drive.
[0036] Level 3 represents the data storage method of shared hard drives on shared physical hosts, which has the worst access performance but the largest storage space.
[0037] Based on the characteristics of the above three-level data storage methods, in this embodiment, the application installed in the confidential container can, according to a preset data storage strategy, store the first type of data generated using hardware encryption in the dedicated memory space corresponding to the confidential container, store the second type of data generated using encryption in the dedicated NVMe hard drive corresponding to the confidential container, and store the third type of data generated in a shared hard drive after encryption. The data heat of the first, second, and third types of data decreases in that order.
[0038] The higher the data popularity, the more frequently the machine container reads and writes data during operation.
[0039] In this embodiment of the application, a three-level data storage structure is proposed to achieve a better balance between the security of the confidential container itself and the data processing performance of the confidential container in big data processing scenarios. This structure is used to store different types of data to meet the data storage needs of different types of data.
[0040] The encryption of memory access by confidential containers can effectively prevent malicious programs from stealing memory data. However, the encryption also means that the memory space on the host machine is divided into fixed intervals by each confidential container. Each confidential container can only use the address of this fixed interval and cannot share memory like traditional containers.
[0041] Once allocated at startup, this memory space cannot be changed during container operation. Because the number of memory slots on a single machine is limited, and the capacity of a single memory module is also limited (e.g., a maximum of 64GB), the memory space allocated to each confidential container is also very limited. For example, if a single physical host has a total memory space of 1TB and is running 10 confidential containers with the same amount of data, each confidential container will receive less than 100GB (some must be reserved for the host machine). Although memory has a high throughput (approximately 50GB / s), the limited memory space available to confidential containers significantly reduces their performance in scenarios with large data volumes.
[0042] To improve the performance of the confidential container, this application introduces a second layer, allowing one or more NVMe hard drives to be directly connected to the confidential container using a device pass-through method. Traditional virtual machines require two layers of address translation to access NVMe hard drives, which cannot fully utilize the performance of NVMe hard drives (NVMe hard drive performance is approximately 5GB / s).
[0043] In this embodiment of the application, it is proposed to pass through NVMe hard drives into a confidential container based on hardware-assisted virtualization technology. Passing through can realize hardware acceleration of the original two-layer address translation, thereby achieving near-native performance of NVMe hard drive access in the virtualization environment of the confidential container.
[0044] However, since the NVMe drive can only be used by one confidential container after passthrough and cannot be used by other containers, the capacity of the NVMe drive that can be passthrough is also limited (the capacity of an NVMe drive is usually 8TB / 15TB).
[0045] Furthermore, to address the limitations of storage space under large data volumes, this application proposes sharing the remaining disk space on the physical host among multiple containers, i.e., using it as a shared hard drive. However, to ensure data security, the confidential container encrypts the data that needs to be stored on the shared hard drive, and the encryption key is generated by the confidential container itself and can be stored in a trusted execution environment, making it inaccessible to outsiders.
[0046] Based on the aforementioned container data storage system, this application also provides a container data storage method, applied to applications installed in confidential containers, such as... Figure 3 As shown, it includes: Step 31: During the execution of data processing tasks, data to be stored is generated; Step 32: When the data to be stored is the first type of data, hardware encryption is used to store the first type of data in the dedicated memory space corresponding to the confidential container. When the data to be stored is the second type of data, an encryption method is used to store the second type of data in the dedicated NVMe hard drive corresponding to the confidential container. When the data to be stored is a third type of data, the third type of data is encrypted and then stored in the shared hard drive; The popularity of the first, second, and third types of data decreases in that order.
[0047] Using the container data storage method provided in this application embodiment, the data to be stored generated by applications in the confidential container can be stored in corresponding levels of data storage methods according to the different characteristics of data popularity. That is, the first type of data with the highest popularity is stored in dedicated memory space, the second type of data with the next highest popularity is stored in dedicated NVMe hard disk, and the third type of data with the lowest popularity is stored in shared hard disk of physical host. In this way, while ensuring data security, the flexibility and rationality of data storage are improved, the efficiency of data storage is further improved, and the overall performance of confidential container is improved.
[0048] The container data storage system and container data storage method provided in this application embodiment can be applied to various application scenarios involving big data processing. For example, it can be applied to the training and inference of AI large models. In this scenario, the application installed in the confidential container can be an AI large model application. The data generated by the AI large model application during the training and inference process of the AI large model can be stored using the corresponding level of data storage method based on the popularity.
[0049] It can also be applied to big data analysis scenarios, such as those based on the distributed computing engine Apache Spark, which will be described in detail below with reference to the accompanying diagram.
[0050] Figure 4 The container data storage method based on the distributed computing engine Apache Spark proposed in this application includes: Step 41: The user submits a Spark job through the client, that is, submits a data analysis request.
[0051] In this step, users can submit Spark jobs through the user program (spark-submit script) on the client.
[0052] Step 42: After receiving the data analysis request sent by the client, the Kubernetes Server Scheduler, the management scheduler of the server-side Kubernetes cluster, coordinates the container resources of the cluster, determines the confidential containers that will participate in executing the data analysis task corresponding to the data analysis request, and starts the Spark Driver.
[0053] Step 43: The server-side Spark Driver parses the data analysis request sent by the client.
[0054] In this step, Spark Driver can parse the user code carried in the data analysis request, such as the submitted Python file, to obtain the overall data analysis task to be executed.
[0055] Step 44: The Spark Driver divides the overall data analysis task into individual data analysis tasks and assigns them to the execution modules (Executors) of each confidential container involved in the execution.
[0056] In this step, Spark Driver can first construct a DAG (Directed Acyclic Graph) and then divide the data analysis tasks based on the constructed DAG.
[0057] Because data analysis involves a large amount of computation, it is necessary to use distributed computing to execute in parallel, thereby improving the efficiency of big data analysis.
[0058] Step 45: After receiving the data analysis task to be executed, the execution module in the confidential container can first read the encrypted data from HDFS (Hadoop Distributed File System) and then decrypt it after reading it into the confidential container.
[0059] In HDFS, data storage is performed on a shared hard drive at level 3.
[0060] If the data analysis task does not require reading encrypted data from HDSF beforehand, then this step is not necessary.
[0061] Step 46: The execution module in the confidential container begins to execute the assigned data analysis task.
[0062] Step 47: During the execution of this data analysis task, the execution module uses hardware encryption to store the non-disk data (equivalent to the first type of data) in the dedicated memory space corresponding to the confidential container.
[0063] Non-disk data stored in dedicated memory space can be read and used as needed, and will not be described in detail here.
[0064] Step 48: During the execution of this data analysis task, the execution module uses encryption to store the intermediate result data (equivalent to the second type of data) in the dedicated NVMe hard drive corresponding to the confidential container.
[0065] Intermediate result data stored on a dedicated NVMe hard drive can be read and used as needed.
[0066] A core operation of the Executor module during computation is Spark Shuffle. In this embodiment, the intermediate results generated by the Spark Shuffle operation are stored on a dedicated NVMe hard drive within a confidential container.
[0067] Based on hardware-assisted virtualization technology, NVMe disks are directly connected to the confidential container. Spark Shuffle is the core process of data redistribution in distributed computing, triggered when tasks need to be regrouped by key (such as reduceByKey, join). The capacity of dedicated memory space is insufficient to handle the intermediate data results generated by Spark Shuffle operations with large data volumes, which will require the intermediate results of Spark Shuffle operations to be persisted to disk.
[0068] The shuffle operation has two phases: Shuffle write and Shuffle read. Shuffle write is when the upstream task (such as the Map task) partitions the data and writes it to a temporary file, while Shuffle read is when the downstream task (such as the Reduce task) pulls the data from the corresponding partition and processes it.
[0069] Frequent Shuffle write and Shuffle read operations generate hot data, which is stored on a dedicated NVMe disk in a confidential container, thus improving performance during Spark computation.
[0070] Step 49: After the execution module completes the data analysis task, it encrypts the calculation results (equivalent to the third type of data) and stores them in the HDFS storage node (equivalent to a shared hard drive).
[0071] In this step, the calculation results obtained after completing the data analysis task are infrequently accessed data, i.e., cold data. Therefore, they can be stored in a shared hard drive with low performance and large capacity in a tier 3 representation.
[0072] Based on the same inventive concept, and according to the container data storage method provided in the above embodiments of this application, another embodiment of this application also provides a container data storage device based on any of the above container data storage systems, applied to an application installed in a confidential container, the structural schematic diagram of which is shown below. Figure 5 As shown, it specifically includes: The data generation module 51 is used to generate data to be stored during the execution of data processing tasks; The data storage module 52 is used to store the first type of data in the dedicated memory space corresponding to the confidential container when the data to be stored is the first type of data using hardware encryption. When the data to be stored is the second type of data, an encryption method is used to store the second type of data in the dedicated NVMe hard drive corresponding to the confidential container. When the data to be stored is a third type of data, the third type of data is encrypted and then stored in the shared hard drive; The popularity of the first type of data, the second type of data, and the third type of data decreases in that order.
[0073] Furthermore, the application installed in the confidential container is the Executor module in the distributed computing engine Apache Spark; The data generation module 51 is specifically used to generate data to be stored during the execution of big data analysis tasks; The data storage module 52 is specifically configured to, when the data to be stored is non-disk data, use hardware encryption to store the non-disk data in the dedicated memory space corresponding to the confidential container. When the data to be stored is intermediate result data generated by performing Spark Shuffle operation, the intermediate result data is stored in the dedicated NVMe hard disk corresponding to the confidential container using an encrypted method. When the data to be stored is the calculation result data generated after the task calculation is completed, the calculation result data is encrypted and then stored in the shared hard disk.
[0074] The functions of the above modules can be corresponding to Figures 1 to 4 The corresponding processing steps in the process shown will not be repeated here.
[0075] The container data storage device provided in the embodiments of this application can be implemented by a computer program. Those skilled in the art should understand that the above-described module division method is only one of many module division methods. Whether it is divided into other modules or not divided into modules, as long as the container data storage device has the above-described functions, it should be within the protection scope of this application.
[0076] This application also provides an electronic device, such as... Figure 6 As shown, it includes a processor 61 and a machine-readable storage medium 62, the machine-readable storage medium 62 storing machine-executable instructions that can be executed by the processor 61, the processor 61 being prompted by the machine-executable instructions to implement any of the container data storage methods described above.
[0077] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the container data storage methods described above.
[0078] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the container data storage methods described above.
[0079] The machine-readable storage medium in the aforementioned electronic device may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0080] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0081] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of devices, electronic devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0082] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0083] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0084] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0085] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0086] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A container data storage system, characterized in that, include: Multiple confidential containers, each confidential container having its own dedicated memory space, each confidential container having its own dedicated NVMe hard disk, and the physical host to which the multiple confidential containers belong having a shared hard disk; The dedicated memory space is used to store data from the corresponding confidential container using hardware encryption. The dedicated NVMe hard drive is directly connected to the corresponding confidential container via a device pass-through method, and is used to store data from the corresponding confidential container in an encrypted manner; The shared hard drive is used to store encrypted data from each of the confidential containers. The encrypted data generated by each confidential container is obtained by encrypting the data using the key generated by the confidential container.
2. The container data storage system as described in claim 1, characterized in that, The memory key used to store data in the dedicated memory space is a derived key generated by obfuscating the hardware root of trust with the container ID of the corresponding confidential container.
3. The container data storage system as described in claim 1, characterized in that, The dedicated NVMe hard drive uses hardware-assisted virtualization technology and a device pass-through method to directly access the corresponding confidential container.
4. The container data storage system as described in claim 1, characterized in that, The confidential container contains the application; The application is used to store the first type of data generated in the dedicated memory space corresponding to the confidential container using hardware encryption, according to a preset data storage strategy; to store the second type of data generated in the dedicated NVMe hard drive corresponding to the confidential container using encryption; and to store the third type of data generated in the shared hard drive after encryption, wherein the data heat of the first type of data, the second type of data and the third type of data decreases in that order.
5. The container data storage system as described in claim 1, characterized in that, It includes multiple physical hosts and a container management module, which is connected to the multiple confidential containers of the multiple physical hosts; The container management module is used to control the operation of the multiple confidential containers.
6. A container data storage method based on the container data storage system according to any one of claims 1-5, characterized in that, Applications installed in confidential containers include: During the execution of data processing tasks, data to be stored is generated; When the data to be stored is the first type of data, hardware encryption is used to store the first type of data in the dedicated memory space corresponding to the confidential container. When the data to be stored is the second type of data, an encryption method is used to store the second type of data in the dedicated NVMe hard drive corresponding to the confidential container. When the data to be stored is a third type of data, the third type of data is encrypted and then stored in the shared hard drive; The popularity of the first type of data, the second type of data, and the third type of data decreases in that order.
7. The method as described in claim 6, characterized in that, The application installed in the confidential container is the Executor module in the distributed computing engine Apache Spark; During the execution of data processing tasks, data to be stored is generated, including: During the execution of big data analysis tasks, data is generated that needs to be stored; When the data to be stored is a first type of data, hardware encryption is used to store the first type of data in the dedicated memory space corresponding to the confidential container, including: When the data to be stored is non-disk data, hardware encryption is used to store the non-disk data in the dedicated memory space corresponding to the confidential container. When the data to be stored is the second type of data, an encryption method is used to store the second type of data in the dedicated NVMe hard drive corresponding to the confidential container, including: When the data to be stored is intermediate result data generated by performing Spark Shuffle operation, the intermediate result data is stored in the dedicated NVMe hard disk corresponding to the confidential container using an encrypted method. When the data to be stored is a third type of data, encrypting the third type of data and storing it in the shared hard drive includes: When the data to be stored is the calculation result data generated after the task calculation is completed, the calculation result data is encrypted and then stored in the shared hard disk.
8. A container data storage device based on the container data storage system according to any one of claims 1-5, characterized in that, Applications installed in confidential containers include: The data generation module is used to generate data to be stored during the execution of data processing tasks; The data storage module is used to store the first type of data in the dedicated memory space corresponding to the confidential container when the data to be stored is the first type of data using hardware encryption. When the data to be stored is the second type of data, an encryption method is used to store the second type of data in the dedicated NVMe hard drive corresponding to the confidential container. When the data to be stored is a third type of data, the third type of data is encrypted and then stored in the shared hard drive; The popularity of the first type of data, the second type of data, and the third type of data decreases in that order.
9. An electronic device, characterized in that, The method includes a processor and a machine-readable storage medium storing machine-executable instructions that can be executed by the processor, the processor being prompted by the machine-executable instructions to perform the method of any one of claims 6-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 6-7.
Citation Information
Patent Citations
NVMe acceleration system, method and device and readable medium
CN115562574A
Data hierarchical storage method and device
CN116414296A
Method for storage, electronic device, program product
CN119652909A
Lightweight confidential container construction method based on hardware trusted environment
CN120216409A
Encrypted disk creation method and data persistence storage method and device
CN120449219A