Kernel mode parallel file system based on secure container
By containerizing the kernel-state parallel file system and combining the thermal upgrade mechanism and the fault fast recovery system, the problems of thermal upgrade difficulties and high resource consumption of traditional kernel-state parallel file systems are solved, and efficient and flexible infrastructure support is achieved to meet the needs of high-performance computing and artificial intelligence applications.
Patent Information
- Application Number
- CN202510504016.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional kernel-state parallel file systems have problems such as difficulty in thermal upgrade, large resource consumption, slow failure recovery speed, and difficulty in high-density deployment, which is difficult to meet the high requirements of high-performance computing and artificial intelligence applications.
Secure container technology is used to containerize the kernel-state parallel file system, combining thermal upgrade mechanism, rapid fault recovery system, resource consumption optimization and ecological compatibility design to achieve service isolation, rapid start-up, thermal upgrade, rapid fault recovery and simplicity.
It improves the high availability, scalability and operation and maintenance management efficiency of the system, reduces operation and maintenance costs, supports high-density deployment and seamless integration of cloud-native ecosystems.
Smart Images

Figure CN120371458A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer storage technology, and particularly relates to a kernel-mode parallel file system based on a secure container and its implementation method, which is applicable to the fields of artificial intelligence and high-performance computing. Background Art
[0002] With the booming development of artificial intelligence and Kubernetes (K8s) cloud-native technology, containerization and parallel file systems have become the core infrastructure to support artificial intelligence applications. However, existing parallel file systems (such as Lustre, BeeGFS, GPFS, and GlusterFS) are mainly designed based on the kernel mode and have the following limitations: 1. Difficult-to-fix BUGs and complex upgrade processes: BUGs in kernel-mode programs are usually difficult to fix, and the upgrade process is complex and time-consuming.
[0003] 2. Wide scope of fault impact: A fault in the file system service may cause a fault in the entire physical node, affecting all services running on that node.
[0004] 3. Complex operation and maintenance management: Traditional parallel file systems require dedicated storage head hardware, which has high hardware costs and complex operation and maintenance management.
[0005] 4. High resource consumption: Traditional virtual machines consume a large amount of resources and it is difficult to achieve high-density deployment.
[0006] 5. No high-availability solution: The file system depends on the high-availability solution of the underlying block storage and does not provide its own high-availability solution.
[0007] 6. Difficult hot upgrade: It is impossible to upgrade only the file system service. The upgrade of the entire node usually requires a downtime of more than several hours.
[0008] 7. Low deployment flexibility: File system services are often deployed on physical machines or separate virtual machines and cannot be easily deployed into a cloud-native environment.
[0009] Specifically, as Figures 1 - 4 shown, traditional kernel-mode parallel file systems mainly adopt the following two deployment schemes: I. Separate deployment of block storage and file system In this deployment scheme, file systems such as GPFS, Lustre, and BeeGFS are separated from block storage. The block storage adopts technologies such as dual-controller arrays to be responsible for providing the high availability and reliability of persistent storage. The file system service is deployed on the head of the storage array or an independent server and can run directly in a physical machine or a virtual machine.
[0010] (1) Physical machine deployment When the file system service is directly deployed on a physical machine, in case of software BUGs or the need for upgrade operations, hot upgrade becomes very difficult. Usually, the entire physical node needs to be restarted to complete the upgrade, resulting in slow service recovery. In addition, the failure of the file system service may trigger the failure of the entire physical node, thereby affecting all other services running on that node.
[0011] (2) Virtual machine deployment Although virtual machines can be managed through virtualization technologies such as qemu and kubevirt, which improves service isolation, fault control, deployment flexibility, and maintenance convenience, traditional virtual machines have the following problems: High resource consumption: Virtual machines require additional resource overhead for operation, reducing the overall resource utilization of the system.
[0012] Slow fault recovery speed: In case of failures, the recovery time of virtual machines is relatively long.
[0013] Difficult to achieve high-density deployment: Due to high resource consumption, it is difficult to deploy a large number of virtual machine instances on a single physical node, affecting the scalability of the system.
[0014] II. File system service directly deployed on a general server In this deployment solution, file system services such as GPFS and GlusterFS directly run on general servers, providing high availability and reliability of persistent storage. However, this deployment method also faces the following problems: (1) Physical machine deployment Similar to the physical machine deployment situation with separate block storage deployment, when directly deploying the file system service on a general server, hot upgrade also faces huge challenges. When BUGs occur or upgrades are needed, usually the entire physical node needs to be restarted, resulting in a slow service recovery process. In addition, the failure of the file system service may trigger the failure of the entire physical node, thereby affecting all other services running on that node.
[0015] In summary, the deployment solutions of traditional kernel-mode parallel file systems have obvious limitations, including difficult hot upgrade, high resource consumption, slow fault recovery speed, and difficult high-density deployment. These problems to a certain extent restrict the high availability, scalability, and operation and maintenance management efficiency of the system, and it is difficult to meet the high requirements of modern high-performance computing and artificial intelligence applications for storage infrastructure. Summary of the Invention
[0016] To address the above challenges, this application proposes a new kernel-mode parallel file system based on secure containers.
[0017] The purpose of this application is to provide a kernel - mode parallel file system based on secure containers and its implementation method, so as to overcome the limitations of traditional kernel - mode parallel file systems and improve the high availability, scalability and operation and maintenance management efficiency of the system.
[0018] This solution integrates the kernel - mode parallel file system with secure container technology. This integration not only provides key advantages such as service isolation, fast startup, ecological compatibility, hot upgrade, fast fault recovery, support for multiple deployment architectures, and ease of maintenance, but also significantly improves the availability, performance and usability of the kernel - mode parallel file system in intelligent computing scenarios. In addition, this solution can effectively reduce the complexity and cost of operation and maintenance, and provide a more powerful, flexible and cost - effective infrastructure solution for the fields of artificial intelligence and high - performance computing.
[0019] The main technical strategies adopted in this solution are as follows: 1. Secure containerized parallel file system: Develop a kernel - mode parallel file system and deploy it through secure containerization technology. This technology can achieve service isolation and fast startup while maintaining compatibility and performance with traditional kernel - mode file systems. By containerizing the file system service, this solution solves the extensive impact of service failures on physical nodes in traditional deployment methods and improves the stability and availability of the system.
[0020] 2. Hot upgrade mechanism Design a hot upgrade mechanism that allows the kernel - mode parallel file system to be upgraded without restarting the physical node. This mechanism significantly shortens the upgrade recovery time and reduces the system downtime caused by upgrades, thus greatly improving the availability and operation and maintenance efficiency of the system.
[0021] 3. Fast fault recovery system Build a fast fault recovery system that can quickly migrate and restart services through cloud - native tools such as Kubernetes when the kernel - mode parallel file system service fails, reducing system downtime. This system utilizes the elastic scheduling ability of cloud - native technology to reduce the impact of faults on services and further improve the high availability of the system.
[0022] 4. Resource consumption optimization Implement a resource consumption optimization mechanism to reduce the resource occupancy of traditional virtual machines through lightweight containerized deployment. This mechanism supports high - density deployment, improves resource utilization, and solves the problems of large resource consumption and low deployment efficiency of traditional virtual machines.
[0023] 5. Ease of maintenance Provide a solution for easy maintenance, which simplifies the maintenance and upgrade processes of the kernel-mode parallel file system through containerization technology. This solution utilizes the tools and processes of the cloud-native ecosystem, reducing the complexity of operation and maintenance management, improving operation and maintenance efficiency, and reducing operation and maintenance costs.
[0024] 6. Ecological Compatibility Ensure seamless integration of the kernel-mode parallel file system with the existing cloud-native ecosystem. Through compatibility design with cloud-native tools such as Kubernetes, this solution can utilize existing tools and processes for management and optimization, further enhancing the usability and scalability of the system.
[0025] To achieve the above objectives, this application provides the following technical solutions: The first aspect of this application provides a kernel-mode parallel file system based on secure containers. The parallel file system includes: A computing service cluster, including multiple client nodes and at least one master node; A Lustre service cluster, which consists of the following service components: A VIP mapping service component, used to map a virtual IP (VIP) to a physical server IP address; A cluster planning service component, used for service deployment and resource allocation; A high-availability architecture component, including a master-slave service architecture, used to achieve the high availability of the Lustre service; A VIP server component, used to simplify the high-availability configuration of Lustre and provide a translation service from VIP to the IP address of the backend server; A block storage cluster, containing multiple object storage targets, and each object storage target is associated with a storage engine, used to provide data storage services.
[0026] Furthermore, in the parallel file system of this application, the master service architecture in the high-availability architecture component is responsible for detecting the server status, migrating the containers of the downed nodes to a new server, updating the management server configuration at the same time, and notifying the client to access the new server.
[0027] Furthermore, in the parallel file system of this application, the master service architecture realizes the hot upgrade and failover recovery of KataContainer containers through Kubernetes; Each container in the Kata Container runs in an independent virtual machine to achieve secure isolation; the Kata Container supports the OCI (Open Container Initiative) standard and the CRI (Container Runtime Interface) interface, and is seamlessly integrated into the Kubernetes (K8s) platform to achieve ecological compatibility.
[0028] Furthermore, in the parallel file system of the present application, when the IP address of the VIP server component changes, the client can access the service without changing the configuration.
[0029] Furthermore, in the parallel file system of the present application, when a node in the block storage cluster fails, the main service architecture is used to re-obtain the data shard location and reconstruct the data on a new node.
[0030] Furthermore, the parallel file system of the present application further includes a data redundancy and automatic data recovery component for automatically recovering data when a storage node fails.
[0031] Furthermore, the parallel file system of the present application implements "near-memory computing" through Kubernetes and supports strong tenant isolation.
[0032] Furthermore, the cluster and computing nodes of the parallel file system of the present application are separately deployed.
[0033] Furthermore, in the fault recovery process of the parallel file system of the present application, when the master node detects a failure of the management server or the object storage server, it will be scheduled to a new server and the VIP mapping table will be updated for other services to access.
[0034] The second aspect of the present application provides an electronic device, including: a memory and a processor; Memory: for storing computer programs; Processor: for executing the computer program to implement the functions of the aforementioned kernel-mode parallel file system based on secure containers.
[0035] In summary, by introducing secure containerization technology, a hot upgrade mechanism, a fast fault recovery system, a resource consumption optimization mechanism, a maintenance simplicity solution, and an ecological compatibility design, the present application significantly improves the availability, performance, and ease of use of the kernel-mode parallel file system in artificial intelligence and high-performance computing scenarios, providing users with a more efficient, reliable, and easy-to-manage infrastructure support. This solution has the following advantages: (1) Improve the high availability and scalability of the system; (2) Shorten the upgrade and recovery time and reduce the system downtime; (3) Reduce resource consumption and support high-density deployment; (4) Simplify the operation and maintenance management process and reduce the operation and maintenance cost; (5) Seamlessly integrate with the existing cloud-native ecosystem and improve the ease of use.
[0036] Other features and advantages of the present application will be elaborated in detail in the subsequent description, or can be understood by implementing the relevant technical solutions of the present application. The objectives and other advantages of the present application can be achieved through the technical features and means clearly pointed out in the description, claims, and drawings, and can be obtained through the implementation process of these technical contents. Brief Description of the Drawings
[0037] To more clearly elaborate the background technology and technical solutions of the embodiments of the present application, the drawings involved in the description of the background technology and embodiments will be briefly introduced below. It should be noted that the drawings only show some embodiments of the present application. For those skilled in the art, other relevant drawings can be derived based on these drawings without creative efforts.
[0038] Figure 1 It is a schematic diagram of the architecture of a typical intelligent computing parallel file storage solution.
[0039] Figure 2 is a schematic diagram of a typical Lustre architecture.
[0040] Figure 3 is a schematic diagram of a typical Lustre HA solution.
[0041] Figure 4 is a schematic diagram of the complexity of typical Lustre hot upgrade and service failure recovery.
[0042] Figure 5 is a schematic diagram of the high-availability architecture of Lustre and Kata security containers of the present application.
[0043] Figure 6 is a schematic diagram of the internal architecture of Lustre and Kata security containers of the present application.
[0044] Figure 7 is an example diagram of the high-availability architecture of Lustre and Kata security containers of the present application.
[0045] Figure 8 is a schematic diagram of the highly reliable architecture of Lustre and Kata security containers of the present application.
[0046] Figure 9 is an example diagram of the highly reliable architecture of Lustre and Kata security containers of the present application.
[0047] Figure 10 is a schematic diagram of the architecture integrating the parallel file system storage of the present application into the intelligent computing platform.
[0048] Figure 11 It is the overall design architecture diagram of the kernel-mode parallel file system based on security containers of the present application.
[0049] Figure 12 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed Embodiments
[0050] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer and more understandable, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. It should be clear that the described embodiments are only part of the embodiments of this application, not all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the protection scope of this application.
[0051] In this document, the term "including" and any form of its deformation (such as "including", "including") are open expressions and should be understood as "including but not limited to", that is, the listed content is not an exhaustive list and may also include other content not explicitly mentioned. The term "based on" should be understood as "at least partially based on", that is, the referred basis or condition may not be the only factor and may also involve other relevant factors. The term "an embodiment" should be understood as "at least one embodiment", that is, the described embodiment is not the only possible implementation, and there may be other similar embodiments.
[0052] In this application, when the terms "a" and "multiple" are used to modify related elements or features, their expressions are illustrative rather than restrictive. Unless otherwise clearly stated in the context, "a" should be understood as "at least one", and "multiple" should be understood as "at least two". Those skilled in the art should reasonably interpret these terms according to the semantic and logical relationships in the context to ensure that they cover the possibility of "one or more".
[0053] Figure 11 The overall design architecture of the kernel-mode parallel file system based on a secure container provided by this application is shown, including: A computing service cluster, including multiple client nodes and at least one master node; A Lustre service cluster, which consists of the following service components: A VIP mapping service component, used to map a virtual IP (VIP) to a physical server IP address; A cluster planning service component, used for service deployment and resource allocation; A high-availability architecture component, including a master-slave service architecture, used to achieve the high availability of the Lustre service; A VIP server component, used to simplify the high-availability configuration of Lustre and provide a translation service from VIP to the IP address of the backend server; A block storage cluster, containing multiple object storage targets, and each object storage target is associated with a storage engine, used to provide data storage services.
[0054] To more clearly elaborate on the technical solution of this application, the following will be further described through embodiments of specific scenarios.
[0055] Figure 5 The following shows a highly available architecture of a kernel-level parallel file system (Lustre service cluster) based on secure containers proposed in this application, including: 1. Computing service cluster The computing service cluster is shown at the top of the picture, which includes multiple Client nodes and one Master node.
[0056] 2. Lustre service cluster, which is composed of multiple service components, including: VIP Mapping: used to map the virtual IP (VIP) to the physical server IP address; Cluster Plan: cluster planning, which involves service deployment and resource allocation strategies.
[0057] Analysis of the highly available architecture Master-Slave architecture: to achieve the high availability of the Lustre service.
[0058] Master service: responsible for detecting the server status and migrating the Kata Container of the downed node to a new server. In addition, the Master service is also responsible for updating the MGS (management server) configuration and notifying the client Client to access the new server.
[0059] VIP Server: simplifies the high availability configuration of Lustre.
[0060] VIP service: provides translation service from VIP to the IP address of the backend server.
[0061] Client side: only the IP of the MGS server group is visible. When the Client needs to communicate with the MGS, it queries the VIP Server to obtain the IP address of the MGS. When the MGS IP address changes, the Client is not aware and does not need to change the configuration.
[0062] The following service components are included in this architecture: MGS: management server, responsible for managing the metadata of the file system.
[0063] MDT: metadata target, storing the metadata of the file system.
[0064] OST: object storage target, storing the actual data of the file system.
[0065] Client Agent: The client agent is responsible for communicating with the file system.
[0066] The fault recovery process of this architecture is as follows: 1. When the Master detects a failure of the MGS, it will schedule the MGS to a new server and update the VIP mapping table for other services to access.
[0067] 2. When the Master detects a failure of the OSS (Object Storage Server), it will schedule the OST to a new server, and the OST will resume providing services.
[0068] 3. In the fault recovery of the block storage cluster nodes, after the Master detects a failure of a cluster node, the block storage client will obtain the data shard location from the Master again and access the data shards again, and then the block storage cluster will reconstruct the data.
[0069] Figure 6 The hot upgrade and fault switch recovery architecture adopted in this application is shown, which details the architecture features, hot upgrade process, and fault switch recovery mechanism of Kata Containers 3.0.
[0070] The features of Kata Container are as follows: Security isolation: Each container runs in an independent virtual machine, providing stronger security and isolation guarantees.
[0071] Performance: By optimizing the design, the performance overhead of virtualization is reduced, making the performance close to that of standard containers.
[0072] Lightweight: It achieves a startup time of hundreds of milliseconds and has low resource occupancy.
[0073] Ecosystem compatibility: It supports the OCI (Open Container Initiative) standard and CRI (Container Runtime Interface) interface, and is seamlessly integrated into the Kubernetes (K8s) platform.
[0074] Multi-architecture and virtualization technologies: It supports multiple CPU architectures and virtualization solutions.
[0075] Open source community: It has become the standard security container solution for top cloud providers.
[0076] Analysis of the hot upgrade and fault switch recovery architecture: 1. Hot upgrade Using Kubernetes (K8s) to achieve gray-scale and rolling upgrades greatly reduces the operation and maintenance difficulty.
[0077] 2. Second-level fault switch recovery When the Master node detects a failure of the Kata Container, Kubernetes quickly schedules the Kata Container to a new node, and the failover recovery time is close to the performance of standard containers.
[0078] Figure 7 The figure shows the high-availability architecture of the Lustre service cluster in a computing service cluster, including the interaction process between the client, the Lustre service cluster, and the block storage cluster. The high-availability processing flow of the Lustre service cluster in the face of server failures is described in detail, including key steps such as failure detection, server migration, cluster replanning, and VIP mapping update.
[0079] This architecture includes: 1. Computing service cluster It contains multiple clients that interact with the Lustre service cluster.
[0080] 2. Lustre service cluster It consists of a master server and slave agents, and the slave agents run in Kata containers.
[0081] The main components include: VIP Server: Responsible for managing virtual IP (VIP) mapping.
[0082] MGS / MDS Server: Management server (metadata server), running in a Kata container but currently in a failed state (indicated by a red cross in the figure).
[0083] OSS Server: Object storage server, also running in a Kata container.
[0084] Client: The client that interacts with the Lustre service cluster.
[0085] 3. Block storage cluster It contains multiple object storage targets (storage nodes, OSTs), and each OST is associated with a Rock solid engine (storage engine) to provide data storage services.
[0086] High-availability process: 1. Failure detection and migration: The master server detects the failure of the MGS / OSS server.
[0087] The Master migrates the Kata containers on the faulty server to a new physical node (the Master schedules the MGS / OST to the new server).
[0088] 2. Cluster re-planning: The Master performs cluster re-planning and schedules the OSS3 server to take over the responsibilities of the OST2.
[0089] 3. VIP mapping update: If necessary, the VIP Server updates the VIP mapping to ensure that the client can access the new server instance.
[0090] Figure 8 The figure shows a highly reliable architecture for Lustre and Kata secure containers proposed in this application, including the interaction process between the computing service cluster, the Lustre service cluster, and the block storage cluster, as well as the data redundancy and automatic data recovery mechanisms.
[0091] Analysis of the data highly reliable architecture: 1. Server availability detection The Master service: responsible for detecting the server status.
[0092] When the server is unavailable, the Master service will schedule a new server to replace it, and the data will be automatically recovered by other backup nodes.
[0093] The CeaStor Block Client automatically switches to other backup nodes for reading and writing operations.
[0094] 2. Bad disk detection The server monitors the disk status. When a bad disk is found, the server automatically isolates the bad disk.
[0095] The Master service schedules a new server to replace it, and the data will be automatically recovered by other backup nodes.
[0096] The CeaStor Block Client automatically switches to the backup node of the data for reading and writing.
[0097] Figure 9 The figure shows an example of a highly reliable architecture for Lustre and Kata secure containers proposed in this application, which details how the computing service cluster interacts with the block storage cluster through the Lustre service cluster and how to achieve automatic data recovery in case of failures. This architecture includes: 1. Computing service cluster Contains multiple clients, which interact with the Lustre service cluster.
[0098] 2. Lustre service cluster It includes the following components: MGS-MDS Server On VM: The metadata server (management server) running on the virtual machine is used to process client requests and manage the metadata of the file system.
[0099] OST Server On VM: The object storage server running on the virtual machine is used to process data storage requests and interact with the CeaStor Block Client.
[0100] CeaStor Block Client: The block storage client communicates with the block storage cluster to read and write data.
[0101] 3. Block Storage Cluster It includes multiple storage nodes, and the RockEngine (or storage management software) runs on each node.
[0102] The storage nodes are distinguished by digital identifiers (1, 2, 3), and there are multiple replicas (represented by different colors) to achieve data redundancy.
[0103] Technical Process: Data Redundancy: Data is represented by squares of different colors on multiple storage nodes, showing the redundant storage of data.
[0104] Block Storage Cluster Node Failure Recovery: When a storage node fails (represented by a red cross), the system performs failure recovery through the following steps: Step 1: The Master detects a cluster node failure.
[0105] Step 2: The Master migrates the data of the failed node to other nodes, and the CeaStor Block Client updates the data access path.
[0106] Step 3: The block storage client re-obtains the data shard location from the Master, and the block storage client re-accesses the data shards.
[0107] Step 4: The block storage cluster reconstructs the data on the new node to ensure data redundancy and availability.
[0108] Figure 10 The figure shows the integration of the parallel file system storage of this application into the intelligent computing platform architecture. This architecture realizes "near-memory computing" through Kubernetes (K8s), and emphasizes strong tenant isolation, performance, security, and availability.
[0109] The picture shows three virtual private cloud (VPC) environments: VPC1, VPC2, and VPC3.
[0110] Each VPC contains computing nodes (represented by green boxes in the figure), which run on Kubernetes (K8s).
[0111] The parallel file system cluster is deployed into the computing nodes of each VPC through K8s.
[0112] The parallel file system cluster consists of multiple parallel file system instances, each running on an independent computing node. The cluster is managed by CeaStor, and the CeaStor cluster is responsible for the management of the file system and data storage.
[0113] The technical advantages of this architecture are as follows: In-memory computing: By deploying the file system cluster into the computing nodes, tight integration of data storage and computing is achieved, reducing data transfer latency and improving computing efficiency.
[0114] Strong tenant isolation: The file system clusters within each VPC run independently, ensuring data isolation between different tenants and enhancing security.
[0115] Performance, security, and availability: The cluster design supports high-performance computing, data security protection, and high availability, meeting the requirements of enterprise-level applications.
[0116] Separate deployment: The parallel file system cluster and computing nodes are separately deployed, providing a more flexible architecture design and reducing deployment and operation and maintenance costs at the same time.
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementations of the devices, methods, and computer program products according to various embodiments of the present application, including the architecture, functions, and operations. In these figures, each block may represent a module, a program segment, or a part of the code, which contains one or more executable instructions for implementing the specified logical function. It should be noted that each block in the block diagram and / or flowchart, as well as the combinations of these blocks, can be implemented by a dedicated hardware-based system to implement the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0118] As Figure 12 shown, an embodiment of the present application also discloses an electronic device, including: a processor 310, a communication interface 320, a memory 330 for storing computer programs executable by the processor, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 runs the executable computer program to implement the functions of the above-mentioned kernel-mode parallel file system based on secure containers.
[0119] It can be understood that, in addition to including a memory and a processor, the electronic device may further include an input device (such as a keyboard), an output device (such as a display), and other communication modules. These input devices, output devices, and other communication modules all communicate with the processor through an I / O interface (i.e., an input / output interface).
[0120] The operations of the present application can be implemented by writing computer program code using one or more programming languages or combinations thereof. The programming languages include, but are not limited to, the following types: Object-oriented programming languages, such as Java, Smalltalk, C++, etc.; Conventional procedural programming languages, such as the "C" language or similar programming languages.
[0121] The execution modes of the program code include, but are not limited to: Executed entirely on the user's computer; Partially executed on the user's computer and partially executed on a remote computer; Executed as an independent software package; Executed entirely on a remote computer or server.
[0122] In scenarios involving a remote computer, the remote computer can be connected to the user's computer through any type of network connection. The network includes, but is not limited to, a local area network (LAN) or a wide area network (WAN). In addition, the remote computer can also be connected to an external computer through an Internet service provider, for example, by using the Internet for connection.
[0123] Furthermore, the present application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can implement the various functions of the kernel-state parallel file system based on a secure container disclosed in the present application.
[0124] In the context of the present application, a computer-readable storage medium refers to a tangible medium that can store computer program code and related data. Specific examples include, but are not limited to, the following: (1) Portable computer disks: Removable magnetic storage media such as floppy disks.
[0125] (2) Hard disks: Fixed storage devices including mechanical hard disks and solid-state drives, etc.
[0126] (3) Random access memory (RAM): Volatile storage media for temporarily storing data and program code.
[0127] (4) Read-only memory (ROM): Non-volatile storage media for storing fixed programs and data.
[0128] (5) Erasable Programmable Read-Only Memory (EPROM) or Flash Memory: A non-volatile storage medium that supports multiple erasures and programming.
[0129] (6) Fiber optic storage device: A storage medium based on fiber optic technology.
[0130] (7) Portable Compact Disc Read-Only Memory (CD-ROM): A read-only medium that stores data in the form of optical discs.
[0131] (8) Optical storage device: Storage media based on optical principles such as DVDs, Blu-ray discs, etc.
[0132] (9) Magnetic storage device: Storage media based on magnetic principles such as magnetic tapes, magnetic disks, etc.
[0133] (10) Any suitable combination of the above: For example, combining multiple storage media to meet different storage requirements.
[0134] These computer-readable storage media can be used to store the program code and related data described in this application to support the operation of the program and the persistent storage of data.
[0135] In particular, according to the embodiments of this application, the processes described in the flowcharts can be implemented as computer software programs. For example, the embodiments of this application relate to a computer program product that includes a computer program carried on a non-transitory computer-readable medium. This computer program contains program code for executing the kernel-mode parallel file system based on a secure container disclosed in this application. When this computer program is executed by a processing device, it can implement the above functions defined in the embodiments of this application.
[0136] Although there are several specific implementation details in the above discussion, these details should not be construed as limitations on the scope of this application. The above description is only for the preferred embodiments of this application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combinations of the above technical features. At the same time, this application should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept.
[0137] Those skilled in the art should also understand that they can modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features, without departing from the spirit and scope of the technical solutions of the embodiments of this application. These modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the core spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A kernel-mode parallel file system based on secure containers, characterized in that The parallel file system includes: A computing service cluster, including multiple client nodes and at least one master node; A Lustre service cluster, which consists of the following service components: A VIP mapping service component, used to map virtual IPs to physical server IP addresses; A cluster planning service component, used for service deployment and resource allocation; A high-availability architecture component, including a master-slave service architecture, used to achieve the high availability of Lustre services; A VIP server component, used to simplify the high-availability configuration of Lustre and provide translation services from VIPs to backend server IP addresses; A block storage cluster, containing multiple object storage targets, and each object storage target is associated with a storage engine, used to provide data storage services.
2. The parallel file system according to claim 1, wherein The master service architecture in the high-availability architecture component is responsible for detecting the server status, migrating the containers of the downed nodes to new servers, updating the management server configuration at the same time, and notifying the clients to access the new servers.
3. The parallel file system according to claim 1, wherein The master service architecture realizes the hot upgrade and failover recovery of Kata Container containers through Kubernetes; Each container in the Kata Container runs in an independent virtual machine to achieve secure isolation; the Kata Container supports the OCI standard and CRI interface, and is seamlessly integrated into the Kubernetes platform to achieve ecological compatibility.
4. The parallel file system according to claim 1, wherein When the IP address of the management server changes in the VIP server component, the clients can access the service without changing the configuration.
5. The parallel file system according to claim 1, wherein When a node fails in the block storage cluster, the master service architecture is used to re-obtain the data shard location and reconstruct the data on the new node.
6. The parallel file system according to claim 1, characterized in that The parallel file system also includes data redundancy and automatic data recovery components, used to automatically recover data when a storage node fails.
7. The parallel file system according to claim 1, characterized in that, The parallel file system realizes "in-memory computing" through Kubernetes and supports strong tenant isolation.
8. The parallel file system according to claim 1, characterized in that, The clusters and computing nodes of the parallel file system are separately deployed.
9. The parallel file system according to claim 1, wherein In the fault recovery process of the parallel file system, when the master node identifies a failure of the management server or the object storage server, it will be scheduled to a new server and the VIP mapping table will be updated for other services to access.
10. An electronic device, characterized in that, Including: A memory and a processor; Memory: used to store computer programs; Processor: used to execute the computer programs to implement the functions of the kernel-mode parallel file system based on secure containers as described in any one of claims 1-9.
Citation Information
Cited By
Cluster system based on domestic VPX server and method for attaching VIP to application service
CN120639771A