Cluster task container management method and system

Through a lightweight state database and port pool mechanism, combined with a double-layer SSH springboard mechanism, the port conflict and state tracking problems in traditional container management are solved, efficient task status recording and automatic port scheduling and recovery are achieved, and the stability of the cluster environment and user experience are improved.

CN120704796AActive Publication Date: 2025-09-26COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
6 Cites 0 Cited by

Patent Information

Application Number
CN202510778288.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-26
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

In a cluster environment with multi-user, concurrent, and high-frequency job scheduling, traditional container management methods have problems such as serious port conflicts, lack of task status tracking mechanism, complex container access path configuration, and lack of failure recovery mechanism.

Method used

A lightweight state database is used to record the status of the entire life cycle of the task, a bitmap-based port pool mechanism is designed, and the concurrent uniqueness of port applications is ensured through memory locks. A double-layer SSH springboard mechanism is used to establish a secure communication channel, implement container lifecycle management, and automatically trigger status checks and tunnel reconstruction in the event of an exception.

Benefits of technology

It achieves persistent recording of task-level status, automatic port scheduling and recovery, improves platform stability, user experience and resource utilization efficiency, and supports safe and reliable container management in a multi-user concurrent environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704796A_ABST
    Figure CN120704796A_ABST
Patent Text Reader

Abstract

A cluster task container management method is applied to a super-computing cluster, the super-computing cluster comprises a management node and a plurality of computing nodes, and the method comprises the steps that a lightweight state database is constructed, the full life cycle state of a task is recorded, and the life cycle comprises submitting, scheduling, container starting, port mapping and recycling stages; designing a port pool mechanism based on a bitmap, generating an available port sequence according to a fixed step length, and ensuring the concurrency uniqueness of port application through a memory lock; establishing an SSH tunnel at the management node, and forwarding the local port to the SSH port of the target computing node; container life cycle management is implemented, tunnel states are scanned regularly, and container state inspection and tunnel reconstruction are automatically triggered when abnormity occurs; all state events are written into a lightweight state database through an API, and the lightweight state database supports task query and connection health check. According to the method, persistent recording of task level states and automatic scheduling and recovery of ports can be supported, and the platform stability and the resource utilization efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high performance computing (HPC) technology, and in particular to a cluster task container management method and system. Background Art

[0002] With the rapid development of high-performance computing (HPC) and artificial intelligence platforms, container technologies (such as Singularity) are widely used in scientific computing environments to isolate software dependencies and facilitate task deployment. However, in cluster environments with multiple users, concurrent workloads, and high-frequency job scheduling, traditional container management methods face the following major issues:

[0003] Serious port conflicts: When multiple container tasks concurrently request host ports for remote communication (such as SSH tunnel forwarding), port resources cannot be allocated efficiently, which can easily lead to conflicts or port leaks.

[0004] Lack of task status tracking mechanism: Traditional script-based container scheduling methods cannot persistently record task status and lifecycle events, making subsequent visualization, diagnosis, and recovery impossible.

[0005] Complex container access path configuration: Users need to manually configure port forwarding, local mapping, jump server connections, and other steps, which has a high barrier to entry and a high error rate.

[0006] Lack of failure recovery mechanism: Once the container exits abnormally or the port mapping is disconnected, the system is often unable to detect and automatically repair it, seriously affecting task continuity and platform availability. Summary of the Invention

[0007] In order to solve the problems existing in the prior art, the embodiments of the present application provide a method, system, computing device, computer storage medium and product containing a computer program for cluster task container management, which can support persistent recording of task-level status, automatic port scheduling and recovery, and improve platform stability, user experience and resource utilization efficiency.

[0008] In the first aspect, an embodiment of the present application provides a cluster task container management method, which is applied to a supercomputing cluster, wherein the supercomputing cluster includes a management node and multiple computing nodes, including: constructing a lightweight state database to record the status of the entire life cycle of the task, and the life cycle includes submission, scheduling, container startup, port mapping and recycling stages; designing a bitmap-based port pool mechanism to generate a discretely distributed sequence of available ports according to a fixed step size, and ensure the concurrent uniqueness of port applications through memory locks; establishing an SSH tunnel in the management node to forward the local port to the SSH port of the target computing node; implementing container lifecycle management, regularly scanning the tunnel status, and automatically triggering container status checks and tunnel reconstruction when an abnormality occurs; writing all status events to the lightweight state database through an API, and the lightweight state database supports task queries and connection health checks.

[0009] In some possible implementations, the lightweight status database is implemented using PyDbLite, the task record uses the task's uuid as the primary key, and clusterName_user_jobId as the business index, and includes user information, node information, job status, container status, and port mapping status fields.

[0010] In some possible implementations, in the port pool mechanism, the memory lock is generated by an MD5 hash algorithm, and the hash input includes the cluster name, user name, job ID, node IP, and port.

[0011] In some possible implementations, the SSH tunnel is established through a double-layer springboard, including: connecting to the cluster login node through the Paramiko library; transparently transmitting from the login node to the target computing node, using Ed25519 key authentication and the open_channel method to create an encrypted channel.

[0012] In some possible implementations, the lifecycle management of containers is based on the Singularity runtime, and the container instance naming convention is service_auto_ <baseport>, dynamically inject the overlay file system, mount path and environment variables at startup.

[0013] In some possible implementations, container status check and tunnel reconstruction are automatically triggered when an exception occurs, including: when an SSH tunnel interruption is detected, querying the lightweight status database to obtain the status of the associated container; if the container is alive, automatically rebuilding the tunnel; if the container fails, marking the exception and releasing port resources.

[0014] In some possible implementations, the method is compatible with the Slurm or Kubernetes scheduling system, and the lightweight state database is deployed as an independent module without modifying the original scheduling logic of Slurm or Kubernetes.

[0015] In the second aspect, an embodiment of the present application provides a cluster task container management system, which is deployed in a supercomputing cluster, and the supercomputing cluster includes a management node and multiple computing nodes, including: a lightweight state database module, which is used to record the status of the entire life cycle of the task, including status information of the submission, scheduling, container startup, port mapping and recycling stages; a remote springboard connection module, which is used to establish a secure communication channel with the computing node through a double-layer SSH springboard mechanism; a container lifecycle control module, which is used to implement lifecycle management of container startup, stop, and status monitoring based on the Singularity runtime; a port scheduling and forwarding module, which is used to allocate ports through a bitmap mechanism and establish an SSH tunnel to realize local access forwarding to the container; a port lifecycle monitoring and abnormal recovery mechanism module, which is used to periodically detect the port and tunnel status and automatically trigger the recovery process in case of an abnormality.

[0016] In a third aspect, an embodiment of the present application provides a computer-readable storage medium comprising computer-readable instructions. When a computer reads and executes the computer-readable instructions, the computer executes the method as described in any one of the first aspects.

[0017] In a fourth aspect, an embodiment of the present application provides a computing device comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the method as described in any one of the first aspects is executed.

[0018] In a fifth aspect, an embodiment of the present application provides a product comprising a computer program, which, when the computer program product runs on a processor, enables the processor to execute the method as described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 This is a schematic diagram of a cluster task container management architecture provided by an embodiment of the present application;

[0021] Figure 2 This is a flow chart of a cluster task container management method provided by an embodiment of the present application;

[0022] Figure 3 This is a schematic diagram of the overall process of cluster task execution provided by an embodiment of the present application. DETAILED DESCRIPTION

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0024] The term "and / or" in this document describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " in this document indicates that the related objects are in an "or" relationship. For example, A / B means either A or B.

[0025] The terms "first," "second," and the like in the specification and claims herein are used to distinguish between different objects, rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish between different response messages, rather than to describe a specific order of response messages.

[0026] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0027] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0028] To facilitate understanding of the embodiments of the present application, further explanation will be given below with reference to specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation on the embodiments of the present invention.

[0029] For example, Figure 1 The diagram of a cluster task container management architecture provided by an embodiment of the present application is shown. The system is deployed in a supercomputing cluster, which includes a management node and multiple computing nodes. Figure 1 As shown in the figure, the cluster task container management architecture includes a lightweight state database module, a remote springboard connection module, a container lifecycle control module, a port scheduling and forwarding module, and a port lifecycle monitoring and abnormal recovery mechanism module.

[0030] The lightweight state database module is the storage and record keeping module in the entire system architecture. It is deployed on the cluster management node and can be implemented using pydblite (a lightweight Python library). The lightweight state database does not rely on any external services and works directly with local files. It records the entire lifecycle of each task, including every key step and state change from start to finish, including submission, scheduling, container startup, port mapping, and task recycling. Each record in the database represents a separate task. Each task has a globally unique uuid as its primary key. To facilitate daily queries, a business index called clusterName_user_jobId is designed, allowing users to quickly find relevant records by user account or job ID. Each record contains fields including, but not limited to, basic user information, the cluster and specific compute node to which the task is assigned, the current job status, the associated container running status, the assigned port number and tunnel status, and the last update time. These fields are continuously updated as the task progresses. Through encapsulated interface functions, the system can timely update the state database at each stage, including task submission, node allocation, container startup, port request, forwarding establishment, anomaly detection, and resource recycling, enabling tracking and management of the entire task lifecycle. For example, when a user submits a task, a new record is created and marked "Submitted." Once the scheduler allocates a compute node, the node information and status are updated to "Scheduled." When a container is started, the startup time and process information are recorded; when a port is allocated, the port number is recorded; when a tunnel is established, the tunnel status is recorded. When the system detects an operational anomaly, error details and recovery actions are recorded. This full lifecycle tracking ensures accurate information about the task's progress at all times. The lightweight state database does not require a separate database server; data is stored directly in local files on the management node and can be directly accessed through a Python interface. This makes it easy to integrate into existing systems without adding additional maintenance burdens. The file-based storage also facilitates backup and migration, allowing the entire system to be copied to other machines as needed.

[0031] The remote jump-board connection module is used for communication between the entire system and compute nodes. On the cluster resource node side, the system establishes a secure communication channel with the node through the remote jump-board connection module. Implemented based on Paramiko (a Python SSH library), the remote jump-board connection module supports a two-tier SSH jump-board mechanism to establish and maintain secure connections to the cluster's compute nodes. This two-tier jump-board design first connects to the cluster's login node, then transparently connects to the target compute node through the login node. This approach adapts to the security architecture of most high-performance computing clusters, which typically prohibit direct login to compute nodes and require a redirection through a designated login node. The connection process uses a private key in Ed25519 format for authentication. The remote jump-board connection module creates a special transparent channel, using the open_channel('direct-tcpi p') method, to establish a communication channel between the login node and the compute node. All management commands and data are transmitted through this encrypted communication channel, ensuring stable and secure remote command issuance and response. It is worth noting that for security reasons, this module does not permanently save SSH keys. Instead, it is only used temporarily during the current session. After the session ends, the key is cleared from memory, which greatly reduces the risk of key leakage. When the system needs to perform an operation on the computing node, such as starting a container or checking the running status, this module will send the command through the established secure channel and obtain the execution results in real time. If the connection is accidentally interrupted, the module will automatically attempt to re-establish the connection. Throughout the process, all connection status and operation records will be updated in real time to the lightweight state database to ensure that every step of the operation is traceable.

[0032] The container lifecycle control module is used to manage the entire lifecycle of containers on computing nodes, and implements container start, stop, preparation, and status query based on Singularity containers. When the system receives a task request submitted by the user, the container lifecycle control module starts working and dynamically assembles a complete container startup command, which contains various necessary parameter configurations. For example, setting up the overlay file system to provide writable space within the container, mounting the user's home directory to ensure access to personal files, and injecting specific environment variables to configure the operating environment. All these parameters will be automatically generated according to the task requirements, without the need for manual specification by the user. Then, the singularity instance start command is executed on the target node through a remote session, and parameters such as the overlay file system, mount path, and environment variables are automatically injected to complete the container instantiation. In order to facilitate subsequent identification and management, each started container will be named according to a unified rule, such as ng inx_auto_ <baseport>In this format, basePort is the port number assigned to this task, allowing the corresponding container to be quickly found by name. After the start command is executed, the module checks the operation result, which indicates whether the container startup was successful or not. This is determined by parsing the instance list returned remotely. Specifically, the module confirms the successful startup of the container by parsing the output of the singularity instance list command. If a startup failure is detected, the module immediately logs the error message and attempts to recover the operation. While the container is running, the module collaborates with other components to continuously monitor the container status. When the task is completed or an exception occurs and requires termination, the module executes the singularity instance stop command to shut down the container instance and clean up related resources. Throughout the entire process, from container startup, running status check, to final shutdown, the module updates the lightweight status database in real time at each key node, ensuring that the system has accurate information about each container at all times.

[0033] The port scheduling and forwarding module facilitates user access to remote containers, resolving port conflicts and remote access issues in multi-user environments. This module uses a bitmap structure to initialize a local pool of available ports. When a user task requires remote access, the port scheduling and forwarding module allocates an available port from a pre-defined pool. This pool of ports is not simply a sequential sequence of numbers, but rather a regularly distributed sequence of ports generated at a fixed step size (for example, at fixed intervals, such as 8090, 8100, and 8110, where every tenth port is selected). This discrete distribution significantly reduces the probability of port conflicts. When multiple users simultaneously request a port, the system incorporates a memory lock (memLock) mechanism. The port scheduling and forwarding module generates an MD5 hash value based on information such as the cluster name, user, job number, node IP address, and base port number, acting as a temporary lock. This ensures that only one request can access a specific port resource at a time. Once the port is successfully allocated, the module immediately establishes a local SSH tunnel, mapping the localhost:basePort port on the management node to the containerNode:22 port on the target compute node, enabling seamless connectivity between the local and remote containers. Users simply connect to the local port, and data is automatically forwarded to the remote compute node via an encrypted tunnel. To ensure security, all tunnel connections utilize encrypted transmission, and the system records each port allocation and tunnel status. Upon task completion, the module automatically reclaims the port resource, marking it free for use by subsequent tasks.

[0034] The port lifecycle monitoring and exception recovery mechanism module periodically scans the status of local SSH tunnels to address port failures or container anomalies. It verifies whether any processes are listening on each port and whether each SSH tunnel remains open. When a port forwarding failure anomaly is detected, such as the unexpected termination of a tunnel process or a port becoming unresponsive, the exception handling process is immediately triggered. First, the lightweight state database is queried to identify the specific tasks, containers, and compute nodes associated with the port, determining whether the corresponding jobs and containers are still alive. If the container is still functioning but the tunnel connection is faulty, the mechanism automatically triggers the tunnel reconstruction process to re-establish the SSH forwarding channel from the management node to the compute node. If the container itself is found to have stopped running, detailed information about the anomaly, including the time of occurrence, anomaly type, and associated tasks, is recorded and updated to the state database. The port is also marked as abnormal to prevent subsequent tasks from accidentally reusing it. In certain configurations, resource recovery can also be triggered to clean up any remaining port usage information. All of these operations generate detailed logs, providing a complete clue for subsequent problem analysis.

[0035] The above describes the cluster task container management architecture. Key system events, including job submission, node allocation, container startup, port application, tunnel establishment, exception recovery, and resource recycling, are written to the state database through encapsulated API interfaces, forming a complete lifecycle record chain. The platform also supports information retrieval based on the state database, such as querying user tasks by email, querying port status, and checking expired jobs, providing data support for subsequent system visualization, operation and maintenance management, and troubleshooting.

[0036] Based on this system architecture, the present application embodiment provides a cluster task container management method. For example, Figure 2 A schematic diagram of a cluster task container management method provided by an embodiment of the present application is shown. The method is applied to Figure 1 As shown in the system architecture. Figure 2 As shown, the method may include the following steps:

[0037] S21: Receive the task submitted by the user, the system generates a unique ID and records the initial state, and waits for the scheduler to allocate a computing node.

[0038] In this embodiment, when a user submits a new task to the management node, the system first generates a unique uuid for the task as an index. To facilitate subsequent queries, a search identifier is also created, consisting of the cluster name, user name, and job ID. This information, along with the user's basic information and submission time, is recorded in a lightweight state database. The task's status is then marked as "submitted." The system then transfers the task to the underlying scheduler, such as a cluster scheduling system like Slurm or Kubernetes. The scheduler assigns the task to a specific compute node based on current resource availability. Once the assignment is complete, the system immediately updates the state database, recording information such as the target node's IP address and host name, and changes the task's status to "scheduled." If any exceptions occur during this process, such as a scheduling failure or a database update error, the system captures these errors, records the detailed error information and the time of occurrence in the state database, and marks the task's status as abnormal.

[0039] S22: Start the container on the target node through SSH, monitor the running status in real time and update the database.

[0040] In this embodiment, after a task is successfully scheduled to the target compute node, the system starts the container instance. The entire process begins by establishing a secure connection. The system connects to the cluster's login node through a preconfigured SSH jump server, and then connects from the login node to the assigned compute node. Once the connection is established, the system automatically assembles a complete container startup command based on the task requirements submitted by the user. This uses the Singularity container's instance start mode, which is equivalent to running the container as a background service. The startup command includes many automatically generated parameters, such as mounting the user's home directory to allow the container to access personal files, setting up the overlay file system to provide writable space, and injecting various environment variables based on the task requirements. To facilitate subsequent management, each container is assigned a consistent name, typically by embedding the task's assigned port number, such as in the format of nginx_auto_8080. After the command is issued to the compute node and executed, the system runs the singularity instance list command and checks the returned result to confirm that the newly started container appears in the running list. Whether the container startup succeeds or fails, this status information is fed back to the system in real time. If successful, the container status field in the status database is updated to "Running," and the entire task status is changed to "Executing." If it fails, specific error information is recorded, such as whether the image pull failed or the startup command failed, and the task status is marked as abnormal. This comprehensive monitoring approach ensures that the system always knows the true status of each container. All status changes are accurately timestamped, providing comprehensive data support for subsequent troubleshooting and performance analysis.

[0041] S23: Dynamically allocates discrete ports and establishes an encrypted SSH tunnel for users to access remote containers.

[0042] In this embodiment, when a user needs to access a running container, the system initiates the port allocation and tunnel establishment process. First, an available port number is determined from a preconfigured port pool. The ports in this port pool are not sequential numbers, but rather a regularly distributed sequence of ports, such as 8090, 8100, and 8110, generated at a fixed step size. This design significantly reduces the possibility of port conflicts. After determining the port, the system generates a lock identifier. This lock identifier is calculated by combining the cluster name, username, job ID, node address, and port number to create an MD5 hash, ensuring that only one task can use this port resource at a time. After obtaining the port, the system initiates an SSH tunnel on the management node, mapping the locally allocated port to port 22 on the target compute node. The system records the allocated port number and the established tunnel status in detail in a status database, including information such as the time of allocation, the compute node in use, and the associated container. Subsequently, whenever a user wishes to access a container, they simply connect to this port on the management node, and the request will be automatically forwarded via an encrypted tunnel to the node hosting the target container, providing both convenience and security. If the port is no longer needed, the system will automatically reclaim it and mark it as idle, waiting to be allocated to other tasks next time.

[0043] S24: Regularly check the port and container status, automatically recover in case of abnormality, and clean up resources after the task is completed.

[0044] In this embodiment, the system periodically checks all active ports and corresponding SSH tunnels, regularly checking whether each communication line is unobstructed. The system determines the health of the connection by checking the local port listening status and the viability of the SSH tunnel process. If the tunnel corresponding to a port is disconnected, the system first checks which task and container the port is associated with. If the container is still operating normally, but the tunnel connection is experiencing a problem, the system attempts to reestablish the SSH tunnel to restore communication. However, if the container itself is stopped, the system records the anomaly, including the time and specific cause, and marks the port as abnormal to prevent other tasks from using it incorrectly. When a task is ultimately completed, whether normally or abnormally, the system initiates a resource recovery process. The previously allocated port number is returned to the resource pool and the port is remarked as available. Finally, all associated SSH connections are closed. All of these operations are recorded in detail in the status database, including information such as the resource recovery time and operation results, ensuring that the entire task lifecycle is traceable.

[0045] Figure 3 This is a schematic diagram of the overall process of cluster task execution provided by the embodiment of this application. Figure 3 As shown, users request API calls for start, stop, or prepare actions, triggering different processes through the instanceAction scheduler. During startup, the system first connects to the target instance via a two-tiered SSH springboard (connecting to the login node and then transparently transmitting to the compute node). The system then calls startInstance and constructs the Singularity startup command using the dynamic parameters generated by dynamicScriptVariables. After the container is started, the state is updated to a lightweight state database (based on PyDbLite, using uuid as the primary key and clusterName_user_jobId as the index to record the entire lifecycle state). During the preprocessing phase, after connecting to the login node, the prepare.sh script is executed to copy the model data to the preprocessing component. For port management, a Bitmap port pool is used to generate available port segments based on the step size. An MD5 hash lock is generated using memLock to ensure port uniqueness in a concurrent environment. After application, an SSH tunnel is established from localhost:basePort to containerNode:22 and the forwarding rules are recorded. During shutdown, the system connects to the instance, executes the Singularity stop command, and updates the state to stopped / revoked. The system records all key events through the status database module (including interfaces such as createJobSubmit and checkExpiredRecord). Administrators or inspectors regularly scan the port forwarding status and automatically rebuild tunnels or recycle resources when an abnormality occurs. The entire mechanism does not change the original Slurm / K8s scheduling logic and has security isolation, multi-user concurrency support and abnormal self-healing capabilities.

[0046] The above is an introduction to the cluster task container management method provided by the embodiments of this application. PyDbLite is used to build a lightweight state database, using UUID as the primary key to record the status of the entire task lifecycle, including submission, scheduling, execution, and recycling. Compute nodes are connected via a two-layer SSH springboard, and containers are dynamically started using the Singularity runtime, automatically injecting parameters such as the overlay file system and environment variables. A port pool management mechanism is designed, allocating ports using discrete steps and using MD5 hash locks to ensure resource uniqueness during multi-user concurrency. An SSH tunnel is also established to provide secure local access to the container. The system has an exception recovery capability, regularly scanning ports and container status, automatically reestablishing connections or reclaiming resources in the event of an anomaly. This method has the following advantages: a high degree of automation, eliminating the need for manual user configuration of ports and mapping paths; the system fully automatically allocates, schedules, and reclaims tasks. Full traceability of task status: The lightweight state library supports real-time updates and historical queries of all task statuses. Robust exception recovery: Any container or port anomaly is automatically identified and repaired. The platform is highly non-invasive: The state library and forwarding logic can be deployed without changing the original Slurm / K8s scheduling logic, ensuring high compatibility. It supports multi-user concurrency: The bitmap + hash lock scheduling mechanism inherently guarantees concurrent access.

[0047] It is understandable that the size of the sequence number of each step in the above-mentioned embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in some possible implementations, the steps in the above-mentioned embodiments can be selectively executed according to actual conditions, and can be partially executed or fully executed, which is not limited here. All or part of any features of any embodiment of the present application can be freely and arbitrarily combined without contradiction. The combined technical solution is also within the scope of the present application.

[0048] Based on the methods in the above embodiments, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the methods in the above embodiments.

[0049] Based on the methods in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the methods in the above embodiments.

[0050] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0051] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0052] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0053] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.< / baseport> < / baseport>

Claims

1. A cluster task container management method, characterized in that: Applied to a supercomputing cluster, the supercomputing cluster includes a management node and multiple computing nodes, the method includes: Build a lightweight state database to record the status of the entire task life cycle, including submission, scheduling, container startup, port mapping, and recycling stages; Design a bitmap-based port pool mechanism to generate a discretely distributed sequence of available ports at a fixed step size, and ensure the concurrent uniqueness of port applications through memory locks; Establish an SSH tunnel on the management node and forward the local port to the SSH port of the target computing node; Implement container lifecycle management, regularly scan tunnel status, and automatically trigger container status checks and tunnel reconstruction when an exception occurs; All state events are written to the lightweight state database through the API, and the lightweight state database supports task query and connection health check.

2. The method according to claim 1, characterized in that The lightweight status database is implemented using PyDbLi te. The task record uses the task's uuid as the primary key and clusterName_user_jobId as the business index. It contains user information, node information, job status, container status, and port mapping status fields.

3. The method according to claim 1, characterized in that In the port pool mechanism, the memory lock is generated by the MD5 hash algorithm, and the hash input includes the cluster name, user name, job ID, node IP and port.

4. The method according to claim 1, wherein The SSH tunnel is established through a double-layer jump board, including: Connect to the cluster login node through the Paramiko library; Transmit data from the login node to the target compute node, using Ed25519 key authentication and the open_channel method to create an encrypted channel.

5. The method according to claim 1, wherein The container lifecycle management is based on the Singularity runtime, and the container instance naming rule is service_auto_ <baseport> , dynamically inject the overlay file system, mount path and environment variables at startup.< / baseport> 6. The method according to claim 1, characterized in that Automatically trigger container status checks and tunnel reconstruction when an exception occurs, including: When the SSH tunnel is detected to be interrupted, query the lightweight state database to obtain the associated container status; If the container survives, the tunnel is automatically rebuilt; if the container fails, the exception is marked and the port resources are released.

7. The method according to claim 1, characterized in that The method is compatible with the Slurm or Kubernetes scheduling system, and the lightweight state database is deployed as an independent module without modifying the original scheduling logic of Slurm or Kubernetes.

8. A cluster task container management system, characterized in that: Deployed in a supercomputing cluster, the supercomputing cluster includes a management node and multiple computing nodes. The system includes: A lightweight state database module, used to record the status of the entire life cycle of a task, including status information during submission, scheduling, container startup, port mapping, and recycling. Remote springboard connection module, used to establish a secure communication channel with the computing node through a double-layer SSH springboard mechanism; The container lifecycle control module is used to implement lifecycle management of containers, including starting, stopping, and status monitoring, based on the Singularity runtime. The port scheduling and forwarding module is used to allocate ports through a bitmap mechanism and establish an SSH tunnel to achieve access forwarding from the local machine to the container; The port lifecycle monitoring and abnormal recovery mechanism module is used to regularly detect the port and tunnel status and automatically trigger the recovery process when an abnormality occurs.

Citation Information

Patent Citations

  • Access method and device for virtual machine

    CN107193634A

  • Container starting method and device

    CN110554905A

  • Browser-based container remote login method and device

    CN111221665A

  • Construction task processing method and device based on Jenkins, electronic equipment and medium

    CN115202818A

  • Code property right management system based on cloud intelligent contract compiling

    CN117786622A