An architecture system, a deployment method and an application method for remote access and management of high-performance computing data
Patent Information
- Application Number
- CN202611230862.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-25
AI Technical Summary
[0013]为此,本发明为解决现有技术中数据管理碎片化、效率低、集成性差等缺点,提供一种高性能计算数据远程访问与管理的架构系统、部署方法及应用方法,通过统一、高效、灵活的远程数据管理解决方案,部署HPC集群,实现并行存储与二级存储的透明访问、远程文件全生命周期管理及批量数据高效迁移
[0066]本发明提供一套统一、高效、灵活的远程数据管理解决方案,部署在HPC集群,实现并行存储与二级存储的透明访问、远程文件全生命周期管理及批量数据高效迁移。
Smart Images

Figure CN122824751A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote data access and management technology. Specifically, it relates to an architecture system, deployment method, and application method for remote access and management of high-performance computing data. Background Technology
[0002] Data management in a high-performance computing (HPC) environment involves: a two-tier storage architecture of "parallel storage + secondary storage (NAS storage)," where parallel storage stores data for real-time computation, and secondary storage is used for archiving, backup, and non-real-time data storage; the lack of a unified access tool necessitates different access methods depending on the storage type, with parallel storage accessed locally and secondary storage accessed remotely via NFS mounting or SSH / FTP / SFTP protocols, resulting in completely different access commands and path formats for the two types of storage; file upload, download, deletion, movement, and permission modification operations require different command tools (such as scp, ftp, rm, chmod, etc.) without unified command semantics, leading to fragmented operation processes; existing tools cannot be deeply integrated with job workflows, and modular jobs cannot seamlessly access remote data, requiring the manual writing of complex scripts for data acquisition and migration, resulting in high adaptation costs.
[0003] Therefore, it has the following drawbacks:
[0004] (1) Fragmented access methods and high learning and usage costs: The access commands and path formats of parallel storage and secondary storage are not uniform. Users need to remember multiple commands and operation logics. Moreover, the operation habits of different users are very different, which makes the scripts unusable and increases the auxiliary workload of pattern development.
[0005] (2) Lack of transparent access capability and cumbersome operation: Users need to manually distinguish between the paths of parallel storage and secondary storage, and cannot achieve seamless switching between the two types of storage. For example, when switching from parallel storage to secondary storage to access data, the mount command or remote access address needs to be re-entered, and the operation process is redundant.
[0006] (3) Low data migration efficiency and poor stability: Conventional scp, ftp and other transmission tools do not support recursive directory transmission, compressed transmission, permission retention and breakpoint resume functions. The data is mostly large files (single files can reach several GB) and mostly organized in the form of directories, which leads to slow batch data migration speed and easy interruption. Furthermore, after migration, information such as file permissions and timestamps are lost, affecting data availability.
[0007] (4) Weak remote file management capabilities and high operation and maintenance costs: There is no unified command set to realize the full life cycle management of remote files. Operations such as remote directory switching, file search, and permission modification require the combination of multiple commands. Furthermore, it is impossible to quickly view the disk usage of remote storage. It is difficult for operation and maintenance personnel to uniformly manage the data access permissions and storage resources of multiple users.
[0008] (5) Poor integration and inability to adapt to the workflow: Existing tools do not have standardized interfaces and cannot be seamlessly embedded into the workflow. Data upload and download operations need to be manually triggered during the operation of the workflow, which cannot realize the automatic acquisition, migration and management of data, thus restricting the automation level of the workflow development.
[0009] The methods in the existing technology cannot solve the above-mentioned shortcomings and problems:
[0010] For example, using SSH-based script tools: By writing shell scripts to encapsulate commands such as scp and ftp, some remote file operation functions can be implemented. However, it can only implement simple command combinations, cannot implement a unified abstract interface, cannot shield storage differences, has low transmission efficiency, poor stability, high maintenance costs, and cannot adapt to the needs of large files and multiple storage scenarios.
[0011] For example, commercial distributed file systems (such as GlusterFS and Ceph): These systems can achieve unified management of multiple storage systems, but they are costly, complex to deploy, and cannot be adapted to the existing environment of HPC clusters. They are difficult to integrate with job workflows, cannot provide standardized command-line tools, and have high user learning costs, which does not meet the R&D requirements of low cost and high adaptability.
[0012] For example, existing open-source file management tools such as rsync and lftp can only implement some data transfer or file management functions. They lack a unified command set, cannot achieve transparent access and multi-mode deployment, have limited functionality, and cannot cover all the core requirements of this invention. They can only serve as auxiliary tools for this invention. Summary of the Invention
[0013] To address the shortcomings of existing technologies, such as fragmented data management, low efficiency, and poor integration, this invention provides an architecture system, deployment method, and application method for remote access and management of high-performance computing data. Through a unified, efficient, and flexible remote data management solution, HPC clusters are deployed to achieve transparent access to parallel storage and secondary storage, remote file lifecycle management, and efficient batch data migration.
[0014] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0015] A high-performance computing data remote access and management architecture system adopts a four-layer layered architecture, including a bottom adaptation layer, a configuration management layer, a command-line tool layer, and a user interaction layer. The bottom adaptation layer provides a unified access interface for accessing different storage systems. The configuration management layer is responsible for the environment configuration, storage information configuration, and authentication information configuration of the high-performance computing data remote access and management architecture system. The command-line tool layer includes a standardized command set module and a command parsing module. The command parsing module parses user-input commands according to the standardized command set module. The user interaction layer receives user-input commands and configuration instructions for remote access to parallel storage and secondary storage, transmits the user-input commands and configuration instructions to the command-line tool layer, and simultaneously feeds back the execution results of the standardized commands corresponding to the user-input commands and configuration instructions to the user.
[0016] The aforementioned architecture system for remote access and management of high-performance computing data uses Python to encapsulate the parallel storage interface and the secondary storage interface into a unified interface for accessing parallel storage and secondary storage in the underlying adaptation layer; it uses the Lustre file system API interface to access parallel storage, and uses the NFS protocol interface and NAS gateway interface to access secondary storage.
[0017] The aforementioned architecture system for remote access and management of high-performance computing data encapsulates the parallel storage interface and the secondary storage interface into a unified abstract interface using Python. Specifically, it uses the `def read_file(storage_type, path)` or `def write_file(storage_type, local_path, remote_path)` language to generate configuration files, storing the configuration parameters corresponding to remote access to secondary and parallel storage. The access commands for the parallel and secondary storage interfaces are abstracted into a unified command interface using Linux or native UNIX access methods. This allows a single command to access both parallel and secondary storage interfaces simultaneously in a high-performance computing system. The configuration parameters include the access node's IP address or domain name, the secondary storage root directory path `nas_path`, the username `user`, and the password `password`. Linux or native UNIX commands include `cd`, `df`, `cp`, `is`, `rm`, `chmod`, `get`, `put`, and `mv`.
[0018] The aforementioned architecture system for remote access and management of high-performance computing data includes a configuration management layer that stores configuration parameters in a configuration information storage file `cmaConfig`. The configuration script file `pycmafs.py` is used to read and modify these parameters. When a user configures access to parallel storage and secondary storage, the `pycmafs.py` file first verifies the MD5 encoding of the `cmaConfig` file. If they are inconsistent, the `cmaConfig` file generated by `pycmafs.py` replaces the contents of the old `cmaConfig` file. This verifies the validity of the user-input access node IP address or domain name and the secondary storage root directory path `nas_path`. If the configuration parameters are incorrect, a clear prompt is given to ensure the correctness of the configuration parameters. The configuration parameters include the access node IP address or domain name, the secondary storage root directory path `nas_path`, the username `user`, and the password `password`. These parameters are automatically written to the `cmaConfig` file after the user completes the configuration.
[0019] The aforementioned architecture system for remote access and management of high-performance computing data includes a standardized command set module in the command-line tool layer. This module contains 14 standardized commands, each prefixed with "cma" and supporting the "-help" parameter. Each standardized command defines explicit parameter rules and supports positional and optional parameters. The command parsing module uses the "argparse" module to parse the parameters of the 14 standardized commands, handling registration, parameter parsing, syntax validation, and help documentation generation. The command parsing module extracts the standardized command name from the user-input command and sends this information to parallel or secondary storage, where the corresponding operation is then performed.
[0020] The 14 standardized commands include:
[0021] (1) cma: Used to list all commands supported by the system and a brief description of their functions;
[0022] (2) cmacd: Moves the remote working directory, supports seamless switching between parallel storage and secondary storage directories, without the need to manually distinguish storage paths; the calling format is: cmacd [parameters] [path];
[0023] (3) cmals: View files under a specific path on a remote server. Supports multiple parameters and displays detailed file information. The calling format is: cmals [parameters] [path].
[0024] (4) cmapwd: Get the current remote working directory path, automatically distinguish the storage type, and return path information in a uniform format; the calling format is: cmapwd [parameters];
[0025] (5) cmaput: Uploads local files or directories to a remote server, i.e., parallel storage or secondary storage. It supports recursive upload, compressed upload, permission preservation, and resume upload. The calling format is: cmaput [parameters] [local path] [target path].
[0026] (6) cmaget: Downloads files or directories from a remote server to the local machine. It supports recursive download, compressed download, permission preservation, and resume download. The parameters are consistent with cmaput, reducing the learning cost. The calling format is: cmaget [parameter] [target path] [local path];
[0027] (7) cmarm: Deletes files or directories on a remote server. It supports recursive deletion, confirmation deletion with the -i parameter, and forced deletion with the -f parameter to avoid accidental operations. The calling format is: cmarm [parameter] [path];
[0028] (8) cmamv: Moves files or directories on a remote server. It supports the overwrite prompt -i parameter and the forced overwrite -f parameter, skips non-existent files, and does not display redundant information. The calling format is: cmamv [parameter] [file path] [target path];
[0029] (9) cmamkdir: Creates a directory on a remote server, supports recursive creation with the -p parameter to ensure complete directory hierarchy; the calling format is: cmamkdir [parameter] [path];
[0030] (10) cmatouch: Creates or modifies files on a remote server, supports multiple time configuration parameters; the calling format is: cmatouch [parameters] [file / directory];
[0031] (11) cmalocate: Searches for documents that meet the conditions on a remote server. It supports regular expression matching, case-insensitive matching, and matching quantity limit parameters to improve search efficiency. The calling format is: cmalocate [parameter] [matching mode];
[0032] (12) cmacp: Copy files or directories on a remote server. It supports recursive copying, permission preservation, and overwrite prompts, and is suitable for remote file backup needs. The calling format is: cmacp [parameters] [source file] [target file];
[0033] (13) cmachmod: Modifies the permissions of files or directories on a remote server. It supports recursive modification with the -R parameter and permission change prompts with the -c parameter, and is compatible with multi-user permission management. The calling format is: cmachmod [parameter] [permission mode] [file / directory];
[0034] (14) cmadf: Statistical analysis of disk usage in remote file systems. It supports multiple display formats (GB / MB / KB) and specifies storage type statistical parameters, which is convenient for operation and maintenance monitoring. The calling format is: cmadf [parameters] [file / directory].
[0035] The aforementioned architecture system for remote access and management of high-performance computing data includes a command-line interaction module, a modal job flow module, and a help documentation module at the user interaction layer; wherein:
[0036] The command-line interaction module supports terminal command input and real-time return of execution results;
[0037] The pattern job flow module provides interfaces for 14 standardized commands; these standardized command interfaces are directly embedded in the job flow script, supporting the automatic invocation of the 14 standardized commands prefixed with cma by the job.
[0038] The help documentation module provides support for the -help parameter for each standardized command. The -help parameter displays the function, calling format, and parameter description of the standardized command.
[0039] The deployment method of the above-mentioned high-performance computing data remote access and management architecture system includes, in the global deployment of the system users:
[0040] Step 11: Obtain the system installation package cmas.tar.gz and upload it to the specified directory on the HPC cluster server;
[0041] Step 12: Unzip the installation package. Execute the command "tar -zxvf cmas.tar.gz -C / opt / cmas" to unzip the tar package to the / opt / cmas directory, and then switch to the unzipped directory cd / opt / cmas;
[0042] Step 13: Configure storage information. Enter the . / config directory, cd / opt / cmas / config, and execute the command "python pycmafs.py". Enter the configuration parameters according to the terminal prompts. The configuration parameters include the access node IP address, the secondary storage name and the root directory path nas_path. Confirm after entering, and the configuration parameters will be automatically written to the configuration information storage file cmaConfig.
[0043] Step 14: Configure system environment variables. Edit the / etc / profile file and add two lines of configuration at the end of the file: "export CMA_HOME= / opt / cmas" and "export PATH=$PATH:$CMA_HOME / bin";
[0044] Step 15: Make the environment variables take effect by executing the command "source / etc / profile";
[0045] Step 16: Deployment verification. Enter "cma" in the terminal. If a list of all commands supported by the system is returned, the global deployment is complete, and all users can use the cma standardized command set in the terminal.
[0046] The deployment method of the above-mentioned high-performance computing data remote access and management architecture system, in a single-user deployment, includes:
[0047] Step 21: Obtain the system installation package cmas.tar.gz and upload it to the home directory of this single user, / home / xxx, where xxx is the username;
[0048] Step 22: Unzip the installation package. Execute the command "tar -zxvf cmas.tar.gz -C / home / xxx / cmas" to unzip the tar package to the cmas directory under the user's home directory, and then switch to the unzipped directory cd / home / xxx / cmas;
[0049] Step 23: Configure storage information. Edit the configuration information storage file cmaConfig, i.e., vi / home / xxx / cmas / cmaConfig, or execute the command "python pycmafs.py". Enter the configuration parameters as prompted. The configuration parameters include the access node IP address, username user, password password, secondary storage name and root directory path nas_path. Save after configuration. The configuration parameters will be automatically written to the configuration information storage file cmaConfig.
[0050] Step 24: Configure system environment variables. Edit the ~ / .bash_profile file and add two lines at the end of the file: "export CMA_HOME= / home / xxx / cmas" and "export PATH=$PATH:$CMA_HOME / bin";
[0051] Step 25: Make the environment variables take effect by executing the command "source ~ / .bash_profile";
[0052] Step 26: Deployment verification. Enter "cma" in the terminal. If a list of all commands supported by the system is returned, the single-user deployment is complete. Only that single user can use the cma standardized command set in the terminal.
[0053] The deployment method of the above-mentioned high-performance computing data remote access and management architecture system includes, in the module deployment:
[0054] Step 31: The administrator completes the setting of the module environment variables and writes the CMAFS environment configuration into the module file;
[0055] Step 32: The administrator configures the environment variables for the module by adding CMA_HOME and PATH configurations to the module file;
[0056] Step 33: The user checks all software environment variables and executes the command "module avail". If the output contains "apps / cmafs / cmafs", then the module environment configuration is complete.
[0057] Step 34: The user imports the cmafs environment variables and executes the command "module load apps / cmafs / cmafs". After loading is complete, proceed to the next step to continue execution.
[0058] Step 35: Deployment verification. Enter "cma". If a list of commands is returned, the module deployment is complete. Users can directly use the cma standardized command set to uninstall the environment by executing "module unload apps / cmafs / cmafs".
[0059] The application method of the above-mentioned high-performance computing data remote access and management architecture system includes the following steps:
[0060] Step A: Loading the user's operating environment: The user logs into the HPC node and selects the corresponding environment loading method according to the deployment method: Global deployment and single-user deployment do not require additional loading and can be done directly using the terminal; Module deployment requires executing "moduleload apps / cmafs / cmafs" to load the environment;
[0061] Step B: User Operation Configuration Initialization: Upon first use, the user executes "python pycmafs.py" and enters configuration parameters including the storage node IP, storage path, and authentication information as prompted. After configuration, the configuration parameters are automatically saved to the configuration information storage file cmaConfig.
[0062] Step C: User operation command input: The user enters the standardized command set related to CMA in the terminal, along with the corresponding parameters, and presses Enter to submit the command after completing the input;
[0063] Step D: System command parsing and execution: The command line tool layer receives the command input by the user, parses the command, then calls the configuration management layer to obtain the storage configuration information, and then accesses the corresponding parallel storage or secondary storage through the unified interface of the underlying adaptation layer to perform the corresponding operation;
[0064] Step 5: System feedback results: After the system completes the execution, the results will be returned to the terminal in real time. The user can judge whether the operation was successful based on the results. If it fails, the user can modify the command or configuration according to the prompts.
[0065] The technical solution of the present invention achieves the following beneficial technical effects:
[0066] This invention provides a unified, efficient, and flexible remote data management solution, deployed on an HPC cluster, enabling transparent access to parallel storage and secondary storage, remote file lifecycle management, and efficient batch data migration.
[0067] 1. The system architecture adopts a layered architecture design, consisting of four layers. Each layer is independently encapsulated and works collaboratively, completely solving the problem of fragmented storage access in existing technologies.
[0068] (1) Achieve transparent access and simplify operation process:
[0069] In the underlying adaptation layer, access interfaces for different storage systems are encapsulated to achieve a unified abstract interface: the parallel storage interface and the secondary storage interface are encapsulated into a unified abstract interface using Python, and unified interface functions are defined, such as def read_file(storage_type, path) and def write_file(storage_type, local_path, remote_path), etc., where storage_type is automatically identified by the system based on the configuration, without requiring user input, thus achieving transparent access;
[0070] Through the unified abstract interface of the underlying adaptation layer, users do not need to manually distinguish between the paths of parallel storage and secondary storage. They can achieve seamless switching between the two types of storage (parallel storage and secondary storage) through a unified command. For example, users can directly view the data directory of the secondary storage by executing "cmals / nas / data" without first mounting the NAS storage. The operation process is greatly simplified, improving the convenience of data access.
[0071] (2) In the configuration management layer, configuration and operation are separated and implemented using separately developed cmaConfig and pycmafs.py. The configuration information storage file cmaConfig is stored in JSON format, and the configuration script file pycmafs.py manages the configuration by reading and modifying the JSON file. It is separated from the core running code of the system, and the system does not need to be recompiled when modifying the configuration parameters, thus improving flexibility.
[0072] (3) This invention solves the problem of fragmented access methods and reduces learning and usage costs: This system provides a unified set of command-line tools (14 standard command sets) at the command-line tool layer. All standardized commands adopt a unified naming convention and calling format, which shields the underlying differences between parallel storage and secondary storage. Users do not need to memorize multiple commands and operation logic. They only need to learn a set of CMA standardized command sets to complete all remote data operations. The scripts can be reused, which greatly reduces the user's learning cost and the auxiliary workload of mode development. According to actual testing, the user's operation efficiency has been improved by more than 60%.
[0073] (4) Improve data migration efficiency and stability to meet the needs of large meteorological files: The cmaput and cmaget commands in the standard command set support recursive transmission, compressed transmission, and breakpoint resume transmission. For large files (several GB) and directory-level data, the transmission speed is 30%–50% faster than existing scp / ftp tools. At the same time, it avoids the problem of retransmission after transmission interruption, and the transmission failure rate is reduced to less than 1%, ensuring the efficiency and stability of data migration. Moreover, the permissions, timestamps and other information of the migrated files can be completely preserved to ensure data availability.
[0074] (5) Enhance remote file management capabilities and reduce operation and maintenance costs: 14 core standard command sets cover the entire lifecycle management of remote files, support remote directory switching, file search, permission modification, disk statistics and other operations. Operation and maintenance personnel can realize multi-user data access permissions and storage resource management through unified commands without the need to combine multiple tools, improving operation and maintenance efficiency by more than 50% and reducing the operation and maintenance error rate.
[0075] (6) Command parsing implementation in the command-line tool layer: The argparse module is used to parse standardized command parameters. argparse is an official Python standard library and plays a key role in command registration, parameter parsing, syntax validation, and help documentation generation. Each standardized command defines clear parameter rules, supports positional parameters and optional parameters, and provides clear prompts when parsing errors occur, improving the user experience.
[0076] (7) Enhance integration and adapt to the automation requirements of the mode operation flow: This system provides a standardized command interface through the mode operation flow module in the user interaction layer. The operation flow script can be directly embedded to realize the automatic acquisition, migration and management of data without the need for manual user triggering. This promotes the automation level of R&D, adapts to the construction requirements of the national meteorological big data cloud platform, and realizes the unified storage management and automatic migration of mode test data.
[0077] 2. This invention provides three standardized deployment methods, covering three scenarios: global user, single-user, and module environments. The deployment process is clear and repeatable. It can adapt to the needs of multiple HPC scenarios and improve deployment flexibility: The three deployment methods (global, single-user, and module) can be flexibly selected according to the actual operation and maintenance needs of the HPC cluster. Global deployment allows all users to share the same space, single-user deployment achieves permission isolation between different users, and module deployment adapts to the module environment specifications of the HPC cluster, with environment variable configurations meeting supercomputing operation and maintenance requirements. The three deployment methods cover the needs of different user groups and support multi-user permission isolation (the configuration of single-user deployment only takes effect for that user), avoiding configuration conflicts between users. It achieves a standardized and repeatable deployment process, reducing deployment and maintenance costs. Attached Figure Description
[0078] Figure 1 A schematic diagram of the architecture system for remote access and management of high-performance computing data according to the present invention;
[0079] Figure 2 The application access flowchart of the architecture system for remote access and management of high-performance computing data of this invention. Detailed Implementation
[0080] Example 1
[0081] In this embodiment, a high-performance computing data remote access and management architecture system is constructed, consisting of four layers: a bottom adaptation layer, a configuration management layer, a command-line tool layer, and a user interaction layer. These four layers are independently encapsulated yet work collaboratively, thereby completely solving the problem of storage access fragmentation in existing technologies. The specific implementation of each layer is as follows:
[0082] 1. Low-level adapter layer: As a bridge connecting the system and external storage devices, its core function is to encapsulate the access interfaces of different storage systems, shield the underlying differences between parallel storage (Lustre) and secondary storage (NAS), and provide a unified access interface to the outside world.
[0083] (1) Parallel storage: By calling the API interface of the Lustre file system, local mounting, file reading and writing, and permission control of parallel storage can be implemented, supporting large-scale parallel access and adapting to the data reading and writing needs of real-time computing.
[0084] (2) Secondary storage: Remote access to secondary storage is achieved through the NFS protocol interface and NAS gateway interface, supporting file upload, download, deletion and other operations, and adapting to data archiving and backup needs.
[0085] (3) Unified encapsulation of access commands: The underlying parallel storage interface and secondary storage interface are encapsulated using Python to form a unified abstract interface. The access commands for the parallel storage interface and secondary storage interface are abstracted into a unified command interface through the access methods provided by Linux or native UNIX, using the def read_file(storage_type, path) or def write_file(storage_type, local_path, remote_path) language. This allows the use of a single command to access both parallel storage and secondary storage interfaces simultaneously. The command-line tool layer does not need to distinguish between storage types; it only needs to call the interface to call the system command to access different storage types, achieving transparent access.
[0086] 2. Configuration Management Layer: Responsible for system environment configuration, storage information configuration, and authentication information configuration, resolving the problems of cumbersome deployment and chaotic configuration in existing technologies. Specific implementation includes:
[0087] (1) Configuration file: It includes two core files: configuration information storage file cmaConfig and configuration script file pycmafs.py. The configuration management layer stores configuration parameters (such as storage node IP, storage path, user authentication information, etc.) in the configuration information storage file cmaConfig, and uses the configuration script file pycmafs.py to read and modify configuration parameters, supporting interactive configuration by users.
[0088] The architecture design achieves separation of configuration and operation: independent configuration files (cmaConfig) and configuration script files (pycmafs.py) are used to separate system configuration from core running code. It supports dynamic modification of parameters such as storage node IP, storage path, and authentication information without recompiling the system, thereby improving the system's flexibility and maintainability.
[0089] (2) Configuration parameters: The core configurable parameters include ip (access node IP address or domain name), nas_path (secondary storage root directory path), user (username), and password (password). Users can configure these parameters flexibly according to the actual deployment scenario. After configuration, the configuration information is automatically written to the configuration information storage file cmaConfig, without the need for manual modification.
[0090] (3) Configuration verification: When users configure access to parallel storage and secondary storage, the built-in verification logic in the configuration script file pycmafs.py verifies the validity of the IP address and storage path entered by the user (such as checking whether the IP is reachable and whether the path exists). If the configuration is incorrect, a clear prompt will be given to ensure that the configuration information is correct. If the cluster user access is configured in the module way, only the information storage file cmaConfig needs to be configured.
[0091] 3. Command line tool layer: Set up a standardized command set module and a command parsing module.
[0092] As the core functional layer of the system architecture of this invention, the standardized command set module provides 14 standardized command sets, covering all core requirements such as remote file management, data transmission, and system statistics. The command parsing module interprets the relevant parameters of these 14 standardized commands, solving the problems of fragmented and semantically ambiguous commands in existing technologies. Specifically, the implementation is as follows:
[0093] (1) Standardized command set module:
[0094] Each standardized command is named with the prefix "cma", and each command supports the "-help" parameter. Each standardized command has clearly defined parameter rules, supports positional parameters and optional parameters, making it easy for users to see how to use it.
[0095] The specific implementations of each standardized command are shown in Table 1 below:
[0096]
[0097]
[0098]
[0099]
[0100] This invention designs a set of 14 standardized commands prefixed with "cma," encapsulating the access interfaces for parallel storage (Lustre) and secondary storage (NAS). It provides a unified command semantics and calling format, shielding the underlying storage differences and enabling transparent access to both types of storage, thus solving the core problem of fragmented access in existing technologies. Each command defines clear parameter rules, supports positional and optional parameters, and provides clear error messages when parsing errors occur, improving the user experience.
[0101] The `cmaput` and `cmaget` commands support recursive transmission, compressed transmission, and resume interrupted transmission, targeting large files (several gigabytes) and directory-level data. They are particularly well-suited for transmitting large meteorological files and directory-level data such as GRIB and NetCDF, improving data migration efficiency and stability, and resolving the issues of slow and easily interrupted data transmission in existing technologies.
[0102] (2) Command parsing module: The argparse module is used to parse command parameters and is responsible for command registration, parameter parsing, syntax verification and help document generation.
[0103] 4. User Interaction Layer: Responsible for receiving user-input commands and configuration instructions, passing the commands to the command-line tool layer, and providing feedback on the execution results to the user, adapting to the operating habits of HPC users. Specific implementation details:
[0104] (1) Command line interaction module: used to support terminal command input and return execution results in real time (such as command execution success prompts, error prompts, file list, disk usage information, etc.).
[0105] (2) Pattern job flow module: provides interfaces for 14 standardized commands. The interfaces of standardized commands are directly embedded in the job flow script, and the job can automatically call standardized commands with the prefix cma to realize data acquisition, migration and management without manual intervention.
[0106] (3) Help document module: It is used to provide -help parameter support for each standardized command. After entering the -help parameter, the function, calling format and parameter description of the command can be displayed, which makes it convenient for users to learn and use it quickly.
[0107] Example 2, Deployment Method
[0108] To address the shortcomings of existing technologies, such as cumbersome deployment and incompatibility with HPC cluster management standards, this system provides three standardized deployment methods, covering three scenarios: global user, single user, and module environment. The deployment process is clear and repeatable, as detailed below:
[0109] (1) Global deployment for system users (applicable to all users of HPC cluster for shared use)
[0110] Step 1: Obtain the system installation package cmas.tar.gz and upload it to the specified directory (such as / opt directory) of the server in the "Pa-Shuguang" HPC cluster using scp or other transfer tools.
[0111] Step 2: Unzip the installation package. Execute the command "tar -zxvf cmas.tar.gz -C / opt / cmas" to unzip the tar package to the / opt / cmas directory (you can modify the unzip directory according to your actual needs), and switch to the unzipped directory (cd / opt / cmas).
[0112] Step 3: Configure storage information. Enter the . / config directory (cd / opt / cmas / config), execute the command "python pycmafs.py", and enter the configuration parameters according to the terminal prompts. The specific configuration parameters include ip (access node IP address) and nas_path (secondary storage name and root directory path). Confirm after entering, and the configuration parameters will be automatically written to the cmaConfig file.
[0113] Step 4: Configure system environment variables. Edit the / etc / profile file (vi / etc / profile) and add two lines at the end of the file: "export CMA_HOME= / opt / cmas" (CMA_HOME is the decompression path) and "exportPATH=$PATH:$CMA_HOME / bin" (add the system command directory to the environment variables).
[0114] Step 5: Make the environment variables take effect by executing the command "source / etc / profile". The environment variables will take effect immediately without restarting the server.
[0115] Step 6: Deployment verification. Enter "cma" in the terminal. If a list of all supported commands is returned, the global deployment is complete, and all users can use the cma command set in the terminal.
[0116] (2) Single-user deployment (suitable for independent use by a single user, with configurations isolated from other users)
[0117] Step 1: Obtain the system installation package cmas.tar.gz and upload it to the user's home directory (e.g., / home / xxx directory, where xxx is the username);
[0118] Step 2: Unzip the installation package. Execute the command "tar -zxvf cmas.tar.gz -C / home / xxx / cmas" to unzip the tar package to the cmas directory under the user's home directory, and then switch to the unzipped directory (cd / home / xxx / cmas).
[0119] Step 3: Configure storage information. Edit the cmaConfig file (vi / home / xxx / cmas / cmaConfig), or execute the command "python pycmafs.py". Enter the configuration parameters as prompted. The specific parameters include ip (access node IP address), user (username), password (password), and nas_path (secondary storage name and root directory path). Save the configuration after completion.
[0120] Step 4: Configure user environment variables. Edit the ~ / .bash_profile file (vi ~ / .bash_profile) and add two lines at the end of the file: "export CMA_HOME= / home / xxx / cmas" (extraction path) and "export PATH=$PATH:$CMA_HOME / bin";
[0121] Step 5: Make the environment variables take effect by executing the command "source ~ / .bash_profile";
[0122] Step 6: Deployment verification. Enter "cma". If a list of commands is returned, the single-user deployment is complete. Only that user can use the cma command set, and other users are not affected.
[0123] (3) Module deployment (suitable for standardized operation and maintenance of HPC clusters, users can load the environment with one click)
[0124] Step 1: The administrator completes the setting of the module environment variables and writes the cmafs environment configuration into the module file. The specific path is / g1 / app / modules / apps / cmafs / cmafs (which is consistent with the existing module environment path of the HPC cluster).
[0125] Step 2: The administrator configures the environment variables of the module by adding CMA_HOME and PATH configurations to the module file to ensure that the cma command can be automatically recognized after the module is loaded;
[0126] Step 3: The user checks all available software environment variables and executes the command "module avail". If the output contains "apps / cmafs / cmafs", then the module environment configuration is complete.
[0127] Step 4: The user imports the cmafs environment variables by executing the command "module load apps / cmafs / cmafs". After loading, no additional environment variable configuration is required. Proceed to the next step.
[0128] Step 5: Deployment verification. Enter "cma". If a list of commands is returned, the module deployment is complete. Users can directly use the cma command set. To uninstall the environment, execute "module unload apps / cmafs / cmafs".
[0129] Example 3, Application Method
[0130] The core workflow consists of 5 steps, covering the entire process from environment loading to command execution. The workflow is clear and easy to use, solving the problems of cumbersome operation and inconsistent processes in existing technologies. The specific workflow is as follows (in conjunction with...). Figure 2 Core workflow diagram):
[0131] Step 1: Environment Loading (User Operation) - Users log in to the "Pai-Shuguang" HPC node and select the corresponding environment loading method according to the deployment method: global deployment and single-user deployment do not require additional loading and can be done directly using the terminal; module deployment requires executing "module load apps / cmafs / cmafs" to load the environment;
[0132] Step 2: Configuration Initialization (User Operation) - When using it for the first time, the user executes "python pycmafs.py" and enters the storage node IP, storage path, authentication information and other configuration parameters as prompted. After the configuration is completed, the system automatically saves it to the cmaConfig file. Subsequent use does not require repeated configuration. The configuration can be modified by re-executing the command.
[0133] Step 3: Command Input (User Operation) – The user enters cma-related commands (such as cmals, cmaput, etc.) in the terminal, and can include corresponding parameters. After entering the commands, the user presses Enter to submit them.
[0134] Step 4: Command parsing and execution (system processing) – The command-line tool layer receives commands input by the user, parses the commands (identifies command type, parameters, path, etc.), then calls the configuration management layer to obtain storage configuration information, and then accesses the corresponding storage device (parallel storage or secondary storage) through the unified interface of the underlying adaptation layer to perform the corresponding operations (such as viewing files, uploading files, etc.).
[0135] Step 5: Result Return (System Feedback) - After the system completes the execution, the execution results (success message, error message, file list, disk usage information, etc.) will be returned to the terminal in real time. Users can judge whether the operation was successful based on the results. If it fails, they can modify the command or configuration according to the prompt.
[0136] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of the claims of this patent application.
Claims
1. An architecture system for remote access and management of high-performance computing data, characterized in that, A four-layer architecture is adopted, including a bottom adaptation layer, a configuration management layer, a command-line tool layer, and a user interaction layer. The bottom adaptation layer provides a unified access interface for accessing different storage systems. The configuration management layer is responsible for the environment configuration, storage information configuration, and authentication information configuration of the architecture system for remote access and management of high-performance computing data. The command-line tool layer has a standardized command set module and a command parsing module. The command parsing module parses the commands input by the user according to the standardized command set module. The user interaction layer is responsible for receiving the commands input by the user and the configuration instructions for remote access to parallel storage and secondary storage, and transmitting the commands and configuration instructions input by the user to the command-line tool layer. At the same time, it feeds back the execution results of the standardized commands corresponding to the commands and configuration instructions input by the user to the user.
2. The architecture system for remote access and management of high-performance computing data according to claim 1, characterized in that, In the underlying adaptation layer, the parallel storage interface and the secondary storage interface are encapsulated into a unified interface for accessing parallel storage and secondary storage using the Python language; the parallel storage is accessed through the Lustre file system API interface, and the secondary storage is accessed through the NFS protocol interface and the NAS gateway interface.
3. The architecture system for remote access and management of high-performance computing data according to claim 2, characterized in that, The parallel storage interface and secondary storage interface are encapsulated into a unified abstract interface using Python. Specifically, a configuration file is generated using the `def read_file(storage_type, path)` or `def write_file(storage_type, local_path, remote_path)` language to store the configuration parameters corresponding to remote access to secondary and parallel storage. The access commands for the parallel and secondary storage interfaces are abstracted into a unified command interface using Linux or native UNIX access methods. This allows a single command to access both parallel and secondary storage interfaces simultaneously in a high-performance computer system. The configuration parameters include the access node's IP address or domain name, the secondary storage root directory path `nas_path`, the username `user`, and the password `password`. Linux or native UNIX commands include `cd`, `df`, `cp`, `is`, `rm`, `chmod`, `get`, `put`, and `mv`.
4. The architecture system for remote access and management of high-performance computing data according to claim 1, characterized in that, The configuration management layer stores configuration parameters in the configuration information storage file `cmaConfig` and uses the configuration script file `pycmafs.py` to read and modify these parameters. When a user configures access to parallel storage and secondary storage, the configuration script file `pycmafs.py` first verifies the MD5 encoding of the configuration information storage file `cmaConfig`. If they are inconsistent, the old configuration information storage file `cmaConfig` is replaced with the one generated by the configuration script file `pycmafs.py`. This verifies the validity of the user-input access node IP address or domain name and the secondary storage root directory path `nas_path`. If the configuration parameters are incorrect, a clear prompt is given to ensure the correctness of the configuration parameter information. The configuration parameters include the access node IP address or domain name, the secondary storage root directory path `nas_path`, the username `user`, and the password `password`. The configuration parameters are automatically written to the configuration information storage file `cmaConfig` after the user completes the configuration.
5. The architecture system for remote access and management of high-performance computing data according to claim 1, characterized in that, In the command-line tool layer, the standardized command set module includes 14 standardized commands. Each standardized command is named with the prefix "cma", and each standardized command supports the "-help" parameter. Each standardized command defines explicit parameter rules and supports positional parameters and optional parameters. The command parsing module uses the argparse module to parse the 14 standardized command parameters in the standardized command set module. It is used for registering the 14 standard commands, parsing parameters, validating syntax, and generating help documentation. The command parsing module extracts the standardized command name from the user's input command and sends the information containing the standardized command name to the parallel storage or secondary storage, which then performs the corresponding operation. The 14 standardized commands include: (1) cma: Used to list all commands supported by the system and a brief description of their functions; (2) cmacd: Moves the remote working directory, supports seamless switching between parallel storage and secondary storage directories, without the need to manually distinguish the storage path; the calling format is: cmacd [parameter] [path]; (3) cmals: View files under a specific path on a remote server. Supports multiple parameters and displays detailed file information. The calling format is: cmals [parameters] [path]. (4) cmapwd: Get the current remote working directory path, automatically distinguish the storage type, and return path information in a uniform format; the calling format is: cmapwd [parameters]; (5) cmaput: Uploads local files or directories to a remote server, i.e., parallel storage or secondary storage. It supports recursive upload, compressed upload, permission preservation, and resume upload. The calling format is: cmaput [parameters] [local path] [target path]. (6) cmaget: Downloads files or directories from a remote server to the local machine. It supports recursive download, compressed download, permission preservation, and resume download. The parameters are consistent with cmaput, reducing the learning cost. The calling format is: cmaget [parameter] [target path] [local path]. (7) cmarm: Deletes files or directories on a remote server. It supports recursive deletion, confirmation deletion with the -i parameter, and forced deletion with the -f parameter to avoid accidental operations. The calling format is: cmarm [parameter] [path]; (8) cmamv: Move files or directories on the remote server. It supports the overwrite prompt -i parameter and the forced overwrite -f parameter. It skips non-existent files and does not display redundant information. The calling format is: cmamv [parameter] [file path] [target path]; (9) cmamkdir: Creates a directory on a remote server, supports recursive creation with the -p parameter to ensure complete directory hierarchy; the calling format is: cmamkdir [parameters] [path]; (10) cmatouch: Creates or modifies files on a remote server, supports multiple time configuration parameters; the calling format is: cmatouch [parameters] [file / directory]; (11) cmalocate: Searches for documents that meet the conditions on a remote server. It supports regular expression matching, case-insensitive matching, and matching quantity limit parameters to improve search efficiency. The calling format is: cmalocate [parameter] [matching mode]; (12) cmacp: Copy files or directories on a remote server. It supports recursive copying, permission preservation, and overwrite prompts, and is suitable for remote file backup needs. The calling format is: cmacp [parameters] [source file] [target file]; (13) cmachmod: Modifies the permissions of files or directories on a remote server. It supports recursive modification with the -R parameter and permission change prompts with the -c parameter, and is compatible with multi-user permission management. The calling format is: cmachmod [parameter] [permission mode] [file / directory]; (14) cmadf: Statistical analysis of disk usage in remote file systems. It supports multiple display formats (GB / MB / KB) and specifies storage type statistical parameters, which is convenient for operation and maintenance monitoring. The calling format is: cmadf [parameters] [file / directory].
6. The architecture system for remote access and management of high-performance computing data according to claim 5, characterized in that, The user interaction layer includes a command-line interaction module, a modal job flow module, and a help documentation module; among which: The command-line interaction module supports terminal command input and real-time return of execution results; The pattern job flow module provides interfaces for 14 standardized commands; these standardized command interfaces are directly embedded in the job flow script, supporting the automatic invocation of the 14 standardized commands prefixed with cma by the job. The help documentation module provides support for the -help parameter for each standardized command. The -help parameter displays the function, calling format, and parameter description of the standardized command.
7. A deployment method for a high-performance computing data remote access and management architecture system as described in any one of claims 1-6, characterized in that, In the global deployment of system users, it includes: Step 11: Obtain the system installation package cmas.tar.gz and upload it to the specified directory on the HPC cluster server; Step 12: Unzip the installation package. Execute the command "tar -zxvf cmas.tar.gz -C / opt / cmas" to unzip the tar package to the / opt / cmas directory, and then switch to the unzipped directory cd / opt / cmas; Step 13: Configure storage information. Enter the . / config directory cd / opt / cmas / config, execute the command "pythonpycmafs.py", and enter the configuration parameters according to the terminal prompts. The configuration parameters include the access node IP address, the secondary storage name and the root directory path nas_path. Confirm after entering, and the configuration parameters will be automatically written to the configuration information storage file cmaConfig. Step 14: Configure system environment variables. Edit the / etc / profile file and add two lines at the end of the file: "export CMA_HOME= / opt / cmas" and "export PATH=$PATH:$CMA_HOME / bin"; Step 15: Make the environment variables take effect by executing the command "source / etc / profile"; Step 16: Deployment verification. Enter "cma" in the terminal. If a list of all commands supported by the system is returned, the global deployment is complete, and all users can use the cma standardized command set in the terminal.
8. A deployment method for a high-performance computing data remote access and management architecture system as described in any one of claims 1-6, characterized in that, In a single-user deployment, this includes: Step 21: Obtain the system installation package cmas.tar.gz and upload it to the home directory of this single user, / home / xxx, where xxx is the username; Step 22: Unzip the installation package. Execute the command "tar -zxvf cmas.tar.gz -C / home / xxx / cmas" to unzip the tar package to the cmas directory under the user's home directory, and then switch to the unzipped directory cd / home / xxx / cmas; Step 23: Configure storage information. Edit the configuration information storage file cmaConfig, i.e., vi / home / xxx / cmas / cmaConfig, or execute the command "python pycmafs.py". Enter the configuration parameters as prompted. The configuration parameters include the access node IP address, username user, password password, secondary storage name and root directory path nas_path. Save the configuration after configuration. The configuration parameters will be automatically written to the configuration information storage file cmaConfig. Step 24: Configure system environment variables. Edit the ~ / .bash_profile file and add two lines at the end of the file: "export CMA_HOME= / home / xxx / cmas" and "export PATH=$PATH:$CMA_HOME / bin"; Step 25: Make the environment variables take effect by executing the command "source ~ / .bash_profile"; Step 26: Deployment verification. Enter "cma" in the terminal. If a list of all commands supported by the system is returned, the single-user deployment is complete. Only that single user can use the cma standardized command set in the terminal.
9. A deployment method for a high-performance computing data remote access and management architecture system as described in any one of claims 1-6, characterized in that, In module deployment, the following are included: Step 31: The administrator completes the setting of the module environment variables and writes the CMAFS environment configuration into the module file; Step 32: The administrator configures the environment variables for the module by adding CMA_HOME and PATH configurations to the module file; Step 33: The user checks all software environment variables and executes the command "module avail". If the output contains "apps / cmafs / cmafs", then the module environment configuration is complete. Step 34: The user imports the cmafs environment variables and executes the command "module load apps / cmafs / cmafs". After loading is complete, proceed to the next step to continue execution. Step 35: Deployment verification. Enter "cma". If a list of commands is returned, the module deployment is complete. Users can directly use the cma standardized command set to uninstall the environment by executing "module unload apps / cmafs / cmafs".
10. The application method of a high-performance computing data remote access and management architecture system as described in any one of claims 1-6, characterized in that, Includes the following steps: Step A: Loading the user's operating environment: The user logs into the HPC node and selects the corresponding environment loading method according to the deployment method: Global deployment and single-user deployment do not require additional loading and can be done directly using the terminal; Module deployment requires executing "module loadapps / cmafs / cmafs" to load the environment; Step B: User Operation Configuration Initialization: Upon first use, the user executes "python pycmafs.py" and enters configuration parameters including the storage node IP, storage path, and authentication information as prompted. After configuration, the configuration parameters are automatically saved to the configuration information storage file cmaConfig. Step C: User operation command input: The user enters the standardized command set related to CMA in the terminal, along with the corresponding parameters, and presses Enter to submit the command after completing the input; Step D: System command parsing and execution: The command line tool layer receives the command input by the user, parses the command, then calls the configuration management layer to obtain the storage configuration information, and then accesses the corresponding parallel storage or secondary storage through the unified interface of the underlying adaptation layer to perform the corresponding operation; Step 5: System feedback results: After the system completes the execution, the results will be returned to the terminal in real time. The user can judge whether the operation was successful based on the results. If it fails, the user can modify the command or configuration according to the prompts.