Autonomous OS Deployment in High-Performance Computer Nodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for changing the operating system in high-performance computer service nodes are cumbersome and prone to operational disruptions, especially when preparing storage disks, as all operations must be synchronized and any errors can render nodes non-operational.
Innovation Solution
A control method and device that allows for autonomous deployment of a new tree-type node software image by transferring a reduced version of the operating system, boot kernel, and launch module to service nodes, which then locally install and reboot, ensuring that nodes remain operational even if errors occur during instantiation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If all operations are synchronized on selected service nodes, then deployment control is improved, but operation complexity and difficulty of monitoring increase
Solution Approach 1:
The patent divides the deployment process into distinct phases: preparation phase (defining reduced version, boot kernel, launch module) and execution phase (local installation on service nodes). This segmentation allows centralized control of deployment parameters while enabling distributed autonomous execution, reducing operational complexity.
Solution Approach 2:
The patent performs preliminary actions by pre-defining the reduced version of the operating system, boot kernel, and launch module before deployment. These preparation steps are completed in advance on the management node, allowing the actual deployment to proceed more smoothly with reduced real-time complexity.
2Productivity
If storage disks are prepared early in the process, then deployment efficiency is improved, but node operational reliability deteriorates as errors can render nodes non-operational
Solution Approach 1:
The patent prepares the reduced version of the operating system in advance as a cushioning measure. This reduced version serves as a fallback that can be deployed if the full deployment process fails, ensuring that nodes remain operational even when storage disk preparation or subsequent steps encounter errors.
Solution Approach 2:
The launch module acts as an intermediary between the boot kernel and the full operating system installation. It provides a controlled interface that can manage the deployment process and revert to previous states if errors occur, maintaining node reliability while enabling efficient deployment.
3Extent of automation
If a reduced version of the operating system is transferred and installed locally, then deployment autonomy is improved, but communication network dependency increases
Solution Approach 1:
The patent segments the operating system into a reduced version that can be deployed autonomously and a complete version that follows. This allows the initial deployment to proceed with minimal network coordination, improving automation while managing network dependency through phased installation.
Data Source
Figure 1~2
AI summary
A method controls the change in operating system in selected service nodes (NH-NNM(N)) of a high-performance computer (CHP). Said method includes: - a step (i) of defining, for the selected service nodes, a reduced version of a new operating system to be installed, a boot kernel, a so-called "reference" tree node software image suitable for the new operating system and comprising a definition of an instantiation to be established in the service nodes, and an activation module (ML) capable of locally installing the reference image in each service node; - a step (ii) wherein the defined reference image, boot kernel, activation module, and reduced operating system version are transferred into the service nodes; and - a step (iii) wherein the transferred activation module (ML) is used in each service node in order to locally install the transferred reference image.