Autonomous OS Deployment in High-Performance Computer Nodes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for changing the operating system in high-performance computer service nodes are cumbersome and prone to operational disruptions, especially when preparing storage disks, as all operations must be synchronized and any errors can render nodes non-operational.

Innovation Solution

A control method and device that allows for autonomous deployment of a new tree-type node software image by transferring a reduced version of the operating system, boot kernel, and launch module to service nodes, which then locally install and reboot, ensuring that nodes remain operational even if errors occur during instantiation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If all operations are synchronized on selected service nodes, then deployment control is improved, but operation complexity and difficulty of monitoring increase

Engineering Contradiction:
Improvedeployment controlVSAvoidoperation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent divides the deployment process into distinct phases: preparation phase (defining reduced version, boot kernel, launch module) and execution phase (local installation on service nodes). This segmentation allows centralized control of deployment parameters while enabling distributed autonomous execution, reducing operational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-defining the reduced version of the operating system, boot kernel, and launch module before deployment. These preparation steps are completed in advance on the management node, allowing the actual deployment to proceed more smoothly with reduced real-time complexity.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If storage disks are prepared early in the process, then deployment efficiency is improved, but node operational reliability deteriorates as errors can render nodes non-operational

Engineering Contradiction:
Improvedeployment efficiencyVSAvoidnode operational reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent prepares the reduced version of the operating system in advance as a cushioning measure. This reduced version serves as a fallback that can be deployed if the full deployment process fails, ensuring that nodes remain operational even when storage disk preparation or subsequent steps encounter errors.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

The launch module acts as an intermediary between the boot kernel and the full operating system installation. It provides a controlled interface that can manage the deployment process and revert to previous states if errors occur, maintaining node reliability while enabling efficient deployment.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Extent of automation

If a reduced version of the operating system is transferred and installed locally, then deployment autonomy is improved, but communication network dependency increases

Engineering Contradiction:
Improvedeployment autonomyVSAvoidnetwork coordination complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent segments the operating system into a reduced version that can be deployed autonomously and a complete version that follows. This allows the initial deployment to proceed with minimal network coordination, improving automation while managing network dependency through phased installation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3123314B1Method and device for controlling the change in operating system in service nodes of a high-performance computer
Publication Date: 2024.05.01 BULL SA
  • EP3123314B1 patent drawingFigure 1~2

AI summary

A method controls the change in operating system in selected service nodes (NH-NNM(N)) of a high-performance computer (CHP). Said method includes: - a step (i) of defining, for the selected service nodes, a reduced version of a new operating system to be installed, a boot kernel, a so-called "reference" tree node software image suitable for the new operating system and comprising a definition of an instantiation to be established in the service nodes, and an activation module (ML) capable of locally installing the reference image in each service node; - a step (ii) wherein the defined reference image, boot kernel, activation module, and reduced operating system version are transferred into the service nodes; and - a step (iii) wherein the transferred activation module (ML) is used in each service node in order to locally install the transferred reference image.