An in-band and out-of-band cooperative operation and maintenance management system for supercomputing clusters

By employing technologies such as minimum weight priority scheduling, consistent hashing, digital twin algorithms, and knowledge graphs in the operation and maintenance of supercomputing clusters, the problems of data heterogeneity between in-band and out-of-band and unbalanced task scheduling have been solved. This has enabled efficient collaborative operation and maintenance with unified data in time and space and fault response, thereby improving operation and maintenance efficiency and energy utilization efficiency.

CN121283868BActive Publication Date: 2026-07-21JIANGSU TIN TIE HUITONG TECHNOLOGY CO LTD
2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU TIN TIE HUITONG TECHNOLOGY CO LTD
Filing Date
2025-09-22
Publication Date
2026-07-21

Smart Images

  • Figure CN121283868B_ABST
    Figure CN121283868B_ABST
Patent Text Reader

Abstract

The present application relates to the field of supercomputer cluster operation and maintenance, in particular to an in-band and out-of-band collaborative operation and maintenance management system for supercomputer clusters, which comprises data acquisition, processing, collaborative decision-making and execution control modules; the data acquisition module allocates task queues using the minimum weight priority scheduling algorithm, matches tasks and devices using the consistent hashing algorithm, and acquires in-band and out-of-band parameter states; the data processing module maps data to the twin base through the digital twin algorithm, introduces spatiotemporal correlation anomaly detection, supercomputer operation and maintenance knowledge graph and entity linking algorithm, and forms a global data set unified in space, time and semantics; the collaborative decision-making module adjusts the in-band and out-of-band weights based on a dynamic weight model, constructs energy consumption prediction and fault risk assessment functions, and optimizes them through reinforcement learning; the execution control module generates inspection tasks according to risk classification, plans paths using an improved A* algorithm, optimizes air conditioner load and server power in combination with energy consumption prediction, and improves operation and maintenance efficiency and energy utilization efficiency.
Need to check novelty before this filing date? Find Prior Art