Cluster equipment health evaluation management method and equipment and readable computer storage medium

By acquiring fault work order data to calculate health indicators, the health status of devices and clusters is generated, solving the problem of health assessment of cluster devices and realizing global cognition and efficient operation and maintenance decision-making.

CN121414232APending Publication Date: 2026-01-27ZHEJIANG LAB
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511988874.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing technologies cannot effectively assess the health status of cluster devices, making it difficult for operators to accurately allocate maintenance resources and make optimal hardware procurement strategies, thus affecting business continuity.

Method used

By acquiring fault work order data, multiple health indicators (such as equipment online rate, number of faults, average repair time, etc.) are calculated to generate equipment health status, and the status of equipment within the cluster is aggregated to output the cluster health status.

Benefits of technology

It enables a global health assessment of cluster devices, shortens fault location time, provides data-driven operational and maintenance decision-making basis, and improves operational and maintenance efficiency and the accuracy of device health assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121414232A_ABST
    Figure CN121414232A_ABST
Patent Text Reader

Abstract

The invention provides a cluster equipment health evaluation management method and device and a readable computer storage medium, and the method comprises the steps: obtaining fault work order data of target equipment in a cluster, calculating at least one health index used for evaluating the operation and maintenance health degree of the target equipment based on the fault work order data, discrete and passive fault data are converted into quantitative evaluation on long-term reliability and maintenance requirements of equipment. And the equipment health state of the target equipment is generated according to the at least one health index, a clear health conclusion is provided for each piece of target equipment, and the accuracy of equipment health state evaluation is improved through comprehensive evaluation of the health indexes of multiple dimensions. Afterwards, the device health states of the devices in the cluster are aggregated, the cluster health state is generated and output, the devices in the cluster do not need to be checked one by one, the fault positioning time is greatly shortened, meanwhile, a user has global cognition on the health situation of the macroscopic cluster, and a direct data-driven decision basis is provided for subsequent operation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of system integration technology, and in particular to a method, device and readable computer storage medium for health evaluation and management of cluster equipment. Background Technology

[0002] The Wanka cluster is a complex system engineering project. During operation, both software and hardware failures can occur, especially hardware failures. When failures occur, on-site maintenance personnel are required for inspection and repair. During maintenance, we have found that some devices, such as the head unit or modules, experience multiple failures, and even multiple devices experiencing coordinated failures. Currently, there are no mature and effective methods for evaluating the equipment. This makes it difficult for operators to accurately grasp the health status of the cluster from a macro perspective, and to make optimal data-driven decisions regarding the precise allocation of maintenance resources, hardware procurement strategies, and business continuity assurance. Summary of the Invention

[0003] To address the aforementioned technical problems, this application provides the following technical solutions: Firstly, a method for health assessment and management of cluster devices is provided, the method comprising: Obtain fault work order data of target devices in the cluster. The fault work order data records the operation and maintenance events of the device from the occurrence of the fault to the completion of the repair. Based on the work order data of the aforementioned fault, calculate at least one health indicator for evaluating the operational health of the target equipment. The device health status of the target device is generated based on the at least one health indicator; Aggregate the device health status of each device in the cluster, generate and output the cluster health status.

[0004] According to the health assessment and management method for cluster devices provided in this application, before obtaining the fault work order data of the target device in the cluster, the method further includes: creating a fault work order; The creation of the fault work order includes: In response to an abnormal alarm received from the monitoring system, a fault work order is automatically created and stored; wherein the fault work order contains fault data based on the abnormal alarm.

[0005] According to the cluster device health assessment management method provided in this application, the creation of the work order includes: In response to a user-triggered work order creation operation, create a framework for the fault work order and display the corresponding interactive interface; The fault data entered through the interactive interface is filled into the frame of the fault work order to generate a complete fault work order.

[0006] According to the cluster device health assessment and management method provided in this application, after creating a fault work order and before obtaining the fault work order data of the target device in the cluster, the method further includes: In response to the user's editing operation, update the device status and repair information of the target device associated with the fault work order; and / or, In response to the closing command for the fault work order, the status of the fault work order is marked as closed.

[0007] According to the health evaluation and management method for cluster devices provided in this application, the health status of the cluster or the health status of at least one device in the cluster is displayed through a visual interface, and the visual interface includes at least a cluster arrangement area. The method further includes: At least one cluster identifier is generated and displayed in the cluster arrangement area; wherein the cluster identifier is used to indicate the health status information of the cluster, and the health status information is determined based on the device health status of each device in the corresponding cluster.

[0008] According to the health assessment and management method for cluster devices provided in this application, the method further includes: Based on the health level represented by the health status information of the cluster, the identifiers of each cluster are displayed differently in the cluster arrangement area.

[0009] According to the health assessment and management method for cluster equipment provided in this application, the visualization interface further includes a rack expansion area; The method further includes: Real-time detection of touch operations received by the cluster identifier corresponding to the target cluster, and generation of a query command corresponding to the touched cluster identifier; In response to the query command, the distribution of each device in the target cluster is displayed in the rack expansion area, and each device in the target cluster is displayed differently based on its corresponding device health status.

[0010] According to the cluster device health evaluation and management method provided in this application, the visualization interface further includes a device information area; The method further includes: In response to the selection of the target device in the rack expansion area, at least one health indicator of the target device's health status is displayed in the device information area.

[0011] According to the health evaluation and management method for cluster equipment provided in this application, the health indicators include at least one of the following: equipment online rate, number of equipment failures, mean time to repair, mean time between failures, and mean time to failure; and / or, The fault work order data includes at least one of the following: equipment failure time, repair time, and fault type.

[0012] Secondly, a product application is provided for implementing the cluster device health assessment and management method as described in any of the first aspects.

[0013] Thirdly, an electronic device is provided, including a memory, a processor, and a cluster device health assessment management program stored in the memory and executable on the processor, wherein the processor, when executing the cluster device health assessment management program, implements the cluster device health assessment management method as described in any one of the second aspects.

[0014] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the program is executed by a processor, it performs the cluster device health evaluation and management method described in any one of the first aspects.

[0015] The cluster device health assessment and management method, device, and computer-readable storage medium in the embodiments of this application have the following beneficial effects: The system acquires fault work order data for target devices within the cluster. Based on this data, it calculates at least one health indicator to evaluate the operational health of the target devices, transforming discrete, reactive fault data into a quantitative assessment of long-term device reliability and maintenance requirements. Then, it generates the device health status of each target device based on the at least one health indicator, providing a clear health conclusion for each device. Furthermore, the comprehensive evaluation of multiple health indicators improves the accuracy of device health status assessment. Subsequently, it aggregates the device health statuses of all devices within the cluster to generate and output the cluster health status. This eliminates the need to check each device individually, significantly shortening fault location time and providing users with a global understanding of the cluster's overall health, offering direct data-driven decision-making support for subsequent operations and maintenance.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0018] Figure 1 This is a flowchart illustrating a health assessment and management method for cluster devices provided in this application; Figure 2 This is a schematic diagram of a computer-readable storage medium provided in this disclosure; Figure 3 This is a schematic diagram of the structure of a computing device provided in this disclosure. Detailed Implementation

[0019] This application describes multiple technical solutions with different concepts. Each concept has one or more embodiments, and these different concepts can be combined to form more embodiments. Those skilled in the art, after reading this application, can combine these different concepts to obtain new technical solutions, and these new technical solutions should also fall within the scope of this application.

[0020] The technical solutions of these different concepts will be introduced in turn below. Some concepts may appear in multiple technical solutions of different concepts. For these concepts, this article will explain them when they first appear and will not repeat them in the following text.

[0021] Existing cluster device health management primarily relies on metrics or models to automatically connect and disconnect devices from the perspectives of fault detection and automatic recovery. The focus is on ensuring business stability, but it lacks a holistic analysis based on actual device fault and repair information from the perspective of device maintenance. This prevents operators from accurately grasping the cluster's health status at a macro level, hindering optimal data-driven decisions regarding the precise allocation of maintenance resources, hardware procurement strategies, and business continuity assurance.

[0022] To address the aforementioned technical issues, this application provides a method for health assessment and management of cluster devices.

[0023] The system aims to collect a large amount of equipment fault repair records (i.e., fault work order data) and calculate multiple health indicators to achieve a comprehensive assessment of equipment health, solving the current problem of focusing only on business stability while ignoring the equipment's own operational data. Simultaneously, it aggregates and elevates these discrete, equipment-level operational events into a unified, global understanding of the overall health of the cluster. This provides a holistic understanding of the health of cluster equipment, effectively enabling optimal data-driven decisions for subsequent precise allocation of operational resources, hardware procurement strategies, and business continuity assurance.

[0024] This application provides an embodiment of a health assessment and management method for cluster devices, referring to... Figure 1 , Figure 1 This is a flowchart illustrating a cluster device health assessment and management method provided in an embodiment of this application.

[0025] The cluster device health assessment and management method includes: Step S100: Obtain fault work order data of the target device in the cluster. The fault work order data records the operation and maintenance events of the device from the occurrence of the fault to the completion of the repair. Step S200: Based on the fault work order data, calculate at least one health indicator for evaluating the operational health of the target equipment. Step S300: Generate the device health status of the target device based on the at least one health indicator; Step S400: Aggregate the device health status of each device in the cluster, generate and output the cluster health status.

[0026] In this embodiment, a 1000-card cluster refers to a cluster composed of GPU cards, which includes at least one device.

[0027] Equipment refers to the hardware devices in a multi-card cluster, including but not limited to GPU cards, compute node hardware, network equipment, and environmental facilities. Equipment failures cover three main categories: machine head, modules, and environment, further subdivided into 20 sub-items such as CPU, network card, GPU, memory, hard drive, network cable, and environmental stability. Different failure types are assigned to different responsible units and have different repair cycles. Therefore, hardware failures in the computing cluster require continuous tracking and prompting of the corresponding nodes to address them in order to minimize losses. The longer the repair time, the greater the unit loss. Therefore, a complete cluster equipment health assessment system is needed to comprehensively understand the health status of the equipment, detect failures in a timely manner, and improve the repair efficiency of faulty equipment as much as possible.

[0028] In this embodiment, the cluster device health evaluation management method is applied to the cluster device health evaluation management system, hereinafter referred to as "the system" or "this system".

[0029] In step S100, fault work order data of the target device in the cluster is obtained. The fault work order data records the operation and maintenance events of the device from the occurrence of the fault to the completion of the repair.

[0030] In this embodiment, fault work order data represents the basic information of a work order used to reflect the fault status of the equipment when a fault occurs. The basic information includes, but is not limited to, fault information and basic information of the faulty equipment.

[0031] In this embodiment, the fault information includes at least one of the following: equipment failure time, repair time, and fault type.

[0032] The target device refers to the current device among multiple devices in the cluster that is being evaluated for health.

[0033] In this embodiment, a fault ticket is created before acquiring fault ticket data for the target device in the cluster. The created fault ticket can then be provided to the user as fault ticket data.

[0034] In particular, when a device failure occurs during the operation of the Wanka cluster (any situation other than normal device operation can be considered a device failure), the system has built a dual-mode work order creation mechanism to meet the needs of rapid response in different scenarios.

[0035] The ways to create a work order are as follows: Method 1: In response to an abnormal alarm received from the observation system, a fault work order is automatically created and stored; wherein, the fault work order contains fault data based on the abnormal alarm.

[0036] Leveraging the automation capabilities of the observation system, when an equipment malfunction is detected, the system immediately invokes its own work order creation interface. Based on the detected fault data and equipment information, a fault work order is generated.

[0037] The device information includes, but is not limited to, the device's Serial Number (SN, a unique identifier used for device authentication; in large-scale deployment environments such as multi-card clusters, recording and maintaining the serial number of each device allows the operations and maintenance team to effectively monitor device status and quickly locate problematic devices), the corresponding module manufacturer, and device status. Basic device information is automatically entered when a fault ticket is created.

[0038] Method 2: In response to a user-triggered work order creation operation, create a framework for a fault work order and display the corresponding interactive interface; fill the framework of the fault work order with the fault data entered through the interactive interface to generate a complete fault work order.

[0039] The system also supports users manually creating fault work orders. When on-site users (such as maintenance personnel, administrators, and other personnel with relevant permissions) discover anomalies during inspections, they can use the system's interactive interface to input required fields such as the device serial number and fault type. This input will then fill the fault data into the fault work order frame, and clicking the "OK" button will trigger the work order generation process. The fault data entered by the user and the fields in the fault work order frame are correlated to ensure the accuracy of the generated work orders.

[0040] After creating a fault work order using the above method, the user maintains the fault information of the Wanka cluster device in the system.

[0041] For example, after creating the fault work order and before obtaining the fault work order data of the target device in the cluster, the method further includes: In response to the user's editing operation, update the device status and repair information of the target device associated with the fault work order.

[0042] To ensure that the fault data recorded on fault work orders is consistent with the actual physical state of the equipment, the system has established a linkage mechanism between fault work orders and the status of the target equipment. When a user performs an editing operation through the system interface, the system not only updates the work order content but also corrects the core status of the target equipment associated with the work order. The core status includes, but is not limited to, the equipment status and repair information of the target equipment.

[0043] For example, in the work order editing interface, users will record key repair information, such as the root cause of the fault (e.g., PU card damage, power module failure, etc.), the repair actions performed (e.g., component replacement, firmware upgrade), repair time, and repair personnel information.

[0044] At the same time, the system provides a status selector (as shown in the drop-down menu) to record the device status. For example, a status change could be the switching of the device from a "faulty" or "offline" status to a "normal" or "online" status.

[0045] The data in the aforementioned fault work orders will be used for subsequent equipment health assessments. For example, if the work order records the fault occurrence event and the repair completion time, the average repair time and equipment downtime can be calculated based on these two parameters.

[0046] In this embodiment, after creating a fault work order and before obtaining the fault work order data of the target device in the cluster, the method further includes: In response to the closing command for the fault work order, the status of the fault work order is marked as closed.

[0047] Once equipment repair is completed and verified to be error-free, authorized users (such as maintenance personnel or team leaders) can execute a "close work order" operation in the system. In response to this closure command, the system updates the work order's status field from "incomplete" (e.g., "processing" or "pending verification") to "closed." Closing a work order is a confirmation action, marking the completion of this maintenance event loop and triggering the system to mark the faulty work order data as a signal usable for health assessment analysis.

[0048] Only closed work orders are captured and processed by the subsequent health assessment system's data collection module. This avoids including incomplete or still-processing maintenance events in the calculations, ensuring the accuracy of health indicators.

[0049] In step 200, based on the fault work order data, at least one health indicator is calculated to evaluate the operational health of the target equipment.

[0050] In this embodiment, the original, discrete fault records are transformed into quantifiable indicators for assessing equipment reliability and maintenance requirements.

[0051] For example, based on fault work order data for target devices in the cluster, multiple health indicators for evaluating the operational health of the target devices are calculated using a predefined algorithm model.

[0052] In this embodiment, the health indicators include at least one of the following: equipment online rate, number of equipment failures, mean repair time, mean time between failures, and mean time to failure.

[0053] The following health indicators (including but not limited to) can be calculated using a predefined algorithm model: Equipment online rate: Within a specified time range, (Presidential time length - Total downtime) / Presidential time length * 100%.

[0054] The total downtime is the sum of the repair times for all fault work orders for the equipment. This metric comprehensively reflects the availability of the equipment.

[0055] Equipment failure frequency: The number of failure work orders for the target equipment within a specified time range. This indicator directly reflects the frequency of equipment failures.

[0056] Average number of work orders for faulty equipment = Total number of work orders for faulty equipment / Total number of faulty equipment.

[0057] Relative average (percentage) = (Target equipment failure work order data - average number of failure equipment work orders) / average number of work orders in the entire cluster.

[0058] Relative average (absolute value) = Current equipment failure work order data - Average number of failure equipment work orders.

[0059] Mean Time To Repair (MTTR): Within a specified time range, AVG (time to normal machine status - time to abnormal machine status).

[0060] For example, calculate the average repair time for all fault work orders. Repair time = Fault end time - Fault start time. This metric measures how quickly equipment returns to normal; a lower value indicates better maintainability.

[0061] Equipment downtime: within a specified time range, SUM (time when machine status changes to normal - time when machine status changes to abnormal).

[0062] Mean Time To Failure (MTTF): The average fault-free runtime of all work orders within a specified time range. TTF = Start time of the current fault - End time of the previous fault. For the first fault, the calculation can start from the beginning of the statistics. This metric measures equipment reliability; a higher value indicates greater stability.

[0063] Mean Time Between Failures (MTBF) = Mean Time to Repair (MTTR) + Mean Time Without Failure (MTTF)

[0064] In step S300, the device health status of the target device is generated based on the at least one health indicator.

[0065] In this embodiment, multiple quantitative indicators are combined into a single, easy-to-understand, and easy-to-operate device-level health conclusion.

[0066] For example, the system has a built-in state mapping engine that maps the above set of health indicators into a comprehensive device health status according to preset rules. This process can be as follows: Method 1: Set thresholds for each health indicator. For example, if the number of failures is >3 and the mean time between failures (MTTF) is <100 hours, the equipment health status is judged as "sub-healthy"; if the number of failures is >10, it is judged as "failed"; otherwise, it is "healthy".

[0067] Method 2: Assign weights to each indicator and calculate a comprehensive score (e.g., a percentage system), then determine the equipment health status based on the score range. For example, 90-100 points is "healthy", 60-89 points is "sub-healthy" (there are abnormalities but no obvious impact on equipment operation), and below 60 points is "faulty".

[0068] Through the above process, the health status of the target device can be output based on at least one health indicator. This device health status is a structured data object, which at least includes the device ID and status level (such as healthy / sub-healthy / faulty, or a specific score), and it serves as the direct basis for performing precise device-level maintenance.

[0069] In step S400, the health status of each device in the cluster is aggregated, and the cluster health status is generated and output.

[0070] In this embodiment, the health status of the cluster is determined based on the health status of each device in the cluster, thereby achieving a global understanding of operation and maintenance management.

[0071] The system can aggregate device health status data for all devices within the target cluster in the following ways: Method 1: Calculate the number and percentage of devices in the cluster that are in "healthy", "sub-healthy", or "faulty" states.

[0072] Method 2: If the device health status is a score, then calculate the average health score of the cluster.

[0073] Method 3: The health status of the cluster is determined by the N devices with the worst health or the devices that are faulty. For example, if there is a device in the cluster that is in a "faulty" state, the health status of the cluster is "abnormal".

[0074] In this embodiment, the device health status and cluster health status are output as a structured data object (such as JSON) for other business systems (such as resource scheduling systems and operation and maintenance management platforms) to call.

[0075] Through the above embodiments, on the one hand, the cluster device health evaluation and management system can uniformly evaluate the cluster devices by accumulating fault data. On the other hand, the operation and maintenance party can achieve precise maintenance of the devices and reduce operation and maintenance costs by using different operation and maintenance plans (such as personnel allocation, daily inspection frequency settings, etc.) for devices with different fault frequencies.

[0076] On the other hand, through a unified equipment evaluation system, operators can have a comprehensive understanding of the health of the cluster. In particular, for intelligent computing clusters, relevant contingency plans can be made in various aspects such as business integration and hardware and software procurement, thereby supporting business continuity, security and sustainable development.

[0077] In this embodiment, the health status of the cluster or the health status of at least one device in the cluster is displayed through a visual interface, which includes at least a cluster arrangement area, a rack expansion area, and a device information area. In some examples, the order of viewing is: cluster arrangement area, rack expansion area, and device information area.

[0078] The method further includes: At least one cluster identifier is generated and displayed in the cluster arrangement area; wherein the cluster identifier is used to indicate the health status information of the cluster, and the health status information is determined based on the device health status of each device in the corresponding cluster.

[0079] In this embodiment, the cluster identifier is used to refer to the cluster, and its style can be an icon similar to the shape of the device.

[0080] In this embodiment, the arrangement of the cluster arrangement areas can follow a specific layout strategy to reflect the logical relationships of the underlying infrastructure. It can be: Topology mapping: Arranged according to the logical view of the physical data center. Each cluster identifier represents an independent cluster within a data center. The relative positions of the cluster identifiers can reflect the row and column layout of the racks within the data center or the network topology between clusters, enabling maintenance personnel to quickly locate the physical clusters.

[0081] Health Status Sorting: Provides a view sorted by health status, automatically arranging all "abnormal" clusters at the top or left of the view for priority display, ensuring that critical issues can be detected immediately.

[0082] The generation of each cluster identifier in the cluster arrangement area is not based on a single data point, but rather on the result of real-time aggregation calculation of the health status of all devices within the cluster it represents.

[0083] In this embodiment, based on the health level represented by the health status information of the cluster, the identifiers of each cluster are displayed differently in the cluster arrangement area.

[0084] For example, the cluster arrangement area displays 6 clusters, namely a1, b1, b2, b3, c1, and c2. Among them, b1, b2, and b3 are clusters with the same health status (classified into the same level in the health status evaluation level or located within the same health status range).

[0085] Similarly, c1 and c2 are clusters with the same health status.

[0086] If cluster health status of a1 is evaluated as "faulty", the cluster identifier is rendered in red. If cluster health status of b1, b2, and b3 is evaluated as "healthy", the cluster identifier is rendered in green. Similarly, if cluster health status of c1 and c2 is evaluated as "sub-healthy" (sub-healthy devices exist, but no faulty devices), the cluster identifier is rendered in yellow.

[0087] For example, to facilitate differentiation, only two colors are used to distinguish the cluster identifiers. If the cluster health status is divided into "normal" and "fault", then red is used to render the identifier of the "fault" cluster, while the "normal" cluster can be rendered with a different color than red, or it can remain in its initial state.

[0088] In this embodiment, based on the health level represented by the health status information of the clusters, each cluster identifier is displayed differently in the cluster arrangement area. Differentiation can also include differences in the shape and display mode of the cluster identifiers.

[0089] For example, different identification shapes can be assigned to clusters of different health levels. This method can be used alone or in combination with color coding to improve recognition and assist color-disabled users.

[0090] For example, clusters that are "faulty" are identified by a circle, while clusters that are "normal" are identified by a triangle.

[0091] For example, the smaller the health level of a cluster, the larger the display size of its identifier can be to make it stand out more visually and attract the attention of maintenance personnel.

[0092] For example, dynamic effects can be used to enhance the ability to capture attention in specific states. For "normal" clusters, the cluster identifier has no dynamic effects, while for "faulty" clusters, the cluster identifier uses highlights or color saturation to change slowly and periodically, or it can use high-frequency flashing or color switching.

[0093] The above embodiments increase the recognizability of cluster health status, making it easier for users to locate problematic clusters in a timely manner.

[0094] In this embodiment, when a user hovers the cursor over a cluster identifier, a tooltip dynamically appears on the interface, displaying key summary metrics for that cluster, such as: cluster online rate, total number of devices, number of currently faulty devices, and cumulative number of faults this week. This allows users to obtain core data without having to navigate to another page.

[0095] The method further includes: Real-time detection of touch operations received by the cluster identifier corresponding to the target cluster, and generation of a query command corresponding to the touched cluster identifier; In response to the query command, the distribution of each device in the target cluster is displayed in the rack expansion area, and each device in the target cluster is displayed differently based on its corresponding device health status.

[0096] In this embodiment, touch operations include multiple types of selection operations, including but not limited to: mouse click, touchscreen click, stylus click, and confirmation by pressing Enter after keyboard navigation.

[0097] The system's front-end interface framework (such as a web browser or client application) continuously listens for user interaction events on the cluster arrangement area.

[0098] When a selection operation is applied to the cluster identifier of a target cluster, the UI logic layer generates a query command. The query command is a data request containing a unique identifier for the target cluster (such as a cluster ID). This request is sent to the backend server via an internal system interface (such as an API call).

[0099] After receiving the request, the backend service queries the database to obtain a detailed list of all devices in the target cluster and their corresponding device health status data, and then returns this data set to the frontend.

[0100] In this embodiment, the rack deployment area presents the returned device health status in a way that closely approximates physical reality. Based on preset device information (such as rack number and U-position coordinates within the rack), the system simulates and draws the physical arrangement of the racks and the device slots and layout within each rack within the rack deployment area. Each device is represented by a device identifier at its corresponding physical location. For example, in cluster X, in rack Z of column Y, device number 15U.

[0101] In this embodiment, each device identifier is visually rendered differently based on its device health status. This differentiation can take the form of differentiated cluster display within the cluster arrangement area.

[0102] For example, the shade of color can be used to indicate the average health of the equipment in the cabinet or the density of faulty equipment.

[0103] The physical layout view greatly shortens the fault location time. At the same time, maintenance personnel can intuitively find out whether faulty devices are clustered in space (such as concentrated in the same cabinet or under the same switch), thus helping to determine whether there is a common cause fault.

[0104] The method further includes: In response to the selection of the target device in the rack expansion area, at least one health indicator of the target device's health status is displayed in the device information area.

[0105] When a user selects a target device in the rack's unfolded area via clicking, touch, or other methods, the system immediately captures this selection event. Based on the target device's unique identifier (such as a device ID or IP address), the system sends a request to the backend data service to retrieve all relevant raw and derived data used to generate the device's health status. The device information area is dynamically refreshed to display the newly loaded device details.

[0106] In this embodiment, the information displayed in the device information area includes, but is not limited to: displaying the device's basic identity information, such as device model, serial number, IP address, and physical location (computer room / rack / U-space).

[0107] And data on at least one health indicator upon which the health status of the target equipment depends. For example, equipment online rate, number of equipment failures, average number of work orders for failed equipment, relative average (percentage), relative average (absolute value), mean repair time, equipment downtime, mean time between failures, mean time to failure, mean interval between failures, etc.

[0108] In some examples, the equipment information area also provides a clickable list displaying all historical fault work orders related to that equipment within the statistical time frame. Clicking on any work order allows you to view its detailed records, such as fault description, repair personnel, replaced parts, and repair logs.

[0109] In other examples, the device information area also provides a trend chart showing the trend of the device's health score or key indicators over a period of time (such as the most recent month), helping to determine changes in the device's status.

[0110] The detailed information displayed provides a direct basis for developing differentiated operation and maintenance solutions. For example, for equipment with an unusually high failure rate, a decision can be made to replace it in advance.

[0111] In this embodiment, the visualization interface includes, from left to right, a device information area, a cluster arrangement area, and a rack expansion area. The display content of the rack expansion area is updated in response to a selection event in the cluster arrangement area; the display content of the device information area is updated in response to a selection event in the cluster arrangement area or the rack expansion area.

[0112] Through the above embodiments, the health status of the cluster or the health status of at least one device in the cluster is displayed through a visual interface. The device information area, cluster arrangement area, and rack expansion area of ​​the visual interface together form a complete and efficient human-computer interaction loop for global overview, problem location, and detailed diagnosis, which greatly improves the construction, operation and maintenance efficiency and accuracy of the health evaluation system for large-scale clusters.

[0113] The cluster device health assessment management method, equipment, and computer-readable storage medium provided in this application acquire fault work order data of target devices in the cluster. Based on the fault work order data, at least one health indicator is calculated to evaluate the operational health of the target devices, thereby transforming discrete, passive fault work order data into a quantitative assessment of the long-term reliability and maintenance requirements of the devices. Then, based on the at least one health indicator, the device health status of the target devices is generated, providing a clear health conclusion for each target device. Furthermore, the accuracy of the device health status assessment is improved through a comprehensive evaluation of multiple dimensions of health indicators. Subsequently, the device health status of each device in the cluster is aggregated to generate and output the cluster health status, eliminating the need to check each device in the cluster individually, greatly shortening the fault location time, and enabling users to have a global understanding of the overall health status of the cluster, providing a direct data-driven decision-making basis for subsequent operation and maintenance.

[0114] Based on the same concept as the methods described above, this application also proposes a cluster device health evaluation and management system. The system includes: The data acquisition module is used to acquire fault work order data of target devices in the cluster. The fault work order data records the operation and maintenance events of the device from the occurrence of the fault to the completion of the repair. The health indicator calculation module is used to calculate at least one health indicator for evaluating the operational health of the target equipment based on the fault work order data. The first health status evaluation module is used to generate the device health status of the target device based on the at least one health indicator. The second health status evaluation module is used to aggregate the device health status of each device in the cluster, generate and output the cluster health status.

[0115] The implementation process of the functions and roles of each unit in the above system is detailed in the implementation process of the corresponding steps in the above method, which can achieve the same technical effect, and will not be repeated here.

[0116] This application also provides a product application for implementing the cluster device health assessment and management method as described in any of the above claims.

[0117] This application also provides an electronic device, including a memory, a processor, and a cluster device health assessment management program stored in the memory and executable on the processor. When the processor executes the cluster device health assessment management program, it implements the cluster device health assessment management method as described in any of the above claims.

[0118] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the cluster device health assessment and management method described in any of the preceding claims.

[0119] This specification also discloses an intelligent agent platform, including one or more processors, for implementing the cluster device health assessment and management method described above.

[0120] Figure 2 This is a schematic diagram of a computer-readable storage medium 140 provided in this disclosure, on which a computer program is stored, which, when executed by a processor, implements the method of any embodiment of this disclosure.

[0121] This disclosure also provides a computing device, including a memory and a processor; the memory is used to store computer instructions that can be executed on the processor, and the processor is used to implement the methods of any embodiment of this disclosure when executing the computer instructions.

[0122] Figure 3 This is a schematic diagram of the structure of a computing device provided in this disclosure, such as... Figure 3As shown, the computing device 15 may include, but is not limited to: a processor 151, a memory 152, and a bus 153 connecting different system components (including the memory 152 and the processor 151).

[0123] The memory 152 stores computer instructions that can be executed by the processor 151, enabling the processor 151 to perform the training method of the aesthetic image generation model according to any embodiment of this disclosure. The memory 152 may include a random access memory unit (RAM) 1521, a cache memory unit (Cache) 1522, and / or a read-only memory unit (ROM) 1523. The memory 152 may also include a program tool 1525 having a set of program modules 1524, including but not limited to: an operating system, one or more application programs, other program modules, and program data. One or more combinations of these program modules may include an implementation of a network environment.

[0124] Bus 153 may include, for example, a data bus, an address bus, and a control bus. The computing device 15 can also communicate with external devices 155 via I / O interface 154, such as a keyboard or a Bluetooth device. The computing device 15 can also communicate with one or more networks via network adapter 156, such as a local area network (LAN), a wide area network (WAN), or a public network. As shown in the figure, the network adapter 156 can also communicate with other modules of the computing device 15 via bus 153.

[0125] Furthermore, although the operations of the methods disclosed herein are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0126] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A method for health assessment and management of cluster devices, characterized in that, The method includes: Obtain fault work order data of target devices in the cluster. The fault work order data records the operation and maintenance events of the device from the occurrence of the fault to the completion of the repair. Based on the work order data of the aforementioned fault, calculate at least one health indicator for evaluating the operational health of the target equipment. The device health status of the target device is generated based on the at least one health indicator; Aggregate the device health status of each device in the cluster, generate and output the cluster health status.

2. The cluster device health evaluation and management method as described in claim 1, characterized in that, Before obtaining the fault work order data of the target device in the cluster, the method further includes: creating a fault work order; The creation of the fault work order includes: In response to an abnormal alarm received from the monitoring system, a fault work order is automatically created and stored; wherein the fault work order contains fault data based on the abnormal alarm.

3. The cluster device health evaluation and management method as described in claim 2, characterized in that, The work order creation includes: In response to a user-triggered work order creation operation, create a framework for the fault work order and display the corresponding interactive interface; The fault data entered through the interactive interface is filled into the frame of the fault work order to generate a complete fault work order.

4. The cluster device health evaluation and management method as described in claim 2 or 3, characterized in that, After creating the fault work order and before obtaining the fault work order data of the target device in the cluster, the method further includes: In response to the user's editing operation, update the device status and repair information of the target device associated with the fault work order; and / or, In response to the closing command for the fault work order, the status of the fault work order is marked as closed.

5. The cluster device health evaluation and management method as described in claim 1, characterized in that, The health status of the cluster or the health status of at least one device in the cluster is displayed through a visual interface, which includes at least a cluster arrangement area. The method further includes: At least one cluster identifier is generated and displayed in the cluster arrangement area; wherein the cluster identifier is used to indicate the health status information of the cluster, and the health status information is determined based on the device health status of each device in the corresponding cluster.

6. The cluster device health assessment and management method as described in claim 5, characterized in that, The method further includes: Based on the health level represented by the health status information of the cluster, the identifiers of each cluster are displayed differently in the cluster arrangement area.

7. The cluster device health assessment and management method as described in claim 5, characterized in that, The visual interface also includes a rack unfolding area; The method further includes: Real-time detection of touch operations received by the cluster identifier corresponding to the target cluster, and generation of a query command corresponding to the touched cluster identifier; In response to the query command, the distribution of each device in the target cluster is displayed in the rack expansion area, and each device in the target cluster is displayed differently based on its corresponding device health status.

8. The cluster device health assessment and management method as described in claim 7, characterized in that, The visual interface also includes a device information area; The method further includes: In response to the selection of the target device in the rack expansion area, at least one health indicator of the target device's health status is displayed in the device information area.

9. The cluster device health evaluation and management method as described in claim 1, characterized in that, The health indicators include at least one of the following: equipment online rate, number of equipment failures, mean time to repair, mean time between failures, and mean time to failure; and / or, The fault work order data includes at least one of the following: equipment failure time, repair time, and fault type.

10. A product application, characterized in that, Used to implement the cluster device health evaluation management method as described in any one of claims 1 to 9.

11. An electronic device, characterized in that, The system includes a memory, a processor, and a cluster device health assessment management program stored in the memory and executable on the processor. When the processor executes the cluster device health assessment management program, it implements the cluster device health assessment management method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, performs the cluster device health evaluation and management method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Equipment cluster health state evaluation method based on industrial big data

    CN107358347A

  • Wind turbine generator health degree and service quality evaluation method

    CN113610443A

  • Industrial equipment health monitoring and fault early warning system based on edge computing

    CN119249274A

  • Fan fault guidance system and fan fault guidance method

    CN120407590A

  • Cluster fault management method and device, electronic equipment and medium

    CN120528774A