Test scheduling method for health monitoring, electronic equipment and storage medium
By constructing test tasks and evaluating their priorities, and combining them with predictive models for dynamic scheduling, the problems of resource waste and delayed fault detection in the health monitoring of computing facilities have been solved, achieving efficient and proactive health monitoring and improving the reliability and stability of the system.
Patent Information
- Application Number
- CN202511786220.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-12-01
AI Technical Summary
Existing health monitoring methods for computing facilities suffer from high testing frequency, high resource consumption, and difficulty in early detection of potential faults, resulting in faults or anomalies being triggered only at a critical stage, thus reducing the initiative and foresight of health management.
By constructing test tasks, evaluating the priority of test tasks based on a preset set of indicators, and using a predictive model to predict the failure probability of test tasks, the test tasks are dynamically scheduled for execution, thereby achieving proactive health monitoring of computing facilities.
It effectively reduced operational risks, improved testing efficiency, ensured the coverage of health monitoring and the efficiency of resource utilization, promptly identified potential problems, and enhanced the reliability and stability of the system.
Smart Images

Figure CN121233271A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the technical field of computing facility testing, and more specifically, to a test scheduling method for health monitoring, an electronic device, and a non-transitory computer-readable storage medium. Background Technology
[0002] Computing infrastructure refers to the hardware and software resources system capable of performing computing tasks, and it is widely used in scenarios such as artificial intelligence, big data, high-performance computing, cloud computing, and distributed systems. To ensure the stability, reliability, and sustainable operation of computing infrastructure, continuous and efficient health monitoring is crucial. However, current mainstream health monitoring methods mainly rely on periodic polling detection or fixed threshold alarm strategies, both of which have significant limitations.
[0003] On the one hand, traditional periodic polling methods typically perform full tests on all nodes in the computing facility at fixed time intervals. While this method can comprehensively cover the system status, the high testing frequency and wide coverage often result in significant computing power and resource overhead, especially in large-scale computing clusters where polling costs are difficult to ignore. On the other hand, fixed threshold alarm strategies monitor hardware indicators such as temperature, memory usage, and power consumption, triggering alarms or tests when these indicators exceed preset thresholds. This method reduces unnecessary testing overhead to some extent, but it is difficult to detect potential faults or performance degradation in advance. This often leads to faults or anomalies developing to a more serious stage by the time alarms are actually triggered, increasing system operational risks and reducing the initiative and foresight of health management. Summary of the Invention
[0004] One objective of this disclosure is to provide a new technical solution for health monitoring of computing facilities, enabling proactive scheduling and efficient execution of test tasks.
[0005] According to a first aspect of this disclosure, a test scheduling method for health monitoring is provided, the method comprising: Multiple test tasks are constructed based on multiple set test items and multiple monitored nodes in the computing power facility; each test task consists of a test item and a node. For each of the aforementioned test tasks, a test indicator value for the corresponding test task is obtained based on a preset indicator set; wherein, the test indicator value represents the test priority of the corresponding test task. The test tasks are scheduled and executed based on the test metric values of each test task.
[0006] Optionally, the preset indicator set includes failure probability indicators for test tasks, and determining the failure probability indicator value for the test task includes: Obtain the associated data of the test task; wherein, the associated data includes data associated with at least one of the nodes and test items in the test task; Generate the feature vector of the test task based on the associated data; The prediction model is invoked to predict the failure probability index value of the test task based on the feature vector, wherein the prediction model is trained to establish a mapping relationship from the feature vector of the test task to the failure probability of the test task.
[0007] Optionally, the feature vector includes structured features and temporal features, and the invocation prediction model predicts the failure probability index value of the test task based on the feature vector, including: The first prediction model is invoked to generate a first failure probability based on the structured features; The second prediction model is invoked to generate a second failure probability based on the aforementioned temporal features; By combining the first failure probability and the second failure probability, the failure probability index value of the test task is obtained.
[0008] Optionally, the plurality of test tasks includes cold start tasks that meet cold start conditions, wherein at least one of the test items and nodes does not have historical test data; generating the feature vector of the cold start task based on the associated data of the cold start task includes: Invoke at least one of the feature extraction model and the knowledge graph to generate the feature vector of the cold start task based on the associated data of the cold start task; The feature extraction model is trained to generate at least some features of the cold start task based on the associated data of the cold start task; the knowledge graph is constructed based on the association between node hardware configuration, test items and fault information, and the target path in the knowledge graph that matches the cold start task is used to generate at least some features of the cold start task.
[0009] Optionally, the preset indicator set includes at least two of the following: a first type of indicator, a second type of indicator, and a third type of indicator; the first type of indicator is used to evaluate the contribution of the combination of test items and nodes to test priority, the second type of indicator is used to evaluate the contribution of the inherent attributes of test items to test priority, and the third type of indicator is used to evaluate the contribution of the inherent attributes of nodes to test priority.
[0010] Optionally, the first type of indicator includes at least some of the indicators among the failure probability indicator and the timeliness indicator of the test task, wherein the timeliness indicator represents the length of time since the most recent successful test; or, The second category of indicators includes at least some of the business impact indicators and resource cost indicators of the test items; the business impact indicators represent the degree of impact of the node capabilities verified by the test item on the business, and the resource cost indicators represent the proportion of resources consumed in executing the test item; or, The third type of indicator includes the reliability degradation index of the node, which is related to the service life of the node.
[0011] Optionally, obtaining the test index value for each test task based on a preset index set includes: Determine the individual indicator value for each indicator in the preset indicator set for the test task; Determine the weight of each indicator in the preset indicator set; wherein, the weights of at least some indicators in the preset indicator set are obtained by sampling based on the weight probability distribution of the corresponding indicators, and the weight probability distribution of the corresponding indicators is updated based on historical test data. The test index value of the test task is obtained based on the individual index value and weight of each index.
[0012] Optionally, scheduling and executing test tasks based on the test metric values of each test task includes: Based on the test index values of each test task, the test tasks are distributed and executed under the constraints of the set scheduling conditions.
[0013] Optionally, after scheduling and executing the test task, the method further includes: Obtain test data from the executed test tasks; Based on the test data, update the adaptive parameters related to the test metric values of the test task, and / or update the adaptive parameters related to scheduling.
[0014] According to a second aspect of this disclosure, an electronic device is also provided, the electronic device comprising: At least one processor; and A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method according to the first aspect of this disclosure.
[0015] According to a third aspect of this disclosure, a non-transitory computer-readable storage medium is also provided, the non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method according to the first aspect of this disclosure.
[0016] This disclosure generates test tasks by combining test items with computing facility nodes, and evaluates the test priority of each test task based on a preset set of indicators, thereby scheduling and executing test tasks based on test priorities. On the one hand, by performing proactive testing on nodes in the computing facility, this disclosure enables continuous monitoring of the health status of the computing facility, effectively reducing operational risks. On the other hand, by dynamically evaluating the priority of test tasks and scheduling them accordingly, unnecessary resource consumption can be reduced while ensuring the effectiveness of health monitoring, thus improving overall testing efficiency.
[0017] The features and advantages of the embodiments of this specification will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of this specification and, together with their description, serve to explain the principles of these embodiments.
[0019] Figure 1 This is a schematic diagram of the composition structure of a computing facility provided in this disclosure; Figure 2 This is a flowchart illustrating a test scheduling method according to some embodiments; Figure 3 This is a flowchart illustrating the probability of test task failure based on some embodiments; Figure 4 This is a flowchart illustrating a test scheduling method according to some other embodiments; Figure 5 This is a schematic diagram of the composition structure of a health monitoring system according to some embodiments; Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to some embodiments. Detailed Implementation
[0020] Various exemplary embodiments of this specification will now be described in detail with reference to the accompanying drawings.
[0021] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the embodiments of this specification or their application or use.
[0022] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0023] This disclosure relates to a technical solution for health monitoring of nodes in a computing facility. For example... Figure 1As shown, the computing facility includes multiple nodes, denoted as node n1, node n2, ..., node nX, where X is a positive integer. These nodes are primarily computing nodes with computational capabilities, and each node contains at least one type of processor, such as at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), or Neural Processing Unit (NPU). Furthermore, each node is typically configured with a communication interface for inter-node data communication, such as an Ethernet interface, fiber optic communication interface, or other high-speed interconnect interface. Some nodes may further include local memory for storing temporary or long-term data required for node operation. For example, ... Figure 1 As shown, node n2 includes at least one processor 201, at least one memory 202, and a communication interface 203. When node n2 contains multiple processors, these processors can be of the same type or different types to adapt to different computing task requirements.
[0024] In some embodiments, the computing facility may also include a shared storage system, in which storage nodes may be shared by some or all of the nodes in the computing facility, for providing centralized or distributed storage services.
[0025] In this embodiment of the disclosure, the computing power facility further includes a health monitoring system 100, which is used to monitor and evaluate the health status of nodes in the computing power facility. The deployment of the health monitoring system 100 is highly flexible; for example, it can be deployed on one or more specific nodes in the computing power cluster; it can be deployed independently of the computing power cluster, such as on a dedicated monitoring node or edge device; or it can adopt a hybrid deployment mode, where some functions are deployed on the nodes of the computing power cluster, such as data acquisition, while other functions are deployed in the cloud, such as the cloud handling health scoring and task scheduling.
[0026] As shown in the figure, the health monitoring system 100 may include at least one processor 101 and at least one memory 102. The processor 101 executes computer programs, which may be written based on various instruction set architectures. The memory 102 stores the executable program and may be of a non-volatile storage medium such as read-only memory (ROM), random access memory (RAM), or hard disk. Furthermore, the health monitoring system 100 may also include a communication interface 103 for data communication with nodes within the computing facility and with external systems, supporting the acquisition, transmission, and feedback of monitoring data and analysis results.
[0027] The following is combined with Figure 1 The schematic computing facilities illustrate various embodiments of this disclosure.
[0028] <First Embodiment> Figure 2 A flowchart illustrating a test scheduling method for health monitoring according to some embodiments is shown. This method can be... Figure 1 The health monitoring system is implemented in 100% of the country. For example... Figure 2 As shown, the method of this embodiment may include the following steps S210 to S230.
[0029] Step S210: Construct multiple test tasks based on the set multiple test items and multiple nodes in the computing power facility.
[0030] In this embodiment, a test set for health monitoring can be set up for the computing facility based on its business domain and structural characteristics. The test set includes multiple test items, which are executable detection units designed to address potential failure modes or performance bottlenecks in the computing facility. Examples of test items include: test items for monitoring the correctness of node computing functions; test items for monitoring the normality of the node's Peripheral Component Interconnect Express (PCIe) links; test items for monitoring the normality of the node's temperature reading function; test items for performing GPU memory diagnostic tests; test items for stress testing the GPU; test items for monitoring network bandwidth and latency; disk I / O throughput tests, etc., and are not limited here.
[0031] In this embodiment, multiple test tasks can be constructed by combining test items from the test set with nodes in the computing facility that need to be monitored. That is, each test task consists of a test item from the test set and a node from the computing facility. In this embodiment, a node refers to an entity in the computing facility that needs to be monitored. These nodes can be nodes with computing capabilities in the computing cluster, such as servers equipped with GPUs or CPUs. These nodes can also include other types of nodes in the computing facility, such as storage nodes, etc., and are not limited here.
[0032] Each node can be combined with test items related to that node's function in the test set to obtain a set of test tasks that need to be scheduled and executed.
[0033] For example, node n1 can be combined with test item t1 and test item t2 to construct the first test task T1(n1,t1) and the second test task T2(n1,t2); node n2 can be combined with test item t1, test item t2 and test item t3 to construct the third test task T3(n2,t1), the fourth test task T4(n2,t2) and the fifth test task T5(n2,t3), etc.
[0034] Step S220: For each test task, obtain the test index value of the corresponding test task based on the preset index set.
[0035] The test index value of a test task indicates the test priority of the corresponding test task. In step S220, the health monitoring system converts the status of each test task into a test index value quantified according to the same standard, thereby providing an accurate and objective decision-making basis for subsequent test task scheduling and realizing proactive testing of nodes.
[0036] In this embodiment, the preset index set includes at least one index, and the indexes in the preset index set are used to evaluate the urgency of the test task.
[0037] In some examples, the preset metric set is a multi-dimensional metric set, which includes metrics from multiple dimensions to evaluate the test priority of test tasks from different dimensions, avoiding the limitations of a single dimension, and thus improving the effectiveness and robustness of test metric values.
[0038] In these examples, the preset metric set may include at least two of the following three categories of metrics: the first category, the second category, and the third category. The first category of metrics is used to evaluate the real-time risk of the combination of test items and nodes, focusing on the specific state of the test task at the current moment. The second category of metrics is used to evaluate the contribution of the inherent attributes of test items to the test priority, reflecting the inherent importance and cost of test items. The third category of metrics is used to evaluate the contribution of the inherent attributes of nodes to the test priority.
[0039] In another example, the preset metric set may also include only the first type of metrics mentioned above, that is, the test priority of the test task is evaluated by the first type of metrics.
[0040] For example, the first type of metric may include test task T. The failure probability metric, where test task T It consists of any node n and any test item t. The failure probability index represents the probability that node n cannot pass test item t in the current state. From the perspective of failure probability, the failure probability index value... The larger the value, the higher the health risk of node n. (Test task T) The higher the test priority or urgency, the better.
[0041] For example, the first type of metric could include test task T. The timeliness metric represents the distance from the current time to the test task T. The length of time since the most recent successful test. From a timeliness perspective, timeliness metric values. The larger the value, the longer the test item at that node has been neglected, increasing uncertainty and the corresponding test task T. The higher the test priority, the better.
[0042] For example, the second type of metric can include the business impact metric for test item t. The business impact metric represents the degree to which the node capabilities verified by test item t affect business continuity. From the perspective of business impact, the business impact metric value... The larger the value, the higher the test priority of the corresponding test task. For example, the test item verifying the network link has a greater service impact and a higher priority than the test item verifying the temperature sensor. Here, a corresponding service impact index value can be configured for each test item based on the importance of the node capability corresponding to each test item in the test set. For example, the service impact index value of the PCIe link test item can be set to 0.8, while the service impact index value of the temperature reading test item can be set to 0.2, etc.
[0043] For example, the second type of metric could include the resource cost metric for test item t. The resource cost metric represents the amount of resources consumed in executing the test item. This metric covers resources including computing resources, and may further include storage resources and / or network resources. Resource cost metric values... The priority of a test can be determined based on its resource consumption percentage and execution time. Resource cost metrics negatively impact test priority; from a resource cost perspective, resource cost metric values... The larger the value, the higher the execution cost, and its scheduling priority will be appropriately reduced when resources are limited.
[0044] The third category of indicators includes, for example, the reliability degradation index of node n, which is related to the service life of node n. Related, service duration This can be measured in days; in this case, the service duration is the number of days between the current time and the timestamp of node n going online. Reliability degradation metric value. This can be modeled as a function of node service life, such as an exponential function. From the perspective of reliability degradation metrics, the longer the service life, the greater the probability of hardware aging in the node, and the higher the reliability degradation metric value. The larger the value, the higher the priority of the corresponding test task.
[0045] For example: reliability degradation index value ,in, The set attenuation coefficient.
[0046] In an example where the preset metric set includes multiple metrics, test task T Test index value It can be based on test task T The weighted summation method for each indicator's individual value is determined. For example, the preset indicator set includes the aforementioned failure probability indicator, business impact indicator, resource cost indicator, reliability degradation indicator, and timeliness indicator, with the following weights for each indicator: , , Test Task T The test metric value can be expressed as:
[0047] Through this weighted model, the health monitoring system can detect test tasks that are "high-risk, have high business impact, have controllable resource costs, and have not been tested for a long time," and prioritize the execution of these test tasks in the test scheduling. This ensures that limited test resources are invested in the areas that can best improve the overall reliability of the computing facilities, thereby maximizing operational efficiency and system stability.
[0048] In some examples, a fixed weight can be set for each metric in a pre-defined set of metrics.
[0049] In other examples, to further enhance the adaptability and long-term optimization of the health monitoring system, the system supports dynamic adjustment and online learning of the weights of each indicator in a preset indicator set during the monitoring process. This approach is more adaptable to complex and ever-changing operational scenarios, and by driving weight optimization through data feedback, it makes the guidance of test indicator values for test scheduling more accurate.
[0050] In these examples, step S220, which involves obtaining the test metric value for each test task based on a preset metric set, may further include the following steps S221 to S223: Step S221: For each test task, determine the individual indicator value of each indicator in the preset indicator set for the test task.
[0051] For example, test task T Failure probability index value Timeliness index value Business impact index values These are all individual indicator values.
[0052] Step S222: Determine the weight of each indicator in the preset indicator set.
[0053] In this example, the weights of at least some indicators are obtained by sampling based on the weight probability distribution of the corresponding indicators, which is dynamically updated based on historical test data.
[0054] This example can employ a reinforcement learning approach based on exploration and exploitation, maintaining a weight probability distribution (such as a Beta distribution) for each indicator that requires dynamic adjustment. Before each round of scheduling decisions, the health monitoring system samples a set of target weights from these weight probability distributions. Through weight sampling, it can utilize the currently optimal weight configuration presented by the weight probability distribution while retaining a certain degree of exploration capability to try potentially better weight configurations.
[0055] In this example, strategies such as Bayesian optimization algorithms (e.g., Thompson Sampling), Upper Confidence Bound (UCB), or Softmax can be used to manage and update the weight probability distribution.
[0056] Taking a preset set of indicators including the aforementioned five indicators—failure probability indicator, business impact indicator, resource cost indicator, reliability degradation indicator, and timeliness indicator—as an example, the health monitoring system maintains a weighted probability distribution for the failure probability indicator and the business impact indicator, while the other indicators use fixed initial weights. , In a scheduling round, based on the weight distribution probability of the failure probability index, the initial weight of the index in this scheduling round is sampled. And based on the weight distribution probability of the business impact indicator, the initial weight of the indicator in this scheduling round is sampled. Then the initial weights Initial weights Initial weights , Normalization is performed so that the sum of the weights of all indicators equals 1, thus obtaining the target weight of each indicator in this scheduling round.
[0057] Step S223: Based on the individual indicator value of each indicator and the weight of each indicator, obtain the test indicator value of the test task.
[0058] Step S230: Schedule and execute test tasks according to the test index values of each test task.
[0059] The health monitoring system 100 can distribute and execute test tasks based on the test index values of all test tasks calculated in step S220, ensuring that test tasks with high test index values are given priority in this round of scheduling.
[0060] To adapt to the continuous changes in the state of computing facilities, the health monitoring system 100 can re-evaluate the test index values of all test tasks at fixed time intervals or based on event triggers, and update the scheduling strategy accordingly. This enables the scheduling strategy to respond to changes in the state of computing facilities in real time, achieving persistent and accurate scheduling.
[0061] In some examples, the health monitoring system can select a subset of test tasks to be executed in the current round from all test tasks based on a preset strategy. The preset strategy could be to select test tasks whose test indicator values exceed a set threshold, or it could be to select the test tasks whose test indicator values rank in the top K, etc.
[0062] For the selected subset of test tasks, they can be sorted in descending order according to their test metric values to generate an ordered sequence of test task execution, ensuring that test tasks with higher test metric values are executed first. For example, Celery can be used as a distributed task queue framework, combined with a Redis Sorted Set (ZSET) to implement a priority queue. Test tasks are stored in the ZSET as scores based on their test metric values, and the health monitoring system retrieves tasks from highest to lowest score, ensuring that high-priority tasks are executed first.
[0063] In some examples, scheduling conditions can be set to constrain the scheduling strategy. These conditions can include limiting at least one of the following: the maximum concurrent execution of test tasks, the upper limit of GPU memory usage, network bandwidth usage, and node CPU utilization. When scheduling test tasks, the health monitoring system distributes them based on the test metric values and the scheduling conditions, ensuring that the resource consumption of nodes executing test items minimizes the impact on the normal operation of the computing facilities. For example, the health monitoring system can prioritize assigning test tasks with low resource consumption (such as GPU memory) and high test metric values to appropriate nodes.
[0064] In some examples, to achieve the optimal balance between exploration and utilization of existing test metric values, health monitoring systems can use reinforcement learning algorithms such as Thompson sampling to drive the scheduling and execution of test tasks. To this end, the health monitoring system can, on the one hand, select the test task with the highest expected reward (i.e., the highest test metric value) based on its test metric value; on the other hand, it can randomly select a subset of test tasks with a certain probability to collect new test data, avoiding getting trapped in local optima. Based on this, the health monitoring system can use Bayesian methods to update its scheduling strategy based on the reward feedback obtained after scheduling execution from both aspects, achieving continuous optimization of the scheduling strategy.
[0065] According to steps S210 to S230, in this embodiment, a set of test tasks is constructed, consisting of nodes in the computing facility and test items in a test set. Each test task is prioritized based on a preset set of indicators. This enables the health monitoring system to accurately identify combinations of high-risk nodes and key test items, thereby achieving proactive and targeted health monitoring and testing of nodes in the computing facility. The method in this embodiment facilitates the timely detection of potential problems before a failure occurs or in the early stages of performance degradation, significantly reducing system operational risks and improving overall reliability and stability.
[0066] On the other hand, this embodiment evaluates the priority of test tasks and implements intelligent scheduling and task distribution based on priority scores, thereby enabling the priority execution of high-value test tasks under limited resource conditions. This strategy not only ensures the effectiveness and coverage of health monitoring, but also effectively avoids the ineffective occupation of computing resources by low-priority or redundant tests, thus improving the utilization efficiency of test resources and reducing the overall system overhead.
[0067] <Second Embodiment> In this embodiment, the preset indicator set includes at least the failure probability indicator of the test task. The health monitoring system predicts the failure probability indicator value of each test task by calling a pre-trained prediction model.
[0068] In this embodiment, as Figure 3 As shown, determining the failure probability index value for any test task may include the following steps S321 to S323: Step S321: Obtain the associated data of the test task.
[0069] In this embodiment, the associated data of the test task includes data associated with at least one of the nodes and test items in the test task. To ensure data privacy and transmission security, all associated data can be transmitted using an encrypted channel.
[0070] The associated data for a test task may include at least some of the following: hardware status data of the corresponding node, historical test data related to the test task, and business load data of the corresponding node.
[0071] Historical test data related to the test task includes historical test data of nodes within the test task for test items within the test task. Historical test data related to the test task may also include at least one of the following: historical test data of the corresponding node for other test items, and historical test data of the corresponding test item on other nodes.
[0072] The node's hardware status data can include hardware configuration data and real-time status data. Hardware configuration data includes, for example, at least some of the static attributes such as GPU model, CPU model, driver version, and firmware version. Real-time status data includes, for example, GPU status data and / or CPU status data. GPU status data can include GPU memory usage, temperature, power consumption, ECC error count, etc., while CPU status data can include CPU load rate, PCIe bandwidth, InfiniBand / RoCE link latency, etc.
[0073] Historical test data is used to record the historical execution status of test tasks. The health monitoring system can maintain time-series data sequences of tests from multiple dimensions. For example, it can maintain corresponding time-series data sequences from the dimensions of test tasks (combinations of test items and nodes), nodes, and test items, respectively, to support the rapid retrieval of historical test data related to test tasks as needed. Each data point in the time-series data sequence can include test execution timestamps, test results (pass or fail), performance metric values obtained from the test, and other test data.
[0074] Load data reflects the current workload on a node and can be obtained through the Kubernetes API. Load data includes, for example, the types of tasks running on the node (such as model training and model inference) and resource request specifications. Load data helps health monitoring systems assess the impact of test items on the node's running capabilities.
[0075] Step S322: Generate feature vectors for the test task based on the associated data.
[0076] Step S322 transforms the original associated data into a set of structured, computable feature vectors, providing high-quality input for the prediction model.
[0077] In some examples, associated data can be transformed into three types of features: static attribute features, time-series features, and contextual features.
[0078] Static attribute features describe the relatively stable inherent properties of nodes or test items. Examples of static attribute features include: hardware model, driver version, firmware version, node service life, etc. Static attribute features are primarily derived from hardware status data.
[0079] Time-series features are dynamic trends and statistical characteristics extracted from historical test data, including at least one of the following: Failure rate sliding window: Calculates the failure rate of the test task in the most recent M historical tests; Performance degradation trend: For example, using methods such as Exponentially Weighted Moving Average (EWMA) or linear fitting, calculate the rate of decline of key performance indicators (such as the computing power throughput of half-precision floating-point FP16). Anomaly detection features: For example, using statistical methods such as Z-score and IQR (interquartile range) to identify abnormal patterns (such as sudden increases in video memory usage) exhibited by nodes in historical tests.
[0080] Contextual features introduce task-related environmental and management information derived from historical test data and workload data, including at least one of the following: Behavior of nodes in the same batch: Calculate the average failure rate of nodes of the same model or batch when executing the same test item; Task priority context: Whether the current node is executing a high-priority business task; Related test item results: The test result characteristics of nodes in the test task in the related test items. The related test items are other test items that are related to the test items in the test task. For example, the health status of the PCIe link affects the multi-card communication test, so the PCIe link test is related to the multi-card communication test.
[0081] To construct a complete and predictive feature vector, it is typically required that the nodes and test items involved in the test task have prior historical test data. However, in real-world scenarios, some test tasks may be cold-start tasks, meaning that at least one of the nodes and test items in the task has not yet accumulated sufficient historical test data, thus failing to provide adequate associated data for feature extraction. Such cold-start tasks, due to limited associated data, can affect the completeness of the feature vector and the accuracy of the prediction results to some extent.
[0082] To improve the predictive effectiveness of cold start tasks, in some examples, the health monitoring system 100 can call at least one of the pre-trained feature extraction module and the knowledge graph to generate feature vectors for cold start tasks based on existing associated data of the cold start task, mainly the hardware status data of the nodes.
[0083] In this example, the feature extraction model is trained to generate representative feature representations based on available associated data (mainly hardware status data) of nodes in the cold start task.
[0084] In this example, the knowledge graph is constructed based on the relationships between node hardware configurations, test items, and fault information. By searching the knowledge graph for target paths that match the cold start task, fault information on the target paths is extracted as at least partial features.
[0085] For example, generating a feature vector for a cold start task may include the following steps: calling a feature extraction model to generate a first part of the cold start task's features based on the associated data of the cold start task; searching for a target path that matches the cold start task in the knowledge graph and using the fault information on the target path as the second part of the cold start task's features; and fusing the first part of the features and the second part of the features to form the feature vector of the cold start task.
[0086] Step S323: Call the prediction model to predict the failure probability index value of the corresponding test task based on the feature vector.
[0087] In this embodiment, the prediction model is trained to establish a mapping relationship from the feature vector of the test task to the failure probability of the test task.
[0088] Predictive models can be trained on a sample set. For real samples, the number of positive samples that pass the test is far greater than the number of negative samples that fail. To address the imbalance between positive and negative samples, a weighted loss function can be used during model training to reduce the influence of the majority of positive samples and improve the model's ability to identify minority negative samples.
[0089] In examples where feature vectors include both temporal and structured features, a model fusion strategy can be employed to fully utilize information from different types of features. This involves using a first prediction model adept at handling structured features, such as XGBoost or LightGBM, and a second prediction model adept at handling temporal features, such as LSTM or Transformer, to jointly predict the failure probability of the test task. Structured features refer to data features without inherent temporal order or sequential dependencies. These features can be independent, discrete, or static attributes whose values do not depend on the order of appearance of other features, including, for example, the aforementioned static attribute features. Structured features can be stored in tabular form, with each row representing a sample and each column representing a feature. Temporal features, on the other hand, are features extracted from a series of data points arranged in chronological order. These data points have sequential dependencies and can be used to capture patterns, trends, periodicity, and dynamic changes in data over time.
[0090] In this example, calling the prediction model to predict the failure probability index value of the test task based on the feature vector can further include the following steps: calling the first prediction model to generate a first failure probability based on structured features; calling the second prediction model to generate a second failure probability based on temporal features; and fusing the first failure probability and the second failure probability to obtain the failure probability index value of the test task.
[0091] In this example, a weighted summation method can be used to fuse the first and second failure probabilities. The first and second prediction models can have the same weight or different weights; for example, the weight of the first prediction model can be set to be greater than the weight of the second prediction model. Alternatively, a fusion strategy such as taking the maximum value can be used to obtain the failure probability index of the test task; no specific limitation is imposed here.
[0092] According to steps S321 to S323, this embodiment uses a pre-trained prediction model to evaluate the failure probability of the test task, which can reduce the dependence on expert experience and thus adapt to the complex and ever-changing operating environment, providing a reliable and objective risk quantification basis for scheduling execution.
[0093] In addition, those skilled in the art will understand that, in other embodiments, the health monitoring system 100 may also establish a mapping between associated data and failure associations based on an expert knowledge base.
[0094] <Third Embodiment> In this embodiment, by constructing a closed loop for health monitoring, the execution results of the test tasks are fed back into the system to dynamically optimize the priority evaluation strategy and / or scheduling strategy, thereby significantly improving the long-term stability and adaptability of the health monitoring system.
[0095] like Figure 4 As shown, the test scheduling method of this embodiment may include the following steps S410 to S450: Step S410: Construct multiple test tasks based on the set multiple test items and multiple nodes in the computing power facility.
[0096] Step S420: For each test task, obtain the test index value of the corresponding test task based on the preset index set.
[0097] Step S430: Schedule and execute test tasks according to the test index values of each test task.
[0098] Step S440: Obtain test data for executing the test task.
[0099] After completing the test task, the health monitoring system acquires the actual test data, including test results (whether the test passed or failed), performance metrics of the tested node, execution time, and any anomalies. This test data reflects the node's true health status in its current state and provides an objective basis for system optimization.
[0100] Step S450: Update the adaptive parameters related to the test metric values of the obtained test task based on the test data, and / or update the adaptive parameters related to scheduling.
[0101] The health monitoring system 100 uses the test data obtained in step S440 to dynamically update and optimize the adaptive parameters that affect the calculation of test index values or task scheduling decisions.
[0102] In some examples where individual metric values are obtained based on models, such as predicting the failure probability metric value of a test task using a prediction model, the adaptive parameters related to obtaining the test metric value of the test task may include model parameters used for metric value calculation, and may also include model parameters of the feature extraction model for the cold start task.
[0103] In an example where the metric weights can be dynamically adjusted, the adaptive parameters related to obtaining the test metric values for the test task can include distribution parameters of the weight probability distribution, such as parameters defining the Beta distribution.
[0104] When the scheduling strategy for a test task is built based on reinforcement learning, heuristic algorithms, etc., the adaptive parameters related to scheduling may include exploration and exploitation strategy parameters, task ranking threshold parameters, and parameters used in scheduling conditions.
[0105] like Figure 5 As shown, in some examples, a health monitoring system may include a data acquisition module for collecting associated data, a feature engineering module for generating feature vectors based on associated data, a risk prediction module for predicting the failure probability of test tasks based on feature vectors, a priority scoring module for prioritizing based on a set of predicted indicators, a scheduling module for executing scheduling strategies, a test execution module for executing test tasks, and a feedback loop module for adaptive parameter updates based on test data. The feedback loop module can be used to update the parameters of the feature engineering module, the risk prediction module, the priority scoring module, the scheduling module, etc.
[0106] <Fourth Embodiment> This disclosure also provides an electronic device, such as... Figure 6 As shown, the electronic device 600 includes at least one processor 610 and at least one memory 620, the memory 620 being used to store computer program instructions, which, when executed by the processor 610, cause the electronic device 600 to perform a test scheduling method according to any embodiment of the present disclosure.
[0107] The electronic device 600 deploys the aforementioned health monitoring system to execute test scheduling methods; it can be a server or other types of equipment.
[0108] This disclosure also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements any of the memory management methods or compilation methods described in the foregoing embodiments of this disclosure. Optionally, the computer-readable storage medium may be a non-transitory storage medium, but is not limited thereto; it may also be a temporary storage medium.
[0109] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0110] This disclosure may be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement any of the methods in the foregoing embodiments of this disclosure.
[0111] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media may include, for example, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), compact disc-read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any combination thereof. The computer-readable storage medium used herein is not to be interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0112] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include one or more of copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to computer-readable storage media in the respective computing / processing device.
[0113] The computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source or object programs written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network (e.g., a local area network or a wide area network), or it may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays, or programmable logic arrays, can execute computer-readable program instructions to implement various aspects of the embodiments of this disclosure by utilizing state information from the computer-readable program instructions.
[0114] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0115] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0116] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It should be noted that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are all equivalent.
[0118] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of this disclosure is defined by the appended claims.
Claims
1. A test scheduling method for health monitoring, characterized in that, The method comprises: constructing a plurality of test tasks according to a plurality of test items and a plurality of nodes monitored in the computing power facility; each test task is composed of a test item and a node; for each test task, obtaining a test index value of the corresponding test task based on a preset index set; the test index value represents a test priority of the corresponding test task; performing scheduling execution of the test task according to the test index value of each test task.
2. The method of claim 1, wherein, The preset index set includes a failure probability index of the test task, and the failure probability index value of the test task is determined by: obtaining associated data of the test task; the associated data includes data associated with at least one of the test item and the node in the test task; generating a feature vector of the test task based on the associated data; calling a prediction model to predict the failure probability index value of the test task based on the feature vector, wherein the prediction model is trained to establish a mapping relationship from the feature vector of the test task to the failure probability of the test task.
3. The method of claim 2, wherein, The feature vector includes a structured feature and a time sequence feature, and the prediction model is called to predict the failure probability index value of the test task based on the feature vector, comprising: calling a first prediction model to generate a first failure probability based on the structured feature; calling a second prediction model to generate a second failure probability based on the time sequence feature; fusing the first failure probability and the second failure probability to obtain the failure probability index value of the test task.
4. The method of claim 2, wherein, The plurality of test tasks includes a cold start task satisfying a cold start condition, the cold start condition includes that at least one of the test item and the node does not exist historical test data; the feature vector of the cold start task is generated based on the associated data of the cold start task, comprising: calling at least one of a feature extraction model and a knowledge graph to generate the feature vector of the cold start task based on the associated data of the cold start task; wherein the feature extraction model is trained to generate at least part of the feature of the cold start task based on the associated data of the cold start task; the knowledge graph is constructed based on the association relationship between the node hardware configuration, the test item and the fault information, and a target path in the knowledge graph matched with the cold start task is used to generate at least part of the feature of the cold start task.
5. The method according to any one of claims 1 to 4, characterized in that, The preset index set includes at least two types of indexes in the first type of index, the second type of index and the third type of index; the first type of index is used to evaluate the contribution degree of the combination of the test item and the node to the test priority, the second type of index is used to evaluate the contribution degree of the inherent attribute of the test item to the test priority, and the third type of index is used to evaluate the contribution degree of the inherent attribute of the node to the test priority.
6. The method of claim 5, wherein: the first type of index includes at least part of the failure probability index of the test task and the timeliness index of the test task, and the timeliness index represents the length of time from the current time to the nearest test success time; or The second type of index includes at least part of a business impact index of the test item and a resource cost index of the test item; the business impact index represents an influence degree of a node capability verified by the test item on a business, and the resource cost index represents a proportion of resource consumption of executing the test item; or The third type of index includes a reliability attenuation index of the node, and the reliability attenuation index is related to a service time length of the node.
7. The method according to any one of claims 1 to 4, characterized in that, The test index value of each test task is obtained based on a preset index set, and the method includes: Respectively determining a single index value of each index in the preset index set for the test task; Determining a weight of each index in the preset index set; wherein the weight of at least part of the indexes in the preset index set is obtained based on a weight probability distribution of the corresponding index, and the weight probability distribution of the corresponding index is updated based on historical test data; The test index value of the test task is obtained based on the single index value of each index and the weight of each index.
8. The method according to any one of claims 1 to 4, characterized in that, The scheduling execution of the test task is performed according to the test index value of each test task, and the method includes: The test task is distributed and executed according to the test index value of each test task and a set scheduling condition as a constraint.
9. The method according to any one of claims 1 to 4, characterized in that, After the scheduling execution of the test task is performed, the method further includes: Obtaining test data of the test task; Updating an adaptive parameter related to the test index value of the test task based on the test data, and / or updating an adaptive parameter related to scheduling.
10. An electronic device, comprising: It includes: At least one processor; And A memory connected in communication with the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1-9.
11. A non-transitory computer readable storage medium, comprising: The non-transitory computer readable storage medium stores computer instructions for causing the computer to execute the method of any one of claims 1-9.
Citation Information
Patent Citations
Intelligent operation and maintenance system and method for AI server
CN119883848A
Computer self-checking period optimization method based on Internet of Things
CN120560915A
Test task scheduling method and device, electronic equipment and storage medium
CN120687371A
Data processing method and device, computing equipment, storage medium and program product
CN120743337A
Parallel fuzzy testing method and system based on variation strategy dynamic adjustment
CN120850296A