Computing power subgroup state monitoring method and system taking task as monitoring unit

By deploying standardized probe tasks in the computing power subgroups and performing intelligent analysis to generate a comprehensive health score, the problem of traditional monitoring methods being unable to assess the overall service capability of the cluster is solved, and accurate, dynamic and global monitoring of the status of the computing power subgroups is achieved.

CN121144136APending Publication Date: 2025-12-16SHANGHAI JIAJIA ZHIYUN DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511138818.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Traditional methods for monitoring computing subgroups are insufficient to comprehensively and accurately assess the overall service capabilities of a cluster and identify complex performance bottlenecks. They also fail to fully reflect the complex state of a subgroup, especially when individual node metrics are normal but internal components have coordination issues.

Method used

By deploying standardized probe tasks to computing power subgroups, simulating real-world applications and collecting performance and status indicators, a subgroup health status aggregation and evaluation mechanism is introduced to deeply integrate and intelligently analyze multi-dimensional, time-series probe data to generate a comprehensive health score.

Benefits of technology

It enables precise, dynamic, and global monitoring of the state of computing power subgroups, allowing for early detection of complex performance bottlenecks and improving the efficiency of interpreting monitoring information and the timeliness of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121144136A_ABST
    Figure CN121144136A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computing power monitoring, and particularly discloses a computing power subgroup state monitoring method and system taking a task as a monitoring unit, which actively deploys and executes a standardized probe task to a target computing power subgroup, and transfers a monitoring focus from an underlying resource index to an actual task execution efficiency. The probe tasks simulate real applications, and execution results of the probe tasks can directly reflect the real ability of the computing power subgroups to provide services to the outside. After the discrete original probe results are collected, a subgroup health state aggregation evaluation mechanism is introduced, and deep integration and intelligent analysis are carried out on multi-dimensional and sequential probe data from a plurality of probe tasks. The dynamic aggregation assessment process can fully consider complex association, historical baselines and real-time dynamic change trends among probe results, so that a comprehensive health score capable of comprehensively, accurately and dynamically representing the current overall health condition of the computing power subgroup is generated, and a health state report of the computing power subgroup is generated accordingly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computing power monitoring, and more specifically, to a method and system for monitoring the status of computing power subgroups with tasks as monitoring units. Background Technology

[0002] With the rapid development of information technology, computing resources are evolving towards large-scale, clustered, and distributed architectures, forming numerous computing subgroups of varying functions and sizes. As the core infrastructure supporting various applications and services, the stability and efficiency of these subgroups are crucial. However, traditional monitoring methods, which simply judge the success or failure of probe tasks, are often insufficient to fully reflect the complex state of the subgroup (e.g., performance degradation, component failure). In other words, the binary state of a single probe task's success / failure, or a single performance metric (such as execution time), cannot comprehensively characterize the complex health status of a computing subgroup composed of multiple components (computing, storage, network, scheduler, etc.). Furthermore, existing computing monitoring methods often focus on monitoring basic resource metrics such as CPU, memory, and network of individual physical or virtual nodes. While this approach can reflect the operation of individual nodes, it is difficult to comprehensively and accurately assess the true capability and health status of the entire computing subgroup as a whole in providing services. For example, individual node metrics may be normal, but bottlenecks or failures in network communication, task scheduling coordination, or specific service components within the subgroup may lead to a decline in the overall service quality of the subgroup.

[0003] Therefore, an optimized computing power subgroup state monitoring scheme is desired. Summary of the Invention

[0004] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide a method and system for monitoring the status of a computing power subgroup using tasks as monitoring units. By proactively deploying and executing standardized probe tasks to the target computing power subgroup, the monitoring focus shifts from underlying resource indicators to actual task execution performance. These probe tasks simulate real-world applications, and their execution results (including performance and status indicators) directly reflect the actual service capabilities of the computing power subgroup. After collecting these discrete raw probe results, instead of simple data listing or isolated analysis, a subgroup health status aggregation and evaluation mechanism is introduced to deeply integrate and intelligently analyze probe data from multiple probe tasks, across multiple dimensions, and with temporal sequence. This dynamic aggregation and evaluation process fully considers the complex correlations between probe results, historical baselines, and real-time dynamic trends, thereby generating a comprehensive health score that comprehensively, accurately, and dynamically characterizes the current overall health status of the computing power subgroup. This strategy, which uses the actual performance of tasks as the monitoring unit and combines intelligent aggregation and evaluation, effectively overcomes the technical difficulties of traditional monitoring methods that only focus on individual resources, are difficult to evaluate the overall service capabilities of the cluster, and are difficult to detect complex performance bottlenecks in the early stages. It achieves accurate, dynamic, and global monitoring of the status of computing power subgroups.

[0005] According to one aspect of this application, a method for monitoring the state of a computing power subgroup with tasks as monitoring units is provided, comprising: Based on the probe task template list, generate a schedule of probe instances to be scheduled for the first computing power subgroup; The task scheduler submits each probe task instance in the probe instance plan to the first computing power subgroup in sequence according to the planned time in the probe instance plan and the current time. The result collector obtains the raw probe results from the first computing power subgroup to obtain a set of raw probe results, which includes performance indicators and status indicators. Based on the set of original probe results, a subgroup health status aggregation assessment is performed to obtain the comprehensive health score of the first computing power subgroup. A health status report for the first computing power subgroup is generated based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold.

[0006] According to another aspect of this application, a computing power subgroup status monitoring system with tasks as monitoring units is provided, comprising: The scheduling plan generation module is used to generate a schedule of probe instances to be scheduled for the first computing power subgroup based on the probe task template list. The instance planning and scheduling module is used to use the task scheduler to submit each probe task instance in the probe instance plan to be scheduled to the first computing power subgroup in sequence according to the planned time in the probe instance plan to be scheduled and the current time. The raw probe result acquisition module is used to obtain raw probe results from the first computing power subgroup using a result acquisition device to obtain a set of raw probe results, wherein the raw probe results include performance indicators and status indicators. The subgroup health status comprehensive assessment module is used to perform subgroup health status aggregation assessment based on the set of original probe results to obtain the comprehensive health score of the first computing power subgroup. The subgroup health status report generation module is used to generate a health status report for the first computing power subgroup based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold.

[0007] Compared with existing technologies, this application provides a method and system for monitoring the status of computing power subgroups using tasks as monitoring units. By proactively deploying and executing standardized probe tasks to the target computing power subgroup, it shifts the monitoring focus from underlying resource indicators to actual task execution performance. These probe tasks simulate real-world applications, and their execution results (including performance and status indicators) directly reflect the actual service capabilities of the computing power subgroup. After collecting these discrete raw probe results, instead of simple data listing or isolated analysis, a subgroup health status aggregation and evaluation mechanism is introduced to deeply integrate and intelligently analyze probe data from multiple probe tasks, across multiple dimensions, and with temporal sequence. This dynamic aggregation and evaluation process fully considers the complex correlations between probe results, historical baselines, and real-time dynamic trends, thereby generating a comprehensive health score that can comprehensively, accurately, and dynamically characterize the current overall health status of the computing power subgroup. This strategy, which uses the actual performance of tasks as the monitoring unit and combines intelligent aggregation and evaluation, effectively overcomes the technical difficulties of traditional monitoring methods that only focus on individual resources, are difficult to evaluate the overall service capabilities of the cluster, and are difficult to detect complex performance bottlenecks in the early stages. It achieves accurate, dynamic, and global monitoring of the status of computing power subgroups. Attached Figure Description

[0008] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0009] Figure 1 This is a flowchart of a computing power subgroup status monitoring method based on tasks as monitoring units, according to an embodiment of this application. Figure 2 This is a schematic diagram of data flow in a computing power subgroup status monitoring method based on tasks as monitoring units, according to an embodiment of this application. Figure 3 This is a flowchart illustrating the process of performing a comprehensive health score assessment of a first computing power subgroup based on the set of original probe results in a computing power subgroup status monitoring method with tasks as monitoring units, according to an embodiment of this application. Figure 4 This is a flowchart illustrating the dynamic aggregation analysis of the set of original probe result encoding vectors to obtain the dynamic aggregation encoding vector of the original probe results, according to the computing power subgroup state monitoring method based on tasks as monitoring units in this application embodiment. Figure 5 This is a block diagram of a computing subgroup status monitoring system based on tasks as monitoring units, according to an embodiment of this application. Detailed Implementation

[0010] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0011] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0012] While this application makes various references to certain modules of the systems according to embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The modules described are merely illustrative, and different aspects of the systems and methods may use different modules.

[0013] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0014] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0015] With the large-scale clustering of computing resources, traditional monitoring methods focus on single-node resources or simple probe task success or failure judgments, making it difficult to comprehensively assess the overall service capabilities and complex health status of computing subgroups. The normal operation of individual nodes does not mean that the subgroup as a whole is safe, because the coordination problems of its internal components (such as network and scheduling) may lead to a decline in service quality. Therefore, there is an urgent need for an optimized monitoring scheme that can accurately reflect the comprehensive status of the subgroup.

[0016] To address the aforementioned technical issues, this application proposes a method for monitoring the status of computing power subgroups using tasks as monitoring units. By deploying standardized probe tasks to the computing power subgroups, the monitoring focus shifts from underlying resources to actual task execution efficiency, with the results directly reflecting the subgroup's true service capability. The core of this method lies in introducing a subgroup health status aggregation and evaluation mechanism. This mechanism deeply integrates and intelligently analyzes multi-dimensional, time-series probe data, comprehensively considering the correlation of various results, historical baselines, and dynamic trends to generate a comprehensive subgroup health score. This achieves accurate, dynamic, and global monitoring of the computing power subgroup status, overcoming the shortcomings of traditional methods in assessing overall service capability and identifying complex performance bottlenecks.

[0017] Figure 1 This is a flowchart of a computing power subgroup status monitoring method based on tasks as monitoring units, according to an embodiment of this application. Figure 2 This is a schematic diagram of data flow in a computing power subgroup state monitoring method based on tasks as monitoring units, according to an embodiment of this application. Figure 1 and Figure 2 As shown, the computing power subgroup status monitoring method based on tasks as monitoring units according to an embodiment of this application includes the following steps: S100, generating a plan of probe instances to be scheduled for a first computing power subgroup based on a probe task template list; S200, the task scheduler submits each probe task instance in the plan of the probe instances to be scheduled to the first computing power subgroup in sequence according to the planned time in the plan of the probe instances to be scheduled and the current time; S300, the result collector obtains the original probe results from the first computing power subgroup to obtain a set of original probe results, the original probe results including performance indicators and status indicators; S400, performing a subgroup health status aggregation evaluation based on the set of original probe results to obtain a comprehensive health score for the first computing power subgroup; S500, generating a health status report for the first computing power subgroup based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold.

[0018] Specifically, in step S100, a schedule of probe instances for the first computing power subgroup is generated based on the probe task template list. It should be understood that the probe task templates predefine various representative computing, storage, or network load patterns, ensuring the consistency of the monitoring baseline; customization allows for the selection of appropriate templates from the template list and configuration of parameters based on the specific characteristics of the first computing power subgroup (such as the type of application it supports, resource allocation, etc.) and monitoring needs (such as monitoring frequency, specific performance dimensions of interest, etc.), generating targeted probe instances; by generating the schedule, a foundation is laid for the subsequent task scheduler to automatically execute monitoring tasks on time and as needed, reducing manual intervention and improving monitoring efficiency and coverage.

[0019] Specifically, in a concrete example of this application, firstly, a probe task template library is constructed and maintained. Each template defines in detail the type of probe task, its execution logic, the type and range of required input parameters, and the expected output metrics (such as CPU-intensive task templates, IO-intensive task templates, network latency test templates, etc.). Secondly, when generating a monitoring plan for a specific computing power subgroup (i.e., the first computing power subgroup), the configuration module or user selects one or more suitable probe task templates from the template library based on the characteristics of the subgroup and the monitoring strategy. For example, if the first computing power subgroup mainly undertakes AI training tasks, then probe templates simulating deep learning model training segments are preferred. Next, the selected templates are instantiated and configured, that is, specific values ​​are assigned to the parameters defined in the template or a data source is specified, and scheduling-related attributes such as the execution frequency, execution time window, and timeout threshold of the probe instance are set. For example, a target IP address list is specified for the network latency test template, and it is set to execute once every 5 minutes. Finally, all configured probe instances and their scheduling information are integrated to form a structured plan of probe instances to be scheduled. The plan is presented as a time-series task list, where each entry clearly indicates when and with what configuration a probe task will be executed. This plan is then submitted to the task scheduler, which is responsible for distributing the corresponding probe task instances to the first computing power subgroup for execution at the specified times, dynamically adapting to changes in the state of the computing power subgroup and monitoring needs.

[0020] Specifically, in step S200, the task scheduler submits each probe task instance in the planned probe instance to the first computing power subgroup sequentially based on the planned time in the planned probe instance and the current time. It should be understood that by comparing the planned time in the planned probe instance with the current time and submitting probe task instances sequentially, the timeliness and continuity of monitoring data can be ensured, providing a reliable and timestamped basis for subsequent health status assessments. This not only avoids the inefficiency and potential errors of manual intervention and automates the monitoring process, but also effectively manages the probe load on the computing power subgroup through fine-grained control of the task submission process, preventing unnecessary interference to normal business operations due to over-probing, thereby achieving harmonious coexistence between monitoring and business operations.

[0021] More specifically, in a concrete example of this application, the core responsibility of this service, built as a continuously running background service or module, is to drive the lifecycle of probe tasks based on a pre-generated scheduler plan for probe instances. The service first loads and parses the scheduler plan, transforming it into an internally processable task queue or data structure containing a unique identifier for each probe instance, its planned execution time, associated probe template parameters, and target computing power subgroup information. The scheduler maintains a high-precision time synchronization mechanism, continuously monitoring the current system time and comparing it in real-time with the scheduled execution times of each probe instance in the plan. When the planned execution time of a probe instance reaches or exceeds the current time, the scheduler marks it as pending submission. Subsequently, the scheduler interacts with the control plane or API interface of the first computing power subgroup, submitting the specific execution instructions (including probe type, parameters, execution environment, etc.) of the probe task instance to that subgroup. This is typically achieved through standard communication protocols such as Remote Procedure Call (RPC), message queues (e.g., Kafka, RabbitMQ), or HTTP / HTTPS requests. To ensure orderly submissions and system stability, the scheduler incorporates concurrency control and retry mechanisms to handle network fluctuations or sudden spikes in subgroup load. After a task is successfully submitted, the scheduler updates the instance's status and awaits feedback from the subsequent result collector. In more advanced implementations, the task scheduler can integrate a machine learning-based decision-making model. This model doesn't directly execute tasks but acts as an intelligent decision-making unit, dynamically adjusting the submission frequency or priority of probe tasks based on multiple factors such as the real-time load of the computing subgroup, historical performance trends, and the importance of the probe tasks. It can even optimize submission batches to maximize monitoring efficiency while minimizing the impact on the production environment. For example, a reinforcement learning model can learn the optimal probe task submission strategy under different system states, thereby achieving adaptive scheduling.

[0022] Specifically, in step S300, the result collector obtains the raw probe results from the first computing power subgroup to obtain a set of raw probe results, which includes performance indicators and status indicators. It should be understood that the execution of the probe task itself cannot directly reveal the status of the computing power subgroup; the data produced by its execution, i.e., the raw probe results, must be effectively collected and aggregated to provide factual basis for subsequent analysis and evaluation. Therefore, in the technical solution of this application, the raw probe results are further obtained from the first computing power subgroup to obtain a set of raw probe results.

[0023] Specifically, in this embodiment, the result collector obtains raw probe results from the first computing power subgroup to obtain a set of raw probe results. This includes: the result collector obtaining status indicators and raw performance indicators from the first computing power subgroup through a periodic polling or callback mechanism. The status indicators include success or failure, and the raw performance indicators include actual task execution time, average network latency, network packet loss rate, throughput, number of processed entries, and error rate of internal operations. This is primarily to ensure the integrity and originality of the monitoring data. Through the periodic polling or callback mechanism, the result collector can timely and systematically pull or receive the output of each completed probe task from the first computing power subgroup. These outputs are defined as raw probe results, containing key status indicators (such as task success or failure) and a series of raw performance indicators (such as actual task execution time, average network latency, network packet loss rate, throughput, number of processed entries, and error rate of internal operations). This comprehensive data collection strategy aims to capture the multi-dimensional performance of the computing power subgroup in responding to standardized tasks, avoiding the one-sidedness of relying solely on simple success / failure judgments or single performance indicators. Secondly, forming a "set of raw probe results" lays the foundation for subsequent batch processing and aggregation analysis, and unifies the management of discrete, multi-source probe feedback.

[0024] More specifically, in this embodiment of the application, the result collector obtains the original probe results from the first computing power subgroup to obtain a set of original probe results, and further includes: The performance indicators in the original performance indicators are normalized to obtain the performance indicators, wherein the normalization process is expressed as follows: ;in, Baseline mean For baseline standard deviation, These are the various performance indicators in the original performance metrics, and N is the tolerance factor. This represents the original deviation ratio. For the initial score, For performance metrics.

[0025] Specifically, in step S400, a subgroup health status aggregation assessment is performed based on the set of original probe results to obtain the comprehensive health score of the first computing power subgroup. It should be understood that the original probe results are discrete and multi-dimensional, and directly interpreting this scattered data makes it difficult to form an intuitive and accurate judgment of the overall health status of the computing power subgroup. Furthermore, the success / failure of a single probe or a single performance indicator cannot comprehensively characterize the health status of a complex computing power subgroup. Therefore, a systematic approach is necessary to integrate and refine information from different probe tasks, reflecting different aspects. This transforms the massive, heterogeneous original probe data into a single quantitative indicator—the comprehensive health score—that can comprehensively, accurately, and dynamically reflect the overall operating status of the first computing power subgroup. This is not only to simplify the presentation of monitoring information, but more importantly, through aggregation assessment, it can reveal potential correlations and systemic problems that are difficult to detect with a single indicator. For example, a slight decline in multiple performance indicators simultaneously may indicate a decline in overall service capability, a trend that is easily overlooked when observed in isolation. Through aggregation assessment, the aim is to grasp the true service capability and potential risks of the computing power subgroup at a macro level, providing a scientific basis for operation and maintenance decisions. This comprehensive health score intuitively reflects the current health level of the first computing power subgroup, enabling administrators to quickly determine whether it is in a healthy, degraded, alarmed, or critical alarm state. This not only greatly improves the efficiency of interpreting monitoring information and the timeliness of decision-making, but also provides key input for achieving automated operation and maintenance, trend prediction, and capacity planning. In this way, this technical solution effectively solves the problems of traditional monitoring methods, which are difficult to comprehensively assess the overall service capabilities of the cluster and difficult to detect complex and interconnected performance bottlenecks in the early stages.

[0026] Figure 3 This is a flowchart illustrating the process of a task-based monitoring method for monitoring the status of a computing power subgroup according to an embodiment of this application. The method involves aggregating and evaluating the health status of the subgroup based on the set of original probe results to obtain a comprehensive health score for the first computing power subgroup. (See flowchart for example.) Figure 3 As shown, according to an embodiment of this application, the method for monitoring the status of a computing power subgroup with tasks as monitoring units includes step S400, which includes: S410, performing one-hot embedding encoding on the status indicators of each original probe result in the set of original probe results to obtain a one-hot embedding encoding vector for the status indicators; S420, forming a performance indicator vector from the performance indicators of each original probe result in the set of original probe results; S430, forming an original probe result encoding vector from the one-hot embedding encoding vector for the status indicators and the performance indicator vector to obtain a set of original probe result encoding vectors; and S440, performing a subgroup health status aggregation evaluation based on the set of original probe result encoding vectors to obtain a comprehensive health score for the first computing power subgroup.

[0027] Specifically, in step S410, one-hot embedding encoding is performed on the state indicators of each original probe result in the set of original probe results to obtain a one-hot embedding encoding vector for the state indicators. It should be understood that state indicators (such as "success" or "failure") are essentially categorical data rather than numerical data. Directly inputting these categorical labels (e.g., assigning "success" as 0 and "failure" as 1) into many machine learning models or performing numerical calculations may introduce misinterpretations of the order or magnitude relationships that do not exist between these labels. For example, the model might mistakenly assume that "failure" (1) is numerically greater than "success" (0) and make inappropriate weight assignments or distance calculations accordingly. Therefore, the one-hot embedding encoding technique can convert these discrete, unordered categorical state indicators into a high-dimensional sparse numerical vector form, so that each state category is independently represented as a dimension, thereby eliminating potential ordinal misunderstandings and ensuring that the model can treat each state fairly. This encoding method allows state information to exist in a format that is more compatible with subsequent numerical analysis and machine learning models, facilitating effective integration and calculation with other numerical performance indicators. For example, when constructing a comprehensive health score, different states can be assigned independent weights or influence factors.

[0028] Specifically, in step S420, the performance metrics of each original probe result in the set of original probe results are combined into a performance metric vector. It should be understood that each probe task generates multiple performance metrics of different dimensions (such as actual task execution time, average network latency, network packet loss rate, throughput, number of processed entries, and internal operation error rate). These metrics collectively reflect the specific performance of the probe task when executed on a computing power subgroup. In order to systematically and uniformly process this multi-dimensional performance information and use it as input to the subsequent aggregation evaluation model, it is necessary to organize them into a structured vector. By arranging and combining the various performance metrics into a vector, the originally scattered multiple scalar values ​​can be integrated into a single mathematical object. The benefits of this approach include: First, it facilitates unified mathematical operations and transformations, such as subsequent normalization, feature scaling, or concatenation with other vectors (e.g., one-hot encoded vectors of state indicators); second, vectorized representation is the standard input format for most machine learning algorithms and data analysis models, enabling these advanced analytical tools to be directly applied to performance data; and third, vectorization allows for a clearer definition and computation of similarities or differences between different probe results, laying the foundation for more complex pattern recognition and trend analysis.

[0029] Specifically, in a concrete example of this application, firstly, for each raw probe result, the types and order of all raw performance metrics it contains are clearly defined. For example, it is agreed that the first element of the performance metric vector represents the actual execution time of the task, the second element represents the average network latency, and so on, until the last performance metric (such as the error rate of internal operations). Once this order is determined, it should be kept consistent when processing all probe results to ensure the uniformity of the meaning of the corresponding dimensions of the vector.

[0030] Secondly, for each raw probe result, the values ​​of each raw performance metric are extracted according to a predefined order. Then, these extracted values ​​are arranged sequentially to form a numerical vector. For example, if a probe task has an execution time of 100ms, an average network latency of 20ms, a packet loss rate of 0.1%, a throughput of 1000 packets / second, a processing count of 5000 items, and an internal error rate of 0.01%, then its corresponding performance metric vector can be represented as [100, 20, 0.001, 1000, 5000, 0.0001]. Finally, the above operation is performed on each probe result in the raw probe result set, resulting in a set of multiple performance metric vectors, each vector representing a multi-dimensional performance profile of a probe task.

[0031] Specifically, in step S430, the one-hot embedding encoding vector of the state index and the performance index vector are combined to form the original probe result encoding vector to obtain a set of original probe result encoding vectors. It should be understood that this integrates two different types of closely related information about the probe task—the final success or failure state of the task (classification information) and the specific performance during task execution (numerical information)—into a unified and comprehensive feature representation. These two types of information reflect the state of the computing power subgroup when executing the probe task from different perspectives; considering either one in isolation may lead to a one-sided understanding of the subgroup's state. For example, a task may be successfully completed, but its performance index may be far below expectations, which also indicates a potential problem.

[0032] More specifically, in this embodiment, the process of combining the one-hot embedding encoding vector of the state indicator and the performance indicator vector to form a set of original probe result encoding vectors includes: concatenating the one-hot embedding encoding vector of the state indicator and the original probe result encoding vector to obtain the original probe result encoding vector. By combining these two encoded vectors, the model can simultaneously learn and utilize the interactions and complex dependencies between state information and performance information. For example, the model can learn that under a specific combination of performance indicators, the probability of task failure is higher, or that even if the task succeeds, certain patterns in some performance indicators may predict a decline in subgroup health. This comprehensive feature representation helps improve the accuracy and robustness of the evaluation model, enabling it to more accurately capture subtle state changes in the computing power subgroup.

[0033] Specifically, in step S440, a subgroup health status aggregation assessment is performed based on the set of original probe result encoding vectors to obtain the comprehensive health score of the first computing power subgroup. It should be understood that although each original probe result encoding vector integrates the status and performance information of a single probe, these are still discrete, microscopic observation points. Traditional methods struggle to assess the overall service capability of the entire computing power subgroup and identify complex performance bottlenecks. Therefore, a further subgroup health status aggregation assessment is performed based on the set of original probe result encoding vectors to extract a unified judgment of the overall subgroup health status from a macroscopic perspective, rather than simply examining the performance of individual probes.

[0034] More specifically, in a specific example of this application, the subgroup health status aggregation assessment based on the set of original probe result encoding vectors to obtain the comprehensive health score of the first computing power subgroup includes: performing dynamic aggregation analysis on the set of original probe result encoding vectors to obtain a dynamic aggregation encoding vector of original probe results; and performing sequence decoding regression on the dynamic aggregation encoding vector of original probe results to obtain the comprehensive health score of the first computing power subgroup. Through subgroup health status aggregation assessment, the complex spatiotemporal dependencies and mutual influences between different probe result encoding features can be captured, key feature combinations and changing trends can be identified, thereby generating a deeper feature representation that can highly summarize the overall state of the current subgroup, namely, the dynamic aggregation encoding vector of original probe results. Subsequently, the dynamic aggregation encoding vector of original probe results is subjected to sequence decoding regression to obtain the comprehensive health score of the first computing power subgroup. This can map this highly condensed, possibly sequence-characteristic dynamic aggregation encoding vector to a continuous, standardized health score value through a regression model (such as support vector regression, gradient boosting regression tree, or the regression head of a deep neural network). This score intuitively quantifies the health level of the subgroup. More specifically, performing sequence decoding regression on the dynamically aggregated encoding vector of the original probe results to obtain the comprehensive health score of the first computing power subgroup includes: passing the dynamically aggregated encoding vector of the original probe results through a decoder-based computing power subgroup state analyzer to obtain the decoded value of the comprehensive health score of the first computing power subgroup.

[0035] Specifically, a dynamic aggregation analysis is performed on the set of original probe result encoding vectors to obtain a dynamically aggregated encoding vector for the original probe results. It should be understood that although each original probe result encoding vector integrates the state and multi-dimensional performance information of a single probe task, these vectors are still discrete and independent observations at different time points. The health status of a computing power subgroup is a dynamically evolving whole influenced by multiple complex factors. Its true service capability and potential problems are often reflected in the collaborative performance and interrelationships of multiple probe tasks at different time points, rather than being fully revealed by isolated individual data. Therefore, in the technical solution of this application, a steady-state expression capable of capturing distributed feature consensus and simultaneously capturing the uniqueness of individual features is constructed by performing dynamic aggregation analysis on the set of original probe result encoding vectors. Specifically, an original feature reference baseline encoding vector that can characterize the consensus of data distribution is established, and the real-time offset of each input probe result encoding vector relative to this baseline is quantified. By dynamically calculating probability weights and asymmetric correction coefficients, this analysis process can adaptively differentiate and enhance the baseline encoding of the original probe results. This allows the analysis to highlight key offset features that indicate anomalies or performance degradation while preserving the overall performance consensus of the computing power subgroup, thus enabling the early detection of performance bottlenecks with complex correlations.

[0036] Figure 4 This is a flowchart illustrating the dynamic aggregation analysis of the set of original probe result encoding vectors to obtain dynamically aggregated encoding vectors for a computing subgroup state monitoring method based on tasks as monitoring units, according to an embodiment of this application. For example... Figure 4 As shown, the computing subgroup status monitoring method based on tasks as monitoring units according to an embodiment of this application performs dynamic aggregation analysis on the set of original probe result encoding vectors to obtain a dynamic aggregated encoding vector of the original probe results, including: S441, performing feature baseline learning on the set of original probe result encoding vectors to obtain a reference base encoding vector of the original probe results; S442, calculating the real-time offset of the set of original probe result encoding vectors relative to the reference base encoding vector of the original probe results to obtain a set of asymmetric correction coefficients for the original probe result encoding; S443, based on the set of asymmetric correction coefficients for the original probe result encoding and the set of original probe result encoding vectors, performing dynamic compensation on the reference base encoding vector of the original probe results to obtain the dynamic aggregated encoding vector of the original probe results.

[0037] More specifically, in step S441, feature baseline learning is performed on the set of original probe result encoding vectors to obtain the original probe result reference base encoding vector, expressed by the formula: ; in, The set of encoded vectors for the original probe results. These are respectively the 1st, 2nd, and 3rd elements in the set of encoded vectors of the original probe results. The and the first A raw probe result encoding vector, For learnable parameter matrix, for Activation function This is the reference base encoding vector for the original probe results.

[0038] It is understandable that to accurately assess the current health status of the first computing power subgroup, a reference standard for its "normal" or "expected" operating state must first be established. While the original probe result encoding vectors integrate the status and performance information of individual probes, these vectors themselves are dynamically changing, and directly interpreting this real-time data makes it difficult to determine whether they deviate from the normal state. Traditional monitoring struggles to comprehensively reflect the overall health status, partly due to the lack of a dynamically adaptable reference that represents the cluster's "health baseline." Therefore, in the technical solution of this application, feature baseline learning is performed on the set of original probe result encoding vectors to extract and generate an original probe result reference baseline encoding vector that represents the probe task response characteristics of the first computing power subgroup under healthy and stable operating conditions. This reference baseline encoding vector is not a simple average value, but rather obtained through feature learning, capable of capturing the core distribution characteristics, statistical equilibrium points, or a kind of "distribution consensus" of probe results under normal operating conditions. It aims to provide a stable and representative comparison base for subsequent dynamic aggregation analysis, serving as an "anchor" for measuring the degree of deviation of real-time probe results. By establishing such an adaptive benchmark, a foundation is laid for subsequent identification of minor but critical performance degradation or state changes (i.e., calculating the offset from the benchmark), thereby more effectively discovering the complex performance bottlenecks and overall service capability changes mentioned in the technical context.

[0039] Accordingly, according to an embodiment of this application, step S442, calculating the real-time offset of the set of original probe result encoding vectors relative to the original probe result reference base encoding vector to obtain the set of original probe result encoding asymmetric correction coefficients, includes: calculating the offset quantization factor of each original probe result encoding vector in the set of original probe result encoding vectors relative to the original probe result reference base encoding vector to obtain the set of original probe result encoding offset quantization factors; and performing regularization processing based on the Softmax activation function on the set of original probe result encoding offset quantization factors to obtain the set of original probe result encoding asymmetric correction coefficients.

[0040] More specifically, the offset quantization factor of each original probe result encoding vector in the set of original probe result encoding vectors relative to the original probe result reference base encoding vector is calculated to obtain the set of original probe result encoding offset quantization factors, expressed by the formula: ;in, Let L be the L2 norm of the vector. Let the first norm of the vector be 1. For positional product, for Activation function for function, The set of offset quantization factors encoding the original probe results. The original probe result encodes an offset quantization factor.

[0041] It is understandable that simply having a benchmark is insufficient; the key lies in how to utilize this benchmark to interpret each real-time generated probe data. Each new "raw probe result encoding vector" carries the instantaneous state and performance information of the computing power subgroup under a specific probe task. Therefore, by calculating the offset quantization factor of each raw probe result encoding vector in the set of raw probe result encoding vectors relative to the raw probe result reference benchmark encoding vector, and comparing each new raw probe result encoding vector with the established raw probe result reference benchmark encoding vector, the degree of deviation and difference of the current state from "ideal" or "normal" can be quantified, and these differences can be transformed into meaningful offset quantization factors or raw probe result encoding offset quantization factors. These factors aim to capture the real-time offset of the current probe result relative to the reference benchmark in a multi-dimensional feature space. Here, the offset quantization factor is essentially a numerical expression of the difference between the current probe performance and the benchmark performance. This makes subsequent dynamic aggregation analysis no longer a simple averaging or summarizing of raw data, but rather enables more intelligent and targeted fusion based on these quantified offset information. For example, a quantization factor with a large negative offset (significant performance degradation) might be given higher attention or weight during aggregation. This allows for dynamic adjustment of the evaluation model's requirements through real-time compensation, making the system more sensitive to minor performance fluctuations or abnormal patterns in computing power subgroups. This enables more effective identification of potential risks and performance bottlenecks, ultimately improving the accuracy and timeliness of overall health status assessments.

[0042] More specifically, the set of original probe result encoding offset quantization factors is regularized based on the Softmax activation function to obtain the set of original probe result encoding asymmetric correction coefficients, expressed by the formula: ;in, For preset parameter values, Let e ​​be the value of an exponential function with the natural constant e as its base. The set of asymmetric correction coefficients encoded for the original probe results. The original probe results encode asymmetric correction coefficients.

[0043] It is understandable that although the offset quantization factor quantifies the offset of each original probe result encoding vector relative to the reference baseline, these original offsets themselves may have inconsistent scales, lack a direct measure of relative importance, and may not be able to effectively distinguish which offsets are more critical to evaluating the overall health of the computing power subgroup when directly used for subsequent feature fusion or correction. Therefore, the set of original probe result encoding offset quantization factors, which may be widely distributed, is further transformed into a set of original probe result encoding asymmetric correction coefficients with probability distribution characteristics through the Softmax activation function. The Softmax function can transform the input set of original probe result encoding offset quantization factors into a vector with element values ​​between 0 and 1 and a sum of all elements of 1. This allows the output correction coefficients to be understood as the "relative importance" or "contribution weight" of each probe result's offset to the overall state evaluation. More importantly, Softmax has the characteristic of amplifying differences. It assigns higher correction coefficients to probe results with larger (more significant) compensation factors while suppressing probe results with smaller compensation factors. This allows the system to adaptively focus on feature offsets with high discriminative power for the current task, while suppressing redundant feature interference, thus achieving dynamic calibration of feature importance. These initial probe results encode asymmetric correction coefficients, providing clear, data-driven weighting guidance for subsequent dynamic aggregation analysis or benchmark adjustments. Furthermore, their asymmetric nature ensures the system can differentiate its focus and impact based on the actual deviations of different probe results from the benchmark. This means the system can more intelligently focus on probe results and their offset characteristics that best indicate changes in the health status of computing power subgroups, thereby more effectively extracting key information from massive probe data. This directly solves the problem of traditional methods struggling to capture complex correlation performance bottlenecks, making the final comprehensive health score more accurately reflect the true and subtle changes in the operational status of computing power subgroups.

[0044] More specifically, in step S443, based on the set of asymmetric correction coefficients for the original probe results and the set of encoding vectors for the original probe results, dynamic compensation is performed on the reference encoding vector for the original probe results to obtain the dynamic aggregate encoding vector for the original probe results, expressed by the formula: ;in, The original probe results are dynamically aggregated into an encoded vector.

[0045] It is understandable that the original probe result reference base encoding vector represents the historical "normal" or "consensus" health level of the computing power subgroup, while the real-time original probe result encoding vector reflects the current specific performance, which may deviate from the base. To obtain a comprehensive characterization that accurately reflects the overall health of the current subgroup, we cannot rely entirely on historical bases (which may be outdated or unable to reflect instantaneous shocks), nor can we simply aggregate real-time original probe result encoding data (which may ignore historical experience and noise interference). Therefore, we further utilize the already calculated asymmetric correction coefficients of the original probe result encoding that reflect the importance of the deviations of each probe result, as well as the original probe result encoding vector containing specific performance and status information, to dynamically and purposefully correct the original probe result reference base encoding vector. Here, "dynamic compensation" means that the reference benchmark is no longer a static anchor point, but is adjusted according to the weighted influence of the real-time probe result encoded data (the weights are asymmetric correction coefficients, the influence of which comes from the original probe result encoded vector itself or its difference from the benchmark). This generates a "dynamically aggregated encoded vector of the original probe result" that includes both historical experience (the "distribution consensus" of the benchmark) and can keenly capture current key changes (the "unique contribution" and "high-discrimination feature shift" of individual features). This process aims to form a more comprehensive, accurate, and timely deep feature representation of the current health status of the computing power subgroup.

[0046] Preferably, in a dynamically evolving computing cluster environment, to accurately capture subtle changes in the health status of subgroups, it is necessary to ensure that the mapping from the original probe result encoding vector to the original probe result reference encoding vector has mathematical smoothness. This continuity requirement stems from the complexity of real-world business scenarios—frequent instantaneous interferences such as network jitter and load spikes occur, while a smooth mapping ensures the generalization ability of the correction coefficient in fluctuating environments. Therefore, a differentiable gradient vector of the offset quantization factor is first constructed: This formula uses a sigmoid transform based on the binorm difference to represent the real-time probe features. With health benchmarks Deviations are transformed into smooth gradients. This design enables the system to detect abnormal patterns such as spikes in network latency and drops in computational throughput, while avoiding noise interference.

[0047] Then calculate the covariance matrix. ,in It is a row vector, the covariance matrix It can describe the geometric relationship between gradients, and the covariance matrix can... F-norm It can serve as a key parameter for representing the steady state of a cluster, quantifying the steady-state strength of the cluster's characteristic space. When the load is balanced, the norm value tends to stabilize; if a local failure occurs (such as an anomaly in a compute node), the norm fluctuation will trigger a correction mechanism, affecting the correction coefficient in the following ways: .

[0048] This operation provides dual safeguards: it ensures the stability of the correction process through differentiable query mapping, while utilizing compressed steady-state norm constraints to prevent minor perturbations from causing abnormal query space coefficient mappings and suppressing feature distortion caused by transient jitter. For example, when a probe task experiences anomalies due to network interruptions, this mechanism can automatically reduce its weight to avoid misjudging the cluster status.

[0049] Therefore, by constructing asymmetric correction coefficients based on a continuously differentiable gradient distribution and performing dynamic correction on the original feature vector based on asymmetry, asymmetric feature space projection can be used to weight and fuse input features through asymmetric correction coefficients, unlike simple linear combinations. This allows for the adaptive fusion of the unique contributions of each feature while preserving distributional consensus. This fusion mechanism, through the geometric correlation constraints of the covariant relation matrix, further ensures that the corrected features inherit both the global characteristics of the benchmark and dynamically incorporate offset discriminative feature information, ultimately forming a dimensionality-reduced steady-state representation with strong representational capabilities.

[0050] Specifically, in step S500, a health status report for the first computing power subgroup is generated based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold. It should be understood that the comprehensive health score of the first computing power subgroup is itself a continuous value. While it can accurately quantify the health level of the subgroup, mapping it to a set of predefined, discrete status levels with clear business implications is crucial for rapid decision-making by operations personnel and for the response of automated systems.

[0051] More specifically, in a specific example of this application, a health status report for the first computing power subgroup is generated based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold. This includes: setting a health status threshold of 0.9, a degradation threshold range of 0.7-0.9, an alarm threshold range of 0.5-0.7, and a severe alarm threshold of 0.5; if the comprehensive health score of the first computing power subgroup is greater than the health status threshold, the first computing power subgroup is in a healthy state; if the comprehensive health score of the first computing power subgroup is within the degradation threshold range, the first computing power subgroup is in a downgraded state; if the comprehensive health score of the first computing power subgroup is within the alarm threshold range, the first computing power subgroup is in an alarm state; and if the comprehensive health score of the first computing power subgroup is less than the severe alarm threshold, the first computing power subgroup is in a severe alarm state. To concretize and contextualize the abstract comprehensive health score, the system compares it with preset thresholds (such as a health status threshold of 0.9, a degradation threshold range of 0.7-0.9, an alarm threshold range of 0.5-0.7, and a critical alarm threshold of 0.5) to clearly determine the current health status of the first computing power subgroup—whether it is healthy, degraded, alarmed, or critically alarmed. This classification not only facilitates human understanding but, more importantly, provides a clear logical basis for triggering automated operation and maintenance strategies (such as alarm escalation, resource scheduling, and fault isolation). The purpose of generating a health status report is to present this determination result and possible supporting data (such as the current comprehensive health score and key impact indicators) in a structured manner, making it a standardized, transferable, and archiveable information unit. This allows the operation and maintenance team to quickly grasp the overall operating status of the computing power subgroup without needing to deeply interpret raw data or complex model outputs to determine the severity of problems and take corresponding measures. For example, when the report shows a status of "alarm," contingency plans can be immediately activated to investigate potential risks. At the same time, these structured health status reports also provide reliable data support for long-term trend analysis, performance bottleneck tracing, and service level agreement (SLA) assessment.

[0052] In summary, the computing power subgroup status monitoring method based on tasks as monitoring units according to the embodiments of this application is explained. It shifts the monitoring focus from underlying resource indicators to actual task execution performance by proactively deploying and executing standardized probe tasks to the target computing power subgroup. These probe tasks simulate real-world applications, and their execution results (including performance and status indicators) directly reflect the actual service capabilities of the computing power subgroup. After collecting these discrete raw probe results, instead of simple data listing or isolated analysis, a subgroup health status aggregation and evaluation mechanism is introduced to deeply integrate and intelligently analyze probe data from multiple probe tasks, multiple dimensions, and time-series data. This dynamic aggregation and evaluation process fully considers the complex correlations between probe results, historical baselines, and real-time dynamic trends, thereby generating a comprehensive health score that comprehensively, accurately, and dynamically characterizes the current overall health status of the computing power subgroup. This strategy, which uses actual task execution performance as the monitoring unit and combines it with intelligent aggregation and evaluation, effectively overcomes the technical difficulties of traditional monitoring methods that only focus on individual resources, struggle to evaluate the overall service capabilities of the cluster, and fail to detect complex performance bottlenecks early on. It achieves accurate, dynamic, and global monitoring of the computing power subgroup status.

[0053] Furthermore, a computing subgroup status monitoring system with tasks as monitoring units is also provided.

[0054] Figure 5 This is a block diagram of a computing subgroup status monitoring system based on tasks as monitoring units, according to an embodiment of this application. Figure 5 As shown, the computing power subgroup status monitoring system 500 based on tasks as monitoring units according to an embodiment of this application includes: a scheduling plan generation module 510, used to generate a scheduler plan for a first computing power subgroup based on a list of probe task templates; an instance plan scheduling module 520, used to use a task scheduler to sequentially submit each probe task instance in the scheduler plan to the first computing power subgroup according to the scheduled time and the current time; an original probe result acquisition module 530, used to use a result acquisition device to obtain original probe results from the first computing power subgroup to obtain a set of original probe results, the original probe results including performance indicators and status indicators; a subgroup health status comprehensive evaluation module 540, used to perform subgroup health status aggregation evaluation based on the set of original probe results to obtain a comprehensive health score for the first computing power subgroup; and a subgroup health status report generation module 550, used to generate a health status report for the first computing power subgroup based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold.

[0055] As described above, the computing subgroup status monitoring system 500 based on tasks as monitoring units according to embodiments of this application can be implemented in various wireless terminals, such as servers with computing subgroup status monitoring algorithms based on tasks as monitoring units. In one possible implementation, the computing subgroup status monitoring system 500 based on tasks as monitoring units according to embodiments of this application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the computing subgroup status monitoring system 500 based on tasks as monitoring units can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the computing subgroup status monitoring system 500 based on tasks as monitoring units can also be one of many hardware modules of the wireless terminal.

[0056] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for monitoring the status of a computing subgroup using tasks as monitoring units, characterized in that, include: Based on the probe task template list, generate a schedule of probe instances to be scheduled for the first computing power subgroup; The task scheduler submits each probe task instance in the probe instance plan to the first computing power subgroup in sequence according to the planned time in the probe instance plan and the current time. The result collector obtains the raw probe results from the first computing power subgroup to obtain a set of raw probe results, which includes performance indicators and status indicators. Based on the set of original probe results, a subgroup health status aggregation assessment is performed to obtain the comprehensive health score of the first computing power subgroup. A health status report for the first computing power subgroup is generated based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold.

2. The method for monitoring the status of computing power subgroups with tasks as monitoring units according to claim 1, characterized in that, The result collector obtains the raw probe results from the first computing power subgroup to obtain a set of raw probe results, including: The results collector obtains status indicators and raw performance indicators from the first computing power subgroup through a periodic polling or callback mechanism. The status indicators include success or failure, and the raw performance indicators include actual task execution time, average network latency, network packet loss rate, throughput, number of processed items, and error rate of internal operations.

3. The method for monitoring the status of computing power subgroups with tasks as monitoring units according to claim 2, characterized in that, The result collector obtains the raw probe results from the first computing power subgroup to obtain a set of raw probe results, which also includes: The performance indicators in the original performance indicators are normalized to obtain the performance indicators, wherein the normalization process is expressed as follows: ;in, Baseline mean For baseline standard deviation, These are the various performance indicators in the original performance metrics, and N is the tolerance factor. This represents the original deviation ratio. For the initial score, For performance metrics.

4. The method for monitoring the status of computing power subgroups with tasks as monitoring units according to claim 3, characterized in that, Based on the set of original probe results, a subgroup health status aggregation assessment is performed to obtain the comprehensive health score of the first computing power subgroup, including: One-hot embedding encoding is performed on the state index of each original probe result in the set of original probe results to obtain the state index one-hot embedding encoding vector. The performance metrics of each original probe result in the set of original probe results are combined into a performance metric vector. The state index one-hot embedding encoding vector and the performance index vector are combined to form the original probe result encoding vector to obtain a set of original probe result encoding vectors; Based on the set of encoded vectors of the original probe results, a subgroup health status aggregation assessment is performed to obtain the comprehensive health score of the first computing power subgroup.

5. The method for monitoring the status of computing power subgroups with tasks as monitoring units according to claim 4, characterized in that, Combining the state indicator one-hot embedding encoding vector and the performance indicator vector to form the original probe result encoding vector to obtain a set of original probe result encoding vectors includes: concatenating the state indicator one-hot embedding encoding vector and the original probe result encoding vector to obtain the original probe result encoding vector.

6. The method for monitoring the status of computing power subgroups with tasks as monitoring units according to claim 4, characterized in that, Based on the set of encoded vectors from the original probe results, a subgroup health status aggregation assessment is performed to obtain the comprehensive health score of the first computing power subgroup, including: Dynamic aggregation analysis is performed on the set of original probe result encoding vectors to obtain the dynamic aggregation encoding vector of the original probe results; The original probe results are dynamically aggregated and encoded vectors are subjected to sequence decoding regression to obtain the comprehensive health score of the first computing power subgroup.

7. The method for monitoring the status of computing power subgroups with tasks as monitoring units according to claim 6, characterized in that, Dynamic aggregation analysis is performed on the set of original probe result encoding vectors to obtain dynamic aggregated encoding vectors of the original probe results, including: Feature baseline learning is performed on the set of original probe result encoding vectors to obtain the original probe result reference baseline encoding vector; Calculate the real-time offset of the set of original probe result encoding vectors relative to the original probe result reference base encoding vector to obtain the set of original probe result encoding asymmetric correction coefficients; Based on the set of asymmetric correction coefficients for the original probe results and the set of encoding vectors for the original probe results, dynamic compensation is performed on the reference encoding vector of the original probe results to obtain the dynamic aggregate encoding vector of the original probe results.

8. The method for monitoring the status of computing power subgroups with tasks as monitoring units according to claim 7, characterized in that, Calculate the real-time offset of the set of original probe result encoding vectors relative to the original probe result reference encoding vector to obtain the set of original probe result encoding asymmetric correction coefficients, including: The offset quantization factor of each original probe result encoding vector in the set of original probe result encoding vectors relative to the original probe result reference base encoding vector is calculated to obtain the set of original probe result encoding offset quantization factors; The set of offset quantization factors encoded by the original probe results is regularized based on the Softmax activation function to obtain the set of asymmetric correction coefficients encoded by the original probe results.

9. The method for monitoring the status of computing power subgroups with tasks as monitoring units according to claim 1, characterized in that, Based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold, a health status report for the first computing power subgroup is generated, including: Set the health status threshold to 0.9, the degradation threshold range to 0.7-0.9, the alarm threshold range to 0.5-0.7, and the critical alarm threshold to 0.5; If the overall health score of the first computing power subgroup is greater than the health status threshold, the health status of the first computing power subgroup is healthy. If the overall health score of the first computing power subgroup is within the downgrade threshold, the health status of the first computing power subgroup is downgraded. If the overall health score of the first computing power subgroup is within the alarm threshold range, the health status of the first computing power subgroup is alarmed. If the overall health score of the first computing power subgroup is less than the critical alarm threshold, the health status of the first computing power subgroup is critical alarm.

10. A computing power subgroup status monitoring system with tasks as monitoring units, characterized in that, include: The scheduling plan generation module is used to generate a schedule of probe instances to be scheduled for the first computing power subgroup based on the probe task template list. The instance planning and scheduling module is used to use the task scheduler to submit each probe task instance in the probe instance plan to be scheduled to the first computing power subgroup in sequence according to the planned time in the probe instance plan to be scheduled and the current time. The raw probe result acquisition module is used to obtain raw probe results from the first computing power subgroup using a result acquisition device to obtain a set of raw probe results, wherein the raw probe results include performance indicators and status indicators; The subgroup health status comprehensive assessment module is used to perform subgroup health status aggregation assessment based on the set of original probe results to obtain the comprehensive health score of the first computing power subgroup. The subgroup health status report generation module is used to generate a health status report for the first computing power subgroup based on a comparison between the comprehensive health score of the first computing power subgroup and a preset threshold.