Heterogeneous resource unified monitoring and multi-level alarm method and device for data annotation cloud platform
By performing real-time performance data processing and multi-level threshold management on the heterogeneous resources of the data annotation cloud platform, the problems of incomplete resource monitoring and non-level alarms in the existing technology have been solved, realizing unified monitoring and multi-level alarms for heterogeneous resources, and ensuring the efficient and stable operation of the platform.
Patent Information
- Application Number
- CN202511461400.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-13
AI Technical Summary
Existing technologies cannot achieve real-time monitoring and anomaly alerts for heterogeneous resources across the entire data annotation cloud platform. This makes it difficult for platform administrators to fully grasp the resource status. Furthermore, traditional alarm systems lack a tiered response mechanism, which can easily lead to redundant alarm information or the neglect of critical issues, affecting task continuity and stability.
By acquiring real-time performance data of heterogeneous resources, multi-dimensional analysis and trend prediction are performed, multi-level thresholds and dynamic performance management are set, multi-level alarm information is generated, and resource expansion, load migration and fault recovery measures are automatically executed. Combined with machine learning for resource planning, unified monitoring and multi-level alarms are achieved.
It enables comprehensive monitoring and efficient management of heterogeneous resources on the data annotation cloud platform, ensuring the platform's continuous and stable operation, reducing manual intervention, and improving alarm response efficiency and platform stability.
Smart Images

Figure CN121333997A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and in particular to a method and apparatus for unified monitoring and multi-level alarm of heterogeneous resources for a data annotation cloud platform. Background Technology
[0002] Existing technologies for resource monitoring and alerting solutions for data annotation cloud platforms have the following limitations: Most monitoring solutions on the market can only collect and alert on monitoring data for hosts and commonly used middleware, failing to complete real-time data collection and anomaly alerting across the entire chain from hardware resources (such as computing, storage, and network devices) to business applications (such as annotation tasks and user interactions) on a single platform. This makes it difficult for platform administrators to fully grasp the overall resource status. Data annotation cloud platforms involve heterogeneous resources, including computing resources (CPU, GPU, virtual machines, containers, etc.), storage resources (disks, cloud storage services, etc.), network resources (bandwidth, latency, etc.), and business operation indicators (number of online users, annotation task time, etc.). Traditional solutions are mostly designed for single resource types, lacking a unified monitoring mechanism for heterogeneous resources, and are difficult to adapt to the complex resource composition of the platform.
[0003] Traditional alarm systems typically employ a single alarm level, failing to provide tiered responses based on the severity of resource usage (e.g., approaching threshold, exceeding threshold, fault status). In complex environments with multiple resources collaborating, this can easily lead to redundant alarm information or the overlooking of critical issues, affecting administrators' ability to prioritize problems. Existing solutions often rely on manual intervention after an alarm is triggered. However, data annotation cloud platforms frequently handle large-scale tasks, making manual responses insufficient to address sudden resource issues (e.g., node overload, insufficient storage), potentially causing task delays or even platform crashes, and compromising the continuity and stability of annotation tasks.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore includes information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0006] According to one aspect of this application, a method for unified monitoring and multi-level alarming of heterogeneous resources for a data annotation cloud platform is provided, comprising: acquiring real-time performance data of heterogeneous resources for the data annotation cloud platform, including CPU utilization, GPU load, storage space utilization, disk I / O, bandwidth utilization, network latency, number of online users, and annotation task time; processing the acquired real-time performance data of heterogeneous resources to capture the usage status and changing trends of various resources, and generating real-time data and preliminary assessment information of the overall resource usage; processing the real-time data and preliminary assessment information of the overall resource usage, setting multi-level thresholds based on a dynamic performance threshold management module, combined with resource usage patterns and load conditions, applying characteristic constraints of different resource types and platform stable operation objective functions, and generating alarm information and detailed resource status reports of corresponding levels; The system processes alarm information and detailed resource status reports at different levels, and automatically executes response measures such as resource expansion, load migration, and fault recovery based on alarm levels. It dynamically adjusts resource configuration and generates automatic response plans to ensure the continuous and stable operation of the platform. The system further processes these automatic response plans, utilizing data analysis and trend prediction modules to analyze historical data through machine learning and data mining techniques, identify resource usage patterns, and generate future resource demand trend predictions and resource planning suggestions. Finally, it processes real-time data and preliminary assessment information on overall resource usage, multi-level alarm information and detailed resource status reports, automatic response plans, future resource demand trend predictions, and resource planning suggestions to generate comprehensive evaluation information on the unified monitoring of heterogeneous resources and the effectiveness of multi-level alarms for the data-annotated cloud platform. This evaluation includes resource monitoring coverage, alarm response efficiency, and platform stability improvement rate.
[0007] Another aspect of this application discloses a unified monitoring and multi-level alarm device for heterogeneous resources on a data annotation cloud platform, comprising: an acquisition module for acquiring real-time performance data of heterogeneous resources on the data annotation cloud platform; a processing module for processing the acquired real-time performance data of heterogeneous resources, capturing the usage status and changing trends of various resources, and generating real-time data and preliminary assessment information of the overall resource usage; processing the real-time data and preliminary assessment information of the overall resource usage, setting multi-level thresholds based on a dynamic performance threshold management module, combining resource usage patterns and load conditions, applying characteristic constraints of different resource types and a platform stability operation objective function, and generating alarm information and detailed resource status reports of corresponding levels; and processing the alarm information and detailed resource status reports of different levels. The system processes alarms and automatically executes response measures such as resource expansion, load migration, and fault recovery, dynamically adjusting resource configurations to generate automatic response plans that ensure the platform's continuous and stable operation. These automatic response plans are then processed using data analysis and trend prediction modules. Machine learning and data mining techniques are employed to analyze historical data, identify resource usage patterns, and generate future resource demand trend predictions and resource planning suggestions. Real-time data and preliminary assessment information on overall resource usage, multi-level alarm information and detailed resource status reports, automatic response plans, and future resource demand trend predictions and resource planning suggestions are processed to generate comprehensive evaluation information on the unified monitoring of heterogeneous resources and the effectiveness of multi-level alarms for the data-annotated cloud platform. This evaluation includes resource monitoring coverage, alarm response efficiency, and platform stability improvement rates.
[0008] According to another aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a second processor, implements the above-described method for unified monitoring and multi-level alarm of heterogeneous resources for a data annotation cloud platform.
[0009] This application provides a method and apparatus for unified monitoring and multi-level alarming of heterogeneous resources for a data annotation cloud platform. The aim is to ensure the efficient and stable operation of data annotation tasks by integrating real-time monitoring and multi-level alarming of hosts, application services, and middleware resources across multiple clusters. A unified resource data acquisition module communicates with various resource nodes through multiple protocols to obtain resource performance data in real time. Then, the monitoring module evaluates resources based on preset thresholds. If resource usage approaches or exceeds the set threshold, corresponding alarms are issued through a multi-level alarm mechanism (including warning, critical, and emergency levels). The alarm system supports timely delivery of alarm information to administrators through various communication methods (such as email, SMS, and dashboard notifications). When an alarm reaches the critical or emergency level, the system can automatically execute response measures, such as automatically expanding computing resources, migrating loads, and repairing faulty nodes, thereby ensuring the continuous and stable operation of the platform. Through unified resource monitoring, dynamic threshold management, multi-level alarming, automatic response, and trend prediction technologies, comprehensive control and efficient management of heterogeneous resources are achieved, ensuring the efficient and stable operation of data annotation tasks.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0011] Figure 1 This document illustrates a flowchart of a method for unified monitoring and multi-level alarming of heterogeneous resources for a data annotation cloud platform, provided in an embodiment of this application.
[0012] Figure 2 This paper presents a schematic diagram of a heterogeneous resource unified monitoring and multi-level alarm device for a data annotation cloud platform, according to an embodiment of this application. Detailed Implementation
[0013] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0014] The following is combined Figure 1 This application describes a method for unified monitoring and multi-level alarming of heterogeneous resources for a data annotation cloud platform, based on exemplary embodiments thereof. In one embodiment, this application also proposes a method and apparatus for unified monitoring and multi-level alarming of heterogeneous resources for a data annotation cloud platform:
[0015] S101 acquires real-time performance data of heterogeneous resources for the data annotation cloud platform.
[0016] In one implementation, real-time performance data of heterogeneous resources on a data annotation cloud platform is acquired to comprehensively understand the real-time operating status of various resources on the platform, providing basic data support for subsequent monitoring, alarms, and resource adjustments. Communication is established with computing nodes (such as virtual machines and containers) in the cloud platform, and standardized protocols (such as SNMP and WMI) are used to collect CPU usage in real time. For example, for a virtual machine running annotation tasks, its CPU usage is collected every 10 seconds. If the virtual machine is currently processing a large number of image annotation tasks, its CPU usage may be collected as 75%. For computing resources equipped with GPUs, GPU monitoring tools (such as NVIDIA System Management Interface) are used to obtain GPU load data, including GPU core utilization and memory utilization. For example, a GPU server used to accelerate deep learning model inference may have a GPU core utilization of 80% when processing high-resolution image annotation tasks, meaning the GPU load is 80%. Storage space utilization: Interaction with storage devices (such as disk arrays and cloud storage services) is conducted, and the total capacity and used capacity of the storage space are obtained through interfaces such as REST APIs, thereby calculating the utilization rate. For example, if a cloud storage service has a total capacity of 1000GB and currently stores 600GB of various labeled data and intermediate files, then its storage space utilization rate is 60%.
[0017] By monitoring disk read / write operations, data such as the number of bytes read and written per unit time and the number of read / write operations are collected to reflect disk I / O performance. For example, if a storage node is performing a large amount of annotation data writing operations and the data shows that it writes 50MB per second, then the disk I / O write rate is 50MB / s. Bandwidth utilization: Using network monitoring tools, data is collected from network devices (such as switches and routers) in the cloud platform to obtain the total network bandwidth capacity and the currently used bandwidth capacity, and the utilization rate is calculated. For example, if a network link connecting annotation task nodes and storage nodes has a total bandwidth of 100Mbps, and currently uses 60Mbps due to a large amount of annotation data transmission, then the bandwidth utilization rate is 60%.
[0018] Network latency is measured by sending probe packets (such as ICMP packets) to target network nodes and recording the round-trip time of the packets. For example, if the round-trip time of a probe packet sent from the node initiating the annotation task to the storage node is 20ms, then the network latency is 20ms. Online user count: The number of users currently logged in and using the platform for annotation tasks is counted in real time through the user management system or application service interface of the cloud platform. For example, if 150 users are performing data annotation operations on the platform at 10:00 AM on a weekday, then the online user count is 150. Annotation task time: The start and end times of each annotation task are obtained from the platform's task management module; the difference between the two is the task's time. For example, if a text annotation task starts at 9:00 AM and completes at 9:10 AM, then the annotation task took 10 minutes.
[0019] S102 processes the acquired real-time performance data of heterogeneous resources, captures the usage status and changing trends of various resources, and generates real-time data and preliminary assessment information on the overall usage of resources.
[0020] In one implementation, real-time aggregation and multi-dimensional analysis are performed on computing resource data, storage resource data, network resource data, and business operation data from heterogeneous resource real-time performance data to generate computing resource load characteristics, storage resource occupancy characteristics, network resource transmission characteristics, and business operation status characteristics, which are then summarized to form resource usage status characteristic parameters. For the collected computing, storage, network, and business operation data, a streaming computing framework (such as Flink) is used for real-time aggregation, analyzing dimensions such as resource load intensity, distribution balance, and response efficiency to extract characteristic parameters. Data such as CPU utilization and GPU load are aggregated to calculate the average load of a single node, the standard deviation of the cluster load, and the percentage of peak loads. For example, a labeling task cluster contains 10 computing nodes with CPU utilization rates of 70%, 75%, 80%,...90%. The calculated average cluster load is 78%, the standard deviation is 5.2%, and 30% of the nodes have a load above 80%. These indicators collectively constitute the computing resource load characteristics.
[0021] By integrating data such as storage space utilization and disk I / O, we analyze the distribution of storage utilization, hot read / write areas, and I / O response latency. For example, in the platform's storage cluster, the utilization rate of the hot data area is 85%, the utilization rate of the cold data area is 40%, and the average disk I / O response time is 15ms. Among these, I / O requests for image annotation data writes account for 60%, forming the characteristics of storage resource occupancy. Based on data such as bandwidth utilization and network latency, we statistically analyze the correlation between peak link load, cross-node transmission latency distribution, and packet loss rate. For example, the peak bandwidth utilization rate of the link connecting the annotation node and the storage node is 90%, the average latency is 25ms, and the packet loss rate rises to 1% when the bandwidth utilization rate exceeds 80%. These data constitute the characteristics of network resource transmission.
[0022] By aggregating business metrics such as the number of online users and the time spent on annotation tasks, we can analyze user concurrency, task type distribution, and average completion time. For example, if the platform currently has 200 online users, 60% of them are processing text annotation tasks, with an average task completion time of 8 minutes; 40% of them are processing image annotation tasks, with an average completion time of 15 minutes, forming characteristics of the business operation status. These four types of characteristic parameters are then summarized to form resource usage status characteristic parameters, fully reflecting the real-time operational status of the current resources.
[0023] The system dynamically tracks and mines periodic patterns of various resources to generate trend characteristics for computing resources, storage resources, network resources, and business metrics, integrating these into resource trend characteristic parameters. Time series analysis tools (such as Prophet) are used to continuously track various resource data, uncovering fluctuation patterns over daily, weekly, and monthly periods, and extracting trend characteristic parameters. Hourly changes in CPU and GPU load are tracked to identify peak periods and fluctuation amplitudes. For example, analysis of data from the past 7 days reveals that peak computing resource load occurs between 9:00-11:00 and 14:00-16:00 daily, with peak values 30% higher than the average. GPU load fluctuations are more significant during periods of concentrated image annotation tasks, forming a computing resource fluctuation trend characteristic. The system tracks the daily growth rate of storage space utilization, fitting a growth curve and predicting saturation time. For example, with an average daily growth of 5GB in storage resources over the past 30 days and a current utilization rate of 70%, this trend predicts a critical threshold of 85% in 15 days. Image annotation data accounts for 75% of the total growth, forming a storage resource growth trend characteristic.
[0024] Track the periodic changes in bandwidth utilization and network latency, correlating them with peak business periods. For example, Mondays and Wednesdays from 9:00 to 10:00 AM are peak times for network bandwidth usage, exceeding the daily average by 40%. During this time, network latency increases by an average of 10ms, forming a trend characteristic of network resource load. Track the weekly changes in the number of online users and the time spent on annotation tasks, analyzing the correlation between user activity and task complexity. For example, data from the past four weeks shows that the number of online users on weekends is 20% lower than on weekdays, but the average time spent on a single annotation task increases by 15% (due to an increase in the proportion of complex tasks), forming a trend characteristic of business indicator changes. Integrate the above four types of trend characteristic parameters to form resource trend characteristic parameters, reflecting the dynamic changes in resource usage.
[0025] The system matches and compares resource usage status characteristics and trend characteristics with a pre-set resource baseline status database. Through resource usage assessment and potential risk identification, it generates real-time data and preliminary assessment information on the overall resource usage. The system also matches resource usage status characteristics and trend characteristics with a pre-set resource baseline status database (including normal operating ranges and risk warning thresholds), using deviation calculations to assess the rationality of resource usage and identify potential risks. For example, comparing the current average computing resource load of 78% with the "normal load range of 50%-80%" in the baseline status database, it is determined to be within a reasonable range. However, the storage resource growth trend indicates that it will reach a critical threshold in 15 days, deviating from the baseline requirement of "30-day advance warning," thus assessing it as "storage resources need attention."
[0026] For example, in network resource transmission characteristics, when bandwidth utilization reaches 80%, the packet loss rate rises to 1%. Compared with the baseline of "a packet loss rate exceeding 0.5% indicates a transmission risk," this identifies a "data transmission stability risk under high bandwidth load." In business operation status, the image annotation task time increased by 20% compared to the average of the previous week. Combined with the fluctuation trend of computing resources, this indicates a "risk of task delay due to insufficient GPU resource allocation." By comprehensively evaluating the results and risk points, real-time data on overall resource usage (such as current resource load and growth rate) and preliminary assessment information (such as "computing resources are normal, storage resources are about to reach a critical value, and there is a stability risk in network transmission") are generated, providing a basis for subsequent threshold judgments and alarm triggering.
[0027] S103 processes real-time data and preliminary assessment information on the overall resource usage. Based on the dynamic performance threshold management module, it sets multi-level thresholds by combining resource usage patterns and load conditions, applies characteristic constraints of different resource types and platform stability operation objective functions, and generates corresponding alarm information and detailed resource status reports.
[0028] In one implementation, dynamic threshold calculations and usage pattern matching are performed on the current resource utilization rate, peak load, and response time in real-time data of overall resource usage. This generates multi-level threshold features for computing resources, critical interval features for storage resources, and early warning line features for network resources, forming the basic information for threshold setting. For key indicators (current utilization rate, peak load, and response time) in the real-time data of overall resource usage, multi-level thresholds are calculated through statistical analysis and algorithm models, combining historical usage patterns and real-time load characteristics, to extract threshold features for various resources.
[0029] Based on real-time data and historical peak values of CPU utilization and GPU load, a percentile method (such as P80, P90, P100) is used to divide the threshold into multiple levels, and these are matched with the load patterns of different task types. For example, the default CPU utilization threshold for text annotation task clusters is: warning level (80%), critical level (90%), and urgent level (100%); while for image annotation tasks, due to higher GPU load, the system adjusts the GPU threshold to: warning level (75%), critical level (85%), and urgent level (95%) through pattern matching, forming a multi-level threshold feature for computing resources.
[0030] Based on the real-time growth rate of storage space utilization and historical saturation periods, different critical intervals are defined and correlated with disk I / O performance. For example, the warning interval for storage resource utilization is 70%-80% (at which point I / O response is normal), the critical interval is 80%-90% (I / O latency begins to rise), and the emergency interval is above 90% (I / O may be interrupted). Combined with a daily average growth rate of 5GB, the time characteristic of "entering the critical interval in 10 days" is marked, forming the critical interval characteristics of storage resources.
[0031] Based on real-time fluctuations in bandwidth utilization and network latency, and combined with load patterns during peak business hours (e.g., 9:00-11:00 daily), dynamic warning thresholds are set. For example, the bandwidth utilization warning threshold is 75% during normal periods, but it is automatically adjusted to 70% during peak business hours through pattern matching; the network latency warning threshold is 100ms, with a critical threshold of 200ms. When a concentration of video annotation tasks is detected, the latency warning threshold is further tightened to 80ms, forming the network resource warning threshold characteristics. Integrating these three types of characteristics forms the basic information for threshold setting, providing a quantitative basis for subsequent alarm level determination.
[0032] The resource characteristic differences and load fluctuation patterns in the preliminary assessment information are transformed into constraints and type boundaries are defined to generate computational resource characteristic constraint parameters, storage resource usage limit parameters, and network resource performance constraint parameters, forming resource constraint information. For the characteristic differences of various resources in the preliminary assessment information (such as the parallel processing capability of computational resources and the read / write characteristics of storage resources) and load fluctuation patterns (such as daytime peaks and nighttime troughs), they are transformed into quantifiable constraint parameters, and type boundaries are defined. Based on the hardware characteristics of CPUs and GPUs (such as CPUs being suitable for multi-task parallelism and GPUs being suitable for graphics rendering), resource allocation constraints are set. For example, it is stipulated that GPU resources are preferentially allocated to image / video annotation tasks, and the load of a single-node GPU should not exceed 90% (to avoid overheating); CPU core allocation must meet the parallel constraint of "single task occupancy not exceeding 20%", forming computational resource characteristic constraint parameters.
[0033] Based on the characteristics of storage media (e.g., SSDs offer fast read / write speeds but have small capacities, while HDDs offer large capacities but high latency) and data types (hot data, cold data), usage limits are defined. For example, SSD storage is only used to store annotation task data from the past 7 days, and its utilization rate must not exceed 85% (to ensure read / write performance); HDD storage is used for archived data, with a maximum utilization rate of 95%, but 5% space must be reserved for defragmentation, thus forming storage resource usage limit parameters.
[0034] Performance constraints are set based on the bandwidth limit of network links and the transmission protocol (e.g., TCP for reliable transmission, UDP for scenarios with high real-time requirements). For example, data transmission for annotation tasks should preferentially use the TCP protocol, and the bandwidth usage of a single link should not exceed 80% of the total bandwidth (to avoid congestion); real-time preview streams should use the UDP protocol, and network latency should not exceed 150ms (to ensure smoothness). These form network resource performance constraint parameters. Integrating these three types of parameters forms resource constraint information, ensuring that the threshold settings conform to the physical characteristics of resources and business requirements.
[0035] For the objective function of platform stable operation, parameters are adapted and the degree of goal achievement is evaluated in conjunction with resource usage status to generate resource load balancing target characteristics and platform stability matching parameters, forming objective function optimization information. Taking the platform stable operation objective function (such as maximizing resource utilization, minimizing task latency, and minimizing failure probability) as the core, and combining it with the current resource usage status, target weights are adjusted through parameter adaptation, and the degree of goal achievement is evaluated to extract optimization characteristics. Based on the balancing objective of "load difference between nodes ≤ 10%", and combined with the current resource distribution (such as node loads of 70% and 90% in a cluster), the load balancing coefficient is calculated (currently 0.8, target value 1.0), and resource scheduling parameters are adapted (such as migrating 20% of tasks to low-load nodes), forming resource load balancing target characteristics.
[0036] With "monthly downtime ≤ 1 hour" as the stability target, and considering the current failure rate (e.g., a cumulative total of 5 minutes of brief outages due to network fluctuations in the past 7 days), the target achievement rate is assessed (currently 92%), and fault tolerance parameters are adjusted (e.g., increasing the number of replicas for critical tasks from 2 to 3) to form platform stability matching parameters. These two types of characteristics are integrated to form objective function optimization information, ensuring that the threshold setting is consistent with the overall platform operation goals.
[0037] Integrate basic threshold setting information, resource constraint information, and objective function optimization information to generate alarm information and detailed resource status reports at corresponding levels. By fusing basic threshold setting information, resource constraint information, and objective function optimization information from multiple dimensions, and comparing the actual resource status with the thresholds, generate alarm information at corresponding levels and status reports including resource details, trend analysis, and suggested measures.
[0038] At a certain moment, the GPU load reached 88%, exceeding its critical threshold (85%), and matched the peak pattern of image annotation tasks (consistent with the multi-level threshold characteristics of computing resources). Simultaneously, the GPU node's load was approaching the constraint of "single node not exceeding 90%" (triggering resource constraint information). Combined with the platform's stability goals, the current load difference reached 15% (failure to achieve the balance target). After comprehensive assessment, a critical level alarm was generated, with the status report detailing, "GPU resource load is too high; it is recommended to migrate 30% of image annotation tasks to idle nodes, which is expected to reduce the load to 75%." This integration process ensures the accuracy and reasonableness of the alarm information, providing a clear basis for subsequent automatic responses.
[0039] S104 processes alarm information and detailed resource status reports at different levels, and automatically executes response measures such as resource expansion, load migration, and fault recovery based on the alarm level. It dynamically adjusts resource configuration and generates automatic response plans to ensure the continuous and stable operation of the platform.
[0040] In one implementation, response strategies are matched and automation levels are graded for different levels of alarm information, including warning-level alerts, critical-level warnings, and emergency-level alarms. This generates warning-level manual intervention suggestion features, critical-level semi-automated processing features, and emergency-level fully automated response features, forming alarm response strategy information. For different levels of alarm information (warning, critical, and emergency), a preset response strategy is matched based on the alarm type and resource characteristics, and processing levels are divided according to automation level, extracting response features for each type of alarm. When an alarm is at the warning level (e.g., CPU utilization close to 80%, storage utilization close to 70%), the system determines that no automatic operation is needed and only generates a manual intervention suggestion. For example, if a text annotation cluster's CPU utilization reaches 78% (warning level), the system matches a "close monitoring + load balancing suggestion" strategy, generating feature information such as "suggest checking node load distribution within 1 hour and manually migrating 5% of low-priority tasks to idle nodes," clarifying the timing and content of manual operation.
[0041] When an alarm reaches a critical level (e.g., CPU utilization reaches 90% or network latency exceeds 200ms), the system initiates a semi-automatic response, automatically generating an execution plan that requires administrator confirmation before execution. For example, if the GPU load on an image annotation node reaches 86% (critical level), the system matches the "resource expansion + load migration" strategy, automatically calculating the need to add two new GPU containers and migrate 30% of tasks to a backup node, and generating a feature message suggesting "confirm and execute the above operations, expected to be completed within 5 minutes, reducing the load to 60%", thus retaining the human decision-making step.
[0042] When an alarm is classified as emergency (e.g., 100% CPU utilization, storage failure), the system triggers a fully automated response, executing preset plans without manual intervention. For example, if a storage node experiences a sudden failure (emergency level), the system matches the "fault recovery + data migration" strategy, automatically starting a backup node, switching data links, and generating characteristic information such as "Backup storage node started, data migration in progress, service expected to resume within 1 minute," recording the operation log for later traceability. These three types of characteristics are integrated to form alarm response strategy information, clearly defining the handling methods and automation levels for different alarm levels.
[0043] The system performs resource scheduling feasibility analysis and adjustment cost assessment on the resource load distribution, node health status, and performance bottleneck location data in the detailed resource status report. This generates load migration path characteristics, resource expansion priority parameters, and faulty node replacement scheme characteristics, forming the basic information for resource adjustment. For key data in the detailed resource status report (load distribution, node health status, performance bottleneck location), the system analyzes the feasibility of resource scheduling (e.g., poor node load, network connectivity), assesses adjustment costs (e.g., migration time, expansion costs), and extracts the core characteristics of resource adjustment. Based on resource load distribution data (e.g., node A is 90% loaded, node B is 40% loaded), the system analyzes the network path, data transmission volume, and time consumption for task migration. For example, migrating 20 image annotation tasks from node A to node B, the system calculates the optimal path as "node A → core switch → node B," transmitting 5GB of data, with an estimated time of 30 seconds, and generates the characteristic information "migration path is smooth, bandwidth meets requirements, it is recommended to prioritize migrating small file tasks."
[0044] Based on the location of performance bottlenecks (such as insufficient GPU resources or saturated storage capacity), the system assesses the priority and cost of expanding different resources. For example, if the platform simultaneously faces high GPU load and insufficient storage capacity, the system analysis concludes that "GPU expansion can immediately alleviate task latency (priority 1), while storage expansion can delay it for 24 hours (priority 2)," and generates the characteristic information "It is recommended to prioritize starting 2 GPU instances (cost 50 yuan / hour), and postpone storage expansion." For the health status data of failed nodes (such as virtual machine crashes or disk damage), the system assesses the matching degree of backup nodes and the replacement time. For example, if node C fails due to disk damage, and the system detects that the hardware configuration of backup node D (same CPU model, same storage capacity) is consistent with node C, it generates the characteristic information "Node C can be immediately replaced by node D, data synchronization takes 5 minutes, and service interruption is expected to last 1 minute." These three types of characteristics are integrated to form basic information for resource adjustment, providing data support for the execution of subsequent response measures.
[0045] Combining alarm levels and dynamic adjustment requirements, the system performs execution sequence planning and conflict resolution on alarm response strategy information and resource adjustment basic information, generating dynamic resource configuration adjustment sequence characteristics and multi-measure collaborative execution coefficients to form response optimization information. Combining alarm levels and real-time adjustment requirements, the system performs time-series planning on alarm response strategy information and resource adjustment basic information (e.g., migration before expansion), resolving conflicts between measures (e.g., simultaneous migration and expansion leading to excessive bandwidth consumption), and extracting optimized execution characteristics. Based on the alarm urgency and operational dependencies, the system plans execution steps and time nodes. For example, in an emergency-level storage failure alarm, the system plans the sequence as "1. Start the backup node (0-30 seconds); 2. Migrate data (30 seconds-5 minutes); 3. Switch service entry (5-5.5 minutes)," and generates the characteristic information that "sequential execution minimizes service interruption time."
[0046] The system assesses the synergistic efficiency of simultaneously implementing multiple measures (such as load migration and resource expansion) and calculates the resource overlap rate. For example, when performing task migration and GPU expansion simultaneously, the system analysis concludes that "the bandwidth overlap rate between the two is 60%, which may lead to a 20% increase in latency," generating the characteristic information "It is recommended to stagger the execution, complete the migration first (30 seconds) and then start the expansion, improving the synergy coefficient to 0.8." These two types of characteristics are integrated to form response optimization information, ensuring that response measures are executed efficiently and collaboratively.
[0047] By integrating alarm response strategy information, basic resource adjustment information, and response optimization information, an automated response plan is generated to ensure the continuous and stable operation of the platform. This multi-dimensional fusion of alarm response strategy information, basic resource adjustment information, and response optimization information generates an automated response plan that includes specific operational steps, execution timing, and expected results. For example, when a platform triggers a critical GPU load alarm (node A load 88%), combining alarm response strategy information (semi-automated processing, recommended expansion + migration), basic resource adjustment information (node B load 40%, 10 tasks can be migrated; starting one GPU instance takes 2 minutes), and response optimization information (migrate first, then expand, coordination coefficient 0.9), the following solution is generated: "1. After manual confirmation, migrate 10 tasks from node A to node B (path unobstructed, time 30 seconds); 2. Start one GPU instance (priority 1, cost 50 yuan / hour); 3. Expected to reduce node A load to 65% within 5 minutes, and improve the overall GPU load balancing rate to 90%." This integration process ensures the feasibility, timeliness, and optimization of the automated response plan, providing a guarantee for the stable operation of the platform.
[0048] S105 processes the automatic response scheme, using the data analysis and trend prediction module to analyze historical data through machine learning and data mining techniques, identify resource usage patterns, and generate future resource demand trend predictions and resource planning suggestions.
[0049] In one implementation, the automatic response scheme is decomposed and analyzed to generate historical resource adjustment features, response measure effectiveness evaluation features, and dynamic resource configuration change correlation information, which together constitute the core elements of the response scheme. The operation records, execution effects, and resource configuration changes in the automatic response scheme are structurally decomposed to extract feature information reflecting resource adjustment patterns, forming the basis data for predictive analysis. Specific operational details of resource expansion, load migration, and fault recovery in the scheme are extracted, including adjustment time, resource types involved, and adjustment magnitude. For example, an automatic response scheme record states: "On June 10th at 9:30 AM, due to GPU load reaching 85%, two GPU containers were started, migrating 15 image annotation tasks to node B." From this, features such as "GPU expansion trigger conditions, number of migration tasks, and node load difference" are extracted to form a historical adjustment pattern dataset.
[0050] The effectiveness of each response measure is quantitatively evaluated, including the load reduction after resource expansion, changes in migration task time, and service recovery rate after fault recovery. For example, after implementing the above solution, the average GPU load decreased from 85% to 62%, the average task migration time was 28 seconds, and the service interruption duration was 0.5 minutes. These indicators constitute the effectiveness evaluation characteristics used to determine the effectiveness of the measures. The correlation between resource adjustments and business load is analyzed, such as the correlation between GPU expansion and the growth in the number of online users, and the matching degree between storage expansion and the amount of labeled tasks. For example, data shows that "for every 50 additional online users, the GPU load increases by about 15%, and the probability of triggering expansion increases by 30%", forming a correlation characteristic between resource configuration and business indicators. The above three types of information are integrated to form the core elements of the response plan, providing data support for subsequent trend prediction.
[0051] For resource-load forecasting models, based on the core elements of response schemes and resource usage fluctuation patterns, time-series pattern recognition algorithms are used to set key model parameters and generate initial forecasting information. Based on historical patterns and resource usage fluctuation characteristics (such as daily peaks and weekly cycles) within the core elements of the response schemes, time-series pattern recognition algorithms (such as ARIMA and LSTM) are used to calibrate key parameters of the forecasting model (such as cycle factors, trend weights, and fluctuation coefficients) to determine the initial model input.
[0052] For the GPU (GPU) prediction model, analysis of response scheme data over the past three months revealed a cyclical pattern in GPU load: peak loads from 9:00-11:00 and 14:00-16:00 on weekdays, with a 40% load decrease on weekends. A time-series pattern recognition algorithm was used to extract a cycle factor of 1440 minutes (1 day), with a trend weight set at 0.7 (recent data has a greater impact), and a fluctuation coefficient set at 0.15 based on the historical maximum load deviation. Simultaneously, by incorporating the historical resource adjustment pattern that "each expansion of one GPU container can support 20 image annotation tasks," task-resource conversion parameters were set to generate initial prediction information for the model, ensuring that the model can capture the cyclical and correlated nature of the load.
[0053] Based on the initial information predicted by the model, a resource load correlation analysis mechanism is used to process the core element information of the response plan, generating initial trend prediction feature information that includes trend characteristics of computing resource demand and correlation characteristics of storage and network resource demand. Based on the initial information predicted by the model, the demand correlation of computing, storage, and network resources (such as the driving effect of computing resource growth on network bandwidth) is analyzed, generating independent trend characteristics and correlation demand characteristics for each type of resource. By analyzing the correlation between GPU load and annotation task type (such as 3D point cloud annotation requiring 5 times more GPUs than text annotation), the computing resource demand under different task proportions is predicted in the future. For example, the model predicts that "the proportion of image annotation tasks will increase from 40% to 55% in the next month, requiring 5 additional GPU containers to meet peak load."
[0054] Based on the correlation that "for every 100GB increase in labeled data, storage demand increases by 120GB (including backups), and network bandwidth demand increases by 5Mbps," and combined with computing resource prediction results, a correlation demand characteristic is generated: "Monthly storage growth is approximately 800GB, and bandwidth needs to be expanded from 100Mbps to 150Mbps." Integrating these characteristics forms initial trend prediction feature information, reflecting the independent and correlated demand trends of various resources.
[0055] The initial trend prediction features are processed based on historical data and real-time status coupling information to generate a fusion trend feature information of resource usage history and real-time. The initial trend prediction features are coupled with real-time resource status (such as the current number of online users and task queue length) to correct deviations in historical trends and generate trend features that fuse real-time dynamics. For example, the initial trend prediction was "GPU load will reach 80% at 9:00 tomorrow," but real-time data showed "Current online users have reached 180 (120 at the same time yesterday), and the task queue has a backlog of 30." The prediction result was corrected through the historical-real-time coupling model: "GPU load will reach 88% at 9:00 tomorrow, an 8% increase from the initial prediction, requiring expansion to be started 1 hour earlier," thus forming a trend feature that fuses real-time dynamics and improves prediction accuracy.
[0056] A multi-resource collaborative prediction mechanism is employed to process historical and real-time integrated trend information on resource usage, generating future resource demand trend predictions and resource planning suggestions. This mechanism (such as a resource association model based on graph neural networks) integrates the integrated trend characteristics of computing, storage, and network resources to predict the total and peak resource demand over a certain period (e.g., one week, one month), and generates specific planning suggestions. The model predicts that "within the next month, the maximum GPU load will reach 92% during weekday morning peak hours, storage will increase by an average of 25GB per day, and network bandwidth will peak at 180Mbps," while also highlighting the special trend that "resource demand is 20% higher than the average in the last week of each month due to task settlement."
[0057] Based on the prediction results, targeted suggestions are generated, such as "automatically start two backup GPU containers before 8:00 AM every Monday to Friday; expand storage resources by 200GB per week in advance, and adopt tiered storage for hot and cold data; upgrade network bandwidth to 200Mbps, and prioritize the data transmission links of marked data." Through multi-resource collaborative prediction, the planning of various resources is ensured to match each other, avoiding imbalances such as "computing expansion but insufficient storage," and providing a decision-making basis for the forward-looking allocation of platform resources.
[0058] S106 processes real-time data and preliminary assessment information on overall resource usage, multi-level alarm information and detailed resource status reports, automatic response plans, future resource demand trend predictions and resource planning suggestions to generate comprehensive evaluation information on the unified monitoring of heterogeneous resources and multi-level alarm effects for data-annotated cloud platforms, including resource monitoring coverage, alarm response efficiency, and platform stability improvement rate.
[0059] In one implementation, the resource coverage breadth and status assessment accuracy in real-time data and preliminary assessment information of overall resource usage are quantitatively analyzed to generate basic quantitative factors for resource monitoring. The basic assessment factors are generated by quantifying the real-time data and preliminary assessment information of overall resource usage from two dimensions: resource coverage breadth (completeness of monitoring indicators) and status assessment accuracy (the degree of consistency between assessment results and actual status). The proportion of resource types (computing, storage, network, services) and specific indicators (such as CPU utilization, GPU load, etc.) monitored by the system to the total resource types and indicators of the platform is statistically analyzed. For example, if the platform has 8 core resource indicators (such as CPU, GPU, storage, etc.), and the system currently monitors 7 of them, with a coverage rate of 87.5%; the sub-indicators under computing resources (such as utilization, load, health status) cover 90%, and the comprehensive computing resource coverage breadth factor is 0.875 × 0.9 = 0.7875.
[0060] The accuracy rate is calculated by comparing the "resource anomaly predictions" in the preliminary assessment information with the actual anomaly events. For example, if the system predicts that "storage resources will reach a critical value in 10 days," but actually triggers a critical alarm 7 days later, the deviation rate is 30%; if it predicts that "GPU resources are normal," but there are no anomalies, the accuracy rate is 100%. Taking into account the assessment deviations of multiple resource types, the status assessment accuracy factor is calculated to be 0.85 (a deviation within 20% is considered accurate). The coverage breadth factor and accuracy factor are weighted and fused (each with a weight of 0.5) to generate the basic quantitative factor for resource monitoring: (0.7875 + 0.85) × 0.5 ≈ 0.8188.
[0061] The accuracy and completeness of alarm reports in multi-level alarm information and detailed resource status reports are quantitatively analyzed to generate alarm effectiveness quantification factors. For multi-level alarm information and detailed resource status reports, alarm effectiveness factors are generated from two dimensions: alarm accuracy (the proportion of actual alarms to total alarms) and status report completeness (the degree to which the report contains key information). The proportion of actual resource anomalies among warning, critical, and emergency alarms is statistically analyzed. For example, the system issued a total of 20 alarms, of which 17 matched the actual anomalies (3 were false alarms), with an alarm accuracy rate of 85%. All 5 emergency alarms were true, 8 out of 10 critical alarms were true, and 4 out of 5 warning alarms were true. The weighted accuracy rate calculated according to the level weights (emergency level 0.4, critical level 0.3, warning level 0.3) is: (5 / 5×0.4)+(8 / 10×0.3)+(4 / 5×0.3)=0.4+0.24+0.24=0.88.
[0062] The check report includes key elements such as the current status of resources, trend analysis, and recommended measures, and scores it on a completeness scale (0-1). For example, if 80% of the reports contain all elements and 20% lack trend analysis, the completeness factor is 0.8×1 + 0.2×0.8 = 0.96. The alarm accuracy rate and report completeness are weighted and combined (weights are 0.6 and 0.4 respectively) to generate an alarm effectiveness quantification factor: 0.88×0.6 + 0.96×0.4 = 0.528 + 0.384 = 0.912.
[0063] Based on fundamental quantitative factors for resource monitoring and quantitative factors for alarm effectiveness, and combined with the hierarchical architecture of the heterogeneous resource end-to-end evaluation model, the execution success rate of automatic response solutions, the prediction accuracy of future resource demand trends, and the prediction accuracy of resource planning suggestions are integrated to generate correlation quantitative features between monitoring, response, and prediction. This is achieved by incorporating the execution success rate of automatic response solutions and the accuracy of future demand predictions into the hierarchical architecture (monitoring layer, response layer, prediction layer) of the heterogeneous resource end-to-end evaluation model.
[0064] The percentage of automated response measures that actually achieve the expected results is calculated. For example, in 10 resource expansions, 9 successfully reduced the load, and in 8 load migrations, 7 balanced the node load. The success rate is (9 / 10 + 7 / 8) × 0.5 = (0.9 + 0.875) × 0.5 = 0.8875. The deviation between predicted and actual resource demands is compared. For example, a prediction of "monthly GPU demand growth of 5 containers" is correct, while the actual growth is 4, a deviation of 20%; a prediction of "monthly storage growth of 800GB" is correct, while the actual growth is 750GB, a deviation of 6.25%. The accuracy factor is (1 - 0.2) × 0.5 + (1 - 0.0625) × 0.5 = 0.4 + 0.46875 = 0.86875.
[0065] Combining the factors from the first two steps (monitoring foundation 0.8188, alarm effectiveness 0.912), and fusing them according to hierarchical weights (monitoring layer 0.3, alarm layer 0.2, response layer 0.3, prediction layer 0.2), the monitoring-response-prediction correlation quantitative feature is generated: 0.8188×0.3+0.912×0.2+0.8875×0.3+0.86875×0.2≈0.2456+0.1824+0.2663+0.1738≈0.8681.
[0066] Based on a heterogeneous resource end-to-end evaluation model, the quantitative characteristics of the monitoring-response-prediction correlation are analyzed and processed to generate comprehensive evaluation information on the unified monitoring and multi-level alarm effects of heterogeneous resources for a data-annotated cloud platform. This evaluation includes resource monitoring coverage, alarm response efficiency, and platform stability improvement rate. The heterogeneous resource end-to-end evaluation model analyzes the quantitative characteristics of the correlation and maps them to three core indicators: resource monitoring coverage, alarm response efficiency, and platform stability improvement rate, forming a comprehensive evaluation result. The resource monitoring coverage is generated by mapping the resource coverage breadth factor, and combined with the coverage of the total resource categories on the platform, it is calculated to be 82% (i.e., the proportion of resource types and indicators monitored by the system is 82%).
[0067] Alarm response efficiency is comprehensively mapped by alarm accuracy, report completeness, and response execution success rate, reflecting the efficiency of the entire process from alarm to resolution. The calculated efficiency is 89% (average response time for emergency alarms < 5 minutes, critical alarms < 30 minutes). Comparing platform downtime and task latency before and after implementing this method, for example, downtime decreased from 120 minutes to 40 minutes per month, and task latency decreased from 15% to 5%, resulting in a stability improvement of ((120-40) / 120 + (15%-5%) / 15%) × 0.5 ≈ (0.6667 + 0.6667) × 0.5 ≈ 66.67%. The final comprehensive evaluation information is: "Resource monitoring coverage 82%, alarm response efficiency 89%, platform stability improvement 67%", fully reflecting the actual effectiveness of the unified monitoring and multi-level alarm system, providing a quantitative basis for subsequent optimization.
[0068] In one implementation, such as Figure 2 As shown, this application also provides a unified monitoring and multi-level alarm device for heterogeneous resources in a data annotation cloud platform, including:
[0069] The acquisition module 201 is used to acquire real-time performance data of heterogeneous resources for the data annotation cloud platform;
[0070] Processing module 202 is used to process the acquired real-time performance data of heterogeneous resources, capture the usage status and changing trends of various resources, and generate real-time data and preliminary assessment information of the overall resource usage. It processes the real-time data and preliminary assessment information of the overall resource usage, and based on the dynamic performance threshold management module, sets multi-level thresholds according to resource usage patterns and load conditions, applies characteristic constraints of different resource types and platform stability operation objective functions, and generates corresponding levels of alarm information and detailed resource status reports. It processes alarm information and detailed resource status reports at different levels, and based on the alarm level, automatically executes response measures such as resource expansion, load migration, and fault recovery. The system dynamically adjusts resource allocation to generate automatic response plans that ensure the platform's continuous and stable operation. These plans are then processed using a data analysis and trend prediction module. Through machine learning and data mining techniques, historical data is analyzed to identify resource usage patterns and generate future resource demand trend predictions and resource planning suggestions. Real-time data and preliminary assessments of overall resource usage, multi-level alarm information and detailed resource status reports, automatic response plans, and future resource demand trend predictions and resource planning suggestions are processed to generate comprehensive evaluation information for the unified monitoring of heterogeneous resources and the effectiveness of multi-level alarms on a data-annotated cloud platform. This evaluation includes resource monitoring coverage, alarm response efficiency, and platform stability improvement rates.
[0071] The various embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for evaluating the unified monitoring and multi-level alarm method, electronic device, electronic device, and readable storage medium for heterogeneous resources of a data annotation cloud platform are relatively simple to describe because they are basically similar to the above-described embodiments of the unified monitoring and multi-level alarm method for heterogeneous resources of a data annotation cloud platform. Relevant parts can be referred to the descriptions of the above-described embodiments of the unified monitoring and multi-level alarm method for heterogeneous resources of a data annotation cloud platform.
Claims
1. A method for unified monitoring and multi-level alarming of heterogeneous resources for data annotation cloud platforms, characterized in that, include: Obtain real-time performance data of heterogeneous resources for the data annotation cloud platform, including CPU utilization, GPU load, storage space utilization, disk I / O, bandwidth utilization, network latency, number of online users, and annotation task time. The acquired real-time performance data of heterogeneous resources is processed to capture the usage status and changing trends of various resources, and to generate real-time data and preliminary assessment information on the overall usage of resources. The system processes real-time data and preliminary assessment information on the overall resource usage. Based on the dynamic performance threshold management module, it sets multi-level thresholds by combining resource usage patterns and load conditions, applies characteristic constraints of different resource types and platform stability operation objective functions, and generates corresponding alarm information and detailed resource status reports. The system processes alarm information and detailed resource status reports at different levels, and automatically executes response measures such as resource expansion, load migration, and fault recovery based on the alarm level. It also dynamically adjusts resource configuration and generates automatic response plans to ensure the continuous and stable operation of the platform. The automatic response scheme is processed, and the data analysis and trend prediction module is used to analyze historical data through machine learning and data mining techniques to identify resource usage patterns and generate future resource demand trend predictions and resource planning suggestions. The system processes real-time data and preliminary assessment information on overall resource usage, multi-level alarm information and detailed resource status reports, automatic response plans, future resource demand trend predictions and resource planning suggestions to generate comprehensive evaluation information on the unified monitoring of heterogeneous resources and multi-level alarm effects for data-annotated cloud platforms. This information includes resource monitoring coverage, alarm response efficiency, and platform stability improvement rate.
2. The method as described in claim 1, characterized in that, The acquired real-time performance data of heterogeneous resources is processed to capture the usage status and changing trends of various resources, generating real-time data and preliminary assessment information on the overall resource usage, including: Real-time aggregation and multi-dimensional analysis of computing resource data, storage resource data, network resource data, and business operation data in heterogeneous resource real-time performance data are performed to generate computing resource load characteristics, storage resource occupancy characteristics, network resource transmission characteristics, and business operation status characteristics, which are then summarized to form resource usage status characteristic parameters. Dynamically track the changing trends of various resources and mine periodic patterns to generate characteristics of computing resource fluctuation trends, storage resource growth trends, network resource load trends, and business indicator changes, and integrate them to form resource trend characteristic parameters; By matching and comparing resource usage status characteristic parameters and resource trend characteristic parameters with a preset resource baseline status database, real-time data and preliminary assessment information on the overall resource usage are generated through resource usage assessment and potential risk identification.
3. The method as described in claim 1, characterized in that, The system processes real-time data and preliminary assessment information on overall resource usage. Based on the dynamic performance threshold management module, it sets multi-level thresholds according to resource usage patterns and load conditions, applies characteristic constraints for different resource types and objective functions for platform stability, and generates corresponding alarm information and detailed resource status reports, including: Dynamically calculate and match usage patterns for the current resource utilization rate, peak load, and response time in real-time data of overall resource usage, and generate multi-level threshold features for computing resources, critical interval features for storage resources, and early warning line features for network resources, forming the basic information for threshold setting. The resource characteristic differences and load fluctuation patterns in the preliminary assessment information are transformed into constraints and the type boundaries are divided to generate computing resource characteristic constraint parameters, storage resource usage limit parameters, and network resource performance constraint parameters, thus forming resource constraint information. For the objective function of platform stable operation, parameters are adapted and the degree of objective achievement is evaluated in combination with resource usage status, generating resource load balancing target characteristics and platform stability matching parameters, thus forming objective function optimization information; Integrate basic threshold setting information, resource constraint information, and objective function optimization information to generate alarm information and detailed resource status reports at the corresponding levels.
4. The method as described in claim 1, characterized in that, The system processes alarm information and detailed resource status reports at different levels, and automatically executes response measures such as resource expansion, load migration, and fault recovery based on the alarm level. It dynamically adjusts resource configuration and generates automated response plans to ensure the platform's continuous and stable operation, including: For different levels of alarm information, such as warning prompts, critical alerts, and emergency alarms, response strategies are matched and the degree of automation is graded to generate warning-level manual intervention suggestion features, critical-level semi-automatic processing features, and emergency-level fully automated response features, thus forming alarm response strategy information. Perform resource scheduling feasibility analysis and adjustment cost assessment on the resource load distribution, node health status, and performance bottleneck location data in the detailed resource status report, generate load migration path characteristics, resource expansion priority parameters, and fault node replacement scheme characteristics, and form basic information for resource adjustment. By combining alarm levels and dynamic adjustment requirements, the execution sequence planning and conflict resolution are carried out on alarm response strategy information and resource adjustment basic information to generate resource configuration dynamic adjustment sequence characteristics and multi-measure collaborative execution coefficients, thus forming response optimization information. By integrating alarm response strategy information, basic resource adjustment information, and response optimization information, an automatic response plan is generated to ensure the continuous and stable operation of the platform.
5. The method as described in claim 4, characterized in that, The automated response plan is processed using a data analysis and trend prediction module. Through machine learning and data mining techniques, historical data is analyzed to identify resource usage patterns and generate future resource demand trend predictions and resource planning recommendations, including: The automatic response plan is decomposed and analyzed to generate historical resource adjustment characteristics, response measure effectiveness evaluation characteristics, and dynamic changes in resource allocation, which together constitute the core elements of the response plan. For the resource-load forecasting model, based on the core element information of the response plan and the fluctuation pattern of resource usage, a time series pattern recognition algorithm is used to set the key parameters of the model and generate the initial information for model forecasting. Based on the initial information predicted by the model, the core element information of the response plan is processed by the resource load correlation analysis mechanism to generate initial trend prediction feature information that includes the trend characteristics of computing resource demand and the correlation characteristics of storage and network resource demand. The initial trend prediction feature information is processed based on the coupling information of historical data and real-time status to generate a fusion trend feature information of resource usage history and real-time. A multi-resource collaborative prediction mechanism is adopted to process historical and real-time trend information of resource use to generate future resource demand trend predictions and resource planning suggestions.
6. The method as described in claim 1, characterized in that, The system processes real-time data and preliminary assessment information on overall resource usage, multi-level alarm information and detailed resource status reports, automatic response plans, future resource demand trend predictions, and resource planning suggestions to generate comprehensive evaluation information on the unified monitoring and multi-level alarm effects of heterogeneous resources on a data-annotated cloud platform. This evaluation includes resource monitoring coverage, alarm response efficiency, and platform stability improvement rates. The real-time data and preliminary assessment information on the overall resource usage are quantitatively analyzed to generate basic quantitative factors for resource monitoring. Quantitative analysis and processing of alarm accuracy and status report completeness characteristics in multi-level alarm information and detailed resource status reports are performed to generate alarm effectiveness quantification factors. Based on the basic quantitative factors of resource monitoring and the quantitative factors of alarm effectiveness, combined with the hierarchical architecture of the heterogeneous resource full-link evaluation model, the execution success rate in the automatic response scheme, the prediction of future resource demand trends and the prediction accuracy in resource planning suggestions are integrated to generate quantitative features related to monitoring, response and prediction. Based on the heterogeneous resource full-link evaluation model, the quantitative characteristics of the monitoring-response-prediction correlation are analyzed and processed to generate comprehensive evaluation information on the unified monitoring and multi-level alarm effects of heterogeneous resources for data annotation cloud platforms, including resource monitoring coverage, alarm response efficiency, and platform stability improvement rate.
7. A unified monitoring and multi-level alarm device for heterogeneous resources in a data annotation cloud platform, characterized in that, The device includes: The acquisition module is used to acquire real-time performance data of heterogeneous resources for the data annotation cloud platform; The processing module processes the acquired real-time performance data of heterogeneous resources, captures the usage status and trends of various resources, and generates real-time data and preliminary assessment information on the overall resource usage. It then processes this data and preliminary assessment information, using a dynamic performance threshold management module to set multi-level thresholds based on resource usage patterns and load conditions. This applies characteristic constraints for different resource types and a platform stability objective function, generating corresponding alarm information and detailed resource status reports. Finally, it processes alarm information and detailed resource status reports at different levels, and based on the alarm level, automatically executes response measures such as resource expansion, load migration, and fault recovery, dynamically... Adjust resource allocation to generate an automatic response plan that ensures the platform's continuous and stable operation. Process the automatic response plan, and use the data analysis and trend prediction module to analyze historical data through machine learning and data mining techniques to identify resource usage patterns and generate future resource demand trend predictions and resource planning suggestions. Process real-time data and preliminary assessment information on overall resource usage, multi-level alarm information and detailed resource status reports, automatic response plans, future resource demand trend predictions and resource planning suggestions to generate comprehensive evaluation information on the unified monitoring of heterogeneous resources and the effectiveness of multi-level alarms for the data-annotated cloud platform. This evaluation includes resource monitoring coverage, alarm response efficiency, and platform stability improvement rate.
8. An electronic device, characterized in that, include: First processor; The processor includes a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the heterogeneous resource unified monitoring and multi-level alarm method for a data annotation cloud platform as described in any one of claims 1 to 6 by executing the executable instructions.