Equipment management method and system based on task priority evaluation

Through the equipment management method based on task priority evaluation in the data warehousing system, using artificial intelligence to predict faults and dynamically adjust resource scheduling, the problems of unreasonable task priority settings and lagging fault processing are solved, and the intelligent and efficient utilization of equipment resources is achieved, and business continuity and data reliability are improved.

CN120448075AActive Publication Date: 2025-08-08PARTNER WISDOM (BEIJING) INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510907584.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-08
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Equipment management in data warehousing systems faces problems such as unreasonable task priority settings, lagging fault processing and rigid scheduling strategies, resulting in low resource utilization efficiency and lagging business response.

Method used

Through the device management method based on task priority evaluation, the artificial intelligence model is used to predict faults, determine the priority of equipment tasks based on business labels and scheduling scenarios, and dynamically adjust resource scheduling and fault handling strategies to achieve intelligent and efficient utilization of equipment.

Benefits of technology

It realizes fine allocation of task priority, ensures that high-priority tasks obtain resources and troubleshoot in a timely manner, improves business response efficiency and data reliability, reduces resource waste, and improves operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448075A_ABST
    Figure CN120448075A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an equipment management method and system based on task priority evaluation. The method comprises the following steps: determining first equipment and a scheduling scene according to a service label of a task; determining a task priority of the first equipment at least according to a service parameter and the scheduling scene; the task priority is at least used for task resource scheduling and fault processing; predicting a fault occurrence condition when the first equipment processes the task based on an artificial intelligence (AI) model, and obtaining fault prediction information; starting a fault processing plan according to the task priority when the fault prediction information indicates that a specified time period meets a specified condition; and when a fault occurs, performing management and fault processing of the first equipment according to the fault processing plan.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data technology, and in particular to a device management method and system based on task priority evaluation. Background Art

[0002] Many enterprises store their data in cloud-based data warehouses. Data warehouses are provided by data storage companies to provide data storage and enterprise-level data decision-making to other enterprises. In data warehouse systems, device management (storage, computing, and network devices) faces challenges such as irrational task prioritization, delayed fault handling, and rigid scheduling strategies. Summary of the Invention

[0003] In view of this, an embodiment of the present invention provides a device management method and system based on task priority evaluation. The technical solution of the present invention is implemented as follows: A first aspect provides a device management method based on task priority evaluation, which is applied to a data warehousing system. The method includes: Determine the first device and the scheduling scenario based on the business tag of the task; Determining a task priority of the first device based at least on the business parameters and the scheduling scenario; the task priority is used at least for resource scheduling and fault handling of the task; Performing fault prediction of the first device based on an artificial intelligence (AI) model to obtain fault prediction information; When the fault prediction information indicates that the specified time period meets the specified conditions, the fault handling plan is initiated according to the task priority; When a fault occurs, management and fault handling of the first device are performed according to the fault handling plan.

[0004] The second party provides a device management system based on task priority evaluation, which is applied to a data warehousing system. The device management system includes: A first determining module, configured to determine a first device and a scheduling scenario according to a service tag of a task; A second determining module is configured to determine a task priority of the first device based at least on the business parameters and the scheduling scenario; the task priority is used at least for resource scheduling and fault handling of the task; An acquisition module, configured to predict, based on an artificial intelligence (AI) model, a failure occurrence condition when the first device processes the task, and obtain failure prediction information; A starting module, configured to start a fault handling plan according to the task priority when the fault prediction information indicates that a specified time period meets a specified condition; The processing module is used to manage and handle the fault of the first device according to the fault handling plan when a fault occurs.

[0005] A third aspect provides a computer-readable storage medium storing computer-executable instructions. After the computer-executable instructions are executed by a processor, the aforementioned device management method based on task priority evaluation can be implemented.

[0006] The technical solution provided by the embodiments of the present disclosure assigns differentiated task priorities to device tasks based on business tags (and scheduling scenarios), achieving refined allocation of task priorities, ensuring that high-priority critical tasks receive priority in resource acquisition and fault handling, and improving the efficiency of core business response. An AI model is used to predict failure probabilities in advance, and emergency plans are initiated in advance based on the predicted probabilities, ensuring timely response when failures occur. Furthermore, by not initiating emergency plans when they are not necessary, the waste of resources corresponding to the emergency plans is reduced. In short, this technical solution, through "priority-driven resource scheduling + AI prediction and proactive prevention and control + scenario-based policy adaptation," addresses the pain points of extensive task priority management and delayed fault response in data warehousing systems, achieving intelligent and efficient utilization of device resources, and significantly improving business continuity, data reliability, and operation and maintenance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a device management method based on task priority evaluation provided by an embodiment of the present invention; Figure 2 A flowchart of another device management method based on task priority evaluation provided by an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a device management system based on task priority evaluation provided by an embodiment of the present invention; Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0008] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0009] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0010] like Figure 1 As shown, an embodiment of the present disclosure provides a device management method based on task priority evaluation, which is applied to a data storage system. The method includes: S1110: Determine a first device and a scheduling scenario based on a service tag of the task; the first device includes at least one of a storage device, a computing device, and a network device; S1120: Determine a task priority of the first device based at least on the service parameters and the scheduling scenario; the task priority is used for resource scheduling and fault handling of the task; S1130: Predicting a fault occurrence when the first device processes a task based on the artificial intelligence (AI) model, and obtaining fault prediction information; S1140: When the fault prediction information indicates that the specified time period meets the specified conditions, the fault handling plan is initiated according to the task priority; S1150: When a fault occurs, manage and handle the fault of the first device according to the fault handling plan.

[0011] In some embodiments, the device management method based on task priority assessment can be used in a management system of a data storage system. The management system includes one or more electronic devices. For example, the device management method based on task priority assessment can be used in one or more electronic devices. In other embodiments, the management system can also be a subsystem integrated into the data storage system.

[0012] Semantic tags are pre-assigned to devices used in the data warehousing system to achieve semantic grouping of devices. For example, data warehousing devices (such as compute nodes, storage clusters, and ETL servers) are tagged according to their business functions (e.g., "Real-time Analysis Cluster," "Historical Data Archiving Storage," and "Data Cleansing Node"). The priority scheduling system automatically matches device groups with corresponding tags based on task types, prioritizing idle resources in device groups belonging to high-priority tasks. This allows the system to quickly identify the first device required for a task based on this semantic grouping. This introduces a business semantic layer, shortening the task-device matching process and improving scheduling efficiency.

[0013] The first device is any device in the data storage system. Classified by device type, the first device includes at least one of a storage device, a computing device, and a network device. Based on the hotness or coldness of the data processed by the device, it can be categorized as either a cold data device or a hot data device. The hotness (access frequency) of data processed by a cold data device is lower than that of data processed by a hot data device.

[0014] In some embodiments, the scheduling scenario can be used to describe the current state of the data warehouse system when scheduling resources for this task, but is not limited to describing the current state of the data warehouse system. Thus, in some embodiments, the scheduling scenario can be determined based on the current state of the data warehouse system.

[0015] In some embodiments, scheduling scenarios can be divided into daily scheduling scenarios and non-daily scheduling scenarios. Daily scheduling scenarios can be further divided into different daily scheduling scenarios based on the proportion of currently executing tasks. For example, if the current storage task accounts for the highest proportion, then the daily scheduling scenario in the current state is a storage-centralized scenario. For another example, if the current computing task accounts for the highest proportion, then the daily scheduling scenario in the current state is a computing-centralized scenario. For another example, if the current access task accounts for the highest proportion, then the daily scheduling scenario in the current state may be an access-centralized scenario. Non-daily scheduling scenarios can also be called abnormal scheduling scenarios. Typical abnormal scheduling scenarios include but are not limited to fault handling scenarios and storage optimization scenarios to avoid storage overload. Storage optimization scenarios are different from storage-centralized scenarios. Storage-centralized scenarios are data storage scenarios triggered by user storage needs. Storage optimization scenarios are storage scenarios related to data migration, compression, or redundancy deletion, etc., which are triggered by the storage needs of the data warehouse system itself. The above are just examples of scheduling scenarios, and the specific implementation is not limited to this example.

[0016] In some embodiments, the business parameters may be used to describe the business corresponding to the task, including but not limited to at least one of the following: Business impact, for example, indicates whether the task affects the generation, update, or deletion of core reports. Business value coefficient, for example, the business value coefficient is used to indicate the importance of the business, for example, it can be divided into core business, mid-level business, marginal business, etc. according to importance; Business urgency, for example, the business urgency is used to indicate the maximum response delay or average response delay allowed for the task corresponding to the business; Failure loss index, for example, is used to indicate the penalty effect caused by the failure of the task response corresponding to the business.

[0017] In the embodiment of the present disclosure, the task priority corresponding to the task is determined in combination with the business parameters and the scheduling scenario.

[0018] This task priority is used for both resource scheduling and troubleshooting. This allows resource scheduling and troubleshooting to share a common task priority, achieving synchronization between the two priorities. This reduces the computational overhead associated with priority calculation compared to setting separate priorities for resource scheduling and troubleshooting. For example, task priority is related to data popularity: devices processing cold data are called cold data devices, while devices processing hot data are called hot data devices. Device priority rules are designed based on the storage characteristics of "hot" (highly accessed) and "cold" (low-frequency archived) data in data warehousing. For tasks like hot data queries and real-time synchronization, SSD storage devices and high-speed compute nodes are prioritized. For batch migration and backup of cold data, mechanical hard drives and low-power devices are used, allowing them to be preempted by higher-priority tasks. Integrating data lifecycle management (DLM) concepts, data characteristics are linked to device priorities to optimize resource utilization.

[0019] In some embodiments, S1120 may include: Determine the objective function for computing task priority based on the scheduling scenario; The task priority of the first device is calculated using the service parameter as a variable of the objective function.

[0020] In the embodiment of the present disclosure, different scheduling scenarios correspond to different objective functions for calculating task priorities. Thus, the task priority of the task is not only related to the business of the task but also to the current scheduling scenario, that is, the business associated with the task and the current status of the data warehouse system are comprehensively considered.

[0021] In some embodiments, the objective function includes a first term and a second term; and the objective function for determining the priority of a computing task according to a scheduling scenario includes: The first score of the first item is determined based on the depth of the task's bloodline hierarchy and the number of downstream affected nodes; Determine the target calculation formula for the second item based on the scheduling scenario; Using the business parameter as a variable in the target calculation formula to obtain a second score; The task priority is determined according to the first score and the second score.

[0022] In some embodiments, ,or, ,in, The first score. The level of a task in the data lineage relationship. For example, raw data collection is level 1, intermediate data processing is level 2, and report generation is level 3. The higher the level and the closer it is to the business endpoint, the higher the task priority.

[0023] Number of downstream affected nodes: The number of downstream tasks / reports affected by the task. For example, affecting 10 reports = 10 points, affecting only 1 report = 1 point. and The weight coefficient can be pre-set according to the importance of the hierarchy in the data lineage relationship and the number of downstream influencing nodes. The weight coefficient can be generated by the AI model through big data collection and processing.

[0024] It can be seen that the complexity of the task is related to the number of downstream nodes associated with the first device. The data involved in the task will affect the hierarchy of the data lineage relationship.

[0025] For example, when a raw data collection task fails, although it's at the bottom of the hierarchy (level = 1), it can potentially impact all downstream levels, resulting in a priority of 1 × N (where N is the number of downstream nodes). If a mid-level task fails, its priority is 2 × M (where M is the number of downstream nodes), typically higher than the bottom level but lower than the terminal task. In this way, task priority can effectively reflect the potential chain reaction of a failure through this data lineage, facilitating the subsequent determination of whether to set up a fault contingency plan and what type of contingency plan to set based on task priority.

[0026] In some embodiments, the target calculation formula for the second item is determined based on the scheduling scenario, including at least one of the following: When the scheduling scenario is a conventional scheduling scenario, determining the target calculation formula to be a first calculation formula, a second calculation formula, a third calculation formula, or a fourth calculation formula, wherein the first calculation formula, the second calculation formula, the third calculation formula, or the fourth calculation formula are all related to resource scheduling, and any two of the first calculation formula, the second calculation formula, the third calculation formula, and the fourth calculation formula have different variables and / or different calculation functions; When the scheduling scenario is a storage optimization scenario, the target calculation formula is determined to be the fifth calculation formula; the fifth calculation formula is related to data access storage; When the scheduling scenario is a fault handling scenario, the target calculation formula is determined to be the sixth calculation formula; the sixth calculation formula is related to fault handling.

[0027] In some embodiments, the first calculation formula may be related to business impact, data sensitivity, time window flexibility, and resource starvation risk.

[0028] , is the business impact. Data sensitivity level, such as whether it involves private data. Time window flexibility, such as whether the task can be delayed. =Resource starvation risk, such as whether long waiting time will cause the device to be idle. In this case, using the first calculation formula, The second score is calculated using the first formula, and this second score takes into account at least four dimensions. Compared with traditional task priorities that are mostly based on a single indicator (such as time urgency), this framework introduces business and data characteristic dimensions for data warehousing scenarios, which is more in line with industry needs. 、 、 、 The weight can be dynamically adjusted.

[0029] In some embodiments, the second calculation formula may be related to the business value coefficient, the failure weight, and the resource consumption coefficient of the task.

[0030] For example, or ,in, is the business value coefficient, is the failure weight. is the resource consumption coefficient. Table 1 may be a description of the business value coefficient, failure weight, and resource consumption coefficient.

[0031] Table 1

[0032] For example, if Task A has a business value of 8 points, a timeliness of 10 points, and a resource consumption of 3 points, then the task priority is (8 × 10) / 3 ≈ 26.67. Task B has a business value of 5 points, a timeliness of 4 points, and a resource consumption of 8 points, then the task priority is (5 × 4) / 8 = 2.5. Based on the calculated task priorities, Task A takes precedence over Task B due to its higher business value, stronger timeliness, and lower resource consumption.

[0033] In some embodiments, the second calculation formula may be related to one or more of business urgency, data consistency, calculation complexity, compliance requirements, and resource preemption risk.

[0034] In some embodiments, ,in, is the second score calculated using the third calculation formula. is the i-th calculated variable. is the weight corresponding to the i-th calculated variable. is the total number of calculated variables. Table 2 is an example of the calculated variables of the third calculation formula.

[0035] Table 2

[0036] In some embodiments, the fourth calculation formula is related to the task urgency, the data locality coefficient, the estimated execution time, and the queue waiting time.

[0037] For example, or ,in, is the second score calculated using the fourth calculation formula. The urgency of the task. is the data locality coefficient. Estimated execution time. The queue waiting time. and Can be a preset weight.

[0038] Data locality coefficient: measures whether the task is executed locally on the data storage node (local = 1.5, remote = 1), optimizing the efficiency of "moving computing to data" in data warehousing.

[0039] Estimated execution time: The task duration is estimated based on historical data, for example, in minutes, normalized to 0.1-10.

[0040] Queue waiting time: The length of time a task has been waiting in the queue. For example, to avoid starvation, the longer the waiting time, the higher the weight.

[0041] Since there are alternative calculation formulas from the first to the fourth in daily scheduling scenarios, a target calculation formula suitable for the working mechanism of the data warehousing system can be selected. For example, the fourth calculation formula can be applied to daily scenarios where the data warehousing system uses queue scheduling tasks.

[0042] In some embodiments, for storage optimization scenarios, the fifth calculation is related to parameters such as access popularity, data capacity, and update frequency. This is because in storage optimization scenarios, both current data access and data optimization need to be considered.

[0043] For example, ,in, is the second score calculated using the fifth calculation formula. Data access heat. is the data capacity. The frequency of data update. 、 、 Can be weight.

[0044] Access popularity: the number of queries / writes per unit time (e.g., normalized to a 0-10 scale, 100 times / day = 10 points, 10 times / day = 3 points).

[0045] Data capacity: data size (e.g. GB, 100GB = 10 minutes, 1GB = 1 minute).

[0046] Update frequency: Number of updates per day (e.g. real-time updates = 10 points, weekly updates = 2 points).

[0047] In some embodiments, user profile data has an access popularity score of 8, a capacity score of 6, and an update frequency score of 7. Priority = 0.5 × 8 + 0.3 × 6 + 0.2 × 7 = 4 + 1.8 + 1.4 = 7.2. Historical log data has an access popularity score of 2, a capacity score of 10, and an update frequency score of 1. Priority = 0.5 × 2 + 0.3 × 10 + 0.2 × 1 = 1 + 3 + 0.2 = 4.2. In this case, when scheduling storage resources, the following storage policy can be adopted based on task priority: user profile data is preferentially stored in high-performance storage tiers (such as SSDs), while historical logs are archived to low-cost storage (such as HDDs or cloud object storage).

[0048] In some embodiments, the fault handling scenario may use the sixth calculation formula, for example, or ,in, is the second score calculated using the sixth calculation formula. For compliance risk. is the fault impact coefficient. is the recovery complexity index. Table 3 shows the variables in the fifth calculation formula.

[0049] Table 3

[0050] For example, Task C has a compliance risk of 8 points, a failure impact of 10 points, and a recovery complexity of 9 points. Priority = 8 × 10 + 9 = 89. Task D has a compliance risk of 3 points, a failure impact of 4 points, and a recovery complexity of 4 points. Priority = 3 × 4 + 4 = 16. Therefore, when scheduling resources, Task C (such as real-time synchronization of core transaction data) requires immediate resource allocation, while Task D (such as a non-critical report generation failure) can be deferred.

[0051] In some embodiments, calculating the task priority according to the objective function may be performed when the task is received or before initiating task execution.

[0052] In some embodiments, the method further includes: adjusting the task priority up or down upon detecting at least one of a business period, a device status, and an external event.

[0053] That is, after the initial task priority calculation is performed according to the objective function, the priority will be dynamically adjusted according to the dynamic priority mechanism so that the currently calculated task priority conforms to the dynamic changes of the data warehouse system.

[0054] For example, based on the business period, for example, the priority of a batch loading task (P2) at night can be temporarily increased to P1 and then reduced back during the day.

[0055] As another example, based on the device status: when a storage array approaches a capacity threshold, the I / O priority of non-critical tasks (such as P3 archiving) is automatically reduced to ensure the writing of P1 tasks.

[0056] Also illustratively, based on external events: when a network attack is detected, the traffic cleaning task of the security protection device (such as a firewall) is automatically upgraded to P1 to block the attack traffic.

[0057] Dynamic Priority Adjustment Mechanism: Priority changes during task execution are designed. For example, when a high-priority task is triggered, lower-priority tasks can automatically enter a "suspended-cached" state (preserving their current progress and releasing some resources) rather than being terminated, thus avoiding data processing interruptions. Based on real-time data traffic fluctuations (such as sudden queries causing a surge in I / O load), the priority of associated tasks is dynamically increased to ensure business continuity. Traditional scheduling systems typically set priorities statically; this mechanism increases runtime flexibility and reduces resource waste.

[0058] In some embodiments, as Figure 2 As shown, S1130 may include: S1131: Determine key monitoring indicators and collect at least one of the following dimensional data for data storage devices (storage / computing / network devices) and task processing processes: backup operation data, task business data, and external environment data.

[0059] Equipment operation data includes but is not limited to at least one of the following: Storage device IOPS, throughput, latency, disk error rate (SMART indicator), and remaining capacity; Compute the device's CPU utilization, memory usage, process status, and task queue length; Bandwidth utilization, packet loss rate, port error number, and routing delay of network devices.

[0060] Task business data includes but is not limited to at least one of the following: Task priority tags (such as P1-P4) and business tags (such as real-time analysis and batch loading); Task execution parameters: data volume, processing time, input and output paths, and dependencies (such as data lineage hierarchy); Task result data: success rate, failure type (such as timeout, data format error, insufficient resources).

[0061] External environmental data includes but is not limited to at least one of the following: Business hours (such as peak / off-peak hours), system load (such as overall cluster utilization), and external events (such as network attacks and power failure warnings).

[0062] S1132: Data collection.

[0063] For example, monitoring tools such as Prometheus and Zabbix can be used to capture device indicators at a frequency of seconds; and the APIs of data warehousing platforms (such as Hadoop and Spark) can be used to obtain task metadata.

[0064] S1133: Data cleaning and feature engineering, specifically including: Missing value processing: For abnormal missing data (such as indicator loss caused by network interruption), use interpolation methods (such as linear interpolation, time series prediction) or delete invalid samples. Feature extraction may include at least one of the following: Time series features: Calculate sliding window statistics of device metrics (such as average I / O latency and CPU utilization peak over the past 5 minutes); Correlation characteristics: The correlation between task priority and device load (for example, whether storage device latency increases significantly when high-priority tasks are executed); Derived feature: Device health index (generated by combining multiple indicators, such as storage device health index = 0.6 × I / O latency + 0.3 × remaining capacity + 0.1 × error rate).

[0065] Data normalization: Scale numerical features to the [0,1] range (such as Min-Max normalization) to facilitate model training.

[0066] Through data cleaning and feature engineering, the authenticity and validity of data used for fault prediction can be ensured.

[0067] S1134: Failure warning of AI models.

[0068] In some cases, a data warehousing system may have multiple failure prediction models.

[0069] For example, the model type selection may include: selecting the following models based on the fault prediction target (such as classification of fault type, regression of fault probability): Time Series Forecasting Model: LSTM / GRU: Suitable for predicting single-device failure trends (e.g., predicting hard drive life based on historical disk error rates); Transformer: It handles long sequence dependencies and captures correlated failures between multiple devices (e.g., network device failure causing compute node task failure).

[0070] Also exemplarily, the fault prediction AI model may be a classification model, which may include but is not limited to at least one of the following: Random Forest (RF): used for fault type classification based on multi-feature fusion (e.g., distinguishing hardware faults from software faults); XGBoost / LightGBM: Efficiently process large-scale heterogeneous data and identify key failure influencing factors (such as the correlation between task data volume and storage device failure).

[0071] As another example, the AI model for predicting faults may be an anomaly detection model, which may include, but is not limited to: Isolation Forest: detecting anomalies in device indicators (such as sudden spikes in IO latency); Autoencoder: identifying device status anomalies through reconstruction errors (such as the memory usage pattern of a computing node deviating from the normal range).

[0072] Deploy the trained model as an API service (e.g., via Flask, TensorFlow Serving) and connect it to the monitoring platform of the data warehouse system.

[0073] In some other embodiments, dedicated prediction models are trained for different scheduling scenarios (such as storage optimization and fault handling): Storage optimization scenario: focuses on predicting storage device capacity bottlenecks and IO performance degradation; Fault handling scenarios: Optimize the ability to predict cascading failures (e.g., a switch failure causing multiple server task failures). In this case, select the fault prediction AI model based on the scheduling scenario, making the AI model's fault prediction more accurate.

[0074] In some embodiments, the prediction frequency of the AI model may include, but is not limited to, at least one of the following: High-priority devices (such as primary storage clusters): perform predictions once per second; Low-priority devices (such as archive storage nodes): predictions are performed once every minute.

[0075] In some embodiments, the fault prediction information may include various relevant information describing possible faults of the first device when processing a task. For example, the fault prediction information may include, but is not limited to, at least one of the following: the probability of a fault occurring, the possible type of fault, the severity of the possible fault, the expected time of the fault occurrence, the possible cause of the fault, and the like.

[0076] In some embodiments, failure warning of the AI model.

[0077] Fault warnings are performed based on dynamic threshold strategies. For example, thresholds are set based on historical data quantiles (e.g., a warning is triggered when the device health index is >0.8). Thresholds are adjusted based on task priorities: high-priority tasks correspond to lower device failure probability thresholds (e.g., the threshold for P1 devices is set to 10%, and for P4 devices is set to 30%).

[0078] Warning levels may include but are not limited to at least one of the following: Yellow warning: The failure probability exceeds the threshold but is lower than the emergency value, triggering preventive maintenance (such as increasing spare parts inventory and dispatching idle resources); Red alert: If the failure probability is ≥ the emergency value (e.g., >50%), the disaster recovery plan is immediately initiated (e.g., hot migration tasks to backup devices).

[0079] In some embodiments, the task priority includes a first priority and a second priority; the first priority is higher than the second priority. S1140 may include at least one of the following: when the task priority is the first priority, generating a first backup or a second backup based on the recoverability of the task; the first backup includes a first disaster recovery link; the first disaster recovery link is used to migrate the task from the first device to the second device in a first migration mode; the first device and the second device are backup devices of the same type; the second backup includes: scheduling a compensation task of the first priority; When the task priority is the second priority, a third or fourth record is generated according to the recoverability of the task; the third record includes a second disaster recovery link; the second disaster recovery link is used to migrate the task from the first device to the second device in a second migration mode; the fourth record includes: compensation task scheduling of the second priority.

[0080] For example, when the task priority is the first priority, the task can resume generating the first record, otherwise it can generate the second record. When the task priority is the second priority, the task can resume generating the third record, otherwise it can generate the fourth record.

[0081] The first migration mode and the second migration mode may be different types of migration modes. For example, the migration rate of the first migration mode is higher than the migration rate of the second migration mode. The user perception of the first migration mode is lower than the user perception of the second migration mode.

[0082] In some embodiments, the first migration mode may be hot migration and the second migration mode may be cold migration. Alternatively, the first migration mode may be jump migration and the second migration mode may be step-by-step migration. Detailed descriptions of these four migration modes are shown in Table 4 below.

[0083] Table 4

[0084] In some embodiments, the first migration mode and the second migration mode are a combination of multiple migration modes.

[0085] For example, the first migration method can be: hot migration + step-by-step migration: For example, in a microservice architecture, first hot migrate a single service (keeping the business running), and then migrate other modules step by step according to service dependencies.

[0086] For example, the second migration method can be: cold migration + jump migration: first complete data backup or system construction through cold migration, and then use jump migration to switch traffic at one time (for example, complete deployment in the new environment first, and then switch DNS).

[0087] For example, the first migration method might be to use hot migration for core business (to ensure continuity) and cold migration + gradual migration for non-core data (to reduce costs and risks). The second migration method might be to use cold migration + gradual migration for all business data.

[0088] In some embodiments, both the first disaster recovery link and the second disaster recovery link may include information about the second device, including an address of the second device. In some embodiments, the first disaster recovery link has fewer hops than the second disaster recovery link, thereby achieving rapid disaster recovery for the first-priority task.

[0089] In some embodiments, when the fault prediction information indicates that a specified period meets a specified condition, a fault handling plan is initiated based on the task priority. For example, the threshold value may be pre-set, for example, 50%.

[0090] In some embodiments, the fault prediction information indicates that a specified period meets a specified condition, including but not limited to at least one of the following: The fault prediction information indicates that the fault occurrence within the specified time period is higher than a first probability threshold; The fault prediction information indicates that a serious fault may occur within a specified time under a second probability threshold.

[0091] In some embodiments, the first probability threshold for the same task to have priority is higher than the second probability threshold.

[0092] In some embodiments, serious failures include but are not limited to failures corresponding to the aforementioned red warnings.

[0093] In some embodiments, different task priorities may correspond to different first probability thresholds and / or first probability thresholds, thereby enabling early activation of the fault handling plan corresponding to the task priority. For example, the threshold corresponding to the first priority is less than the first probability threshold and / or first probability threshold corresponding to the second priority. In some embodiments, the task priority also includes a third priority and a fourth priority; the task priority also includes a third priority and a fourth priority; the third priority is higher than the fourth priority, and the third priority is lower than the second priority. At this time, the first probability threshold and / or first probability threshold corresponding to the third priority is higher than the first probability threshold and / or first probability threshold of the second priority. The first probability threshold and / or first probability threshold corresponding to the fourth priority is higher than the first probability threshold and / or first probability threshold of the third priority.

[0094] In some embodiments, the fourth priority level may not require the pre-initiation of the fault handling plan. Alternatively, the third priority level or the fourth priority level may not require the pre-initiation of the fault handling plan.

[0095] In some embodiments, a priority-aware failover mechanism: When a device carrying a high-priority task fails, the system automatically triggers a "fast disaster recovery link." It prioritizes redundant nodes from the same-priority group to take over the task (rather than following the normal failover process). For tasks that cannot be quickly recovered, fault compensation tasks are automatically generated (e.g., raising their priority to the highest level during rescheduling), minimizing business impact. Traditional disaster recovery mechanisms do not prioritize tasks, but this design achieves "priority-sensitive" fault handling, reducing the interruption losses of high-value tasks.

[0096] When a fault occurs, the first device is managed and the fault is handled according to the fault handling plan. If the fault handling plan is not activated, conventional fault handling is performed. For example, if the first device fails while processing a priority 4 task, the fault handling plan may not be activated, so conventional fault handling is performed.

[0097] Troubleshooting can be performed based on work orders. For example, automated work orders can be generated for tasks of different priorities: P1 faults trigger an SMS / phone alert within 10 seconds, automatically assigning a senior engineer to arrive on-site within 30 minutes; P2 faults respond within 1 hour and are repaired within 4 hours; P3 / P4 faults are handled according to the work calendar and closed within 24 hours.

[0098] like Figure 3 As shown, an embodiment of the present disclosure provides a device management system based on task priority evaluation, which is applied to a data warehousing system. The device management system includes: A first determining module 3101 is configured to determine a first device and a scheduling scenario according to a service tag of a task; A second determining module 3102 is configured to determine a task priority of the first device based at least on the service parameters and the scheduling scenario; the task priority is used for resource scheduling and fault handling of the task; An acquisition module 3103 is configured to predict, based on an artificial intelligence (AI) model, a fault condition when the first device processes a task, and obtain fault prediction information; The starting module 3104 is used to start the fault handling plan according to the task priority when the fault prediction information indicates that the specified time period meets the specified conditions; The processing module 3105 is used to manage and handle the fault of the first device according to the fault handling plan when a fault occurs.

[0099] In some embodiments, the second determination module is specifically configured to determine an objective function for calculating the priority of a task according to a scheduling scenario; and calculate the priority of the task of the first device using a business parameter as a variable of the objective function.

[0100] In some embodiments, the objective function includes a first term and a second term; the second determination module is also used to determine the first score of the first term based on the depth of the bloodline hierarchy associated with the task and the number of downstream affected nodes; determine the target calculation formula of the second term based on the scheduling scenario; use the business parameters as variables in the target calculation formula to obtain the second score; and determine the task priority based on the first score and the second score.

[0101] In some embodiments, the second determination module is specifically used to execute at least one of the following: when the scheduling scenario is a conventional scheduling scenario, determining the target calculation formula to be the first calculation formula, the second calculation formula, the third calculation formula or the fourth calculation formula, wherein the first calculation formula, the second calculation formula, the third calculation formula or the fourth calculation formula are all related to resource scheduling, and any two of the first calculation formula, the second calculation formula, the third calculation formula and the fourth calculation formula have different variables and / or different calculation functions; when the scheduling scenario is a storage optimization scenario, determining the target calculation formula to be the fifth calculation formula; the fifth calculation formula is related to data access storage; when the scheduling scenario is a fault handling scenario, determining the target calculation formula to be the sixth calculation formula; the sixth calculation formula is related to fault handling.

[0102] In some embodiments, the task priority includes a first priority and a second priority; the first priority is higher than the second priority; the processing module is specifically used to generate a first record or a second record according to the recoverability of the task when the task priority is the first priority; the first record includes a first disaster recovery link; the first disaster recovery link is used to migrate the task from the first device to the second device in a first migration mode; the first device and the second device are backup devices of the same type; the second record includes: compensation task scheduling of the first priority; when the task priority is the second priority, a third record or a fourth record is generated according to the recoverability of the task; the third record includes a second disaster recovery link; the second disaster recovery link is used to migrate the task from the first device to the second device in a second migration mode; the fourth record includes: compensation task scheduling of the second priority.

[0103] In some embodiments, the task priority further includes a third priority and a fourth priority; the task priority further includes a third priority and a fourth priority; the third priority is higher than the fourth priority, and the third priority is lower than the second priority.

[0104] In some embodiments, the second determination module is specifically configured to determine an objective function for calculating the priority of a task according to a scheduling scenario; and calculate the priority of the task of the first device using a business parameter as a variable of the objective function.

[0105] In some embodiments, the second determination module is further configured to detect at least one of a business period, a device status, and an external event, and adjust the task priority up or down.

[0106] An embodiment of the present disclosure provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions; after the computer-executable instructions are executed by a processor, the device management method based on task priority evaluation provided by any of the aforementioned technical solutions can be implemented.

[0107] Combine Figure 4As shown, an embodiment of the present application provides an electronic device, which can be a component device of a device management system based on task priority evaluation, including a processor 10 and a memory 11. Optionally, the device may also include a communication interface 12 and a bus 9. The processor 10, the communication interface 12, and the memory 11 can communicate with each other through the bus 9. The communication interface 12 can be used for information transmission. The processor 10 can call the logic instructions in the memory 11 to execute the queue-based voiceprint data processing method of the above embodiment.

[0108] In addition, the logic instructions in the memory 11 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0109] Memory 11, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present application. Processor 10 executes the program instructions / modules stored in memory 11 to execute functional applications and data processing, thereby implementing the device management method based on task priority evaluation in the above-mentioned embodiments.

[0110] The memory 11 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function; the data storage area may store data generated based on the use of the electronic device. Furthermore, the memory 11 may include a high-speed random access memory and a non-volatile memory.

[0111] The technical solutions of the embodiments of the present application can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method of the embodiments of the present application. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, an optical disk, and other media that can store program code, or a transient storage medium.

[0112] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0113] The embodiments or examples disclosed in this application are not exhaustive, but are merely illustrations of some embodiments or examples, and are not intended to be specific limitations on the scope of protection of this disclosure. In the absence of contradiction, each step in a certain embodiment or example can be implemented as an independent example, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment or example can also be implemented as an independent example, and the order of the steps in a certain embodiment or example can be arbitrarily exchanged. In addition, the optional methods or optional examples in a certain embodiment or example can be arbitrarily combined; in addition, the various embodiments or examples can be arbitrarily combined. For example, some or all of the steps in different embodiments or examples can be arbitrarily combined, and a certain embodiment or example can be arbitrarily combined with the optional methods or optional examples of other embodiments or examples.

[0114] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0115] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.

[0116] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0117] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, and other media that can store program code.

[0118] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A device management method based on task priority evaluation, applied to a data warehousing system, characterized in that: The method comprises: Determine the first device and the scheduling scenario based on the business tag of the task; Determining a task priority of the first device based at least on the business parameters and the scheduling scenario; the task priority is used at least for resource scheduling and fault handling of the task; Predicting, based on an artificial intelligence (AI) model, a failure occurrence condition when the first device processes the task, and obtaining failure prediction information; When the fault prediction information indicates that the specified time period meets the specified conditions, the fault handling plan is initiated according to the task priority; When a fault occurs, management and fault handling of the first device are performed according to the fault handling plan.

2. The method according to claim 1, characterized in that Determining the task priority of the first device based at least on the service parameter and the scheduling scenario includes: Determining an objective function for calculating the task priority according to the scheduling scenario; The task priority of the first device is calculated using the service parameter as a variable of the objective function.

3. The method according to claim 2, characterized in that The objective function includes a first item and a second item; and determining the objective function for calculating the task priority according to the scheduling scenario includes: Determining a first score for the first item based on the depth of the bloodline hierarchy associated with the task and the number of downstream affected nodes; Determining a target calculation formula for the second item according to the scheduling scenario; Using the business parameter as a variable of the target calculation formula to obtain a second score; The task priority is determined according to the first score and the second score.

4. The method according to claim 3, characterized in that Determining a target calculation formula for the second item according to the scheduling scenario includes at least one of the following: When the scheduling scenario is a normal scheduling scenario, determining the target calculation formula to be a first calculation formula, a second calculation formula, a third calculation formula, or a fourth calculation formula, wherein the first calculation formula, the second calculation formula, the third calculation formula, or the fourth calculation formula are all related to resource scheduling, and any two of the first calculation formula, the second calculation formula, the third calculation formula, and the fourth calculation formula have different variables and / or different calculation functions; When the scheduling scenario is a storage optimization scenario, determining that the target calculation formula is a fifth calculation formula; the fifth calculation formula is related to data access storage; When the scheduling scenario is a fault handling scenario, the target calculation formula is determined to be a sixth calculation formula; the sixth calculation formula is related to fault handling.

5. The method according to any one of claims 1 to 4, characterized in that The task priority includes a first priority and a second priority; the first priority is higher than the second priority; The initiating of a fault handling plan according to the task priority includes: When the task priority is the first priority, a first record or a second record is generated according to the recoverability of the task; the first record includes a first disaster recovery link; the first disaster recovery link is used to migrate the task from the first device to the second device in a first migration mode; the first device and the second device are backup devices of the same type; the second record includes: compensation task scheduling of the first priority; When the task priority is the second priority, a third record or a fourth record is generated according to the recoverability of the task; the third record includes a second disaster recovery link; the second disaster recovery link is used to migrate the task from the first device to the second device in a second migration mode; the fourth record includes: compensation task scheduling of the second priority.

6. The method according to claim 5, characterized in that The task priority also includes a third priority and a fourth priority; the third priority is higher than the fourth priority, and the third priority is lower than the second priority.

7. The method according to claim 5, characterized in that The method further comprises: At least one of a business period, a device state, and an external event is detected, and the task priority is increased or decreased.

8. A device management system based on task priority evaluation, applied to a data warehousing system, characterized in that: The equipment management system includes: A first determining module, configured to determine a first device and a scheduling scenario according to a service tag of a task; A second determining module is configured to determine a task priority of the first device based at least on the business parameters and the scheduling scenario; the task priority is used at least for resource scheduling and fault handling of the task; An acquisition module, configured to predict, based on an artificial intelligence (AI) model, a failure occurrence condition when the first device processes the task, and obtain failure prediction information; A starting module, configured to start a fault handling plan according to the task priority when the fault prediction information indicates that a specified time period meets a specified condition; The processing module is used to manage and handle the fault of the first device according to the fault handling plan when a fault occurs.

9. The system according to claim 8, characterized in that The second determination module is specifically configured to determine an objective function for calculating the task priority according to the scheduling scenario; and calculate the task priority of the first device using the business parameter as a variable of the objective function.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions; after the computer-executable instructions are executed by a processor, the method according to any one of claims 1 to 7 can be implemented.

Citation Information

Patent Citations

  • Task mode selection and task execution method and device, equipment and storage medium

    CN109725996A

  • Disaster recovery method and device of cloud host and medium

    CN115426251A

  • Data disaster recovery method and device

    CN115718674A

  • Link adaptive fault tolerance method and device for storing multi-control cluster, and server

    CN116436839A

  • Transform-based task fault prediction method in cloud data center environment

    CN116471197A

Cited By

  • Database mirror image resource allocation method and device, electronic equipment and storage medium

    CN121166654A