Method and system for device management based on task priority assessment
By adopting an equipment management method based on task priority assessment, and using artificial intelligence models to predict faults and dynamically adjust priorities, the problems of crude task priority management and delayed fault response in data warehouse systems have been solved, thereby achieving efficient utilization of equipment resources and improved business continuity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2026-04-10
AI Technical Summary
In data warehouse systems, equipment management suffers from problems such as unreasonable task priority settings, leading to delayed fault handling and rigid scheduling strategies.
By using a device management method based on task priority assessment, the system uses artificial intelligence models to predict faults, determines the task priority of devices based on business tags and scheduling scenarios, and activates fault handling plans when the fault prediction information meets specified conditions, dynamically adjusting priorities to prioritize high-priority tasks.
It enables fine-grained allocation of task priorities, ensuring timely processing of high-priority tasks, reducing resource waste, improving the intelligent utilization of equipment resources, and enhancing business continuity and operational efficiency.
Smart Images

Figure CN120448075B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data, in particular to a device management method and system based on task priority evaluation. BACKGROUND
[0002] Many enterprises store data in data warehouses. Data warehouses can be provided by data storage enterprises to provide data storage and enterprise-level data decision-making to other enterprises. In a data storage system, device (storage device, computing device, network device) management faces problems such as unreasonable task priority setting, lagging fault handling, and rigid scheduling strategy. SUMMARY
[0003] Therefore, an embodiment of the present application provides a device management method and system based on task priority evaluation. The technical solution of the present application is implemented as follows.
[0004] The first aspect provides a device management method based on task priority evaluation, applied to a data storage system, and the method comprises the following steps.
[0005] A first device and a scheduling scenario are determined according to a business label of a task.
[0006] A task priority of the first device is determined according to at least a business parameter and the scheduling scenario; the task priority is used for at least resource scheduling and fault handling of the task.
[0007] Fault prediction of the first device is performed based on an artificial intelligence (AI) model to obtain fault prediction information.
[0008] When the fault prediction information indicates that a specified period meets a specified condition, a fault handling plan is started according to the task priority.
[0009] When a fault occurs, the first device is managed and the fault is handled according to the fault handling plan.
[0010] The second aspect provides a device management system based on task priority evaluation, applied to a data storage system, and the device management system comprises the following modules.
[0011] A first determination module is configured to determine a first device and a scheduling scenario according to a business label of a task.
[0012] A second determination module is configured to determine a task priority of the first device according to at least a business parameter and the scheduling scenario; the task priority is used for at least resource scheduling and fault handling of the task.
[0013] The acquisition module is configured to acquire failure prediction information by predicting, based on an artificial intelligence (AI) model, a failure occurrence condition of the first device when processing the task.
[0014] The starting module is configured to start a failure processing plan according to the task priority when the failure prediction information indicates that a specified period meets a specified condition.
[0015] The processing module is configured to perform management and failure processing of the first device according to the failure processing plan when a failure occurs.
[0016] The third aspect provides a computer-readable storage medium storing computer executable instructions; after the computer executable instructions are executed by a processor, the device management method based on task priority evaluation can be implemented.
[0017] The technical solution provided by the embodiments of the present disclosure allocates differentiated task priorities for device tasks according to business labels (and scheduling scenarios), realizes fine allocation of task priorities, ensures that high-priority key tasks are preferentially acquired resources and failure processing, and improves core business response efficiency. The AI model is used to predict the failure probability in advance, and the plan is started in advance according to the predicted probability, so as to ensure that the failure can be responded to in time when the failure occurs, and the waste of resources corresponding to the plan is reduced when the plan is not started. In short, the technical solution solves the problems of extensive task priority management and slow failure response in the data warehouse system by "priority-driven resource scheduling + AI prediction and active prevention + scenario-based strategy adaptation", realizes intelligent and efficient use of device resources, and significantly improves business continuity, data reliability and operation and maintenance efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which are included to provide a further understanding of the application, constitute a part of this application, illustrate the illustrative embodiments of the application and its description, and do not constitute an improper limitation of the application. In the drawings:
[0019] Figure 1 A flowchart of a device management method based on task priority evaluation provided by an embodiment of the present disclosure is shown in the figure.
[0020] Figure 2 A flowchart of another device management method based on task priority evaluation provided by an embodiment of the present disclosure is shown in the figure.
[0021] Figure 3 A structural diagram of a device management system based on task priority evaluation provided by an embodiment of the present disclosure is shown in the figure.
[0022] Figure 4 A structural diagram of an electronic device provided by an embodiment of the present disclosure is shown in the figure. DETAILED DESCRIPTION
[0023] In order to make the personnel in the technical field better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person of ordinary skill in the art without creative labor should belong to the protection scope of the present application.
[0024] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] As shown in Figure 1 The embodiments of the present disclosure provide a device management method based on task priority evaluation, applied to a data warehouse system, and the method comprises the following steps:
[0026] S1110: determining a first device and a scheduling scene according to a business label of a task; the first device comprises at least one of a storage device, a computing device and a network device;
[0027] S1120: determining a task priority of the first device according to at least a business parameter and the scheduling scene; the task priority is used for resource scheduling and fault handling of the task;
[0028] S1130: predicting a fault occurrence condition when the first device handles the task based on an artificial intelligence AI model, and obtaining fault prediction information;
[0029] S1140: starting a fault handling plan according to the task priority when the fault prediction information indicates that a specified period meets a specified condition;
[0030] S1150: managing and handling the fault of the first device according to the fault handling plan when the fault occurs.
[0031] In some embodiments, the task priority evaluation based device management method can be used in a management system of a data warehouse system. The management system includes one or more electronic devices. Illustratively, the task priority evaluation based device management method can be used in the one or more electronic devices. In other embodiments, the management system can also be a subsystem integrated in the data warehouse system.
[0032] The devices used by the data warehouse system are pre-labeled with semantic labels to achieve semantic grouping of the devices. Illustratively, the data warehouse devices (such as computing nodes, storage clusters, ETL servers) are labeled according to business functions (such as “real-time analysis cluster”, “historical data archiving storage”, “data cleaning node”), and the priority scheduling system automatically matches the device groups with corresponding labels according to the task type, and preferentially calls the idle resources of the device groups to which high-priority tasks belong. In this way, when determining the devices required by a task, the first device can be quickly found based on such semantic grouping, i.e., a business semantic layer is introduced, the matching path of the task and the device is shortened, and the scheduling efficiency is improved.
[0033] The first device is any device in the data warehouse system. According to the type of the device, the first device includes at least one of a storage device, a computing device, and a network device. According to the degree of coldness and hotness of the data processed by the device, the device can be divided into a cold data device and a hot data device. The data processed by the cold data device has a lower degree of hotness (access frequency) than the data processed by the hot data device.
[0034] In some embodiments, the scheduling scenario can be used to describe the current state of the data warehouse system when the resources of the task are scheduled, but is not limited to describing the current state of the data warehouse system. In this way, in some embodiments, the scheduling scenario can be determined according to the current state of the operation of the data warehouse system.
[0035] In some embodiments, the scheduling scenarios can be divided into daily scheduling scenarios and non-daily scheduling scenarios. The daily scheduling scenarios can be further divided into different daily scheduling scenarios according to the proportion of tasks currently being executed, for example, if the proportion of storage tasks is the highest, the daily scheduling scenario in the current state is a storage intensive scenario. For another example, if the proportion of computing tasks is the highest, the daily scheduling scenario in the current state is a computing intensive scenario. For another example, if the proportion of access tasks is the highest, the daily scheduling scenario in the current state can be an access intensive scenario. The non-daily scheduling scenario can also be referred to as an abnormal scheduling scenario. Typical abnormal scheduling scenarios include, but are not limited to, fault handling scenarios, and storage optimization scenarios for avoiding storage overload. The storage optimization scenario is different from the storage intensive scenario, and the storage intensive scenario is a data storage scenario triggered by user storage demand. The storage optimization scenario is a data migration, compression or redundancy deletion related storage scenario initiated by the storage demand of the data warehouse system. The above are only examples of scheduling scenarios, and the specific implementation is not limited to the examples.
[0036] In some embodiments, the business parameters can be used to describe the business corresponding to the task, including but not limited to at least one of the following:
[0037] Business impact, for example, the business impact is used to indicate whether the corresponding task affects the generation, update, deletion, etc. of the core report;
[0038] Business value coefficient, for example, the business value coefficient is used to indicate the importance of the business, for example, according to the importance, it can be divided into core business, waist business, edge business, etc.;
[0039] Business emergency degree, for example, the business emergency degree is used to indicate the maximum response delay or average response delay allowed by the task corresponding to the business;
[0040] Failure loss index, for example, the failure loss index is used to indicate the punishment effect, etc. caused by the failure of the task corresponding to the business.
[0041] In the embodiments of the present disclosure, the task priority corresponding to the task is determined in combination with the business parameters and the scheduling scenario.
[0042] The task priority is used for resource scheduling of the task on one hand and fault processing of the task on the other hand, so that the resource scheduling and the fault processing share the task priority, the priority of the resource scheduling and the fault processing is synchronized, and the calculation amount caused by priority calculation can be reduced relative to setting two priorities of the resource scheduling and the fault processing respectively. Exemplarily, the task priority is related to the heat of data, a device processing cold data is referred to as a cold data device, and a device processing hot data is referred to as a hot data device. According to the storage characteristics of "hot data" (high-frequency access) and "cold data" (low-frequency archiving) in a data warehouse, a device priority rule is designed: when processing tasks such as query and real-time synchronization of hot data, an SSD storage device and a high-frequency computing node are preferentially called; when processing batch migration and backup tasks of cold data, a mechanical hard disk and a low-power device are used, and the resources are allowed to be preempted by high-priority tasks. Combined with the concept of data life cycle management (DLM), the data characteristics and the device priority are bound, and the resource utilization efficiency is optimized.
[0043] In some embodiments, S1120 can include:
[0044] According to the scheduling scenario, a target function for calculating the task priority is determined.
[0045] A business parameter is taken as a variable of the target function, and the task priority of the first device is calculated.
[0046] In the embodiments of the present disclosure, different scheduling scenarios correspond to different target functions for calculating the task priority, so that the task priority of the task is not only related to the business of the task but also related to the current scheduling scenario, that is, the business associated with the task and the current status of the data warehouse system are comprehensively considered.
[0047] In some embodiments, the target function includes a first term and a second term; according to the scheduling scenario, the target function for calculating the task priority is determined, including:
[0048] According to the bloodline level depth and the number of downstream influence nodes associated with the task, a first score of the first term is determined;
[0049] According to the scheduling scenario, a target calculation formula of the second term is determined;
[0050] A business parameter is taken as a variable of the target calculation formula to obtain a second score;
[0051] According to the first score and the second score, the task priority is determined.
[0052] In some embodiments, , or wherein, is the first score. This refers to the task's hierarchy within the data lineage. For example, raw data acquisition is layer 1, intermediate processing is layer 2, and report generation is layer 3. The higher the layer, the closer it is to the business endpoint, and the higher the task's priority.
[0053] Number of downstream affected nodes: The number of downstream tasks / reports affected by this task. For example, affecting 10 reports = 10 points, affecting only 1 report = 1 point. and Weighting coefficients can be pre-set based on the importance of hierarchy and the number of downstream influential nodes in the data lineage. These weighting coefficients can be generated by the AI model through big data collection and processing.
[0054] It is evident that the complexity of a task is related to the number of downstream nodes associated with the first device. The data involved in the task will affect the hierarchy in the data lineage.
[0055] For example, when the raw data acquisition task fails, although it is at the bottom of the hierarchy (level = 1), it may affect all downstream levels, with a priority of 1 × N (where N is the number of downstream nodes). If an intermediate layer task fails, its priority is 2 × M (where M is the number of downstream nodes), typically higher than the bottom layer but lower than the terminal task. In this way, task priority can reflect the potential chain reaction of a failure through this data hierarchy, facilitating subsequent determination of whether a fault contingency plan needs to be set up, or what type of backup plan to implement, based on task priority.
[0056] In some embodiments, the target calculation formula for the second term is determined based on the scheduling scenario, including at least one of the following:
[0057] When the scheduling scenario is a regular scheduling scenario, the target calculation formula is determined to be the first calculation formula, the second calculation formula, the third calculation formula, or the fourth calculation formula. Among them, the first calculation formula, the second calculation formula, the third calculation formula, or the fourth calculation formula are all related to resource scheduling, and any two of the first calculation formula, the second calculation formula, the third calculation formula, and the fourth calculation formula have different variables and / or different calculation functions.
[0058] When the scheduling scenario is a storage optimization scenario, the target computation formula is determined to be the fifth computation formula; the fifth computation formula is related to data access storage.
[0059] When the scheduling scenario is a fault handling scenario, the target calculation formula is determined to be the sixth calculation formula; the sixth calculation formula is related to fault handling.
[0060] In some embodiments, the first calculation may be related to business impact, data sensitivity, time window flexibility, and resource starvation risk.
[0061] ,
[0062] Business Impact. For example, whether the task is related to a business-critical process. Data Sensitivity. For example, whether the task involves private data. Time Window Elasticity. For example, whether the task allows for delayed execution. Resource Starvation Risk. For example, whether long waiting time leads to device idling. Thus, using the first calculation formula, the second score is calculated as The second score calculated using the first calculation formula takes into account at least four dimensions, as opposed to traditional task priority which is based on a single indicator (e.g., time urgency). This framework introduces business and data characteristics dimensions for data warehouse scenarios, which are more aligned with industry needs. , , , The weights can be dynamically adjusted.
[0063] In some embodiments, the second calculation formula can be related to a business value coefficient, a failure weight, and a resource consumption coefficient of the task.
[0064] Exemplarily, or wherein, is the business value coefficient, is the failure weight. is the resource consumption coefficient. The second score calculated using the second calculation formula. Table 1 can provide a description of the business value coefficient, the failure weight, and the resource consumption coefficient.
[0065] Table 1
[0066]
[0067] Exemplarily, for task A: business value 8 points, time urgency 10 points, resource consumption 3 points, then task priority = (8x10) / 3 ≈ 26.67. For task B: business value 5 points, time urgency 4 points, resource consumption 8 points, then task priority = (5x4) / 8 = 2.5. According to the calculated task priority, task A is executed in priority over task B, as it has high business value, strong time urgency, and low resource consumption.
[0068] In some embodiments, the second calculation formula can be related to one or more of business urgency, data consistency, calculation complexity, compliance requirement, and resource preemption risk.
[0069] In some embodiments, wherein, is a second score calculated using a third calculation formula. is the i-th calculation variable. is the weight corresponding to the i-th calculation variable. is the total number of calculation variables. Table 2 is an example of calculation variables of the third calculation formula.
[0070] Table 2
[0071]
[0072] In some embodiments, the fourth calculation formula is related to task urgency, data locality coefficient, predicted execution time, and queue waiting time.
[0073] Exemplarily, or wherein, is a second score calculated using a fourth calculation formula. is the task urgency. is the data locality coefficient. is the predicted execution time. is the queue waiting time. and may be a preset weight.
[0074] Data locality coefficient: measures whether the task is executed locally at the data storage node (local = 1.5, remote = 1), optimizes the efficiency of "moving computation to data" in the data warehouse.
[0075] Predicted execution time: predicts the task duration through historical data, for example, in minutes, normalized to 0.1-10.
[0076] Queue waiting time: the length of time the task has been waiting in the queue, for example, to avoid starvation, the longer the weight is higher.
[0077] Since there are alternatives of the first calculation formula to the fourth calculation formula in the daily scheduling scenario, the target calculation formula suitable for the working mechanism of the data warehouse system can be selected according to the working mechanism of the data warehouse system. For example, the fourth calculation formula can be suitable for the daily scenario of the data warehouse system using queue scheduling tasks.
[0078] In some embodiments, for the storage optimization scenario, the fifth calculation is related to parameters such as data access frequency, data capacity, and update frequency. Because in the storage optimization scenario, both current data access and data optimization need to be considered.
[0079] Exemplarily, , wherein, is a second score calculated using a fifth calculation formula. is the data access frequency. Data Volume. Data Update Frequency. 、 、 Weight.
[0080] Access Hotness: Query / Write times per unit time (e.g. normalized to 0-10 points, 100 times / day = 10 points, 10 times / day = 3 points).
[0081] Data Volume: Data size (e.g. in GB, 100 GB = 10 points, 1 GB = 1 point).
[0082] Update Frequency: Number of updates per day (e.g. real-time updates = 10 points, weekly updates = 2 points).
[0083] In some embodiments, User Profile Data: Access Hotness 8 points, Volume 6 points, Update Frequency 7 points → Priority = 0.5x8 + 0.3x6 + 0.2x7 = 4 + 1.8 + 1.4 = 7.2. Historical Log Data: Access Hotness 2 points, Volume 10 points, Update Frequency 1 point → Priority = 0.5x2 + 0.3x10 + 0.2x1 = 1 + 3 + 0.2 = 4.2. In this case, when performing storage resource scheduling, the following storage strategy can be adopted according to the task priority: User Profile Data is preferentially stored in a high-performance storage layer (e.g. SSD), and Historical Log is archived to a low-cost storage (e.g. HDD or cloud object storage).
[0084] In some embodiments, the fault handling scenario can adopt a sixth calculation formula, for example, or wherein, is a second score calculated by using the sixth calculation formula. is a compliance risk. is a fault impact coefficient. is a recovery complexity index. Table 3 is a variable related description of the fifth calculation formula.
[0085] Table 3
[0086]
[0087] For example, task C: compliance risk 8 points, failure impact 10 points, recovery complexity 9 points -> priority = 8x10 + 9 = 89. Task D: compliance risk 3 points, failure impact 4 points, recovery complexity 4 points -> priority = 3x4 + 4 = 16. Thus, when scheduling resources, task C (such as real-time synchronization of core transaction data) needs to be immediately scheduled for processing, and task D (such as failure of non-critical report generation) can be delayed.
[0088] In some embodiments, calculating the task priority according to the target function can be when the task is received or before starting the task execution.
[0089] In some embodiments, the method further comprises: detecting at least one of a business period, a device state, and an external event, and increasing or decreasing the task priority.
[0090] That is, after the initial task priority calculation according to the target function, the dynamic priority adjustment according to the dynamic priority mechanism is also made, so that the currently calculated task priority conforms to the dynamic changes of the data warehouse system.
[0091] For example, based on the business period: for example, the priority of the night batch loading task (P2) can be temporarily increased to P1, and decreased to P1 during the day.
[0092] For another example, based on the device state: when a certain storage array approaches the capacity threshold, the I / O priority of the non-critical task (such as P3 archiving) is automatically reduced to ensure P1 task writing.
[0093] For another example, based on the external event: when a network attack is detected, the traffic cleaning task of the security protection device (such as a firewall) is automatically increased to P1 to block the attack traffic.
[0094] Dynamic priority adjustment mechanism: priority change rules in the task execution process are designed, for example: when a high-priority task is triggered, a low-priority task can automatically enter a "suspend-cache" state (retaining the current progress, releasing part of the resources), instead of being directly terminated, avoiding data processing interruption; based on real-time data flow fluctuation (such as sudden query leading to IO load surge), the priority of the associated task is dynamically increased to ensure business continuity. The priority of the traditional scheduling system is usually statically set, and this mechanism increases the flexibility during running and reduces resource waste.
[0095] In some embodiments, as shown in Figure 2 S1130 can include:
[0096] S1131: Determine key monitoring indicators, collect at least one of the following dimensions of data for data storage devices (storage / computing / network devices) and task processing processes: backup running data, task business data, external environment data.
[0097] Device running data includes but is not limited to at least one of the following:
[0098] IOPS, throughput, delay, disk error rate (SMART indicators), remaining capacity of storage devices;
[0099] CPU utilization, memory usage, process state, task queue length of computing devices;
[0100] Bandwidth utilization, packet loss rate, port error number, routing delay of network devices.
[0101] Task business data includes but is not limited to at least one of the following:
[0102] Task priority label (such as P1-P4), business label (such as real-time analysis, batch loading);
[0103] Task execution parameters: data volume, processing time, input / output path, dependency relationship (such as data blood relationship level);
[0104] Task result data: success rate, failure type (such as timeout, data format error, resource shortage).
[0105] External environment data includes but is not limited to at least one of the following:
[0106] Business period (such as peak / valley period), system load (such as overall utilization of cluster), external event (such as network attack, power failure warning).
[0107] S1132: Data collection.
[0108] For example, through Prometheus, Zabbix and other monitoring tools, device indicators are captured at a frequency of seconds; task metadata is obtained by using the API of data storage platform (such as Hadoop, Spark).
[0109] S1133: Data cleaning and feature engineering, which can specifically include: missing value processing: for abnormally missing data (such as missing indicators caused by network interruption), interpolation method (such as linear interpolation, time series prediction) or deletion of invalid samples. Feature extraction can include at least one of the following:
[0110] Time series features: calculate sliding window statistics of device indicators (such as average IO delay in the past 5 minutes, CPU utilization peak);
[0111] Correlation features: relevance of task priority to device load (e.g., whether storage device latency significantly increases when high-priority tasks are executed);
[0112] Derived features: device health index (generated by combining multiple indicators, e.g., storage device health index = 0.6 x IO latency + 0.3 x remaining capacity + 0.1 x error rate).
[0113] Data normalization: scaling numerical features to the [0, 1] interval (e.g., Min-Max normalization) to facilitate model training.
[0114] Through data cleaning and feature engineering, the authenticity and effectiveness of the data used for fault prediction can be ensured.
[0115] S1134: Fault warning of AI model.
[0116] In some cases, the data warehouse system can have multiple fault prediction models.
[0117] Exemplarily, model type selection can include: according to the fault prediction target (such as classifying fault types, regressing fault probabilities), the following models are selected:
[0118] Time series prediction model:
[0119] LSTM / GRU: suitable for single-device fault trend prediction (e.g., predicting hard disk life based on historical disk error rate);
[0120] Transformer: handles long sequence dependencies and captures associated faults between multiple devices (e.g., network device failure leading to task failure of computing nodes).
[0121] Exemplarily, the fault prediction AI model can be a classification model, which can include but is not limited to at least one of the following:
[0122] Random Forest (RF): used for multi-feature fusion fault type classification (e.g., distinguishing between hardware and software faults);
[0123] XGBoost / LightGBM: efficiently handles large-scale heterogeneous data and identifies key fault influencing factors (e.g., correlation between task data volume and storage device failure).
[0124] Exemplarily, the fault prediction AI model can be an anomaly detection model, which can exemplarily include but is not limited to: Isolation Forest: detects outliers in device indicators (e.g., sudden IO latency spikes); Autoencoder: identifies device state anomalies through reconstruction error (e.g., computing node memory usage pattern deviates from normal range).
[0125] Deploy the trained model as an API service (e.g. via Flask, TensorFlow Serving) to access the monitoring platform of the data warehouse system.
[0126] In some embodiments, special prediction models are trained for different scheduling scenarios (e.g. storage optimization, failure handling):
[0127] Storage optimization scenario: focus on predicting the capacity bottleneck and IO performance degradation of storage devices;
[0128] Failure handling scenario: optimize the prediction ability of cascading failures (e.g. one switch failure causes multiple server task failures). In this case, the AI model for failure prediction is selected according to the scheduling scenario, so that the failure prediction of the AI model is more accurate.
[0129] In some embodiments, the prediction frequency of the AI model can include but is not limited to at least one of the following:
[0130] High priority devices (e.g. primary storage cluster): perform prediction every second;
[0131] Low priority devices (e.g. archival storage nodes): perform prediction every minute.
[0132] In some embodiments, the failure prediction information can include various related information describing the possible failure of the first device when processing the task. For example, the failure prediction information can include but is not limited to at least one of the following: the probability of failure, the type of possible failure, the severity of possible failure, the time of expected failure, the possible cause of failure, etc.
[0133] In some embodiments, the AI model provides failure warning.
[0134] Failure warning based on dynamic threshold strategy, for example, set threshold based on historical data quantile (e.g. trigger warning when device health index > 0.8); adjust threshold combined with task priority: higher priority task corresponds to lower failure probability threshold for the device (e.g. P1 device threshold set to 10%, P4 device set to 30%).
[0135] Warning classification can include but is not limited to at least one of the following:
[0136] Yellow warning: failure probability exceeds threshold but is lower than emergency value, trigger preventive maintenance (e.g. increase spare parts inventory, schedule idle resources);
[0137] Red warning: failure probability ≥ emergency value (e.g. > 50%), immediately start disaster recovery plan (e.g. hot migration tasks to backup devices).
[0138] In some embodiments, the task priority comprises a first priority and a second priority; the first priority is higher than the second priority. S1140 can comprise at least one of: when the task priority is the first priority, generating a first backup record or a second backup record according to the recoverability of the task; the first backup record comprises a first disaster recovery link; the first disaster recovery link is used for migrating the task from the first device to the second device in a first migration manner; the first device and the second device are backup devices of the same type; the second backup record comprises a compensation task schedule of the first priority;
[0139] When the task priority is the second priority, generating a third backup record or a fourth backup record according to the recoverability of the task; the third backup record comprises a second disaster recovery link; the second disaster recovery link is used for migrating the task from the first device to the second device in a second migration manner; the fourth backup record comprises a compensation task schedule of the second priority.
[0140] For example, when the task priority is the first priority, and the task is recoverable, the first backup record is generated, otherwise the second backup record can be generated. When the task priority is the second priority, and the task is recoverable, the third backup record is generated, otherwise the fourth backup record can be generated.
[0141] The first migration manner and the second migration manner can be different types of migration manners. For example, the migration rate of the first migration manner is higher than that of the second migration manner. The user perception degree of the first migration manner is lower than that of the second migration manner.
[0142] In some embodiments, the first migration manner can be hot migration and the second migration manner is cold migration. Alternatively, the first migration manner is jump migration and the second migration manner is step-by-step migration. The specific descriptions of the four migration manners are shown in Table 4 below.
[0143] Table 4
[0144]
[0145] In some embodiments, the first migration manner and the second migration manner are combinations of multiple migration manners.
[0146] For example, the first migration manner can be: hot migration + step-by-step migration: for example, in a micro-service architecture, first perform hot migration on a single service (keep the business running), and then perform step-by-step migration on other modules according to service dependency.
[0147] For example, the second migration manner can be: cold migration + jump migration: first complete data backup or system building through cold migration, and then switch traffic in one time through jump migration (such as first deploying in a new environment, and then switching DNS).
[0148] For another example, the first migration mode can be: hot migration is adopted for core services (guarantee continuity), and cold migration + step-by-step migration is adopted for non-core data (reduce cost and risk). The second migration mode is: cold migration + step-by-step migration is adopted for all service data.
[0149] In some embodiments, the first disaster recovery link and the second disaster recovery link can each include information of the second device, such as an address of the second device. In some embodiments, the first disaster recovery link has a smaller number of hops than the second disaster recovery link, so as to realize fast disaster recovery for tasks of the first priority.
[0150] In some embodiments, the failure prediction information indicates that a specified period meets a specified condition, and a failure handling plan is started according to the priority of the task. For example, the threshold can be pre-set, such as 50%.
[0151] In some embodiments, the failure prediction information indicating that a specified period meets a specified condition includes but is not limited to at least one of the following:
[0152] The failure prediction information indicates that the failure in the specified period is higher than a first probability threshold;
[0153] The failure prediction information indicates that a second probability threshold condition occurs in the specified period.
[0154] In some embodiments, the first probability threshold of the same task has a higher priority than the second probability threshold.
[0155] In some embodiments, the malignant failure includes but is not limited to the failure corresponding to the aforementioned red alert.
[0156] In some embodiments, different task priorities can correspond to different first probability thresholds and / or first probability thresholds, so as to realize early start of the failure handling plan corresponding to the task priority. For example, the threshold corresponding to the first priority is smaller than the first probability threshold and / or the first probability threshold corresponding to the second priority. In some embodiments, the task priority further includes a third priority and a fourth priority; the third priority is higher than the fourth priority, and the third priority is lower than the second priority. At this time, the first probability threshold and / or the first probability threshold corresponding to the third priority is higher than the first probability threshold and / or the first probability threshold of the second priority. The first probability threshold and / or the first probability threshold corresponding to the fourth priority is higher than the first probability threshold and / or the first probability threshold of the third priority.
[0157] In some embodiments, the fourth priority can not start the failure handling plan in advance. Alternatively, neither the third priority nor the fourth priority can start the failure handling plan in advance.
[0158] In some embodiments, the priority-aware failover mechanism: when the device carrying high-priority tasks fails, the system automatically triggers the "fast disaster recovery link". The redundant node is preferentially called to take over the task from the same priority device group (instead of switching step by step according to the conventional failover process);
[0159] For tasks that cannot be quickly recovered, automatically generate fault compensation tasks (such as raising the priority to the highest when rescheduling), to ensure that the business impact is minimized. The traditional disaster recovery mechanism does not distinguish between task priorities. This design realizes "priority-sensitive" fault handling, reducing the interruption loss of high-value tasks.
[0160] When a fault occurs, the first device is managed and fault handled according to the fault handling plan. If the fault handling plan is not started, the fault is handled conventionally. For example, when the first device fails for a task of the fourth priority, the fault handling plan may not be started at present, and the conventional fault handling can be performed.
[0161] When fault handling is performed, it can be based on a work order. For example, automatic work orders are generated for different priority tasks: P1 corresponds to a fault: trigger an SMS / phone alarm within 10 seconds, automatically assign a senior engineer, and arrive at the scene within 30 minutes; P2 corresponds to a fault: respond within 1 hour and repair within 4 hours; P3 / P4 corresponds to a fault: handle according to the work calendar and close the loop within 24 hours.
[0162] As shown in Figure 3 The embodiments of the present disclosure provide a device management system based on task priority evaluation, applied to a data warehouse system, and the device management system comprises:
[0163] A first determination module 3101 is configured to determine a first device and a scheduling scenario according to a business tag of a task;
[0164] A second determination module 3102 is configured to determine a task priority of the first device according to at least a business parameter and the scheduling scenario; the task priority is used for resource scheduling and fault handling of the task;
[0165] An acquisition module 3103 is configured to predict a fault occurrence condition of the first device when handling the task based on an artificial intelligence (AI) model, and obtain fault prediction information;
[0166] A starting module 3104 is configured to start a fault handling plan according to the task priority when the fault prediction information indicates that a specified period meets a specified condition.
[0167] A processing module 3105 is configured to manage and handle a fault of the first device according to the fault handling plan when the fault occurs.
[0168] In some embodiments, the second determining module is specifically configured to determine a target function for calculating the task priority according to the scheduling scenario; and take the service parameter as a variable of the target function to calculate the task priority of the first device.
[0169] In some embodiments, the target function includes a first term and a second term; the second determining module is further configured to determine a first score of the first term according to a blood relationship level depth and a number of downstream influence nodes associated with the task; determine a target calculation formula of the second term according to the scheduling scenario; take the service parameter as a variable of the target calculation formula to obtain a second score; and determine the task priority according to the first score and the second score.
[0170] In some embodiments, the second determining module is specifically configured to perform at least one of the following: when the scheduling scenario is a regular scheduling scenario, determine that the target calculation formula is a first calculation formula, a second calculation formula, a third calculation formula or a fourth calculation formula, wherein the first calculation formula, the second calculation formula, the third calculation formula or the fourth calculation formula are all related to resource scheduling, and variables and / or calculation functions of any two of the first calculation formula, the second calculation formula, the third calculation formula and the fourth calculation formula are different; when the scheduling scenario is a storage optimization scenario, determine that the target calculation formula is a fifth calculation formula; the fifth calculation formula is related to data access storage; when the scheduling scenario is a fault handling scenario, determine that the target calculation formula is a sixth calculation formula; the sixth calculation formula is related to fault handling.
[0171] In some embodiments, the task priority includes a first priority and a second priority; the first priority is higher than the second priority; the processing module is specifically configured to generate a first backup record or a second backup record according to the recoverability of the task when the task priority is the first priority; the first backup record includes a first disaster recovery link; the first disaster recovery link is used to migrate the task from the first device to the second device in a first migration manner; the first device and the second device are backup devices of the same type; the second backup record includes a compensation task scheduling of the first priority; generate a third backup record or a fourth backup record according to the recoverability of the task when the task priority is the second priority; the third backup record includes a second disaster recovery link; the second disaster recovery link is used to migrate the task from the first device to the second device in a second migration manner; the fourth backup record includes a compensation task scheduling of the second priority.
[0172] In some embodiments, the task priority further includes a third priority and a fourth priority; the task priority further includes the third priority and the fourth priority; the third priority is higher than the fourth priority, and the third priority is lower than the second priority.
[0173] In some embodiments, the second determining module is specifically configured to determine a target function for calculating the task priority according to the scheduling scenario; and take the service parameter as a variable of the target function to calculate the task priority of the first device.
[0174] In some embodiments, the second determining module is further configured to increase or decrease the task priority according to at least one of a service period, a device state, and an external event.
[0175] Embodiments of the present disclosure provide a computer readable storage medium, which stores computer executable instructions; after the computer executable instructions are executed by a processor, a task priority evaluation based device management method provided by any of the foregoing technical solutions can be implemented.
[0176] In combination with Figure 4 As shown in the accompanying drawings, the embodiments of the present application provide an electronic device, which can be a component device of the task priority evaluation based device management system, including a processor 10 and a memory 11. Optionally, the apparatus can further include a communication interface 12 and a bus 9. The processor 10, the communication interface 12, and the memory 11 can communicate with each other through the bus 9. The communication interface 12 can be used for information transmission. The processor 10 can invoke the logic instructions in the memory 11 to execute the queue based voiceprint data processing method of the above-mentioned embodiments.
[0177] In addition, the logic instructions in the memory 11 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium.
[0178] The memory 11 as a computer readable storage medium can be used to store software programs, computer executable programs, such as program instructions / modules corresponding to the method in the embodiments of the present application. The processor 10 executes the function application and data processing by running the program instructions / modules stored in the memory 11, that is, implements the task priority evaluation based device management method in the above-mentioned embodiments.
[0179] The memory 11 can include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required by a function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 11 can include a high-speed random access memory and can also include a non-volatile memory.
[0180] The technical solutions of the embodiments of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiments of the present application. The foregoing storage medium can be a non-transitory storage medium, including a U disk, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, etc. a variety of media that can store program codes, or a transitory storage medium.
[0181] In the above-described embodiments of the present application, the description of each embodiment focuses on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0182] The disclosed embodiments or examples are not exhaustive, and only illustrate some embodiments or examples, and are not specific limitations on the protection scope of the present disclosure. In the case of no contradiction, each step in a certain embodiment or example can be implemented as an independent embodiment, and the steps can be combined arbitrarily, for example, the scheme after removing some steps in a certain embodiment or example can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment or example can be exchanged arbitrarily, in addition, the optional ways or optional examples in a certain embodiment or example can be combined arbitrarily; in addition, the embodiments or examples can be combined arbitrarily, for example, the steps of different embodiments or examples can be combined arbitrarily, a certain embodiment or example can be combined with the optional ways or optional examples of other embodiments or examples.
[0183] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between units or modules, which can be electrical or other forms.
[0184] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.
[0185] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.
[0186] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0187] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.
Claims
1. A device management method based on task priority evaluation, applied to a data warehouse system, characterized in that, The method comprises: The first device performing the task is determined according to a business label of the task, and a scheduling scene is determined, and a target function for calculating the task priority is determined according to the scheduling scene; the target function includes a first term and a second term; ; A first score of the first term is a first score of the first term; A level of the task in a data blood relationship is 1, a level of intermediate layer processing is 2, and a level of report generation is 3; And The importance of the level and the number of downstream influence nodes in the data blood relationship is pre-set weight coefficients, respectively. Before starting the task execution, a task priority of a first device is determined according to a business parameter and the scheduling scene and using a target function, comprising: determining a first score of the first item according to a blood relationship level depth and a downstream influence node number associated with the task; determining a second score using a target calculation formula of the second item determined according to the scheduling scene; and determining the task priority according to the first score and the second score; wherein, determining the second score using the target calculation formula of the second item determined according to the scheduling scene comprises: when the scheduling scene is a regular scheduling scene, determining that the target calculation formula is a fourth calculation formula; the fourth calculation formula is or ;, the second score calculated using the fourth calculation formula; is a task urgency degree; is a data locality coefficient; is a predicted execution time; is a queue waiting time; and may be a preset weight; When the scheduling scenario is the storage optimization scenario, it is determined that the target calculation formula is a fifth calculation formula; the fifth calculation formula is: ; is a second score calculated by using the fifth calculation formula; is a data access frequency; is a data capacity; is a data update frequency; , , may be a weight; When the scheduling scenario is a fault handling scenario, it is determined that the target calculation formula is a sixth calculation formula; the sixth calculation formula is related to fault handling; the sixth calculation formula is or ; a second score calculated by using the sixth calculation formula; a compliance risk; a fault impact coefficient; a recovery complexity index; obtaining fault prediction information based on an artificial intelligence (AI) model predicting a fault occurrence condition when the first device processes the task; when the fault prediction information indicates that a specified period meets a specified condition, starting a fault handling plan in advance according to the task priority; when the fault prediction information indicates that a specified period does not meet a specified condition, not starting a fault handling plan in advance; the fault prediction information indicating that a specified period meets a specified condition comprises at least one of the following: the fault prediction information indicating that a fault occurrence probability in the specified period is higher than a first probability threshold; the fault prediction information indicating that a specified time has a second probability threshold condition of malignant fault; the first probability threshold is higher than the second probability threshold; the first probability threshold and / or the second probability threshold of tasks of different priorities are different; when a fault occurs and a fault handling plan is started in advance, managing and handling the fault of the first device according to the fault handling plan.
2. The method of claim 1, wherein, The task priority comprises a first priority and a second priority; the first priority is higher than the second priority; starting a fault handling plan in advance according to the task priority comprises: when the task priority is the first priority, generating a first backup or a second backup according to the recoverability of the task; the first backup comprises a first disaster recovery link; the first disaster recovery link is used to migrate the task from the first device to a second device in a first migration manner; the first device and the second device are backup devices of the same type; the second backup comprises a compensation task schedule of the first priority; when the task priority is the second priority, generating a third backup or a fourth backup according to the recoverability of the task; the third backup comprises a second disaster recovery link; the second disaster recovery link is used to migrate the task from the first device to a second device in a second migration manner; the fourth backup comprises a compensation task schedule of the second priority.
3. The method of claim 2, wherein, The task priority further comprises a third priority and a fourth priority; the third priority is higher than the fourth priority, and the third priority is lower than the second priority; when the task priority is the third priority or the fourth priority, the fault handling plan is not started in advance.
4. The method of claim 3, wherein, The method further comprises: detecting at least one of a business period, a device state, and an external event, and adjusting the task priority up or down.
5. A device management system based on task priority evaluation, applied to a data warehouse system, characterized in that, The device management system comprises: The first determining module is configured to determine a first device for executing a task and a scheduling scenario according to a business tag of the task, and determine the target function for calculating the priority of the task according to the scheduling scenario; the target function comprises a first term and a second term; ; is a first score of the first term; is a level of the task in a data blood relationship; wherein a level of original data collection is 1, a level of intermediate layer processing is 2, and a level of report generation is 3; and are weight coefficients pre-set according to the level in the data blood relationship and the importance of the number of downstream influence nodes, respectively. The second determining module is configured to determine the task priority of the first device according to the business parameter and the scheduling scenario and using a target function before starting the task execution, including: determining a first score of the first item according to the blood relationship level depth and the number of downstream influence nodes associated with the task; determining a second score using a target calculation formula of the second item determined according to the scheduling scenario; and determining the task priority according to the first score and the second score; wherein determining the second score using the target calculation formula of the second item determined according to the scheduling scenario includes: when the scheduling scenario is a regular scheduling scenario, determining that the target calculation formula is a fourth calculation formula; the fourth calculation formula is or ;, the second score calculated using the fourth calculation formula; is a task urgency degree; is a data locality coefficient; is a predicted execution time; is a queue waiting time; and may be a preset weight; When the scheduling scenario is the storage optimization scenario, it is determined that the target calculation formula is a fifth calculation formula; the fifth calculation formula is: ; is a second score calculated by using the fifth calculation formula; is a data access frequency; is a data capacity; is a data update frequency; 、 、 may be a weight; When the scheduling scenario is the fault handling scenario, it is determined that the target calculation formula is a sixth calculation formula; the sixth calculation formula is related to fault handling; the sixth calculation formula is or ; a second score calculated by using the sixth calculation formula; a compliance risk; a fault impact coefficient; a recovery complexity index; an acquisition module configured to obtain fault prediction information based on an artificial intelligence (AI) model predicting a fault occurrence condition when the first device processes the task; The starting module is configured to start the fault handling plan in advance according to the task priority when the fault prediction information indicates that a specified period meets a specified condition, and not to start the fault handling plan in advance when the fault prediction information indicates that the specified period does not meet the specified condition. The fault prediction information indicating that the specified period meets the specified condition includes at least one of the following: the fault prediction information indicating that a fault occurrence probability in the specified period is higher than a first probability threshold; the fault prediction information indicating that a second probability threshold case occurs a malignant fault in the specified period; the first probability threshold is higher than the second probability threshold; and the first probability threshold and / or the second probability threshold of different priority tasks are different. The processing module is configured to perform management and fault handling of the first device according to the fault handling plan when a fault occurs and the fault handling plan is started in advance.
6. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions; and the computer executable instructions, when executed by a processor, implement the method in any one of claims 1 to 4.
Citation Information
Patent Citations
Task mode selection and task execution method and device, equipment and storage medium
CN109725996A
Transform-based task fault prediction method in cloud data center environment
CN116471197A
Multi-cluster distributed training-oriented disaster recovery drill and performance evaluation method and system
CN120104453A
Intelligent operation and maintenance method fusing multi-modal data and active learning
CN120198106A