Task control method and device, equipment and medium
The method enhances task control in distributed systems by collecting event data, calculating evaluation metrics, and using a rules engine to ensure precise and timely management of tasks, addressing the challenges of inaccurate and delayed control in existing systems.
Patent Information
- Application Number
- CN202510460088.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-15
AI Technical Summary
Existing task control methods are difficult to control faulty tasks in a timely and accurate manner in distributed systems, resulting in unstable execution and waste of resources.
By obtaining the execution status data of distributed tasks and resource usage data, calculating the indicator values of the evaluation indicators, and matching them with the rule conditions in the rule engine, precise control of the task is achieved.
It improves the accuracy and timeliness of control of distributed tasks, reduces the risk of resource waste and system crashes, and enhances the stability and efficiency of the system.
Smart Images

Figure CN120315840A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer data processing, and particularly relates to a task control method, device, equipment and medium. Background Art
[0002] With the continuous development of computer technology, more and more tasks can be developed and used online, such as transfer services, payment services, model training services, etc. In order to ensure that each service can be executed normally, data during the service execution process can be collected, and based on the data, it can be determined whether the service needs to be controlled. However, the executed services are scattered and executed at different nodes, and have a long link, making it difficult to control a faulty task in a timely and accurate manner.
[0003] Therefore, how to control a faulty distributed task in a timely and accurate manner is a technical problem to be solved urgently. Summary of the Invention
[0004] Embodiments of the present specification provide a task control method, device, equipment and medium to solve the problem of inaccurate control of faulty tasks in existing task control methods.
[0005] To solve the above technical problem, an embodiment of the present specification provides a task control method, which is applied to a distributed task control system and includes: Obtain event data generated during the execution of a distributed task; the event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task; According to the event data, calculate the value of an evaluation index corresponding to the distributed task; the evaluation index is used to evaluate the execution situation and resource usage situation of the distributed task; Judge whether the index value meets the rule conditions set in the rule engine; the rule conditions are used to judge whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements; If the index value meets the rule conditions, control the distributed task based on the control information set by the rule engine.
[0006] An embodiment of the present specification also provides a task control device, which is applied to a distributed task control system and includes: A data acquisition module, configured to obtain event data generated during the execution of a distributed task; the event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task; An index calculation module, configured to calculate an index value of an evaluation index corresponding to the distributed task according to the event data; the evaluation index is used to evaluate the execution situation and resource usage situation of the distributed task; A judgment module, configured to judge whether the index value meets a rule condition set in a rule engine; the rule condition is used to judge whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements; A control module, configured to, if the evaluation index value meets the rule condition, control the distributed task based on the control information set by the rule engine.
[0007] An embodiment of this specification further provides a task control device, which is applied to a distributed task control system and includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: Obtain event data generated during the execution of the distributed task; the event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task; Calculate an index value of an evaluation index corresponding to the distributed task according to the event data; the evaluation index is used to evaluate the execution situation and resource usage situation of the distributed task; Judge whether the index value meets a rule condition set in a rule engine; the rule condition is used to judge whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements; If the evaluation index value meets the rule condition, control the distributed task based on the control information set by the rule engine.
[0008] An embodiment of this specification further provides a computer-readable storage medium, on which a computer program or instruction is stored, and the computer program or instruction can be executed by a processor to implement the steps of a task control method.
[0009] At least one embodiment of this specification can achieve the following beneficial effects: By obtaining the event data generated during the execution of a distributed task, the metric values of the evaluation metrics used to evaluate the execution status and resource usage of the distributed task can be calculated based on the event data. Furthermore, when it is determined that the metric values meet the rule conditions set in the rule engine, the distributed task can be controlled in a timely manner based on the control information set in the rule engine. Thus, the distributed task control system can be used to query the various event data of the distributed task, eliminating the need to separately obtain data from other different systems; it can also accurately control the distributed task based on the control information in the rule engine that matches the metric values, providing the accuracy for controlling the distributed task. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 is a flowchart of a task control method provided by an embodiment of this specification; Figure 2 is a flowchart of rule execution provided by an embodiment of this specification; Figure 3 is an overall framework diagram of a task control method provided by an embodiment of this specification; Figure 4 corresponds to that provided by an embodiment of this specification Figure 1 is a schematic structural diagram of a task control device; Figure 5 corresponds to that provided by an embodiment of this specification Figure 1 is a schematic structural diagram of a task control device. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] To make the objectives, technical solutions, and advantages of one or more embodiments of this specification clearer, the following will clearly and completely describe the technical solutions of one or more embodiments of this specification in conjunction with the specific embodiments and corresponding accompanying drawings of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by one or more embodiments of this specification.
[0013] The following will, in conjunction with the accompanying drawings, elaborate in detail on the technical solutions provided in each embodiment of this specification.
[0014] In the prior art, monitoring systems such as Prometheus and Grafana can be used to monitor tasks. By collecting task data and setting alarm threshold rules, alarm notifications can be sent after the alarm threshold rules are triggered. However, it only notifies users that a task has failed and does not control the task, nor can it achieve precise management and control.
[0015] To address the deficiencies in the prior art, the following embodiments are provided in this solution: Next, a task control method provided in the embodiments of the specification will be specifically described in conjunction with the accompanying drawings: Figure 1 It is a schematic flowchart of a task control method provided in the embodiments of this specification. From a program perspective, the execution entity of the process can be a program running on an application server, an application client, or a distributed task control system.
[0016] As Figure 1 shown, the method may include the following steps.
[0017] Step 102: Obtain event data generated during the execution of the distributed task; the event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task.
[0018] In the embodiments of this specification, the distributed task may be a task for training a model, a task for resource transfer, a task for commodity trading, a task for using a model, a task executed by a business R & D platform, and so on. The distributed task may include one or more subtasks. The event data may represent data related to the distributed task or data of other tasks that can affect the distributed task. The event data may be generated during the execution of the distributed task by the business R & D platform.
[0019] In the embodiments of this specification, the execution status data may include data such as task start data indicating task start, task failure data indicating task failure, and task completion data indicating task completion. The task start data may include the number of task starts and the start time corresponding to each task, etc. The task failure data may include the number of task failures, the reasons for failure, etc. The task completion data may include the number of task completions and the task completion time, etc. The resource usage data may include data such as CPU usage rate, memory occupancy, occupancy of hardware resources, task response time, and task throughput.
[0020] In practical applications, the event data may further include code detection data of a distributed task, and the code detection data may include data such as code detection result data and code performance analysis data. The code detection result data may be data indicating whether the code corresponding to the distributed task is abnormal, and the code performance analysis data may indicate the quality of the code corresponding to the distributed task and the performance of running the distributed task.
[0021] Step 104: Calculate the metric value of the evaluation metric corresponding to the distributed task according to the event data; the evaluation metric is used to evaluate the execution situation and resource usage situation of the distributed task.
[0022] In the embodiments of this specification, the evaluation metric may be a metric used to evaluate a distributed task. The evaluation metrics may include evaluation metrics such as the success rate metric of distributed task execution and the failure rate metric of distributed task execution for evaluating the execution situation of the distributed task. The evaluation metrics may also include evaluation metrics such as the resource utilization rate metric of the distributed task and the CPU occupancy rate metric of the distributed task for evaluating the resource usage situation of the distributed task.
[0023] In the embodiments of this specification, the distributed task control system may calculate the metric values of multiple evaluation metrics through multiple different dimensions, so as to subsequently control the distributed task based on the metric values of the multiple evaluation metrics, improving the accuracy of controlling the distributed task.
[0024] In practical applications, the evaluation metric may also be used to evaluate the code quality of the distributed task. The metric value of the evaluation metric may also indicate the quality of the code of the distributed task. For example, if the metric value is 1, it may indicate poor code quality; if the metric value is 2, it may indicate relatively poor code quality; if the metric value is 3, it may indicate average code quality; if the metric value is 4, it may indicate relatively good code quality; if the metric value is 5, it may indicate good code quality.
[0025] In practical applications, various plugins for calculating evaluation metrics may be preset, and the plugins may be called based on actual needs to calculate the metric values of the evaluation metrics of the distributed task, thereby eliminating the need to repeatedly set the calculation method of the evaluation metrics, improving the efficiency of calculating the metric values of the evaluation metrics for the distributed task, and reducing the calculation cost.
[0026] Step 106: Determine whether the metric value meets the rule conditions set in the rule engine; the rule conditions are used to determine whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements.
[0027] In the embodiments of this specification, the rule conditions in the rule engine can be set based on user requirements, or can be set based on user requirements and the performance of the distributed task control system, or can also be set based on expert experience. The occupied hardware resources can be resources such as CPU, memory, and hard disk; the execution results can be completion status, failure status, delay status, etc.
[0028] In practical applications, the evaluation metrics can be matched with the rule engine to determine the rule engine that matches the evaluation metrics; then, based on the rule conditions set in the rule engine that matches the evaluation metrics, it can be determined whether the metric value meets the requirements, and further, it can be determined whether the distributed task meets the requirements, and further determine whether it is necessary to control the distributed task.
[0029] Step 108: If the metric value meets the rule conditions, control the distributed task based on the control information set by the rule engine.
[0030] In the embodiments of this specification, the rule engine may include rule conditions and control information corresponding to the rule conditions. The control information can represent the actions that can be performed on the distributed task when the metric value meets the rule conditions. The control information can include information such as stopping the execution of the distributed task, deleting the distributed task, adjusting the resources occupied by the distributed task, etc., which can manage and control the distributed task.
[0031] In practical applications, the distributed task control system can execute the task control method on one or more distributed tasks. The distributed task control system can carry a big data platform to enable large-scale computing and storage of data.
[0032] It should be understood that the order of some steps of the method described in one or more embodiments of this specification can be interchanged according to actual needs, or some of the steps can also be omitted or deleted.
[0033] Figure 1 In the method, by obtaining the event data generated during the execution of the distributed task, the metric value of the evaluation metrics used to evaluate the execution status and resource usage of the distributed task can be calculated based on the event data. Furthermore, when it is determined that the metric value meets the rule conditions set in the rule engine, the distributed task can be controlled in a timely manner based on the control information set in the rule engine. Thus, the distributed task control system can be used to query the various event data of the distributed task, without having to obtain data from other different systems separately; it can also accurately control the distributed task based on the control information in the rule engine that matches the metric value, providing the accuracy of controlling the distributed task.
[0034] Based on Figure 1For the method, some specific implementation manners of the method are further provided in the embodiments of this specification and will be described below.
[0035] Optionally, in order to facilitate querying the control of distributed tasks, in the embodiments of this specification, the method may further include: Recording control log information; the control log information includes metric values, rule conditions, and control result information; Receiving a control log information query request; Responding to the query request and returning the control log information.
[0036] In the embodiments of this specification, the control log information is used to record the log information of controlling distributed tasks whose metric values meet the rule conditions. The control log information may further include identification information of the distributed tasks, the duration of controlling the distributed tasks, the start time of the control, the end time of the control, and other information. The identification information of the distributed tasks can be used to uniquely identify the distributed tasks; it can be any one of the numbers, IDs, and names of the distributed tasks.
[0037] In the embodiments of this specification, the query request may include the identification information of the distributed tasks that the user wants to query, or may include the time period that the user wants to query. The distributed task control system can respond to the query request, search for the table storing the control log information of the distributed tasks based on the identification information of the distributed tasks in the query request, and can also filter the control log information in the table based on the time period in the query request, and return the filtered control log information of the distributed tasks within this time period to the terminal interface, facilitating the user to view the relevant control records, so that the user can understand the running status of the distributed tasks.
[0038] As an implementation manner, optionally, in the embodiments of this specification, the event data includes the execution status data of the distributed tasks and the resource usage data of the distributed tasks; calculating the metric value of the evaluation metric corresponding to the distributed tasks according to the event data may specifically include: calculating the failure rate and the latency rate based on the execution status data; calculating the resource waste rate based on the resource usage data; obtaining the first weight value corresponding to the failure rate, the second weight value corresponding to the latency rate, and the third weight value corresponding to the resource waste rate; calculating the comprehensive score of the distributed tasks based on the failure rate and the first weight value, the latency rate and the second weight value, and the resource waste rate and the third weight value.
[0039] In the embodiments of this specification, the failure rate can represent the ratio of the number of failures of a distributed task to the number of starts of the distributed task. The latency rate can represent the ratio of the number of latencies in the execution of a distributed task to the number of starts of the distributed task. Latency can mean that the distributed task has been completed, but the execution duration is longer than the preset execution duration. The resource waste rate can represent the ratio of the difference between the allocated resources for a distributed task and the actually occupied resources of the distributed task to the allocated resources.
[0040] In the embodiments of this specification, the first weight value, the second weight value, and the third weight value can be set based on expert experience, or based on the type or importance of the distributed task, or based on user requirements. For example, if the user pays more attention to the failure rate of the distributed task, the first weight value corresponding to the failure rate can be set higher than the second weight value and the third weight value.
[0041] In the embodiments of this specification, it is possible to determine whether to control a distributed task based on the comprehensive score of the distributed task. Specifically, it can be determined whether the comprehensive score of the distributed task is greater than a preset score. If the comprehensive score of the distributed task is greater than the preset score, the execution of the distributed task can be stopped; if the comprehensive score of the distributed task is less than or equal to the preset score, there is no need to control the distributed task. For example: the first weight value can be set to 0.4, the second weight value can be set to 0.3, the third weight value can be set to 0.3, and the preset score can be set to 0.3. If the failure rate of the distributed task is 0.15, the latency rate is 0.1, and the resource volume rate is 0.25, then the comprehensive score can be determined to be 0.15 * 0.4 + 0.1 * 0.3 + 0.25 * 0.3 = 0.165. Thus, it can be determined that the comprehensive score is less than the preset score, and there is no need to control the distributed task. Through the combined calculation of multi-dimensional indicators, while realizing the integration of cross-system indicators, it is also possible to precisely control distributed tasks based on multi-dimensional indicators.
[0042] In practical applications, other systems can also be connected to the distributed task control system for control, so that there is no need to redesign the distributed task control system for other systems, reducing the cost of controlling the corresponding tasks of other systems. For example, the logistics scheduling system can be connected to the distributed task control system, so that the combined score can be obtained based on the on-time rate, vehicle empty driving rate, and route optimization coefficient of the logistics tasks in the logistics scheduling system. Among them, the route optimization coefficient can be calculated based on information such as transportation distance, transportation time, and transportation cost; when the combined score is less than the preset logistics score, it is possible to perform control such as route optimization and vehicle allocation adjustment on the logistics task and notify the dispatcher, thereby improving the task accuracy rate, reducing the empty driving rate, and fuel cost.
[0043] To enhance the flexibility of controlling distributed tasks, the weights used for the above comprehensive calculation of scores can be adjusted. Optionally, before calculating the comprehensive score of the distributed task, it may further include: determining whether the current time is within a specified time period; the specified time period includes a time period during which the quantity of tasks to be processed is greater than a specified threshold; if the current time is within the specified time period, then increase the first weight value and the second weight value to obtain an adjusted first weight value and an adjusted second weight value; calculating the comprehensive score of the distributed task, specifically including: calculating the comprehensive score of the distributed task according to the failure rate, the adjusted first weight value, the delay rate, and the adjusted second weight value.
[0044] In the embodiments of this specification, the specified time period may represent a peak period, and the peak period may represent a time period during which the number of distributed tasks that the distributed task control system needs to process is greater than a first preset quantity of tasks; or it may represent a time period during which the amount of resources occupied by the distributed task control system is greater than a first preset amount of resources. The quantity of tasks to be processed may represent the number of distributed tasks to be processed.
[0045] In the embodiments of this specification, if it is within the specified time period, the number of distributed tasks to be processed is relatively large, the amount of resources occupied is relatively large, and the resource waste rate is relatively low. The comprehensive score of the distributed task can be calculated based on the failure rate and the delay rate without considering the resource waste rate. Continuing with the first weight value and the second weight value involved in the above example, if the distributed task control system determines that the current time period is within the specified time period, the first weight value can be adjusted from 0.4 to 0.6; the second weight value can be adjusted from 0.3 to 0.4, or the weight value of the resource waste rate can be adjusted from 0.3 to 0, or the resource waste rate index in the indicators for calculating the comprehensive score can be deleted.
[0046] In practical applications, if the current time period is within a low - peak time period, the third weight value can be increased, and then the comprehensive score of the distributed task is calculated based on the failure rate and the first weight value, and the resource waste rate and the adjusted third weight value. The low - peak time period may represent a time period during which the number of distributed tasks that the distributed task control system needs to process is less than a second preset quantity of tasks; or it may represent a time period during which the amount of resources occupied by the distributed task control system is less than a second preset amount of resources, where the first preset quantity of tasks is greater than the second preset quantity of tasks; the first preset amount of resources is greater than the second preset amount of resources; the specified threshold may represent a value that is greater than or equal to the second preset quantity of tasks and less than or equal to the first preset quantity of tasks.
[0047] In practical applications, if it is in the low-peak time period, the amount of tasks to be processed is small, and the delay rate is relatively low. The comprehensive score of the distributed task can be calculated based on the failure rate and the resource waste rate, without considering the delay rate. Continuing with the third weight value involved in the above example, if the distributed task control system determines that the current time period is in the low-peak time period, the third weight value can be adjusted from 0.3 to 0.6, the second weight value can be adjusted to 0, or the delay rate index in the indicators for calculating the comprehensive score can be deleted. If the current time period is in a stable time period, that is, the amount of tasks to be processed is within the specified threshold, or the resource usage is within the range of the second preset resource amount and the first preset resource amount, it can be regarded as a stable time period, and then the first weight value, the second weight value, and the third weight value do not need to be adjusted.
[0048] In the embodiments of this specification, based on preset conditions, such as the time period in which the current time period is located, the weight values corresponding to each indicator in the combined calculation can be determined, and then the comprehensive score can be determined. Thus, the indicator weights can be adjusted according to the system state, and then the distributed task can be controlled more accurately based on the comprehensive score calculated by the dynamically adjusted indicator weights.
[0049] As another implementation manner, optionally, in the embodiments of this specification, the event data includes the execution status data of the distributed task and the resource usage data of the distributed task; calculating the index value of the evaluation index corresponding to the distributed task according to the event data may specifically include: calculating the success rate based on the execution status data; calculating the resource utilization rate based on the resource usage data; calculating the resource efficiency index of the distributed task based on the resource utilization rate and the success rate; wherein, the resource efficiency index is positively correlated with the logarithm of the resource utilization rate; the resource efficiency index is positively correlated with the success rate.
[0050] In the embodiments of this specification, the success rate can represent the proportion of the number of successful executions of the distributed task to the number of starts of the distributed task. The resource utilization rate can represent the proportion of the allocated resources to the distributed task to the allocated resources.
[0051] In the embodiments of this specification, the resource efficiency index can be calculated based on the formula: resource efficiency index = success rate * ln(1 + resource utilization rate) * 100. Among them, the higher the success rate, the larger the resource efficiency index; the higher the resource utilization rate, the larger the resource efficiency index.
[0052] In the embodiments of this specification, it is possible to determine whether the resource efficiency index is less than or equal to a preset index. If the resource efficiency index is less than or equal to the preset index, the amount of resources allocated to the distributed task can be reduced; if the resource efficiency index is greater than the preset index, there is no need to control the distributed task. For example, the preset index is set to 50. Assuming that the success rate of the distributed task is 0.85 and the resource utilization rate is 0.7, the resource efficiency index = 0.85 * ln(1 + 0.7) * 100 = 45. It can be determined that the resource efficiency index is less than the preset index, so 30% can be recovered from the amount of resources allocated to the distributed task. Thus, the performance of the distributed task can be determined based on the success rate and the resource utilization rate, and then it can be determined whether control is needed, thereby improving the accuracy of the control of the distributed task.
[0053] As another implementation, it is also possible to perform combined calculation again based on the combined calculated metrics to further improve the accuracy of control. Specifically, the resource efficiency index can be adjusted based on the comprehensive score to obtain the adjusted resource efficiency index, so that it can be determined whether to control the distributed task based on the adjusted resource efficiency index, and the control actions can also be adjusted correspondingly. The preset index can also be adjusted, which can improve the accuracy of the evaluation criteria, and then improve the accuracy of the control result, reduce the error rate of controlling the distributed task, and thus reduce the resource consumption caused by improper control.
[0054] In practical applications, the adjusted resource efficiency index can be calculated based on the formula: adjusted resource efficiency index = original resource efficiency index * (1 - comprehensive score). For example, the preset index is 0.4, assuming that the original resource efficiency index is 0.6 and the comprehensive score is 0.4, it can be determined that the adjusted resource efficiency index is 0.36, and it is less than the preset index 0.4, so the resources occupied by the distributed task can be readjusted. Specifically, the proportion of resources occupied by each subtask can be adjusted.
[0055] As another implementation, it is also possible to predict the failure rate of the distributed task in the future time period based on the failure rate within a preset time period, and then determine whether to control the distributed task based on the predicted failure rate. Specifically, obtain the failure rate of the current cycle, determine the adjustment coefficient based on the predicted cycle, determine the failure rate of the predicted cycle based on the adjustment coefficient and the failure rate of the current cycle, and judge whether the failure rate of the predicted cycle is greater than the preset predicted value. If the failure rate of the predicted cycle is greater than the preset predicted value, an alarm message can be sent to the user terminal; if the failure rate of the predicted cycle is less than or equal to the preset predicted value, there is no need to send an alarm message.
[0056] In practical applications, the current cycle may include the current moment and the start moment of the current cycle; the initial moment may be determined based on the current moment and the preset duration corresponding to the cycle. For example, if the current moment is 15:59 on January 10, 2025, and assuming the preset duration of the cycle is 24 hours, then the start moment of the current cycle can be determined as 16:00 on January 9, 2025.
[0057] In practical applications, the failure rate of the predicted cycle can be calculated based on the formula: failure rate of the predicted cycle = failure rate of the current cycle + (1 + 0.1 * the number of cycles between the predicted cycle and the current cycle). For example, if the preset prediction value is set to 0.3, and assuming the failure rate of the current cycle is 0.25, if it is necessary to predict the failure rate of the cycle three cycles away from the current moment, then the failure rate of the predicted cycle = 0.25 + (1 + 0.1 * 3) = 0.325. If it is determined that the failure rate of the predicted cycle is greater than the preset prediction value, then an alarm message can be generated based on the distributed task and sent to the user terminal, so that the user terminal can be alerted in a timely manner, minimizing the failure rate of the distributed task, increasing the probability of successful execution of the distributed task, and reducing the resource consumption caused by the failure of the distributed task execution.
[0058] In practical applications, a visualization window can also be provided for the user, enabling the user to combine and configure multiple metrics by dragging the metrics in the visualization window, and then generating a corresponding formula based on the interactive formula editor to calculate the metric values; so as to evaluate the distributed task from multiple different dimensions and then perform control operations.
[0059] In practical applications, the above method of calculating metrics can be completed based on the called metric calculation plug-in. The server can pre-generate metric calculation plug-ins based on various single metrics or combined metrics, so as to flexibly call the corresponding metric calculation plug-ins based on the type or calculation requirements of the distributed task, avoiding re-editing the metric calculation method for each distributed task, reducing the repetition of the same work, and also saving costs.
[0060] As an implementation method, optionally, in the embodiments of this specification, the event data includes the execution status data of the distributed task; the metric value includes the failure rate calculated based on the execution status data; determining whether the metric value meets the rule conditions set in the rule engine may specifically include: determining whether the failure rate is greater than the first preset failure rate; if the metric value meets the rule conditions, then controlling the distributed task based on the control information set in the rule engine, specifically including: if the failure rate is greater than the first preset failure rate, then stop executing the distributed task.
[0061] In the embodiments of this specification, during the process of stopping the execution of a distributed task, an alarm message for stopping the execution of the distributed task can also be sent to the user terminal, so that after the user obtains the alarm message, the distributed task can be processed in a timely manner, and then the distributed task can be executed normally as soon as possible, reducing the resource loss caused by stopping the execution of the distributed task.
[0062] In the embodiments of this specification, the server can set the stop duration for stopping the execution of the distributed task. After the stop duration is reached, the distributed task can continue to be executed, thereby reducing the amount of distributed tasks to be processed and avoiding the impact caused by stopping the execution of the distributed task for too long. For example, if the distributed task is a transaction task, if the transaction task is stopped for too long, it is likely to cause inconvenience to transaction users and also cause some losses to users.
[0063] In practical applications, if the failure rate is less than or equal to the first preset failure rate, the distributed task can continue to be executed.
[0064] In practical applications, in the distributed task control system, the metric layer and the rule engine layer can be executed in parallel to improve the execution efficiency, increase the CPU utilization rate, and reduce the average latency. The control execution layer can also be separated from the rule engine layer to improve the execution efficiency through an asynchronous framework.
[0065] To clearly illustrate the relationship between the metric layer, the rule engine layer, and the control execution layer, Figure 2 is a rule execution flow chart provided by the embodiments of this specification. As Figure 2 shown, the metric layer can calculate the metric value of the evaluation metric based on the metric plugin and the event data; then, the metric value of the evaluation metric can be sent to the rule engine layer. The rule engine layer can match the metric value with the rule condition. If the metric value matches the rule condition successfully, the control information contained in the rule engine with the rule condition can be determined, and the control information indicating the triggering of the control action can be sent to the control layer. The control layer can execute the corresponding control action based on the control information.
[0066] In practical applications, a priority can also be set for the control action or the rule engine. If the metric values of multiple evaluation metrics corresponding to the distributed task are matched with multiple rule engines in the rule engine layer, and multiple rule conditions are hit, then based on the priority of the rule engine, the control action corresponding to the control information in the rule engine with the highest priority can be executed.
[0067] As an implementation, the first preset failure rate can be adjusted based on the execution environment where the distributed task is actually executed and the user's requirements, etc. Optionally, in the embodiments of this specification, the event data further includes the resource usage data of the distributed task; the metric value further includes the resource utilization rate calculated based on the resource usage data; before determining whether the failure rate is greater than the first preset failure rate, it may further include: determining whether the resource utilization rate is greater than the first resource utilization rate threshold; if the resource utilization rate is greater than the first resource utilization rate threshold, multiplying the first preset failure rate by a first preset coefficient to obtain a first adjusted failure rate; the first preset coefficient is a positive number less than 1; determining whether the failure rate is greater than the first preset failure rate specifically includes: determining whether the failure rate is greater than the first adjusted failure rate.
[0068] In the embodiments of this specification, if the resource utilization rate is greater than the first resource utilization rate threshold, it can be determined that the current resource utilization of the distributed task is relatively high, and it is in a relatively optimal and high-load working state. The value of the first preset failure rate can be reduced, thereby improving the execution requirements for the distributed task, reducing the error rate of the distributed task, and avoiding problems such as system crashes caused by errors in the distributed task control system under high load.
[0069] For example, the first resource utilization rate threshold can be set to 0.8, the first preset failure rate can be set to 0.2, and the first preset coefficient is 0.7; assuming the resource utilization rate is 0.9, the first preset failure rate can be adjusted to 0.14 to obtain the first adjusted failure rate, so that when the failure rate of the distributed task is greater than the first adjusted failure rate, the distributed task can be controlled in a timely manner, avoiding system crashes and reducing waste of resources.
[0070] As another implementation, optionally, in the embodiments of this specification, the event data further includes the resource usage data of the distributed task; before determining whether the failure rate is greater than the first preset failure rate, it may further include: determining whether the resource utilization rate is less than the second resource utilization rate threshold; the second resource utilization rate threshold is less than the first resource utilization rate threshold; if the resource utilization rate is less than the second resource utilization rate threshold, multiplying the first preset failure rate by a second preset coefficient to obtain a second adjusted failure rate; the first preset coefficient is greater than 1; determining whether the failure rate is greater than the first preset failure rate specifically includes: determining whether the failure rate is greater than the second adjusted failure rate.
[0071] In the embodiments of this specification, if the resource utilization rate is less than the second resource utilization rate threshold, it can be determined that the current resource utilization of the distributed task is relatively low, and it is in a low-load working state. The value of the first preset failure rate can be appropriately relaxed. Furthermore, it can be determined that the current distributed task has a relatively small impact on the distributed task control system, and the error rate of the distributed task is allowed to be relatively high. Thus, flexible and real-time adjustment of the rule conditions can be achieved, providing an intelligent decision-making for the control of the distributed task.
[0072] For example, the second resource utilization rate threshold can be set to 0.3, the first preset failure rate can be set to 0.2, and the second preset coefficient is 1.2. Assuming that the resource utilization rate is 0.25, the first preset failure rate can be adjusted to 0.24 to obtain the second adjusted failure rate. Thus, when the failure rate of the distributed task is greater than the second adjusted failure rate, the distributed task can be timely controlled to avoid a relatively high error rate and affect the normal operation of the distributed task.
[0073] In the embodiments of this specification, if the resource utilization rate is greater than or equal to the second resource utilization rate threshold and less than or equal to the first resource utilization rate threshold, the first preset failure rate may not be adjusted.
[0074] As an implementation manner, optionally, in the embodiments of this specification, determining whether the index value meets the rule conditions set in the rule engine may specifically include: determining whether the index value of the evaluation index is greater than the first preset index value corresponding to the evaluation index; the evaluation index includes the resource occupancy rate or the number of failures; controlling the distributed task based on the control information set by the rule engine specifically includes: if the index value of the evaluation index is greater than the first preset index value corresponding to the evaluation index, sending an alarm message to the user terminal; determining whether the index value is greater than the second preset index value corresponding to the evaluation index; the second preset index value is greater than the first preset index value; if the index value is greater than the second preset index value, reducing the execution priority of the distributed task.
[0075] In the embodiments of this specification, the resource occupancy rate may represent the ratio of the resource amount occupied by the distributed task to the total resource amount. For example, the ratio of the memory occupied by the distributed task to the total memory, or the ratio of the CPU occupied by the distributed task to the total CPU, etc. The number of failures may represent the number of failures that occur during the execution of the distributed task by the distributed task control system. The failures may include interface exceptions, network failures, code errors, and other failures. The user may refer to the R & D personnel who develop the distributed task or the management personnel who manage the distributed task.
[0076] In the embodiments of this specification, the reasons for failures can be determined based on a variety of different methods. For example, a full-link tracing system can trace based on the execution link of distributed tasks to determine the cause of the failure; an intelligent diagnosis engine can generate an intelligent diagnosis engine based on pre-set failure diagnosis rules to trace the execution failures of distributed tasks, or other methods can be used to find the cause of the failure, so as to quickly and accurately locate the cause of the failure, assist users in handling the failure, and improve the handling efficiency.
[0077] In the embodiments of this specification, if the evaluation index is the resource occupancy rate, it can be determined whether: the resource occupancy rate is greater than the first preset occupancy rate, and the duration for which the resource occupancy rate is greater than the first preset occupancy rate is greater than or equal to the first preset duration; if the resource occupancy rate is greater than the first preset occupancy rate, and the duration for which the resource occupancy rate is greater than the first preset occupancy rate is greater than or equal to the first preset duration, an alarm message indicating that the distributed task continuously occupies an excessive resource occupancy rate can be sent to the user terminal, so as to be able to notify the user in a timely manner and prevent the user from being unaware of the status of the distributed task. For example, if the first preset occupancy rate is set to 70% and the first preset duration is set to 60s, and if the CPU occupancy rate of the distributed task > 70% and lasts for 60s, an alarm message indicating that the CPU occupancy rate is too high and suggesting capacity expansion can be sent to the user terminal.
[0078] After sending the alarm message, it can also be determined whether: the resource occupancy rate is greater than the second preset occupancy rate, and the duration for which the resource occupancy rate is greater than the second preset resource occupancy rate is greater than or equal to the second preset duration, and the second preset duration is less than or equal to the first preset duration; if the resource occupancy rate is greater than the second preset occupancy rate, and the duration for which the resource occupancy rate is greater than the second preset resource occupancy rate is greater than or equal to the second preset duration, the execution priority of the distributed task can be reduced. Thus, by reducing the execution priority of the distributed task, the amount of resources occupied by the distributed task can be reduced. Continuing with the above example, if the second preset occupancy rate is set to 80% and the second preset duration is set to 30s, the user has performed capacity expansion, but after the capacity expansion, the occupancy rate of the distributed task reaches 85% and the duration has reached 30s, then the execution priority of the distributed task can be reduced, thereby reducing the quota of the distributed task. At the same time, task requests for non-core tasks can also be rejected to ensure that core tasks can continue to execute without being affected.
[0079] In the embodiments of this specification, if the evaluation index is the number of failures, it can be determined whether the number of failures that occur during the execution of the distributed task by the distributed task control system is 1. If the number of failures that occur in the distributed task control system is 1, a failure log can be recorded, and an alarm message indicating that the distributed task control system has failed can be sent to the user terminal.
[0080] After sending the alarm information, it is also possible to determine whether the number of failures that occur during the execution of the distributed task by the distributed task control system is greater than a second preset number of failures; if the number of failures of the distributed task control system is greater than the second preset number of failures, the execution priority of the distributed task can be reduced, where the second preset number of failures can be calculated based on the number of starts of the distributed task and a first preset ratio. For example, if the first preset ratio is 10% and the number of starts is 100 times, the second preset number of failures can be determined to be 10. If the number of failures that occur during the execution of the distributed task by the distributed task control system is 11 times, while reducing the execution priority of the distributed task, the call volume of the sub-tasks that fail in the distributed task can also be reduced, such as reducing the call volume to 80%.
[0081] In the embodiments of this specification, the above method can be used to perform progressive control on the distributed task, thereby avoiding controlling the distributed task based on a single threshold, which affects the operation of other tasks, or reducing the probability of failures in the distributed task system.
[0082] As another implementation, optionally, in the embodiments of this specification, after reducing the priority of executing the distributed task, it may further include: determining whether the metric value is greater than a third preset metric value corresponding to the evaluation metric; the third preset metric value is greater than the second preset metric value; if the metric value is greater than the third preset metric value, stop executing the distributed task.
[0083] In the embodiments of this specification, if the evaluation metric is the resource occupancy rate, it can be determined whether: the resource occupancy rate is greater than a third preset occupancy rate, and the duration for which the resource occupancy rate is greater than the third preset occupancy rate is greater than or equal to a third preset duration, and the third preset duration is less than or equal to the second preset duration; if the resource occupancy rate is greater than the third preset occupancy rate, and the duration for which the resource occupancy rate is greater than the third preset occupancy rate is greater than or equal to the third preset duration, the execution of the distributed task can be stopped. For example, if the third preset occupancy rate is set to 90% and the third preset duration is 10s, if the CPU occupancy rate of the distributed task > 90% and lasts for 10s, the execution of the distributed task can be stopped to avoid affecting the distributed tasks for executing core services in the distributed task control system; it can be understood that if the distributed task is a distributed task for executing core tasks, traffic limiting can be performed on other distributed tasks in the distributed task control system to give priority to ensuring that the distributed tasks for executing core services are not affected and can continue to run.
[0084] In the embodiments of this specification, if the evaluation index is the number of failures, it is also possible to determine whether the number of failures that occur during the execution of the distributed task by the distributed task control system is greater than a third preset failure number; if the number of failures that occur in the distributed task control system is greater than the third preset failure number, the execution of the distributed task can be stopped. Among them, the third preset failure number can be calculated based on the number of starts of the distributed task and a second preset ratio, and the second preset ratio is greater than the first preset ratio. For example, if the second preset ratio is 30% and the number of starts is 200 times, the third preset failure number can be determined to be 60. If the number of failures that occur during the execution of the distributed task by the distributed task control system is 61 times, the distributed task can be isolated from the distributed task control system and the execution of the distributed task can be stopped.
[0085] In practical applications, progressive control can also be performed based on other evaluation indexes. For example, evaluation indexes such as the latency of the distributed task and the failure rate of the distributed task are not listed one by one here due to space limitations.
[0086] In the embodiments of this specification, multiple distributed tasks can be executed in the distributed task control system, which can quickly respond to the control of the distributed task and can also ensure the continuity of the service corresponding to the distributed task to the greatest extent. Moreover, the above-mentioned progressive control can adopt the method of increasing the threshold to gradually control each distributed task in the distributed task control system, thereby reducing the misinterception rate, reducing the number of manual interventions, reducing the scope of business impact, and improving the rate of fault recovery. For example, using the traditional method of controlling with one threshold, the misinterception rate is 2.1%, and the misinterception rate of the progressive control is 0.4%, which can reduce the misinterception rate by 81%.
[0087] As an implementation manner, optionally, in the embodiments of this specification, the event data includes the resource usage data of the distributed task; the method may further include: determining whether the event data meets a preset first trigger condition; the first trigger condition includes at least one of the CPU usage rate exceeding the preset CPU usage rate and the memory usage rate being greater than the preset memory usage rate; the CPU usage rate and the memory usage rate are calculated based on the resource usage data of the distributed task; if the event data meets the preset first trigger condition, perform a first preset action; the first preset action includes at least one of restricting new task scheduling, adjusting the resource allocation rule, and triggering the capacity expansion process.
[0088] In the embodiments of this specification, the CPU usage rate can represent the ratio of the CPU occupied during the execution of the distributed task to the CPU allocated to the distributed task; the memory usage rate can represent the ratio of the memory occupied by the distributed task to the memory allocated to the distributed task.
[0089] In the embodiments of the present specification, there is a corresponding relationship between the first trigger action and the first preset action, and the first preset action can be configured based on the first trigger action by the distributed task control system. Restricting new task scheduling can mean reducing the number of distributed tasks newly added to the distributed task control system for execution or stopping the execution of the distributed tasks newly added to the distributed task control system. Adjusting the resource allocation rule can mean adjusting the amount of resources allocated to each distributed task in the distributed task control system. Triggering the expansion process can mean increasing the operating space and storage space of the distributed task control system, etc., and thus increasing the amount of resources allocated to the distributed tasks, such as adding a new server. Thus, the rule conditions can be dynamically adjusted or new rules can be added based on the event data to improve the efficiency of the distributed task control system in processing distributed tasks.
[0090] As another implementation manner, optionally, in the embodiments of the present specification, the event data includes the execution status data of the distributed task; the method may further include: determining whether the event data meets a preset second trigger condition; the second trigger condition includes at least one of a task request growth rate greater than a preset request growth rate and a failure rate greater than a second preset failure rate; the task request growth rate and the failure rate are calculated based on the execution status data of the distributed task; if the event data meets the preset second trigger condition, then execute a second preset action; the second preset action includes at least one of starting a current limiting rule and reducing the priority of selected tasks.
[0091] In the embodiments of the present specification, the task request growth rate can represent the ratio of the difference between the task requests of the distributed tasks received per second by the distributed task control system in the current time period and the task requests of the distributed tasks received per second in the historical time period to the task requests of the distributed tasks received per second in the historical time period. The failure rate can represent the ratio of the number of failures of the distributed task to the number of starts.
[0092] In the embodiments of the present specification, there can be a corresponding relationship between the second trigger condition and the second preset action, and the second preset action can be configured by the distributed task control system based on the second trigger condition. Starting the current limiting rule can mean restricting the rate at which the distributed task control system receives task requests based on the token bucket method. Reducing the priority of selected tasks can mean reducing the execution priority of the distributed tasks belonging to non-core tasks among several distributed tasks processed in the distributed task control system.
[0093] In practical applications, the second preset action may further include enabling a verification code verification rule. For example, a user can request to generate a verification code, store it in the cache of the distributed task control system after generating the first verification code, and set a valid duration for the first verification code. As a result, when the user sends a task request to the distributed task control system using the task R & D platform or the user terminal, the second verification code can be carried. The distributed task control system can determine whether the second verification code carried in the task request is the same as the first verification code in the cache and within the valid period based on the first verification code and the valid duration in the cache. If so, the distributed task system can process the task request, thereby effectively reducing the situation of malicious requests or automated attacks on the distributed task control system.
[0094] In practical applications, when there is a major promotion event, the rule conditions can also be adjusted within a preset duration before the event starts. For example, resource expansion can be performed, and the throttling threshold can be adjusted, etc. The dynamic adjustment of the rule conditions can be adjusted based on actual needs, and no specific limitation is made here.
[0095] As an implementation manner, optionally, in the embodiments of this specification, after controlling the distributed task based on the control information set by the rule engine, it may further include: generating a first notification message and a second notification message based on the control result of the distributed task; sending the first notification message to the server so that the server records the execution status of the distributed task based on the first notification message; sending the second notification message to the user terminal so that the user terminal displays the control result of the distributed task to the user.
[0096] In the embodiments of this specification, the first notification message can be used to notify the distributed task control system to record the control log information for the distributed task, specifically including information such as the execution status of the distributed task, the identification information of the distributed task, and the control type for the distributed task. The first notification message can be specifically sent through Webhook or a message queue. Thus, the synchronization of the control status and the traceability of the control process of the distributed task can be achieved, etc.
[0097] In the embodiments of this specification, the second notification message can be used to notify the user terminal of the control result of the distributed task, can also send information that requires manual intervention to the user terminal, or can also send information about the abnormality of the distributed task to the user, etc.
[0098] As an implementation manner, optionally, in the embodiments of this specification, the sending the second notification message to the user terminal may specifically include: determining the notification level of the second notification message based on the control result; using a notification method corresponding to the notification level to send the second notification message.
[0099] In the embodiment of this specification, if the notification level of the second notification information is higher, the frequency of sending the second notification information is higher, the more notification methods of sending the second notification information are, and the effectiveness of sending the second notification information is higher, so that the user can check the second notification information in time, and then solve the problems existing in distributed tasks.
[0100] For example, the first level is an emergency level, which requires users to handle distributed tasks immediately. Voice broadcast, SMS sending and the first instant messaging method can be used for notification. The first instant messaging method can be a user's internal communication method, and the notification is performed at a first preset frequency. If it is not resolved within the first preset notification duration, a notification can be sent to the supervisor user so that the supervisor can solve the problem in time, or assign a more professional user to solve the problem. The second level can be an alarm level, which requires users to handle distributed tasks within the first time period. Email and the second instant messaging method can be used for notification. The second instant messaging method can be a third-party communication method, and the notification is performed at a second preset frequency. The second preset frequency is greater than the first preset frequency. If it is not resolved within the second preset notification duration, the alarm level can be upgraded to an emergency level, and the second preset notification duration is greater than the first preset notification duration. The third level can be a notification level, which does not require user processing and is only used to notify the user. The user can be notified once by SMS and APP push. If the first instant messaging method in the emergency level fails to send, the email method in the alarm level can be used to send, so that the notification arrival rate can be improved, and the user can handle the faults in the distributed task in time, avoiding excessive resource loss caused by too long a time.
[0101] In actual applications, the second notification information can be generated based on a notification template. For example, for an order processing timeout alarm, the template may include order status information, order waiting time, order details link information, etc.; for an emergency payment exception notification, the content in the SMS template may include payment core service exceptions, error code information, duration of the exception, etc. For inventory alarms, the email template may include subject, product ID, product inventory, suggestion information, etc. Thus, the second notification information can be quickly generated based on the notification template, improving the efficiency of generating the second notification information, thereby improving the efficiency of notifying users, and further improving the rate at which distributed tasks return to normal.
[0102] As an implementation mode, optionally, in the embodiments of the present specification, the acquisition of event data generated during the execution of distributed tasks may specifically include: using a polling plug-in to collect the event data according to a preset period; or, using a monitoring plug-in to continuously monitor the distributed tasks to obtain the event data; or, using a polling plug-in to collect the event data according to a preset period, and using a monitoring plug-in to continuously monitor the distributed tasks to obtain the event data.
[0103] In the embodiments of this specification, the polling plugin can be a component for regularly checking the system status and can obtain event data from multiple different systems; the monitoring plugin can be a system component for capturing event changes in the system and can obtain event data from the subscribed message middleware.
[0104] In practical applications, it is also possible to obtain task requests from different business systems, such as event data representing requests like payment requests, detection requests, and exception handling requests. To facilitate the query and use of event data, the collected event data can be aggregated and processed together and stored in a preset database.
[0105] In practical applications, the collected event data can be subjected to data formatting processing so that the event data has a unified format, and it is determined whether there are abnormalities in the event data, such as abnormal collection time, inconsistent status of distributed tasks in the event data with the actual status, etc. The event data can be deleted, and the remaining event data with a unified format after deletion can be aggregated into a preset database so that when calculating metrics, data can be queried from the preset database.
[0106] In practical applications, the data interfaces for different scenarios are different, and the methods of collecting data are also different. The pre-set collection plugins can be flexibly used to collect event data in different scenarios.
[0107] To clearly illustrate the task control method in the embodiments of this specification, Figure 3 is an overall framework diagram of a task control method provided by the embodiments of this specification.
[0108] As Figure 3 shown, in the data plugin layer, the monitoring plugin can obtain first data related to distributed tasks from the task R & D platform based on the message middleware; the polling plugin can obtain second data related to distributed tasks from an external system. The data service layer can receive the first data and the second data from the data plugin layer; the diagnostic system in the data service layer diagnoses and processes the first data to obtain diagnostic data including execution status data and resource usage data; the code detection callback in the data service layer detects the code information contained in the second data to obtain code detection data; the approval callback in the data service layer obtains approval data, and the approval data can include data that needs to be system-approved, such as the identifier of the distributed task corresponding to the task request received by the distributed task control system, the response time, etc.; if the distributed task processing fails as determined by the diagnostic system in the data service layer, it can be given to the work order system, so that the work order system can generate work order information based on the failed distributed task and send it to the user terminal for the user to view.
[0109] The data service layer can transmit the acquired data to the data detail layer in a zero-copy transfer manner, so that the data detail layer can organize and clean the data. The data service layer can store the cleaned data in a preset database for aggregation. The metric layer can obtain event data (including diagnostic data, code detection data, and approval data) for calculating metric values for evaluating metrics from the preset database, calculate the metric values using a metric calculation plugin or a preset metric calculation method, and provide the metric values of the evaluation metrics to the rule engine layer; the rule engine layer can match based on the rule conditions and the metric values of the evaluation metrics, and when it is determined that the metric values meet the rule conditions, send the control information corresponding to the rule conditions to the control layer; the control layer can perform corresponding control actions on the distributed task based on the control information. The control layer can send the control result to the result recording layer, and the result recording layer can record the control log information for the distributed task. The notification layer can generate notification information based on the control log information recorded by the result recording layer and send it to the user terminal through an interface. Among them, data processing can be performed asynchronously between the various layers in the distributed task control system.
[0110] In practical applications, the task R & D platform can send a distributed task detection request to the distributed task control system provided in the embodiments of this specification, and then can execute the task control method based on the above process. The distributed task control system can also include an algorithm platform and a big data platform, so as to be able to calculate the data of the distributed task using the algorithm capabilities provided by the algorithm platform and the big data platform; store the data using the storage services provided by the big data platform.
[0111] Through the above method, multi-dimensional metric evaluation can be performed on the distributed task, so that accurate and fast control of the distributed task can be performed based on metrics from multiple dimensions, reducing the false interception rate and improving the repair efficiency of distributed task anomalies.
[0112] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method. Figure 4 For the corresponding to the embodiments of this specification Figure 1 structural schematic diagram of a task control device. The task control device can be applied to a distributed task control system, as Figure 4 shown, the device can include: A data acquisition module 402, configured to acquire event data generated during the execution of the distributed task; the event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task; An index calculation module 404, configured to calculate an index value of an evaluation index corresponding to the distributed task according to the event data; the evaluation index is used to evaluate the execution status and resource usage of the distributed task; A judgment module 406, configured to judge whether the index value meets a rule condition set in a rule engine; the rule condition is used to judge whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements; A control module 408, configured to control the distributed task based on control information set by the rule engine if the evaluation index value meets the rule condition.
[0113] Based on Figure 4 For the device, embodiments of this specification also provide some specific implementation schemes of this method, which will be described below.
[0114] Optionally, the index calculation module may specifically be configured to: calculate a failure rate and a latency rate based on the execution status data; calculate a resource waste rate based on the resource usage data; obtain a first weight value corresponding to the failure rate, a second weight value corresponding to the latency rate, and a third weight value corresponding to the resource waste rate; calculate a comprehensive score of the distributed task based on the failure rate and the first weight value, the latency rate and the second weight value, and the resource waste rate and the third weight value.
[0115] Optionally, the index calculation module includes a weight adjustment unit, which may specifically be configured to: judge whether the current time is within a specified time period; the specified time period includes a time period when the quantity of tasks to be processed is greater than a specified threshold; if the current time is within the specified time period, increase the first weight value and the second weight value to obtain an adjusted first weight value and an adjusted second weight value; calculate the comprehensive score of the distributed task, specifically including: calculating the comprehensive score of the distributed task according to the failure rate and the adjusted first weight value, and the latency rate and the adjusted second weight value.
[0116] Optionally, the index calculation module may specifically be configured to: calculate a success rate based on the execution status data; calculate a resource utilization rate based on the resource usage data; calculate a resource efficiency index of the distributed task based on the resource utilization rate and the success rate; wherein, the resource efficiency index is positively correlated with the logarithm value of the resource utilization rate; the resource efficiency index is positively correlated with the success rate.
[0117] Optionally, the metric value includes a failure rate calculated based on the execution status data; the judgment module may specifically be configured to: judge whether the failure rate is greater than a first preset failure rate; if the metric value meets the rule condition, then control the distributed task based on the control information set by the rule engine, specifically including: if the failure rate is greater than the first preset failure rate, then stop executing the distributed task.
[0118] Optionally, the metric value further includes a resource utilization rate calculated based on the resource usage data; the judgment module may specifically be configured to: judge whether the resource utilization rate is greater than a first resource utilization rate threshold; if the resource utilization rate is greater than the first resource utilization rate threshold, then multiply the first preset failure rate by a first preset coefficient to obtain a first adjusted failure rate; the first preset coefficient is a positive number less than 1; the judgment of whether the failure rate is greater than the first preset failure rate specifically includes: judging whether the failure rate is greater than the first adjusted failure rate.
[0119] Optionally, the judgment module may specifically be configured to: judge whether the resource utilization rate is less than a second resource utilization rate threshold; the second resource utilization rate threshold is less than the first resource utilization rate threshold; if the resource utilization rate is less than the second resource utilization rate threshold, then multiply the first preset failure rate by a second preset coefficient to obtain a second adjusted failure rate; the first preset coefficient is greater than 1; the judgment of whether the failure rate is greater than the first preset failure rate specifically includes: judging whether the failure rate is greater than the second adjusted failure rate.
[0120] Optionally, the judgment module may specifically be configured to: judge whether the metric value of the evaluation metric is greater than a first preset metric value corresponding to the evaluation metric; the evaluation metric includes a resource occupancy rate or the number of failures; the control of the distributed task based on the control information set by the rule engine specifically includes: if the metric value of the evaluation metric is greater than the first preset metric value corresponding to the evaluation metric, then send an alarm message to the user terminal; judge whether the metric value is greater than a second preset metric value corresponding to the evaluation metric; the second preset metric value is greater than the first preset metric value; if the metric value is greater than the second preset metric value, then reduce the execution priority of the distributed task.
[0121] Optionally, the device may further be configured to: judge whether the metric value is greater than a third preset metric value corresponding to the evaluation metric; the third preset metric value is greater than the second preset metric value; if the metric value is greater than the third preset metric value, then stop executing the distributed task.
[0122] Optionally, the device can also be used to: determine whether the event data meets a preset first trigger condition; the first trigger condition includes at least one of the CPU usage rate exceeding a preset CPU usage rate and the memory usage rate being greater than a preset memory usage rate; the CPU usage rate and the memory usage rate are calculated based on the resource usage data of the distributed task; if the event data meets the preset first trigger condition, then perform a first preset action; the first preset action includes at least one of restricting new task scheduling, adjusting the resource allocation rule, and triggering an expansion process.
[0123] Optionally, the device can also be used to: determine whether the event data meets a preset second trigger condition; the second trigger condition includes at least one of the task request growth rate being greater than a preset request growth rate and the failure rate being greater than a second preset failure rate; the task request growth rate and the failure rate are calculated based on the execution status data of the distributed task; if the event data meets the preset second trigger condition, then perform a second preset action; the second preset action includes at least one of starting a current limiting rule and reducing the priority of selected tasks.
[0124] Optionally, the device can also be used to: generate a first notification message and a second notification message based on the control result of the distributed task; send the first notification message to the server so that the server records the execution status of the distributed task based on the first notification message; send the second notification message to the user terminal so that the user terminal displays the control result of the distributed task to the user.
[0125] Optionally, the device can also be used to: determine the notification level of the second notification message based on the control result; use a notification method corresponding to the notification level to send the second notification message.
[0126] Optionally, the data acquisition module can specifically be used to: collect the event data using a polling plugin at a preset period; or, continuously monitor the distributed task using a monitoring plugin to obtain the event data; or, collect the event data using a polling plugin at a preset period and continuously monitor the distributed task using a monitoring plugin to obtain the event data.
[0127] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method.
[0128] Figure 5 For the corresponding one provided by the embodiments of this specification Figure 1 is a schematic structural diagram of a task control device. The task control device can be applied to a distributed task control system. As Figure 5 shown, the device 500 may include: at least one processor 510; and, a memory 530 communicatively connected to the at least one processor; wherein, the memory 530 stores instructions 520 executable by the at least one processor 510, and when the instructions are executed by the at least one processor 510, the at least one processor 510 is enabled to: obtain event data generated during the execution of a distributed task; the event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task; calculate a metric value of an evaluation metric corresponding to the distributed task according to the event data; the evaluation metric is used to evaluate the execution situation and resource usage of the distributed task; determine whether the metric value meets a rule condition set in a rule engine; the rule condition is used to determine whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements; if the evaluation metric value meets the rule condition, control the distributed task based on the control information set by the rule engine.
[0129] Based on the same idea, an embodiment of this specification also provides a computer-readable storage medium corresponding to the above method. A computer program or instructions are stored on the computer-readable storage medium, and the computer program or instructions can be executed by a processor to implement the steps of the above task control method.
[0130] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for Figure 5 the device shown, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0131] In the 1990s, it was quite obvious to distinguish whether an improvement in a technology was an improvement in hardware (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement in method flows). However, with the development of technology, many improvements in method flows today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method flows into the hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user's programming of the device. The designer can program by himself to "integrate" a digital system on a single PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). And there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0132] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0133] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0134] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0135] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0136] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart Figure 1 for one or more flows and / or blocks Figure 1 for one or more blocks.
[0137] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart Figure 1 for one or more flows and / or blocks Figure 1 for one or more blocks.
[0138] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart Figure 1 for one or more flows and / or blocks Figure 1 for one or more blocks.
[0139] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0140] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0141] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0142] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0143] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0144] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0145] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A task control method applied to a distributed task control system, comprising: Obtaining event data generated during the execution of a distributed task; The event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task; Calculating the value of an evaluation index corresponding to the distributed task according to the event data; the evaluation index is used to evaluate the execution situation and resource usage of the distributed task; Judging whether the index value meets the rule conditions set in the rule engine; the rule conditions are used to judge whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements; If the index value meets the rule conditions, controlling the distributed task based on the control information set by the rule engine.
2. The method according to claim 1, wherein the event data includes the execution status data of the distributed task and the resource usage data of the distributed task; calculating the value of the evaluation index corresponding to the distributed task according to the event data specifically includes: Calculating the failure rate and the delay rate based on the execution status data; Calculating the resource waste rate based on the resource usage data; Obtaining a first weight value corresponding to the failure rate, a second weight value corresponding to the delay rate, and a third weight value corresponding to the resource waste rate; Calculating the comprehensive score of the distributed task based on the failure rate and the first weight value, the delay rate and the second weight value, and the resource waste rate and the third weight value.
3. The method according to claim 2, before calculating the comprehensive score of the distributed task, further comprising: Judging whether the current time is within a specified time period; the specified time period includes a time period when the quantity of tasks to be processed is greater than a specified threshold; If the current time is within the specified time period, increasing the first weight value and the second weight value to obtain an adjusted first weight value and an adjusted second weight value; Calculating the comprehensive score of the distributed task specifically includes: Calculating the comprehensive score of the distributed task according to the failure rate and the adjusted first weight value, the delay rate and the adjusted second weight value.
4. The method according to claim 1, wherein the event data includes the execution status data of the distributed task and the resource usage data of the distributed task; calculating the value of the evaluation index corresponding to the distributed task according to the event data specifically includes: Calculating the success rate based on the execution status data; Calculating the resource utilization rate based on the resource usage data; Calculating the resource efficiency index of the distributed task based on the resource utilization rate and the success rate; wherein, the resource efficiency index is positively correlated with the logarithm of the resource utilization rate; the resource efficiency index is positively correlated with the success rate.
5. According to the method described in claim 1, the event data includes the execution status data of the distributed task; the metric value includes the failure rate calculated based on the execution status data; determining whether the metric value meets the rule conditions set in the rule engine specifically includes: Determining whether the failure rate is greater than a first preset failure rate; If the metric value meets the rule conditions, then controlling the distributed task based on the control information set by the rule engine, specifically including: If the failure rate is greater than the first preset failure rate, then stop executing the distributed task.
6. According to the method described in claim 5, the event data further includes the resource usage data of the distributed task; the metric value further includes the resource utilization rate calculated based on the resource usage data; before determining whether the failure rate is greater than the first preset failure rate, it further includes: Determining whether the resource utilization rate is greater than a first resource utilization rate threshold; If the resource utilization rate is greater than the first resource utilization rate threshold, then multiply the first preset failure rate by a first preset coefficient to obtain a first adjusted failure rate; The first preset coefficient is a positive number less than 1; The determining whether the failure rate is greater than the first preset failure rate specifically includes: Determining whether the failure rate is greater than the first adjusted failure rate.
7. According to the method described in claim 6, the event data further includes the resource usage data of the distributed task; before determining whether the failure rate is greater than the first preset failure rate, it further includes: Determining whether the resource utilization rate is less than a second resource utilization rate threshold; The second resource utilization rate threshold is less than the first resource utilization rate threshold; If the resource utilization rate is less than the second resource utilization rate threshold, then multiply the first preset failure rate by a second preset coefficient to obtain a second adjusted failure rate; The first preset coefficient is greater than 1; The determining whether the failure rate is greater than the first preset failure rate specifically includes: Determining whether the failure rate is greater than the second adjusted failure rate.
8. According to the method described in claim 1, the determining whether the metric value meets the rule conditions set in the rule engine specifically includes: Determining whether the metric value of the evaluation metric is greater than a first preset metric value corresponding to the evaluation metric; The evaluation metric includes the resource occupancy rate or the number of failures; The controlling the distributed task based on the control information set by the rule engine specifically includes: If the metric value of the evaluation metric is greater than the first preset metric value corresponding to the evaluation metric, then send an alarm message to the user terminal; Determining whether the metric value is greater than a second preset metric value corresponding to the evaluation metric; the second preset metric value is greater than the first preset metric value; If the metric value is greater than the second preset metric value, then reduce the execution priority of the distributed task.
9. According to the method described in claim 8, after reducing the priority of executing the distributed task, it further includes: Determining whether the metric value is greater than a third preset metric value corresponding to the evaluation metric; The third preset index value is greater than the second preset index value; If the index value is greater than the third preset index value, the execution of the distributed task is stopped.
10. According to the method described in claim 1, the event data includes resource usage data of the distributed task; The method further includes: Determining whether the event data meets a preset first trigger condition; the first trigger condition includes at least one of the CPU usage rate exceeding a preset CPU usage rate and the memory usage rate being greater than a preset memory usage rate; the CPU usage rate and the memory usage rate are calculated based on the resource usage data of the distributed task; If the event data meets the preset first trigger condition, a first preset action is performed; the first preset action includes at least one of restricting new task scheduling, adjusting resource allocation rules, and triggering an expansion process.
11. The method according to claim 1, wherein the event data includes the execution status data of the distributed task; the method further includes: Determining whether the event data meets a preset second trigger condition; The second trigger condition includes at least one of a task request growth rate being greater than a preset request growth rate and a failure rate being greater than a second preset failure rate; the task request growth rate and the failure rate are calculated based on the execution status data of the distributed task; If the event data meets the preset second trigger condition, a second preset action is performed; the second preset action includes at least one of starting a current limiting rule and reducing the priority of selected tasks.
12. The method according to claim 1, after controlling the distributed task based on the control information set by the rule engine, further includes: Generating a first notification message and a second notification message based on the control result of the distributed task; Sending the first notification message to the server so that the server records the execution status of the distributed task based on the first notification message; Sending the second notification message to the user terminal so that the user terminal displays the control result of the distributed task to the user.
13. The method according to claim 12, wherein sending the second notification message to the user terminal specifically includes: Determining the notification level of the second notification message based on the control result; Sending the second notification message by using a notification method corresponding to the notification level.
14. The method according to claim 1, wherein obtaining the event data generated during the execution of the distributed task includes: Collecting the event data by using a polling plugin at a preset period; Or, Continuously monitoring the distributed task by using a monitoring plugin to obtain the event data; Or, Collecting the event data by using a polling plugin at a preset period and continuously monitoring the distributed task by using a monitoring plugin to obtain the event data.
15. A task control device, applied to a distributed task control system, includes: A data acquisition module, configured to acquire event data generated during the execution of a distributed task; The event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task; An index calculation module, configured to calculate the index value of the evaluation index corresponding to the distributed task according to the event data; The evaluation index is used to evaluate the execution situation and resource usage situation of the distributed task; A judgment module, configured to judge whether the index value meets the rule conditions set in the rule engine; the rule conditions are used to judge whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements; A control module, configured to, if the evaluation index value meets the rule conditions, control the distributed task based on the control information set by the rule engine.
16. A task control device, applied to a distributed task control system, includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Obtain event data generated during the execution of the distributed task; the event data includes at least one of the execution status data of the distributed task and the resource usage data of the distributed task; Calculate the index value of the evaluation index corresponding to the distributed task according to the event data; the evaluation index is used to evaluate the execution situation and resource usage situation of the distributed task; Judge whether the index value meets the rule conditions set in the rule engine; the rule conditions are used to judge whether the hardware resources occupied by the distributed task and the execution result of the distributed task meet the requirements; If the evaluation index value meets the rule conditions, control the distributed task based on the control information set by the rule engine.
17. A computer-readable storage medium, on which a computer program or instruction is stored, and the computer program or instruction is executed by a processor to implement the steps of the task control method according to any one of claims 1 to 14.