Big data task exception detection method and device, electronic equipment and storage medium

By acquiring time and duration samples of big data tasks and using a box algorithm to calculate boundary values, the problem of not being able to detect task anomalies in advance in existing technologies is solved, enabling accurate judgment of task health and early warning of potential risks.

CN118642930BActive Publication Date: 2025-11-18CHERY AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410684689.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-11-18
Estimated Expiration
2044-05-29

AI Technical Summary

Technical Problem

Existing technologies cannot detect potential anomalies in big data tasks in advance, resulting in uncertain task execution times and a potential backlog of tasks.

Method used

By obtaining the start time, end time, and cycle time of multiple target tasks, and using a pre-defined box algorithm to calculate the boundary values ​​of execution time samples and time consumption samples, the execution time and time consumption of the tasks are judged to be abnormal.

Benefits of technology

It enables early detection of potential risks in big data tasks, improves the accuracy of task health assessment, and avoids uncertainty and backlog in task execution time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118642930B_ABST
    Figure CN118642930B_ABST
Patent Text Reader

Abstract

The application relates to a big data task exception detection method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring the start time, end time and cycle time of a plurality of target tasks; acquiring a plurality of target task execution time samples based on the start time and the cycle time, acquiring a plurality of target task execution time consumption samples based on the start time and the end time, introducing the execution time samples and the execution time consumption samples into a preset box type algorithm for calculation to obtain a first boundary value of the execution time samples and a second boundary value of the execution time consumption samples; acquiring the execution time boundary value and the execution time consumption boundary value of a plurality of new big data tasks, and determining the execution time exception and the execution time consumption exception of the plurality of new big data tasks based on the execution time boundary value and the first boundary value and the execution time consumption boundary value and the second boundary value. Thus, the potential risk of the task cannot be detected, the execution time of the task is inconsistent, and the task accumulation problem is easily caused.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data anomaly detection technology, and in particular to a method, apparatus, electronic device and storage medium for detecting anomalies in big data tasks. Background Technology

[0002] With the development of technology, big data technology has been widely applied in many fields. Among them, anomaly detection for big data tasks is essential to ensure the accuracy of big data task execution.

[0003] In related technologies, the main method for judging anomalies in big data tasks is to detect abnormal task timing, such as task execution timeout, task failure, and task backlog.

[0004] However, most of the above methods only perform analysis after a substantial failure occurs during task execution, and cannot detect some potentially risky tasks. As a result, they cannot determine the health status of the task in advance, reducing the detection accuracy, which urgently needs to be addressed. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and storage medium for detecting anomalies in big data tasks, in order to solve the potential risks in related technologies that cannot detect task execution time and execution time, thus making it impossible to accurately determine the health status of tasks, and causing uncertain task execution time, which can easily lead to task backlog and other problems.

[0006] The first aspect of this application provides a method for detecting anomalies in big data tasks, comprising the following steps:

[0007] Obtain the start time, end time, and cycle time of multiple target tasks;

[0008] Based on the start time and cycle time of the multiple target tasks, multiple target task execution time samples are obtained within a first preset time period. Simultaneously, based on the start time and end time of the target tasks, multiple target task execution time samples are obtained within a second preset time period. The multiple target task execution time samples and the multiple target task execution time samples are introduced into a preset box algorithm, and the multiple target task execution time samples and the multiple target task execution time samples are calculated according to the preset box algorithm to obtain the first boundary value of the multiple target task execution time samples and the second boundary value of the multiple target task execution time samples.

[0009] Obtain the execution time boundary value and execution time boundary value of multiple new big data tasks, determine the execution time anomaly of the multiple new big data tasks based on the execution time boundary value and the first boundary value, and determine the execution time anomaly of the multiple new big data tasks based on the execution time boundary value and the second boundary value.

[0010] According to one embodiment of this application, before obtaining the start time, end time, and cycle time of the target task, the method further includes:

[0011] Retrieve multiple executed big data tasks;

[0012] Among the multiple big data tasks, identify multiple target tasks based on a preset time period.

[0013] According to one embodiment of this application, the step of determining the execution time anomalies of the new plurality of big data tasks based on the execution time boundary value and the first boundary value, and determining the execution time anomalies of the new plurality of big data tasks based on the execution time boundary value and the second boundary value, includes:

[0014] Determine whether the execution time boundary value is within the first boundary value range;

[0015] If the execution time boundary value is within the first boundary value range, then determine whether the execution time boundary value is within the second boundary value range.

[0016] According to one embodiment of this application, determining whether the execution time boundary value is within the second boundary value range includes:

[0017] If the execution time boundary value is within the second boundary value range, then it is determined that the new multiple big data are not in an abnormal state;

[0018] If the execution time boundary value is not within the second boundary value range, then the new multiple big data are determined to be in an abnormal execution time state.

[0019] According to one embodiment of this application, after determining whether the execution time boundary value is within the first boundary value range, the method further includes:

[0020] If the execution time boundary value is not within the first boundary value range, then the new multiple big data are determined to be in an abnormal execution time state.

[0021] According to the big data task anomaly detection method of this application embodiment, the start time, end time, and cycle time of multiple target tasks are obtained; multiple target task execution time samples and multiple target task execution time samples are obtained based on the start time and cycle time; these samples are then introduced into a preset box algorithm to calculate the execution time samples and execution time samples, obtaining a first boundary value for the execution time samples and a second boundary value for the execution time samples; new execution time boundary values ​​and execution time boundary values ​​of multiple big data tasks are obtained; and new execution time anomalies and execution time anomalies of multiple big data tasks are determined based on the execution time boundary values ​​and the first boundary value, and the execution time boundary values ​​and the second boundary value, respectively. This solves the potential risks of not being able to detect task execution time and execution time in related technologies, thus failing to accurately determine the health status of tasks, and causing uncertain task execution times, which can easily lead to task backlog.

[0022] A second aspect of this application provides a device for detecting anomalies in big data tasks, comprising:

[0023] The first acquisition module is used to acquire the start time, end time, and cycle time of multiple target tasks;

[0024] The calculation module is used to obtain multiple target task execution time samples within a first preset time period based on the start time and cycle time of the multiple target tasks, and simultaneously obtain multiple target task execution time samples within a second preset time period based on the start time and end time of the target tasks. The multiple target task execution time samples and the multiple target task execution time samples are introduced into a preset box algorithm, and the multiple target task execution time samples and the multiple target task execution time samples are calculated according to the preset box algorithm to obtain a first boundary value of the multiple target task execution time samples and a second boundary value of the multiple target task execution time samples.

[0025] The second acquisition module is used to acquire the execution time boundary value and execution time boundary value of multiple new big data tasks, determine the execution time anomaly of the multiple new big data tasks based on the execution time boundary value and the first boundary value, and determine the execution time anomaly of the multiple new big data tasks based on the execution time boundary value and the second boundary value.

[0026] According to one embodiment of this application, before obtaining the start time, end time, and cycle time of the target task, the first obtaining module is further configured to:

[0027] Retrieve multiple executed big data tasks;

[0028] Among the multiple big data tasks, identify multiple target tasks based on a preset time period.

[0029] According to one embodiment of this application, the second acquisition module is specifically used for:

[0030] Determine whether the execution time boundary value is within the first boundary value range;

[0031] If the execution time boundary value is within the first boundary value range, then determine whether the execution time boundary value is within the second boundary value range.

[0032] According to one embodiment of this application, the second acquisition module is specifically used for:

[0033] If the execution time boundary value is within the second boundary value range, then it is determined that the new multiple big data are not in an abnormal state;

[0034] If the execution time boundary value is not within the second boundary value range, then the new multiple big data are determined to be in an abnormal execution time state.

[0035] According to one embodiment of this application, the second acquisition module is specifically used for:

[0036] If the execution time boundary value is not within the first boundary value range, then the new multiple big data are determined to be in an abnormal execution time state.

[0037] The big data task anomaly detection device according to an embodiment of this application acquires the start time, end time, and cycle time of multiple target tasks; it acquires multiple target task execution time samples based on the start time and cycle time, and acquires multiple target task execution time samples based on the start time and end time. Simultaneously, it introduces these samples into a preset box algorithm to calculate the execution time samples and execution time samples, obtaining a first boundary value for the execution time samples and a second boundary value for the execution time samples; it acquires the execution time boundary values ​​and execution time boundary values ​​of new big data tasks, and determines the execution time anomalies and execution time anomalies of the new big data tasks based on the execution time boundary values ​​and the first boundary value, and the execution time boundary values ​​and the second boundary value, respectively. This solves the potential risks in related technologies where task execution time and execution time cannot be detected, leading to an inability to accurately determine the health status of tasks, and causing uncertain task execution times, which can easily result in task backlog.

[0038] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for detecting anomalies in big data tasks as described in the above embodiments.

[0039] A fourth aspect of this application provides a computer-readable storage medium storing computer instructions for causing the computer to perform the big data task anomaly detection method as described in the above embodiments.

[0040] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0041] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0042] Figure 1 This is a flowchart of a method for detecting anomalies in a big data task according to an embodiment of this application;

[0043] Figure 2 An internal logic diagram according to an embodiment of this application;

[0044] Figure 3 A schematic diagram of a box-type calculation model according to an embodiment of this application;

[0045] Figure 4 A schematic diagram of the normal distribution of a box-type computational model according to an embodiment of this application;

[0046] Figure 5 This is an example diagram of a big data task anomaly detection device according to an embodiment of this application;

[0047] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0048] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0049] The following describes a method, apparatus, electronic device, and storage medium for detecting anomalies in big data tasks according to embodiments of this application, with reference to the accompanying drawings. Addressing the potential risks of failing to detect task execution time and execution duration in the related technologies mentioned in the background, thus failing to accurately determine the health of tasks and leading to uncertain task execution times and task backlog, this application provides a method for detecting anomalies in big data tasks. In this method, the start time, end time, and cycle time of multiple target tasks are obtained; multiple target task execution time samples and multiple target task execution duration samples are obtained based on the start time and cycle time; these samples are then introduced into a preset binning algorithm to calculate the execution time samples and execution duration samples, obtaining a first boundary value for the execution time samples and a second boundary value for the execution duration samples; new execution time boundary values ​​and execution duration boundary values ​​of multiple big data tasks are obtained; and execution time anomalies and execution duration anomalies of the new multiple big data tasks are determined based on the execution time boundary values ​​and the first boundary value, and the execution duration boundary values ​​and the second boundary value, respectively. This solves the potential risks of not being able to detect task execution time and execution duration in related technologies, which leads to the inability to accurately determine the health status of tasks and causes uncertain task execution time, thus easily resulting in task backlog.

[0050] Specifically, Figure 1 This is a flowchart illustrating a method for detecting anomalies in big data tasks provided in an embodiment of this application.

[0051] like Figure 1 As shown, the anomaly detection method for this big data task includes the following steps:

[0052] In step S101, the start time, end time, and cycle time of multiple target tasks are obtained.

[0053] According to one embodiment of this application, before obtaining the start time, end time, and cycle time of the target task, the method further includes: obtaining multiple executed big data tasks; and determining multiple target tasks based on a preset time period among the multiple big data tasks.

[0054] The preset time can be set by those skilled in the art according to actual testing needs, and no specific limitation is made here.

[0055] Specifically, such as Figure 2As shown, to enable the detection of potential risks during the execution of big data tasks, this embodiment first obtains multiple big data tasks that have been successfully executed. After the big data tasks have been run 10 times, more than 10 target tasks from the past month are extracted. Secondly, based on the extracted more than 10 target tasks from the past month, the basic execution information of the target tasks is obtained, mainly including the start time, end time, and cycle time of the target tasks. The cycle time is the periodic interval between repeated scheduling of the target tasks. Thus, the extracted target tasks are analyzed based on their start time, end time, and cycle time, thereby providing an analytical basis for the anomaly detection of subsequent big data tasks.

[0056] In step S102, multiple target task execution time samples are obtained within a first preset time period based on the start time and cycle time of multiple target tasks. At the same time, multiple target task execution time samples are obtained within a second preset time period based on the start time and end time of the target tasks. The multiple target task execution time samples and multiple target task execution time samples are introduced into a preset box algorithm, and the multiple target task execution time samples and multiple target task execution time samples are calculated according to the preset box algorithm to obtain the first boundary value of the multiple target task execution time samples and the second boundary value of the multiple target task execution time samples.

[0057] The first preset time, the second preset time, and the preset box type algorithm can all be set by those skilled in the art according to actual testing needs, and no specific limitation is made here. The first preset time can be the same as the second preset time.

[0058] Specifically, in this embodiment of the application, after obtaining the start time, end time, and cycle time of multiple target tasks, firstly, multiple target task execution time samples can be obtained within a first preset time period based on the start time and cycle time of the multiple target tasks. For example, more than 10 target task execution time samples can be obtained from the operation of target tasks within the past month. At the same time, multiple target task execution time samples can be obtained within a second preset time period based on the start time and end time of the target tasks. For example, more than 10 target task execution time samples can be obtained from the operation of target tasks within the past month. The obtained multiple target task execution time samples and multiple target task execution time samples are introduced into a preset box algorithm, and the multiple target task execution time samples and multiple target task execution time samples are calculated according to the preset box algorithm to obtain the first boundary value of the multiple target task execution time samples and the second boundary value of the multiple target task execution time samples.

[0059] Specifically, such as Figure 3 and Figure 4As shown, this application calculates the execution time samples and execution duration samples of multiple target tasks based on a preset binning algorithm. In the preset binning algorithm, the upper quartile U, the lower quartile L, the upper bound (first boundary value) and the lower bound (second boundary value) are defined respectively. The execution time samples and execution duration samples of multiple target tasks are imported into the model for calculation. It is assumed that the data from the upper quartile to the lower quartile accounts for 50% of the total data, and the data from each quartile to the upper and lower bounds is 25%, thus forming the effect of a normal distribution.

[0060] In order to ensure the normal distribution of the first boundary value and the second boundary value, the embodiment of this application defines the first boundary value as U+1.5LQR (Linear Quadratic Regulator) and the second boundary value as L-1.5LQR, where LQR = UL. In order to ensure the requirement of normal distribution and improve the calculation accuracy of the boundary value, the embodiment of this application uses 1.5 times for calculation.

[0061] For example, such as Figure 4 As shown, according to the requirements of the normal distribution, approximately 68% of the total time data is within one standard deviation (<1σ) of the mean (μ) (both sides), approximately 95% of the total time data is within two standard deviations (2σ) of the mean (μ) (both sides), approximately 99.7% of the time data is within three standard deviations (<3σ) of the mean (μ) (both sides), and the remaining 0.3% of the time data is outside three standard deviations (>3σ) of the mean (μ) (both sides). Q1 and Q3 are located at -0.675σ and +0.675σ of the mean, respectively.

[0062] If "1" is used as the first boundary value and the second boundary value for calculation, then

[0063] Second boundary value = Q1-1*LQR = q1-1*(q3-q1) = -0.675σ-1*(0.675-[-0.675])

[0064] σ=-0.675σ-1*1.35σ=-2.025σ;

[0065] First boundary value=Q3+1*LQR=Q3+1*(Q3-Q1)=0.675σ+1*(0.675-[-0.675])σ=0.675σ+1*1.35σ=2.025σ.

[0066] Therefore, when using 1, according to the LQR method, any time data that exceeds 2.025σ of the average value (μ) should be considered an outlier on either side. This would make the decision range too exclusive, and it would also mean that nearly 5% of the valid data would be considered outliers. Therefore, this application embodiment does not use "1" for calculation.

[0067] If "2" is used as the first boundary value and the second boundary value for calculation, then

[0068] Second boundary value = Q1-2*LQR = q1-2*(q3-q1) = -0.675σ-2*(0.675-[-0.675])σ = -0.675σ-2*1.35σ = -3.375σ.

[0069] First boundary value = Q3 + 2 * LQR = Q3 + 2 * (Q3 - Q1) = 0.675σ + 2 * (0.675 - [-0.675])σ = 0.675σ + 2 * 1.35σ = 3.375σ.

[0070] Therefore, when using 2, according to the LQR method, any time data that exceeds 3.375σ of the average value (μ) should be considered an outlier. This would make the decision range too broad, meaning that even if there are abnormal situations or data, they will not be defined as outliers. Therefore, this application embodiment does not use "2" for calculation.

[0071] If "1.5" is used as the first boundary value and the second boundary value for calculation, then

[0072] Second boundary value = q1 - 1.5 * LQR = q1 - 1.5 * (q3 - q1) = -0.675σ - 1.5 * (0.675 - [-0.675])σ = -0.675σ - 1.5 * 1.35σ = -2.7σ.

[0073] First boundary value = q3 + 1.5 * LQR = q3 + 1.5 * (q3 - q1) = 0.675σ + 1.5 * (0.675 - [-0.675])σ = 0.675σ + 1.5 * 1.35σ = 2.7σ.

[0074] Therefore, when using 1.5, according to the LQR method, if any data exceeds 2.7σ of the mean (μ) at any time, it should be considered an anomaly on either side. Thus, the decision range 3σ = 99.72% of the data that is closest to the normal distribution can be obtained. Therefore, the embodiments of this application use "1.5" for calculation to improve the calculation accuracy.

[0075] It should be noted that the first and second boundary values ​​calculated by the preset binning algorithm in this application are time intervals. Therefore, by comparing the execution time of each subsequent task with the time interval, it can be determined whether a time anomaly has occurred in the task. At the same time, by using the preset binning algorithm and the relevant logic added for anomaly detection, some potential risks in the execution time of big data scheduling tasks can be accurately revealed.

[0076] In step S103, the execution time boundary value and execution time boundary value of the new multiple big data tasks are obtained. The execution time anomaly of the new multiple big data tasks is determined based on the execution time boundary value and the first boundary value, and the execution time anomaly of the new multiple big data tasks is determined based on the execution time boundary value and the second boundary value.

[0077] According to one embodiment of this application, determining the execution time anomalies of multiple new big data tasks based on execution time boundary values ​​and a first boundary value, and determining the execution time anomalies of multiple new big data tasks based on execution time boundary values ​​and a second boundary value, includes: determining whether the execution time boundary value is within the first boundary value range; if the execution time boundary value is within the first boundary value range, then determining whether the execution time boundary value is within the second boundary value range.

[0078] According to one embodiment of this application, determining whether the execution time boundary value is within the second boundary value range includes: if the execution time boundary value is within the second boundary value range, then determining that the new multiple big data are not in an abnormal state; if the execution time boundary value is not within the second boundary value range, then determining that the new multiple big data are in an abnormal execution time state.

[0079] According to one embodiment of this application, after determining whether the execution time boundary value is within the first boundary value range, the method further includes: if the execution time boundary value is not within the first boundary value range, then determining that the new multiple big data are in an abnormal execution time state.

[0080] Specifically, this application can use the first boundary value of the multiple target task execution time samples and the second boundary value of the multiple target task execution time samples obtained above as the detection basis for the execution of new multiple big data tasks. Thus, this application obtains the execution time boundary value and execution time boundary value of the new multiple big data tasks, compares the execution time boundary value of the new multiple big data tasks with the first boundary value, and compares the execution time boundary value of the new multiple big data tasks with the second boundary value, thereby detecting whether there are any abnormalities in the execution time and execution time of the new multiple big data tasks.

[0081] Specifically, firstly, the execution time boundary values ​​of multiple new big data tasks are compared with the first boundary value, and it is determined whether the execution time boundary value is within the first boundary value range. If the execution time boundary value is not within the first boundary value range, it indicates that the new big data task has an execution time anomaly, and the new big data task is determined to be in an abnormal state. If the execution time boundary value is within the first boundary value range, it indicates that the new big data task does not have an execution time anomaly. At the same time, it is further determined whether the execution time boundary value is within the second boundary value range.

[0082] Furthermore, if the execution time boundary value is within the second boundary value range, it is determined that the new big data task does not have an execution time anomaly, and thus it can be determined that the new big data task is not in an abnormal state; if the execution time boundary value is not within the second boundary value range, it indicates that the new big data task does not have an execution time anomaly, but it does have an execution time anomaly, and thus it can also be determined that the new big data task is in an abnormal state.

[0083] It should be noted that, according to the analysis box algorithm model, the embodiments of this application determine the risks of multiple new big data tasks and issue timely warnings through the task monitoring system. At the same time, the calculated first boundary value and second boundary value are updated in real time every day, and corresponding diagnoses can be performed every day to ensure the accuracy of the data. In addition, after obtaining the diagnostic results of multiple new big data tasks, this application can use programming languages ​​such as JAVA and big data clusters to optimize the multiple new big data tasks.

[0084] Therefore, this application embodiment can determine whether the execution time and execution time of the new multiple big data tasks are abnormal based on the judgment of the execution time and execution time consumption of the new multiple big data tasks. Thus, it can determine whether the new multiple big data tasks have potential risks and further determine the health status of the new multiple big data tasks. At the same time, when the new multiple big data tasks have potential risks, it can provide early warning and make corrections. For example, when the new multiple big data tasks have potential risks, it can send warning information to the target technical personnel by calling a robot, i.e., a webhook. After receiving the warning information, the target technical personnel can optimize the new multiple big data tasks in a targeted manner according to the diagnosed potential risks, thereby eliminating some potential risks in the execution process of big data tasks to a large extent at the source.

[0085] According to the big data task anomaly detection method of this application embodiment, the start time, end time, and cycle time of multiple target tasks are obtained; multiple target task execution time samples and multiple target task execution time samples are obtained based on the start time and cycle time; these samples are then introduced into a preset box algorithm to calculate the execution time samples and execution time samples, obtaining a first boundary value for the execution time samples and a second boundary value for the execution time samples; new execution time boundary values ​​and execution time boundary values ​​of multiple big data tasks are obtained; and new execution time anomalies and execution time anomalies of multiple big data tasks are determined based on the execution time boundary values ​​and the first boundary value, and the execution time boundary values ​​and the second boundary value, respectively. This solves the potential risks of not being able to detect task execution time and execution time in related technologies, thus failing to accurately determine the health status of tasks, and causing uncertain task execution times, which can easily lead to task backlog.

[0086] Next, referring to the accompanying drawings, a detection device for big data task anomalies according to an embodiment of this application is described.

[0087] Figure 5 This is a block diagram of a big data task anomaly detection device according to an embodiment of this application.

[0088] like Figure 5 As shown, the big data task anomaly detection device 10 includes: a first acquisition module 100, a calculation module 200, and a second acquisition module 300.

[0089] The first acquisition module 100 is used to acquire the start time, end time, and cycle time of multiple target tasks.

[0090] The calculation module 200 is used to obtain multiple target task execution time samples within a first preset time period based on the start time and cycle time of multiple target tasks, and simultaneously obtain multiple target task execution time samples within a second preset time period based on the start time and end time of the target tasks. The multiple target task execution time samples and multiple target task execution time samples are introduced into a preset box algorithm, and the multiple target task execution time samples and multiple target task execution time samples are calculated according to the preset box algorithm to obtain the first boundary value of the multiple target task execution time samples and the second boundary value of the multiple target task execution time samples.

[0091] The second acquisition module 300 is used to acquire the execution time boundary value and execution time boundary value of multiple new big data tasks, determine the execution time anomaly of multiple new big data tasks based on the execution time boundary value and the first boundary value, and determine the execution time anomaly of multiple new big data tasks based on the execution time boundary value and the second boundary value.

[0092] According to one embodiment of this application, before obtaining the start time, end time, and cycle time of the target task, the first acquisition module 100 is further configured to:

[0093] Retrieve multiple executed big data tasks;

[0094] Identify multiple target tasks within a preset timeframe from a range of big data tasks.

[0095] According to one embodiment of this application, the second acquisition module 300 is specifically used for:

[0096] Determine whether the execution time boundary value is within the first boundary value range;

[0097] If the execution time boundary value is within the first boundary value range, then determine whether the execution time boundary value is within the second boundary value range.

[0098] According to one embodiment of this application, the second acquisition module 300 is specifically used for:

[0099] If the execution time boundary value is within the second boundary value range, then the new multiple big data are determined not to be in an abnormal state;

[0100] If the execution time boundary value is not within the second boundary value range, then the new multiple big data are determined to be in an abnormal execution time state.

[0101] According to one embodiment of this application, the second acquisition module 300 is specifically used for:

[0102] If the execution time boundary value is not within the first boundary value range, then the new multiple big data are determined to be in an abnormal execution time state.

[0103] The big data task anomaly detection device according to an embodiment of this application acquires the start time, end time, and cycle time of multiple target tasks; it acquires multiple target task execution time samples based on the start time and cycle time, and acquires multiple target task execution time samples based on the start time and end time. Simultaneously, it introduces these samples into a preset box algorithm to calculate the execution time samples and execution time samples, obtaining a first boundary value for the execution time samples and a second boundary value for the execution time samples; it acquires the execution time boundary values ​​and execution time boundary values ​​of new big data tasks, and determines the execution time anomalies and execution time anomalies of the new big data tasks based on the execution time boundary values ​​and the first boundary value, and the execution time boundary values ​​and the second boundary value, respectively. This solves the potential risks in related technologies where task execution time and execution time cannot be detected, leading to an inability to accurately determine the health status of tasks, and causing uncertain task execution times, which can easily result in task backlog.

[0104] Figure 6A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:

[0105] The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.

[0106] When the processor 602 executes the program, it implements the big data task anomaly detection method provided in the above embodiments.

[0107] Furthermore, electronic devices also include:

[0108] Communication interface 603 is used for communication between memory 601 and processor 602.

[0109] The memory 601 is used to store computer programs that can run on the processor 602.

[0110] The memory 601 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0111] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0112] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.

[0113] The processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0114] This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described method for detecting anomalies in big data tasks.

[0115] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0116] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0117] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0118] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0119] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0120] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiments.

[0121] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0122] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for detecting anomalies in big data tasks, characterized in that, Includes the following steps: Obtain the start time, end time, and cycle time of multiple target tasks; Based on the start time and cycle time of the multiple target tasks, multiple target task execution time samples are obtained within a first preset time period. Simultaneously, based on the start time and end time of the target tasks, multiple target task execution time samples are obtained within a second preset time period. The multiple target task execution time samples and the multiple target task execution time samples are introduced into a preset box algorithm, and the multiple target task execution time samples and the multiple target task execution time samples are calculated according to the preset box algorithm to obtain the first boundary value of the multiple target task execution time samples and the second boundary value of the multiple target task execution time samples. Obtain the execution time boundary value and execution time boundary value of multiple new big data tasks, determine the execution time anomaly of the multiple new big data tasks based on the execution time boundary value and the first boundary value, and determine the execution time anomaly of the multiple new big data tasks based on the execution time boundary value and the second boundary value.

2. The method according to claim 1, characterized in that, Before obtaining the start time, end time, and cycle time of the target task, the process also includes: Retrieve multiple executed big data tasks; Among the multiple big data tasks, identify multiple target tasks based on a preset time period.

3. The method according to claim 1, characterized in that, The step of determining the execution time anomalies of the new multiple big data tasks based on the execution time boundary value and the first boundary value, and determining the execution time anomalies of the new multiple big data tasks based on the execution time boundary value and the second boundary value, includes: Determine whether the execution time boundary value is within the first boundary value range; If the execution time boundary value is within the first boundary value range, then determine whether the execution time boundary value is within the second boundary value range.

4. The method according to claim 3, characterized in that, The step of determining whether the execution time boundary value is within the second boundary value range includes: If the execution time boundary value is within the second boundary value range, then it is determined that the new multiple big data are not in an abnormal state; If the execution time boundary value is not within the second boundary value range, then the new multiple big data are determined to be in an abnormal execution time state.

5. The method according to claim 3, characterized in that, After determining whether the execution time boundary value is within the first boundary value range, the method further includes: If the execution time boundary value is not within the first boundary value range, then the new multiple big data are determined to be in an abnormal execution time state.

6. A device for detecting anomalies in big data tasks, characterized in that, include: The first acquisition module is used to acquire the start time, end time, and cycle time of multiple target tasks; The calculation module is used to obtain multiple target task execution time samples within a first preset time period based on the start time and cycle time of the multiple target tasks, and simultaneously obtain multiple target task execution time samples within a second preset time period based on the start time and end time of the target tasks. The multiple target task execution time samples and the multiple target task execution time samples are introduced into a preset box algorithm, and the multiple target task execution time samples and the multiple target task execution time samples are calculated according to the preset box algorithm to obtain a first boundary value of the multiple target task execution time samples and a second boundary value of the multiple target task execution time samples. The second acquisition module is used to acquire the execution time boundary value and execution time boundary value of multiple new big data tasks, determine the execution time anomaly of the multiple new big data tasks based on the execution time boundary value and the first boundary value, and determine the execution time anomaly of the multiple new big data tasks based on the execution time boundary value and the second boundary value.

7. The apparatus according to claim 6, characterized in that, Before acquiring the start time, end time, and cycle time of the target task, the first acquisition module is further configured to: Retrieve multiple executed big data tasks; Among the multiple big data tasks, identify multiple target tasks based on a preset time period.

8. The apparatus according to claim 6, characterized in that, The second acquisition module is specifically used for: Determine whether the execution time boundary value is within the first boundary value range; If the execution time boundary value is within the first boundary value range, then determine whether the execution time boundary value is within the second boundary value range.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the method for detecting anomalies in big data tasks as described in any one of claims 1-5.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method for detecting anomalies in big data tasks as described in any one of claims 1-5.

Citation Information

Patent Citations

  • KPI anomaly detection method and device, computing equipment and computer storage medium

    CN114095337A

  • Task data processing method and device, electronic equipment and storage medium

    CN114238055A