A method and device for abnormal task warning
By constructing and decomposing the running time matrix of each task in the task group, and performing correlation analysis with preset thresholds, the problem of high missed and false alarms of task operation timeout exception alarms in the existing technology is solved, and more accurate abnormal detection and alarms are achieved.
Patent Information
- Application Number
- CN202111574740.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The prior art requires a large amount of operation and maintenance labor when detecting and alarm tasks running timeout abnormally, resulting in frequent false alarm and false alarms.
By constructing a matrix of the current running time and historical running time of each task in the task group, the matrix decomposition is performed to determine the running time estimate matrix, and correlation analysis is performed based on the preset threshold, to determine whether there is an operation timeout exception in the task group and alarm.
This method can more accurately determine whether there are tasks in the task group that have run timeout abnormalities, reduce the missed and false alarm rates, reduce the dependence on operation and maintenance manpower, and improve the accuracy and efficiency of alarms.
Smart Images

Figure CN114328095B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of financial technology (Fintech), and in particular to a method and device for abnormal task alarm. Background Art
[0002] With the development of computer technology, more and more technologies are applied in the financial field. Traditional finance is gradually transforming into financial technology. However, due to the security and real-time requirements of the financial industry, higher requirements are also placed on technology. In the field of financial big data, the calculation, analysis and processing of big data are often composed of multiple task processing units, and each task processing unit completes its own data processing logic. Unlike general data processing tasks, these big data task processing units usually have strong dependencies. These dependencies generally rely on the processing order of data. For example, after the upstream task completes the output data, the downstream task can continue to execute after obtaining the data output by the upstream task. In order to meet this requirement, a scheduling system is usually used to run and manage these task processing units. The scheduling system regularly checks whether the task meets the running requirements, including time requirements and upstream dependency requirements. At the same time, the scheduling system needs to perform some specific operations when the task fails, such as retry or alarm. Unlike general data processing tasks, the big data tasks executed on the scheduling system need to rely on many related big data systems. It is inevitable that each system will have exceptions during operation, which will lead to task failure and timeout. Based on this, in order to ensure that each task can be processed normally and smoothly, how to detect and alarm task anomalies has become an urgent problem to be solved by the scheduling system.
[0003] Existing task anomaly detection and alarm solutions are usually based on scanning log files. By merging and unifying the errors recorded in the log files, frequent alarms for the same error are avoided. Specifically, for example, after obtaining the errors from the log files, the errors are clustered, the alarm summaries within a period of time are sorted out, and then the alarms are issued. The main purpose of this clustering method is to classify the errors. In this way, after the classification process, the operation and maintenance personnel can more conveniently classify and handle task anomalies. Alternatively, the processing records of historical alarms in the log files are first annotated, and based on these annotated data, the sorting model is trained using a multi-feature fusion method. Then the trained model is used to sort the alarm streams arriving online in real time, and the sorting results are used as the severity classification results. In this way, after such processing, the operation and maintenance personnel can prioritize the problems according to their severity, greatly improving the operation and maintenance efficiency. However, the existing solution is mainly for processing deterministic task operation anomalies, while uncertain task operation anomalies require a certain degree of operation and maintenance manpower to handle. At the same time, for uncertain task operation anomalies, manually setting the task operation anomaly alarm time for detection requires relying on the experience of the operation and maintenance personnel, which is highly subjective. Due to the different experiences of different operation and maintenance personnel, the set abnormal alarm time is also different, which leads to frequent false alarms and / or missed alarms for uncertain task operation anomalies.
[0004] In summary, there is an urgent need for a task abnormality alarm method to effectively reduce the missed alarm rate and false alarm rate of task operation timeout abnormality alarm. Summary of the invention
[0005] The embodiment of the present invention provides a task abnormality alarm method and device, which are used to effectively reduce the missed alarm rate and false alarm rate of task operation timeout abnormality alarm.
[0006] In a first aspect, an embodiment of the present invention provides a method for abnormal task warning, comprising:
[0007] When an abnormality detection request for any task group is detected, task information of the task group is obtained from a task database; the task information is used to indicate the current running time of each of the k first tasks in the task group;
[0008] Constructing a first running time matrix according to the current running time of each of the k first tasks;
[0009] Acquire the historical running time of each of the k first tasks in the first preset historical period from the task database, and construct a second running time matrix according to the historical running time of each of the k first tasks in the first preset historical period;
[0010] Performing matrix decomposition on the second running time matrix to determine a first running time estimation value matrix for characterizing a normal completion running process of the k first tasks;
[0011] Determine a first correlation value between the first running time matrix and the first running time estimate value matrix, and determine whether the first correlation value is less than a first preset threshold; the first preset threshold is any preset threshold randomly selected from a preset threshold interval determined based on the historical running time of each of the plurality of second tasks within a second preset historical period;
[0012] When the first correlation value is less than the first preset threshold, it is determined that at least one first task has a running timeout exception, and a timeout exception alarm is issued for the at least one first task.
[0013] In the above technical scheme, since the existing technical scheme detects the abnormal operation of uncertain tasks by manually setting the abnormal alarm time of task operation, it needs to rely on the experience of the operation and maintenance personnel, which is relatively subjective. Therefore, due to the different experiences of different operation and maintenance personnel, the false alarm and / or missed alarm of the abnormal operation of uncertain tasks often occur. Based on this, the technical scheme in the present invention determines the running time estimation value matrix (i.e., the running time expectation value matrix) for each task group by the historical running time of each task in the task group in each time period within the preset time period. The running time estimation value matrix determined in this way is more in line with reality and more in line with the actual running conditions of each task in the task group, and the running time estimation value matrix is used as a benchmark for judging whether the running of each task is overtime, so that it can be more truly and accurately determined whether there is a running timeout abnormality in the task group. Specifically, when an abnormal detection request for any task group is detected, the task information of the task group is obtained from the task database, and the first running time matrix can be constructed by the current running time of each of the k first tasks. Then, the historical running time of each of the k first tasks in the first preset historical period is obtained from the task database, and the second running time matrix can be constructed through the historical running time of each of the k first tasks in the first preset historical period, and the first running time estimation value matrix can be accurately determined by performing matrix decomposition on the second running time matrix. Then, when it is determined that the first correlation value between the first running time matrix and the first running time estimation value matrix is less than the first preset threshold value, it is determined that at least one of the k first tasks has a running timeout exception, and a timeout exception alarm is issued for the at least one first task. In this way, the running time estimation value matrix for each task determined by the scheme through the historical running time of each task in the task group in each time period within the preset historical time period is more in line with reality and more in line with the actual operation status of the task group. Therefore, by comparing the first correlation value between the running time estimation value matrix and the first running time matrix with the first preset threshold, it is possible to more accurately determine whether there are tasks in the task group that have running timeout exceptions, thereby effectively avoiding the high missed alarm rate and false alarm rate of task running timeout exception alarms due to manually setting the abnormal alarm time, thereby effectively reducing the missed alarm rate and false alarm rate of task running timeout exception alarms.
[0014] Optionally, constructing a first running time matrix according to the current running time of each of the k first tasks includes:
[0015] Normalizing the current running time of each of the k first tasks to obtain k normalized current running times;
[0016] Constructing the first running time matrix through the k normalized current running times;
[0017] The second running time matrix is constructed by using the historical running time of each of the k first tasks in the first preset historical period, including:
[0018] For each sub-period within the first preset historical period, an initial second running time matrix is constructed using the historical running time of each of the k first tasks belonging to the sub-period as a matrix column and the k first tasks as a matrix row;
[0019] For the k matrix values in each column of the initial second running time matrix, normalize the k matrix values in the column to obtain k normalized matrix values in the column;
[0020] The second runtime matrix is constructed by using the k normalized matrix values in each column.
[0021] In the above technical solution, since the running times of different tasks may be quite different, with some tasks having a long running time and some tasks having a short running time, the dimensional data can be converted into dimensionless data by normalizing the current running time of the k first tasks and the historical running time of the k first tasks belonging to each sub-period in the first preset historical period, that is, the running times of different tasks are normalized to the same dimension for corresponding processing, such as mapping to the interval [0,1] or [-1,1]. This facilitates the subsequent timely and accurate data calculation and processing (such as comparison between data or calculation of correlation between data, etc.) under the same dimension, and can avoid large errors in data calculation and processing due to different dimensions of data.
[0022] Optionally, performing matrix decomposition on the second running time matrix to determine a first running time estimation value matrix for characterizing a normal completion running process of the k first tasks includes:
[0023] Decomposing the second runtime matrix by a singular value decomposition algorithm to determine a plurality of singular values;
[0024] The multiple singular values are compared to determine a maximum singular value, and a left singular matrix corresponding to the maximum singular value is determined as the first runtime estimation value matrix.
[0025] In the above technical scheme, if there are many tasks configured for the calculation, analysis or processing of a certain type of big data, the constructed running time matrix is high-dimensional, which is not conducive to the subsequent calculation of the correlation value between the current running time matrix of each task (i.e., the first running time matrix) and the historical running time matrix of each task in the preset historical period (i.e., the second running time matrix), which takes a lot of time, resulting in low detection efficiency for task running timeout anomalies. Therefore, the scheme uses the structural characteristics of the second running time matrix to perform low-rank estimation on the second running time matrix (i.e., use a matrix with a lower rank to approximate the original matrix), and maps the second running time matrix from a high-dimensional space to a low-dimensional space. Specifically, through the singular value decomposition algorithm, the error between the second running time matrix and the first running time estimation value matrix is minimized, thereby determining the first running time estimation value matrix, that is, by performing matrix decomposition on the second running time matrix, and shrinking the singular values of the second running time matrix according to the rank, some singular values can be set to zero, thereby achieving the purpose of low-rank estimation, and the first running time estimation value matrix can be calculated. In addition, since the big data platform supports the singular value decomposition algorithm, the practicability of the solution in practical applications can be effectively ensured. At the same time, the singular value decomposition algorithm can complete the matrix low-rank estimation in a timely and accurate manner, thereby avoiding the task running timeout detection process taking too long due to placing the matrix low-rank estimation operation on an ordinary server.
[0026] Optionally, performing matrix decomposition on the second runtime matrix by a singular value decomposition algorithm to determine a plurality of singular values includes:
[0027] Converting the second runtime matrix into a low-rank matrix and an error matrix; each error value in the error matrix conforms to a normal distribution;
[0028] The low-rank matrix is decomposed by the singular value decomposition algorithm to determine the multiple singular values.
[0029] In the above technical solution, due to the structural characteristics of the second running time, the second running time matrix is actually composed of the expected value matrix of the task running time and the error value matrix of the normal distribution with zero mean. Among them, the fluctuation of the task running time and the error caused by the artificial setting of the abnormal alarm time are fitted to the normal distribution, so that the real scene can be better restored when performing the operation timeout abnormal analysis. Therefore, the second running time matrix is first converted into a low-rank matrix (that is, the expected value matrix of the task running time) and an error matrix, but the low-rank matrix and the error matrix are unknown, so it is necessary to estimate the estimated value matrix of the approximate low-rank matrix through the matrix low-rank estimation method, that is, to perform matrix decomposition on the low-rank matrix through the singular value decomposition algorithm, so as to determine multiple singular values, and determine the estimated value matrix of the approximate low-rank matrix through the multiple singular values, and then it can be convenient to judge whether there is a timeout abnormality in the task operation more accurately based on the estimated value matrix.
[0030] Optionally, the k first tasks are tasks that are currently running and are not marked as abnormal; and the k first tasks have correlations that meet set requirements.
[0031] In the above technical solution, in order to reduce the impact of tasks that are not currently running or marked as abnormal on whether there are tasks in the current task group that have run timeout abnormalities, the tasks that are not currently running or marked as abnormal are screened out, thereby improving the accuracy of judging whether there are tasks in the current task group that have run timeout abnormalities. In addition, since the calculation, analysis or processing of a type of big data usually involves multiple tasks, and the multiple tasks have a certain dependency (i.e., have a certain correlation), multiple tasks with a certain correlation (i.e., correlation that meets the set requirements) are integrated into the same task group in advance, so that it is convenient to detect in the form of a group whether there is a run timeout abnormality in the calculation, analysis or processing of such big data, that is, to detect in parallel whether there is a run timeout abnormality in each task in the task group, so that it can more accurately detect in which task link the run timeout abnormality specifically occurs, and to a certain extent, it can improve the detection efficiency of task run timeout abnormalities.
[0032] Optionally, the preset threshold interval range is determined by:
[0033] According to the Monte Carlo simulation method, m unrelated second tasks that all conform to the normal distribution are regarded as a task group, and a plurality of different second preset thresholds are set;
[0034] Obtaining the historical running time of each of the m second tasks within the second preset historical period, and constructing a third running time matrix according to the historical running time of each of the m second tasks within the second preset historical period;
[0035] Performing matrix decomposition on the third running time matrix to determine a second running time estimation value matrix for characterizing a normal completion running process of the m second tasks;
[0036] Set, for each second preset threshold, to run each of the m second tasks multiple times in the current period, and construct a fourth running time matrix according to the running time of each of the m second tasks in each running;
[0037] Determine a second correlation value between the fourth running time matrix and the second running time estimation value matrix, and determine whether there is a false alarm and / or missed alarm of running timeout among the m second tasks by determining whether the second correlation value is less than the second preset threshold, until multiple runs of each of the m second tasks in the current time period are traversed and completed at the second preset threshold, thereby determining the running timeout missed alarm rate and running timeout false alarm rate of each of the m second tasks running multiple times in the current time period at the second preset threshold;
[0038] Generate a Monte Carlo simulation graph by using the multiple different second preset thresholds and the operation timeout missed alarm rates and the operation timeout false alarm rates corresponding to the multiple different second preset values;
[0039] Through the Monte Carlo simulation diagram, the second preset thresholds corresponding to the operation timeout omission rate being less than or equal to the first set value and the operation timeout false alarm rate being less than or equal to the second set value are determined, and based on the second preset thresholds, a preset threshold interval range for detecting whether the task operation has timed out is constructed.
[0040] In the above technical scheme, in order to more accurately judge whether there is a timeout exception in the task operation, it is necessary to determine a preset threshold interval range that can more truly and accurately judge whether there is a timeout exception in the task operation, so as to randomly select a preset threshold from the preset threshold interval range to judge whether there is a timeout exception in the task operation, so that the flexibility of selecting the preset threshold is higher. Moreover, in order to further achieve a lower operation timeout false alarm rate and operation timeout missed alarm rate, the scheme needs to set a reasonable preset threshold. If the preset threshold is set too large, although the operation timeout missed alarm rate can be reduced, it will cause the operation timeout false alarm rate to increase significantly; if the preset threshold is set too small, although the operation timeout false alarm rate can be reduced, it will increase the operation timeout missed alarm rate. Therefore, the scheme simulates the task operation based on m completely unrelated and normal distribution second tasks according to the Monte Carlo simulation method, thereby constructing a Monte Carlo simulation diagram, and through the Monte Carlo simulation diagram, the preset threshold interval range with low overall task operation timeout false alarm rate and operation timeout missed alarm rate can be intuitively and clearly obtained.
[0041] Optionally, after determining that at least one of the first tasks has a running timeout exception, the method further includes:
[0042] For each of the k first tasks, determining a deviation between a current running time of the first task and a running time estimate of the first task in the first running time estimate matrix;
[0043] sorting the deviations corresponding to the k first tasks in descending order, and determining the first i tasks in the sorting order as tasks with a running timeout exception;
[0044] The first task ranked in the first i positions is sent to an exception handling personnel for manual processing, and the first task ranked in the first i positions is marked as an exception.
[0045] In the above technical solution, for each first task, by calculating the deviation between the current running time of the first task and the running time estimate of the first task in the first running time estimate matrix, the deviation corresponding to the first task can be determined, so that the deviations corresponding to the first tasks can be sorted in order from large to small, and the first tasks ranked in the first i can be accurately determined to have a running timeout exception, so that the inaccurate task running abnormality alarm time set due to the manual estimation of the expected value of the task running time can be avoided, and the first tasks ranked in the first i are sent to the abnormality handling personnel for corresponding manual processing, so that the specific problem of the running timeout exception of the first tasks ranked in the first i can be accurately located through manual processing, and the specific problem can be handled by taking corresponding solutions. At the same time, the first tasks ranked in the first i are marked as abnormal, so as to avoid the i first tasks marked as abnormal from interfering with the running timeout exception judgment of the task group when the running timeout exception judgment is performed next time for the task group to which the first tasks ranked in the first i belong, so as to effectively ensure that each running timeout exception judgment for any task group can be performed accurately and normally.
[0046] In a second aspect, an embodiment of the present invention further provides a task abnormality alarm device, comprising:
[0047] an acquisition unit, configured to acquire task information of any task group from a task database when an abnormality detection request for the task group is detected; the task information is used to indicate the current running time of each of the k first tasks in the task group;
[0048] A processing unit is used to construct a first running time matrix through the current running time of each of the k first tasks; obtain the historical running time of each of the k first tasks within a first preset historical period from the task database, and construct a second running time matrix through the historical running time of each of the k first tasks within the first preset historical period; perform matrix decomposition on the second running time matrix to determine a first running time estimation value matrix used to characterize the normal completion of the running process of the k first tasks; determine a first correlation value between the first running time matrix and the first running time estimation value matrix, and determine whether the first correlation value is less than a first preset threshold; the first preset threshold is any preset threshold randomly selected from a preset threshold interval determined based on the historical running time of each of the multiple second tasks within the second preset historical period; when the first correlation value is less than the first preset threshold, it is determined that at least one of the first tasks has a running timeout exception, and a timeout exception alarm is issued for the at least one first task.
[0049] Optionally, the processing unit is specifically configured to:
[0050] Normalizing the current running time of each of the k first tasks to obtain k normalized current running times;
[0051] Constructing the first running time matrix through the k normalized current running times;
[0052] The processing unit is specifically used for:
[0053] For each sub-period within the first preset historical period, an initial second running time matrix is constructed using the historical running time of each of the k first tasks belonging to the sub-period as a matrix column and the k first tasks as a matrix row;
[0054] For the k matrix values in each column of the initial second running time matrix, normalize the k matrix values in the column to obtain k normalized matrix values in the column;
[0055] The second runtime matrix is constructed by using the k normalized matrix values in each column.
[0056] Optionally, the processing unit is specifically configured to:
[0057] Decomposing the second runtime matrix by a singular value decomposition algorithm to determine a plurality of singular values;
[0058] The multiple singular values are compared to determine a maximum singular value, and a left singular matrix corresponding to the maximum singular value is determined as the first runtime estimation value matrix.
[0059] Optionally, the processing unit is specifically configured to:
[0060] Converting the second runtime matrix into a low-rank matrix and an error matrix; each error value in the error matrix conforms to a normal distribution;
[0061] The low-rank matrix is decomposed by the singular value decomposition algorithm to determine the multiple singular values.
[0062] Optionally, the k first tasks are tasks that are currently running and are not marked as abnormal; and the k first tasks have correlations that meet set requirements.
[0063] Optionally, the processing unit is specifically configured to:
[0064] According to the Monte Carlo simulation method, m unrelated second tasks that all conform to the normal distribution are regarded as a task group, and a plurality of different second preset thresholds are set;
[0065] Obtaining the historical running time of each of the m second tasks within the second preset historical period, and constructing a third running time matrix according to the historical running time of each of the m second tasks within the second preset historical period;
[0066] Performing matrix decomposition on the third running time matrix to determine a second running time estimation value matrix for characterizing a normal completion running process of the m second tasks;
[0067] Set, for each second preset threshold, to run each of the m second tasks multiple times in the current period, and construct a fourth running time matrix according to the running time of each of the m second tasks in each running;
[0068] Determine a second correlation value between the fourth running time matrix and the second running time estimation value matrix, and determine whether there is a false alarm and / or missed alarm of running timeout among the m second tasks by determining whether the second correlation value is less than the second preset threshold, until multiple runs of each of the m second tasks in the current time period are traversed and completed at the second preset threshold, thereby determining the running timeout missed alarm rate and running timeout false alarm rate of each of the m second tasks running multiple times in the current time period at the second preset threshold;
[0069] Generate a Monte Carlo simulation graph by using the multiple different second preset thresholds and the operation timeout missed alarm rates and the operation timeout false alarm rates corresponding to the multiple different second preset values;
[0070] Through the Monte Carlo simulation diagram, the second preset thresholds corresponding to the operation timeout omission rate being less than or equal to the first set value and the operation timeout false alarm rate being less than or equal to the second set value are determined, and based on the second preset thresholds, a preset threshold interval range for detecting whether the task operation has timed out is constructed.
[0071] Optionally, the processing unit is further configured to:
[0072] After determining that at least one of the first tasks has a running timeout exception, determining, for each of the k first tasks, a deviation between a current running time of the first task and a running time estimate of the first task in the first running time estimate matrix;
[0073] sorting the deviations corresponding to the k first tasks in descending order, and determining the first i tasks in the sorting order as tasks with a running timeout exception;
[0074] The first task ranked in the first i positions is sent to an exception handling personnel for manual processing, and the first task ranked in the first i positions is marked as an exception.
[0075] In a third aspect, an embodiment of the present invention provides a computing device, comprising at least one processor and at least one memory, wherein the memory stores a computer program, and when the program is executed by the processor, the processor executes any task abnormality alarm method described in the first aspect.
[0076] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program executable by a computing device, wherein when the program runs on the computing device, the computing device executes any of the task abnormality alarm methods described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0078] Figure 1 A flowchart of a task abnormality alarm method provided by an embodiment of the present invention;
[0079] Figure 2 A schematic diagram of a probability distribution provided by an embodiment of the present invention;
[0080] Figure 3 A schematic diagram of performing matrix decomposition on a matrix provided by an embodiment of the present invention;
[0081] Figure 4 A schematic diagram of Monte Carlo simulation results for a task running timeout missed alarm rate and a task running timeout false alarm rate provided by an embodiment of the present invention;
[0082] Figure 5 A schematic diagram of the structure of a task abnormality alarm device provided by an embodiment of the present invention;
[0083] Figure 6 A schematic diagram of the structure of a computing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0084] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0085] Figure 1 The process of a task abnormality alarm method provided by an embodiment of the present invention is exemplarily shown, and the process can be executed by a task abnormality alarm device.
[0086] like Figure 1 As shown in the figure, the process specifically includes:
[0087] Step 101 : when an abnormality detection request for any task group is detected, task information of the task group is obtained from a task database.
[0088] In the embodiment of the present invention, since the calculation, analysis or processing of a type of big data usually involves multiple tasks, and the multiple tasks have a certain dependency (i.e., have a certain correlation), multiple tasks with a certain correlation (i.e., a correlation that meets the set requirements) are integrated into the same task group in advance, so that it is convenient to detect whether the calculation, analysis or processing of this type of big data has a running timeout exception in the form of a group, that is, to detect in parallel whether each task in the task group has a running timeout exception, so that it can more accurately detect which task link the running timeout exception specifically occurs in, and to a certain extent, it can improve the detection efficiency of the task running timeout exception. In this way, the various tasks configured for the calculation, analysis or processing of various types of big data in the big data platform are integrated and processed, and multiple task groups can be integrated, and the tasks in each task group have a correlation that meets the set requirements, such as requiring the correlation value between the tasks in any task group to meet a certain set threshold, and integrating the tasks whose correlation values meet the set threshold into one task group. After integrating multiple task groups, the multiple task groups are stored in the task database. In this way, each time a task operation timeout anomaly detection is performed, the corresponding pre-configured task group that needs to be checked can be obtained from the task database. For example, when detecting an anomaly detection request from a user for any task group, the anomaly detection request contains the identification number of the task group that needs to be detected, or when receiving an anomaly detection instruction from a user for any task group, the anomaly detection instruction contains the identification number of the task group that needs to be detected. In this way, the task information matching the identification number of the task group can be obtained from the task group database according to the identification number of the task group that needs to be detected, that is, the task information of the task group can be obtained, wherein the task information of the task group is used to indicate the current running time of each of the k first tasks in the task group; and the k first tasks are currently running and not marked as abnormal tasks, and there is a correlation between the k first tasks that meets the set requirements. Wherein, k is an integer greater than 1. That is to say, the task information of the task group includes the number of tasks currently running in the task group, such as k currently running tasks, and the current running time of each of the k currently running tasks. In addition, it should be noted that the k currently running tasks are also tasks that are not marked as abnormal. In this way, the influence of currently non-running tasks or tasks marked as abnormal on the current determination of whether there is a task in the task group with a running timeout abnormality can be effectively reduced, thereby improving the accuracy of the current determination of whether there is a task in the task group with a running timeout abnormality.
[0089] Step 102: construct a first running time matrix according to the current running time of each of the k first tasks.
[0090] In an embodiment of the present invention, after obtaining the current running time of each of the k first tasks, a first running time matrix can be constructed according to the current running time of each of the k first tasks. Specifically, the current running time of each of the k first tasks is normalized to obtain k normalized current running times, and the first running time matrix can be constructed through the k normalized current running times. Since the running times of different tasks may vary greatly, some tasks have a long running time, and some tasks have a short running time, the dimensional data can be converted into dimensionless data by normalizing the current running time of each of the k first tasks, that is, the running time of different tasks is normalized to the same dimension for corresponding processing, such as mapping to the interval [0,1] or [-1,1], so that it is convenient for the subsequent timely and accurate data calculation processing under the same dimension. For example, the current running time of each of the k first tasks is m 1 、m 2 、m 3 ,…,m k By normalizing the current running time of each of the k first tasks, that is, In this way, the first running time matrix M = [m 1 ′,m 2 ′,m 3 ′,…,m k ′].
[0091] Step 103 , obtaining the historical running time of each of the k first tasks in the first preset historical period from the task database, and constructing a second running time matrix through the historical running time of each of the k first tasks in the first preset historical period.
[0092] In the embodiment of the present invention, in order to more conveniently introduce the technical solution in the embodiment of the present invention, it is necessary to establish a model for the business scenario to which the technical solution in the embodiment of the present invention is applied. The model is based on the following assumptions:
[0093] (1) Since the running time of any task will fluctuate within a certain range with the amount of data, it can be assumed that the probability distribution of its fluctuation conforms to the characteristics of the normal distribution:
[0094]
[0095] Among them, t 1 is the task running time of any task, μ 1 is the expected value of the task running time of any task, σ 1 is the variance of the normal distribution.
[0096] (2) Accordingly, the set task running timeout alarm time can also be assumed to have the characteristics of normal distribution:
[0097]
[0098] Among them, t 2 The task running timeout warning time set for any task, μ 2 The expected value of the task running timeout warning time set for any task, σ 2 is the variance of the normal distribution.
[0099] (3) Generally, μ 2 >μ 1 At this time, the probability distribution of the two (i.e., the probability distribution of task running time fluctuation and the probability distribution of the set task running timeout alarm time) can be as follows Figure 2 Based on Figure 2 ,Should Figure 2 The three dashed lines from left to right represent the task running time t 1 , Task running timeout alarm time t 2 And the task abnormal time t 3 Among them, t 3 Indicates that the task has become abnormal if the time exceeds this time. From a statistical point of view, when the task running time satisfies the normal distribution, t 3 Generally in μ 1 +3σ 1 It is more appropriate to be at this point, when the task running time of any task exceeds μ 1 +3σ 1 The probability of t is less than 2%. However, in actual scenarios, the expected value of the task running time is usually unknown, so in actual scenarios, t 3 If the expected value of the task running time is unknown, the setting of the task running timeout alarm time is usually random. However, if the expected value of the task running time can be estimated, the error caused by human factors (i.e., manually setting the task running timeout alarm time) can be avoided.
[0100] (4) Ideally, t 2 =t 3 At this time, the task running timeout alarm time can well warn the task abnormality, but in actual situation, this situation will occur, that is, when t 2 <t 1 <t 3 , the task generates a running timeout alarm, but in fact the task is still within the normal running time. This situation is manifested as a false alarm. At this time, the false alarm probability for the task is:
[0101]
[0102] Among them, erfc() is the complementary error function, and its expression is:
[0103]
[0104] (5) For a set of unrelated tasks, the overall false positive rate is (1-ξ n ), n is the number of tasks. For a group of tasks with certain correlation on a big data platform, the false alarm rate will be between (ξ,1-ξ n )between.
[0105] (6) If the task timeout alarm time needs to be delayed to avoid false alarms, there is a risk of missed alarms, especially when t 3 <t 1 <t 2 When there is a false negative, the false negative probability for the task is:
[0106]
[0107] Based on the above assumptions about the model, the calculation operation of the expected value of the running time of each task in any task group can be started. When it is determined that anomaly detection of running timeout alarms needs to be performed for k tasks in a task group, after obtaining the task information of the task group, the historical running time of each of the k first tasks within the first preset historical period (for example, for the k first tasks, a window period before the current period is set, such as within 10 days, 20 days or 30 days, etc.) can be obtained from the task database based on the k first tasks. For each sub-period within the first preset historical period, the historical running time of the k first tasks belonging to the sub-period is used as the matrix column, and the k first tasks are used as the matrix row to construct the initial second running time matrix. For example, the running time of the k first tasks on the jth day in the first preset historical period is r. 1,j 、r 2,j 、r 3,j ,…,r k,j , the running time of the k first tasks on the jth day can form a vector r j =[r 1,j ,r 2,j ,r 3,j ,…,r k,j ] T Among them, r k,jIt is used to represent the running time of the kth task on the jth day. According to the assumption of the above model, the running time of the k first tasks on the jth day conforms to the normal distribution. It should be noted that due to the characteristics of the big data platform business, with the expansion of the business, the amount of data processed by the task is mostly fluctuating. If the task operation timeout abnormal alarm time is set manually, it will often cause an erroneous task operation timeout abnormal alarm as the amount of data increases, and the operation and maintenance personnel need to readjust the task operation timeout abnormal alarm time. By using the historical running time of each task in a certain task group within a preset historical period to estimate the expected value of the running time of each task, it can be convenient for subsequent operation and maintenance personnel to accurately set the task operation timeout abnormal alarm time. Then, as time goes by, the historical running time of each task in the task group is also constantly updated. Therefore, the technical solution in the example of the present invention can ensure that the task operation timeout abnormal alarm time is also updated accordingly, so as to better meet the needs of actual application scenarios. In order to determine the task running time estimation matrix more timely and accurately through the singular value decomposition algorithm, the number k of each first task is usually set to be less than the first preset historical period j, and the larger the value of j, the better the subsequent task running time estimation matrix will be. In this way, the running time of each of the k first tasks from the 1st day to the jth day can form an initial matrix, that is, R = [r 1 ,r 2 ,r 3 ,…,r j ]. The initial matrix can be expanded as:
[0108]
[0109] Each column of the matrix R represents the running time of each task in the task group within a day.
[0110] According to the assumptions of the above model, the task running time of each task within j days can be regarded as conforming to the normal distribution. In this way, the matrix R can be decomposed into a task running time expected value matrix and an error value matrix with normal distribution characteristics. At this time, the matrix R can be expressed as:
[0111]
[0112] In which, each column of the matrix X is the same, its rank is 1, [x 1 ,x 2 ,x 3 ,…,x k ] TIt is the expected running time of tasks 1 to k, that is, the expected running time of each task. After each task runs normally for the corresponding expected running time, it can be considered that the task is in the task completion state. Each item in the matrix N is a normal distribution error value with zero mean.
[0113] In addition, since the running time of different tasks may vary greatly, some tasks have a long running time, and some tasks have a short running time. In order to facilitate the subsequent correlation between the current running time of each task in the task group and the estimated running time of each task in the task group, it is necessary to normalize the matrix R, so as to convert the dimensional data into dimensionless data, that is, to normalize the running time of different tasks to the same dimension for corresponding processing, such as mapping to the interval [0,1] or [-1,1], so as to facilitate the subsequent timely and accurate data operations under the same dimension. Specifically, for the k matrix values of each column in the initial second running time matrix, the k matrix values of the column are normalized to obtain the k normalized matrix values of the column, and the second running time matrix R′=[r 1 ′,r 2 ′,r 3 ′,…,r j ′], that is
[0114] Step 104 , performing matrix decomposition on the second running time matrix to determine a first running time estimation value matrix for characterizing a process in which the k first tasks complete their running normally.
[0115] In the embodiment of the present invention, based on the assumption of the above model, according to the characteristics of the task running time, the problem of obtaining the expected value of the task running time can be converted into the problem of low-rank estimation of the matrix. That is, after obtaining the second running time matrix, the first running time estimation value matrix for k first tasks can be obtained by performing matrix decomposition on the second running time matrix. Exemplarily, taking the use of the singular value decomposition algorithm (SVD) to perform matrix decomposition on the second running time matrix as an example, since there are many tasks configured for the calculation, analysis or processing of big data, the constructed running time matrix is high-dimensional, which is not conducive to the subsequent calculation of the correlation value between the current running time matrix of each task (i.e., the first running time matrix) and the historical running time matrix of each task in the preset historical period (i.e., the second running time matrix), so the scheme uses the structural characteristics of the second running time matrix to perform low-rank estimation on the second running time matrix (i.e., use a matrix with a lower rank to approximate the original matrix), and map the second running time matrix from high-dimensional space to low-dimensional space. Specifically, the error between the second runtime matrix and the first runtime estimate matrix is minimized through a singular value decomposition algorithm, thereby determining the first runtime estimate matrix. That is, the second runtime matrix is decomposed through a singular value decomposition algorithm to determine multiple singular values, and the multiple singular values are compared to determine the largest singular value, and the left singular matrix corresponding to the largest singular value is determined as the first runtime estimate matrix. In this way, the scheme can set some singular values to zero by performing matrix decomposition on the second runtime matrix and shrinking the singular values of the second runtime matrix according to the rank, thereby achieving the purpose of low-rank estimation, and the first runtime estimate matrix can be calculated.
[0116] Among them, when the second running time matrix is decomposed by the singular value decomposition algorithm, due to the structural characteristics of the second running time, the second running time matrix is actually composed of the expected value matrix of the task running time and the error value matrix of the normal distribution with zero mean. Among them, the fluctuation of the task running time and the error caused by the artificial setting of the abnormal alarm time are fitted to the normal distribution, so that the real scene can be better restored when performing the operation timeout abnormal analysis. Therefore, the second running time matrix is first converted into a low-rank matrix (that is, the expected value matrix of the task running time) and an error matrix, wherein each error value in the error matrix conforms to the normal distribution, but the low-rank matrix and the error matrix are unknown, so it is necessary to estimate the estimated value matrix of the approximate low-rank matrix through the matrix low-rank estimation method, that is, to decompose the low-rank matrix through the singular value decomposition algorithm, so as to determine multiple singular values.
[0117] Exemplarily, from the task running time matrix obtained above (i.e., the second running time matrix, such as matrix R), it can be known that the structure of this matrix has particularity. For example, each column of matrix X is the same and its rank is 1. Generally, the rank of a matrix can characterize the correlation between data. For a set of correlated data, the rank of the matrix composed of it is much smaller than the number of its columns, that is, it has the low-rank property. Then, in this case, for the principal component analysis of such correlated data, it is equivalent to performing low-rank estimation on the data matrix, projecting it from a high-dimensional space to a low-dimensional space to obtain the component with the maximum correlation. In addition, in the problem of matrix low-rank estimation, a function is first defined, and through this function, under the constraint of low rank, the error between the original matrix and the estimated matrix is minimized. For example, for a matrix D, this matrix D is approximately estimated as the product of U×V T where rank(U×V T ) < rank(D). The low-rank estimation of a matrix is to approximate the original matrix with a matrix of lower rank, and the condition it needs to satisfy is that the error between the estimated matrix and the original matrix is the smallest, that is:
[0118] min‖D - U×V T ‖ F
[0119] where ‖‖ F is used to represent the Frobenius norm.
[0120] Among them, the low-rank estimation of a matrix can mainly be used for data compression. For example, as shown in Figure 3 , matrix D can be approximated by two smaller matrices (i.e., matrix U and matrix V). Among them, the mathematical model for using matrix low-rank estimation for matrix D is D = X + E. Among them, matrix X is a low-rank matrix, matrix E is a noise or error matrix, and matrix D is the data matrix that can be obtained in the actual scenario. This model reflects that in the actual situation, the original data matrix X is a low-rank matrix, but due to the influence of noise or measurement error E, the rank of the matrix D that can be obtained in the actual scenario is much larger than the rank of the original low-rank matrix X.
[0121] Then, in order to determine the original low-rank matrix X from the matrix D that can be obtained in the actual scenario, the following optimization problem needs to be solved:
[0122]
[0123] Among them, the most commonly used method to solve this optimization problem is the singular value decomposition algorithm, that is, by using the singular value decomposition algorithm to decompose matrix D, and then according to the rank, shrink its singular values, that is, set some singular values to zero, so as to achieve the purpose of low-rank estimation, and at the same time, a low-rank matrix approximating matrix D can also be obtained
[0124] Through the above analysis, it can be seen that for the matrix R, the actual running time of each task in the matrix R can also constitute such a low-rank estimation model. Through the decomposition of the above matrix R, it can be seen that the matrix R can actually be divided into two parts, one part is the task running time expected value matrix, and the expected values of each task running time in the task running time expected value matrix are fixed values, and the other part is the fluctuation of the running time due to factors such as the amount of data, which presents a standard normal distribution characteristic. Comparing the matrix R = X + N with the matrix D = X + E, it can be found that the data distribution characteristics of the task running time are completely consistent with the model of low-rank estimation denoising, so the low-rank estimation method can be used to reduce the impact of the running time fluctuation caused by factors such as the amount of data, thereby estimating the expected value of the task running time of each task in the task group. The expected value of the task running time of each task in the task group can accurately set the running timeout alarm time, rather than manually estimating the expected value of the task running time of each task to configure the running timeout alarm time.
[0125] Based on this, for the matrix R, by using the matrix low rank estimation to reduce its dimension to 1, that is, the hard threshold method is used for the rank selection, then a set of task running times can be obtained, which is the time with the greatest correlation with the task running time in the past j days. Specifically, the low rank estimation of the matrix R can also be performed in the same way as the matrix D, so that the low rank matrix used to characterize the expected value of the task running time can be determined from the matrix R, that is:
[0126]
[0127] Among them, the low-rank matrix The rank is 1, a low-rank matrix It can be expressed as:
[0128]
[0129] in, It is the time series with the greatest correlation with the task running time in the past j days, that is, the expected value of the estimated running time of each task, v = [v 1 ,v 2 ,v 3 ,…,v j ] is the daily task running time in the past j days and Among them, for the low-rank estimation of the matrix, in the process of decomposing the low-rank matrix using the singular value decomposition algorithm, only the matrix corresponding to the largest singular value among the obtained multiple singular values is retained, that is, the left singular matrix corresponding to the largest singular value is retained, and the result can be obtained In this way, the expected running time of each of the k tasks can be determined from the historical running time of each of the k tasks within the preset historical period.
[0130] Step 105: determine a first correlation value between the first running time matrix and the first running time estimation value matrix, and determine whether the first correlation value is less than a first preset threshold.
[0131] In an embodiment of the present invention, after determining the estimated running time values for k first tasks (i.e., the first running time estimated value matrix), the first running time matrix and the first running time estimated value matrix can be used to perform correlation calculation, and then the correlation is used to determine whether to issue a running timeout alarm for at least one task in the task group. Specifically, the first correlation value between the first running time matrix and the first running time estimated value matrix is determined by a correlation value calculation formula, and it is determined whether the first correlation value is less than a first preset threshold value, so as to determine whether it is necessary to issue a running timeout alarm for at least one task in the task group. The first preset threshold value is any preset threshold value randomly selected from a preset threshold value interval range determined based on the historical running time of each of the multiple second tasks in a second preset historical period.
[0132] Exemplarily, the correlation calculation formula between the first running time matrix and the first running time estimation value matrix is as follows:
[0133]
[0134] Wherein, M is used to represent the first running time matrix (i.e., the matrix constructed after normalizing the current running time of each of the k first tasks). When calculating the first correlation value between the first running time matrix and the first running time estimation value matrix, the calculated first correlation value λ∈[-1,1] can be obtained by normalizing both the first running time matrix and the first running time estimation value matrix. At this time, a first preset threshold ε can be set, and when it is determined that the calculated first correlation value λ is less than the first preset threshold ε, a task running timeout alarm is issued. For example, the first preset threshold ε is set to 0.9. If the calculated first correlation value λ<0.9, a task running timeout alarm can be issued.
[0135] Among them, in order to more accurately determine whether there is a timeout exception in the task operation, it is necessary to determine a preset threshold interval range that can more truly and accurately determine whether there is a timeout exception in the task operation, so as to randomly select a preset threshold from the preset threshold interval range to determine whether there is a timeout exception in the task operation. In this way, the flexibility of selecting the preset threshold is higher. Moreover, in order to further achieve a lower operation timeout false alarm rate and operation timeout missed alarm rate, the scheme needs to set a reasonable preset threshold. If the preset threshold is set too large, although the operation timeout missed alarm rate can be reduced, it will cause the operation timeout false alarm rate to increase significantly; if the preset threshold is set too small, although the operation timeout false alarm rate can be reduced, it will increase the operation timeout missed alarm rate. Therefore, this scheme simulates the task operation status based on m completely unrelated second tasks that conform to the normal distribution according to the Monte Carlo simulation method, so as to determine the preset threshold interval range in which the overall task operation timeout false alarm rate and the operation timeout omission rate are low, so that when judging whether there is an operation timeout alarm in the task group, a preset threshold can be randomly selected from the preset threshold interval range in time as the preset threshold for accurately judging the operation timeout alarm.
[0136] Specifically, the preset threshold interval range can be determined in the following manner: according to the Monte Carlo simulation method, m unrelated second tasks that all conform to the normal distribution are taken as a task group, and multiple different second preset thresholds are set, and the historical running time of each of the m second tasks in the second preset historical period (for example, for the m second tasks, a window period before the current period is set, such as within 10 days, 20 days or 30 days, etc.) is obtained, and the third running time matrix is constructed through the historical running time of each of the m second tasks in the second preset historical period. Then, the third running time matrix is decomposed by the singular value decomposition algorithm to determine the second running time estimation value matrix used to characterize the normal completion of the running process of the m second tasks, and for each second preset threshold, the m second tasks are each run multiple times in the current period, and the running time of each of the m second tasks in each run in the current period is set, and the fourth running time matrix is constructed through the running time of each of the m second tasks in each run. Then, determine the second correlation value between the fourth running time matrix and the second running time estimation value matrix, and by determining whether the second correlation value is less than the second preset threshold, determine whether there is a false alarm and / or omission of running timeout in any of the m second tasks, until the multiple runs of each of the m second tasks in the current time period are traversed and completed at the second preset threshold, thereby determining the running timeout omission rate and running timeout false alarm rate of each of the m second tasks running multiple times in the current time period at the second preset threshold. Finally, generate a Monte Carlo simulation graph through multiple different second preset thresholds and the running timeout omission rate and running timeout false alarm rate corresponding to each of the multiple different second preset values, and through the Monte Carlo simulation graph, it can be determined that the running timeout omission rate is less than or equal to the first set value and the running timeout false alarm rate is less than or equal to the second set value. And according to each second preset threshold, the preset threshold interval range for detecting whether the task has timed out can be accurately constructed. Among them, the first setting value and the second setting value can be set according to the experience of technical personnel in this field, or can be set according to the needs of actual application scenarios, or can be obtained through multiple experiments based on historical data, and the embodiments of the present invention are not limited to this.
[0137] For example, according to the Monte Carlo simulation method, a group of 30 completely unrelated tasks that conform to the normal distribution are used. The mean running time of the 30 tasks is 2 hours, the variance is 0.2, and the running time history is 60 days. Assuming that the task running time exceeds the mean 3σ, that is, if the running time of a task exceeds 2.6 hours, then the task is determined to be an abnormal task. In this simulation, the following is obtained: Figure 4 The Monte Carlo simulation results are shown in the figure. Figure 4It can be seen that with the increase of the preset threshold, the false alarm rate of task timeout is increasing, but the missed alarm rate of task timeout is decreasing, which is consistent with the above judgment. Therefore, according to the simulation results, it can be concluded that when the preset threshold is between 0.988-0.990, the overall false alarm rate of task timeout and the missed alarm rate of task timeout are both low.
[0138] It should be noted that in actual situations, since the tasks in a task group are generally related, the obtained operation timeout false alarm rate and operation timeout omission rate are lower than the above simulation results, and the selection range of appropriate preset thresholds will also be larger.
[0139] Step 106: When the first correlation value is less than the first preset threshold, it is determined that at least one first task has a running timeout exception, and a timeout exception alarm is issued for the at least one first task.
[0140] In an embodiment of the present invention, if it is determined that the first correlation value is less than the first preset threshold, it can be determined that at least one first task has an operation timeout exception, and a timeout exception alarm is issued for the at least one first task; if it is determined that the first correlation value is greater than or equal to the first preset threshold, it can be determined that there is no operation timeout exception among the k first tasks, and therefore there is no need to issue an operation timeout alarm.
[0141] Among them, after determining that at least one first task has a running timeout exception, for each of the k first tasks, the deviation corresponding to the first task can be determined by calculating the deviation between the current running time of the first task and the running time estimate of the first task in the first running time estimate matrix, and the deviations corresponding to the k first tasks are sorted in order from large to small, so that it can be accurately determined that the first tasks ranked in the first i are the tasks with the running timeout exception. In this way, it can be avoided that the set task running abnormality alarm time is inaccurate due to manual estimation of the expected value of the task running time, and the first tasks ranked in the first i are sent to the exception handling personnel for corresponding manual processing, so that the specific problem of the running timeout exception of the first tasks ranked in the first i can be accurately located through manual processing, and the corresponding solution is adopted to deal with the specific problem. At the same time, the first task ranked in the first i is marked as abnormal in order to avoid the i first tasks marked as abnormal from interfering with the task group's timeout abnormality judgment when the task group to which the first task ranked in the first i is assigned is judged for timeout abnormality next time, thereby effectively ensuring that each timeout abnormality judgment for any task group can be performed accurately and normally. The deviation between the current running time of the first task and the running time estimation value of the first task in the first running time estimation value matrix is determined in the following manner, namely:
[0142]
[0143] in, It is used to represent the deviation corresponding to any task, t is used to represent the current running time of the task, and t′ is used to represent the estimated running time of the task.
[0144] The above embodiments show that, since the existing technical solution for uncertain task operation anomalies is detected by artificially setting the task operation anomaly alarm time, it needs to rely on the experience of the operation and maintenance personnel, which is relatively subjective. Therefore, due to the different experiences of different operation and maintenance personnel, the false alarm and / or missed alarm of uncertain task operation anomalies often occur. Based on this, the technical solution in the present invention is for any task group, and the running time estimation value matrix (i.e., the running time expectation value matrix) for each task is determined by the historical running time of each task in the task group in each time period within the preset time period. The running time estimation value matrix determined in this way is more in line with reality and more in line with the actual running conditions of each task in the task group, and the running time estimation value matrix is used as a benchmark for judging whether the running of each task is overtime, so that it can be more truly and accurately determined whether there is a running timeout anomaly in the task group. Specifically, when an abnormality detection request for any task group is detected, the task information of the task group is obtained from the task database, and the first running time matrix can be constructed by the current running time of each of the k first tasks. Then, the historical running time of each of the k first tasks in the first preset historical period is obtained from the task database, and the second running time matrix can be constructed through the historical running time of each of the k first tasks in the first preset historical period, and the first running time estimation value matrix can be accurately determined by performing matrix decomposition on the second running time matrix. Then, when it is determined that the first correlation value between the first running time matrix and the first running time estimation value matrix is less than the first preset threshold value, it is determined that at least one of the k first tasks has a running timeout exception, and a timeout exception alarm is issued for the at least one first task. In this way, the running time estimation value matrix for each task determined by the scheme through the historical running time of each task in the task group in each time period within the preset historical time period is more in line with reality and more in line with the actual operation status of the task group. Therefore, by comparing the first correlation value between the running time estimation value matrix and the first running time matrix with the first preset threshold, it is possible to more accurately determine whether there are tasks in the task group that have running timeout exceptions, thereby effectively avoiding the high missed alarm rate and false alarm rate of task running timeout exception alarms due to manually setting the abnormal alarm time, thereby effectively reducing the missed alarm rate and false alarm rate of task running timeout exception alarms.
[0145] Based on the same technical concept, Figure 5An exemplary task abnormality alarm device provided by an embodiment of the present invention is shown, and the device can execute the process of the task abnormality alarm method.
[0146] like Figure 5 As shown, the device comprises:
[0147] The acquisition unit 501 is used to acquire task information of any task group from a task database when an abnormality detection request for the task group is detected; the task information is used to indicate the current running time of each of the k first tasks in the task group;
[0148] The processing unit 502 is used to construct a first running time matrix through the current running time of each of the k first tasks; obtain the historical running time of each of the k first tasks within the first preset historical time period from the task database, and construct a second running time matrix through the historical running time of each of the k first tasks within the first preset historical time period; perform matrix decomposition on the second running time matrix to determine a first running time estimation value matrix used to characterize the normal completion of the running process of the k first tasks; determine a first correlation value between the first running time matrix and the first running time estimation value matrix, and determine whether the first correlation value is less than a first preset threshold; the first preset threshold is any preset threshold randomly selected from a preset threshold interval determined based on the historical running time of each of the multiple second tasks within the second preset historical time period; when the first correlation value is less than the first preset threshold, it is determined that at least one of the first tasks has a running timeout exception, and a timeout exception alarm is issued for the at least one first task.
[0149] Optionally, the processing unit 502 is specifically configured to:
[0150] Normalizing the current running time of each of the k first tasks to obtain k normalized current running times;
[0151] Constructing the first running time matrix through the k normalized current running times;
[0152] The processing unit 502 is specifically used for:
[0153] For each sub-period within the first preset historical period, an initial second running time matrix is constructed using the historical running time of each of the k first tasks belonging to the sub-period as a matrix column and the k first tasks as a matrix row;
[0154] For the k matrix values in each column of the initial second running time matrix, normalize the k matrix values in the column to obtain k normalized matrix values in the column;
[0155] The second runtime matrix is constructed by using the k normalized matrix values in each column.
[0156] Optionally, the processing unit 502 is specifically configured to:
[0157] Decomposing the second runtime matrix by a singular value decomposition algorithm to determine a plurality of singular values;
[0158] The multiple singular values are compared to determine a maximum singular value, and a left singular matrix corresponding to the maximum singular value is determined as the first runtime estimation value matrix.
[0159] Optionally, the processing unit 502 is specifically configured to:
[0160] Converting the second runtime matrix into a low-rank matrix and an error matrix; each error value in the error matrix conforms to a normal distribution;
[0161] The low-rank matrix is decomposed by the singular value decomposition algorithm to determine the multiple singular values.
[0162] Optionally, the k first tasks are tasks that are currently running and are not marked as abnormal; and the k first tasks have correlations that meet set requirements.
[0163] Optionally, the processing unit 502 is specifically configured to:
[0164] According to the Monte Carlo simulation method, m unrelated second tasks that all conform to the normal distribution are regarded as a task group, and a plurality of different second preset thresholds are set;
[0165] Obtaining the historical running time of each of the m second tasks within the second preset historical period, and constructing a third running time matrix according to the historical running time of each of the m second tasks within the second preset historical period;
[0166] Performing matrix decomposition on the third running time matrix to determine a second running time estimation value matrix for characterizing a normal completion running process of the m second tasks;
[0167] Set, for each second preset threshold, to run each of the m second tasks multiple times in the current period, and construct a fourth running time matrix according to the running time of each of the m second tasks in each running;
[0168] Determine a second correlation value between the fourth running time matrix and the second running time estimation value matrix, and determine whether there is a false alarm and / or missed alarm of running timeout among the m second tasks by determining whether the second correlation value is less than the second preset threshold, until multiple runs of each of the m second tasks in the current time period are traversed and completed at the second preset threshold, thereby determining the running timeout missed alarm rate and running timeout false alarm rate of each of the m second tasks running multiple times in the current time period at the second preset threshold;
[0169] Generate a Monte Carlo simulation graph by using the multiple different second preset thresholds and the operation timeout missed alarm rates and the operation timeout false alarm rates corresponding to the multiple different second preset values;
[0170] Through the Monte Carlo simulation diagram, the second preset thresholds corresponding to the operation timeout omission rate being less than or equal to the first set value and the operation timeout false alarm rate being less than or equal to the second set value are determined, and based on the second preset thresholds, a preset threshold interval range for detecting whether the task operation has timed out is constructed.
[0171] Optionally, the processing unit 502 is further configured to:
[0172] After determining that at least one of the first tasks has a running timeout exception, determining, for each of the k first tasks, a deviation between a current running time of the first task and a running time estimate of the first task in the first running time estimate matrix;
[0173] sorting the deviations corresponding to the k first tasks in descending order, and determining the first i tasks in the sorting order as tasks with a running timeout exception;
[0174] The first task ranked in the first i positions is sent to an exception handling personnel for manual processing, and the first task ranked in the first i positions is marked as an exception.
[0175] Based on the same technical concept, the embodiment of the present invention also provides a computing device, such as Figure 6 As shown, it includes at least one processor 601 and a memory 602 connected to the at least one processor. The specific connection medium between the processor 601 and the memory 602 is not limited in the embodiment of the present invention. Figure 6 For example, the processor 601 and the memory 602 are connected via a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0176] In the embodiment of the present invention, the memory 602 stores instructions that can be executed by at least one processor 601. The at least one processor 601 can execute the steps included in the aforementioned task abnormality alarm method by executing the instructions stored in the memory 602.
[0177] Among them, the processor 601 is the control center of the computing device, and can use various interfaces and lines to connect various parts of the computing device, and realize data processing by running or executing instructions stored in the memory 602 and calling data stored in the memory 602. Optionally, the processor 601 may include one or more processing units, and the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes the issuance of instructions. It is understandable that the above-mentioned modem processor may not be integrated into the processor 601. In some embodiments, the processor 601 and the memory 602 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on independent chips.
[0178] The processor 601 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the task abnormality alarm method may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0179] The memory 602 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 602 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 602 is any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 602 in the embodiment of the present invention can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0180] Based on the same technical concept, an embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program executable by a computing device. When the program runs on the computing device, the computing device executes the steps of the above-mentioned task abnormality alarm method.
[0181] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0182] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0183] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0184] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0185] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0186] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of this application and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A task abnormality alarm method, It is characterized in that include: When an abnormality detection request for any task group is detected, obtaining task information of the task group from a task database; The task information is used to indicate the current running time of each of the k first tasks in the task group; Constructing a first running time matrix according to the current running time of each of the k first tasks; Acquire the historical running time of each of the k first tasks in the first preset historical period from the task database, and construct a second running time matrix according to the historical running time of each of the k first tasks in the first preset historical period; Performing matrix decomposition on the second running time matrix to determine a first running time estimation value matrix for characterizing a normal completion running process of the k first tasks; Determine a first correlation value between the first run time matrix and the first run time estimate value matrix, and determine whether the first correlation value is less than a first preset threshold; The first preset threshold is any preset threshold randomly selected from a preset threshold interval determined based on the historical running time of each of the plurality of second tasks within the second preset historical period; The preset threshold interval range is obtained by simulating the task operation status based on m completely unrelated second tasks that conform to the normal distribution according to the Monte Carlo simulation method; When the first correlation value is less than the first preset threshold, it is determined that at least one first task has a running timeout exception, and a timeout exception alarm is issued for the at least one first task.
2. The method according to claim 1, It is characterized in that The constructing a first running time matrix according to the current running time of each of the k first tasks includes: Normalizing the current running time of each of the k first tasks to obtain k normalized current running times; Constructing the first running time matrix through the k normalized current running times; The second running time matrix is constructed by using the historical running time of each of the k first tasks in the first preset historical period, including: For each sub-period within the first preset historical period, an initial second running time matrix is constructed using the historical running time of each of the k first tasks belonging to the sub-period as a matrix column and the k first tasks as a matrix row; For the k matrix values in each column of the initial second running time matrix, normalize the k matrix values in the column to obtain k normalized matrix values in the column; The second runtime matrix is constructed by using the k normalized matrix values in each column.
3. The method according to claim 1, It is characterized in that The performing matrix decomposition on the second running time matrix to determine a first running time estimation value matrix for characterizing the normal completion of the running process of the k first tasks includes: Decomposing the second runtime matrix by a singular value decomposition algorithm to determine a plurality of singular values; The multiple singular values are compared to determine a maximum singular value, and a left singular matrix corresponding to the maximum singular value is determined as the first runtime estimation value matrix.
4. The method according to claim 3, It is characterized in that The step of performing matrix decomposition on the second runtime matrix by using a singular value decomposition algorithm to determine a plurality of singular values includes: Converting the second runtime matrix into a low-rank matrix and an error matrix; each error value in the error matrix conforms to a normal distribution; The low-rank matrix is decomposed by the singular value decomposition algorithm to determine the multiple singular values.
5. The method according to claim 1, It is characterized in that The k first tasks are tasks that are currently running and are not marked as abnormal; and there are correlations between the k first tasks that meet set requirements.
6. The method according to claim 1, It is characterized in that The preset threshold interval range is determined by: According to the Monte Carlo simulation method, m unrelated second tasks that all conform to the normal distribution are regarded as a task group, and a plurality of different second preset thresholds are set; Obtaining the historical running time of each of the m second tasks within the second preset historical period, and constructing a third running time matrix according to the historical running time of each of the m second tasks within the second preset historical period; Performing matrix decomposition on the third running time matrix to determine a second running time estimation value matrix for characterizing a normal completion running process of the m second tasks; Set, for each second preset threshold, to run each of the m second tasks multiple times in the current period, and construct a fourth running time matrix according to the running time of each of the m second tasks in each running; Determine a second correlation value between the fourth running time matrix and the second running time estimation value matrix, and determine whether there is a false alarm and / or missed alarm of running timeout among the m second tasks by determining whether the second correlation value is less than the second preset threshold, until multiple runs of each of the m second tasks in the current time period are traversed and completed at the second preset threshold, thereby determining the running timeout missed alarm rate and running timeout false alarm rate of each of the m second tasks running multiple times in the current time period at the second preset threshold; Generate a Monte Carlo simulation graph by using the multiple different second preset thresholds and the operation timeout missed alarm rates and the operation timeout false alarm rates corresponding to the multiple different second preset values; Through the Monte Carlo simulation diagram, the second preset thresholds corresponding to the operation timeout omission rate being less than or equal to the first set value and the operation timeout false alarm rate being less than or equal to the second set value are determined, and based on the second preset thresholds, a preset threshold interval range for detecting whether the task operation has timed out is constructed.
7. The method according to any one of claims 1 to 6, It is characterized in that After determining that at least one of the first tasks has a running timeout exception, the method further includes: For each of the k first tasks, determining a deviation between a current running time of the first task and a running time estimate of the first task in the first running time estimate matrix; sorting the deviations corresponding to the k first tasks in descending order, and determining the first i tasks in the sorting order as tasks with a running timeout exception; The first task ranked in the first i positions is sent to an exception handling personnel for manual processing, and the first task ranked in the first i positions is marked as an exception.
8. A task abnormality alarm device, It is characterized in that include: an acquisition unit, configured to acquire task information of any task group from a task database when an abnormality detection request for the task group is detected; The task information is used to indicate the current running time of each of the k first tasks in the task group; A processing unit, configured to construct a first running time matrix according to the current running time of each of the k first tasks; Acquire the historical running time of each of the k first tasks in the first preset historical period from the task database, and construct a second running time matrix according to the historical running time of each of the k first tasks in the first preset historical period; Performing matrix decomposition on the second running time matrix to determine a first running time estimation value matrix for characterizing the normal completion of the running process of the k first tasks; determining a first correlation value between the first running time matrix and the first running time estimation value matrix, and determining whether the first correlation value is less than a first preset threshold; The first preset threshold is any preset threshold randomly selected from a preset threshold interval determined based on the historical running time of each of the plurality of second tasks within the second preset historical period; The preset threshold interval range is obtained by simulating the task operation status based on m completely unrelated second tasks that all conform to the normal distribution according to the Monte Carlo simulation method; when the first correlation value is less than the first preset threshold, it is determined that at least one first task has an operation timeout exception, and a timeout exception alarm is issued for the at least one first task.
9. A computing device, It is characterized in that The method comprises at least one processor and at least one memory, wherein the memory stores a computer program, and when the program is executed by the processor, the processor executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, It is characterized in that It stores a computer program executable by a computing device, and when the program is run on the computing device, the computing device executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data processing method and device
CN110795324A
Abnormality detection method and device
CN113556258A