Alarm configuration method and device of data operation link
By building a directed acyclic graph in the data processing job link and using a multi-source breadth priority search algorithm to calculate the recommended time limit threshold, the problem of complex dependence of data processing job links and low alarm configuration efficiency is solved, and efficient time calculation and accurate alarm configuration are achieved.
Patent Information
- Application Number
- CN202510128441.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-05
AI Technical Summary
In the data processing operation scenarios of large commercial banks, due to the diversification of business and the complexity of data processing operations, the dependencies of the data processing operation link are complex, resulting in low time-based computing efficiency and high operation and maintenance pressure. The existing alarm configuration scheme is prone to alarm storms or the upstream delay cannot be detected in time.
By constructing a directed acyclic graph, using a multi-source breadth-first search algorithm, the recommended time limit for jobs is calculated, and key nodes are determined, and the alarm threshold is automatically configured to improve the accuracy and efficiency of alarms.
It reduces the communication cost between the terminal operating system and the upstream operating system, improves the efficiency of operation and maintenance, reduces the pressure of operation and maintenance, avoids alarm storms, and improves the operation and maintenance effect.
Smart Images

Figure CN119996154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of alarm configuration, and in particular to an alarm configuration method and device for a data operation link. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the invention recited in the claims. No admission is made that the description herein is prior art by inclusion in this section.
[0003] With the development of financial technology in large commercial banks, the application of big data in bank IT systems has become more and more extensive. This trend has led to an exponential growth in the amount of data that needs to be processed by the bank's back-end, and the timeliness requirements for data processing operations have also increased accordingly.
[0004] In the development scenario of massive data processing jobs, due to the diversification of business and the complexity of data processing jobs, the link dependencies of data processing jobs become intricate, specifically manifested in long links and multiple cross-system node dependencies. Therefore, in the actual operation and maintenance process, a lot of communication costs are required to determine the dependencies between data processing jobs, and the efficiency of data processing job time calculation is low.
[0005] The existing data processing job timeliness monitoring system adopts two alarm configuration schemes: one scheme is to configure alarms for nodes in the entire data operation link. If the alarm threshold configuration is unreasonable, when the upstream operation alarms, the downstream operation will also alarm synchronously, resulting in an alarm storm and many useless alarms, which will cause great interference to the analysis of operation and maintenance personnel; the other scheme is to configure alarms only for terminal operations, which cannot detect delays in the upstream operation system in time, causing the terminal operation system to bear most of the operation and maintenance pressure, and the operation and maintenance effect is poor. It is often found that the delay has missed the best time for emergency response. Summary of the invention
[0006] The embodiment of the present invention provides a method for configuring an alarm of a data operation link, which is used to reduce the communication cost between a terminal operation system and an upstream operation system, improve the efficiency of operation time calculation, improve the accuracy and efficiency of data operation link alarms, and reduce the operation and maintenance pressure of the operation. The method includes:
[0007] A directed acyclic graph is constructed with jobs as nodes and dependencies between jobs as edges. The direction of the edge is from the predecessor job to the corresponding successor job.
[0008] Initialize the value of the terminal job node to the timeliness threshold of the terminal job; determine the average execution time of the predecessor job corresponding to the edge as the weight of the edge; where the terminal job node is a node with an out-degree of 0, and the timeliness threshold of the terminal job indicates the time allowed for the terminal job to be delayed;
[0009] According to the data processing systems to which the jobs belong and the dependencies between the jobs, job nodes that belong to different data processing systems and have dependencies are identified as key nodes;
[0010] Starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph, and the value of each node is updated according to the following rules: the weight of the edge pointing to the current search node is added to the value of the current search node; if the sum is greater than the value of the starting node of the edge, or the starting node of the edge is not assigned a value, the value of the starting node of the edge is updated to the sum; the updated value of each node is the recommended timeliness threshold of the job corresponding to each node;
[0011] Configure alarm thresholds for jobs corresponding to key nodes based on the recommended timeliness thresholds for key nodes.
[0012] The embodiment of the present invention further provides an alarm configuration device for a data operation link, which is used to reduce the communication cost between the terminal operation system and the upstream operation system, improve the operation time efficiency calculation efficiency, improve the accuracy and efficiency of the data operation link alarm, and reduce the operation and maintenance pressure of the operation. The device includes:
[0013] The link construction module is used to construct a directed acyclic graph with jobs as nodes and dependencies between jobs as edges. The direction of the edge is from the predecessor job to the corresponding successor job.
[0014] The link initialization module is used to initialize the value of the terminal job node to the timeliness threshold of the terminal job; the average execution time of the predecessor job corresponding to the edge is determined as the weight of the edge; the terminal job node is a node with an out-degree of 0, and the timeliness threshold of the terminal job indicates the time allowed for the terminal job to be delayed;
[0015] A key node analysis module is used to determine the job nodes that belong to different data processing systems and have dependencies as key nodes according to the data processing systems to which the jobs belong and the dependencies between the jobs;
[0016] The timeliness threshold analysis module is used to perform a multi-source breadth-first search on the directed acyclic graph starting from the terminal job node, and update the value of each node according to the following rules: add the weight of the edge pointing to the current search node to the value of the current search node; if the sum of the added values is greater than the value of the starting node of the edge, or the starting node of the edge is not assigned a value, then update the value of the starting node of the edge to the sum; the updated value of each node is the recommended timeliness threshold of the job corresponding to each node;
[0017] The alarm configuration module is used to configure the alarm threshold for the job corresponding to the key node according to the recommended time threshold of the key node.
[0018] An embodiment of the present invention further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned alarm configuration method for the data operation link when executing the computer program.
[0019] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the alarm configuration method for the data operation link is implemented.
[0020] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the alarm configuration method of the data operation link is implemented.
[0021] Existing data processing job timeliness monitoring systems usually adopt a solution that configures alarms for all nodes in the entire data job chain, or only configures alarms for terminal jobs. Compared with the existing technical solutions, in the embodiment of the present invention, a directed acyclic graph is constructed based on the execution time of the job, the timeliness threshold of the terminal job, and the job dependency. The recommended timeliness threshold of the job is efficiently solved through the multi-source breadth-first search algorithm of the graph, which reduces the communication cost between the terminal job system and the upstream job system and improves the efficiency of job timeliness calculation. In the embodiment of the present invention, jobs with dependencies between different systems are also identified as key nodes, and all key nodes on the data job chain are automatically identified. Alarms are configured according to the recommended timeliness threshold of the key nodes, which improves the accuracy and efficiency of data job chain alarms, avoids the generation of alarm storms, improves operation and maintenance effects, and reduces operation and maintenance pressure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0023] Figure 1 It is a flow chart of the alarm configuration method of the data operation link in an embodiment of the present invention;
[0024] Figure 2 It is a job dependency diagram of a data job link in an embodiment of the present invention;
[0025] Figure 3 A weighted job dependency graph of a data job link in an embodiment of the present invention;
[0026] Figure 4A flowchart of the key operation time efficiency analysis in an embodiment of the present invention;
[0027] Figure 5 Schematic diagram of an alarm configuration device for a data operation link in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] To make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0029] First, the relevant technical terms in the embodiments of the present invention are introduced:
[0030] Job: A collection of programs, usually implemented in a programming language based on business logic. The developed programs are deployed on production servers and automatically executed at a fixed time or frequency through a scheduling management platform. Jobs in big data applications are usually used to load, copy, process, and clean data in the database.
[0031] Job dependency: The relationship between jobs. The successor job can only start executing after the predecessor job is completed.
[0032] Data job link: a directed acyclic graph (DAG) consisting of jobs and job dependencies;
[0033] Topological sorting: A sorting algorithm for directed acyclic graphs that follows the front-end and back-end dependencies of directed graphs to build an ordered node sequence and ensure that the job nodes that are sorted first are completed first.
[0034] Figure 1 FIG. 1 is a flow chart of an alarm configuration method for a data operation link in an embodiment of the present invention. Figure 1 As shown, the method can be implemented by the following steps:
[0035] Step 101: construct a directed acyclic graph with jobs as nodes and dependencies between jobs as edges; wherein the direction of the edge is from the predecessor job to the corresponding successor job;
[0036] Step 102: Initialize the value of the terminal job node as the timeliness threshold of the terminal job; determine the average execution time of the predecessor job corresponding to the edge as the weight of the edge; wherein the terminal job node is a node with an out-degree of 0, and the timeliness threshold of the terminal job indicates the time length allowed for the terminal job to be delayed in execution;
[0037] Step 103: According to the data processing systems to which the jobs belong and the dependency relationships between the jobs, job nodes that belong to different data processing systems and have dependency relationships are determined as key nodes;
[0038] Step 104: Starting from the terminal operation node, perform a multi-source breadth-first search on the directed acyclic graph, and update the value of each node according to the following rules: add the weight of the edge pointing to the current search node to the value of the current search node; if the sum of the added values is greater than the value of the starting node of the edge, or the starting node of the edge is not assigned a value, then update the value of the starting node of the edge to the sum; wherein the updated value of each node is the recommended timeliness threshold of the operation corresponding to each node;
[0039] Step 105: According to the recommended timeliness threshold of the key node, configure the alarm threshold for the job corresponding to the key node.
[0040] Depend on Figure 1 It can be seen from the process that in the embodiment of the present invention, a directed acyclic graph is constructed according to the execution time of the job, the timeliness threshold of the terminal job, and the job dependency. The recommended timeliness threshold of the job is efficiently solved through the multi-source breadth-first search algorithm of the graph, and the jobs with dependencies between different systems are determined as key nodes, and alarm configuration is performed according to the recommended timeliness threshold of the key nodes. Compared with the scheme of configuring alarms for nodes of the entire data operation link or the scheme of configuring alarms only for the terminal job in the prior art, the recommended timeliness threshold of the job is efficiently solved through the multi-source breadth-first search algorithm of the graph in the embodiment of the present invention, the communication cost between the terminal operation system and the upstream operation system is reduced, the efficiency of the operation timeliness calculation is improved, and all key nodes on the data operation link can be automatically identified, and alarm configuration is performed according to the recommended timeliness threshold of the key nodes, thereby improving the accuracy and efficiency of the data operation link alarm, avoiding the generation of alarm storms, improving the operation and maintenance effect, and reducing the operation and maintenance pressure.
[0041] With the rapid development of banking business, more and more data processing jobs need to be operated and managed in big data development scenarios. Moreover, these data processing jobs have many tasks, complex dependencies, and span multiple systems, which brings great challenges to the timeliness monitoring of jobs. In order to simplify the job dependencies in the development scenario, the inventor uses graph theory knowledge to abstract data processing jobs into data job links, which not only clearly shows the dependencies between jobs, but also greatly simplifies the communication and collaboration between downstream and upstream job systems.
[0042] In the embodiment of the present invention, a directed acyclic graph is constructed with jobs as nodes and dependencies between jobs as edges; wherein the direction of the edge is from the predecessor job to the corresponding successor job.
[0043] In an embodiment of the present invention, the value of the terminal job node is initialized as the timeliness threshold of the terminal job; the average execution time of the predecessor job corresponding to the edge is determined as the weight of the edge; wherein the terminal job node is a node with an out-degree of 0, and the timeliness threshold of the terminal job represents the time length allowed for delayed execution of the terminal job.
[0044] Figure 2 FIG. 1 is a job dependency diagram of a data job link in an embodiment of the present invention. For example, Figure 2 As shown, there are 3 terminal operation nodes (terminal operation 1-terminal operation 3), 14 upstream operation nodes (operation 1-operation 14), and 4 data operation systems (system A, system B, system C, system D); taking the red terminal operation node in the figure as an example, the time limit threshold of terminal operation 1 is 32 hours, and the time limit threshold of terminal operation 2 is 30 hours, then the value of the node corresponding to terminal operation 1 is initialized to the operation time limit threshold -32, and the value of the node corresponding to terminal operation 2 is initialized to -30 (the format of the node value can also be: T+N hh:mm:ss, then the value of the node corresponding to terminal operation 1 is T+1 8:00:00, and the value of the node corresponding to terminal operation 2 is T+1 6:00:00; wherein T represents the current date, N represents the number of days allowed for delay, and hh:mm:ss represents the hours: minutes: seconds allowed for delay).
[0045] In one embodiment, before determining the average execution time of the predecessor job corresponding to the edge as the weight of the edge, it also includes: collecting the historical execution time of the job; removing the outliers of the historical execution time of the job; calculating the average value of the historical execution time of the job after removing the outliers, to obtain the average execution time of the job.
[0046] The inventors improved the multi-source breadth-first search process, so that while performing a multi-source breadth-first search on a directed acyclic graph, the values of each node in the directed acyclic graph are calculated, and the communication cost between the downstream operation system and the upstream operation system is greatly reduced by using a mature graph search algorithm. To this end, the inventors designed the weight of the directed edge in the directed acyclic graph to be the execution time of the corresponding operation at the starting point of the directed edge.
[0047] In one embodiment, before determining the average execution time of the predecessor job corresponding to the edge as the weight of the edge, it also includes: starting from the terminal job node, performing a multi-source breadth-first search on the directed acyclic graph to obtain the upstream dependent job set of the terminal job; retaining the nodes in the directed acyclic graph that belong to the upstream dependent job set, and deleting the nodes in the directed acyclic graph that do not belong to the upstream dependent job set.
[0048] Since existing big data development scenarios usually involve multiple terminal jobs, if the entire data job chain is searched when only part of the terminal jobs are analyzed, it will cause a waste of computing power. To this end, in this embodiment, before assigning weights to edges, a multi-source breadth-first search algorithm of the graph is first used to identify all upstream dependent job sets of the terminal job to be analyzed, and then the job nodes that do not belong to the upstream dependent job set of the terminal job to be analyzed are deleted, thereby reducing the waste of computing power.
[0049] In the embodiment of the present invention, job nodes that belong to different data processing systems and have dependency relationships are determined as key nodes based on the data processing systems to which the jobs belong and the dependency relationships between the jobs.
[0050] In the existing big data development scenarios, it is of great significance to ensure that each job can be completed within the time limit, including: ensuring data timeliness to meet business needs; reasonably allocating computing resources and storage resources according to job time requirements and resource requirements; locating performance bottlenecks, and optimizing performance in a targeted manner to improve the overall performance of the development scenario. However, in the data link, the link is usually long and has many cross-system node dependencies, and the communication cost between nodes is high. Therefore, how to locate the problem job in time is a major problem that needs to be solved urgently.
[0051] In one embodiment, based on the data processing system to which the job belongs and the dependency relationship between jobs, job nodes that belong to different data processing systems and have dependencies are determined as key nodes, including: based on the data processing system to which the job belongs and the dependency relationship between jobs, a job that has a different data processing system from that of a subsequent dependent job is determined as an exit job, and a job that has a different data processing system from that of a preceding dependent job is determined as an entry job; and nodes corresponding to the exit job and the entry job are determined as key nodes.
[0052] Big data development scenarios usually involve multiple data processing systems. Since the communication cost between cross-system operations is much higher than the communication cost between operations within a system, in order to quickly locate the operation node with problems, the inventors identify the operations that interact between different systems as key nodes. According to the operation time limit of the key node, the data processing system where the problem operation is located can be quickly locked, thereby locating the problem node in time and reducing the operation and maintenance pressure of the operation. Specifically, Figure 2 As shown, the yellow nodes in the figure are key nodes in the upstream job set of terminal job 1 and terminal job 2; among them, job 2, job 4, job 5, job 10, and job 12 are entry jobs, and job 6, job 7, and job 13 are exit jobs.
[0053] Figure 3 It is a weighted job dependency graph of a data job link in an embodiment of the present invention. Figure 3The data operation link in is terminal operation 1, terminal operation 2, and their corresponding upstream dependent operation set. In this example, the value of the node corresponding to terminal operation 1 is -31 (T+1 7:00:00), and the value of the node corresponding to terminal operation 2 is -34 (T+1 10:00:00). There are 6 operation nodes in the upstream dependent operation set of the two terminal operation nodes. Figure 3 It can be seen that the average execution time of job 1 is 2 hours, the average execution time of job 2 is 5 hours, the average execution time of job 3 is 2 hours, the average execution time of job 4 is 8 hours, the average execution time of job 5 is 2 hours, and the average execution time of job 6 is 3 hours.
[0054] In an embodiment of the present invention, starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph, and the value of each node is updated according to the following rules: the weight of the edge pointing to the current search node is added to the value of the current search node; if the sum is greater than the value of the starting node of the edge, or the starting node of the edge is not assigned a value, the value of the starting node of the edge is updated to the sum; wherein the updated value of each node is the recommended timeliness threshold of the job corresponding to each node.
[0055] In one embodiment, before performing a multi-source breadth-first search on the directed acyclic graph starting from the terminal job node, it also includes: performing a reverse topological sorting on the directed acyclic graph to obtain a job order; starting from the terminal job node, performing a multi-source breadth-first search on the directed acyclic graph according to the job order.
[0056] For example, first, the job set is sorted in the reverse topological order of the directed acyclic graph to obtain a job order (the job order obtained here follows the front-to-back dependency relationship of the directed graph to ensure that the job nodes sorted later are completed first); then, according to the obtained job order, a multi-source breadth-first search is performed on the directed acyclic graph; in the process of multi-source breadth-first search, the value of each node is updated according to the predetermined rules, and the updated node value is the recommended timeliness threshold of the job corresponding to the node.
[0057] exist Figure 3 An inverse topology sequence of the embodiment shown is: terminal job 1, terminal job 2, job 1, job 4, job 2, job 3, job 5, job 6. The specific update process of the value of each node is as follows:
[0058] Step 1 (current search node: terminal job 1): -31+2=-29, the value of the node corresponding to job 1 is -29 (T+15:00:00);
[0059] Step 2 (current search node: terminal job 2): -34+8=-26, the value of the node corresponding to job 4 is -26 (T+12:00:00);
[0060] Step 3 (current search node: Job 1): -29+5=-24, the value of the node corresponding to Job 2 is -24 (T+10:00:00); -29+2=-27, the value of the node corresponding to Job 3 is -27 (T+1: 3:00:00);
[0061] Step 4 (current search node: job 4): -26+2=-24, the value of the node corresponding to job 5 is -24 (T+024:00:00);
[0062] Step 5 (Current search node: Job 2): Job 2 has no upstream jobs and does not need to be calculated;
[0063] Step 6 (current search node: job 3): -27+2=-25<-24, the value of the node corresponding to job 5 remains unchanged at -24 (T+024:00:00);
[0064] Step 7 (current search node: job 5): -24+3=-20, the value of the node corresponding to job 6 is -20 (T+020:00:00);
[0065] Step 8 (Current search node: Job 6): Job 6 has no upstream jobs and does not need to be calculated.
[0066] Figure 4 Flow chart of the key operation time efficiency analysis in the embodiment of the present invention. Figure 4 As shown in the figure, the key task time efficiency analysis can be implemented by following the steps below:
[0067] Step 401, operation threshold configuration, corresponding to step 102, initializing the data link;
[0068] Step 402: link operation analysis, using a multi-source breadth-first search algorithm to simplify the data link;
[0069] Step 403, key node analysis, corresponding to step 103, determines the key nodes in the data link;
[0070] Step 404, timeliness threshold analysis, corresponds to step 104, and obtains the recommended timeliness threshold of each job in the data link.
[0071] In the embodiment of the present invention, an alarm threshold is configured for a job corresponding to the key node according to the recommended timeliness threshold of the key node.
[0072] For example, by performing a timeliness analysis on the data link in the above manner, we can eventually obtain the recommended timeliness thresholds of all upstream dependent jobs of the terminal job, and then filter out the recommended timeliness thresholds of the key nodes from the recommended timeliness thresholds of all nodes, and automatically generate alarm configuration information for the data processing jobs corresponding to the key nodes for monitoring.
[0073] The present invention also provides an alarm configuration device for a data operation link, as described in the following embodiments. Since the principle of solving the problem by the device is similar to the alarm configuration method for a data operation link, the implementation of the device can refer to the implementation of the alarm configuration method for a data operation link, and the repeated parts will not be repeated.
[0074] Figure 5 FIG. 1 is a schematic diagram of an alarm configuration device for a data operation link in an embodiment of the present invention. Figure 5 As shown, the device comprises:
[0075] The link construction module 501 is used to construct a directed acyclic graph with jobs as nodes and dependencies between jobs as edges; wherein the direction of the edge is from the predecessor job to the corresponding successor job;
[0076] The link initialization module 502 is used to initialize the value of the terminal job node as the timeliness threshold of the terminal job; the average execution time of the predecessor job corresponding to the edge is determined as the weight of the edge; wherein the terminal job node is a node with an out-degree of 0, and the timeliness threshold of the terminal job indicates the time length allowed for the terminal job to be delayed in execution;
[0077] The key node analysis module 503 is used to determine the job nodes that belong to different data processing systems and have dependencies as key nodes according to the data processing systems to which the jobs belong and the dependencies between the jobs;
[0078] The timeliness threshold analysis module 504 is used to perform a multi-source breadth-first search on the directed acyclic graph starting from the terminal operation node, and update the value of each node according to the following rules: add the weight of the edge pointing to the current search node to the value of the current search node; if the sum of the added values is greater than the value of the starting node of the edge, or the starting node of the edge is not assigned a value, then update the value of the starting node of the edge to the sum; wherein the updated value of each node is the recommended timeliness threshold of the operation corresponding to each node;
[0079] The alarm configuration module 505 is used to configure an alarm threshold for the job corresponding to the key node according to the recommended timeliness threshold of the key node.
[0080] In one embodiment, the link initialization module 502 is further configured to, before determining the average execution time of the predecessor job corresponding to the edge as the weight of the edge:
[0081] The historical execution time of the collection operation;
[0082] Remove outliers in the historical execution time of jobs;
[0083] Calculate the average of the historical execution time of the job after removing outliers to obtain the average execution time of the job.
[0084] In one embodiment, the link initialization module 502 is further configured to, before determining the average execution time of the predecessor job corresponding to the edge as the weight of the edge:
[0085] Starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph to obtain the upstream dependent job set of the terminal job;
[0086] The nodes in the directed acyclic graph that belong to the upstream dependent job set are retained, and the nodes in the directed acyclic graph that do not belong to the upstream dependent job set are deleted.
[0087] In one embodiment, the key node analysis module 503 is specifically used to:
[0088] According to the data processing system to which the job belongs and the dependency relationship between the jobs, the job that has a different data processing system from the post-dependent job is determined as the exit job, and the job that has a different data processing system from the pre-dependent job is determined as the entry job;
[0089] The nodes corresponding to the export operation and the import operation are determined as key nodes.
[0090] In one embodiment, the time threshold analysis module 504 is specifically used to:
[0091] Perform reverse topological sorting on the directed acyclic graph to obtain the order of operations;
[0092] Starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph according to the job order.
[0093] An embodiment of the present invention further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned alarm configuration method for the data operation link when executing the computer program.
[0094] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the alarm configuration method for the data operation link is implemented.
[0095] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the alarm configuration method of the data operation link is implemented.
[0096] To summarize, in an embodiment of the present invention, a directed acyclic graph is constructed with jobs as nodes and dependencies between jobs as edges; the value of the terminal job node is initialized as the timeliness threshold of the terminal job; a multi-source breadth-first search is performed with the terminal job node as the starting point to obtain a set of upstream dependent jobs and simplify the directed acyclic graph; after denoising the execution time of the job, the average execution time of the job is calculated; the average execution time of the predecessor job corresponding to the edge is determined as the weight of the edge; according to the data processing system to which the job belongs and the dependency relationship between the jobs, job nodes that belong to different data processing systems and have dependency relationships are determined as key nodes; starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph in the order of inverse topological sorting to update the value of each node. Compared with the technical solutions in the prior art that configure alarms for nodes in the entire data operation chain, or the technical solutions that configure alarms only for terminal operations, in the embodiments of the present invention, a directed acyclic graph is constructed based on operations and operation dependencies, and graph theory knowledge is used to reduce the communication cost between downstream systems and upstream systems, and improve the efficiency of operation time calculation. In view of the problem of multiple cross-system node dependencies in existing big data development scenarios, data processing operations that interact between systems are identified as key nodes, all key nodes on the data operation chain are automatically identified, and alarms are configured according to the recommended time thresholds of the key nodes, thereby improving the accuracy and efficiency of data operation chain alarms, avoiding alarm storms, improving operation and maintenance effects, and reducing operation and maintenance pressure.
[0097] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0098] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0099] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0101] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for configuring an alarm for a data operation link, characterized in that: include: A directed acyclic graph is constructed with jobs as nodes and dependencies between jobs as edges. The direction of the edge is from the predecessor job to the corresponding successor job. Initialize the value of the terminal job node as the timeliness threshold of the terminal job; determine the average execution time of the predecessor job corresponding to the edge as the weight of the edge; where the terminal job node is a node with an out-degree of 0, and the timeliness threshold of the terminal job indicates the time allowed for the terminal job to be delayed; According to the data processing systems to which the jobs belong and the dependencies between the jobs, job nodes that belong to different data processing systems and have dependencies are identified as key nodes; Starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph, and the value of each node is updated according to the following rules: the weight of the edge pointing to the current search node is added to the value of the current search node; if the sum is greater than the value of the starting node of the edge, or the starting node of the edge is not assigned a value, the value of the starting node of the edge is updated to the sum; the updated value of each node is the recommended timeliness threshold of the job corresponding to each node; Configure alarm thresholds for jobs corresponding to key nodes based on the recommended timeliness thresholds for key nodes.
2. The method according to claim 1, characterized in that Before determining the average execution time of the predecessor job corresponding to the edge as the weight of the edge, the following is also included: The historical execution time of the collection operation; Remove outliers from the historical execution time of jobs; Calculate the average of the historical execution time of the job after removing outliers to obtain the average execution time of the job.
3. The method according to claim 1, characterized in that Before determining the average execution time of the predecessor job corresponding to the edge as the weight of the edge, the following is also included: Starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph to obtain the upstream dependent job set of the terminal job; The nodes in the directed acyclic graph that belong to the upstream dependent job set are retained, and the nodes in the directed acyclic graph that do not belong to the upstream dependent job set are deleted.
4. The method according to claim 1, characterized in that According to the data processing systems to which the jobs belong and the dependencies between the jobs, the job nodes that belong to different data processing systems and have dependencies are identified as key nodes, including: According to the data processing system to which the job belongs and the dependency relationship between the jobs, the job that has a different data processing system from the post-dependent job is determined as the exit job, and the job that has a different data processing system from the pre-dependent job is determined as the entry job; The nodes corresponding to the export operation and the import operation are determined as key nodes.
5. The method according to claim 1, characterized in that Starting from the terminal operation node, a multi-source breadth-first search is performed on the directed acyclic graph, including: Perform reverse topological sorting on the directed acyclic graph to obtain the order of operations; Starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph according to the job order.
6. An alarm configuration device for a data operation link, characterized in that: include: The link construction module is used to construct a directed acyclic graph with jobs as nodes and dependencies between jobs as edges. The direction of the edge is from the predecessor job to the corresponding successor job. The link initialization module is used to initialize the value of the terminal job node to the timeliness threshold of the terminal job; the average execution time of the predecessor job corresponding to the edge is determined as the weight of the edge; the terminal job node is a node with an out-degree of 0, and the timeliness threshold of the terminal job indicates the time allowed for the terminal job to be delayed; A key node analysis module is used to determine the job nodes that belong to different data processing systems and have dependencies as key nodes according to the data processing systems to which the jobs belong and the dependencies between the jobs; The timeliness threshold analysis module is used to perform a multi-source breadth-first search on the directed acyclic graph starting from the terminal job node, and update the value of each node according to the following rules: add the weight of the edge pointing to the current search node to the value of the current search node; if the sum is greater than the value of the starting node of the edge, or the starting node of the edge is not assigned a value, then update the value of the starting node of the edge to the sum; the updated value of each node is the recommended timeliness threshold of the job corresponding to each node; The alarm configuration module is used to configure the alarm threshold for the job corresponding to the key node according to the recommended time threshold of the key node.
7. The device according to claim 6, characterized in that The link initialization module is also used to determine the average execution time of the predecessor job corresponding to the edge as the weight of the edge: The historical execution time of the collection operation; Remove outliers from the historical execution time of jobs; Calculate the average of the historical execution time of the job after removing outliers to obtain the average execution time of the job.
8. The device according to claim 6, characterized in that The link initialization module is also used to determine the average execution time of the predecessor job corresponding to the edge as the weight of the edge: Starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph to obtain the upstream dependent job set of the terminal job; The nodes in the directed acyclic graph that belong to the upstream dependent job set are retained, and the nodes in the directed acyclic graph that do not belong to the upstream dependent job set are deleted.
9. The device according to claim 6, characterized in that Key node analysis module, specifically used for: According to the data processing system to which the job belongs and the dependency relationship between the jobs, the job that has a different data processing system from the post-dependent job is determined as the exit job, and the job that has a different data processing system from the pre-dependent job is determined as the entry job; The nodes corresponding to the export operation and the import operation are determined as key nodes.
10. The device according to claim 6, characterized in that The time threshold analysis module is specifically used for: Perform reverse topological sorting on the directed acyclic graph to obtain the order of operations; Starting from the terminal job node, a multi-source breadth-first search is performed on the directed acyclic graph according to the job order.
11. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
13. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Data operation adjustment method and device, computer equipment and storage medium
CN115481109A
Operation timeliness alarm method, device and system and medium
CN115934485A
Data processing method and device, equipment and storage medium
CN118153931A
Scheduling software jobs having dependencies
US20200073727A1