Data Warehouse Task Monitoring Method, System, Device and Storage Medium
By judging the criticality of the underlying tasks in the data warehouse, and using the dependency and blood relationship map, only monitoring of critical tasks is solved, the problem of errors in the execution of underlying tasks is achieved, and the precise delivery of resources and real-time data adjustment is achieved.
Patent Information
- Application Number
- CN202210862696.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-07-21
AI Technical Summary
The existing technology cannot know the impact of upstream tasks in data warehouses on the underlying tasks in real time, resulting in errors in the execution of underlying tasks and excessive resource consumption for monitoring all tasks.
By obtaining the preset table parameter information of the underlying task, we judge whether it is a critical task, and using the dependency map and blood relationship map of the data warehouse, we only monitor the critical tasks, obtain the upstream blood relationship tasks, and realize real-time monitoring of data changes.
It improves the security and reliability of key data, reduces resource consumption, realizes timely adjustments to critical tasks, and reduces resource costs.
Smart Images

Figure CN115098336B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a method, system, device and storage medium for monitoring data warehouse tasks. Background Art
[0002] The lineage of existing big data warehouse tasks includes five layers from top to bottom according to the data flow direction, namely ODS (Operational Data Store), DWD (Data Warehouse Detail), DWM (Data Warehouse Middle), DWS (Data Warehouse Service), and ADS (Application Data Service). After the data in the data source undergoes extraction, cleaning, and transmission, that is, the ETL process, it enters the ODS layer. After the ODS performs operations such as denoising and field cleaning on the data, it enters the data warehouse layer, which includes three layers: DWD, DWM, and DWS. After a series of operations such as isolation, aggregation, and analysis, data that can be directly used by the application software in the ADS layer is obtained. Therefore, since the data chain is from top to bottom, changes in the upper-layer data will have a greater impact on the lower-layer data. There may be more than a dozen upstream lineage tasks with dependencies and lineage relationships from the topmost ODS layer task to the bottommost ADS layer task, resulting in a low awareness of the importance of the top layer to the bottom layer tasks. When the upper-layer data changes, the bottommost tasks should also synchronize the changes in a timely manner. However, in the prior art, it is impossible to know the bottom layer tasks that have a lineage relationship with the upstream tasks, resulting in errors when the bottom layer tasks are executed. Summary of the Invention
[0003] The present invention provides a method, system, device and storage medium for monitoring data warehouse tasks. Its main purpose is to extract the tasks in the entire data warehouse that have a lineage relationship with the bottom layer tasks and monitor each task to know the data change situation in real time.
[0004] In a first aspect, the present invention provides a method for monitoring data warehouse tasks, including:
[0005] For the bottom layer tasks, obtain the preset table parameter information in the bottom layer tasks, where the preset table parameter information represents the relevant information of the target database and target table in the data warehouse;
[0006] Match with the keyword field information table according to the preset table parameter information to determine whether the bottom layer task is a critical task;
[0007] If the underlying task is a critical task, all upstream lineage tasks related to the underlying task are obtained according to the underlying task, the dependency graph and the lineage graph of the data warehouse;
[0008] Monitor the underlying task and the upstream lineage tasks.
[0009] Preferably, the keyword field information is obtained in the following manner:
[0010] Extract initial parameter information according to the application logs of the databases in the data warehouse;
[0011] Obtain the score of the initial parameter information according to the initial parameter information and the target classification neural network;
[0012] Filter the keyword field information according to the score of the initial parameter information.
[0013] Preferably, the extracting the initial parameter information according to the application logs of the databases in the data warehouse includes:
[0014] Extract key metric information according to the application logs of the databases;
[0015] Input the key metric information into the target classification neural network to obtain the level corresponding to the key metric information;
[0016] Filter out the initial parameter information according to the level corresponding to the key metric information.
[0017] Preferably, the extracting the key metric information according to the application logs of the databases includes:
[0018] Synchronize the application logs of the databases to the big data cluster and perform data cleaning on the application logs;
[0019] Split the cleaned application logs to obtain field information;
[0020] Obtain the key metric information corresponding to each database according to the field information.
[0021] Preferably, the keyword field information includes the number of table calls.
[0022] Preferably, the preset table parameter information includes the target database type, the target table size, the number of calls of the target table, and the target table PK value in the underlying task.
[0023] Preferably, the underlying task represents the task of the ADS layer in the data warehouse.
[0024] In a second aspect, an embodiment of the present invention provides a data warehouse task monitoring system, including:
[0025] A bottom layer module, which is used for a bottom layer task to obtain preset table parameter information in the bottom layer task, and the preset table parameter information represents relevant information of a target database and a target table in a data warehouse;
[0026] A matching module, which is used to match with a keyword field information table according to the preset table parameter information to determine whether the bottom layer task is a key task;
[0027] A derivation module, which is used to, if the bottom layer task is a key task, obtain all upstream blood relationship tasks related to the bottom layer task according to the bottom layer task, a dependency relationship graph and a blood relationship graph of the data warehouse;
[0028] A monitoring module, which is used to monitor the bottom layer task and the upstream blood relationship tasks.
[0029] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned data warehouse task monitoring method are implemented.
[0030] In a fourth aspect, an embodiment of the present invention provides a computer storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned data warehouse task monitoring method are implemented.
[0031] A data warehouse task monitoring method, system, device and storage medium provided by the present invention extract preset table parameter information in a bottom layer task, match it with keyword field information, and determine whether the bottom layer task is a key task. Since there are many tasks involved in a data warehouse, a lot of resources are required to monitor each task. Therefore, only key tasks are monitored. After selecting key tasks, all upstream blood relationship tasks related to the key bottom layer task are obtained by using the dependency relationship graph and blood relationship graph of the data warehouse, and then monitored. Once relevant data changes, the bottom layer task can be adjusted in a timely manner, thus ensuring the security and reliability of key data. Description of the Drawings
[0032] Figure 1 It is a schematic diagram of an application scenario of a data warehouse task monitoring method provided by an embodiment of the present invention;
[0033] Figure 2 It is a flowchart of a data warehouse task monitoring method provided by an embodiment of the present invention;
[0034] Figure 3 It is a schematic diagram of the structure of a data warehouse task monitoring system provided by an embodiment of the present invention;
[0035] Figure 4 This is a schematic structural diagram of a computer device provided in an embodiment of the present application.
[0036] The realization, functional characteristics, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0037] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0038] Figure 1 This is a schematic diagram of an application scenario of a data warehouse task monitoring method provided in an embodiment of the present invention. As Figure 1 shown, the user determines the underlying tasks of the data warehouse on the client side and sends the underlying tasks to the server side. After receiving the underlying tasks, the server side executes the data warehouse task monitoring method to monitor the underlying tasks and related upstream lineage tasks.
[0039] It should be noted that the server side can be implemented by an independent server or a server cluster composed of multiple servers. The client side can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The client side and the server side can be connected through Bluetooth, USB (Universal Serial Bus), or other communication connection methods, and the embodiments of the present invention do not limit this here.
[0040] The embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI for short) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, method, technology, and application system.
[0041] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning, deep learning, etc.
[0042] Figure 2 This is a flowchart of a data warehouse task monitoring method provided in an embodiment of the present invention. As Figure 2 shown, the method includes:
[0043] S210. For the underlying task, obtain the preset table parameter information in the underlying task, where the preset table parameter information represents the relevant information of the target database and target table in the data warehouse;
[0044] A data warehouse, also known as a data storehouse, is a strategic collection that provides all types of data support for the decision-making processes at all levels of an enterprise. It is a single data storage created for analytical reporting and decision support purposes, providing guidance for business process improvement, monitoring time, cost, quality, and control for enterprises that require business intelligence.
[0045] Underlying task: In the embodiments of the present invention, the underlying task can be determined by the user himself. Generally, it refers to extracting data from the ADS layer to the application software, or extracting data from the upper layer and synchronizing it to the ADS layer. The specific meaning can be determined according to the actual situation, and the embodiments of the present invention do not make specific limitations here. For example, for a certain underlying task of synchronizing data from database A to database C, database A can be any layer in the upper layer, and database C is located in the ADS layer. However, since database C is located at the bottom layer, it may be necessary to complete certain tasks before synchronizing data from database A to database C. It is also possible that some operations occur upstream, changing the data in database A, which will affect the execution of this underlying task. Only by monitoring all tasks in the entire data chain related to the underlying task can the data situation be known in a timely manner. However, if all underlying tasks are monitored, the task volume is too large, and the resources consumed by the server are also too much. In order to balance the calculation volume and monitoring accuracy, it is necessary to screen out the key underlying tasks and only monitor the key underlying tasks.
[0046] In order to monitor the underlying task, first, it is necessary to determine whether the underlying task is a key task, and it is necessary to first extract the preset table parameter information of the underlying task. The preset parameter table information represents the relevant information of the target database and target table in the data warehouse. Specifically, if the underlying task is to synchronize data from database A to table D in database C, then database C is the target database of this underlying task, and table D is the target table of this underlying task. The relevant information can be the type of the target database, the name of the target database, the location of the target database, the size of the target table, the location of the target table, the PK value of the target table, etc. The types of target databases include oracle, mysql, tidb, pg, etc. Since the underlying task will synchronize and write the final result data into the target database used by the data application, there are multiple types of this database and diverse usage methods, so the type of the target database may affect the determination of this key task. It can be specifically determined according to the actual situation, and the embodiments of the present invention do not make specific limitations here.
[0047] S220. According to the preset table parameter information, match it with the keyword field information table to determine whether the underlying task is a critical task;
[0048] Then, according to the preset table parameter information obtained in the previous step, match the budget table parameter information with the keyword field information table to determine whether the underlying task is a critical task. In the embodiments of the present invention, the keyword field information stores the field information corresponding to the critical task, which may include the call times of the table, keyword field information, the access times of the database, etc. The higher these indicators are, the more important this table is, and thus the higher the possibility that the underlying task is a critical task. Matching the preset table parameter information with the keyword field information table, the specific matching methods include keyword matching, exact matching, phrase matching, broad matching, and negative matching, etc., which can be determined according to the actual situation, and the embodiments of the present invention do not make specific limitations in this regard. For example, in the embodiments of the present invention, by matching the table parameters included in the preset table parameter information with the keyword field information one by one, to see whether the keyword field information contains the table parameter. If it contains, it means that the underlying task is a critical task. If it does not contain, it means that the underlying task is not a critical task.
[0049] S230. If the underlying task is a critical task, then according to the underlying task, the dependency graph and the lineage graph of the data warehouse, obtain all the upstream lineage tasks related to the underlying task;
[0050] If this underlying task is a critical task, then according to the dependency graph and the lineage graph of the data warehouse, extract all the upstream lineage tasks related to this underlying task. In the embodiments of the present invention, the dependency graph of the data warehouse represents the hierarchical dependency relationship of big data tasks. For example, for two hierarchical big data tasks a and b, only when the upper-level a task is executed completely, the lower-level b task can be executed. Such a series of dependency relationships from the topmost task to the bottommost task is the dependency graph. The dependency graph of the database can be obtained specifically through the following method: Initialize the subject table of the data warehouse; Obtain the data stream to be warehoused in real time from different business databases; Determine whether the fact object described by the data stream to be warehoused contains the primary key corresponding to the target field of the target subject table to be warehoused. If it contains, it is determined that the data stream to be warehoused is the main table. If it does not contain, it is determined that the data stream to be warehoused is the slave table; Warehousing the main table for the data stream to be warehoused determined as the main table, and warehousing the slave table for the data stream to be warehoused determined as the slave table; Writing the warehoused data stream to the database to obtain a data warehouse with a unified subject layer. This method realizes the construction of a reliable dependency relationship for the data streams of multiple data tables in different business databases, and obtains the dependency graph of the data warehouse.
[0051] A blood relationship graph refers to a relationship graph derived based on information such as tables and fields in a database queried for big data tasks. The blood relationship graph can be used for data traceability. The data used for analysis and processing may come from a wide range of sources, including government data, Internet data, data obtained from third parties through data transactions, and data owned by itself. Data from different sources has uneven quality, and the impact on the results of analysis and processing is also different. When data anomalies occur, it is necessary to be able to trace the reasons for the anomalies and control the risks at an appropriate level. The blood relationship of data reflects the context of the data, which can help trace the source of the data and the data processing process. On the visual graph of the blood relationship of data, the data source nodes are on the left side of the main node, which is very clear and obvious at a glance. It can also be seen from the visual graph which conversions the data has undergone, which is very helpful for analyzing the reasons for the generation of abnormal data.
[0052] Specifically, when using the dependency graph and the blood relationship graph to extract the upstream blood relationship tasks related to the underlying task, it is necessary to first find all the tasks with blood relationships related to the underlying task according to the blood relationship graph. The so-called tasks with blood relationships refer to those that have a blood relationship with the source data in the underlying task. Then, based on all the tasks with blood relationships, according to the dependency graph, find the upstream tasks of each task with a blood relationship, and then use the upstream task and the task with a blood relationship as the upstream blood relationship tasks of the underlying task.
[0053] In the embodiment of the present invention, according to the dependency graph and the blood relationship graph, all the upstream blood relationship tasks related to the underlying task are extracted. Through the blood relationship graph, it is possible to quickly and conveniently find the tasks with a blood relationship with the underlying task. Then, according to the dependency graph, it is possible to quickly and conveniently find the upstream tasks related to the blood relationship tasks, so as to obtain the upstream blood relationship tasks of the underlying task. Since the dependency graph and the blood relationship graph can clearly describe the structural relationship and dependency relationship of the data in the data warehouse, the upstream blood relationship tasks related to the underlying task can be accurately extracted through the dependency graph and the blood relationship graph.
[0054] The embodiment of the present invention complements and supplements the top-down blood relationship and dependency relationship, ensuring the robustness of the data warehouse system, thereby promoting the reliability of the overall big data task.
[0055] S240, monitor the underlying task and the upstream blood relationship tasks.
[0056] Monitor the underlying tasks and upstream lineage tasks. Common monitoring methods include collection monitoring. Each time the collector starts, record the server IP, start time, and collector ID, etc. After obtaining the task set each time, record the start time, end time, identification set of tasks to be collected, collector ID, etc. During the task execution process, record the start time of a single task, download start time, request return code, download end time, parsing time consumption, amount of data parsed, and the current task ID, etc. When all tasks are completed, record the start time, end time, and total amount of data parsed for the current batch of tasks. For the above four aspects of monitoring, a lot of log information will be generated every day. To ensure that log persistence does not affect the collection efficiency, the data is temporarily stored in the Redis cluster. Then, analyze the log information daily and clear the historical logs from a week ago at the same time.
[0057] By monitoring the upstream lineage tasks, once some data in the upstream lineage tasks changes, it will affect the target data or source data in the underlying tasks. The underlying tasks can be immediately refreshed to change the source data in the underlying tasks, so that the underlying tasks can keep up with the changes in the data warehouse in a timely manner. In the existing data warehouses, big data tasks change rapidly. Manually determining the key underlying tasks has a high cost, poor effect, and takes a long time. The embodiments of the present invention reduce costs and increase efficiency because key tasks can be sorted out more quickly, resource tilting can be done well to ensure key tasks, and at the same time non-key tasks can be cleaned up in a timely manner, saving resource costs and achieving precise resource allocation.
[0058] A data warehouse task monitoring method proposed by the present invention determines whether the underlying task is a key task by extracting the preset table parameter information in the underlying task and matching it with the keyword field information. Since there are many tasks involved in a data warehouse, a lot of resources are required to monitor each task. Therefore, only key tasks are monitored. After selecting the key tasks, use the dependency graph and lineage graph of the data warehouse to obtain all upstream lineage tasks related to the key underlying task, and then monitor them. Once the relevant data changes, the underlying task can be adjusted immediately, thus ensuring the security and reliability of the key data.
[0059] Based on the above embodiments, preferably, the keyword field information is obtained through the following method:
[0060] Extract the initial parameter information according to the application logs of the database in the data warehouse;
[0061] Obtain the score of the initial parameter information according to the initial parameter information and the target classification neural network;
[0062] Screen the keyword field information according to the score of the initial parameter information.
[0063] Specifically, the key field information in the embodiments of the present invention can be extracted in the following manner: First, extract the application logs of each database in the data warehouse. The application logs record all relevant operations in the database. In the embodiments of the present invention, the application logs of the database are stored in the form of slices, and keyword extraction is performed on the application logs. This application log refers to the log information of the associated applications for performing insert, delete, update, and query operations using the database into which data is synchronously written by the underlying tasks. This application log can record the SQL information for performing insert, delete, update, and query operations on this database. The query times of the tables in the database can be roughly obtained through the number of SQL statements. The keyword can be preset, generally including the query times of each table in the database and some key fields. The extracted table query times and key fields are used as the initial parameter information. Preset scores for information such as the query times in the tag data. For example, for many, medium, and few query times, many is 2 points, medium is 1 point, and few is 0 points; then, sum up the scores of each index in the tag data; the tag data with a higher score is the tag with the highest level, and vice versa for the low-level tags. When screening the key fields, try to screen the tags with a higher level as much as possible.
[0064] Then, input the initial parameter information into the target classification neural network to obtain the score of each initial parameter information. The higher the score, the more important the initial parameter information is, and this initial parameter information can be used as a key field information. The target classification neural network is a type of neural network. Before using this target classification neural network, it needs to be trained first.
[0065] The target classification neural network in the embodiments of the present invention belongs to a type of neural network. Before using this target classification neural network, it also needs to be trained. The target classification neural network is trained with the pre-obtained samples and tags. The training process of this target classification neural network can be divided into three steps: Define the structure of the target classification neural network and the output result of the forward propagation; Define the loss function and the algorithm for backpropagation optimization; Finally, generate a session and repeatedly run the backpropagation optimization algorithm on the training data.
[0066] Among them, a neuron is the smallest unit that constitutes a neural network. A neuron can have multiple inputs and one output. The input of each neuron can be either the output of other neurons or the input of the entire neural network. The output of this neural network is the weighted sum of the inputs of all neurons. The weights of different inputs are the neuron parameters. The optimization process of the neural network is the process of optimizing the values of the neuron parameters.
[0067] The effect and optimization goal of a neural network are defined by a loss function. The loss function gives a formula for calculating the gap between the output result of the neural network and the true label. Supervised learning is a way of training a neural network. The idea is that on a labeled dataset with known answers, the result given by the neural network should be as close as possible to the true answer (i.e., the label). By adjusting the parameters in the neural network to fit the training data, the neural network can provide prediction ability for unknown samples.
[0068] The backpropagation algorithm implements an iterative process. At the beginning of each iteration, a portion of the training data is taken, and the prediction result of the neural network is obtained through the forward propagation algorithm. Since the training data all have correct answers, the gap between the prediction result and the correct answer can be calculated. Based on this gap, the backpropagation algorithm will correspondingly update the values of the neural network parameters to make them closer to the true answer.
[0069] After completing the training process through the above method, the trained target classification neural network can be used for applications.
[0070] Based on the above embodiments, preferably, extracting the initial parameter information according to the application logs of the databases in the data warehouse includes:
[0071] Extracting key metric information according to the application logs of the database;
[0072] Inputting the key metric information into the target classification neural network to obtain the level corresponding to the key metric information;
[0073] Filtering out the initial parameter information according to the level corresponding to the key metric information.
[0074] Model the application log information for the production and operation using the data in the database. Modeling steps: 1. Extract key metric information; 2. Select a classification model (neural network) for the extracted key information (such as the table name, field name, query times, query period, etc. of the queried database) to model the data; 3. Train the model, continuously adjust the parameters, and determine the optimal parameters according to the training results; 4. Evaluate the model results. The evaluation criterion is to accurately classify the tables in the database into levels and form label information, for example, the label of table a is (table level: important table, table use: user information, query times: 10000, number of table rows: 1000).
[0075] Specifically, the process of extracting initial parameter information from the application logs of the database is somewhat similar to the process of extracting key field information described above. Both start by extracting key metric information from the application logs of the database, then input the key metric information into the target classification neural network to obtain the level corresponding to the key metric information. The level corresponding to the key metric information describes the classification based on factors such as the number of queries, time periods, and corresponding services. Finally, through shuffling and discretization, these data feature labels are arranged and combined to form a feature label feature example with levels and arrangement methods, and the result is output to obtain the initial parameter information.
[0076] Shuffling is to prevent the situation where the classification is inaccurate due to the over-concentration of certain classes of data before the data set enters the target classification neural network for classification. Therefore, the data set of the classified data labels will be shuffled here. For example, through the Python code np.random.shuffle, the label data in the data set can be split into an irregular set. Discretization is a manifestation when these labeled data are added to the target classification neural network for classification. Discretization re-classifies the labels with relatively close scoring ranges in the label data through discretization encoding. For example, 100 - 90 points and 90 - 80 points are two score ranges, but both belong to the high-score range, and they can be classified and mapped to a feature area.
[0077] Based on the above embodiments, preferably, the extracting of key metric information according to the application logs of the database includes:
[0078] Synchronize the application logs of the database to the big data cluster and perform data cleaning on the application logs;
[0079] Split the cleaned application logs to obtain field information;
[0080] Based on the field information, obtain the key metric information corresponding to each database.
[0081] The process and role of extracting key metric information from the application logs of the database are as follows: Information extraction process and role:
[0082] Synchronize the application logs in the database to the big data Hadoop cluster. Use synchronization tools such as Filebeat, FileSync, and ETL to ingest the logs into the Hadoop cluster and filter and clean the log-format data therein. Write code according to the format for the logs to split the log information and form field information (such as date, executed SQL, execution time consumption, returned result data volume, whether there is an error, error message, etc.). These data can be stored in the Hadoop cluster or other relational databases to obtain the key metric information corresponding to each database.
[0083] Based on the above embodiments, preferably, the keyword field information includes the number of table calls.
[0084] Specifically, the keyword field information includes the number of table calls, that is, by including the number of table calls in the extracted preset table parameter information. By extracting the number of calls of the target table, if the number of calls is greater than the preset call threshold, it indicates that the table is relatively important, that is, the underlying task is a key task. If the number of calls is not greater than the preset call threshold, it indicates that the table is not important, that is, the underlying task is not a key task. The number of calls can indicate the frequency of the table being called. If the call is relatively frequent and the number of calls of the target table in the underlying task is high, it indicates that the underlying task is relatively critical.
[0085] Based on the above embodiments, preferably, the preset table parameter information includes the target database type, target table size, number of calls of the target table, and target table PK value in the underlying task.
[0086] In the embodiments of the present invention, the budget table parameter information includes the target database type, target table size, number of calls of the target table, and target table PK value in the underlying task. The PK value comes from the PK value to be written into the database table in the underlying task. The PK value is the primary key information (primary key) of this table, and the PK values of different tables are different.
[0087] Figure 3 For a structural schematic diagram of a data warehouse task monitoring system provided by an embodiment of the present invention, as Figure 3 shown, the system includes an underlying task 310, a matching module 320, a derivation module 330, and a monitoring module 340, where:
[0088] The underlying module 310 is used for the underlying task to obtain the preset table parameter information in the underlying task, and the preset table parameter information represents the relevant information of the target database and target table in the data warehouse;
[0089] The matching module 320 is used to match with the keyword field information table according to the preset table parameter information, and determine whether the underlying task is a critical task;
[0090] The derivation module 330 is used to, if the underlying task is a critical task, obtain all upstream lineage tasks related to the underlying task according to the underlying task, the dependency graph and the lineage graph of the data warehouse;
[0091] The monitoring module 340 is used to monitor the underlying task and the upstream lineage tasks.
[0092] This embodiment is a system embodiment corresponding to the above method embodiment, and its specific implementation process is the same as that of the above method embodiment. For details, please refer to the above method embodiment, and this system embodiment will not be elaborated here.
[0093] On the basis of the above embodiment, preferably, the matching module includes an extraction unit, a scoring unit and a screening unit, where:
[0094] The extraction unit is used to extract initial parameter information according to the application logs of the databases in the data warehouse;
[0095] The scoring unit is used to obtain the score of the initial parameter information according to the initial parameter information and the target classification neural network;
[0096] The screening unit is used to screen keyword field information according to the score of the initial parameter information.
[0097] On the basis of the above embodiment, preferably, the extraction unit includes an index unit, a level unit and a parameter unit, where:
[0098] The index unit is used to extract key index information according to the application logs of the databases;
[0099] The level unit is used to input the key index information into the target classification neural network to obtain the level corresponding to the key index information;
[0100] The parameter unit is used to screen out the initial parameter information according to the level corresponding to the key index information.
[0101] On the basis of the above embodiment, preferably, the index unit includes a synchronization unit, a splitting unit and an information unit, where:
[0102] The synchronization unit is used to synchronize the application logs of the databases to the big data cluster and perform data cleaning on the application logs;
[0103] The splitting unit is used to split the cleaned application logs to obtain field information;
[0104] The information unit is used to obtain the key index information corresponding to each database according to the field information.
[0105] Based on the above embodiments, preferably, the keyword field information includes the number of table calls.
[0106] Based on the above embodiments, preferably, the preset table parameter information includes the target database type, target table size, number of calls of the target table, and target table PK value in the underlying task.
[0107] Based on the above embodiments, preferably, the underlying task represents the task of the ADS layer in the data warehouse.
[0108] Each module in the above data warehouse task monitoring system can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0109] Figure 4 The structure diagram of a computer device provided in an embodiment of the present application. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a computer storage medium and an internal memory. The computer storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the computer storage medium. The database of the computer device is used to store the data generated or obtained during the execution of the data warehouse task monitoring method, such as preset table parameter information, keyword field information, dependency relationship graph, and lineage relationship graph. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a data warehouse task monitoring method.
[0110] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the data warehouse task monitoring method in the above embodiments. Or, when the processor executes the computer program, it implements the functions of each module / unit in the embodiment of the data warehouse task monitoring system.
[0111] In one embodiment, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data warehouse task monitoring method in the above embodiment are implemented. Alternatively, when the computer program is executed by a processor, the functions of each module / unit in the above embodiment of the data warehouse task monitoring system are implemented.
[0112] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0113] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0114] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for monitoring data warehouse tasks, characterized in that, Including: For the underlying task, obtain the preset table parameter information in the underlying task, where the preset table parameter information represents the relevant information of the target database and target table in the data warehouse; Match with the keyword field information table according to the preset table parameter information to determine whether the underlying task is a critical task; If the underlying task is a critical task, obtain all upstream blood relationship tasks related to the underlying task according to the underlying task, the dependency relationship graph and the blood relationship graph of the data warehouse; Monitor the underlying task and the upstream blood relationship tasks; The keyword field information is obtained through the following method: Extract the initial parameter information according to the application log of the database in the data warehouse; Obtain the score of the initial parameter information according to the initial parameter information and the target classification neural network; Filter the keyword field information according to the score of the initial parameter information.
2. The data warehouse task monitoring method according to claim 1, wherein The extracting the initial parameter information according to the application log of the database in the data warehouse includes: Extract the key indicator information according to the application log of the database; Input the key indicator information into the target classification neural network to obtain the level corresponding to the key indicator information; Filter out the initial parameter information according to the level corresponding to the key indicator information.
3. The data warehouse task monitoring method according to claim 2, wherein The extracting the key indicator information according to the application log of the database includes: Synchronize the application log of the database to the big data cluster and perform data cleaning on the application log; Split the cleaned application log to obtain field information; Obtain the key indicator information corresponding to each database according to the field information.
4. The data warehouse task monitoring method according to claim 1, wherein The keyword field information includes the table call times.
5. The data warehouse task monitoring method according to any one of claims 1 to 4, characterized in that, The preset table parameter information includes the target database type, target table size, call times of the target table, and target table PK value in the underlying task.
6. The data warehouse task monitoring method according to claim 1, wherein The underlying task represents the task of the ADS layer in the data warehouse.
7. A data warehouse task monitoring system, characterized in that, Including: The underlying module is used to obtain the preset table parameter information in the underlying task for the underlying task, where the preset table parameter information represents the relevant information of the target database and target table in the data warehouse; The matching module is used to match with the keyword field information table according to the preset table parameter information to determine whether the underlying task is a critical task; The derivation module is used to, if the underlying task is a critical task, obtain all upstream blood relationship tasks related to the underlying task according to the underlying task, the dependency relationship graph and the blood relationship graph of the data warehouse; The monitoring module is used to monitor the underlying task and the upstream blood relationship tasks; The keyword field information is obtained through the following method: Extract the initial parameter information according to the application log of the database in the data warehouse; Obtain the score of the initial parameter information according to the initial parameter information and the target classification neural network; Filter the keyword field information according to the score of the initial parameter information.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the data warehouse task monitoring method as described in any one of claims 1 to 6.
9. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the data warehouse task monitoring method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Data processing method and device for data warehouse, medium and computing equipment
CN111966692A