Blood relationship analysis method and system based on log analysis and storage medium

Through the method based on log analysis, a data link model is constructed and Kettle conversion logs are analyzed, which solves the shortcomings of Kettle software in blood analysis and realizes accurate identification and visualization of data processing paths.

CN120029983APending Publication Date: 2025-05-23AISINO CORPORATION +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411915851.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Kettle software lacks blood relationship analysis functions in the field of data governance, making it difficult for users to intuitively understand all links in the data processing process. The existing methods are not accurate enough to analyze the system, and the system is complex and difficult to maintain.

Method used

Using a blood relationship analysis method based on log analysis, the conversion log is recorded by performing Kettle conversion tasks, incremental logs are determined, log data is gathered and cleaned, data link model is constructed, input and output tables are analyzed, and blood relationship data is obtained.

Benefits of technology

It realizes accurate identification and visualization of Kettle software data processing paths, simplifies the complexity of the blood ties tracking system, and makes up for the shortcomings of Kettle software in blood ties analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029983A_ABST
    Figure CN120029983A_ABST
Patent Text Reader

Abstract

The invention discloses a blood relationship analysis method and system based on log analysis and a storage medium, and the method comprises the steps: executing a Kettle conversion task, recording a conversion log in a data conversion process, and obtaining an execution log; determining incremental execution logs according to the execution time of the execution logs, and converging the incremental execution logs; reading an execution log in a log analysis database, performing log cleaning based on a conversion name and execution time of the execution log, and storing the log in an intermediate data table; and analyzing an input table and an output table in the intermediate data table by taking a data chain as a dimension, and removing repeated data to obtain blood relationship data. According to the method, the data consanguinity of the data processed by the Kettle in the data governance process can be accurately identified, so that a convenient method and tool are provided for a user to track a data processing chain, the complexity of a consanguinity tracking system is simplified, and the defects of Kettle in consanguinity analysis are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data governance technology, and more specifically, to a bloodline analysis method, system and storage medium based on log analysis. Background Art

[0002] Kettle is an open source ETL (data extraction, transformation, loading) tool commonly used in the field of data governance, which is used to collect, clean and distribute data. However, Kettle lacks lineage analysis function, and users cannot intuitively understand the various links of the final data processing.

[0003] At present, the method of adding blood relationship analysis function to Kettle is generally to analyze the input and output relationship of the table by analyzing the xml file converted by Kettle. However, due to the complexity of the data processing process, the current method is not accurate enough. At the same time, the blood relationship analysis system also has complex analysis methods and is difficult to implement and maintain.

[0004] Therefore, a lineage analysis method based on log analysis is needed. Summary of the invention

[0005] The present invention proposes a blood relationship analysis method, system and storage medium based on log analysis to solve the problem of how to perform blood relationship analysis based on logs.

[0006] In order to solve the above problem, according to one aspect of the present invention, a blood relationship analysis method based on log analysis is provided, the method comprising:

[0007] Execute Kettle conversion tasks, record conversion logs during data conversion, and obtain execution logs;

[0008] Determine the incremental execution logs according to the execution time of the execution logs, aggregate the incremental execution logs, and save them uniformly in the log analysis database;

[0009] Read the execution log in the log analysis database, clean the log based on the conversion name and execution time of the execution log, and save it in the intermediate data table;

[0010] The input table and the output table in the intermediate data table are analyzed with the data chain as the dimension, and duplicate data are removed to obtain blood relationship data.

[0011] Preferably, the method further comprises:

[0012] Construct a data link model; wherein the data link model includes: node information, database connection information, and connection information and direction of front and rear nodes.

[0013] Preferably, the method further comprises:

[0014] Execute Kettle conversion tasks at preset time intervals, or execute Kettle conversion tasks at specified times according to fixed cycles.

[0015] Preferably, the log cleaning based on the conversion name and execution time of the execution log includes:

[0016] Extract the transformation name and execution time of the execution log, and parse the latest execution log of each transformation based on the transformation name and execution time.

[0017] Based on the latest execution log, the conversion name, data link information, database connection, input table and output table in each conversion are extracted, and steps that do not need to be processed are removed.

[0018] Preferably, the step of parsing the latest execution log of each conversion based on the execution time with the conversion name as a dimension includes:

[0019] For any extracted execution log, compare it with the existing execution log. If the execution time of any execution log is less than or equal to the existing execution log, discard it. If the execution time of any execution log is greater than the existing execution log, replace the existing execution log with any execution log.

[0020] Preferably, the method further comprises:

[0021] The blood relationship data is visualized to show the data processing path from the original table to the final target table, including: the original table, the intermediate table, the target table and the connection path; wherein the path is displayed by directed line segments.

[0022] Preferably, the method further comprises:

[0023] According to the node information selected in the blood relationship data, all the child nodes of the node are traversed and displayed in a visual way to perform influence analysis visualization;

[0024] According to the node information selected in the blood relationship data, all parent nodes of the node are traversed and displayed in a visual graphic form to visualize the blood relationship analysis;

[0025] According to the node information selected in the blood relationship data, all child nodes and parent nodes are traversed at the same time and displayed in a visual graphical form to perform full-chain analysis visualization.

[0026] Preferably, the method further comprises:

[0027] The execution logs obtained during the conversion execution, the incremental execution logs in the log analysis database, the data in the intermediate data table, and the blood relationship data are persisted.

[0028] According to another aspect of the present invention, a blood relationship analysis system based on log analysis is provided, characterized in that the system comprises:

[0029] The execution log acquisition module is used to execute Kettle conversion tasks, record conversion logs during data conversion, and obtain execution logs;

[0030] The log aggregation module is used to determine the incremental execution logs according to the execution time of the execution logs, aggregate the incremental execution logs, and save them uniformly in the log analysis database;

[0031] A log cleaning module, used to read the execution log in the log analysis database, and perform log cleaning based on the conversion name and execution time of the execution log, and save it in an intermediate data table;

[0032] The blood relationship data acquisition module is used to analyze the input table and the output table in the intermediate data table based on the data link as the dimension, and remove duplicate data to obtain the blood relationship data.

[0033] Preferably, the system further comprises:

[0034] The data link model building module is used to build a data link model; wherein the data link model includes: node information, database connection information, and connection information and direction of front and rear nodes.

[0035] Preferably, the execution log acquisition module further includes:

[0036] Execute Kettle conversion tasks at preset time intervals, or execute Kettle conversion tasks at specified times according to fixed cycles.

[0037] Preferably, the log cleaning module performs log cleaning based on the conversion name and execution time of the execution log, including:

[0038] Extract the transformation name and execution time of the execution log, and parse the latest execution log of each transformation based on the transformation name and execution time.

[0039] Based on the latest execution log, the conversion name, data link information, database connection, input table and output table in each conversion are extracted, and steps that do not need to be processed are removed.

[0040] Preferably, the log cleaning module parses the latest execution log of each conversion according to the execution time based on the conversion name as a dimension, including:

[0041] For any extracted execution log, compare it with the existing execution log. If the execution time of any execution log is less than or equal to the existing execution log, discard it. If the execution time of any execution log is greater than the existing execution log, replace the existing execution log with any execution log.

[0042] Preferably, the system further comprises:

[0043] A visualization module is used to visualize the blood relationship data and display the data processing path from the original table to the final target table, including: the original table, the intermediate table, the target table and the connection path; wherein the path is displayed through directed line segments.

[0044] Preferably, the system further comprises: a visualization module, for:

[0045] According to the node information selected in the blood relationship data, all the child nodes of the node are traversed and displayed in a visual way to perform influence analysis visualization;

[0046] According to the node information selected in the blood relationship data, all parent nodes of the node are traversed and displayed in a visual graphic form to visualize the blood relationship analysis;

[0047] According to the node information selected in the blood relationship data, all child nodes and parent nodes are traversed at the same time and displayed in a visual graphical form to perform full-chain analysis visualization.

[0048] Preferably, the system further comprises:

[0049] The persistence processing module is used to perform persistence processing on the execution logs obtained during the conversion execution, the incremental execution logs in the log analysis database, the data in the intermediate data table, and the blood relationship data.

[0050] Based on another aspect of the present invention, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any step of a blood relationship analysis method based on log analysis.

[0051] According to another aspect of the present invention, the present invention provides an electronic device, including:

[0052] The computer-readable storage medium described above; and

[0053] One or more processors are used to execute the program in the computer-readable storage medium.

[0054] The present invention provides a blood relationship analysis method, system and storage medium based on log analysis, including: executing Kettle conversion tasks, and recording conversion logs during data conversion to obtain execution logs; determining incremental execution logs according to the execution time of the execution logs, and aggregating the incremental execution logs, and uniformly saving them in a log analysis database; reading the execution logs in the log analysis database, and performing log cleaning based on the conversion name and execution time of the execution logs, and saving them in an intermediate data table; analyzing the input table and output table in the intermediate data table with data chain as the dimension, and removing duplicate data to obtain blood relationship data. The present invention aggregates and analyzes Kettle software logs by constructing a data chain model, extracts input node and output node information on the data processing path, and obtains blood relationship information of a complete data processing process through integration and analysis. The method of the present invention can accurately identify the data blood relationship of data processed by Kettle in the data governance process, thereby providing a convenient method and tool for users to track the data processing chain, and also simplifies the complexity of the blood relationship tracking system, making up for the shortcomings of Kettle software in blood relationship analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:

[0056] Figure 1 is a flow chart of a blood relationship analysis method 100 based on log analysis according to an embodiment of the present invention;

[0057] Figure 2 A schematic diagram of the whole process of blood relationship analysis according to an embodiment of the present invention;

[0058] Figure 3 Schematic diagram of the structure of a blood relationship analysis system 300 based on log analysis according to an embodiment of the present invention. DETAILED DESCRIPTION

[0059] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.

[0060] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.

[0061] Figure 1 FIG. 1 is a flow chart of a blood relationship analysis method 100 based on log analysis according to an embodiment of the present invention. Figure 1 As shown, the lineage analysis method based on log analysis provided by the embodiment of the present invention, by building a data chain model, aggregating and analyzing the Kettle software log, extracting the input node and output node information on the data processing path, and integrating and analyzing to obtain the lineage information of the complete data processing flow. The method of the present invention can accurately identify the data lineage relationship of the data processed by Kettle in the data governance process, thereby providing a convenient method and tool for users to track the data processing chain, while also simplifying the complexity of the lineage tracking system and making up for the shortcomings of the Kettle software in lineage analysis. The lineage analysis method 100 based on log analysis provided by the embodiment of the present invention starts from step 101. In step 101, the Kettle conversion task is executed, and the conversion log is recorded during the data conversion process to obtain the execution log.

[0062] Preferably, the method further comprises:

[0063] Execute Kettle conversion tasks at preset time intervals, or execute Kettle conversion tasks at specified times according to fixed cycles.

[0064] Preferably, the method further comprises:

[0065] Construct a data link model; wherein the data link model includes: node information, database connection information, and connection information and direction of front and rear nodes.

[0066] In the present invention, a data link model is constructed. The model contains multiple aspects of information involved in data conversion, including node information, database connection information, connection information and direction of previous and next nodes, etc. The data link can be one-to-many, many-to-one, and many-to-many.

[0067] Combination Figure 2As shown, in the present invention, based on the task scheduling module, log aggregation tasks, log cleaning tasks and blood relationship analysis tasks are executed regularly. The task scheduling module can be used to execute Kettle conversion tasks. The module can execute tasks at fixed time intervals, or execute tasks at specified time points in a fixed cycle. The conversion log recording function is turned on in data conversion, and the conversion log, step log and log channel are recorded separately. The logs generated during the conversion execution are persisted.

[0068] In step 102, incremental execution logs are determined according to the execution time of the execution logs, and the incremental execution logs are aggregated and uniformly saved in a log analysis database.

[0069] Combination Figure 2 As shown, in the present invention, the execution logs of each executor are aggregated based on the log aggregation module, and the logs are centrally saved in a unified database. Specifically, the log aggregation module is connected to each log database, and the incremental execution logs are extracted according to the execution time of the execution logs. After the incremental logs are aggregated, they are uniformly persisted.

[0070] In step 103, the execution log in the log analysis database is read, and the log is cleaned based on the conversion name and execution time of the execution log, and saved in the intermediate data table.

[0071] Preferably, the log cleaning based on the conversion name and execution time of the execution log includes:

[0072] Extract the transformation name and execution time of the execution log, and parse the latest execution log of each transformation based on the transformation name and execution time.

[0073] Based on the latest execution log, the conversion name, data link information, database connection, input table and output table in each conversion are extracted, and steps that do not need to be processed are removed.

[0074] Preferably, the step of parsing the latest execution log of each conversion based on the execution time with the conversion name as a dimension includes:

[0075] For any extracted execution log, compare it with the existing execution log. If the execution time of any execution log is less than or equal to the existing execution log, discard it. If the execution time of any execution log is greater than the existing execution log, replace the existing execution log with any execution log.

[0076] Combination Figure 2As shown, in the present invention, execution log cleaning is performed based on the log cleaning module. Specifically, the conversion name and execution time are extracted, and the latest execution log of each conversion is parsed with the conversion name as the dimension, and compared with the existing execution log. If the execution time of the extracted execution log is less than or equal to the existing execution log, it is discarded. If the execution time of the extracted execution log is greater than the existing execution log, the read execution log replaces the existing execution log and performs persistence processing. Then, with the conversion name as the dimension, data extraction and cleaning are performed on the log records, and the conversion name, data link, database connection, input table, and output table in each conversion are extracted. Other steps that do not need to be analyzed and processed are discarded, and the results of log cleaning are processed as data persistence.

[0077] In step 104, the input table and the output table in the intermediate data table are analyzed with the data chain as the dimension, and duplicate data are removed to obtain blood relationship data.

[0078] Preferably, the method further comprises:

[0079] The blood relationship data is visualized to show the data processing path from the original table to the final target table, including: the original table, the intermediate table, the target table and the connection path; wherein the path is displayed by directed line segments.

[0080] Preferably, the method further comprises:

[0081] According to the node information selected in the blood relationship data, all the child nodes of the node are traversed and displayed in a visual way to perform influence analysis visualization;

[0082] According to the node information selected in the blood relationship data, all parent nodes of the node are traversed and displayed in a visual graphic form to visualize the blood relationship analysis;

[0083] According to the node information selected in the blood relationship data, all child nodes and parent nodes are traversed at the same time and displayed in a visual graphical form to perform full-chain analysis visualization.

[0084] Preferably, the method further comprises:

[0085] The execution logs obtained during the conversion execution, the incremental execution logs in the log analysis database, the data in the intermediate data table, and the blood relationship data are persisted.

[0086] Combination Figure 2 As shown, in the present invention, based on the blood relationship analysis module, the output of the log cleaning module is analyzed, the input table and output table on the same data chain are extracted and the duplication is removed. Finally, the blood relationship data for data analysis is obtained, such as: influence analysis, blood relationship analysis and full chain analysis.

[0087] The lineage analysis visualization reads the lineage relationship from the persistence results of the above steps and displays it in a visual graphic. The lineage analysis visualization displays the data processing path from the original table to the final target table, including the original table, intermediate table, target table and connection path, and the path is displayed through directed line segments. Lineage relationship analysis includes: influence analysis, lineage analysis and full chain analysis. The influence analysis visualization traverses all child nodes of the node according to the selected node information and displays them in a visual way. The lineage analysis visualization traverses all parent nodes of the node according to the selected node information and displays them in a visual graphic. The full chain analysis visualization traverses all child nodes and parent nodes at the same time according to the selected node information and displays them in a visual graphic.

[0088] Specifically, by selecting the target table for lineage analysis, you can visualize the data processing path from the original table to the final target table, including the original table, intermediate table, target table, and the paths in between, and all paths are connected by directed lines. By selecting the original table for impact analysis, you can display all data processing paths for reading data from the table, including the original table, intermediate table, target table, target table, and the paths in between, and all paths are connected by directed lines. By selecting a table for full-chain analysis, you can simultaneously display the results of lineage analysis and impact analysis.

[0089] The data lineage analysis method based on Kettle log analysis of the present invention specifically comprises the following steps:

[0090] (1) Design Kettle transformations and build a data chain model, including input tags, output tags, database connection information, and data chain information. Data chains can be one-to-many, many-to-one, and many-to-many. Enable the transformation log recording function to record transformation logs, step logs, and log channels respectively. The logs generated during transformation execution are persisted.

[0091] (2) Connect to the log database server, extract the execution log, extract the incremental execution log according to the execution time of the execution log, and aggregate these logs into a unified log analysis database.

[0092] (3) Read the execution log records from the log analysis database, parse the latest execution log of each transformation with the transformation name as the dimension, and save it in a new data table. Compare the existing execution logs. If the execution log does not exist, insert a new execution log; if the execution time of the extracted execution log is less than or equal to the existing execution log, discard this execution log; if the execution time of the extracted execution log is greater than the existing execution log, delete the existing execution log and then insert the new execution log. Continue to use the transformation name as the dimension to extract the transformation name, data link information, database connection, input table, and output table in each transformation, discard other steps that do not need to be analyzed and processed, and save the results of log cleaning in a new data table.

[0093] (4) Bloodline analysis uses the data chain as a dimension, extracts the input table and output table on the data chain and removes duplications, obtains bloodline relationship data for data analysis, and saves it in the bloodline data table.

[0094] (5) The visualization display of lineage analysis includes influence analysis display, lineage analysis display and full-chain analysis display. The lineage relationship is read from the persistence results of the above steps to display the data processing path from the original table to the final target table, including the original table, intermediate table, target table and intermediate path. The direct path between the tables is represented by directed line segments. The influence lineage analysis display traverses all child nodes of the node according to the selected node information and displays them in a visual graphic format. The lineage analysis display traverses all parent nodes of the node according to the selected node information and displays them in a visual graphic format. The full-chain analysis traverses all child nodes and parent nodes at the same time according to the selected node information and displays them in a visual graphic format.

[0095] The present invention extracts and parses data lineage information and displays it in a visual way by analyzing Kettle's conversion execution log. The present invention can accurately identify the data lineage relationship of data processed by Kettle in the data governance process, thereby providing users with convenient methods and tools to track data processing paths, while also simplifying the complexity of the lineage tracking system and reducing the requirements for users and system maintenance personnel. Visual display of data lineage information is easier for users to accept and understand than presentation in text or table form. Users can intuitively see the flow and changes of data, solving the problems of complex data tracing and difficult location and troubleshooting of problematic data.

[0096] Figure 3 FIG. 3 is a schematic diagram of the structure of a blood relationship analysis system 300 based on log analysis according to an embodiment of the present invention. Figure 3As shown, the blood relationship analysis system 300 based on log analysis provided by the embodiment of the present invention includes: an execution log acquisition module 301, a log aggregation module 302, a log cleaning module 303 and a blood relationship data acquisition module 304.

[0097] Preferably, the execution log acquisition module 301 is used to execute the Kettle conversion task, and to record the conversion log during the data conversion process to obtain the execution log.

[0098] Preferably, the execution log acquisition module 301 further includes:

[0099] Execute Kettle conversion tasks at preset time intervals, or execute Kettle conversion tasks at specified times according to fixed cycles.

[0100] Preferably, the log aggregation module 302 is used to determine the incremental execution logs according to the execution time of the execution logs, aggregate the incremental execution logs, and save them uniformly in the log analysis database.

[0101] Preferably, the log cleaning module 303 is used to read the execution log in the log analysis database, and perform log cleaning based on the conversion name and execution time of the execution log, and save it in the intermediate data table.

[0102] Preferably, the log cleaning module 303 performs log cleaning based on the conversion name and execution time of the execution log, including:

[0103] Extract the transformation name and execution time of the execution log, and parse the latest execution log of each transformation based on the transformation name and execution time.

[0104] Based on the latest execution log, the conversion name, data link information, database connection, input table and output table in each conversion are extracted, and steps that do not need to be processed are removed.

[0105] Preferably, the log cleaning module 303 parses the latest execution log of each conversion according to the execution time based on the conversion name, including:

[0106] For any extracted execution log, compare it with the existing execution log. If the execution time of any execution log is less than or equal to the existing execution log, discard it. If the execution time of any execution log is greater than the existing execution log, replace the existing execution log with any execution log.

[0107] Preferably, the blood relationship data acquisition module 304 is used to analyze the input table and the output table in the intermediate data table based on the data link as the dimension, and remove duplicate data to obtain the blood relationship data.

[0108] Preferably, the system further comprises:

[0109] The data link model building module is used to build a data link model; wherein the data link model includes: node information, database connection information, and connection information and direction of front and rear nodes.

[0110] Preferably, the system further comprises:

[0111] A visualization module is used to visualize the blood relationship data and display the data processing path from the original table to the final target table, including: the original table, the intermediate table, the target table and the connection path; wherein the path is displayed through directed line segments.

[0112] Preferably, the system further comprises: a visualization module, for:

[0113] According to the node information selected in the blood relationship data, all the child nodes of the node are traversed and displayed in a visual way to perform influence analysis visualization;

[0114] According to the node information selected in the blood relationship data, all parent nodes of the node are traversed and displayed in a visual graphic form to visualize the blood relationship analysis;

[0115] According to the node information selected in the blood relationship data, all child nodes and parent nodes are traversed at the same time and displayed in a visual graphical form to perform full-chain analysis visualization.

[0116] Preferably, the system further comprises:

[0117] The persistence processing module is used to perform persistence processing on the execution logs obtained during the conversion execution, the incremental execution logs in the log analysis database, the data in the intermediate data table, and the blood relationship data.

[0118] The blood relationship analysis system 300 based on log analysis in the embodiment of the present invention corresponds to the blood relationship analysis method 100 based on log analysis in another embodiment of the present invention, and will not be described in detail here.

[0119] Based on another aspect of the present invention, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any step of a blood relationship analysis method based on log analysis.

[0120] According to another aspect of the present invention, the present invention provides an electronic device, including:

[0121] The computer-readable storage medium described above; and

[0122] One or more processors are used to execute the program in the computer-readable storage medium.

[0123] The invention has been described above with reference to a few embodiments. However, it is readily apparent to a person skilled in the art that other embodiments than the ones disclosed above are equally within the scope of the invention, as defined by the appended patent claims.

[0124] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise therein. All references to "a / said / the [means, components, etc.]" are to be openly interpreted as at least one instance of said means, components, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not necessarily have to be performed in the exact order disclosed, unless explicitly stated otherwise.

[0125] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0127] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.

[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A blood relationship analysis method based on log analysis, characterized in that: The method comprises: Execute Kettle conversion tasks, record conversion logs during data conversion, and obtain execution logs; Determine the incremental execution logs according to the execution time of the execution logs, aggregate the incremental execution logs, and save them uniformly in the log analysis database; Read the execution log in the log analysis database, clean the log based on the conversion name and execution time of the execution log, and save it in the intermediate data table; The input table and the output table in the intermediate data table are analyzed with the data chain as the dimension, and duplicate data are removed to obtain blood relationship data.

2. The method according to claim 1, characterized in that The method further comprises: Construct a data link model; wherein the data link model includes: node information, database connection information, and connection information and direction of front and rear nodes.

3. The method according to claim 1, characterized in that The method further comprises: Execute Kettle conversion tasks at preset time intervals, or execute Kettle conversion tasks at specified times according to fixed cycles.

4. The method according to claim 1, characterized in that: The log cleaning based on the conversion name and execution time of the execution log includes: Extract the transformation name and execution time of the execution log, and parse the latest execution log of each transformation based on the transformation name and execution time; Based on the latest execution log, the conversion name, data link information, database connection, input table and output table in each conversion are extracted, and steps that do not need to be processed are removed.

5. The method according to claim 1, characterized in that The latest execution log of each transformation is parsed based on the transformation name and execution time, including: For any extracted execution log, compare it with the existing execution log. If the execution time of any execution log is less than or equal to the existing execution log, discard it. If the execution time of any execution log is greater than the existing execution log, replace the existing execution log with any execution log.

6. The method according to claim 1, characterized in that The method further comprises: The blood relationship data is visualized to show the data processing path from the original table to the final target table, including: the original table, the intermediate table, the target table and the connection path; wherein the path is displayed by directed line segments.

7. The method according to claim 1, characterized in that The method further comprises: According to the node information selected in the blood relationship data, all the child nodes of the node are traversed and displayed in a visual way to perform influence analysis visualization; According to the node information selected in the blood relationship data, all parent nodes of the node are traversed and displayed in a visual graphic form to visualize the blood relationship analysis; According to the node information selected in the blood relationship data, all child nodes and parent nodes are traversed at the same time and displayed in a visual graphical form to perform full-chain analysis visualization.

8. The method according to claim 1, characterized in that The method further comprises: The execution logs obtained during the conversion execution, the incremental execution logs in the log analysis database, the data in the intermediate data table, and the blood relationship data are persisted.

9. A blood relationship analysis system based on log analysis, characterized in that: The system comprises: The execution log acquisition module is used to execute Kettle conversion tasks, record conversion logs during data conversion, and obtain execution logs; The log aggregation module is used to determine the incremental execution logs according to the execution time of the execution logs, aggregate the incremental execution logs, and save them uniformly in the log analysis database; A log cleaning module, used to read the execution log in the log analysis database, and perform log cleaning based on the conversion name and execution time of the execution log, and save it in an intermediate data table; The blood relationship data acquisition module is used to analyze the input table and the output table in the intermediate data table based on the data link as the dimension, and remove duplicate data to obtain the blood relationship data.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Data blood relationship extraction method and device and electronic equipment

    CN113326261A

  • Data blood relationship analysis display method and device based on ETL (Extract-Transform-Load) and electronic equipment

    CN115145988A

  • Blood relationship data extraction method based on data warehouse log file

    CN116662308A

  • Data blood relationship analysis method and system based on kettle actuator

    CN117931840A

  • Data blood source relationship analysis system and method

    CN118170785A