Method for establishing reusable model based on ETL multi-source heterogeneous data acquisition

By building a data acquisition reusable model for ETL scheduling, executable programs and feedback evaluation modules in the steel industry, the problem of data resource storage is solved, the automatic collection of multi-source heterogeneous data and the establishment of enterprise-level big data information database are realized, and data utilization and decision-making support capabilities are improved.

CN120011439APending Publication Date: 2025-05-16TANGSHAN HUITANG WULIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510156889.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The massive data resources accumulated in the steel industry are stored and cannot be fully mined and efficiently utilized, resulting in data island phenomenon, data utilization rate declines, and lack of in-depth mining and decision-making support.

Method used

By building an ETL scheduling module, an ETL executable program module and an ETL feedback evaluation module, a data acquisition reusable model is established, automatic collection of multi-source heterogeneous data is realized, and an enterprise-level big data information database is established to provide data support for comprehensive lean production management and control.

Benefits of technology

It realizes data synchronization between different types of data sources and target systems, completes the integration of enterprise-level big data, improves data utilization, provides in-depth data analysis and decision-making support, and ensures data stability and high availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011439A_ABST
    Figure CN120011439A_ABST
Patent Text Reader

Abstract

The invention provides a method for establishing a reusable model based on ETL multi-source heterogeneous data acquisition. The method comprises an ETL scheduling module, an ETL executable program module and an ETL feedback evaluation module. The ETL scheduling module is used for controlling starting and running of ETL tasks, including controlling starting time, running cycle, triggering conditions and the like of the ETL tasks, and ensuring high availability of scheduling tasks by setting timed tasks, controlling scheduling execution scripts and triggering running of ETL executable programs and by establishing a data redundancy mechanism and scheduling control task triggering conditions; the ETL executable module establishes an ETL model by researching an ETL executable program mechanism, and a user only needs to establish an ETL execution task by configuring parameters such as a task i d, so that data acquisition from a source table to a target table is realized. According to the method, automatic collection of multi-source heterogeneous data of iron and steel enterprises is realized by building a data collection reusable model, and an enterprise-level big data information base is established. And high availability and data accuracy of data acquisition are both considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data collection and information technology in the metallurgical industry, and in particular to a method for establishing a reusable model for ETL-based multi-source heterogeneous data collection. Background Art

[0002] In the process of production and operation, the steel industry has accumulated a large amount of resource data of multiple subjects, multiple businesses, and multiple levels between various systems, which contains huge mining value. However, it also has the pain points of poor correlation and difficulty in centralized control. Various application systems with their own data storage and access methods will eventually lead to the phenomenon of "data islands", and the data utilization rate will be greatly reduced. The data reflecting the company's operation status, such as process, quality, and production, lacks in-depth mining and decision-making support, which is not conducive to production line production and quality improvement.

[0003] With the development of big data related technologies, ETL technology, with its advantages of high integration efficiency and simplified interface development, provides enterprises with a way to integrate multi-source heterogeneous data and store them in enterprise big data information bases. However, commercial software is expensive and has a low penetration rate, and open source and free tools have problems with system stability.

[0004] In order to realize the integration of enterprise-level big data resources, deeply explore the potential value of data, and improve the cohesion of data decisions, the following problems need to be solved: explore a set of data fusion solutions, realize the collection of multi-source heterogeneous data of the enterprise, build an enterprise-level information database, deeply analyze the production rules of production lines, provide data support for enterprise decision-making, and at the same time ensure the stability, high availability and data accuracy of data fusion technology solutions. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a method for establishing a reusable model for multi-source heterogeneous data acquisition based on ETL, which solves the problem that the massive data resources accumulated by steel enterprises are stored in a scattered manner and cannot be fully mined and efficiently utilized. The method explores the fusion of multi-source heterogeneous data and builds an enterprise-level big data information database to establish a method for establishing a reusable model for data acquisition based on ETL.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for establishing a reusable model for multi-source heterogeneous data collection based on ETL, establishing a reusable data collection model by building an ETL scheduling module, an ETL executable program module and an ETL feedback evaluation module, realizing the automatic collection of multi-source heterogeneous data of steel enterprises including production, process, quality, logistics, etc., establishing an enterprise-level big data information database, and providing data support for the enterprise's comprehensive lean production management and control.

[0007] Preferably, the method specifically comprises the following steps:

[0008] Step S1: Build an ETL scheduling module, set scheduled tasks according to different business scenarios and priorities, flexibly schedule ETL execution scripts, and trigger the operation of ETL executable program modules.

[0009] Step S2: Establish a multi-source heterogeneous data collection model based on ETL, divide it into structured data and real-time data according to the data type, classify the data and integrate it, and realize data collection from the source system to the target system.

[0010] Step S3: Establish an ETL evaluation feedback model, develop an interface with the ETL execution log, obtain the information identifying errors in the execution log, and locate the ETL task with data anomalies, the cause of the anomaly, and the time when the anomaly occurred.

[0011] Preferably, in step S1, it is characterized in that a data redundancy mechanism is established for the ETL scheduling module, and the master-slave scheduling task realizes load balancing of the master-slave machines by setting a startup time difference and scheduling execution status update.

[0012] Preferably, the specific implementation steps of the ETL scheduling module are as follows: (1) Establish an ETL scheduling task table in the Oracle database to configure the scheduling task information such as the scheduling task name, scheduling execution time, execution cycle, execution status, and execution path; (2) Use the task status 0 / 1 value in the ETL scheduling task configuration table to identify whether the task is being executed. (3) Use the scheduling execution time in the ETL scheduling task configuration table to record the execution time of the master and standby machines respectively. (4) Set the logic loop according to the scheduled task to execute the process in steps (2) and (3).

[0013] Preferably, in step S2, the ETL multi-source heterogeneous data acquisition model established is characterized in that corresponding ETL executable programs are established according to different data types of structured data and real-time data. The structured data acquisition model is divided into three branches: full data extraction, incremental data extraction, and associated table incremental data extraction. The real-time data model is divided into two branches: WEBAPI and OPC.

[0014] Preferably, the ETL multi-source heterogeneous data collection model is established to abstract parameters such as number, business line, ETL task id, subtask id, source library type, data source table, target table, write mode, task description, time parameters, etc. from different types of data extraction tasks, write the configured parameters into the Oracle database, and automatically implement the entire ETL processing process. The same type of ETL tasks can be used for data collection through the same model, which is convenient for later development.

[0015] Preferably, in step S3, it is characterized by writing a Java program, establishing an interface with the ETL execution log, obtaining error information in the log, and feeding back the abnormal information to the front-end interface for display, and implementing color identification.

[0016] The present invention provides a method for establishing a reusable model for ETL multi-source heterogeneous data collection.

[0017] It has the following beneficial effects:

[0018] 1. The present invention provides a method for establishing a reusable model for multi-source heterogeneous data collection based on ETL. Through the study of ETL mechanism, a multi-source heterogeneous data collection model is built to achieve data synchronization from different types of data sources to the target system, thus completing the integration of enterprise-level big data.

[0019] 2. The present invention provides a method for establishing a reusable model for multi-source heterogeneous data collection based on ETL, which realizes the identification, feedback and evaluation of abnormal data information by establishing an interface with the execution log and matching key information.

[0020] 3. The present invention provides a method for establishing a reusable model based on ETL multi-source heterogeneous data collection. It comprehensively considers the stability of the model during use and establishes a data redundancy mechanism so that data pressure can be dispersed to different server nodes, thereby improving data synchronization efficiency and the security and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic diagram of the system architecture of the present invention;

[0022] Figure 2 It is a structural diagram of a multi-source heterogeneous data acquisition model;

[0023] Figure 3 It is the ETL data collection model flow chart. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] The embodiment of the present invention provides a method for establishing a reusable model for multi-source heterogeneous data collection based on ETL, Figure 1The figure shows a schematic diagram of the system architecture of the present invention, a method of a reusable model for data collection based on ETL includes three parts: an ETL scheduling module, an ETL executable program module and an ETL feedback evaluation module, which realizes the automatic collection of multi-source heterogeneous data including production, process, quality, logistics, etc. of steel enterprises, establishes an enterprise-level big data information database, and provides data support for the comprehensive management and control of lean production of enterprises.

[0026] The technical solution of the present invention specifically comprises the following steps:

[0027] Step S1: Build an ETL scheduling module, set scheduled tasks according to different business scenarios and priorities, flexibly schedule ETL execution scripts, and trigger the operation of ETL executable program modules.

[0028] Step S2: Establish a multi-source heterogeneous data collection model based on ETL, divide it into structured data and real-time data according to the data type, classify the data and integrate it, and realize data collection from the source system to the target system.

[0029] Step S3: Develop an interface with the ETL execution log, obtain the error information in the execution log, and locate the ETL task with data anomaly, the cause of the anomaly, and the time when the anomaly occurred.

[0030] In step S1, a data redundancy mechanism is established for the ETL scheduling module, and the master and standby machine scheduling tasks achieve load balancing of the master and standby machines by setting the startup time difference and scheduling execution status update.

[0031] The specific implementation steps of the ETL scheduling module are as follows: (1) Establish an ETL scheduling task table in the Oracle database to configure the scheduling task information such as the scheduling task name, scheduling execution time, execution cycle, execution status, and execution path; (2) Use the task status 0 / 1 value in the ETL scheduling task configuration table to identify whether the task is being executed. (3) Use the scheduling execution time in the ETL scheduling task configuration table to record the execution time of the primary and standby machines to preempt the task execution. (4) Set up a logical loop according to the scheduled task to execute the process in steps (2) and (3).

[0032] Taking a group of tasks for scheduling MES data in the iron and steel industry as an example, the ETL scheduling task TQ_15m, the scheduled task is 15 minutes, and the execution script path are set in the ETL scheduling task table. The initial scheduling task status is set to 0, and an initial host scheduling execution time and standby scheduling execution time are given. When the program controls the timing task to start, the primary and standby servers start the scheduling task at the same time. The scheduling task that successfully preempts sets the task status STATUS field in the configuration table to 1. Taking the successful host preemption as an example, the primary server triggers the execution of the ETL executable program and starts data synchronization from the source table to the target table. After the execution is completed, the execution time is written into the host scheduling execution time LAST_RUNT I ME_PRI MARY field. The standby scheduling task that enters later recognizes that the task status STATUS in the configuration table is 1, so it is set to 0 and the task exits.

[0033] Multiple scheduling tasks can be set for the same source system, and the same scheduling task can trigger the execution of multiple ETL executable programs.

[0034] In step S2, the ETL multi-source heterogeneous data acquisition model is established, and the corresponding ETL executable programs are established according to the different data types of structured data and real-time data. The structured data acquisition model is divided into three branches: full data extraction, incremental data extraction, and incremental data extraction of associated tables. The real-time data model is divided into two branches: WEBAP I and OPC.

[0035] The incremental extraction model of associated tables is an ETL model derived from the incremental extraction model to solve the situation where there is a certain relationship between the main table and the sub-table, but the sub-table has no primary key.

[0036] Taking the data collection model with converter-related information of the production and sales system as the data source as an example, the PM table is the main table furnace event table, and the pmprp furnace report, pmess furnace charging information, pmtmp temperature measurement information, etc. are sub-tables associated with the PM table. The specific steps of its ETL executable program are: (1) synchronize the main table PM, (2) synchronize the sub-tables pmprp, pmess...pmtmp, (3) update the main table PM, (4) update the sub-tables pmprp, pmess...pmtmp, and (5) update the timestamp.

[0037] Taking the WEBAPI data collection model with real-time data as the data source as an example, the specific implementation steps are: establish a real-time data parameter table, configure the tag name of the real-time data to be collected and the fields such as collection time and data time, develop an interface program to send a request to the source system, obtain real-time data feedback, and write it to the target table.

[0038] The established ETL multi-source heterogeneous data collection model abstracts the number, business line, ETL task id, subtask id, source library type, data source table, target table, write mode, task description, sql parameters and other parameters from different types of data extraction tasks, writes the configured parameters into the Oracle database, and automatically implements the entire ETL processing process. The same type of ETL tasks can collect data through the same model, which is convenient for later development.

[0039] In step S3, a Java program is written to establish an interface with the ETL execution log, obtain error information in the log, match the error information library, and feed back the exception information to the front-end interface for display and implement color identification.

[0040] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for establishing a reusable model for ETL multi-source heterogeneous data collection, characterized by: By building the ETL scheduling module, ETL executable program module and ETL feedback evaluation module, a reusable data collection model was established, which enabled the automatic collection of multi-source heterogeneous data of steel enterprises including production, process, quality, logistics, etc., and established an enterprise-level big data information database, providing data support for the company's comprehensive management and control of lean production.

2. The method for establishing a reusable model for ETL multi-source heterogeneous data collection according to claim 1 is characterized in that: The specific steps include: Step S1: Build an ETL scheduling module, set scheduled tasks according to different business scenarios and priorities, flexibly schedule ETL execution scripts, and trigger the operation of ETL executable program modules. Step S2: Establish a multi-source heterogeneous data collection model based on ETL, divide it into structured data and real-time data according to the data type, classify the data and integrate it, and realize data collection from the source system to the target system. Step S3: Establish an ETL evaluation feedback model, develop an interface with the ETL execution log, obtain the information identifying errors in the execution log, and locate the ETL task with data anomalies, the cause of the anomaly, and the time when the anomaly occurred.

3. The method for establishing a reusable model for ETL multi-source heterogeneous data collection according to claim 2 is characterized in that: In the step S1, it is characterized in that a data redundancy mechanism is established for the ETL scheduling module, and the master and standby machine scheduling tasks realize load balancing of the master and standby machines by setting the startup time difference and scheduling execution status update.

4. The method for establishing a reusable model for ETL multi-source heterogeneous data collection according to claim 3 is characterized in that: The specific implementation steps of the ETL scheduling module are as follows: (1) Establish an ETL scheduling task table in the Oracle database to configure the scheduling task information such as the scheduling task name, scheduling execution time, execution cycle, execution status, and execution path; (2) Use the task status 0 / 1 value in the ETL scheduling task configuration table to identify whether the task is being executed. (3) Use the scheduling execution time in the ETL scheduling task configuration table to record the execution time of the primary and standby machines respectively. (4) Set up a logical loop according to the scheduled task to execute the process in steps (2) and (3).

5. The method for establishing a reusable model for ETL multi-source heterogeneous data collection according to claim 2, characterized in that: In step S2, the ETL multi-source heterogeneous data acquisition model established is characterized in that corresponding ETL executable programs are established according to different data types of structured data and real-time data. The structured data acquisition model is divided into three branches: full data extraction, incremental data extraction, and associated table incremental data extraction. The real-time data model is divided into two branches: WEB API and OPC.

6. The method for establishing a reusable model for ETL multi-source heterogeneous data collection according to claim 5, characterized in that: The established ETL multi-source heterogeneous data collection model abstracts parameters such as number, business line, ETL task id, subtask id, source library type, data source table, target table, write mode, task description, time parameters, etc. from different types of data extraction tasks, writes the configured parameters into the Oracle database, and automatically implements the entire ETL processing process. The same type of ETL tasks can collect data through the same model, which is convenient for later development.

7. The method for establishing a reusable model for ETL multi-source heterogeneous data collection according to claim 2, characterized in that: The step S3 is characterized by writing a Java program, establishing an interface with the ETL execution log, obtaining error information in the log, and feeding back the abnormal information to the front-end interface for display, and implementing color identification.