Intelligent data integration method, system and equipment for realizing ETL data processing flow optimization and medium
By automatically analyzing metadata, identifying key information, filling in missing data and automatically generating ETL tasks, the problem of ETL task execution relies on manual scripting, and the optimization of ETL data processing process is achieved, and the efficiency and flexibility of data integration are improved.
Patent Information
- Application Number
- CN202411863273.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-23
AI Technical Summary
ETL task execution relies on manual scripting or complex configuration, which leads to high development costs and high maintenance difficulties. In addition, traditional manual ETL methods are difficult to quickly adapt to changes in data sources and target systems, resulting in limited flexibility and response speed of data integration.
By automatically analyzing metadata, identifying key information, filling in missing data, and automatically generating ETL tasks, reducing manual intervention. Use a large model analysis system for in-depth analysis, identify the correlation between data, and flexibly configure and schedule according to business needs and data characteristics.
The optimization of ETL data processing process is achieved, the efficiency and flexibility of data integration is improved, the development cost and maintenance difficulty is reduced, the resource utilization is optimized, and the timeliness of data processing is improved.
Smart Images

Figure CN120030071A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing technology, and specifically relates to an intelligent data integration method, system, equipment and medium for realizing ETL data processing flow optimization. Background Art
[0002] ETL, short for Extract-Transform-Load, refers to the process of extracting, transforming, and loading data from the source to the destination.
[0003] With the development of big data technology, enterprises have an increasing demand for efficient data processing and intelligent application. Although data processing technology, especially ETL technology, plays an important role in data integration, its efficiency and accuracy are crucial to improving data quality. However, there are still some problems in ETL tasks, which limit the efficiency and flexibility of data integration. First, the execution of ETL tasks is highly dependent on manually written scripts or complex configurations. This method not only increases the development cost and maintenance difficulty, but also easily leads to ETL process failure or data quality degradation due to human errors. Secondly, with the continuous changes in data sources and target systems, traditional manual ETL methods are difficult to adapt quickly, resulting in limited flexibility and response speed of data integration. In addition, during the data loading process, a reasonable scheduling strategy is required to optimize resource utilization and improve the timeliness of data processing. However, the existing data loading process often adopts a fixed scheduling strategy, which cannot be flexibly adjusted according to the actual data request and business needs. This scheduling method may not only cause key data to fail to load in time, but also cause resource waste due to unreasonable resource allocation. Summary of the invention
[0004] In a first aspect, the present application provides an intelligent data integration method for optimizing the ETL data processing process, including the following steps: S1. Collect metadata information from the source library, analyze the metadata information, and identify key information; S2. Use the large model analysis system to analyze metadata information and key information, identify missing data in the source library, and then fill in the missing data based on the analysis results; S3. Perform basic configuration of ETL tasks and generate full ETL tasks and incremental ETL tasks; S4. Define the scheduling strategy for data processing and start the execution of ETL tasks according to the scheduling strategy for data processing.
[0005] Furthermore, the specific steps of step S1 are as follows: S11. Start data integration, configure source library parameters, and connect to the source library according to the source library connection information; S12. Automatically scan the tables and fields in the source library and collect metadata information of the source library; the metadata information includes object name, structure, field type and data constraints; S13. Use intelligent recognition algorithms to conduct preliminary analysis on the collected metadata information and identify key information that characterizes the features.
[0006] Furthermore, the specific steps of step S13 are as follows: S131. After cleaning and normalizing the collected metadata information, key fields are extracted, including table name, field name, field type, primary key constraint, and foreign key constraint; S132. Extract features from key fields of metadata information, and extract rules representing primary keys, foreign keys, and indexes as key features; S133. Use a machine learning algorithm to build a classification model, and use the annotated historical metadata information to train to obtain a trained classification model; S134. After data processing and feature extraction of the newly collected metadata information, the metadata information is input into the trained classification model for classification to identify the primary key, foreign key and index.
[0007] Furthermore, the specific steps of step S2 are as follows: S21. Input the collected metadata information and identified key information into the existing large model analysis system for in-depth analysis, and identify the associations between data and the missing data in the source library; S22. Based on the recognition results of the large model analysis system, the missing data in the source library is inferred based on the relationship between the data and combined with historical data or business rules, and then filled with the inferred data.
[0008] Furthermore, the specific steps of step S3 are as follows: S31. Obtain the source database data after missing data filling and load it into the target table; S32. Determine the source database data, target table, target table incremental primary key, target table entry time, source primary table and organization name of the ETL task, and obtain the basic configuration of the ETL task; S33. Determine the data conversion rules and algorithms for full data conversion according to the standard data structure of the target table, and generate the ELT full script; S34. Obtain the user-defined incremental configuration or automatically recommend the incremental configuration based on the source master table; S35. Generate an ETL incremental script based on the ETL full script and incremental configuration; S36. Generate ETL full tasks and ETL incremental tasks based on the basic configuration, incremental configuration, ETL full script and ELT incremental script of the ETL task.
[0009] Furthermore, the specific steps of step S36 are as follows: S361. Review the basic configuration of ETL tasks; If the review is passed, proceed to step S362; If the review fails, return to step S32; S362. Review the full ELT script; If the review is passed, proceed to step S363; If the review fails, return to step S33; S363. Determine whether the ETL incremental script needs to be reviewed; If yes, go to step S364; If not, proceed to step S365; S364. Review ETL incremental scripts; If the review is passed, proceed to step S365; If the review fails, return to step S34; S365. Determine whether the ETL full task is generated; If yes, go to step S4; If not, proceed to step S366; S366. Generate ETL full task according to ETL full script; S367. Determine whether it is necessary to generate an ETL incremental task; If yes, go to step S368; If not, proceed to step S4; S368. Determine whether the ETL incremental script has been confirmed; If yes, go to step S369; If not, return to step S364; S369. Generate ETL incremental tasks according to the ETL incremental script.
[0010] Furthermore, the specific steps of step S4 are as follows: S41. Define the scheduling strategy for source database data based on business needs, importance and urgency; S42. Sort and schedule each ETL task according to the scheduling strategy, and determine the execution time and execution order of each ETL task; S43. Start ETL task execution; S44. Identify whether the current time point satisfies the execution of a certain ETL task; If yes, take the ETL task as the current task and go to step S45; If not, wait for the set time period and return to step S44; S45. Determine the current task type; If it is a full ETL task, go to step S48; If it is an ETL incremental task, go to step S46; S46. Determine whether the ETL full task has been successfully executed; If yes, go to step S47; If not, return to step S44; S47. Execute the ETL incremental script of the current ETL incremental task, and return to step S44; S48. Execute the ETL full script of the current ETL full task and return to step S44.
[0011] In a second aspect, the embodiment of the present application further provides an intelligent data integration system for optimizing the ETL data processing process, including: The metadata extraction and identification module is used to collect metadata information from the source library, analyze the metadata information, and identify key information; The metadata analysis and data filling module is used to analyze metadata information and key information using the large model analysis system, identify missing data in the source library, and then fill in the missing data based on the analysis results; The ETL task configuration module is used to perform basic configuration of ETL tasks, full ETL task configuration, and incremental ETL task configuration, and generate ETL tasks based on the configuration information; The ETL task scheduling execution module is used to define the scheduling strategy for data processing and start the execution of ETL tasks according to the scheduling strategy.
[0012] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the intelligent data integration method for optimizing the ETL data processing flow as described in the first aspect are implemented.
[0013] In a fourth aspect, an embodiment of the present application further provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the intelligent data integration method for optimizing the ETL data processing flow as described in the first aspect.
[0014] It can be seen from the above technical solutions that the present invention has the following advantages: The intelligent data integration method, system, device and medium for realizing ETL data processing flow optimization provided by the present application can automatically analyze metadata, identify key information, fill in missing data, and automatically generate and schedule ETL tasks, thereby reducing manual intervention and error rate. At the same time, it can also be flexibly configured and scheduled according to business needs and data characteristics, thereby optimizing resource utilization, improving the timeliness of data processing, providing strong support for big data processing, and meeting the needs of enterprises for efficient data processing and intelligent application. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0016] Figure 1 The figure is a flow chart of an intelligent data integration method for optimizing the ETL data processing flow according to the present invention.
[0017] Figure 2 A schematic diagram of an intelligent data integration system for optimizing the ETL data processing flow according to the present invention. DETAILED DESCRIPTION
[0018] In the specific steps of the intelligent data integration method for realizing the optimization of the ETL data processing flow, which will be described in detail below, various embodiments of the present disclosure will be described more comprehensively. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.
[0019] For example, with the rapid development of big data technology, enterprises have an increasingly urgent need for efficient and intelligent data processing applications. Although ETL technology plays an important role in the field of data integration, its efficiency and accuracy are crucial to improving data quality. However, current ETL tasks face some challenges that limit the efficiency and flexibility of data integration. The primary problem is that the execution of ETL tasks is overly dependent on manually written scripts or complex configurations. This approach not only greatly increases the development cost and maintenance complexity, but is also prone to interruption of the ETL process or decline in data quality due to human negligence or errors. Secondly, with the frequent changes in data sources and target systems, the traditional manual ETL method seems to be unable to cope with it and is difficult to adapt to these changes quickly. This has led to a significant reduction in the flexibility and response speed of data integration, which cannot meet the growing data processing needs of enterprises. In addition, in the process of data loading, a reasonable scheduling strategy is crucial to optimizing resource utilization and improving the timeliness of data processing. However, the existing data loading process often adopts a fixed scheduling strategy that cannot be flexibly adjusted according to the actual needs of the data and business scenarios. This rigid scheduling method may not only result in the failure to load key data into the system in a timely manner, but may also cause waste of resources due to unreasonable resource allocation.
[0020] In response to the above problems, this embodiment provides an intelligent data integration method for optimizing the ETL data processing process. The method can optimize the ETL data processing process, improve the efficiency and flexibility of data integration, and can automatically analyze metadata, identify key information, fill in missing data, and automatically generate ETL tasks, thereby reducing manual intervention, development costs, and maintenance difficulties.
[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] See also Figure 1 FIG. 1 is a flowchart of an intelligent data integration method for optimizing the ETL data processing process in a specific embodiment, the method comprising: S1. Collect metadata information from the source library, analyze the metadata information, and identify key information; It should be noted that the extraction and identification of metadata is a crucial link, which is directly related to the efficiency and quality of subsequent analysis, storage and utilization of data; S2. Use the large model analysis system to analyze metadata information and key information, identify missing data in the source library, and then fill in the missing data based on the analysis results; It should be noted that the big model analysis technology is used to deeply analyze the data elements and intelligently fill in the missing metadata to ensure the integrity and accuracy of the data; S3. Perform basic configuration of ETL tasks and generate full ETL tasks and incremental ETL tasks; It should be noted that the ETL task script is generated according to the preset rules and algorithms to achieve accurate conversion and deep cleaning of data and improve the accuracy of data integration; S4. Define the scheduling strategy for data processing and start the execution of ETL tasks according to the scheduling strategy for data processing; It should be noted that through customized scheduling strategies, the optimal scheduling interval and order of data loading can be intelligently planned according to the importance, urgency and business needs of the data, ensuring that the data can be transmitted to the target system stably and orderly during the integration process, thereby enhancing the reliability and stability of data integration.
[0023] This embodiment optimizes the ETL data processing flow through an intelligent data integration method, and improves the efficiency and flexibility of data integration. The method can automatically analyze metadata, identify key information, fill in missing data, and automatically generate ETL tasks, reducing manual intervention, development costs, and maintenance difficulties.
[0024] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another intelligent data integration method for optimizing the ETL data processing process is provided, and the method includes: S1. Collect metadata information from the source library, analyze the metadata information, and identify key information; the specific steps of step S1 are as follows: S11. Start data integration, configure source library parameters, and connect to the source library according to the source library connection information; S12. Automatically scan the tables and fields in the source library and collect metadata information of the source library; the metadata information includes object name, structure, field type and data constraints; S13. Use intelligent recognition algorithms to conduct preliminary analysis on the collected metadata information and identify key information that characterizes the features; It should be noted that by collecting metadata information from the source library, the accuracy and completeness of the metadata information are ensured. By automatically scanning the tables and fields in the source library and collecting metadata information, a reliable data foundation is provided for subsequent analysis and processing. S2. Use the large model analysis system to analyze the metadata information and key information, identify the missing data in the source library, and then fill in the missing data based on the analysis results; the specific steps of step S2 are as follows: S21. Input the collected metadata information and identified key information into the existing large model analysis system for in-depth analysis, and identify the associations between data and the missing data in the source library; S22. Based on the identification results of the large model analysis system, the missing data in the source database is inferred based on the association between the data and combined with historical data or business rules, and then filled with the inferred data; It should be noted that S3. Perform basic configuration of ETL tasks and generate full ETL tasks and incremental ETL tasks; the specific steps of step S3 are as follows: S31. Obtain the source database data after missing data filling and load it into the target table; S32. Determine the source database data, target table, target table incremental primary key, target table entry time, source primary table and organization name of the ETL task, and obtain the basic configuration of the ETL task; S33. Determine the data conversion rules and algorithms for full data conversion according to the standard data structure of the target table, and generate the ELT full script; S34. Obtain the user-defined incremental configuration or automatically recommend the incremental configuration based on the source master table; S35. Generate an ETL incremental script based on the ETL full script and incremental configuration; S36. Generate ETL full tasks and ETL incremental tasks according to the basic configuration, incremental configuration, ETL full script and ELT incremental script of the ETL task; It should be noted that the use of large model analysis systems to deeply analyze metadata information and key information, identify the associations and missing data between data, and perform inferences and fill in data based on associations and historical data or business rules, thereby improving data integrity and consistency and providing a more reliable data source for data integration; S4. Define the scheduling strategy for data processing and start the execution of the ETL task according to the scheduling strategy for data processing; the specific steps of step S4 are as follows: S41. Define the scheduling strategy for source database data based on business needs, importance and urgency; S42. Sort and schedule each ETL task according to the scheduling strategy, and determine the execution time and execution order of each ETL task; S43. Start ETL task execution; S44. Identify whether the current time point satisfies the execution of a certain ETL task; If yes, take the ETL task as the current task and go to step S45; If not, wait for the set time period and return to step S44; S45. Determine the current task type; If it is a full ETL task, go to step S48; If it is an ETL incremental task, go to step S46; S46. Determine whether the ETL full task has been successfully executed; If yes, go to step S47; If not, return to step S44; S47. Execute the ETL incremental script of the current ETL incremental task, and return to step S44; S48. Execute the ETL full script of the current ETL full task, and return to step S44; It should be noted that the data processing scheduling strategy is defined according to business needs, importance and urgency, and the ETL tasks are sorted and scheduled according to the scheduling strategy, which optimizes resource utilization, improves the timeliness of data processing, and ensures that key data can be loaded into the target system in a timely manner.
[0025] In an embodiment of the present invention, based on step S13, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0026] The specific steps of step S13 are as follows: S131. After cleaning and normalizing the collected metadata information, key fields are extracted, including table name, field name, field type, primary key constraint, and foreign key constraint; S132. Extract features from key fields of metadata information, and extract rules representing primary keys, foreign keys, and indexes as key features; S133. Use a machine learning algorithm to build a classification model, and use the annotated historical metadata information to train to obtain a trained classification model; Exemplarily, the key features of the metadata information are extracted by combining the schema information and the constraints of the database; The input of the classification model is the feature vector of metadata. These feature vectors can be extracted from the database table structure, including but not limited to table name, field name, field type (such as INT, VARCHAR, DATE, etc.), field constraints (such as PRIMARY KEY, FOREIGN KEY, UNIQUE, INDEX, NOT NULL, etc.); For each field, its features can be encoded as a vector, for example: Table name: "orders" -> encoded as a specific identifier or index Field name: "order_id" -> encoded as a specific identifier or index Field type: "INT" -> encoded as a number or category identifier Field constraints: "PRIMARY KEY" -> encoded as a binary bit or class flag (e.g. 1 for primary key, 0 for not) These feature vectors will serve as input to the classification model; Set the output of the classification model to be the classification label of the field, that is, whether the field is a primary key, foreign key, or index (or none of them). The output can be a discrete set of category labels, for example: Primary Key: 1 Foreign Key: 2 Index: 3 Other: 0 Taking the machine learning algorithm of logistic regression as an example, the parameters of the classification model may include the weight vector W and the bias term b; For each input feature vector x, the output y of the model can be calculated by the following formula: y = softmax(Wx + b) The softmax function converts the output into a probability distribution so that the sum of the probabilities of all output categories is 1. In the training phase, the classification model is trained to find the optimal W and b so that the classification model has the highest prediction accuracy on the training data; S134. After processing and feature extraction of the newly collected metadata information, the newly collected metadata information is input into the trained classification model for classification, and primary keys, foreign keys and indexes are identified; It should be noted that by analyzing metadata information, extracting key fields and features, and building a classification model to identify primary keys, foreign keys, and indexes, the accuracy and efficiency of data processing are improved, providing strong support for subsequent data conversion and loading.
[0027] In an embodiment of the present invention, based on step S36, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0028] The specific steps of step S36 are as follows: S361. Review the basic configuration of ETL tasks; If the review is passed, proceed to step S362; If the review fails, return to step S32; S362. Review the full ELT script; If the review is passed, proceed to step S363; If the review fails, return to step S33; S363. Determine whether the ETL incremental script needs to be reviewed; If yes, go to step S364; If not, proceed to step S365; S364. Review ETL incremental scripts; If the review is passed, proceed to step S365; If the review fails, return to step S34; S365. Determine whether the ETL full task is generated; If yes, go to step S4; If not, proceed to step S366; S366. Generate ETL full task according to ETL full script; S367. Determine whether it is necessary to generate an ETL incremental task; If yes, go to step S368; If not, proceed to step S4; S368. Determine whether the ETL incremental script has been confirmed; If yes, go to step S369; If not, return to step S364; S369. Generate ETL incremental tasks according to the ETL incremental script; It should be noted that through the ETL task configuration process of basic configuration, full ETL task configuration and incremental ETL task configuration, ETL scripts and tasks are automatically generated, which reduces the time for manual script writing and configuration and improves the degree of automation of data integration.
[0029] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0030] like Figure 2 As shown, the following is an embodiment of the intelligent data integration system for realizing ETL data processing flow optimization provided by the embodiment of the present disclosure. The system and the intelligent data integration method for realizing ETL data processing flow optimization in the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the intelligent data integration system for realizing ETL data processing flow optimization, reference can be made to the embodiment of the above-mentioned intelligent data integration method for realizing ETL data processing flow optimization.
[0031] The system includes: The metadata extraction and identification module is used to collect metadata information from the source library, analyze the metadata information, and identify key information; The metadata analysis and data filling module is used to analyze metadata information and key information using the large model analysis system, identify missing data in the source library, and then fill in the missing data based on the analysis results; The ETL task configuration module is used to perform basic configuration of ETL tasks, full ETL task configuration, and incremental ETL task configuration, and generate ETL tasks based on the configuration information; The ETL task scheduling execution module is used to define the scheduling strategy for data processing and start the execution of ETL tasks according to the scheduling strategy.
[0032] This embodiment realizes a fully automated process from metadata collection to ETL task execution through the collaborative work of the metadata extraction and identification module, the metadata analysis and data filling module, the ETL task configuration module and the ETL task scheduling and execution module, thereby improving the efficiency and accuracy of data integration.
[0033] The intelligent data integration method for realizing ETL data processing flow optimization provided by the embodiment of the present application can be applied to electronic devices. It will be appreciated by those skilled in the art that the electronic device structure involved in the embodiment of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or less components than shown, or combine certain components, or different component arrangements. In an embodiment of the present invention, the electronic device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the embodiments of the present application described herein and / or required.
[0034] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, buttons, a camera, a display, and a SIM card interface, etc.
[0035] It is to be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0036] The processor may include one or more processing units, for example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated into one or more processors.
[0037] The processor can be the nerve center and command center of the electronic device. The controller can generate an operation control signal according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0038] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. The memory may store instructions or data that the processor has just used or is cyclically used. If the processor needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor, and thus improves system efficiency.
[0039] The above-mentioned electronic device implements the intelligent data integration method of the present application for optimizing the ETL data processing process, collects metadata information from the source library, analyzes the metadata information, and identifies key information; uses a large model analysis system to parse the metadata information and key information, and identifies the missing data in the source library, and then fills the missing data according to the analysis results; performs basic configuration of ETL tasks, and generates full ETL tasks and incremental ETL tasks; defines the scheduling strategy for data processing, and starts the execution of ETL tasks according to the scheduling strategy for data processing. The technical solution has achieved the ability to automatically analyze metadata, identify key information, fill in missing data, and automatically generate ETL tasks, reducing manual intervention, and reducing the beneficial effects of development costs and maintenance difficulties.
[0040] The storage medium provided in the present application stores a program product of an intelligent data integration method capable of optimizing the ETL data processing flow.
[0041] The intelligent data integration method for optimizing the ETL data processing process includes: collecting metadata information from the source library, analyzing the metadata information, and identifying key information; using the large model analysis system to parse the metadata information and key information, and identify the missing data in the source library, and then fill in the missing data based on the analysis results; performing basic configuration of ETL tasks, and generating full ETL tasks and incremental ETL tasks; defining the scheduling strategy for data processing, and starting the execution of ETL tasks according to the scheduling strategy for data processing.
[0042] In some possible implementations, the intelligent data integration method for optimizing the ETL data processing flow disclosed herein may be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps described in the above “Exemplary Method” section of this specification according to various exemplary implementations of the present disclosure.
[0043] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0044] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent data integration method for optimizing ETL data processing flow, characterized in that: The steps include: S1. Collect metadata information from the source library, analyze the metadata information, and identify key information; S2. Use the large model analysis system to analyze metadata information and key information, identify missing data in the source library, and then fill in the missing data based on the analysis results; S3. Perform basic configuration of ETL tasks and generate full ETL tasks and incremental ETL tasks; S4. Define the scheduling strategy for data processing and start the execution of ETL tasks according to the scheduling strategy for data processing.
2. The intelligent data integration method for optimizing the ETL data processing flow according to claim 1, characterized in that: The specific steps of step S1 are as follows: S11. Start data integration, configure source library parameters, and connect to the source library according to the source library connection information; S12. Automatically scan the tables and fields in the source library and collect metadata information of the source library; the metadata information includes object name, structure, field type and data constraints; S13. Use intelligent recognition algorithms to conduct preliminary analysis on the collected metadata information and identify key information that characterizes the features.
3. The intelligent data integration method for realizing ETL data processing process optimization as claimed in claim 2, characterized in that: The specific steps of step S13 are as follows: S131. After cleaning and normalizing the collected metadata information, key fields are extracted, including table name, field name, field type, primary key constraint, and foreign key constraint; S132. Extract features from key fields of metadata information, and extract rules representing primary keys, foreign keys, and indexes as key features; S133. Use a machine learning algorithm to build a classification model, and use the annotated historical metadata information to train to obtain a trained classification model; S134. After data processing and feature extraction of the newly collected metadata information, the metadata information is input into the trained classification model for classification to identify the primary key, foreign key and index.
4. The intelligent data integration method for optimizing the ETL data processing flow according to claim 2, characterized in that: The specific steps of step S2 are as follows: S21. Input the collected metadata information and identified key information into the existing large model analysis system for in-depth analysis, and identify the associations between data and the missing data in the source library; S22. Based on the recognition results of the large model analysis system, the missing data in the source library is inferred based on the relationship between the data and combined with historical data or business rules, and then filled with the inferred data.
5. The intelligent data integration method for optimizing the ETL data processing flow according to claim 4, characterized in that: The specific steps of step S3 are as follows: S31. Obtain the source database data after missing data filling and load it into the target table; S32. Determine the source database data, target table, target table incremental primary key, target table entry time, source primary table and organization name of the ETL task, and obtain the basic configuration of the ETL task; S33. Determine the data conversion rules and algorithms for full data conversion according to the standard data structure of the target table, and generate the ELT full script; S34. Obtain the user-defined incremental configuration or automatically recommend the incremental configuration based on the source master table; S35. Generate an ETL incremental script based on the ETL full script and incremental configuration; S36. Generate ETL full tasks and ETL incremental tasks based on the basic configuration, incremental configuration, ETL full script and ELT incremental script of the ETL task.
6. The intelligent data integration method for optimizing the ETL data processing flow according to claim 5, characterized in that: The specific steps of step S36 are as follows: S361. Review the basic configuration of ETL tasks; If the review is passed, proceed to step S362; If the review fails, return to step S32; S362. Review the full ELT script; If the review is passed, proceed to step S363; If the review fails, return to step S33; S363. Determine whether the ETL incremental script needs to be reviewed; If yes, go to step S364; If not, proceed to step S365; S364. Review ETL incremental scripts; If the review is passed, proceed to step S365; If the review fails, return to step S34; S365. Determine whether the ETL full task is generated; If yes, go to step S4; If not, proceed to step S366; S366. Generate ETL full task according to ETL full script; S367. Determine whether it is necessary to generate an ETL incremental task; If yes, go to step S368; If not, proceed to step S4; S368. Determine whether the ETL incremental script has been confirmed; If yes, go to step S369; If not, return to step S364; S369. Generate ETL incremental tasks according to the ETL incremental script.
7. The intelligent data integration method for optimizing the ETL data processing flow according to claim 6, characterized in that: The specific steps of step S4 are as follows: S41. Define the scheduling strategy for source database data based on business needs, importance and urgency; S42. Sort and schedule each ETL task according to the scheduling strategy, and determine the execution time and execution order of each ETL task; S43. Start ETL task execution; S44. Identify whether the current time point satisfies the execution of a certain ETL task; If yes, take the ETL task as the current task and go to step S45; If not, wait for the set time period and return to step S44; S45. Determine the current task type; If it is a full ETL task, go to step S48; If it is an ETL incremental task, go to step S46; S46. Determine whether the ETL full task has been successfully executed; If yes, go to step S47; If not, return to step S44; S47. Execute the ETL incremental script of the current ETL incremental task, and return to step S44; S48. Execute the ETL full script of the current ETL full task and return to step S44.
8. An intelligent data integration system for optimizing ETL data processing flow, characterized in that: include: The metadata extraction and identification module is used to collect metadata information from the source library, analyze the metadata information, and identify key information; The metadata analysis and data filling module is used to analyze metadata information and key information using the large model analysis system, identify missing data in the source library, and then fill in the missing data based on the analysis results; The ETL task configuration module is used to perform basic configuration of ETL tasks, full ETL task configuration, and incremental ETL task configuration, and generate ETL tasks based on the configuration information; The ETL task scheduling execution module is used to define the scheduling strategy for data processing and start the execution of ETL tasks according to the scheduling strategy.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the intelligent data integration method for optimizing the ETL data processing flow as described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent data integration method for optimizing the ETL data processing flow as described in any one of claims 1 to 7 are implemented.