A Real-time Data Loading Method and System Based on Flume

Through the technical combination of OGG+Flume+Kafka, the regulatory authorities' demand for real-time data reporting is solved, the rapid synchronization and flexible loading of data are achieved, the implementation costs are reduced, and the implementation costs are applicable to more data recipients.

CN113504950BActive Publication Date: 2025-07-04CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110769047.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-07
Publication Date
2025-07-04
Estimated Expiration
2041-07-07

AI Technical Summary

Technical Problem

The existing technology is difficult to meet the regulatory authorities' requirements for real-time data reporting. The traditional data collection method is low in time, and the data synchronization method of OGG for bigdata is not flexible enough to be flexible with other components.

Method used

Using the technical combination of OGG+Flume+Kafka, bank transaction data is collected in real time through OGG components, and the data is parsed using the Flume component, and the parsed data is loaded into Kafka Topic for use by downstream data processing programs.

Benefits of technology

It realizes the synchronization of data from the transaction system to the data processing and reporting end within a few seconds, reduces the cost of implementing data acquisition, provides more technical choices, and has more flexible data loading methods, and is suitable for more data recipients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113504950B_ABST
    Figure CN113504950B_ABST
Patent Text Reader

Abstract

The present invention provides a real-time data loading method and system based on Flume, which relates to the technical field of big data collection and storage. The method includes: configuring OGG components at the source end and the target end respectively, and configuring Flume components at the target end; the source end collects bank transaction data in real time through the OGG component; the target end reads the bank transaction data through the OGG component and uses the Flume component to parse the bank transaction data; the parsed bank transaction data is loaded into the Kafka Topic for use by downstream data processing programs. The present invention uses the open-source Flume for data extraction, which can reduce the cost of implementing data collection and provide more technical options for the data loading end; and through the parsing of data files by Flume, it can be applied to more data receivers, and the data loading method is more flexible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data collection and storage, and particularly to a real-time data loading method and system based on Flume. Background Art

[0002] Since the regulatory frequency of regulatory agencies has increased the real-time reporting method on the basis of the original common daily, monthly, quarterly, and annual reports, in order to meet the reporting timeliness, each financial institution must adopt a new real-time data collection and processing method. Refer to Figure 5 and Figure 6 as shown, which are two data loading schemes in the prior art:

[0003] One, as Figure 5 shown, is a data loading method based on data files and batch processing in the prior art.

[0004] In order to collect data from each upstream trading system to the data warehouse for data processing and analysis, this scheme generally selects to unload data from the source system at the end of the day without affecting the efficiency of the trading line, and then transmits it to the data warehouse through a file transfer component, and loads it into the database through a source-attached program. The main processes are as follows:

[0005] S51: Unload trading line data, and unload the trading line data into data files at the table level according to the agreed file format and field delimiter.

[0006] S52: Transmit data files, and transmit the trading line data files to the document server of the data warehouse through file transfer components such as NFT or BDE.

[0007] S53: Batch process and load data into the database, read the data from the file by writing a loading script and insert it into the specified database. Generally, a cumulative operation will be performed according to a certain algorithm during loading.

[0008] When collecting data by this scheme, it is necessary to unload data from the production database of the trading line, which will affect the current transactions of the trading line. Generally, the unloading operation is selected at the end of the day (24:00). Therefore, the collection frequency is low, and the timeliness of the warehouse to obtain data is low;

[0009] Since the trading data of a day needs to be synchronized to the data warehouse, generally the data volume is large and the processing time is long. Before the data is finally loaded, the downstream data lines cannot use the data.

[0010] Two, refer to Figure 6 , which is a data loading method based on OGG for bigdata in the prior art.

[0011] This solution uses the OGG technology to synchronize transaction data from the transaction line, and loads the data into the hadoop platform (Kafka, Hdfs, etc.) through the OGG for bigdata plug-in for data consumption. The main process is as follows:

[0012] S61. When insert, delete, or update operations are performed on the original transaction component database, the OGG component will obtain the data by analyzing the logs.

[0013] S62. The OGG client obtains the changes in the OGG server data.

[0014] S63. OGG maps the real-time collected transaction data into the Kafka Topic through the OGG for big data component. After insert and update data, Topic: ZTVOUCHER_INS; before delete and update data, Topic: ZTVOUCHER_DEL.

[0015] This solution must perform data synchronization through OGG for bigdata. OGG for bigdata is not a standard OGG component and needs to be purchased additionally. Moreover, there are certain limitations. For example, the data synchronization method of OGG for bigdata is not flexible enough and cannot be programmed to cooperate with other existing data loading components.

[0016] From the perspective of traditional data collection methods, it is mainly based on unloading data from files, and data synchronization is carried out through file transfer and batch processing for loading into the database. In terms of timeliness, it is generally T+1 day, that is, data is sent and loaded into the database at the end of each day. This method can no longer meet the requirements of regulatory agencies for some real-time reporting tasks.

[0017] In summary, there is an urgent need for a technical solution that can overcome the above defects and meet the requirements of real-time reporting tasks. Summary of the Invention

[0018] To solve the problems existing in the prior art, the present invention proposes a real-time data loading method and system based on Flume. Through the technical combination of OGG+Flume+Kafka, it can complete the synchronization of data from the transaction system within a few seconds after it occurs to the data processing and reporting end. Among them, the use of open-source Flume for data extraction can reduce the cost of implementing data collection, providing more technical options for the data loading end; and through the parsing of data files by Flume, it can be applied to more data receivers, and real-time data loading can be realized by convenient and fast configuration according to actual needs, and the data loading method is more flexible.

[0019] In the first aspect of the embodiments of the present invention, a real-time data loading method based on Flume is proposed, and the method includes:

[0020] Configure OGG components at the source end and the target end respectively, and configure Flume components at the target end;

[0021] The source end collects bank transaction data in real time through the OGG component;

[0022] The target end reads bank transaction data through the OGG component, and uses the Flume component to parse the bank transaction data;

[0023] Load the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs.

[0024] Further, the source end collects bank transaction data in real time through the OGG component, including:

[0025] Collect corporate deposit data of the bank through the OGG component, including: details of corporate demand deposit contract accounts and details of corporate time deposit contract accounts.

[0026] Further, the OGG component includes: Manager process, Extract process, Pump process and Replicat process.

[0027] Further, configuring the OGG component at the source end includes:

[0028] Install OGG software in the database at the source end, including at least: Manager process, Extract process and Pump process; where

[0029] The Manager process is the control process of GoldenGate, which is used to start, monitor, restart the processes of GoldenGate, report errors and time, allocate data storage space, and publish threshold reports;

[0030] The Extract process runs at the database source end and is used to capture data from the source end data table or log;

[0031] The Pump process runs at the database source end and is used to send Trail files in the form of data blocks to the target end through the TCP / IP protocol.

[0032] Further, the database at the source end includes an Oracle database.

[0033] Further, configuring the OGG component at the target end includes:

[0034] Install the OGG component in the database at the target end and receive data in the form of file landing.

[0035] Furthermore, configure the OGG component at the target end, including:

[0036] Configure the Replicat process at the target end; wherein, the Replicat process is used to receive the delivered data, read the content in the Trail file at the target end, and parse it to obtain DML or DDL statements, and apply the parsing results to the target database.

[0037] Furthermore, the database at the target end includes one or more of the following combinations:

[0038] Gbase database, Kudu database, Hive database, Hdfs distributed database, and OLAP database.

[0039] Furthermore, the Flume component includes multiple nodes for logical processing, which are used for distributed log collection, aggregation, and transmission.

[0040] Furthermore, the target end reads bank transaction data through the OGG component and uses the Flume component to parse the bank transaction data, including:

[0041] Set tasks in the distributed message queue according to the bank transaction data;

[0042] Use the Flume component to select matching nodes for the tasks in the distributed message queue through a polling strategy, so that the nodes perform logical processing on the polled tasks, parse the bank transaction data, and write the parsing results into the Kafka data queue for data consumption.

[0043] Furthermore, load the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs, including:

[0044] Create a Topic for the collection object in the Kafka cluster;

[0045] Insert the parsing results obtained by the Flume component parsing the OGG landing file into the Kafka data queue according to the Topic.

[0046] Furthermore, loading the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs also includes:

[0047] Store the parsed bank transaction data in the Kafka data bus for use by downstream data processing programs, including supplying it for data consumption by the Flink computing engine or the Spark cluster.

[0048] In the second aspect of the embodiments of the present invention, a real-time data loading system based on Flume is proposed. The system includes:

[0049] A configuration module, configured to configure OGG components at the source end and the target end respectively, and configure Flume components at the target end;

[0050] An acquisition module, configured to acquire bank transaction data in real time through the OGG component at the source end;

[0051] An analysis module, configured to read bank transaction data through the OGG component at the target end, and analyze the bank transaction data by using the Flume component;

[0052] A loading module, configured to load the analyzed bank transaction data into the Kafka Topic for use by downstream data processing programs.

[0053] Further, the acquisition module is specifically configured to:

[0054] Acquire corporate deposit data of the bank through the OGG component, including: details of corporate current deposit contract accounts and details of corporate time deposit contract accounts.

[0055] Further, the OGG component includes: a Manager process, an Extract process, a Pump process, and a Replicat process.

[0056] Further, the configuration module is specifically configured to:

[0057] Install OGG software in the database at the source end, including at least: a Manager process, an Extract process, and a Pump process; where

[0058] The Manager process is the control process of GoldenGate, used to start, monitor, restart the processes of GoldenGate, report errors and time, allocate data storage space, and issue threshold reports;

[0059] The Extract process runs at the database source end and is used to capture data from the source data table or log;

[0060] The Pump process runs at the database source end and is used to send Trail files in the form of data blocks to the target end through the TCP / IP protocol.

[0061] Further, the configuration module is specifically configured to:

[0062] Install the OGG component in the database at the target end, and receive data in the form of file landing; among them, configure the Replicat process at the target end; the Replicat process is used to receive the delivered data, read the content in the Trail file at the target end, and parse it to obtain DML or DDL statements, and apply the parsing results to the target database.

[0063] Furthermore, the Flume component includes multiple nodes for logical processing, which are used for distributed log collection, aggregation, and transmission.

[0064] Furthermore, the parsing module is specifically used for:

[0065] Set tasks in the distributed message queue according to bank transaction data;

[0066] Use the Flume component to select matching nodes for the tasks in the distributed message queue through a polling strategy, so that the nodes perform logical processing on the polled tasks, parse the bank transaction data, and write the parsing results into the Kafka data queue for data consumption.

[0067] Furthermore, the loading module is specifically used for:

[0068] Create a Topic for the collection object in the Kafka cluster;

[0069] Insert the parsing results obtained by the Flume component parsing the OGG landing file into the Kafka data queue according to the Topic.

[0070] Furthermore, the loading module is specifically used for:

[0071] Store the parsed bank transaction data in the Kafka data bus for use by downstream data processing programs, including supplying it to the Flink computing engine or Spark cluster for data consumption.

[0072] In the third aspect of the embodiments of the present invention, a computer device is proposed, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a real-time data loading method based on Flume.

[0073] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is proposed. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements a real-time data loading method based on Flume.

[0074] The real-time data loading method and system based on Flume proposed by the present invention uses the open-source Flume for data extraction, which can reduce the cost of implementing data collection and provide more technical options for the data loading end. Moreover, through the parsing of data files by Flume, it can be applied to more data receivers, and real-time data loading can be achieved by simply and quickly configuring according to actual needs, making the data loading method more flexible. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0076] Figure 1 It is a schematic flowchart of the real-time data loading method based on Flume according to an embodiment of the present invention.

[0077] Figure 2 It is a schematic diagram of the overall physical deployment relationship of the system according to a specific embodiment of the present invention.

[0078] Figure 3 It is a schematic diagram of the architecture of the real-time data loading system based on Flume according to an embodiment of the present invention.

[0079] Figure 4 It is a schematic diagram of the structure of a computer device according to an embodiment of the present invention.

[0080] Figure 5 It is a schematic diagram of the data loading method based on data files and batch processing in the prior art.

[0081] Figure 6 It is a schematic diagram of the data loading method based on OGG for bigdata in the prior art. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are only provided to enable those skilled in the art to better understand and implement the present invention, and do not limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to convey the scope of the present disclosure to those skilled in the art completely.

[0083] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0084] According to an embodiment of the present invention, a real-time data loading method and system based on Flume are proposed, which relate to the technical field of big data collection and storage.

[0085] In the embodiments of the present invention, the terms to be explained are as follows:

[0086] OGG: Oracle Golden Gate software, which is a log-based structured data replication software. It can achieve real-time capture, transformation and delivery of a large amount of transaction data, realize data synchronization between the source database and the target database, and maintain sub-second data latency. It obtains the incremental changes of data by parsing the online log or archived log of the source database, and then applies these changes to the target database, thereby realizing the synchronization between the source database and the target database. Oracle Golden Gate can achieve sub-second real-time replication of a large amount of data between heterogeneous IT infrastructures (including almost all common operating system platforms and database platforms), and thus can be applied in multiple scenarios such as emergency systems, online reports, real-time data warehouse supply, transaction tracking, data synchronization, centralization / distribution, disaster recovery, database upgrade and transplantation, and dual business centers. At the same time, Oracle Golden Gate can achieve various flexible topological structures such as one-to-one, broadcast (one-to-many), aggregation (many-to-one), two-way, point-to-point, and cascade.

[0087] Flume: A highly available, highly reliable, and distributed system for massive log collection, aggregation, and transmission provided by Cloudera. Flume supports customizing various data senders in the log system for collecting data; at the same time, Flume provides the ability to simply process data and write it to various data receivers (customizable). Flume provides the ability to collect data from data sources such as console (console), RPC (Thrift-RPC), text (file), tail (UNIX tail), syslog (syslog log system), supporting 2 modes of TCP and UDP, and exec (command execution).

[0088] Kafka: An open-source stream processing platform developed by the Apache Software Foundation, written in Scala and Java. Kafka is a high-throughput distributed publish-subscribe messaging system that can handle all action stream data of consumers in a website. Such actions (web browsing, searching, and other user actions) are a key factor in many social functions on modern networks. This data is usually addressed by processing logs and log aggregation due to throughput requirements. For log data and offline analysis systems like Hadoop, but with the limitation of requiring real-time processing, using Kafka is a viable solution. The purpose of Kafka is to unify online and offline message processing through Hadoop's parallel loading mechanism and to provide real-time messages through a cluster.

[0089] Topic: A Topic can be considered a type of message. Each Topic will be divided into multiple partitions, and each partition is an append log file at the storage level. The Kafka cluster can be responsible for the distribution of multiple Topics simultaneously.

[0090] Reference will now be made in detail to several representative embodiments of the present invention to explain the principles and spirit of the present invention.

[0091] Figure 1 It is a schematic flowchart of a real-time data loading method based on Flume according to an embodiment of the present invention. As Figure 1 shown, the method includes:

[0092] Step S101, configure OGG components at the source end and the target end respectively, and configure Flume components at the target end;

[0093] Step S102, the source end collects bank transaction data in real time through the OGG component;

[0094] Step S103, the target end reads the bank transaction data through the OGG component and uses the Flume component to parse the bank transaction data;

[0095] Step S104, load the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs.

[0096] The real-time data loading method based on Flume proposed by the present invention does not use the OGG for bigdata component to load data of Kafka Topic. Instead, it uses the open-source Flume for data extraction, which can reduce the cost of implementing data collection and provide other technical options for the data loading end in addition to OGG for bigdata. Moreover, through the parsing of data files by Flume, it can be applicable to more data receivers, and real-time data loading can be achieved by convenient and fast configuration according to actual requirements, and the data loading method is more flexible.

[0097] In order to explain the above real-time data loading method based on Flume more clearly, the following will be described in detail in combination with each step.

[0098] Step S101, configure the OGG components at the source end and the target end respectively, and configure the Flume component at the target end:

[0099] Among them, the OGG components include: Manager process, Extract process, Pump process and Replicat process.

[0100] In this embodiment, configuring the OGG components at the source end includes:

[0101] Install the OGG software in the database at the source end, including at least: Manager process, Extract process and Pump process; among them,

[0102] The Manager process is the control process of GoldenGate, which is used to start, monitor, restart the processes of GoldenGate, report errors and time, allocate data storage space, and issue threshold reports; GoldenGate software provides a single platform, which can achieve second-level disaster backup for any enterprise environment. GoldenGate is a log-based structured data replication method. It obtains the add, delete, and modify changes of data (the data volume is only about one-fourth of the log) by parsing the online log or archived log of the source database, and then applies these changes to the target database to achieve synchronization and dual-live of the source database and the target database.

[0103] The Extract process runs at the database source end and is used to capture data from the source data table or log;

[0104] The Pump process runs at the database source end and is used to send the Trail file to the target end in the form of data blocks through the TCP / IP protocol.

[0105] The database at the source end includes Oracle database.

[0106] In this embodiment, an OGG component is configured at the target end, including:

[0107] Install the OGG component in the database at the target end, and use the file landing method to receive data.

[0108] Specifically, configure a Replicat process at the target end; wherein, the Replicat process is used to receive the delivered data, read the content in the Trail file at the target end, and parse it to obtain DML or DDL statements, and apply the parsing result to the target database.

[0109] The database at the target end includes one or more of the following combinations:

[0110] Gbase database, Kudu database, Hive database, Hdfs distributed database, and OLAP database.

[0111] Step S102, the source end collects bank transaction data in real time through the OGG component:

[0112] Collect the corporate deposit data of the bank through the OGG component, including: the details of the corporate current deposit contract accounts and the details of the corporate time deposit contract accounts.

[0113] Step S103, the target end reads the bank transaction data through the OGG component and uses the Flume component to parse the bank transaction data:

[0114] Among them, the Flume component contains multiple nodes for logical processing, and is used for distributed log collection, aggregation, and transmission.

[0115] The specific process is as follows:

[0116] Set tasks in the distributed message queue according to the bank transaction data;

[0117] Use the Flume component to select matching nodes for the tasks in the distributed message queue through a polling strategy, so that the nodes perform logical processing on the polled tasks, parse the bank transaction data, and write the parsing result into the Kafka data queue for data consumption.

[0118] Step S104, load the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs:

[0119] Create a Topic for the collection object in the Kafka cluster;

[0120] Insert the parsing result obtained by the Flume component parsing the OGG landing file into the Kafka data queue according to the Topic.

[0121] Specifically, the parsed bank transaction data is stored in the Kafka data bus for use by downstream data processing programs, including supplying data consumption to the Flink computing engine or Spark cluster.

[0122] Through the technical combination of OGG + Flume + Kafka, the present invention can complete the synchronization of data from the occurrence of the transaction system to the data processing and reporting end within a few seconds.

[0123] In actual application scenarios, regulatory agencies require each financial institution to report regulatory data in real time. Therefore, it is necessary to solve data real-time collection, data loading, real-time data processing and reporting at the technical level. If the data reporting fails to be completed within the time limit required by the regulatory agency, it will face regulatory assessment deductions and penalties. In this regard, the present invention can complete the real-time collection of data. After collecting the data from the transaction system, it enters the Kafka data bus for downstream regulatory application processing programs to consume the data.

[0124] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0125] In order to more clearly explain the above real-time data loading method based on Flume, the following will be described in conjunction with a specific embodiment. However, it should be noted that this embodiment is only for better explaining the present invention and does not constitute an improper limitation to the present invention.

[0126] Taking Province A and City B as an example, refer to Figure 2 , which is a schematic diagram of the overall physical deployment relationship of the system of a specific embodiment of the present invention. In Figure 2 , P2: Unified employee channel, as the login interface for internal users, completing functions such as security authentication, permission control, and interface display, and serving as the function of viewing result data in this case. F5: Load balancing. P9: Data integration platform, the main execution platform of the method of this patent, used for data exchange and processing, and serving as the real-time data collection, as well as the work of reading and loading into Kafka in this case. External connection area: The result data after real-time data processing forms an MQ message through the external connection area and is reported to the local financial regulatory agency.

[0127] Specifically, the process of real-time data loading is as follows:

[0128] S21, configure the OGG collection and file receiving methods at the source system and the target end.

[0129] Install the OGG software in the trading oracle database of the source system.

[0130] Install the OGG software at the target end as well, and use the file landing method for data reception.

[0131] In this embodiment, the OGG is used to collect the corporate deposit data of the bank, which is divided into the details of the corporate current deposit contract accounts and the details of the corporate time deposit contract accounts.

[0132] The partial key fields of the source tables are shown in Table 1 (details of corporate current deposit contract accounts) and Table 2 (details of corporate time deposit contract accounts):

[0133] Table 1 Partial key fields of the source table (details of corporate current deposit contract accounts)

[0134]

[0135]

[0136] Table 2 Partial key fields of the source table (details of corporate time deposit contract accounts)

[0137]

[0138]

[0139] S22. Configure Flume at the target end for file parsing.

[0140] Flume is a highly available, highly reliable, and distributed system for massive log collection, aggregation, and transmission provided by Cloudera. In the present invention, Flume is used to perform simple processing on the data and write it to Kafka for subsequent data consumption.

[0141] S23. Perform the operation of loading data to the Kafka Topic at the target end.

[0142] Kafka is an open-source stream processing platform developed by the Apache Software Foundation. It is a high-throughput distributed publish-subscribe messaging system. Since the present invention adopts real-time data collection, the collected real-time transaction data needs to be stored in the data bus and then supplied to downstream data processing programs for consumption, such as Flink or Spark. First, create the Topics for each collection object on Kafka, and then configure Flume to parse the OGG landing files read and insert them into Kafka according to the Topics.

[0143] Refer to Table 3 for the corresponding relationship between the original file names and the file names after being loaded by Flume.

[0144] Table 3 Corresponding Relationship Table between Original File Names and File Names after Loading with Flume

[0145]

[0146]

[0147] The real-time data loading method based on Flume proposed by the present invention does not use the OGG for bigdata component to load data of Kafka Topic. Instead, it uses the open-source Flume for data extraction, which can reduce the cost of implementing data collection and provides other technical options for the data loading end in addition to OGG for bigdata. Through the parsing of data files by Flume, it can be applicable to more data receivers, and the data loading method is more flexible and can be completed through configuration.

[0148] After introducing the method of the exemplary embodiment of the present invention, next, refer to Figure 3 to introduce the real-time data loading system based on Flume of the exemplary embodiment of the present invention.

[0149] The implementation of the real-time data loading system based on Flume can refer to the implementation of the above method, and the repeated parts will not be described again. The terms "module" or "unit" used hereinafter can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and contemplated.

[0150] Based on the same inventive concept, the present invention also proposes a real-time data loading system based on Flume, as Figure 3 shown, the system includes:

[0151] Configuration module 310, configured to configure OGG components at the source end and the target end respectively, and configure Flume components at the target end;

[0152] Collection module 320, configured to collect bank transaction data in real time at the source end through the OGG component;

[0153] Parsing module 330, configured to read bank transaction data at the target end through the OGG component, and parse the bank transaction data by using the Flume component;

[0154] Loading module 340, configured to load the parsed bank transaction data into Kafka Topic for use by downstream data processing programs.

[0155] In this embodiment, the collection module 310 is specifically configured to:

[0156] Collect the corporate deposit data of the bank through the OGG component, including: the details of the corporate demand deposit contract accounts and the details of the corporate time deposit contract accounts.

[0157] In this embodiment, the OGG component includes: a Manager process, an Extract process, a Pump process, and a Replicat process.

[0158] The configuration module 310 is specifically used for:

[0159] Install the OGG software in the database at the source end, including at least: a Manager process, an Extract process, and a Pump process; where

[0160] The Manager process is the control process of GoldenGate, used to start, monitor, and restart the processes of GoldenGate, report errors and time, allocate data storage space, and publish threshold reports;

[0161] The Extract process runs at the database source end and is used to capture data from the source data tables or logs;

[0162] The Pump process runs at the database source end and is used to send the Trail file in the form of data blocks to the target end through the TCP / IP protocol.

[0163] The configuration module 310 is specifically used for:

[0164] Install the OGG component in the database at the target end and receive data in the way of file landing; where, configure the Replicat process at the target end; the Replicat process is used to receive the delivered data, read the content in the Trail file at the target end, and parse it to obtain DML or DDL statements, and apply the parsing results to the target database.

[0165] In this embodiment, the Flume component includes multiple nodes for logical processing, and is used for distributed log collection, aggregation, and transmission.

[0166] The parsing module 330 is specifically used for:

[0167] Set tasks in the distributed message queue according to the bank transaction data;

[0168] Use the Flume component to select matching nodes for the tasks in the distributed message queue through a polling strategy, so that the nodes perform logical processing on the polled tasks, parse the bank transaction data, and write the parsing results into the Kafka data queue for data consumption.

[0169] In this embodiment, the loading module 340 is specifically configured to:

[0170] Create a Topic for the collection object in the Kafka cluster;

[0171] Insert the parsing results obtained by the Flume component parsing the OGG landing file into the Kafka data queue according to the Topic.

[0172] The loading module 340 is specifically configured to:

[0173] Store the parsed bank transaction data in the Kafka data bus for use by downstream data processing programs, including supplying it to the Flink computing engine or Spark cluster for data consumption.

[0174] It should be noted that although several modules of the real-time data loading system based on Flume are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of two or more of the above-described modules can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.

[0175] Based on the foregoing inventive concept, as Figure 4 shown, the present invention also proposes a computer device 400, including a memory 410, a processor 420, and a computer program 430 stored on the memory 410 and executable on the processor 420. When the processor 420 executes the computer program 430, the foregoing real-time data loading method based on Flume is implemented.

[0176] Based on the foregoing inventive concept, the present invention proposes a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing real-time data loading method based on Flume is implemented.

[0177] The real-time data loading method and system based on Flume proposed by the present invention utilize the open-source Flume for data extraction, which can reduce the cost of implementing data collection, provide more technical options for the data loading end; and through the parsing of data files by Flume, it can be applicable to more data receivers, and real-time data loading can be achieved by convenient and quick configuration according to actual needs, and the data loading method is more flexible.

[0178] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0179] The present invention is described with reference to the flowcharts and / or block diagrams of methods and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0180] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0182] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A real-time data loading method based on Flume, characterized in that, This method synchronizes data from the transaction system to the data processing and reporting end through the technical combination of OGG, Flume, and Kafka, including: Configure OGG components at the source end and the target end respectively, and configure Flume components at the target end; The source end collects bank transaction data in real time through the OGG component; The target end reads bank transaction data through the OGG component and uses the Flume component to parse the bank transaction data; Load the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs; Among them, the Flume component includes multiple nodes for logical processing, which are used for distributed log collection, aggregation, and transmission; Among them, the target end reads bank transaction data through the OGG component and uses the Flume component to parse the bank transaction data, including: Set tasks in the distributed message queue according to the bank transaction data; Use the Flume component to select matching nodes for the tasks in the distributed message queue through the polling strategy, so that the nodes perform logical processing on the polled tasks, parse the bank transaction data, and write the parsing results into the Kafka data queue for data consumption; Among them, loading the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs includes: Create a Topic for the collection object in the Kafka cluster; Insert the parsing results obtained by the Flume component parsing the OGG landing file into the Kafka data queue according to the Topic; among them, during the data loading process, the Flume component is used for data extraction and parsing, and the OGG for bigdata component is not used for data loading of the Kafka Topic.

2. The real-time data loading method based on Flume according to claim 1, wherein The source end collects bank transaction data in real time through the OGG component, including: Collect the corporate deposit data of the bank through the OGG component, including: the details of the corporate current deposit contract account and the details of the corporate time deposit contract account.

3. The real-time data loading method based on Flume according to claim 1, wherein The OGG component includes: Manager process, Extract process, Pump process, and Replicat process.

4. The real-time data loading method based on Flume according to claim 3, wherein Configure the OGG component at the source end, including: Install OGG software in the database at the source end, including at least: Manager process, Extract process, and Pump process; among them, The Manager process is the control process of GoldenGate, which is used to start, monitor, restart the GoldenGate process, report errors and time, allocate data storage space, and publish threshold reports; The Extract process runs at the database source end and is used to capture data from the source end data table or log; The Pump process runs at the database source end and is used to send the Trail file to the target end in the form of data blocks through the TCP / IP protocol.

5. The real-time data loading method based on Flume according to claim 4, characterized in that The database at the source end includes an Oracle database.

6. The real-time data loading method based on Flume according to claim 4, wherein Configure the OGG component at the target end, including: Install the OGG component in the database at the target end and receive data in the form of file landing.

7. The real-time data loading method based on Flume according to claim 6, characterized in that Configure the OGG components at the target end, including: Configure the Replicat process at the target end; wherein, the Replicat process is used to receive the delivered data, read the content in the Trail file at the target end, and parse it to obtain DML or DDL statements, and apply the parsing results to the target database.

8. The real-time data loading method based on Flume according to claim 6, wherein The databases at the target end include one or more of the following combinations: Gbase database, Kudu database, Hive database, Hdfs distributed database, and OLAP database.

9. The real-time data loading method based on Flume according to claim 1, characterized in that, Loading the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs further includes: Storing the parsed bank transaction data in the Kafka data bus for use by downstream data processing programs, including supplying it to the Flink computing engine or Spark cluster for data consumption.

10. A real-time data loading system based on Flume, characterized in that, Through the technical combination of OGG, Flume, and Kafka, this system synchronizes data from the transaction system to the data processing and reporting end, including: A configuration module for configuring OGG components at the source end and the target end respectively, and configuring Flume components at the target end; A collection module for collecting bank transaction data in real time through the OGG component at the source end; A parsing module for reading bank transaction data through the OGG component at the target end and parsing the bank transaction data using the Flume component; A loading module for loading the parsed bank transaction data into the Kafka Topic for use by downstream data processing programs; Among them, the Flume component includes multiple nodes for logical processing, which are used for distributed log collection, aggregation, and transmission; Among them, the parsing module is specifically used for: Setting tasks in the distributed message queue according to the bank transaction data; Using the Flume component to select matching nodes for the tasks in the distributed message queue through a polling strategy, enabling the nodes to perform logical processing on the polled tasks, parsing the bank transaction data, and writing the parsing results into the Kafka data queue for data consumption; Among them, the loading module is specifically used for: Creating a Topic for the collection object in the Kafka cluster; Inserting the parsing results obtained by the Flume component parsing the OGG landing file into the Kafka data queue according to the Topic; wherein, during the data loading process, the Flume component is used for data extraction and parsing, and the OGG for bigdata component is not used for data loading of the Kafka Topic.

11. The real-time data loading system based on Flume according to claim 10, characterized in that, The collection module is specifically used for: Collecting corporate deposit data of the bank through the OGG component, including: details of corporate current deposit contract accounts and details of corporate time deposit contract accounts.

12. The real-time data loading system based on Flume according to claim 10, characterized in that The OGG component includes: Manager process, Extract process, Pump process, and Replicat process.

13. The real-time data loading system based on Flume according to claim 12, wherein The configuration module is specifically used for: Installing OGG software in the database at the source end, including at least: Manager process, Extract process, and Pump process; wherein, The Manager process is the control process of GoldenGate, which is used to start, monitor, and restart GoldenGate processes, report errors and time, allocate data storage space, and issue threshold reports; The Extract process runs on the database source side and is used to capture data from source data tables or logs; The Pump process runs on the database source side and is used to send Trail files in the form of data blocks to the target side through the TCP / IP protocol.

14. The real-time data loading system based on Flume according to claim 13, wherein The configuration module is specifically used for: Installing OGG components in the target database and receiving data in a file landing manner; among them, configuring the Replicat process on the target side; the Replicat process is used to receive the delivered data, read the content in the target Trail file, parse it to obtain DML or DDL statements, and apply the parsing results to the target database.

15. The real-time data loading system based on Flume according to claim 10, wherein The loading module is specifically used for: Storing the parsed bank transaction data in the Kafka data bus for use by downstream data processing programs, including supplying data consumption to the Flink computing engine or Spark cluster.

16. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 9.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Power grid operation data sharing system based on large data technology

    CN106339509A

  • Big data-based data access permission control method

    CN112699096A