A method, device and computer equipment for obtaining full amount of real-time data
By adopting the FlinkCdc+doris method in real-time large-width table technology, the problem of cross-base analysis and data time correlation span is solved, and the low-cost full-time large-width table generation and data synchronization is achieved, which improves data timeliness and query efficiency.
Patent Information
- Application Number
- CN202210128596.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-11
AI Technical Summary
The existing real-time large-wide table technology is difficult to effectively deal with the problem of cross-store analysis and data time correlation span, resulting in low data timeliness and inaccurate results.
The real-time data full acquisition method based on FlinkCdc+doris is adopted. The binlog data of the mysql database is monitored and collected through the flinkcdc component, written to the kafka system, and the data processing and storage are used to ensure the orderliness and completeness of the data.
It realizes a low-cost solution for cross-library analysis of mysql, generates a full-time large-wide table, optimizes query speed, improves usage efficiency, and solves the downstream data synchronization problem caused by real-time statistics, and changes to the real-time statistics of the Central Plain table modification.
Smart Images

Figure CN114579614B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data, and in particular to a method, device and computer equipment for acquiring full amount of real-time data. Background Art
[0002] As the development of the Internet enters the second half, the timeliness of data is becoming increasingly important to the refined operations of enterprises. The business world is like a battlefield. How to effectively and in real time extract valuable information from the massive amount of data generated every day will be of great help to the company's decision-making and operational strategy adjustments.
[0003] From the perspective of smart business, data results represent user feedback, and the timeliness of obtaining results is particularly important. Rapidly obtaining data feedback can help decision makers make decisions faster and better iterate corresponding software products. Real-time data warehouses play an irreplaceable role in this process.
[0004] Usually, data warehouses want to have data from the first day a new business goes online, and then record it until now. However, real-time stream processing technology emphasizes the current processing status, and there is a certain contradiction between the two, which will make the data timeliness of the current offline data warehouse very low.
[0005] Specifically:
[0006] (1) Currently, the real-time large wide table (data warehouse) is Flink + ClickHouse. The generation of wide tables depends on the Flink component to associate the results of various tables and write them to ClickHouse. ClickHouse is only responsible for queries, and its OLAP support is insufficient, and the performance of table association is very poor.
[0007] (2) The flink component generates a large wide table with a concept of time (a time window must be added and a time range must be set; otherwise, there will be problems with the task, the state will increase infinitely, and OOM will occur). There is also a problem of data delay and other issues that cannot be associated, which will result in inaccurate data results. This is an objective reality.
[0008] (3) The flink client obtains the binary log and writes it to the kafka side. Since flink can set multiple parallelisms to improve efficiency, and kafka can also have multiple partitions. When multiple threads obtain multiple binlogs with the same primary key, the order of the data must be guaranteed. Otherwise, the disordered data will result in incorrect data being landed.
[0009] Let’s take a simple example, such as business: placing an order + payment.
[0010] Placing an order is a piece of information, and payment is a piece of information. To link the two, Flink SQL Join is used. However, there is a time range. Payment must be made within half an hour. Failure to pay is considered a failure. What if the payment is made one month or one year later? Then the order data message will be waiting in the memory. The amount of data is large, and the memory will store it all the time, making it impossible to join the entire historical data.
[0011] In summary, Flink+clickhouse cannot solve the problem of large time correlation span of data, which is also a pain point. Summary of the invention
[0012] In view of this, the present invention provides a method for obtaining full amount of real-time data based on the construction of data warehouse and real-time stream data processing technology.
[0013] To achieve the above purpose, the present invention proposes a method for obtaining full amount of real-time data, which is based on FlinkCdc+doris and includes the following steps:
[0014] S101: Access real-time binary log binlog data: Use the flinkcdc component to monitor and collect binlog data from the mysql database;
[0015] S102: Write binlog data to the Kafka system: Assign a global sorting ID to the collected binlog data. When writing to the Kafka system, the primary key of the table is used as the partitioning strategy. The primary key value adopts hash partitioning, so that the same primary key data of the same table data is in the same partition;
[0016] S103: Create the doris large wide table: Create the doris large wide table, in which all fields except the primary key are replaced by REPLACE_IF_NOT_NULL;
[0017] S104: Write data to the Doris big wide table: Create multiple Flink stream tasks, write multiple DWD layer or DWS layer data in the Kafka system to the Doris big wide table through stream load, so that the data with the same primary key in the Doris big wide table is overwritten and updated;
[0018] S105: Get full data: query the doris large wide table to get full data according to actual needs.
[0019] Furthermore, in step S105, when querying the doris large wide table, a cross-dimensional association query is performed by using the association fields with other tables as query conditions.
[0020] Furthermore, after the data of the doris large wide table is written, when adding a new field to the table at any time, an asynchronous execution method is adopted, and the specific process is as follows:
[0021] S201: In the doris large wide table, a new field d is created; after the field d is created, the incremental data of the previously existing task a is accessed.
[0022] S202: Create a new offline task b, and import all the data in the doris large wide table before the creation of field d into field d;
[0023] S203: During the import process, field d receives both incremental data and historical full data at the same time, and the two are in disorder. The incremental data is constantly updated with the historical full data with the same primary key until the import is completed.
[0024] S204: Create a new Flink temporary task c and connect to data consumption. At this time, the temporary task c and the original task d write data to the field d at the same time. After the temporary task d runs for a period of time T, the temporary task c is closed. At this time, the doris large wide table completes the field update.
[0025] A real-time data full-volume acquisition device, the device comprising:
[0026] Binlog data acquisition unit: access to real-time binary log binlog data: use flinkcdc component to monitor and collect binlog data from mysql database;
[0027] Data partitioning unit: Assigns a global sorting ID to the collected binlog data. When writing to the Kafka system, the primary key of the table is used as the partitioning strategy. The primary key value adopts hash partitioning so that the same primary key data of the same table data is in the same partition;
[0028] doris large wide table creation unit: create doris large wide table, in which all fields except primary key are replaced by REPLACE_IF_NOT_NULL;
[0029] Doris large wide table data filling unit: create multiple flink stream tasks, write multiple dwd layer or dws layer data in the kafka system to the doris large wide table through stream load, so that the data with the same primary key in the doris large wide table is overwritten and updated;
[0030] Full data acquisition unit: query the doris large wide table to obtain full data according to actual needs.
[0031] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements any step of the method for acquiring full amount of real-time data when executing the computer program.
[0032] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any method for acquiring full amount of real-time data.
[0033] The beneficial effects provided by the present invention are:
[0034] 1. It mainly solves the problem of MySQL cross-database analysis, and this solution is low-cost;
[0035] 2. Generate a full-scale real-time large wide table (full-scale data large wide table) based on Flink+doris, directly perform real-time statistics on the large wide table, optimize query speed, and improve usage efficiency;
[0036] 3. The problem of downstream data synchronization caused by modifying the original table in real-time statistics is solved with minimal cost;
[0037] 4. Based on Doris, the problem of generating large wide tables with full historical data is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flow chart of a method for acquiring full amount of real-time data according to the present invention;
[0039] Figure 2 This is a simple example of creating a doris large wide table;
[0040] Figure 3 It is the process of writing data to a large wide table;
[0041] Figure 4 This is a schematic diagram of data update. DETAILED DESCRIPTION
[0042] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0043] First, the relevant terms are explained uniformly as follows:
[0044] Flink is a top-level open source project of Apache and a computing engine for real-time processing.
[0045] Flinkcdc: is the flink-cdc-connectors component developed by the Flink community. It is a source component that can directly read full data and incremental change data from databases such as MySQL and PostgreSQL. CDC means monitoring and capturing database changes. These changes are fully recorded in the order in which they occur and written to the message middleware for other services to subscribe and consume;
[0046] Mysql Binlog is a binary log file that records changes to the database within Mysql.
[0047] Kafka is a high-throughput distributed publish-subscribe messaging system;
[0048] Doris is an MPP OLAP system that provides high-performance analysis and report query capabilities on large data sets at a low cost;
[0049] This invention provides a method for obtaining full real-time data. For the basic idea, please refer to Figure 1 . Figure 1 is a flow chart of the method of the present invention.
[0050] S101: Access real-time binary log binlog data: Use the flinkcdc component to monitor and collect binlog data from the mysql database;
[0051] It should be noted that the flinkcdc component monitors the binlog logs in the mysql database, then connects the log data to the message middleware kafka, and then develops an application to consume the data. Or you can use other tools in existing commercial libraries, such as Alibaba's logtail tool, and then develop a program for consumption.
[0052] It is easy to understand that in this application, the method of obtaining real-time changing data by monitoring the binlog log file of the MySQL database does not affect the database performance compared to traditional and ordinary data query and export solutions, and solves the data synchronization performance and timeliness problems.
[0053] S102: Write binlog data to the Kafka system: Assign a global sorting ID to the collected binlog data. When writing to the Kafka system, the primary key of the table is used as the partitioning strategy. The primary key value adopts hash partitioning, so that the same primary key data of the same table data is in the same partition;
[0054] It should be noted that the binlog data of the MySQL database is written to the Kafka system in real time. During this process, the binlog is at the second level. However, if the data in the same second is out of order, the final data will be wrong.
[0055] Therefore, when collecting binlog, this application assigns a global sorting ID to the collected binlog. Furthermore, when writing to Kafka, the primary key of the table is used as the partitioning strategy, and the primary key value hash partition is used to ensure that the same primary key data of the same table data is in the same partition, which ultimately ensures the orderliness of the data and also ensures that the data is ordered when Doris data is written and updated.
[0056] It should be further explained that the data of the Kafka system is divided into ODS layer, DWD layer and DWS layer data. The data of the Kafka system is parsed into a standard format after being processed by ETL.
[0057] S103: Create the doris large wide table: Create the doris large wide table, in which all fields except the primary key are replaced by REPLACE_IF_NOT_NULL;
[0058] Please refer to Figure 2 , Figure 2 This is a simple example of creating a doris large wide table;
[0059] Figure 2 In the example, a large wide table named "test" is created, which includes primary keys key1 and key2. All fields except the primary key are REPLACE_IF_NOT_NULL.
[0060] In this application, through such processing method, it can be achieved that when the modification and deletion data of binog is accessed, the corresponding primary key data will be overwritten and updated.
[0061] S104: Write data to the Doris big wide table: Create multiple Flink stream tasks, write multiple DWD layer or DWS layer data in the Kafka system to the Doris big wide table through stream load, so that the data with the same primary key in the Doris big wide table is overwritten and updated;
[0062] It should be noted that when writing multiple Flink stream data to the Doris large wide table at the same time, multiple Flink tasks are created, and data from multiple DWD layers or DWS layers are written to Doris through stream load. The same primary key data will be overwritten and updated.
[0063] Please refer to Figure 3 , Figure 3It is the process of writing data to a large wide table.
[0064] For example, in a certain embodiment, three stream data are included;
[0065] First stream: Key1, key2, value1, value2
[0066] Second stream: Key1, key2, value3, value4
[0067] The third stream: key2, value5, value6;
[0068] The three streams are written simultaneously to the large wide table.
[0069] S105: Get full data: query the doris large wide table to get full data according to actual needs.
[0070] In step S105, when querying the doris large wide table, a cross-dimensional association query is performed by using the association fields with other tables as query conditions.
[0071] It should be noted that when querying the doris large wide table, this table can be treated as a large table, and can be associated with other tables for query (cross-dimensional associated query). The doris table can be optimized for a single label, similar to roll up, to improve the query speed.
[0072] After the data of the doris large wide table is written, when adding new fields to the table at any time, an asynchronous execution method is adopted. The specific process is as follows:
[0073] S201: In the doris large wide table, a new field d is created; after the field d is created, the incremental data of the previously existing task a is accessed.
[0074] S202: Create a new offline task b, and import all the data in the doris large wide table before the creation of field d into field d;
[0075] S203: During the import process, field d receives both incremental data and historical full data at the same time, and the two are in disorder. The incremental data is constantly updated with the historical full data with the same primary key until the import is completed.
[0076] S204: Create a new Flink temporary task c and connect to data consumption. At this time, the temporary task c and the original task d write data to the field d at the same time. After the temporary task d runs for a period of time T, the temporary task c is closed. At this time, the doris large wide table completes the field update.
[0077] The above process, in simple terms, is: if the historical value of a field in the large wide table needs to be modified, this application adds a doris table field for asynchronous execution.
[0078] To add a new field, first import the historical data into the new field of Doris, create a new temporary task, and then access the incremental data volume. Write it to the new and old fields of Doris at the same time as the original task, run for a period of time (set the time, and call the script to kill the task when the time comes), close the temporary task, modify the original incremental task to write the field, restart it, and then consume the data source. This ensures the update and synchronization of the data.
[0079] For a better explanation, please refer to Figure 4 , Figure 4 It is a schematic diagram of data update;
[0080] As mentioned above, in a certain embodiment, three stream data are included;
[0081] First stream: Key1, key2, value1, value2
[0082] Second stream: Key1, key2, value3, value4
[0083] The third stream: key2, value5, value6;
[0084] The third stream only has key2. To update its data, it can be supplemented through the dimension table (cross-dimensional association query), see Figure 4 The upper right corner of the .
[0085] That is, the third flow needs to associate the dimension table to complete the key and becomes:
[0086] Key1, key2, value5, value6. The process adopted is the process described in steps S201 to S204. Finally, the key2 field in the third stream data is supplemented and the data is overwritten and updated through this process.
[0087] A real-time data full-volume acquisition device, the device comprising:
[0088] Binlog data acquisition unit: access to real-time binary log binlog data: use flinkcdc component to monitor and collect binlog data from mysql database;
[0089] Data partitioning unit: Assigns a global sorting ID to the collected binlog data. When writing to the Kafka system, the primary key of the table is used as the partitioning strategy. The primary key value adopts hash partitioning so that the same primary key data of the same table data is in the same partition;
[0090] doris large wide table creation unit: create doris large wide table, in which all fields except primary key are replaced by REPLACE_IF_NOT_NULL;
[0091] Doris large wide table data filling unit: create multiple flink stream tasks, write multiple dwd layer or dws layer data in the kafka system to the doris large wide table through stream load, so that the data with the same primary key in the doris large wide table is overwritten and updated;
[0092] Full data acquisition unit: query the doris large wide table to obtain full data according to actual needs.
[0093] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements any step of the method for acquiring full amount of real-time data when executing the computer program.
[0094] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any method for acquiring full amount of real-time data.
[0095] The beneficial effects of the present invention are:
[0096] 1. It mainly solves the problem of MySQL cross-database analysis, and this solution is low-cost;
[0097] 2. Generate a full-scale real-time large wide table (full-scale data large wide table) based on Flink+doris, directly perform real-time statistics on the large wide table, optimize query speed, and improve usage efficiency;
[0098] 3. The problem of downstream data synchronization caused by modifying the original table in real-time statistics is solved with minimal cost;
[0099] 4. Based on Doris, the problem of generating large wide tables with full historical data is solved.
[0100] In the absence of conflict, the above embodiments and features of the embodiments of the present invention may be combined with each other.
[0101] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for obtaining full amount of real-time data, characterized in that: The following steps are involved: S101: Access real-time binary log binlog data: Use the flinkcdc component to monitor and collect binlog data from the mysql database; S102: Write binlog data to the Kafka system: Assign a global sorting ID to the collected binlog data. When writing to the Kafka system, the primary key of the table is used as the partitioning strategy. The primary key value adopts hash partitioning, so that the same primary key data of the same table data is in the same partition; S103: Create the doris large wide table: Create the doris large wide table, in which all fields except the primary key are replaced by REPLACE_IF_NOT_NULL; S104: Write data to the Doris big wide table: Create multiple Flink stream tasks, write multiple DWD layer or DWS layer data in the Kafka system to the Doris big wide table through stream load, so that the data with the same primary key in the Doris big wide table is overwritten and updated; S105: Get full data: query the doris large wide table to get full data according to actual needs.
2. A method for obtaining full amount of real-time data according to claim 1, characterized in that: In step S105, when querying the doris large wide table, a cross-dimensional association query is performed by using the association fields with other tables as query conditions.
3. A method for obtaining full amount of real-time data according to claim 1, characterized in that: After the data of the doris large wide table is written, when adding new fields to the table at any time, an asynchronous execution method is adopted. The specific process is as follows: S201: In the doris large wide table, a new field d is created; after the field d is created, the incremental data of the original task a is accessed; S202: Create a new offline task b, and import all the data in the doris large wide table before the creation of field d into field d; S203: During the import process, field d receives both incremental data and historical full data at the same time, and the two are in disarray. The incremental data continuously overwrites the historical full data with the same primary key until the import is completed. S204: Create a new Flink temporary task c and connect to data consumption. At this time, the temporary task c and the original task d write data to the field d at the same time. After the temporary task d runs for a period of time T, the temporary task c is closed. At this time, the doris large wide table completes the field update.
4. A real-time data full-volume acquisition device, characterized in that: The device comprises: Binlog data acquisition unit: access to real-time binary log binlog data: use flinkcdc component to monitor and collect binlog data from mysql database; Data partitioning unit: Assigns a global sorting ID to the collected binlog data. When writing to the Kafka system, the primary key of the table is used as the partitioning strategy. The primary key value adopts hash partitioning so that the same primary key data of the same table data is in the same partition; doris large wide table creation unit: create doris large wide table, in which all fields except primary key are replaced by REPLACE_IF_NOT_NULL; Doris large wide table data filling unit: create multiple flink stream tasks, write multiple dwd layer or dws layer data in the kafka system to the doris large wide table through stream load, so that the data with the same primary key in the doris large wide table is overwritten and updated; Full data acquisition unit: query the doris large wide table to obtain full data according to actual needs.
5. A computer device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for acquiring full amount of real-time data as described in any one of claims 1 to 3 are implemented.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for acquiring full real-time data as claimed in any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Device and method for processing online production data in real time in streaming manner
CN111209278A
Order data indexing method and system, computer equipment and storage medium
CN113934713A