Data processing method and device, storage medium and computer program product

Data is collected and corrected through dual-link merging, which solves the problem of single data sources and difficult to guarantee consistency, and realizes the acquisition and use of high-quality data.

CN119961325APending Publication Date: 2025-05-09CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510070793.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, data sources are single and data consistency is difficult to guarantee, resulting in the problems of producer data loss and inconsistent data use.

Method used

The dual-link merging method is adopted to collect quasi-real-time data and offline data from the data source layer, and correct the real-time data through offline data alignment to ensure the integrity and consistency of the data.

Benefits of technology

By increasing data acquisition channels and compensation processing for offline data, the accuracy and consistency of data are improved, and reliable data support is provided to ensure the quality of data read by the application system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961325A_ABST
    Figure CN119961325A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, a storage medium and a computer program product, and relates to the technical field of data processing.The method comprises the steps that quasi-real-time data and offline data are collected and output from a data source layer in a double-link merging mode; taking the off-line data as compensation, and performing correction processing on the quasi real-time data; and sending the off-line data and the corrected quasi-real-time data to a storage layer application system database for an application system to read and use. According to the method, the problems that the data source is single and the data consistency is difficult to guarantee can be effectively solved through double-link combined collected data, correction processing is carried out on offline data by aiming at real-time data, and an effective data storage and reading mechanism, so that high-quality data can be generated to be read and used by an application system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, device, storage medium and computer program product. Background Art

[0002] Real-time data can quickly capture changes in the database, thereby reflecting the latest status of business data in a timely manner and achieving high-performance computing of massive data. To obtain real-time data, OGG and CDC are usually deployed to collect the Binlog log of the database. The Binlog log records all changes to the database. Parsing the Binlog log and synchronizing the original data change information obtained from the analysis to Kafka can timely perceive changes in business data. Although Kafka uses multiple methods to ensure that data is not lost on the producer side, broker side, and consumer side, in actual applications, data consistency still requires a high-reliability architecture to ensure it. The existing real-time data acquisition process cannot solve the problem of producer data loss. And relying only on Kafka data collected by OGG / CDC as the data source is prone to problems such as a single data source and lack of data consistency.

[0003] Therefore, how to effectively solve the problem of single data source and difficulty in ensuring data consistency, and produce high-quality data for application systems to read and use, has become an urgent problem to be solved in this application.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of this application is to provide a data processing method, device, storage medium and computer program product, aiming to solve the problems of single data source and difficulty in ensuring data consistency, and to produce high-quality data for application systems to read and use.

[0006] To achieve the above objectives, the present application proposes a data processing method, which includes:

[0007] Collect and generate quasi-real-time data and offline data from the data source layer in a dual-link merging manner;

[0008] Using the offline data as compensation, correcting the quasi-real-time data;

[0009] The offline data and the corrected quasi-real-time data are sent to the storage layer application system database for the application system to read and use.

[0010] In one embodiment, the dual link includes a quasi-real-time link and a data warehouse link, and the step of collecting and generating quasi-real-time data and offline data from the data source layer in a dual-link merging manner includes:

[0011] Use quasi-real-time links to collect and generate quasi-real-time data from the source system at the data source layer in real time;

[0012] Use the data warehouse link to collect and generate offline data from the database of the source system.

[0013] In one embodiment, the step of using the quasi-real-time link to collect and generate quasi-real-time data from a source system at the data source layer in real time includes:

[0014] Use data change capture and synchronization to collect and parse log files from the source system, and forward the parsed log files to the Kafka message middleware;

[0015] Use the Flink cluster to consume the Kafka message middleware to obtain consumption data;

[0016] The consumption data is aggregated at different granularities through custom operators to obtain quasi-real-time data.

[0017] In one embodiment, the step of using the quasi-real-time link to collect and generate quasi-real-time data from a source system at the data source layer in real time includes:

[0018] Write bypass output logic and capture bypass data from the source system of the data source layer according to the bypass output logic;

[0019] Forward the bypass data to the Kafka message middleware, and use the Flink cluster to consume the Kafka message middleware to obtain consumption data;

[0020] The consumption data is aggregated at different granularities through custom operators to obtain quasi-real-time data.

[0021] In one embodiment, the step of using the Flink cluster to consume the Kafka message middleware to obtain consumption data further includes:

[0022] Maintain the data key value file through HBase, and read whether there is a consumption record of the Kafka message middleware in the data key value file;

[0023] If the data key-value file contains a consumption record of the Kafka message middleware, the consumption record is discarded;

[0024] If the consumption record of the Kafka message middleware does not exist in the data key value file, the consumption record is written.

[0025] In one embodiment, the step of using the data warehouse link to collect and generate offline data from the database of the source system includes:

[0026] Extract data warehouse files from the database of the source system to outside the warehouse using a pre-configured data extraction component;

[0027] The data warehouse file is processed in the warehouse to obtain offline data.

[0028] In one embodiment, before the step of extracting the data warehouse file from the database of the source system to the outside of the warehouse using the pre-configured data extraction component, the step further includes:

[0029] Write configuration data extraction components through configuration jobs.

[0030] In addition, to achieve the above objectives, the present application also proposes a data processing device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data processing method described above.

[0031] In addition, to achieve the above objectives, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data processing method described above are implemented.

[0032] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the data processing method described above are implemented.

[0033] One or more technical solutions proposed in this application have at least the following technical effects:

[0034] Quasi-real-time data and offline data are collected and produced from the data source layer in a dual-link merging manner. The dual-link merging method increases the data collection channels, thereby avoiding the singleness of the data source; offline data is used as compensation to correct the quasi-real-time data. Due to its integrity and historical characteristics, offline data can be used as a supplement and verification of quasi-real-time data. When quasi-real-time data is missing, erroneous or inconsistent, it can be corrected by offline data, thereby improving the accuracy and consistency of the data. Offline data and corrected quasi-real-time data are sent to the storage layer application system database for the application system to read and use, so that the data read by the application system includes both offline data and quasi-real-time data, thereby providing reliable data support for the application system. In summary, through dual-link merging to collect data, offline data to correct quasi-real-time data, and effective data storage and reading mechanisms, the problems of single data source and difficult to ensure data consistency can be effectively solved, thereby producing high-quality data for the application system to read and use. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0036] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0037] Figure 1 This is a flow chart of the first embodiment of the data processing method of the present application;

[0038] Figure 2 This is a flow chart of the second embodiment of the data processing method of the present application;

[0039] Figure 3 This is a flow chart of a fourth embodiment of the data processing method of the present application;

[0040] Figure 4 This is a schematic diagram of the Flink quasi-real-time link for this application to collect and produce quasi-real-time data in real time;

[0041] Figure 5 This is a schematic diagram of the dual-link combined output data for this application;

[0042] Figure 6 This is a schematic diagram of the module structure of the data processing device according to an embodiment of the present application;

[0043] Figure 7Schematic diagram of the device structure of the hardware operating environment involved in the data processing method in the embodiment of the present application.

[0044] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0045] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0046] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0047] The main solution of the embodiment of the present application is: to process and output data in a dual-link parallel manner of data warehouse output + quasi-real-time data processing, which solves the problem of a single data source. The data warehouse link processing provides more reliable offline data as compensation, and the data can be corrected when the real-time link processing is abnormal. The dual-path decoupled parallel solution completes highly reliable data processing and provides data integrity protection. Even if data loss occurs during quasi-real-time data collection, it does not affect the final consistency of the data. Finally, the offline data and the corrected quasi-real-time data are sent to the storage layer application system database for the application system to read and use.

[0048] In this embodiment, for ease of description, the following description is made by identifying Flink and the offline data warehouse collaborative system as the execution entity.

[0049] The embodiments of the present application take into account that: since the change operations occurring in the database can be quickly captured through real-time data, the latest status of the business data can be reflected in a timely manner, and high-performance computing of massive data can be achieved. To obtain real-time data, OGG and CDC are usually deployed to collect the Binlog log of the database, and the Binlog log records all the change operations on the database. Parsing the Binlog log and synchronizing the original data change information obtained by the analysis to Kafka can timely perceive the changes in business data. Although Kafka ensures that data is not lost in a variety of ways on the producer side, the Broker side, and the consumer side, in actual applications, data consistency still requires a high-reliability architecture to ensure. The existing real-time data acquisition process cannot solve the problem of producer data loss. And relying only on Kafka data collected by OGG / CDC as the data source is prone to problems such as a single data source and data consistency cannot be guaranteed.

[0050] Therefore, the present application provides a solution to collect and produce quasi-real-time data and offline data from the data source layer in a dual-link merging manner, increase the data collection channels by dual-link merging, thereby avoiding the singleness of the data source; use offline data as compensation to correct the quasi-real-time data, and offline data can be used as a supplement and verification of quasi-real-time data due to its integrity and historical characteristics. When quasi-real-time data is missing, erroneous or inconsistent, it can be corrected by offline data to improve the accuracy and consistency of the data. Send offline data and corrected quasi-real-time data to the storage layer application system database for the application system to read and use, so that the data read by the application system includes both offline data and quasi-real-time data, thereby providing reliable data support for the application system. In summary, by collecting data through dual-link merging, correcting offline data for quasi-real-time data, and an effective data storage and reading mechanism, the problem of single data source and difficult to ensure data consistency can be effectively solved, thereby producing high-quality data for the application system to read and use.

[0051] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a Flink and offline data warehouse collaborative system, etc. The following takes the Flink and offline data warehouse collaborative system as an example to illustrate this embodiment and the following embodiments.

[0052] Based on this, the present application embodiment provides a data processing method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the data processing method of the present application.

[0053] In this embodiment, the data processing method includes steps S10 to S30:

[0054] Step S10, collecting and producing quasi-real-time data and offline data from the data source layer in a dual-link merging manner;

[0055] Dual-link merging refers to a strategy that combines quasi-real-time processing links and offline processing links, aiming to fully utilize the advantages of both processing modes to meet the requirements of different business scenarios for data timeliness and accuracy.

[0056] In order to better illustrate and explain the data processing method of this application, the following uses the output transaction statistics data as an example to illustrate and explain the data processing method of this application. Offline data specifically refers to the complete data set as of 24:00 on the night of the day before the transaction day, and quasi-real-time data specifically refers to the transaction data set on the transaction day. Specifically, if T is the date of the day, then T-1 represents yesterday's date. Offline data is the transaction data set of T-1 day, and quasi-real-time data is the transaction data of T day.

[0057] The link for processing data from the offline data warehouse is combined with the link for processing data from the quasi-real-time link, and the data output is completed by merging two links. The quasi-real-time link is responsible for consuming quasi-real-time data and performing corresponding data processing to generate T-day transaction statistics; while the offline batch link is responsible for providing T-1 day transaction statistics that are less timely but can be used as a backup.

[0058] Step S20, using the offline data as compensation to correct the quasi-real-time data;

[0059] When data loss or abnormality occurs during the collection or processing of quasi-real-time data on day T, the offline data on day T-1 can be used to correct it. Specifically, the T-1 data processed in the offline data warehouse is compared with the T-day data generated in the quasi-real-time link, and the offline data is used to supplement or correct the inconsistent or missing data.

[0060] The final consistency of data can be ensured through compensation and correction of offline data. Even if problems occur during the quasi-real-time data processing on day T, corrections can be automatically made using the offline data in the data warehouse on day T+1, thus ensuring the integrity and accuracy of the data.

[0061] Step S30, sending the offline data and the corrected quasi-real-time data to the storage layer application system database for the application system to read and use.

[0062] The processed and corrected offline data and quasi-real-time data are loaded into the application system database through data loading, so that the application system can directly read these data from the database and obtain statistical results including both historical data of day T-1 and data of day T through simple query. The read and queried data can be used to display product sales information in multiple dimensions, conduct data analysis and decision support, etc.

[0063] This embodiment provides a data processing method, which collects and produces quasi-real-time data and offline data from the data source layer in a dual-link merging manner, increases the data collection channels by dual-link merging, thereby avoiding the singleness of the data source; offline data is used as compensation to correct the quasi-real-time data, and offline data can be used as a supplement and verification of quasi-real-time data due to its integrity and historical characteristics. When quasi-real-time data is missing, erroneous or inconsistent, it can be corrected by offline data, thereby improving the accuracy and consistency of the data. Offline data and corrected quasi-real-time data are sent to the storage layer application system database for the application system to read and use, so that the data read by the application system includes both offline data and quasi-real-time data, thereby providing reliable data support for the application system. In summary, by collecting data through dual-link merging, correcting offline data for quasi-real-time data, and an effective data storage and reading mechanism, the problem of single data source and difficult to ensure data consistency can be effectively solved, thereby producing high-quality data for the application system to read and use.

[0064] Based on the first embodiment of the present application, the second embodiment of the present application is proposed. In the second embodiment of the present application, the same or similar contents as those of the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following.

[0065] On this basis, please refer to Figure 2 , Figure 2 This is a flow chart of the second embodiment of the data processing method of the present application.

[0066] In this embodiment, the dual link includes a quasi-real-time link and a data warehouse link, and the step S10 of collecting and producing quasi-real-time data and offline data from the data source layer in a dual-link merging manner includes steps S11 to S12:

[0067] Step S11, using a quasi-real-time link to collect and generate quasi-real-time data from a source system of a data source layer in real time;

[0068] The data source layer is the lowest level of the Flink and offline data warehouse collaborative system architecture. The data source layer is the original data, which consists of T-1 day data warehouse files and T day quasi-real-time business data. The collection layer and processing layer of the Flink and offline data warehouse collaborative system architecture use quasi-real-time links to collect and produce quasi-real-time data from the source system of the data source layer in real time.

[0069] The quasi-real-time link refers to the use of data change capture and synchronization methods such as CDC / OGG to parse the database Binlog log in the source system for real-time data collection. After forwarding the log file to the Kafka message cluster, the Flink framework is used to consume the data in the Kafka message cluster to finally obtain quasi-real-time data.

[0070] Step S12: Use the data warehouse link to collect and generate offline data from the database of the source system.

[0071] The source of offline data is the database of the source system, which contains various business data in the enterprise operation. The data cutoff time is 24:00 on the previous night, that is, the data of T-1 day. Such data selection helps to ensure the integrity and stability of the data, and facilitates subsequent data analysis and decision support.

[0072] The data warehouse link uses ETL (Extract-Transform-Load) tasks as the main collection method. The ETL task is responsible for extracting data from the database of the source system and performing data cleaning, transformation, and loading on the extracted data.

[0073] Specifically, in a feasible implementation, the step S11 of using a quasi-real-time link to collect and generate quasi-real-time data from a source system of a data source layer in real time may include steps S111 to S113:

[0074] Step S111, collecting and parsing log files from the source system by means of data change capture and synchronization, and forwarding the parsed log files to the Kafka message middleware;

[0075] Collecting and parsing log files from the source system using data change capture and synchronization means collecting and parsing database Binlog logs from the source system using CDC / OGG. Among them, CDC (Change Data Capture) is a technology specifically used to capture and synchronize database changes; OGG (Oracle GoldenGate) is a data change capture and synchronization tool; Binlog logs are binary log files that record all statements that update data in the database, and these statements are saved in binary form on the disk.

[0076] Specifically, the Binlog logs of the MySQL database in the source system are collected in real time through CDC / OGG. The CDC / OGG tool usually establishes a connection with the database and monitors the write operations of the database so as to capture these changes immediately when the data changes.

[0077] It is understandable that the collected Binlog logs need to be parsed in order to extract useful data change information. The parsing process usually involves steps such as identifying the log format, extracting and converting data. The parsed data change information usually exists in the form of stream data and can be consumed by subsequent processing components.

[0078] Kafka is a distributed message publish-subscribe system that allows publishing and subscribing to streaming data in a high-throughput manner. In data processing scenarios, Kafka is often used as a message middleware to transfer data between different components.

[0079] By encapsulating the parsed data change information into Kafka messages and sending them to the specified Kafka topic, the parsed data change information (stream data) is forwarded to the Kafka message middleware. Kafka's distributed architecture and high throughput characteristics enable this process to efficiently process a large amount of data change information. Once the data change information is sent to the Kafka message middleware, it can be consumed by the Flink stream processing framework for subsequent processing and analysis.

[0080] Step S112, using the Flink cluster to consume the Kafka message middleware to obtain consumption data;

[0081] Flink is a distributed stream processing framework that can efficiently process and analyze large-scale data streams. When using Flink to consume Kafka data, you first need to configure and start the Flink cluster. This process includes setting up the necessary job manager (JobManager) and task manager (TaskManager) nodes, as well as configuring related resources (such as memory, CPU, etc.).

[0082] Use the configured Flink cluster to pull data from the Kafka message middleware and consume and process it according to the written processing logic. Consumed data refers to the data records read and processed from the Kafka message middleware.

[0083] It should be noted that in order to ensure the stability and performance of Flink jobs, they also need to be monitored and tuned. This includes monitoring the running status of the job, performance indicators (such as throughput, latency, etc.), and resource usage.

[0084] Step S113, aggregating the consumption data at different granularities through a custom operator to obtain quasi-real-time data.

[0085] In Flink, custom operators are usually used to implement specific logical processing on consumed data. Custom operators can perform operations such as conversion, aggregation, and filtering on input data streams according to business needs.

[0086] In order to obtain quasi-real-time data aggregated at different granularities, you need to write a custom Flink operator. The custom Flink operator reads the data consumed from Kafka, that is, the consumption data, and then groups and aggregates the consumption data according to the aggregation logic preset in the custom operator (such as by time window, by user ID, by product category, etc.), thereby obtaining quasi-real-time data.

[0087] In this implementation, the Binlog logs of the database are collected in real time to generate quasi-real-time data streams. The Flink stream processing architecture is used to efficiently process the generated massive data, achieving low-latency, high-throughput data processing. The Flink custom operator aggregates the data at different granularities, significantly improving the quality of quasi-real-time data.

[0088] Specifically, in another feasible implementation, the step S11 of using a quasi-real-time link to collect and generate quasi-real-time data from a source system of a data source layer in real time may include steps A111 to A113:

[0089] Step A111, writing bypass output logic, and capturing bypass data from a source system of a data source layer according to the bypass output logic;

[0090] In addition to using CDC / OGG and other methods to parse database Binlog logs for real-time data collection, the quasi-real-time link can also use bypass output. Specifically, first write the bypass output logic, which is used to capture the bypass data generated in real time by the source system of the data source layer. The captured bypass data may include transaction records, user behavior logs, etc., and the bypass data will be used for subsequent data processing and analysis.

[0091] Step A112: forward the bypass data to the Kafka message middleware, and use the Flink cluster to consume the Kafka message middleware to obtain consumption data;

[0092] The captured bypass data is forwarded to the Kafka message middleware through network protocols (such as HTTP, TCP, etc.). As a high-performance message queue system, Kafka can efficiently handle a large number of concurrent data write and read requests. As a real-time data processing framework, the Flink cluster can consume messages in Kafka and process them in real time. After the Flink cluster consumes the data in Kafka, it converts it into consumption data that conforms to the internal data structure of Flink for subsequent processing and analysis. During the consumption process, Flink will perform parallel processing based on the characteristics and processing requirements of the data to improve data processing efficiency.

[0093] Step A113, aggregate the consumption data at different granularities through a custom operator to obtain quasi-real-time data.

[0094] Write custom Flink operators based on specific business needs and data characteristics. Custom operators can aggregate, filter, and transform consumption data to meet the needs of subsequent data analysis and applications. Use custom operators to aggregate consumption data. For example, aggregation processing may include aggregation by time window, grouping and aggregation by specific dimensions, etc. Through aggregation processing, consumption data can be converted into a more structured and easy-to-analyze data form. After aggregation processing, quasi-real-time data is obtained.

[0095] In this embodiment, bypass output of quasi-real-time data is adopted, and bypass output of quasi-real-time data can perform additional processing and analysis on the data without affecting the execution of the main business process. This method will not interfere with the normal flow of the business, and at the same time can flexibly obtain and process the data generated in real time. Through bypass output, data can be captured and forwarded to message middleware such as Kafka in real time, and then processed using streaming processing frameworks such as Flink. This method ensures the timeliness of the data and can respond to and analyze the data in a timely manner.

[0096] Specifically, in a feasible implementation manner, step S12 of collecting and generating offline data from the database of the source system using the data warehouse link may include steps S121 to S122:

[0097] Step S121, using a pre-configured data extraction component to extract data warehouse files from the database of the source system to outside the warehouse;

[0098] The "Extract" step in the ETL (Extract-Transform-Load) task is used to extract data from the data source. The database of the source system is the original source of data and contains the data that needs to be processed offline. The pre-configured data extraction component refers to the ETL component, which is responsible for reading data from the database of the source system. The target object of the extraction is the data warehouse file, which contains the data set in the source system that needs to be processed offline.

[0099] The extraction process is completed through the interaction between the ETL component and the source system database, including SQL query and database connection. After the extraction is completed, the data warehouse file is transferred to the external storage area of ​​the offline data warehouse for subsequent processing.

[0100] Step S122, the data warehouse file is stored in the warehouse to obtain offline data.

[0101] The process of loading the data warehouse files extracted and transferred outside the warehouse into the business data storage layer of the offline data warehouse. This process involves operations such as data format conversion, data cleaning, and data integration to ensure that the data meets the storage requirements and business needs of the offline data warehouse. After loading, the offline data can be used in the offline data warehouse for subsequent data analysis and application.

[0102] In this embodiment, by using the pre-configured ETL components, efficient and accurate data extraction can be achieved. The extraction process interacts closely with the database of the source system to ensure the integrity and accuracy of the data. Through the warehouse processing, the original data in the data warehouse file is converted into a data format that can be used by the offline data warehouse. Ensuring the integrity and accuracy of offline data provides a reliable data source for subsequent data analysis.

[0103] In this embodiment, a quasi-real-time link is used to collect and produce quasi-real-time data from the source system of the data source layer in real time; a data warehouse link is used to collect and produce offline data from the database of the source system. By using both quasi-real-time links and data warehouse links, the diversification of data sources is achieved. The quasi-real-time link provides the latest data stream, while the data warehouse link provides a highly reliable historical data set. The two complement each other and together constitute a rich and comprehensive data source, thereby solving the problem of a single data source.

[0104] Based on the first embodiment and / or the second embodiment of the present application, the third embodiment of the present application is proposed. In the third embodiment of the present application, the same or similar contents as those of the above embodiments can be referred to the above introduction, and will not be repeated in the following.

[0105] In this embodiment, before step S121 of extracting data warehouse files from the database of the source system to the outside of the warehouse using a pre-configured data extraction component, step S1 is also included:

[0106] Step S1, writing a configuration data extraction component through a configuration operation.

[0107] In the data warehouse link, in order to extract data from the database of the source system and produce offline data, additional data processing jobs need to be developed. These jobs are usually written in a configurable way so that they can flexibly adapt to different data sources, data formats, and data processing requirements.

[0108] The ETL (Extract-Transform-Load) component is used as the main tool for data extraction. The specific parameters and behaviors of the data extraction component are configured through configurable job writing. The configuration includes data source connection information, data extraction rules, data conversion logic, etc.

[0109] In this embodiment, by configuring the job writing method, the specific parameters and behaviors of the data extraction component can be flexibly configured, thereby realizing flexible adaptation to different data sources and data processing requirements. This method not only improves the efficiency and accuracy of data processing, but also reduces the complexity and maintenance cost of data processing.

[0110] Based on the above embodiments of the present application, a fourth embodiment of the present application is proposed. In the fourth embodiment of the present application, the same or similar contents as those of the above embodiments can be referred to the above introduction, and will not be described in detail later.

[0111] On this basis, reference Figure 3 , Figure 3 A flowchart of the fourth embodiment of the data processing method provided for the present application.

[0112] In this embodiment, the Flink cluster is used to consume the Kafka message middleware, and steps A1 to A3 are included before step S112 of obtaining consumption data:

[0113] Step A1, maintaining a data key value file through HBase, and reading whether there is a consumption record of the Kafka message middleware in the data key value file;

[0114] HBase is an open source non-relational distributed database developed based on Google's Bigtable model. It provides high reliability, high performance, column storage, scalability, and real-time read and write NoSQL database services. HBase is used as a storage system in this application to record the consumption status of Kafka messages.

[0115] The data key-value file actually refers to the data structure stored in HBase, which represents the unique identifier of the Kafka message and whether the message has been consumed. Before consuming Kafka messages, in order to prevent duplicate consumption, you need to query HBase to check whether the Kafka message to be processed currently has a consumption record.

[0116] Step A2: If the data key value file contains a consumption record of the Kafka message middleware, the consumption record is discarded;

[0117] If a consumption record corresponding to the current Kafka message is found in HBase, it means that the message has been processed before. In order to avoid repeated processing, it is decided not to process the message, that is, to "discard" the consumption task of the message. This saves resources and ensures data consistency.

[0118] Step A3: If the consumption record of the Kafka message middleware does not exist in the data key value file, write the consumption record.

[0119] If the consumption record corresponding to the current Kafka message is not found in HBase, it means that this is a new or unprocessed message, and then the data write operation is performed, that is, the record is written to the HBase database. This operation ensures the uniqueness and accuracy of the data and avoids duplicate data writing.

[0120] In this embodiment, HBase is used to record the consumption status of Kafka messages to avoid duplicate consumption. By checking the records in HBase, it is intelligently decided whether to process a Kafka message. HBase's high concurrent read and write and distributed column storage capabilities are used to maintain the global uniqueness of data and ensure efficient and accurate data processing. Even in the face of data replay or retransmission, data consistency can be maintained.

[0121] The overall processing flow of the data processing method of this application is described below in combination with the above embodiments. Figure 4 As shown, Figure 4 This is a schematic diagram of the Flink quasi-real-time link for this application to collect and produce quasi-real-time data in real time. The quasi-real-time processing link is mainly completed by Kafka and CDC collection and forwarding. Taking the statistics of product sales data at different granularities as an example, in the streaming processing stage, the Flink framework is first used to process the product transaction Kafka data generated by the quasi-real-time consumer source system, and a file with the unique key of the data is maintained through HBase. When consuming data, first read the file to see if the unique key value corresponding to the record is maintained. If not, write the data and consume the data, otherwise discard the record. The data is aggregated at different granularities through Flink custom operators, and the aggregated data is stored in the TIDB database through multiple forwarding processing operations. The TIDB database is an open source distributed relational database.

[0122] like Figure 5 As shown, Figure 5 This is a schematic diagram of the dual-link combined output data for this application. The quasi-real-time link is responsible for consuming quasi-real-time product entrustment data and performing corresponding data processing to complete the generation of T-day transaction statistics. The offline batch link, that is, the data warehouse link, is responsible for providing T-1 day transaction statistics that are less time-sensitive but can be used as a backup. The general process of the entire merged link is as follows:

[0123] First, the quasi-real-time data link consumes quasi-real-time Kafka data, processes and forwards it through Flink custom operators, produces T-day data of different dimensions, and loads the data into the application system database as the processing result data of T-day.

[0124] Furthermore, in the offline data link, it is necessary to develop offline jobs to process the sales statistics of day T-1 based on the offline data warehouse, and then load the data into the application system library through the ETL task, compare and correct it with the T-day data in the library to ensure the accuracy of the T-day data.

[0125] Finally, the business system only needs to perform simple queries to read the data, ensuring that the statistical results in the application database contain both historical data and data on day T.

[0126] This application also provides a data processing device, please refer to Figure 6 , the data processing device comprises:

[0127] The processing module 10 is used to collect and generate quasi-real-time data and offline data from the data source layer in a dual-link merging manner, and use the offline data as compensation to perform correction processing on the quasi-real-time data;

[0128] The result data storage module 20 is used to send the offline data and the corrected quasi-real-time data to the storage layer application system database for the application system to read and use.

[0129] The data processing device provided by the present application adopts the data processing method in the above embodiment to solve the technical problem of data processing. Compared with the prior art, the beneficial effects of the data processing device provided by the present application are the same as the beneficial effects of the data processing method provided by the above embodiment, and other technical features in the data processing device are the same as the features disclosed in the above embodiment method, which will not be described in detail here.

[0130] The present application provides a data processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method in the above-mentioned embodiment 1.

[0131] Reference below Figure 7 , which shows a schematic diagram of the structure of a data processing device suitable for implementing an embodiment of the present application. The data processing device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The data processing device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0132] like Figure 7 As shown, the data processing device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the data processing device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the data processing device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a data processing device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0133] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0134] The data processing device provided by the present application adopts the data processing method in the above embodiment to solve the technical problem of data processing. Compared with the prior art, the beneficial effects of the data processing device provided by the present application are the same as the beneficial effects of the data processing method provided by the above embodiment, and other technical features in the data processing device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0135] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0136] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0137] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, wherein the computer-readable program instructions are used to execute the data processing method in the above-mentioned embodiment.

[0138] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0139] The computer-readable storage medium may be included in a data processing device, or may exist independently without being incorporated into a data processing device.

[0140] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the data processing device, the data processing device enables the data processing device to: collect and produce quasi-real-time data and offline data from the data source layer in a dual-link merging manner; use the offline data as compensation to correct the quasi-real-time data; and send the offline data and the corrected quasi-real-time data to the storage layer application system database for the application system to read and use.

[0141] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0142] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0143] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0144] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned data processing method, and can solve the technical problems of data processing. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the data processing method provided in the above-mentioned embodiment, and will not be repeated here.

[0145] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned data processing method when executed by a processor.

[0146] The computer program product provided by this application can solve the technical problem of data processing. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as the beneficial effects of the data processing method provided by the above embodiment, which will not be repeated here.

[0147] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Collect and generate quasi-real-time data and offline data from the data source layer in a dual-link merging manner; Using the offline data as compensation, correcting the quasi-real-time data; The offline data and the corrected quasi-real-time data are sent to the storage layer application system database for the application system to read and use.

2. The method according to claim 1, characterized in that The dual link includes a quasi-real-time link and a data warehouse link. The step of collecting and generating quasi-real-time data and offline data from the data source layer in a dual-link merging manner includes: Use quasi-real-time links to collect and generate quasi-real-time data from the source system at the data source layer in real time; Use the data warehouse link to collect and generate offline data from the database of the source system.

3. The method according to claim 2, characterized in that The step of using the quasi-real-time link to collect and generate quasi-real-time data from the source system of the data source layer in real time includes: Use data change capture and synchronization to collect and parse log files from the source system, and forward the parsed log files to the Kafka message middleware; Use the Flink cluster to consume the Kafka message middleware to obtain consumption data; The consumption data is aggregated at different granularities through custom operators to obtain quasi-real-time data.

4. The method according to claim 2, characterized in that The step of using the quasi-real-time link to collect and generate quasi-real-time data from the source system of the data source layer in real time includes: Write bypass output logic and capture bypass data from the source system of the data source layer according to the bypass output logic; Forward the bypass data to the Kafka message middleware, and use the Flink cluster to consume the Kafka message middleware to obtain consumption data; The consumption data is aggregated at different granularities through custom operators to obtain quasi-real-time data.

5. The method according to claim 3, characterized in that The step of using the Flink cluster to consume the Kafka message middleware and obtaining consumption data also includes: Maintain the data key value file through HBase, and read whether there is a consumption record of the Kafka message middleware in the data key value file; If the data key-value file contains a consumption record of the Kafka message middleware, the consumption record is discarded; If the consumption record of the Kafka message middleware does not exist in the data key value file, the consumption record is written.

6. The method according to claim 2, characterized in that The step of using the data warehouse link to collect and generate offline data from the database of the source system includes: Extract data warehouse files from the database of the source system to outside the warehouse using a pre-configured data extraction component; The data warehouse file is processed in the warehouse to obtain offline data.

7. The method according to claim 6, characterized in that Before the step of extracting data warehouse files from the database of the source system to outside the warehouse using the pre-configured data extraction component, the following step is also included: Write configuration data extraction components through configuration jobs.

8. A data processing device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data processing method according to any one of claims 1 to 7.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.