A data processing method and system
By building a Flink-based data processing framework, the high availability and data real-time problems of HBase and ElasticSearch are solved, efficient data transfer and retrieval are achieved, and disk space utilization is optimized.
Patent Information
- Application Number
- CN202011565888.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-12-25
AI Technical Summary
In the prior art, the data synchronization calling system of HBase and ElasticSearch cannot guarantee high availability, data synchronization has delays and poor real-time performance, and the HBase-Scanner table scanner performance is poor, and data aggregation and grouping is difficult.
Build a big data processing framework, use the Flink distributed processing engine to receive data flow and enter detailed data into HBase, perform real-time summary and periodic query, generate and maintain the mapping relationship between HBase and ElasticSearch. ElasticSearch only stores retrieval information, and HBase provides detailed data support for result.
It realizes high-availability data transfer, reduces data transfer delay, improves data retrieval accuracy and query efficiency, and optimizes disk space utilization.
Smart Images

Figure CN114676160B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing and query, and particularly relates to a data processing method and system. Background Art
[0002] If you want to conveniently implement the full-text search function in the target system, currently the most preferred is the ElasticSearch full-text search engine (which can also be called a full-text search engine). The underlying layer of ElasticSearch is the open-source library Lucene (Lucene is an open-source full-text search engine toolkit, but it is not a complete full-text search engine. Instead, it is an architecture of a full-text search engine, providing a complete query engine, an index engine, and a partial text analysis engine).
[0003] The ElasticSearch full-text search engine is a real-time search and analysis engine that can provide multi-user, distributed and scalable full-text search.
[0004] During the application process of ElasticSearch, in the field of big data, it is often used in combination with databases such as HBase. The HBase database is a distributed NoSQL non-relational database. Usually, both ElasticSearch and HBase adopt the form of cluster application. Among them:
[0005] ElasticSearch is used to provide full-text search functions and is suitable for big data queries;
[0006] HBase is used to provide distributed storage, with excellent characteristics such as high reliability, high performance, column-oriented, and scalability. It is particularly suitable for storing big data (TB-level data).
[0007] The advantage of HBase is that it can achieve high-performance concurrent read and write operations. At the same time, HBase will automatically perform transparent segmentation on the data, so that the storage itself has horizontal scalability. Similarly, HBase also has disadvantages. It does not support conditional queries and only supports queries according to the RowKey (row primary key). It temporarily does not support the failover of the Master server. To solve the problem of fast retrieval, ElasticSearch is generally used to synchronize HBase data to solve the data query problem.
[0008] In the prior art, there are already relatively mature solutions for storage and query based on HBase and ElasticSearch, but these solutions still have some disadvantages.
[0009] Solution 1, such as Figure 1As shown, HBase (cluster) reads (receives) data from the data sender Kafka, then directly stores it in HBase. After that, the data synchronization call system queries the data stored in HBase, and finally the query results are stored in ElasticSearch.
[0010] The first solution has the following disadvantages:
[0011] 1. The data synchronization call system generally runs on a single machine and cannot guarantee high availability. When the data synchronization call system fails, ElasticSearch will not receive data;
[0012] 2. The data synchronization call system synchronizes data periodically, resulting in poor data real-time performance.
[0013] The second solution is as Figure 2 shown. HBase (cluster) reads (receives) data from the data sender Kafka, then directly stores it in HBase. After that, the HBase-Scanner table scanner is used to scan the data and transfer it to ElasticSearch.
[0014] Using the HBase-Scanner table scanner to scan HBase for data is usually a full-table scan.
[0015] The second solution has the following disadvantages:
[0016] 1. The performance of the HBase-Scanner table scanner is poor, and it takes a long time to scan a large amount of data;
[0017] 2. It is difficult to implement data aggregation and grouping.
[0018] The information disclosed in this background art section is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention
[0019] Aiming at the defects existing in the prior art, the purpose of the present invention is to provide a data processing method and system. By constructing a big data processing framework, it ensures high availability of data transfer from HBase to ElasticSearch, reduces the latency of data transfer to ElasticSearch, improves data accuracy, reduces redundant data, and optimizes disk space.
[0020] To achieve the above purpose, the present invention provides a data processing method applicable to a big data processing framework, the framework including: a processing engine, HBase, and ElasticSearch. The method includes the following steps:
[0021] The processing engine receives the data stream from the data sender, processes the data stream, and enters the detailed data included in the data stream into HBase;
[0022] The processing engine performs real-time summarization on the detailed data entered into HBase and periodically queries the data in HBase; HBase returns a data result set to the processing engine, and the data result set is HBase data retrieval information;
[0023] The processing engine stores the queried HBase data retrieval information in ElasticSearch; generates and maintains the mapping relationship between the HBase data retrieval information and the ElasticSearch data.
[0024] The present invention also provides a data processing system, which includes a processing engine, HBase, and ElasticSearch;
[0025] The processing engine is used to receive the data stream from the data sender, process the data stream, and enter the detailed data included in the data stream into HBase;
[0026] The processing engine is also used to perform real-time summarization on the detailed data entered into HBase and periodically query the data in HBase;
[0027] The HBase is used to return a data result set to the processing engine, and the data result set is HBase data retrieval information;
[0028] The processing engine is also used to store the queried HBase data retrieval information in ElasticSearch; generates and maintains the mapping relationship between the HBase data retrieval information and the ElasticSearch data.
[0029] The data processing method and system provided by the present invention have the following beneficial effects:
[0030] 1. Based on the Flink, HBase, and ElasticSearch clusters, a big data processing framework is constructed, which has high availability and can store and retrieve efficiently.
[0031] 2. ElasticSearch is used to store the aggregated data, reducing the latency of transferring data to ElasticSearch;
[0032] 3. The accuracy of data retrieval is improved, the data query efficiency is enhanced, and redundant data is reduced. Description of the Drawings
[0033] The present invention has the following drawings:
[0034] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0035] Figure 1 Flow chart of the existing storage and query solution 1 based on HBase and ElasticSearch.
[0036] Figure 2 Flow chart of the existing storage and query solution 2 based on HBase and ElasticSearch.
[0037] Figure 3 Schematic diagram of the system architecture of the data processing method described in the present invention.
[0038] Figure 4 Schematic of the content stored in HBase.
[0039] Figure 5 Schematic of the content stored in ElasticSearch. Detailed implementation manners
[0040] The present invention will be further described in detail below with reference to the accompanying drawings. The detailed description is made in connection with the exemplary embodiments of the present invention, including various details of the embodiments of the present invention to facilitate understanding, which should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted below.
[0041] As Figure 3 shown, the data processing method described in the present invention is applicable to a big data processing framework, and the framework includes: a processing engine, HBase, and ElasticSearch, and includes the following steps:
[0042] The processing engine receives the data stream from the data sending end, processes the data stream, and enters the detailed data included in the data stream into HBase;
[0043] The processing engine performs real-time summarization on the detailed data entered into HBase, and periodically queries the data in HBase; HBase returns a data result set to the processing engine, and the data result set is HBase data retrieval information;
[0044] The processing engine stores the queried HBase data retrieval information in ElasticSearch; generates and maintains the mapping relationship between the HBase data retrieval information and the ElasticSearch data.
[0045] In the present invention, by storing the HBase data retrieval information in ElasticSearch, ElasticSearch no longer stores the detailed data, but only a small amount of retrieval information, and HBase provides the data support for the result details. See Figure 4 , Figure 5 .
[0046] In the present invention, the retrieval function of ElasticSearch and the storage function of HBase are taken into account, and the data is not in a backup relationship, which reduces the data storage volume, saves disk usage, and improves the data writing performance at the same time. The summary information is stored in ElasticSearch to make up for the shortcoming that HBase cannot summarize data.
[0047] Based on the above technical solution, the processing engine is a Flink distributed processing engine.
[0048] Based on the above technical solution, the specific process of the processing engine for processing the data stream is that the processing engine is responsible for logically processing the data, and the logical processing includes but is not limited to: data aggregation and data reconstruction.
[0049] The processing engine enters the detailed data included in the data stream into HBase, including: the data stream at the data sending end is used as the data input. When the processing engine receives the data input, the detailed data included therein is first bucketed and then entered into HBase.
[0050] The default number of buckets for the bucket processing is an adjustable preset value.
[0051] The processing engine summarizes the detailed data entered into HBase to generate summary data, and periodically transfers the summary data to ElasticSearch.
[0052] Based on the above technical solution, the processing engine receives the data stream from the data sending end and stores it in HBase.
[0053] The data stream at the data sending end is used as the data input. When the processing engine receives the data input, the detailed data included therein is first bucketed and then entered into HBase.
[0054] The default number of buckets for the bucket processing is an adjustable preset value. Figure 4 In the illustrated embodiment, the number of buckets is 8, and the bucket identification ranges from 00 to 07.
[0055] Based on the above technical solution, when the HBase cluster forwards the data stream from the data sender and stores the data on disk in HBase, the row key (RowKey) of HBase adopts the following format: bucketing identifier + time granularity + time series + unique identifier; that is:
[0056] The row key (RowKey) of the HBase uses the time series, with the bucketing identifier as the prefix of the RowKey and a unique identifier as the suffix of the RowKey. The time granularity is further added between the time series and the prefix. The time granularity refers to the time unit for data aggregation, and the time granularity is also called the time grain.
[0057] Among them:
[0058] The bucketing identifier is used for multi-threaded parallel processing.
[0059] The value of the time series is obtained from the data time included in the detailed data, which is used for data sorting and for searching within a time range.
[0060] The time granularity matches the granularity in ElasticSearch. When querying through ElasticSearch, the value is mapped to HBase. The time granularity is dynamically generated and needs to be specifically defined according to the actual situation. The definition method does not belong to the content of this invention and will not be elaborated further.
[0061] The unique identifier is used to avoid data duplication.
[0062] Then: the format of the row key (RowKey) is: bucketing mark + time granularity + data time included in the detailed data + unique identifier; for example: row key (RowKey): 00_01020100_1578668912000_JOULGTKUXIOTX, where the bucketing mark is 00, the time granularity is 01020100, the data time is 1578668912000, and the unique identifier is JOULGTKUXIOTX; querying data in HBase through the RowKey is the basic query method designed for HBase.
[0063] Furthermore, the processing engine performs real-time calculation on the detailed data entered into HBase to obtain summary data. After completing the data for one time granularity, the detailed data in HBase is queried, and data cleaning can be selectively executed according to the setting. The purpose of executing data cleaning is to ensure the consistency and validity of the subsequent data.
[0064] Based on the above technical solution, the data sender is a messaging system using Kafka message queue or Mq message queue. Based on the message queue, the messages are differentiated to form streaming data containing topic information. The data sender outputs the streaming data containing topic information in the form of a data stream.
[0065] Based on the above technical solution, when returning the data result set to the query HBase request,
[0066] it is returned to the data sender of the query HBase request,
[0067] or returned to the specified data receiver in the query HBase request,
[0068] or returned to the specified data source in the query HBase request.
[0069] Based on the above technical solution, the HBase cluster responds to the query HBase data initiated by the processing engine.
[0070] Based on the above technical solution, as Figure 5 shown, the processing engine stores the HBase data retrieval information in ElasticSearch, which means: storing the HBase data retrieval information in the form of a data body in ElasticSearch, and the data body is summary data; and ElasticSearch stores the row primary key RowKey of HBase in the following format: time granularity + data summary time + summary data volume + data summary identifier;
[0071] The processing engine generates and maintains the mapping relationship between the HBase data retrieval information and the ElasticSearch data, which means: establishing the data association between the row primary key RowKey of HBase and ElasticSearch; for example:
[0072] The mapping relationship is: 01020100_1578668912000_10000_JOULGT4723HU2, where the time granularity is 01020100, the data summary time is 1578668912000, the summary data volume is 10000, and the data summary identifier is JOULGT4723HU2.
[0073] Based on the above technical solution, when data retrieval is required:
[0074] If the time granularity of the retrieved data is within the current summary period, the cached data in the processing engine can be directly queried;
[0075] If retrieving historical data, ElasticSearch data can be directly retrieved;
[0076] When larger time granularity data is needed, the summarization ability of ElasticSearch is utilized to directly retrieve ElasticSearch data;
[0077] If detailed data needs to be viewed, the detailed data in HBase can be queried according to the mapping relationship and the retrieved data, that is: according to the RowKey mapping information corresponding to ElasticSearch, assemble the starting value of the row primary key RowKey of HBase (all HBase bucket identifiers + time granularity + time series), and then query the corresponding number of data entries. This avoids full table scans, and multiple buckets can be retrieved in parallel.
[0078] The present invention also discloses a data processing system, the system includes a processing engine, HBase, and ElasticSearch; and is used to execute each step in the above method embodiments;
[0079] The processing engine is used to receive the data stream from the data sender, process the data stream, and enter the detailed data included in the data stream into HBase;
[0080] The processing engine is also used to perform real-time summarization on the detailed data entered into HBase, and periodically query the data in HBase;
[0081] HBase is used to return a data result set to the processing engine, and the data result set is the HBase data retrieval information;
[0082] The processing engine is also used to store the retrieved HBase data information in ElasticSearch; generate and maintain the mapping relationship between the HBase data retrieval information and the ElasticSearch data. The processing engine is a Flink distributed processing engine;
[0083] ElasticSearch is used to store the retrieval information of HBase data.
[0084] The system provided in this embodiment uses a Flink cluster as a data stream receiver. The Flink cluster can ensure the stability of data reception, that is, the high availability of the Flink cluster, and the entire architecture will not become unavailable due to the failure of the Flink master node. The HA method is implemented as follows: Flink integrates ZooKeeper to configure a standby master node, and when the master node fails, the standby node replaces the master node. When Flink reads data, it enables exactly-once to reduce redundant data, and uses asynchronous IO to write to ElasticSearch to reduce latency. The HA high availability is started. At the data sending end, the data is sent to the Flink cluster. The Flink cluster receives the data, stores the data in HBase, and after the HBase storage is successful, the data in HBase is queried for aggregation serialization and other business operations and then stored in ElasticSearch. This improves the high availability of data dumping, improves data accuracy, reduces redundant data to achieve the purpose of optimizing disk space, and at the same time reduces the latency of storing data in ElasticSearch.
[0085] The content not described in detail in this specification belongs to the prior art well known to those skilled in the art.
[0086] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modification or change made by those skilled in the art according to the content disclosed in the present invention shall be included in the protection scope recorded in the claims.
Claims
1. A data processing method, characterized in that, Applicable to a big data processing framework, the framework includes: a processing engine, HBase, and ElasticSearch. The processing engine is a Flink distributed processing engine. The method includes the following steps: The processing engine receives the data stream from the data sender, processes the data stream, and enters the detailed data contained in the data stream into HBase, including: taking the data stream from the data sender as data input. When the processing engine receives the data input, it first performs bucketing processing on the detailed data contained therein, and then enters it into HBase; The default number of buckets for the bucketing processing is an adjustable preset value; The processing engine summarizes the detailed data entered into HBase to generate summary data, and periodically transfers the summary data to ElasticSearch; When the HBase cluster responds to the data stream from the data sender forwarded by the processing engine and stores the data on disk in HBase, the row primary key RowKey of HBase adopts the following format: bucketing identifier + time granularity + time series + unique identifier; the time granularity matches the granularity in ElasticSearch. When querying through ElasticSearch, the value mapped to HBase, the time granularity is dynamically generated and specifically defined according to the actual situation; The processing engine performs real-time calculation on the detailed data entered into HBase to obtain summary data. After completing the data of one time granularity, it queries the detailed data in HBase and can choose to perform data cleaning according to the setting; The processing engine performs real-time summarization on the detailed data entered into HBase and periodically queries the data in HBase; HBase returns a data result set to the processing engine, and the data result set is HBase data retrieval information; The processing engine stores the HBase data retrieval information queried in ElasticSearch; generates and maintains the mapping relationship between the HBase data retrieval information and the ElasticSearch data; The processing engine stores the HBase data retrieval information in ElasticSearch, which means: storing the HBase data retrieval information in the form of a data body in ElasticSearch, and the data body is summary data; and ElasticSearch stores the row primary key RowKey of HBase in the following format: time granularity + data summarization time + summary data volume + data summarization identifier; The processing engine generates and maintains the mapping relationship between the HBase data retrieval information and the ElasticSearch data, which means: establishing the data association between the row primary key RowKey of HBase and the ElasticSearch data.
2. The method according to claim 1, characterized in that, The processing of the data stream by the processing engine includes but is not limited to: data aggregation, data reconstruction.
3. The method according to claim 1, characterized in that The data sending end is a message system using a Kafka message queue or an Mq message queue. Based on the message queue, messages are differentiated to form streaming data containing topic information. The data sending end outputs and sends the streaming data containing topic information to the processing engine in the form of a data stream.
4. The method according to claim 1, characterized in that, The method further includes: When a retrieval request is received and it is confirmed that detailed data is needed, according to the data retrieval information and mapping relationship stored in ElasticSearch, the detailed data in HBase is queried to complete the retrieval request.
5. A data processing system, characterized in that, The system includes a processing engine, HBase, and ElasticSearch; The processing engine is a Flink distributed processing engine. The processing engine is used to receive the data stream from the data sending end, process the data stream, and enter the detailed data contained in the data stream into HBase, including: taking the data stream of the data sending end as data input. When the processing engine receives the data input, the detailed data contained therein is first bucketed and then entered into HBase; The default number of buckets for the bucketing process is an adjustable preset value; The processing engine summarizes the detailed data entered into HBase to generate summary data, and periodically stores the summary data in ElasticSearch; When the HBase cluster responds to the data stream from the data sending end forwarded by the processing engine and stores the data on disk in HBase, the row primary key RowKey of HBase adopts the following format: bucket identifier + time granularity + time series + unique identifier; the time granularity matches the granularity in ElasticSearch. When querying through ElasticSearch, this value is mapped to HBase. The time granularity is dynamically generated and specifically defined according to the actual situation; The processing engine performs real-time calculation on the detailed data entered into HBase to obtain summary data. After completing the data for one time granularity, the detailed data in HBase is queried, and data cleaning can be selectively performed according to the setting; The processing engine is also used to perform real-time summarization on the detailed data entered into HBase and periodically query the data in HBase; HBase is used to return a data result set to the processing engine, and the data result set is HBase data retrieval information; The processing engine is also used to store the queried HBase data retrieval information in ElasticSearch; generate and maintain the mapping relationship between the HBase data retrieval information and the ElasticSearch data; When the processing engine stores the HBase data retrieval information in ElasticSearch, it means: storing the HBase data retrieval information in the form of a data body in ElasticSearch, and the data body is the summary data; and ElasticSearch stores the row primary key RowKey of HBase in the following format: time granularity + data summary time + summary data volume + data summary identifier; The processing engine generates and maintains the mapping relationship between HBase data retrieval information and ElasticSearch data, which means: establishing the data association between the row primary key RowKey of HBase and ElasticSearch data; ElasticSearch is used to store the retrieval information of HBase data.
Citation Information
Patent Citations
Data storage and data search method and device
CN110888839A
HBase index synchronization method and device, computer equipment and storage medium
CN110928954A
Sensitive data processing method and device, equipment and medium
CN112035531A