Data processing method, device and equipment and computer readable storage medium
By setting hot and cold data partitions in the data warehouse, the problem of inefficient data updates in Hive is solved, efficient data updates and queries are achieved, and resource waste and data redundancy are reduced.
Patent Information
- Application Number
- CN202510125033.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
In the context of big data, data update efficiency in Hive is low, especially when processing tens of billions of data. The existing technologies such as INSERT OVERWRITE lead to waste of resources and inefficiency.
By setting up a secondary partition of cold data partition and hot data partition under the first-level partition of the data warehouse, the cold data partition is used to store new transaction data added during the transaction occurrence period, and the hot data partition is used to store transaction data updated during the transaction occurrence period, thereby realizing the update of transaction data to store transaction data in different secondary partitions.
This method can greatly improve the efficiency of data updates, reduce the amount of data involved in the update operation, meet the requirements for update timeliness, and reduce data redundancy without affecting the efficiency of data query.
Smart Images

Figure CN120067121A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of big data technology, and particularly relates to a data processing method, apparatus, device, and computer-readable storage medium. Background Art
[0002] As a data warehouse tool, Hive allows mapping structured data files into a table and provides a Structured Query Language (SQL) query function. Since Hive was originally designed for data warehouse applications and the content of the data warehouse is usually read less and written more, Hive has inherent limitations in data updates. However, with the increasing demand for data processing, updating historical data in Hive has become an important requirement.
[0003] Currently, the historical data in Hive can be updated through the Hive INSERT OVERWRITE method. The INSERT OVERWRITE statement can regenerate the data of the entire table or entire partition based on the historical data and updated data, and replace the original data with the regenerated data.
[0004] However, in the context of big data, the number of data items involved in data updates may reach tens of billions, and the data volume is huge, resulting in low data update efficiency. Summary of the Invention
[0005] Embodiments of this application provide a data processing method, apparatus, device, computer-readable storage medium, and computer program product, which can improve data update efficiency.
[0006] In a first aspect, embodiments of this application provide a data processing method, which includes:
[0007] Obtain transaction data within a target period, where the transaction data includes first transaction data updated within the target period;
[0008] Based on the transaction occurrence period of the first transaction data, determine a target first-level partition corresponding to the first transaction data and a target hot data partition under the target first-level partition from multiple first-level partitions of a data warehouse; the first-level partitions are determined based on the transaction occurrence period, secondary partitions are set under the first-level partitions, the secondary partitions include a cold data partition and a hot data partition, the cold data partition is used to store newly added transaction data within the transaction occurrence period, and the hot data partition is used to store updated transaction data within the transaction occurrence period;
[0009] Use the first transaction data to update second transaction data in the target hot data partition under the target first-level partition.
[0010] In a possible implementation, updating the second transaction data in the target hot data partition under the target first-level partition by using the first transaction data includes:
[0011] When there is no target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data, storing the first transaction data into the target hot data partition.
[0012] In a possible implementation, updating the second transaction data in the target hot data partition under the target first-level partition by using the first transaction data includes:
[0013] When there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data, replacing the second transaction data corresponding to the target business primary key in the target hot data partition with the first transaction data.
[0014] In a possible implementation, before replacing the second transaction data corresponding to the target business primary key in the target hot data partition with the first transaction data when there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data, the method further includes:
[0015] Determining a first key-value pair with the transaction occurrence period and the business primary key in the first transaction data as the key and the data in the first transaction data other than the transaction occurrence period and the business primary key as the value;
[0016] Determining a second key-value pair with the transaction occurrence period and the business primary key in the second transaction data as the key and the data in the second transaction data other than the transaction occurrence period and the business primary key as the value;
[0017] When the keys of the first key-value pair and the second key-value pair are the same, determining that there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data.
[0018] In a possible implementation, the first transaction data and the second transaction data include a business timestamp, and the business timestamp represents the update time of the transaction data; replacing the second transaction data corresponding to the target business primary key in the target hot data partition with the first transaction data includes:
[0019] Determining a target key-value pair with a later update time of the transaction data in the first key-value pair and the second key-value pair according to the business timestamp;
[0020] Store the transaction data corresponding to the target key-value pair in the storage path where the target hot data partition is located;
[0021] In the storage path where the target hot data partition is located, use the transaction data corresponding to the target key-value pair to replace the second transaction data corresponding to the second key-value pair.
[0022] In a possible implementation, before storing the transaction data corresponding to the target key-value pair in the storage path where the target hot data partition is located, the method further includes:
[0023] Output the transaction data corresponding to the target key-value pair to the temporary directory under the target first-level partition to obtain the data file corresponding to the target key-value pair;
[0024] Back up the data file corresponding to the second transaction data;
[0025] The storing the transaction data corresponding to the target key-value pair in the storage path where the target hot data partition is located includes:
[0026] Move the data file corresponding to the target key-value pair to the storage path where the target hot data partition is located.
[0027] In a possible implementation, the transaction data further includes third transaction data newly added during the target period. After obtaining the transaction data during the target period, the method further includes:
[0028] Store the third transaction data in the cold data partition under the first-level partition corresponding to the target period.
[0029] In a possible implementation, the method further includes:
[0030] For the secondary partition, when the storage duration of the transaction data in the cold data partition reaches a preset duration, merge the transaction data in the cold data partition and the transaction data in the hot data partition.
[0031] In a possible implementation, the merging the transaction data in the cold data partition and the transaction data in the hot data partition includes:
[0032] For each fourth transaction data in the cold data partition, when there is a fifth transaction data with the same business primary key as the fourth transaction data stored in the hot data partition, use the fifth transaction data to replace the fourth transaction data.
[0033] In a possible implementation, the method further includes:
[0034] In response to a data query operation, for each of the first-level partitions, perform the following operations:
[0035] Obtain the sixth transaction data stored in the hot data partition and the seventh transaction data stored in the cold data partition under the first-level partition respectively;
[0036] Generate a query view table based on the sixth transaction data and the seventh transaction data;
[0037] In the case where the business primary keys of the sixth transaction data and the seventh transaction data are the same, perform deduplication processing on the transaction data in the query view table based on the business timestamps of the sixth transaction data and the seventh transaction data to obtain an updated query view table.
[0038] In a second aspect, an embodiment of the present application provides a data processing device, and the device includes:
[0039] An acquisition module, configured to acquire transaction data within a target period, where the transaction data includes first transaction data updated within the target period;
[0040] A determination module, configured to determine a target first-level partition corresponding to the first transaction data and a target hot data partition under the target first-level partition from multiple first-level partitions of a data warehouse based on the transaction occurrence period of the first transaction data; the first-level partitions are determined based on the transaction occurrence period, and second-level partitions are provided under the first-level partitions, and the second-level partitions include a cold data partition and a hot data partition, the cold data partition is used to store newly added transaction data within the transaction occurrence period, and the hot data partition is used to store updated transaction data within the transaction occurrence period;
[0041] An update module, configured to update second transaction data in the target hot data partition under the target first-level partition by using the first transaction data.
[0042] In a third aspect, an embodiment of the present application provides an electronic device, and the device includes: a processor and a memory storing computer program instructions;
[0043] When the processor executes the computer program instructions, the method in any possible implementation method in the first aspect above is implemented.
[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, and computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the method in any possible implementation method in the first aspect above is implemented.
[0045] In a fifth aspect, an embodiment of the present application provides a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the method in any possible implementation method in the first aspect described above.
[0046] In the embodiment of the present application, by setting a secondary partition including a cold data partition and a hot data partition under a primary partition of a data warehouse, where the cold data partition is used to store newly added transaction data within a transaction occurrence cycle, and the hot data partition is used to store updated transaction data within a transaction occurrence cycle, it is possible to store transaction data in different secondary partitions based on the update situation of the transaction data. In this way, when the transaction data within a target cycle includes first transaction data updated within the target cycle, first, based on the transaction occurrence cycle of the first transaction data, determine the target primary partition corresponding to the first transaction data and the target hot data partition under the target primary partition from multiple primary partitions of the data warehouse, and then use the first transaction data to update the second transaction data in the target hot data partition under the target primary partition, thereby achieving data update without updating all the data under the target primary partition, and the amount of data involved in the update operation is small, thus greatly improving the efficiency of data update. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0048] Figure 1 is a schematic flowchart of a data processing method provided by an embodiment of the present application;
[0049] Figure 2 is a schematic diagram of data splitting provided by an embodiment of the present application;
[0050] Figure 3 is a schematic diagram of data update provided by an embodiment of the present application;
[0051] Figure 4 is a schematic diagram of implementing data merging based on MapReduce provided by an embodiment of the present application;
[0052] Figure 5 is a schematic diagram of data query provided by an embodiment of the present application;
[0053] Figure 6 is a schematic diagram of data merging provided by an embodiment of the present application;
[0054] Figure 7 is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0055] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0056] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.
[0057] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article or device comprising the said elements.
[0058] In addition, the acquisition, storage, use, processing, etc. of data in the technical solution of the present application all comply with the relevant provisions of national laws and regulations.
[0059] As a data warehouse tool, Hive allows mapping structured data files into a table and provides a Structured Query Language (SQL) query function. Since Hive was originally designed for data warehouse applications and the content of the data warehouse is usually read less and written more, Hive has inherent limitations in data update. Updating historical data in Hive is a relatively complex task because Hive was originally designed for batch processing and analysis of big data, rather than frequent update operations like traditional relational databases. However, with the increasing demand for data processing, updating historical data in Hive has become an important requirement. In order to meet the requirements of business for the accuracy and timeliness of Hive data query, it is necessary to perform efficient data update on row-level historical data of a Hive partitioned table with a data volume of hundreds of billions.
[0060] In the project of building business continuity for the online service system, the transaction system will back up the data generated on the same day in the transaction record table to the big data platform in a batch manner at the end of each day. The big data platform splits and transforms the transaction data, converts the transaction data into a compressed format (rcfile), splits it by the transaction occurrence date, and stores it in different directories of the distributed file system (Hadoop Distributed System, HDFS). An external Hive table is created on this directory to meet the requirements of the business system to query and analyze the transaction table data by day using tools and frameworks such as Hive / Impala in the big data platform.
[0061] With the subsequent development of online business, when the business modifies the transaction status that occurred several days ago, it is necessary to synchronously modify the transaction status in the historical data of the big data platform. This requires first finding the original transaction record in the HDFS file and performing accurate and efficient updates, which is a huge challenge for a data table with a scale of hundreds of billions. Such updates also have the following characteristics:
[0062] (1) The data update time range has a large span. According to business requirements, the time span of transactions involved in updates is up to 1 year at most, so there are many historical partitions that need to be updated daily.
[0063] (2) The number of data updated daily in the historical partitions is much smaller than the number of data in the original partitions. The daily partitions store data in the order of billions, while the number of data that needs to be updated daily is generally less than 10,000.
[0064] (3) Since this data also needs to be shared by multiple tools or frameworks in the big data platform, such as Impala, etc. Therefore, this Hive table needs to be defined as an external table (External Table), and the UPDATE operation of Hive cannot be used.
[0065] (4) There are certain requirements for update timeliness. The big data platform also has subsequent tasks that depend on the transaction table data, and the data needs to be updated within a certain time limit (generally not exceeding 1 hour) to avoid affecting the execution of subsequent tasks.
[0066] Facing the requirement of updating the historical data of the Hive transaction table by row for data in the order of hundreds of billions, there is currently no particularly suitable and mature technical solution in the industry.
[0067] Currently, the mainstream solutions for updating Hive historical data mainly utilize the capabilities of Hive itself. For example, the update / merge using the Hive ACID (Atomicity, Consistency, Isolation, Durability) feature, the method using Apache Hudi, and the method using Hive INSERT OVERWRITE can be used to update historical data in Hive, enabling users to update historical data relatively easily, and it can be completed only using Hive SQL.
[0068] Among them, using the INSERT OVERWRITE statement can regenerate the data of the entire table or partition based on the historical data and the updated data, thus achieving the update effect. Usually, it involves creating a temporary table, merging the data to be updated with the new data, and then overwriting the original table. However, when using this method, it is necessary to regenerate the data files of the historical partitions involved in the update. In the case where there are many historical partitions involved in each data update, the amount of data generated by this data update is huge (when the daily data volume of the transaction table is at the billion level and the update data time range is 1 year, the number of data records involved in generating the data file is at the ten billion level), which does not meet the requirements for update timeliness and will waste a large amount of computing resources, resulting in waste of resources on the big data platform. Therefore, this data update method is relatively inefficient. This method is more suitable for data tables with a small total amount of data and a limited data update range.
[0069] Using the Apache Hudi method to update Hive data can support efficient incremental data writing and reading, and can only process the changed data instead of the full amount of data. This greatly improves the efficiency of data processing and query speed. However, the Apache Hudi method only supports internal tables.
[0070] Hive ACID allows performing row-level update (update) and delete (merge) operations in Hive tables, but it needs to use a specific file format (such as ORC) and storage attributes (such as transactional = true) to achieve, and the ACID feature only supports internal tables.
[0071] In the scenario of multi-party sharing of a large amount of transaction data, in order for multiple users or multiple systems to share the data of the big data platform, the data files need to be stored in the specified directory of HDFS, which is convenient for other tools and applications to access and share, without frequent data migration and data redundancy. And when Hive enables ACID support, maintaining the metadata of transactions during data insertion will increase additional storage and performance overheads, and additional processing is required during the execution of update operations, which may lead to a decline in query performance. To sum up, in order to ensure the efficiency of data storage and query of such Hive tables in the scenario of multi-party sharing of a large amount of transaction data, the Hive table needs to be established as an external table, and the update / merge operations of Hive cannot be used to update the data.
[0072] Thus, to solve the problems of the prior art, the embodiments of the present application provide a data processing method, device, equipment, computer-readable storage medium, and computer program product. Among them, the data processing method can be applied to the scenario of storing transaction data as an external table of a data warehouse (such as Hive). The data processing method can be specifically applied to scenarios such as storing new data in the data warehouse, updating historical data in the data warehouse, and querying data in the data warehouse.
[0073] First, the data warehouse in the embodiments of the present application will be introduced below.
[0074] The data warehouse in the embodiments of the present application may include multiple first-level partitions. Among them, the first-level partitions can be determined based on the transaction occurrence cycle. The transaction occurrence cycle can be the statistical cycle of transaction data. The transaction occurrence cycle can be any one of days, weeks, months, quarters, years, etc. If the transaction occurrence cycle is days, the first-level partitions can be determined based on the transaction occurrence date.
[0075] In addition, second-level partitions can be set under each first-level partition. The second-level partitions can include cold data partitions and hot data partitions. Among them, the cold data partitions can be used to store the newly added transaction data within the transaction occurrence cycle, and the hot data partitions can be used to store the updated transaction data within the transaction occurrence cycle.
[0076] Next, the data processing method provided by the embodiments of the present application will be introduced.
[0077] Figure 1 The flowchart of a data processing method provided by the embodiments of the present application is shown. This data processing method can be executed by a big data platform. As Figure 1 shown, the data processing method provided by the embodiments of the present application includes the following steps:
[0078] S110. Obtain the transaction data within the target cycle, where the transaction data includes the first transaction data updated within the target cycle;
[0079] S120. Determine the target first-level partition corresponding to the first transaction data and the target hot data partition under the target first-level partition from multiple first-level partitions of the data warehouse based on the transaction occurrence cycle of the first transaction data;
[0080] S130. Update the second transaction data in the target hot data partition under the target first-level partition by using the first transaction data.
[0081] In the embodiment of the present application, by setting a second-level partition including a cold data partition and a hot data partition under the first-level partition of the data warehouse, and the cold data partition is used to store the newly added transaction data within the transaction occurrence cycle, and the hot data partition is used to store the updated transaction data within the transaction occurrence cycle, it is possible to store the transaction data in different second-level partitions based on the update situation of the transaction data. In this way, when the transaction data within the target cycle includes the first transaction data updated within the target cycle, first, based on the transaction occurrence cycle of the first transaction data, determine the target first-level partition corresponding to the first transaction data and the target hot data partition under the target first-level partition from multiple first-level partitions of the data warehouse, and then update the second transaction data in the target hot data partition under the target first-level partition by using the first transaction data, so as to achieve data update without updating all the data under the target first-level partition, and the amount of data involved in the update operation is small, thereby greatly improving the efficiency of data update.
[0082] The specific implementation manners of the above steps are introduced below.
[0083] In some embodiments, in S110, the target cycle may be a transaction occurrence cycle including the transaction occurrence date. For example, if the transaction occurrence cycle is in days, the target cycle may be the (T - 1)th day. If the transaction occurrence cycle is in months, the target cycle may be the (T - 1)th month.
[0084] The transaction data may include the third transaction data newly added within the target cycle and the first transaction data updated within the target cycle. For example, if three transaction data A are newly added on the (T - 1)th day and one transaction data B newly added on the (T - 2)th day is modified, then the transaction data A may be the third transaction data, and the transaction data B may be the first transaction data.
[0085] As an example, as Figure 2 shown, if the transaction occurrence cycle is in days, the transaction system can batch-backup the transaction data newly added or updated on the (T - 1)th day to the big data platform HDFS at the end of the day. The big data platform can first perform data cleaning and transformation on the original file (i.e., the transaction data newly added or updated on the (T - 1)th day), and then split the transaction data according to the transaction occurrence date to obtain the third transaction data newly added on the (T - 1)th day and the first transaction data updated on the (T - 1)th day but with different transaction occurrence dates.
[0086] Based on this, in order to ensure that complete transaction information can be queried during subsequent data queries, in some embodiments, after the above S110, the method may further include:
[0087] Storing the third transaction data into the cold data partition under the first-level partition corresponding to the target period.
[0088] Here, the target period may be the transaction occurrence date of the third transaction data. By storing the third transaction data into the cold data partition under the first-level partition corresponding to the target period, the full amount of new transaction information within the target period can be retained to ensure that complete transaction information can be queried during subsequent data queries.
[0089] In some embodiments, in S120, the target first-level partition may be the first-level partition corresponding to the first transaction data. The hot data partition under the target first-level partition may be the target hot data partition. The first transaction data with different transaction occurrence dates may correspond to different target first-level partitions.
[0090] As an example, multiple first-level partitions in the data warehouse may correspond to different storage paths. After splitting the transaction data into the third transaction data and at least one first transaction data according to the transaction occurrence period, the storage paths of the first-level partitions corresponding to the third transaction data and at least one first transaction data can be determined first, and then the third transaction data and at least one first transaction data can be stored into the first-level partitions respectively. If the transaction occurrence period is the transaction occurrence date, the third transaction data and at least one first transaction data can be stored into the directories of different dates in HDFS respectively. Specifically, the third transaction data can be stored into the cold data partition under its corresponding date directory, and the first transaction data can be stored into the hot data partition under its corresponding date directory.
[0091] In some embodiments, in S130, the transaction data already stored in the target hot data partition may be the second transaction data. The second transaction data may be the transaction data that has been updated. For each target first-level partition, there may be no data stored in its target hot data partition, or there may already be second transaction data stored, which is not limited here. If there is no data stored in the target hot partition, it can be determined that the transaction data stored in the target first-level partition has not been updated. If there is second transaction data stored in the target hot data partition, the business primary key of the second transaction data may be the same as or different from the business primary key of the first transaction data. If the business primary key of the second transaction data is the same as the business primary key of the first transaction data, it means that the first transaction data is not the first update.
[0092] Based on this, in order to improve the data update efficiency, in some embodiments, the above S130 may specifically include:
[0093] In the case that there is no target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data, store the first transaction data in the target hot data partition.
[0094] Here, for each target first-level partition, if there is no target business primary key in the target hot partition that is the same as the business primary key of the first transaction data, the first transaction data can be directly stored in the target hot data partition to achieve rapid update of transaction data, thereby improving data update efficiency.
[0095] It should be noted that in the embodiments of the present application, the principle of achieving rapid update of transaction data by storing the first transaction data in the target hot data partition is as follows: when performing data query later, instead of separately accessing the transaction data in the cold data partition or the hot data partition, a query view is first established based on the transaction data in the hot data partition and the cold data partition, and then the transaction data is de-duplicated based on the business timestamp of the transaction data. When there are the same business primary keys in the view, the transaction data corresponding to the business primary key with a newer business timestamp is retained, so as to query the complete and latest transaction data. Among them, the business timestamp can represent the update time of the transaction data.
[0096] In addition, to improve data update efficiency, in some embodiments, the above S130 may specifically include:
[0097] In the case that there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data, use the first transaction data to replace the second transaction data corresponding to the target business primary key in the target hot data partition.
[0098] Here, if the business primary key of the second transaction data is the same as the business primary key of the first transaction data, this business primary key can be determined as the target business primary key. If the business primary key of the second transaction data is the same as the business primary key of the first transaction data, it means that the first transaction data is not updated for the first time. By using the first transaction data to replace the second transaction data corresponding to the target business primary key in the target hot data partition, the second transaction data can be updated again to achieve rapid update of transaction data, thereby improving data update efficiency.
[0099] To better understand the above process, based on the above embodiments, a specific example is given.
[0100] For example, a schematic diagram of data update provided by the embodiments of the present application can be as Figure 3 shown. In Figure 3In it, the transaction data in tbl_1 can be stored in multiple first-level partitions respectively according to the transaction occurrence date, and a hot data partition (i.e., deltadata partition) and a cold data partition (i.e., maindata partition) can be set under each first-level partition. Additionally, the data update can specifically be to merge the first transaction data and the second transaction data in the corresponding target hot data partition thereof.
[0101] Based on Figure 3 , a more specific example is given. Suppose the transaction data newly added or updated on T-1 day includes the third transaction data A newly added on T-1 day, the first transaction data B, the first transaction data C, and the first transaction data D updated on T-1 day, and the transaction occurrence date of the first transaction data B is T-2 day, and the transaction occurrence dates of the first transaction data C and the first transaction data D are T-3 day. Then, the third transaction data A can be stored in the cold data partition under the first-level partition corresponding to T-1 day, and the target first-level partition and the target hot data partition under the target first-level partition corresponding to each of the first transaction data B, the first transaction data C, and the first transaction data D can be determined. Among them, the target first-level partition corresponding to the first transaction data B can be the first-level partition corresponding to T-2 day. If no data is stored in the hot data partition under the first-level partition corresponding to T-2 day, the first transaction data B can be directly stored in the hot data partition under the first-level partition corresponding to T-2 day. Additionally, the target first-level partitions corresponding to the first transaction data C and the first transaction data D can be the same, both being the first-level partition corresponding to T-3 day. If the second transaction data with business primary key 01 and the second transaction data with business primary key 02 are stored in the hot data partition under the first-level partition corresponding to T-3 day, and the business primary key of the first transaction data C is 01 and the business primary key of the first transaction data D is 03, then the first transaction data C can be used to replace the second transaction data with business primary key 01, and the first transaction data D can be stored in the hot data partition under the first-level partition corresponding to T-3 day.
[0102] Based on this, in order to accurately determine whether there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data, in some embodiments, when there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data, before using the first transaction data to replace the second transaction data corresponding to the target business primary key in the target hot data partition, the method may further include:
[0103] Determine a first key-value pair with the transaction occurrence period and the business primary key in the first transaction data as the key and the data in the first transaction data other than the transaction occurrence period and the business primary key as the value;
[0104] Using the transaction occurrence period and business primary key in the second transaction data as keys, and the data in the second transaction data other than the transaction occurrence period and business primary key as values, determine the second key-value pair;
[0105] When the keys of the first key-value pair and the second key-value pair are the same, determine that there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data.
[0106] Here, data merging can be specifically implemented based on a distributed computing framework (MapReduce). The data merging process based on MapReduce can include a Map phase. The Map phase can process the input key-value pairs based on the map function to generate intermediate result key-value pairs.
[0107] Specifically, during the data merging process, first, all directories of the data files to be updated and the hot partition directories in the involved historical partitions can be read simultaneously based on the Map task. Among them, reading the directory of the data files to be updated can obtain the first transaction data, and reading the hot partition directories in the historical partitions involved in the first transaction data can obtain the second transaction data. After obtaining the first transaction data, the business primary key and the transaction occurrence period can be determined in the first transaction data first, and then the business primary key and the transaction occurrence period can be concatenated with "#" to obtain the key (key) of the intermediate result. The value (value) of the intermediate result can be the data in the first transaction data other than the transaction occurrence period and the business primary key. For example, if the transaction occurrence period of the first transaction data is 20240101 and the business primary key is 01, the key of the intermediate result can be recorded as 20240101#01. If the value of the intermediate result is recorded as value1, the first key-value pair can be "20240101#01, value1". The process of determining the second key-value pair based on the second transaction data is the same and will not be elaborated here in detail.
[0108] Since the first key-value pair is determined based on the business primary key and the transaction occurrence period of the first transaction data, the second key-value pair is determined based on the business primary key and the transaction occurrence period of the second transaction data, and the occurrence periods of the first transaction data and the second transaction data are the same. Therefore, if the keys of the first key-value pair and the second key-value pair are the same, it can be determined that there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data.
[0109] By judging whether the keys of the first key-value pair and the second key-value pair are the same in the embodiments of the present application, it can be accurately determined whether there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data.
[0110] In addition, the first transaction data and the second transaction data may include a business timestamp. Among them, the business timestamp may represent the update time of the transaction data. It can be understood that the newer the time of the business timestamp, the later the update time of the transaction data.
[0111] Based on this, in order to improve the data update efficiency, in some embodiments, the above-mentioned use of the first transaction data to replace the second transaction data corresponding to the target business primary key in the target hot data partition may specifically include:
[0112] According to the business timestamp, determine the target key-value pair with a later update time of the transaction data among the first key-value pair and the second key-value pair;
[0113] Store the transaction data corresponding to the target key-value pair under the storage path where the target hot data partition is located;
[0114] Under the storage path where the target hot data partition is located, use the transaction data corresponding to the target key-value pair to replace the second transaction data corresponding to the second key-value pair.
[0115] Here, the data merging process implemented based on MapReduce may further include a shuffle stage and a reduce stage. Among them, the shuffle stage can shuffle the first key-value pair and the second key-value pair with the same key together. In the reduce stage, for the first key-value pair and the second key-value pair with the same key, obtain and compare the business timestamps from the values respectively, and retain one record with a newer business timestamp (i.e., the target key-value pair) when outputting in the reduce. The target key-value pair is usually the first key-value pair.
[0116] In this way, by storing the transaction data corresponding to the target key-value pair under the storage path where the target hot data partition is located, and under the storage path where the target hot data partition is located, using the transaction data corresponding to the target key-value pair to replace the second transaction data corresponding to the second key-value pair, the data update in the target hot data partition can be realized, and the data update efficiency can be improved.
[0117] Based on this, in order to ensure data security during the data update process, in some embodiments, before storing the transaction data corresponding to the target key-value pair under the storage path where the target hot data partition is located, the method may further include:
[0118] Output the transaction data corresponding to the target key-value pair to a temporary directory under the target first-level partition to obtain a data file corresponding to the target key-value pair;
[0119] Back up the data file corresponding to the second transaction data;
[0120] Storing the transaction data corresponding to the target key-value pair under the storage path where the target hot data partition is located includes:
[0121] Move the data file corresponding to the target key-value pair to the storage path where the target hot data partition is located.
[0122] Here, after determining the target key-value pair, the transaction data corresponding to the target key-value pair can be first saved as a data file and temporarily stored in the temporary directory under its corresponding target first-level partition. Then, back up the data file (i.e., the data file corresponding to the second transaction data) in the target hot data partition path under the target first-level partition corresponding to the target key-value pair. Finally, move the data file generated by the MapReduce process (i.e., the data file corresponding to the target key-value pair) to the corresponding hot data partition path, thereby realizing replacing the second transaction data corresponding to the second key-value pair with the transaction data corresponding to the target key-value pair under the storage path where the target hot data partition is located.
[0123] In the embodiment of the present application, if there is a problem in the data update process (i.e., replacing the second transaction data corresponding to the second key-value pair with the transaction data corresponding to the target key-value pair), or in the updated data, the historical data can be restored by restoring the backed-up data file, which can ensure data security during the data update process.
[0124] A schematic diagram of data merging implemented based on MapReduce provided by the embodiment of the present application can be as Figure 4 shown.
[0125] In addition, as described above, in the embodiment of the present application, the cold data partition is used to store the newly added transaction data during the transaction occurrence period, and the hot data partition is used to store the updated transaction data during the transaction occurrence period. That is, for each first-level partition, there will necessarily be transaction data with the same business primary key in the cold data partition as in the hot data partition. If the first-level partition is determined based on the transaction occurrence date, data with the same business primary key on the same day may exist in both the cold and hot partitions.
[0126] Therefore, in order to shield the impact on the use of querying data by external systems and ensure the accuracy of data querying, in some embodiments, the method may further include:
[0127] In response to a data query operation, for each first-level partition, perform the following operations:
[0128] Respectively obtain the sixth transaction data stored in the hot data partition under the first-level partition and the seventh transaction data stored in the cold data partition;
[0129] Generate a query view table based on the sixth transaction data and the seventh transaction data;
[0130] When the business primary keys of the sixth transaction data and the seventh transaction data are the same, deduplicate the transaction data in the query view table based on the business timestamps of the sixth transaction data and the seventh transaction data to obtain an updated query view table.
[0131] Here, when the external system queries Hive, it does not directly access the original data table. Instead, it first obtains the sixth transaction data stored in the hot data partition and the seventh transaction data stored in the cold data partition under the first-level partition respectively, and generates a query view table (view) based on the sixth transaction data and the seventh transaction data. Then, it accesses through the view established in Hive. After establishing the view, the view can use the ROW_NUMBER() OVER statement to deduplicate the transaction records with the same primary key in the cold data partition and the hot data partition according to the business timestamp to shield the impact on external use, and can solve the problem that the data of the same day may be repeated due to the storage of data in the secondary partition, ensuring the accuracy of data query.
[0132] A schematic diagram of data query provided by an embodiment of the present application can be as Figure 5 shown.
[0133] In addition, as described above, in the embodiment of the present application, the cold data partition is used to store the newly added transaction data during the transaction occurrence period, and the hot data partition is used to store the updated transaction data during the transaction occurrence period. That is, there will inevitably be transaction data with the same primary key in the cold data partition as in the hot data partition. Therefore, there will be redundancy in the transaction data stored in the cold data partition and the hot data partition. In actual situations, the data update of transaction data usually occurs within a certain time span (such as 1 year). Beyond this time span, data updates usually will not be performed anymore.
[0134] Therefore, in order to reduce data redundancy and improve data query efficiency without affecting data update efficiency, in some embodiments, the method may further include:
[0135] For the secondary partition, when the storage duration of the transaction data in the cold data partition reaches the preset duration, merge the transaction data in the cold data partition and the transaction data in the hot data partition.
[0136] Here, the preset duration can be the preset storage duration of the transaction data. The preset duration can be, for example, 1 year. For each secondary partition under the first-level partition, if the storage duration of the transaction data in the cold data partition reaches the preset duration, it can be determined that the transaction data has exceeded the maximum time range required for business updates, that is, the data in this partition will no longer be updated. Therefore, during the low period of the cluster business, a hot and cold data merge task can be initiated to merge the transaction data in the hot data partition and the cold data partition.
[0137] In some embodiments, the merging of the transaction data in the cold data partition and the transaction data in the hot data partition may specifically include:
[0138] For each fourth transaction data in the cold data partition, when there is fifth transaction data with the same business primary key as the fourth transaction data stored in the hot data partition, the fourth transaction data is replaced with the fifth transaction data.
[0139] The data merging process in the embodiments of the present application can be implemented based on MapReduce. The specific data merging process can refer to the above description and will not be elaborated here in detail.
[0140] A schematic diagram of data merging provided by the embodiments of the present application can be as Figure 6 shown.
[0141] In the embodiments of the present application, by merging the transaction data in the cold data partition and the transaction data in the hot data partition when the storage duration of the transaction data in the cold data partition reaches a preset duration, data redundancy can be reduced and data query efficiency can be improved without affecting the data update efficiency.
[0142] In summary, to solve the problem of efficient historical data row update in the scenario of multi-party sharing of a large amount of transaction data in Hive, a Hive historical data update / query method based on MapReduce is proposed, which has the following technical effects:
[0143] (1) By directly updating the data storage files on HDFS using MapReduce, without relying on the support of Hive ACID, the access to the data table is not affected during data update. The Hive table can be established as an external table, which can be conveniently accessed and shared by other tools and applications.
[0144] (2) Meeting the data security requirements: Before data update, the historical data files are backed up. If the update needs to be rolled back to the previous data version, only the data file needs to be rolled back to the previous version (i.e., the backup version) to restore the data.
[0145] (3) Meeting the requirements of data update efficiency: By setting up secondary partitions to distinguish cold data and hot data, each update only merges the data files within the hot data partition and regenerates the data files, which can greatly improve the data update efficiency and ensure that the update of a data table with tens of billions of records can be completed within 1 hour.
[0146] Based on the data processing method provided in the above embodiments, correspondingly, the present application also provides a specific implementation manner of a data processing device. Please refer to the following embodiments.
[0147] As Figure 7As shown in the figure, the data processing device 700 provided by the embodiment of the present application includes the following modules:
[0148] An acquisition module 710, configured to acquire transaction data within a target period, where the transaction data includes first transaction data updated within the target period;
[0149] A determination module 720, configured to determine a target first-level partition corresponding to the first transaction data and a target hot data partition under the target first-level partition from multiple first-level partitions of a data warehouse based on the transaction occurrence period of the first transaction data; the first-level partitions are determined based on the transaction occurrence period, and second-level partitions are set under the first-level partitions, and the second-level partitions include a cold data partition and a hot data partition, the cold data partition is used to store newly added transaction data within the transaction occurrence period, and the hot data partition is used to store updated transaction data within the transaction occurrence period;
[0150] An update module 730, configured to update second transaction data in the target hot data partition under the target first-level partition by using the first transaction data.
[0151] The data processing device 700 is described in detail below, as follows:
[0152] In some embodiments, the update module 730 may specifically include:
[0153] A storage sub-module, configured to store the first transaction data in the target hot data partition when there is no target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data.
[0154] In some embodiments, the update module 730 may specifically include:
[0155] A first replacement sub-module, configured to replace the second transaction data corresponding to the target business primary key in the target hot data partition by using the first transaction data when there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data.
[0156] In some embodiments, the update module 730 may further specifically include:
[0157] A first determination sub-module, configured to determine a first key-value pair with the transaction occurrence period and business primary key in the first transaction data as keys and the data other than the transaction occurrence period and business primary key in the first transaction data as values before replacing the second transaction data corresponding to the target business primary key in the target hot data partition by using the first transaction data when there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data;
[0158] A second determination sub-module, configured to determine a second key-value pair by using the transaction occurrence period and the business primary key in the second transaction data as keys, and using the data in the second transaction data other than the transaction occurrence period and the business primary key as values;
[0159] A third determination sub-module, configured to determine that there is a target business primary key in the target hot data partition under the target first-level partition that is the same as the business primary key of the first transaction data when the keys of the first key-value pair and the second key-value pair are the same.
[0160] In some embodiments, the first transaction data and the second transaction data include a business timestamp, and the business timestamp represents the update time of the transaction data. Based on this, the first replacement sub-module may specifically include:
[0161] A determination unit, configured to determine a target key-value pair with a later update time of the transaction data in the first key-value pair and the second key-value pair according to the business timestamp;
[0162] A storage unit, configured to store the transaction data corresponding to the target key-value pair under the storage path where the target hot data partition is located;
[0163] A replacement unit, configured to replace the second transaction data corresponding to the second key-value pair with the transaction data corresponding to the target key-value pair under the storage path where the target hot data partition is located.
[0164] In some embodiments, the first replacement sub-module may specifically further include:
[0165] An output unit, configured to output the transaction data corresponding to the target key-value pair to a temporary directory under the target first-level partition to obtain a data file corresponding to the target key-value pair before storing the transaction data corresponding to the target key-value pair under the storage path where the target hot data partition is located;
[0166] A backup unit, configured to back up the data file corresponding to the second transaction data.
[0167] Based on this, the storage unit may specifically include:
[0168] A moving sub-unit, configured to move the data file corresponding to the target key-value pair to the storage path where the target hot data partition is located.
[0169] In some embodiments, the transaction data further includes third transaction data newly added within a target period. Based on this, the data processing device 700 may further include:
[0170] A storage module, configured to store the third transaction data in a cold data partition under the first-level partition corresponding to the target period after obtaining the transaction data within the target period.
[0171] In some of these embodiments, the data processing device 700 may further include:
[0172] A merging module, configured to merge the transaction data in the cold data partition and the transaction data in the hot data partition for a secondary partition when the storage duration of the transaction data in the cold data partition reaches a preset duration.
[0173] In some of these embodiments, the merging module may specifically include:
[0174] A second replacement sub-module, configured to replace each fourth transaction data in the cold data partition with a fifth transaction data when the fifth transaction data with the same business primary key as the fourth transaction data is stored in the hot data partition.
[0175] In some of these embodiments, the data processing device 700 may further include:
[0176] An execution module, configured to, in response to a data query operation, perform the following operations for each first-level partition:
[0177] Respectively obtain the sixth transaction data stored in the hot data partition and the seventh transaction data stored in the cold data partition under the first-level partition;
[0178] Generate a query view table based on the sixth transaction data and the seventh transaction data;
[0179] When the business primary keys of the sixth transaction data and the seventh transaction data are the same, perform deduplication processing on the transaction data in the query view table based on the business timestamps of the sixth transaction data and the seventh transaction data to obtain an updated query view table.
[0180] In the embodiments of the present application, by setting a secondary partition including a cold data partition and a hot data partition under the first-level partition of the data warehouse, where the cold data partition is used to store the newly added transaction data during the transaction occurrence period, and the hot data partition is used to store the updated transaction data during the transaction occurrence period, it is possible to store the transaction data in different secondary partitions based on the update situation of the transaction data. Thus, when the transaction data in the target period includes the first transaction data updated in the target period, first determine the target first-level partition corresponding to the first transaction data and the target hot data partition under the target first-level partition from multiple first-level partitions of the data warehouse based on the transaction occurrence period of the first transaction data, and then use the first transaction data to update the second transaction data in the target hot data partition under the target first-level partition, so as to achieve data update without updating all the data under the target first-level partition, and the amount of data involved in the update operation is small, thereby greatly improving the efficiency of data update.
[0181] Based on the data processing method provided in the above embodiments, the embodiments of the present application also provide a specific implementation manner of an electronic device. Figure 8 FIG. 800 of an electronic device provided by an embodiment of the present application is shown.
[0182] The electronic device 800 may include a processor 810 and a memory 820 storing computer program instructions.
[0183] Specifically, the above-mentioned processor 810 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0184] The memory 820 may include a mass storage for data or instructions. By way of example and not limitation, the memory 820 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In a suitable case, the memory 820 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 820 may be internal or external to the electronic device 800. In a specific embodiment, the memory 820 is a non-volatile solid-state memory.
[0185] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.
[0186] The processor 810 reads and executes the computer program instructions stored in the memory 820 to implement any one of the data processing methods in the above embodiments.
[0187] In one example, the electronic device 800 may further include a communication interface 830 and a bus 840. Among them, as Figure 8 shown, the processor 810, the memory 820, and the communication interface 830 are connected through the bus 840 and complete communication with each other.
[0188] The communication interface 830 is mainly used to implement the communication between various modules, devices, units, and / or equipment in the embodiments of the present application.
[0189] The bus 840 includes hardware, software, or both, and couples the components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In suitable cases, the bus 840 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0190] Exemplarily, the electronic device 800 may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc.
[0191] The electronic device can execute the data processing method in the embodiments of the present application, so as to implement the combination Figures 1 to 7 the described data processing method and device.
[0192] In addition, in combination with the data processing method in the above embodiments, the embodiments of the present application may provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the data processing methods in the above embodiments is implemented.
[0193] In combination with the data processing method in the above embodiments, the embodiments of the present application may provide a computer program product to implement. When the instructions in the computer program product are executed by the processor of the electronic device, any one of the data processing methods in the above embodiments is implemented.
[0194] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0195] The functional blocks shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0196] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0197] The various aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to generate a machine such that the instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It can also be understood that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0198] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application.
Claims
1. A data processing method, characterized in that: include: Acquire transaction data within a target period, the transaction data including first transaction data updated within the target period; Based on the transaction occurrence cycle of the first transaction data, determine a target first-level partition corresponding to the first transaction data and a target hot data partition under the target first-level partition from multiple first-level partitions of the data warehouse; The first-level partition is determined based on a transaction cycle, and a second-level partition is set under the first-level partition. The second-level partition includes a cold data partition and a hot data partition. The cold data partition is used to store transaction data newly added during the transaction cycle, and the hot data partition is used to store transaction data updated during the transaction cycle. The first transaction data is used to update the second transaction data in the target hot data partition under the target first-level partition.
2. The method according to claim 1, characterized in that The updating of the second transaction data in the target hot data partition under the target first-level partition by using the first transaction data includes: When there is no target business primary key identical to the business primary key of the first transaction data in the target hot data partition under the target first-level partition, the first transaction data is stored in the target hot data partition.
3. The method according to claim 1, characterized in that The updating of the second transaction data in the target hot data partition under the target first-level partition by using the first transaction data includes: When a target business primary key identical to the business primary key of the first transaction data exists in the target hot data partition under the target first-level partition, the first transaction data is used to replace the second transaction data corresponding to the target business primary key in the target hot data partition.
4. The method according to claim 3, characterized in that In the case where a target business primary key identical to the business primary key of the first transaction data exists in the target hot data partition under the target first-level partition, before replacing the second transaction data corresponding to the target business primary key in the target hot data partition with the first transaction data, the method further includes: Determine a first key-value pair by using the transaction period and the business primary key in the first transaction data as a key and using data other than the transaction period and the business primary key in the first transaction data as a value; Determine a second key-value pair by using the transaction period and the business primary key in the second transaction data as a key and using data other than the transaction period and the business primary key in the second transaction data as a value; When the key of the first key-value pair is the same as the key of the second key-value pair, it is determined that a target business primary key that is the same as the business primary key of the first transaction data exists in the target hot data partition under the target first-level partition.
5. The method according to claim 4, characterized in that The first transaction data and the second transaction data include a business timestamp, and the business timestamp represents an update time of the transaction data; The using the first transaction data to replace the second transaction data corresponding to the target business primary key in the target hot data partition includes: Determine, according to the business timestamp, a target key-value pair whose transaction data is updated later in the first key-value pair and the second key-value pair; Store the transaction data corresponding to the target key-value pair in the storage path where the target hot data partition is located; In the storage path where the target hot data partition is located, the second transaction data corresponding to the second key-value pair is replaced with the transaction data corresponding to the target key-value pair.
6. The method according to claim 5, characterized in that Before storing the transaction data corresponding to the target key-value pair in the storage path where the target hot data partition is located, the method further includes: Output the transaction data corresponding to the target key-value pair to a temporary directory under the target first-level partition to obtain a data file corresponding to the target key-value pair; Backing up the data file corresponding to the second transaction data; The storing the transaction data corresponding to the target key-value pair in the storage path where the target hot data partition is located includes: The data file corresponding to the target key-value pair is moved to the storage path where the target hot data partition is located.
7. The method according to claim 1, characterized in that The transaction data also includes third transaction data newly added in the target period. After acquiring the transaction data in the target period, the method further includes: The third transaction data is stored in a cold data partition under the primary partition corresponding to the target period.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: For the secondary partition, when the storage time of the transaction data in the cold data partition reaches a preset time, the transaction data in the cold data partition and the transaction data in the hot data partition are merged.
9. The method according to claim 8, characterized in that The merging of the transaction data in the cold data partition and the transaction data in the hot data partition includes: For each fourth transaction data in the cold data partition, when fifth transaction data having the same business primary key as the fourth transaction data is stored in the hot data partition, the fourth transaction data is replaced by the fifth transaction data.
10. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: In response to the data query operation, for each of the first-level partitions, the following operations are performed: Respectively acquiring the sixth transaction data stored in the hot data partition under the first-level partition and the seventh transaction data stored in the cold data partition; generating a query view table based on the sixth transaction data and the seventh transaction data; In the case that the business primary keys of the sixth transaction data and the seventh transaction data are the same, deduplication processing is performed on the transaction data in the query view table based on the business timestamps of the sixth transaction data and the seventh transaction data to obtain an updated query view table.
11. A data processing device, characterized in that: The device comprises: An acquisition module, configured to acquire transaction data within a target period, wherein the transaction data includes first transaction data updated within the target period; A determination module, configured to determine, based on a transaction cycle of the first transaction data, a target first-level partition corresponding to the first transaction data and a target hot data partition under the target first-level partition from multiple first-level partitions of a data warehouse; the first-level partition is determined based on the transaction cycle, a second-level partition is provided under the first-level partition, the second-level partition includes a cold data partition and a hot data partition, the cold data partition is used to store transaction data newly added during the transaction cycle, and the hot data partition is used to store transaction data updated during the transaction cycle; An update module is used to update the second transaction data in the target hot data partition under the target first-level partition using the first transaction data.
12. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the data processing method according to any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the data processing method according to any one of claims 1 to 10 is implemented.
14. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the data processing method according to any one of claims 1 to 10.
Citation Information
Cited By
Transaction timestamp injection data time sequence reconstruction method compatible with cloud native architecture
CN121280142A
A transaction timestamp injection data timing reconstruction method compatible with cloud native architecture
CN121280142B