Data processing method and device and related equipment

By recording the updated or deleted data locations in the main table in the log file, the performance problems caused by the database system inversely checking the main table when generating binlog is solved, and the effect of improving the log file generation efficiency and database response performance is achieved.

CN120196650APending Publication Date: 2025-06-24HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410171936.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-02-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the process of generating binlog, the database system needs to check the values ​​of the corresponding columns in the main table through the primary key, resulting in the binlog saving and updating records for too long, affecting the response performance of the database system.

Method used

By recording the location of updated or deleted data in the main table in the log file, instead of looking up the main table in reverse, the efficiency of generating log files is improved, and the operation of frequent decompression of the compression unit is avoided.

Benefits of technology

It effectively improves the efficiency of generating log files, reduces the response time of the database system, and improves the performance of data entry into the database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196650A_ABST
    Figure CN120196650A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method which comprises the steps that a first database statement is obtained, the first database statement is used for executing a first target operation on first data in a main table in a database, the first target operation comprises an updating operation or a deleting operation, and the main table stores the data according to columns; in the process of executing the first database statement, the first position of the first data in the main table is recorded in a log file, and the log file further records a first target operation; and persistently storing the log file. In the process of updating or deleting the first data, the first position of the first data in the main table instead of the first data is recorded in the log file, so that the process of reversely finding out the first data from the main table does not need to be executed in the process of generating the log file, the efficiency of generating the log file can be effectively improved, and the user experience is improved. Therefore, the response performance of the database system can be improved. In addition, the invention also provides a corresponding data processing device and related equipment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application titled "Incremental Data Synchronization Method, Apparatus and Related Equipment" with an application number of 202311785285.X and filed with the National Intellectual Property Administration on December 22, 2023. The entire content of which is incorporated herein by reference. Technical Field

[0002] This application relates to the field of storage technology, and in particular, to a data processing method, apparatus and related equipment. Background Art

[0003] Currently, in databases such as Hologres database, GaussDB database, or Postgres database, etc., data can usually be stored by columns, and this data storage method can be called columnar storage (or column storage). In columnar storage, the data of one column will be compressed into one or more compression units (CUs) to improve the data compression rate. Moreover, when querying data, only the column where the data to be read is located needs to be loaded, without loading the entire table, thereby improving the data query efficiency.

[0004] In actual application scenarios, when modifying data in the database, the database system can use the write-ahead log (WAL) to ensure the atomicity, consistency, and isolation of data operations. That is, when modifying data, the database system needs to first successfully record the modification operation for the data in the binlog (binary log, which belongs to a type of WAL log), and then modify the data. Usually, each binlog can record multiple data modification operations. Finally, when the data volume of the binlog reaches the threshold, the database system writes the modified data into the database according to the data modification operations recorded in the binlog. In this way, if the database system fails, the binlog can be used to restore the data in the database to ensure that no data is lost when the database system fails. At the same time, when the data modification operation is successfully recorded in the binlog, the database system can feedback to the front end (such as the client) that the data writing is successful to improve the response efficiency of data writing.

[0005] However, in the process of generating the binlog, for each operation of updating data, the database system usually needs to reverse-lookup the value of the corresponding column in the main table through the primary key and write the reverse-looked-up value into the binlog to save the update record for the data in the binlog. This will cause the time for saving the update record in the binlog to be too long, thus affecting the response performance of the database system, such as resulting in a low efficiency of the database system in responding to updated data, etc. Summary of the Invention

[0006] In view of this, embodiments of the present application provide a data processing method to improve the performance of data storage in a database. The present application also provides a corresponding data processing device, a computing device cluster, a computer-readable storage medium, and a computer program product.

[0007] In a first aspect, embodiments of the present application provide a data processing method. This method can be executed by a corresponding data processing device. Specifically, the data processing device obtains a first database statement. The first database statement can be, for example, an SQL statement or a DML statement, etc. And this first database statement is used to perform a first target operation on the first data in the main table. The first target operation performed includes an update operation or a delete operation. Among them, the database includes this main table, and the main table stores data by columns. During the execution of the first database statement, the data processing device records the first position of the first data in the main table in a log file. The log file also records the first target operation in the first database statement. Finally, the data processing device persistently stores the log file. For example, the log file can be persistently stored on a local disk, etc.

[0008] Since during the process of updating or deleting the first data in the database, the data processing device records the first position of the first data updated or deleted in the main table in the log file, rather than recording the first data in the log file, this makes it unnecessary to perform the process of retrieving the first data from the main table during the generation of the log file, thereby effectively improving the efficiency of generating the log file, and further improving the response performance of the database system.

[0009] In a possible implementation manner, the first position of the first data in the main table is located in the first compression unit corresponding to the main table, that is, the first data is the data in the first compression unit, and the log file also records the second position of the second data in the first compression unit. Then, the data processing device can also access the first compression unit according to the first position and the second position recorded in the log file to obtain the updated first data and second data. That is, during the process of retrieving the main table, only one decompression of the first compression unit needs to be performed to obtain the first data and the second data. Then, the data processing device can update or delete the first data persistently stored in the database according to the first target operation recorded in the log file, and update or delete the second data persistently stored in the database according to the second target operation recorded in the log file. In this way, during the process of respectively performing update or delete operations on different data in the same compression unit, only one decompression process of this compression unit needs to be performed, and there is no need to decompress this compression unit multiple times, thereby avoiding the time consumption and resource consumption caused by decompressing the same compression unit multiple times and improving the efficiency of data storage in the database.

[0010] In a possible implementation, the first target operation may include an update operation. At this time, the first target operation performed by the data processing device is used to update the first data to the third data, and the log file records the third data while recording the first position and the first target operation.

[0011] In a possible implementation, the first target operation includes an update operation, and the update operation includes a pre-update identifier and a post-update identifier. The pre-update identifier is associated with the first position, and the post-update identifier is associated with the third data. In this way, the pre-update identifier and the post-update identifier can be recorded in the log file to indicate that the first data in the main table is updated to the third data.

[0012] In a possible implementation, when the transaction corresponding to the first database statement is successfully committed, the data processing device establishes a mapping relationship between the identifier of the transaction and the global CSN (transaction commit sequence number), and in the log file, records the CSN corresponding to the identifier of the transaction for the third data. In this way, the subsequent data processing device can track the commit status and order of each transaction according to this mapping relationship.

[0013] In a possible implementation, the CSN is used to indicate the order of synchronizing the first target operation, the first data, and the third data to the target storage area. In this way, in a distributed storage system, the incremental data on each data node can be synchronously transferred to the target storage area in sequence according to the global CSN.

[0014] In a possible implementation, the log file records the first position of the first data in the main table, the first target operation, the second position of the second data in the main table, and the second target operation. The first position and the second position are located in the first compression unit corresponding to the main table, that is, the first data and the second data are different data in the same compression unit. Then, the data processing device can also access the first compression unit according to the first position and the second position recorded in the log file to obtain the updated first data and second data. Thus, the data processing device can synchronize the first target operation, the first data, the second target operation, and the second data to the target storage area according to the CSN corresponding to the first data and the CSN corresponding to the second data recorded in the log file. In this way, in a data synchronization scenario (such as incremental data synchronization), for different data in the same compression unit, the data processing device can perform the process of decompressing the compression unit only once to obtain the first data and the second data, so as to synchronize the first data and the second data to the target storage area, and then process the first data (such as updating the first data to the third data, etc.) and process the second data in the target storage area according to the processed first data, second data, first target operation, and second target operation, so as to achieve data synchronization.

[0015] In a possible implementation, the first position of the first data in the main table includes the identifier of the first compression unit where the first data is located, and the number of rows of the first data in the first compression unit.

[0016] In a second aspect, the present application provides a data processing device, including: an acquisition module, configured to acquire a first database statement for performing a first target operation on first data in a main table, where the first target operation includes an update operation or a delete operation, and the database includes a main table that stores data by column; an execution module, configured to record the first position of the first data in the main table in a log file during the execution of the first database statement, and the log file also records the first target operation in the first database statement; and a storage module, configured to persistently store the log file.

[0017] In a possible implementation, the first position is located in a first compression unit corresponding to the main table, and the log file also records the second position of the second data in the first compression unit; the data processing device further includes a data change module, configured to access the first compression unit according to the first position and the second position recorded in the log file to obtain the updated first data and second data; update or delete the first data persistently stored in the database according to the first target operation recorded in the log file; and update or delete the second data persistently stored in the database according to the second target operation recorded in the log file.

[0018] In a possible implementation, the first target operation includes an update operation for updating the first data to third data, and the log file also records the third data.

[0019] In a possible implementation, the first target operation includes an update operation, and the update operation includes a pre-update identifier and a post-update identifier, where the pre-update identifier is associated with the first position, and the post-update identifier is associated with the third data.

[0020] In a possible implementation, the execution module is further configured to: when successfully committing the transaction corresponding to the first database statement, establish a mapping relationship between the identifier of the transaction and the global transaction commit sequence number CSN; and record the CSN corresponding to the identifier of the transaction in the log file for the third data.

[0021] In a possible implementation, the CSN is used to indicate the order of synchronizing the first target operation, the first data, and the third data to the target storage area.

[0022] In a possible implementation, the log file records a first location, a first target operation, a second location, and a second target operation, where the first location and the second location are located in a first compression unit corresponding to the main table; the data processing device further includes a synchronization module configured to: access the first compression unit according to the first location and the second location recorded in the log file to obtain updated first data and second data; and synchronize the first target operation, the first data, the second target operation, and the second data to a target storage area according to the CSN corresponding to the first data and the CSN corresponding to the second data recorded in the log file.

[0023] In a possible implementation, the first location includes an identifier of the first compression unit where the first data is located and the row number of the first data in the first compression unit.

[0024] The data processing device provided in the second aspect corresponds to the data processing method provided in the first aspect. Therefore, for the technical effects of the data processing device provided in the second aspect, reference can be made to the relevant descriptions of the technical effects of the first aspect or any implementation manner in the first aspect, which will not be elaborated herein.

[0025] In a third aspect, the present application provides a computing device cluster, which includes at least one computing device. The at least one computing device includes at least one processor and at least one memory; the at least one memory is configured to store instructions, and the at least one processor executes the instructions stored in the at least one memory to enable the computing device cluster to execute the data processing method in the first aspect or any possible implementation manner in the first aspect. It should be noted that the memory may be integrated into the processor or independent of the processor. The at least one computing device may further include a bus. Among them, the processor is connected to the memory through the bus. Among them, the memory may include a readable memory and a random access memory.

[0026] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are run on at least one computing device, the at least one computing device is enabled to execute the data processing method in the first aspect or any possible implementation manner in the first aspect.

[0027] In a fifth aspect, the present application provides a computer program product containing instructions. When the computer program product is run on at least one computing device, the at least one computing device is enabled to execute the data processing method in the first aspect or any possible implementation manner in the first aspect.

[0028] Based on the implementation manners provided in the above aspects, the present application can be further combined to provide more implementation manners. Description of the Drawings

[0029] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a schematic structural diagram of an exemplary distributed storage system provided by the present application;

[0031] Figure 2 It is a schematic diagram of storing data by columns in the main table;

[0032] Figure 3 It is a schematic flowchart of a data processing method provided by the present application;

[0033] Figure 4 It is a schematic diagram of using multiple columns to record information in a log file provided by the present application;

[0034] Figure 5 It is a schematic diagram of synchronizing incremental data in the distributed storage system 10 to the storage system 20 provided by the present application;

[0035] Figure 6 It is a schematic diagram of the main thread extracting incremental data on multiple data nodes to the incremental data queue according to the CSN;

[0036] Figure 7 It is a schematic diagram of multiple threads respectively extracting incremental data on multiple data nodes to the incremental data queue according to the CSN;

[0037] Figure 8 It is a schematic structural diagram of a data processing device provided by the present application;

[0038] Figure 9 It is a schematic structural diagram of a computing device provided by the present application;

[0039] Figure 10 It is a schematic structural diagram of a computing device cluster provided by the present application. Detailed implementation manners

[0040] The following will describe the solutions in the embodiments provided by the present application in conjunction with the accompanying drawings in the present application.

[0041] The terms "first", "second", etc. in the specification, claims and above-mentioned accompanying drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application.

[0042] See Figure 1 , which is a schematic structural diagram of an exemplary distributed storage system 10. As Figure 1 shown, the distributed storage system 10 includes an application 101, a coordination node 102, and N data nodes, Figure 1 taking the data nodes from data node 1 to data node N as an example, where N is an integer greater than 1, and each data node can be configured with a persistent storage area for persistently storing data locally in a columnar storage manner.

[0043] Among them, a database system can be co-deployed on the coordination node and the N data nodes. The persistent storage areas configured on the N data nodes can be used to build a database. The data stored in the local persistent storage area of each data node can be part of the data in the database, such as one or more data shards, etc., and the local persistent storage areas of different data nodes are used to store different parts (different shards) of the data.

[0044] Moreover, the N data nodes can be distributed in the same region or in different regions. Further, the N data nodes can be distributed in the same availability zone (AZ), or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, usually one region can include multiple AZs. Similarly, the N data nodes can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Among them, usually one VPC is set within one region, and communication gateways need to be set in each VPC for cross-region communication between two VPCs within the same region and between VPCs in different regions, and the interconnection between VPCs is achieved through the communication gateways.

[0045] The application 101 is used to generate database statements, which can be, for example, structured query language (SQL) statements, or data manipulation language (DML) statements, or other types of statements. And, these database statements can be used to update the data in the database, including inserting (insert), deleting (delete), updating (update), etc. the data on each data node. Further, these database statements can also be used to query the data in the database.

[0046] The coordination node 102 is used to receive database statements sent by the application 101, and forward the database statements to the data node where the data to be updated or queried as indicated by the database statements is located, so that the data node can perform corresponding operations to update or query the database. Moreover, the coordination node 102 is also used to feedback the execution results of each data node for the database statements to the application 101. Among them, the execution results for the database statements can be, for example, a response indicating successful data update, or the data requested by the application 101 for query, etc.

[0047] The data node 1 (the same applies to the other data nodes) is used to execute the received database statements, insert, delete, update, query data, etc. in the local persistent storage area, and feedback the execution results of the database statements to the coordination node 102. Among them, a data processing device 200 can be deployed in the data node 1, then, the relevant operations performed by the data node 1 on the database can be specifically executed by the data processing device 200.

[0048] Taking the data processing device 200 executing the database statement as an example, usually, during the execution of the database statement by the data processing device 200, the main table can be read from the persistent storage area into the memory, and corresponding operations can be performed on the main table in the memory, such as updating the data in the main table in the memory. Among them, the main table can be continuously stored in the memory so that the data in the main table to be read can be directly hit from the memory subsequently. Then, the data processing device 200 can save the operation records for the data in the database to a log file (such as binlog, etc.). Usually, after the operation records are successfully saved to the log file, the data processing device 200 can feedback a response indicating successful data update to the application 101 through the coordination node 102, without waiting for the updated main table to be successfully saved from the memory to the local persistent storage area, so as to improve the response efficiency of the distributed data system 10.

[0049] However, if the data processing device 200 saves the updated data a and the updated data A in the log file during the process of updating the data in the main table, then, the data processing device 200 needs to first execute the process of reverse querying the main table to obtain the updated data a. As Figure 2As shown in the figure, the main table stores data by column, and consecutive multiple rows of data in a column of the main table (or all row data included in this column) will be compressed and stored in a CU (compression unit). Therefore, when the data processing device 200 back-checks the data a in the main table, it will first determine the CU where the data a is located, and decompress this CU to obtain all row data stored in this CU, then find the data a from the multiple rows of data obtained by decompression, and finally save the data a to the log file. Similarly, when the data processing device 200 executes the next database statement for updating data in the main table, it will also refer to the above similar method to write the updated data in the main table obtained by back-checking into the log file. In this way, the data processing device 200 needs to frequently execute the process of back-checking the main table, that is, it needs to frequently perform the operation of decompressing the CU, which will cause the delay caused by frequent decompression of the CU to reduce the efficiency of saving data update records in the log file, thereby affecting the performance of the database system.

[0050] Based on this, in Figure 1 In the distributed storage system 10 shown in the figure, during the process of updating data in the main table, the data processing device 200 will record the position of the updated data a in the main table in the log file, rather than saving the data a in the log file by back-checking the main table. This position can be represented by the identifier of the CU where the data a is located and the row number of the data a in this CU, etc. Of course, the log file will also record the updated data A and the update operation, etc. In this way, during the process of generating the log file, the data processing device 200 does not need to execute the process of back-checking the data a from the main table, so as to avoid the data processing device from frequently performing the operation of decompressing the CU, thereby effectively improving the efficiency of saving data update records in the log file, and further improving the response performance of the database system.

[0051] Furthermore, since the updated data A is saved in the log file, therefore, the data processing device 200 usually needs to save the data A in the log file to the local persistent storage area. During this process, the data processing device 200 can back-check the main table to determine the updated data a. At this time, the data processing device 200 can determine multiple updated data located in the same CU (different data are indicated for update based on different database statements), and perform a decompression operation on this CU to obtain multiple updated data (including the data a). In this way, during the process of updating multiple data in the same CU, the data processing device only needs to perform the decompression operation on the CU once, so as to avoid the process of frequently performing the decompression operation on the same CU, thereby further improving the performance of the database system.

[0052] Similarly, when the data processing device 200 deletes data a in the main table, it will record the position of the updated data a in the main table in the log file, instead of saving data a in the log file by reverse querying the main table. The position can be represented by the identifier of the CU where data a is located and the number of rows of data a in the CU. Of course, the log file will also record the deletion operation. In this way, when the data processing device 200 generates the log file, it is not necessary to perform the process of reverse querying data a from the main table, so as to avoid the data processing device frequently performing the operation of decompressing the CU, thereby effectively improving the efficiency of saving data update records in the log file, and further improving the response performance of the database system.

[0053] Regarding the remaining data nodes in the distributed storage system 10, data processing devices may also be deployed. The data processing device deployed on each data node may have the same or similar functions as the data processing device 200, and is responsible for performing corresponding data query, insertion, deletion, and update operations on the local persistent storage data.

[0054] Further, Figure 1 The distributed storage system 10 shown may also include a global transaction management (GTM) 103, which may be used to assign a globally unified transaction commit sequence number (CSN) to transactions submitted in each data node to record the submission order of the transactions.

[0055] It is worth noting that the above Figure 1 The distributed storage system 10 shown is only an implementation example. In other distributed storage systems, other types of devices may also be included to support the distributed storage system with more functions. The present application does not limit the specific architecture of the distributed storage system 10. Figure 1 The distributed storage system 10 shown is used as an example to store data update records in a log file. In other possible storage systems, data update records may also be stored in a log file in accordance with the above method. For example, in other possible storage systems, only application 101 and data node 1 may be included, or in a centralized storage system including data node 1, data node 1 may store the location of updated data in a log file in accordance with the above implementation method to avoid data node 1 frequently performing the process of reverse querying the main table, and this is not limited.

[0056] In addition, the above data processing device 200 may be implemented by software or hardware. As an example, the implementation of the data processing device 200 is described below.

[0057] As an example of a software functional unit, the data processing device 200 may include code running on an instance, which may be at least one of a host, a virtual machine, a container, a thread, and a process. Further, the above instance may be one or more. For example, the data processing device 200 may include code running on multiple hosts / virtual machines / containers.

[0058] As an example of a hardware functional unit, the data processing device 200 may include at least one computing device. Alternatively, the data processing device 200 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), or any combination thereof.

[0059] For ease of understanding, various non-limiting specific embodiments of the process of processing images will be described in detail below.

[0060] Refer to Figure 3 , which is a schematic flowchart of a data processing method in an embodiment of the present application. This method can be applied to the distributed storage system 10 shown above Figure 1 or may also be applied to other applicable data storage systems. The following takes the distributed storage system 10 shown in Figure 1 as an example to introduce the process of the data processing device 200 updating the data in the main table.

[0061] As Figure 3 shown, the data processing method may specifically include:

[0062] S301: The application 101 sends the database statement 1 to the coordination node 102.

[0063] Among them, the database statement 1 sent by the application 101 can be, for example, an SQL statement, a DML statement, etc., and there is no limitation on this. Moreover, the database statement sent by the application 101 can be used to indicate data query, data insertion, data deletion, or data modification of the database. In this embodiment, it is taken as an example that the database statement 1 is specifically used to indicate updating the data 1 in the main table included in the database to the data 2.

[0064] In actual application, the application 101 can be, for example, a client on the user side, and the application 101 can generate the database statement based on the user's operation. When the user instructs to update the data 1 in the database to the data 2, the application 101 can generate the database statement 1 for updating the data 1.

[0065] S302: The coordination node 102 forwards the database statement 1 to the data processing device 200 in the data node 1.

[0066] In the distributed storage system 10, the local storage area of each data node is responsible for the persistent storage of part of the data in the database, and the data node is responsible for the corresponding data reading and writing of the local data. Therefore, after receiving the database statement 1, the coordination node 102 can first determine the data node where the data 1 to be updated is located. In this embodiment, it is set as the data node 1, then the coordination node 102 can forward the database statement 1 to the data node 1 so that the data node 1 can update the data 1 in the database according to the database statement 1.

[0067] S303: During the execution of the database statement 1 by the data processing device 200, the position 1 of the data 1 to be updated in the main table, the updated data 2, and the update operation 1 in the database statement 1 are recorded in the log file.

[0068] In this embodiment, the data node 101 can use the data processing device 200 to execute the data update process.

[0069] Specifically, the data processing device 200 can first parse the database statement 1 to determine the type of the operation to be executed (i.e., the update operation 1) and determine the position 1 of the data 1 to be updated in the main table. Then, the data processing device 200 can query whether there is a main table containing the data 1 in the memory. If so, the data processing device 200 can update the data 1 in the main table to the data 2 according to the database statement 1. If not, the data processing device 200 can first read the main table from the local persistent storage area to the memory, and then update the data 1 included in the main table in the memory.

[0070] Meanwhile, the data processing device 200 can record this location 1 in the log file, and record the update operation 1 for this data 1 and the updated data (i.e., data 2). Among them, the data processing device 200 can serially store data 2 in the log file. At this time, the location 1, the update operation 1, and the data 2 saved in the log file can constitute the update record for data 1.

[0071] In this embodiment, the log file stores data by columns, and the log file can include multiple columns. Exemplarily, as Figure 4 shown, the log file can include column 2 and column 3. Among them, column 2 is used to record the operation type, and column 3 is used to record the value, which can be the updated data 2 and the location 1 of the data 1 being updated in the main table. Among them, the update operation saved by the data processing device 200 in the log file can be an operation code, or can be an identifier of the update operation.

[0072] When the update operation saved in the log file is specifically the identifier of the update operation, the identifier can include a pre-update identifier and a post-update identifier. For example, the pre-update identifier can be "before_update", and the post-update identifier can be "after_update", etc. And, the pre-update identifier can be associated with the location 1 of the data 1 being updated in the main table 1, that is, the pre-update identifier is associated with the data 1 being updated. The post-update identifier can be associated with the updated data 2. As Figure 4 shown, "before_update" (pre-update identifier) and TID (location 1) can be correspondingly saved in the log file; "after_update" (post-update identifier) and the updated data 2 can be correspondingly saved in the log file.

[0073] After the data processing device 200 successfully saves the update record for data 1 in the log file, without waiting for the main table in the memory to complete persistent storage, it can send a response indicating successful data update to the coordination node 102, and the coordination node will feedback this response to the application 101, as Figure 3 shown. In this way, the latency for the database system to feedback successful data update to the application 101 can be shortened, thereby improving the performance of the database system. In other embodiments, the data processing device 200 can also send a response indicating successful data update to the coordination node 102 after successfully completing the persistent storage of the log file.

[0074] S304: The data processing device 200 persistently stores the log file.

[0075] In a first possible implementation, each time the data processing device 200 executes a database statement, it can persistently store the generated log file.

[0076] In a second possible implementation, the data processing device 200 can perform a persistent storage log process only when it has completed executing multiple database statements. For example, the data processing device 200 can create a transaction for each database statement and commit the transaction when the database statement is executed. Then, when the number of committed transactions reaches a preset value, the data processing device 200 can flush the log file to the local persistent storage area. At this time, the log file can save multiple operation records for different data.

[0077] The above is described by taking the execution of one database statement as an example. For other database statements used to update the data in the main table, the above similar process can be referred to to implement the data update of the main table. For example, after updating the data in the second CU, the data processing device 200 can record the position of the updated data in the second CU in the log file, and record the update operation and the updated data.

[0078] Generally, the updated data 2 needs to be saved to the main table in the persistent storage area. Therefore, the data processing device 200 can also update the data 2 to the main table according to the log file at an appropriate time (such as a low-load state).

[0079] As a first implementation example, for each data update record in the log file, the data processing device 200 can search the main table in the persistent storage area according to the position of the updated data recorded in the log file to determine the updated data at that position in the main table, and update the data to the new data included in the data update record.

[0080] As a second implementation example, for multiple data update records in the log file, the data processing device 200 can first analyze multiple positions in the same CU in the main table in the log file. Then, the data processing device 200 can reverse-search the main table once according to the multiple positions to determine the updated data corresponding to the multiple positions in the main table, so that the data processing device 200 can update each updated data. In this way, for multiple updated data located in the same CU, the data processing device 200 only needs to reverse-search the main table once, that is, only needs to perform a decompression operation on the CU once. This can not only avoid the resource consumption caused by repeated decompression of the CU, but also improve the efficiency of the data processing device 200 in updating the data in the main table, that is, improve the performance of data warehousing.

[0081] For ease of understanding, the following takes the example where the data indicated by the two data update records included in the log file is the data in the same CU for illustration. Based on this, this embodiment may further include the following steps.

[0082] S305: The data processing device 200 accesses the first CU in the main table according to the position 1 and position 2 recorded in the log file, and obtains the updated data 1 and data 3, where the position 1 of data 1 in the main table, the update operation 1 for data 1, the updated data 2, the position 2 of data 3 in the main table, the update operation 2 for data 3, and the updated data 4 are recorded in the log file.

[0083] Among them, the specific implementation manner of the data processing device 200 recording the position 2, data 3, and data 4 in the log file is similar to the specific implementation manner of recording the position 1, data 1, and data 2 described above. For details, reference can be made to the relevant descriptions above, and no further elaboration will be provided here. In this embodiment, it is assumed that the position 1 and the position 2 belong to different positions in the first CU.

[0084] In practical applications, the data processing device 200 may execute the process of storing data 2 and data 4 into the database under the condition of meeting a preset condition. For example, when the data processing device 200 detects that the data node 1 is in a low-load state (such as the load is lower than the threshold, etc.), it can execute the data storage process according to the log file. Another example is that when the number of data update records included in the log file reaches a preset quantity, the data storage process can be executed according to the log file.

[0085] In this embodiment, during the process of persistently storing the updated data 2 and data 4 into the database according to the log file, the data processing device 200 may first analyze the positions of multiple updated data recorded in the log file in the main table to determine the positions located in the same CU. For example, the position of the updated data in the main table can be indicated by the CU identifier and the row identifier of the data in the CU, so that the data processing device 200 can determine multiple positions located in the same CU. In this embodiment, it is assumed that the multiple positions located in the first CU include the position 1 corresponding to data 1 and the position 2 corresponding to data 3.

[0086] S306: The data processing device 200 updates the data 1 persistently stored in the database to data 2 according to the update operation 1 recorded in the log file, and updates the data 3 persistently stored in the database to data 4 according to the update operation 2 recorded in the log file.

[0087] After determining position 1 and position 2 in the same CU, the data processing device 200 can decompress the first CU in the main table according to position 1 and position 2 to obtain multiple rows stored in the CU. Thus, the data processing device 200 can determine the updated data 1 indicated by position 1 and the updated data 3 indicated by position 2 according to the positions of the updated data included in position 1 and position 2 in the CU. For example, the updated data 1 and data 3 can be determined according to the row number identifiers included in position 1 and position 3 respectively. In this way, in the process of determining data 1 and data 3 in the main table, the data processing device 200 can perform the decompression process on the first CU only once, without performing the decompression process on the first CU according to position 1 and position 2 respectively.

[0088] After determining data 1 and data 3, the data processing device 200 can update data 1 in the main table to data 3 and update data 2 to data 4 according to the data update records in the log file, so as to write the updated data 2 and data 4 into the persistent storage main table and complete data warehousing. Since only one decompression process needs to be performed on the first CU during the process of warehousing data 2 and data 4, the delay caused by performing multiple decompressions on the first CU can be avoided, thereby improving the performance of warehousing data 2 and data 4.

[0089] The processes of steps S301 to S306 mainly introduce the example of the data processing device 200 updating the data in the database. In actual application, the data processing device 200 can insert new data into the database, or delete the data stored in the database, or query the data stored in the database, etc. The scenarios of data insertion, data deletion, and data query will be introduced in detail below.

[0090] 1. Regarding the data insertion scenario, the application 101 can send the database statement 2 to the coordination node 102. The database statement 2 is used to indicate inserting data 5 into the main table included in the database. Assume that the main table is stored in the data node 1. The coordination node 102 can send the database statement 2 to the data node 1. For the database statement 2, the data processing device 200 in the data node 1 can query that the main table is included in the memory. For example, the data processing device 200 read the main table from the local storage area into the memory during the execution of the database statement 1 before, and insert data 5 into the main table in the memory. At the same time, the data processing device 200 can generate a data insertion record in the log file to indicate inserting data 5 into the main table. The data insertion record can include the identifier of the insertion operation and the inserted data 5. Exemplarily, Figure 4As shown, the data processing device 200 can save the identifier "I" of the insertion operation in column 2 of the log file, and save the serialized data 5 in column 3 of the log file.

[0091] Accordingly, when storing the inserted data into the database, the data processing device 200 can save the data 5 into the main table in the local persistent storage area according to the identifier of the insertion operation and the inserted data 5 recorded in the log file, thereby realizing data warehousing.

[0092] 2. Regarding the data deletion scenario, the application 101 can send the database statement 3 to the coordination node 102. The database statement 3 is used to indicate deleting the data 6 in the main table included in the database. Assume that the main table is saved in the data node 1. The coordination node 102 can send the database statement 3 to the data node 1. For the database statement 3, the data processing device 200 in the data node 1 can query that the main table is included in the memory and delete the data 6 in the main table in the memory. At the same time, the data processing device 200 can generate a data deletion record in the log file to indicate deleting the data 6 in the main table. The data deletion record can include the identifier of the deletion operation and the position 3 of the deleted data 6 in the main table. Exemplarily, as Figure 4 shown, the data processing device 200 can save the identifier "D" of the deletion operation in column 2 of the log file, and save the position 3 of the deleted data 6 in the main table in column 3 of the log file.

[0093] Accordingly, when updating the database according to the data deletion record in the log file, the data processing device 200 can determine the data 6 deleted in the main table according to the position 3 recorded in the log file, and delete the data 6 in the main table according to the identifier of the deletion operation recorded in the log file.

[0094] 3. Regarding the data query scenario, the application 101 can send the database statement 4 to the coordination node 102. The database statement 4 is used to indicate querying the data 7 in the main table included in the database. Assume that the main table is saved in the data node 1. The coordination node 102 can send the database statement 4 to the data node 1. For the database statement 4, the data processing device 200 in the data node 1 can query that the main table is included in the memory and find out the data 7 from the main table. At the same time, the data processing device 200 can generate a data query record in the log file to indicate querying the data 7 in the main table. The data query record can include the identifier of the query operation and the position 4 of the queried data 7 in the main table. Exemplarily, as Figure 4 shown, the data processing device 200 can save the identifier "S" of the query operation in column 2 of the log file, and save the position 4 of the deleted data 7 in the main table in column 3 of the log file.

[0095] Accordingly, when updating the database according to the records in the log file, for data query records, there is no need to adjust the data in the database.

[0096] Furthermore, the data processing device 200 can support a transaction mechanism, that is, during the execution of a database statement, the data processing device 200 can create a transaction, and only when the transaction is successfully committed, it will determine that the database statement execution is completed. Based on this, during the execution of a database statement, the data processing device 200 will create a transaction for the database statement and assign an identifier to the created transaction, and this identifier can be, for example, a transaction number, etc. After the data processing device 200 successfully saves a data operation record (such as a data update record, a data deletion record, etc.) to the log file, it can commit the transaction and assign a globally unified CSN to the committed transaction in the log file, where the value of the CSN increases gradually. As Figure 4 shown, for each row value in column 2 and column 3 of the log file, the data processing device 200 can record the corresponding CSN value in column 1 of the log file.

[0097] In an actual application scenario, since there may be a difference between the order in which the data processing device 200 creates transactions and the order in which it commits transactions. For example, the multiple transactions successively created by the data processing device 200 are transaction 1, transaction 2, transaction 3, and transaction 4 respectively, while the multiple transactions successively committed by the data processing device 200 are transaction 2, transaction 1, transaction 4, transaction 3, etc. This makes the order of the CSN values corresponding to the multiple transactions inconsistent with the creation order of the multiple transactions when the data processing device 200 commits the multiple transactions. Therefore, the data processing device 200 can establish a mapping relationship between the identifier of each transaction and the CSN of the transaction. For example, it can establish a mapping relationship between the transaction number and the CSN through an asynchronous thread in the data processing device 200, so as to track the commit status and order of each transaction in the future.

[0098] In this embodiment, for the database statement sent by application 101 to perform a target operation on the data in the main table, when the target operation is an update operation or a deletion operation, during the process of the data processing device 200 executing the database statement and generating a log file, there is no need to perform the process of retrieving the data from the main table in reverse, but record the position of the updated or deleted data in the main table in the log file, so as to avoid the data processing device 200 from frequently performing the operation of decompressing the CU, thereby effectively improving the efficiency of saving data update records or data deletion records in the log file, and further improving the response performance of the database system.

[0099] Moreover, in the above steps S305 and S306, the example is updating different data in the same CU. When deleting different data in the same CU, or deleting and updating different data in the same CU separately, the data processing device 200 can also record the operation identifier and the updated or deleted position in the log file by referring to the above method. Moreover, when performing persistent update or deletion of data in the database, for different data in the same CU, the operation of decompressing the CU can be executed only once, so as to avoid the latency caused by decompressing the same CU multiple times, thereby improving the performance of data storage in the database.

[0100] In an actual application scenario, the data in the distributed storage system 10 may need to be synchronized to other storage systems. For example, the incremental data in the distributed storage system 10 can be synchronized to other storage systems periodically. Among them, the incremental data refers to the data corresponding to the data insertion, deletion, and update operations executed in the distributed storage system 10 during the time period between the last data synchronization time and the current time. For the sake of understanding, the following will be combined with Figures 5 to 7 , and a detailed description will be given here.

[0101] Refer to Figure 5 , which shows a schematic diagram of synchronizing the incremental data in the distributed storage system 10 to the storage system 20. Among them, the storage system 20 can be a distributed storage system or a centralized storage system, and this is not limited. Among them, the incremental data synchronization process can be executed by the distributed storage system 10, or can be executed by the storage system 20. For the sake of understanding, the following will take the data processing device 200 in the distributed storage system 10 executing the incremental data synchronization process as an example for description. At this time, the functional module for executing incremental data synchronization in the data processing device 200 can be deployed on the data node 1, or can be deployed on the coordination node 102 in the distributed storage system 10 or a separately deployed node, etc., and this is not limited.

[0102] In the distributed storage system 10, each data node can maintain its own synchronization point information, which can be, for example, the CSN corresponding to the data that was last synchronized to the storage system 20. Then, during the process of synchronizing the incremental data in the distributed storage system 10 to the storage system 20, the data processing device 200 can first obtain the synchronization point information maintained by each data node 1, and based on the synchronization point information maintained by each data node 1, synchronize the incremental data on each data node to the storage system 20. Among them, the incremental data on each data node can be determined according to the synchronization point information. Taking the synchronization point information as the CSN specifically, the incremental data on the data node 1 is the data between the CSN at the last synchronization of the data node 1 and the maximum CSN currently stored. For example, assuming that the CSN at the last synchronization of the data is 100 and the maximum CSN currently stored is 150, the incremental data on the data node 1 is the data with CSN from 101 to 150. Among them, since the CSN is globally and uniformly maintained for multiple data nodes, the CSN recorded on the data node 101 may not be continuous. For example, the CSN on the data node 101 can include 101, 102, 106, 110, 150, and the CSN on the data node 2 is 103, 104, 105, 107, 149, etc.

[0103] In this embodiment, the following two non-limiting implementation examples for synchronizing incremental data are provided.

[0104] In the first implementation example, as Figure 6 shown, the data processing device 200 can include a main thread and multiple sub-threads. Among them, each sub-thread is used to query the synchronization point information on a data node and provide the queried synchronization point information to the main thread; among them, multiple sub-threads can query the synchronization point information concurrently. The main thread is used to determine the incremental data on each data node according to the synchronization point information fed back by each sub-thread, and synchronize the incremental data to the storage system 20, specifically, it can be synchronized to the target storage area in the storage system 20 for persistent storage.

[0105] Taking the synchronization point information as the CSN specifically, each sub-thread can obtain the CSN corresponding to the data that has been synchronized on a data node and provide the CSN to the main thread. The main thread can extract the incremental data on each data node to the incremental data queue in ascending order of the CSN according to the CSN fed back by each sub-thread, as Figure 6 shown.

[0106] For example, assume that the distributed storage system includes data node 1 to data node 3, and the CSN corresponding to the data synchronized on data node 1 is 100 (i.e., the synchronization point information), and the incremental data thereon includes the data corresponding to CSN 101 and CSN 105. The CSN corresponding to the data synchronized on data node 2 is 99 (i.e., the synchronization point information), and the incremental data thereon includes the data corresponding to CSN 103 and 104. The CSN corresponding to the data synchronized on data node 3 is 98 (i.e., the synchronization point information), and the incremental data thereon includes the data corresponding to CSN 102 and 106. Then, the main thread can determine that the minimum CSN corresponding to the data that has not been synchronized currently is 101, extract the data corresponding to CSN 101 on data node 1 into the data queue, and can mark the data corresponding to CSN 101 on data node 1 as having completed incremental data synchronization through a child thread. Then, the main thread can determine that the minimum CSN that has not been synchronized currently is 102, extract the data corresponding to CSN 102 on data node 3 into the data queue, and can mark the data corresponding to CSN 102 on data node 3 as having completed incremental data synchronization through a child thread. Next, the main thread can determine that the minimum CSN that has not been synchronized currently is 103, extract the data corresponding to CSN 103 on data node 2 into the data queue, and can mark the data corresponding to CSN 103 on data node 2 as having completed incremental data synchronization through a child thread. And so on, the main thread can successively extract the data corresponding to CSN 101 on data node 1, the data corresponding to CSN 102 on data node 3, the data corresponding to CSN 103 on data node 2, the data corresponding to CSN 104 on data node 2, the data corresponding to CSN 105 on data node 1, and the data corresponding to CSN 106 on data node 3 into the incremental data queue in sequence. At this time, the synchronization point information maintained by data node 1 can be updated to CSN 105, the synchronization point information maintained by data node 2 can be updated to CSN 104, and the synchronization point information maintained by data node 3 can be updated to CSN 106.

[0107] In this way, it is possible to control the consumption progress of the synchronization point information of each data node to have a relatively small difference, and thus can avoid as much as possible the problem of data synchronization disorder caused by a large difference in the synchronization point information of each data node.

[0108] Then, the main thread can sequentially send the data in the incremental data queue to the storage system 20 to synchronize the incremental data in the distributed storage system 10 to the storage system 20. In specific implementation, the incremental data queue can record not only the incremental data, but also the operation type corresponding to the incremental data. Among them, for the data update scenario, the incremental data queue can record the data to be updated, the updated data, and the identifier of the update operation; for the data deletion scenario, the incremental data queue records the data to be deleted and the identifier of the deletion operation; for the data insertion scenario, the incremental data queue records the inserted data and the identifier of the insertion operation. In this way, the data processing device 200 sequentially generates corresponding database statements according to the operation type and data in the incremental data queue, and sends the database statements to the storage system 20, so that the storage system 20 can synchronize the incremental data by sequentially executing the database statements.

[0109] In the second implementation example, as Figure 7 shown, the data processing device 200 may include multiple threads. Among them, each thread is used to query the synchronization point information on a data node, and determine the incremental data on the data node according to the synchronization point information, so as to synchronize the incremental data to the storage system 20.

[0110] Taking the synchronization point information as CSN specifically, each thread can extract the incremental data on the data node it is responsible for into the incremental data queue in ascending order of CSN according to the CSN on the data node it is responsible for. Thus, multiple threads can parallelly extract the incremental data on multiple data nodes into multiple different incremental data queues, as Figure 7 shown.

[0111] For example, assume that the distributed storage system includes data nodes 1 to 3, and the CSN corresponding to the data synchronized on data node 1 is 100 (i.e., the synchronization point information), and the incremental data on it includes the data corresponding to CSN 101 and CSN 105. The CSN corresponding to the data synchronized on data node 2 is 99 (i.e., the synchronization point information), and the incremental data on it includes the data corresponding to CSN 103 and 104. The CSN corresponding to the data synchronized on data node 3 is 98 (i.e., the synchronization point information), and the incremental data on it includes the data corresponding to CSN 102 and 106.

[0112] Thread 1 is responsible for synchronizing the incremental data on data node 1. Moreover, since the minimum CSN corresponding to the data that has not been synchronized yet on data node 1 is 101, thread 1 can extract the data corresponding to CSN 101 into data queue 1, and update the CSN corresponding to the data that has been synchronized on data node 2, that is, update the synchronization point information. Then, thread 1 extracts the data corresponding to CSN 105 into data queue 1 and updates the synchronization point information.

[0113] Meanwhile, thread 2 can determine that the minimum CSN that has not been synchronized yet on data node 2 is 103, and extract the data corresponding to CSN 103 on data node 2 into data queue 2 and update the synchronization point information. Then, thread 2 extracts the data corresponding to CSN 104 into data queue 2 and updates the synchronization point information. Moreover, thread 3 can determine that the minimum CSN that has not been synchronized yet on data node 3 is 102, and extract the data corresponding to CSN 102 on data node 2 into data queue 3 and update the synchronization point information. Then, thread 3 extracts the data corresponding to CSN 106 into data queue 3 and updates the synchronization point information.

[0114] At this time, the synchronization point information maintained by data node 1 can be updated to CSN 105, the synchronization point information maintained by data node 2 can be updated to CSN 104, and the synchronization point information maintained by data node 3 can be updated to CSN 106.

[0115] Then, each thread can sequentially send the data in its respective incremental data queue to storage system 20 to synchronize the incremental data in distributed storage system 10 to storage system 20. Specifically, when implemented, each incremental data queue can record not only the incremental data but also the operation type corresponding to the incremental data. Thus, each thread can sequentially generate corresponding database statements based on the operation type and data in the incremental data queue, and send the database statement to storage system 20 so that storage system 20 can synchronize the incremental data by sequentially executing the database statements.

[0116] In this way, by using multiple threads to parallelly extract the incremental data on multiple data nodes into the corresponding incremental data queues, it can be ensured that the multiple data nodes do not wait for each other and independently execute the process of extracting incremental data, which can help improve the efficiency of synchronizing the incremental data on multiple data nodes in distributed storage system 10 to storage system 20.

[0117] Among them, when performing incremental data synchronization according to the log file, for the operation records of updating or deleting different data in the same CU recorded in the log file, the data processing device 200 can also obtain the deleted or updated data by decompressing the CU once, so as to avoid the latency caused by performing multiple decompressions on the same CU and improve the performance of incremental data synchronization.

[0118] In specific implementation, the log file records the first position, the target operation 1 (delete operation or update operation) performed on the data 1 at the first position in the main table, the second position, and the target operation 2 (delete operation or update operation) performed on the data 2 at the second position in the main table. And, the first position and the second position are located in the same CU corresponding to the main table. Then, the data processing device 200 can access the CU according to the first position and the second position recorded in the log file to obtain the updated data 1 and data 2. Thus, the data processing device 200 can synchronize the data 1, the target operation 1, the data 2, and the target operation 2 to the target storage area in the storage system 20 according to the CSN corresponding to the data 1 and the CSN corresponding to the second data recorded in the log file. Among them, for the specific implementation manner of the data processing device 200 to obtain the updated or deleted data through the process of decompressing the same CU once according to the log file, reference can be made to the relevant descriptions above and will not be elaborated here.

[0119] Furthermore, when both the target operation 1 and the target operation 2 are update operations, the data processing device 200 can also synchronize the data obtained after updating data 1 and the data obtained after updating data 2 to the target storage area.

[0120] Similarly, when the log file also records insertion records for inserting data into the database, the data processing device 200 can also synchronize the insertion operation and the inserted data to the target storage area.

[0121] In this way, the storage system 20 updates or deletes its persistently stored data according to the synchronized data and operations, thereby realizing incremental data synchronization.

[0122] It should be noted that the above takes the data processing device 200 performing the synchronization of incremental data in the distributed storage system 10 to the storage system 20 as an example for illustrative description. In actual application, it can also be the storage system 20 that performs the process of synchronizing incremental data to the storage system 20 with reference to the above implementation manner; or it can be a data processing device independently deployed from the data nodes in the distributed storage system 10 that performs the synchronization process of incremental data; or it can be a third-party software tool that synchronizes the incremental data in the distributed storage system 10 to the storage system 20, etc., and this is not limited.

[0123] Figure 8 shows a schematic structural diagram of the data processing device 200. As Figure 8 shown, the data processing device 200 may include:

[0124] An acquisition module 801, configured to acquire a first database statement, where the first database statement is used to perform a first target operation on first data in a main table, and the first target operation includes an update operation or a delete operation. Among them, the database includes a main table, and the main table stores data by columns;

[0125] An execution module 802, configured to record the first position of the first data in the main table in a log file during the execution of the first database statement, and the log file also records the first target operation in the first database statement;

[0126] A storage module 803, configured to persistently store the log file.

[0127] In a possible implementation manner, the first position is located in a first compression unit corresponding to the main table, and the log file also records the second position of the second data in the first compression unit;

[0128] The data processing device 200 further includes a data change module 804, and the data change module 804 is configured to:

[0129] Access the first compression unit according to the first position and the second position recorded in the log file to obtain the updated first data and second data;

[0130] Update or delete the first data persistently stored in the database according to the first target operation recorded in the log file;

[0131] Update or delete the second data persistently stored in the database according to the second target operation recorded in the log file.

[0132] In a possible implementation manner, the first target operation includes an update operation, and the first target operation is used to update the first data to third data, and the log file also records the third data.

[0133] In a possible implementation manner, the first target operation includes an update operation, and the update operation includes a pre-update identifier and a post-update identifier. The pre-update identifier is associated with the first position, and the post-update identifier is associated with the third data.

[0134] In a possible implementation manner, the execution module 802 is further configured to:

[0135] When successfully submitting the transaction corresponding to the first database statement, establish a mapping relationship between the identifier of the transaction and the global transaction commit sequence number CSN;

[0136] In the log file, the CSN corresponding to the identifier of the third data record transaction.

[0137] In a possible implementation, the CSN is used to indicate the order of synchronizing the first target operation, the first data, and the third data to the target storage area.

[0138] In a possible implementation, the log file records a first location, a first target operation, a second location, and a second target operation, where the first location and the second location are located in a first compression unit corresponding to the main table;

[0139] The data processing device 200 further includes a synchronization module 805, which is used for:

[0140] Access the first compression unit according to the first location and the second location recorded in the log file to obtain the updated first data and second data;

[0141] Synchronize the first target operation, the first data, the second target operation, and the second data to the target storage area according to the CSN corresponding to the first data and the CSN corresponding to the second data recorded in the log file.

[0142] In a possible implementation, the first location includes the identifier of the first compression unit where the first data is located and the row number of the first data in the first compression unit.

[0143] Figure 8 The data processing device 200 shown corresponds to Figure 3 the data processing method shown, so Figure 8 For the specific implementation manner of the data processing device 200 shown and the technical effects it has, reference can be made to the relevant descriptions in the above Figure 3 shown embodiments, which will not be elaborated here.

[0144] The above Figures 1 to 8 In the shown embodiment, the data processing device 200 can be software configured on a computing device or a computing device cluster, and by running this software on the computing device or the computing device cluster, the computing device or the computing device cluster can implement the functions of the data processing device 200. Next, from the perspective of implementing based on hardware devices, the data processing device 200 will be introduced in detail.

[0145] Figure 9 A schematic structural diagram of a computing device is shown. The above data processing device 200 can be deployed on this computing device. The computing device can be a computing device (such as a server) in a cloud environment or a computing device in an edge environment, etc., and can specifically be used to implement the functions of the data processing device 200 in the above Figure 3 shown embodiment.

[0146] As shown Figure 9 in FIG. 1, the computing device 900 includes a processor 910, a memory 920, a communication interface 930, and a bus 940. The processor 910, the memory 920, and the communication interface 930 communicate with each other via the bus 940. The bus 940 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 9 only a thick line is used to represent it in FIG. 1, but it does not mean that there is only one bus or one type of bus. The communication interface 930 is used for external communication, such as obtaining a first image, outputting a second image, etc.

[0147] Among them, the processor 910 can be a central processing unit (CPU), an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits. The processor 910 can also be an integrated circuit chip with the ability to process signals. In the implementation process, the functions of each device in the data processing device 200 can be completed by the integrated logic circuit in the hardware of the processor 910 or instructions in software form. The processor 910 can also be a general-purpose processor, a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, which can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 920, and the processor 910 reads the information in the memory 920 and combines its hardware to complete some or all of the functions in the data processing device 200.

[0148] The memory 920 may include a volatile memory, such as a random access memory (RAM), or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a HDD, or a SSD.

[0149] The memory 920 stores executable codes, and the processor 910 executes the executable codes to execute the method executed by the aforementioned data processing device 200 .

[0150] Specifically, in implementing Figure 3 In the case of the embodiment shown, and Figure 3 In the case where the data processing device 200 described in the illustrated embodiment is implemented by software, the execution Figure 3 The software or program code required for the functions of the data processing device 200 is stored in the memory 920. The interaction between the data processing device 200 and other devices is realized through the communication interface 930. The processor is used to execute the instructions in the memory 920 to implement the method executed by the data processing device 200.

[0151] Figure 10 A schematic diagram of the structure of a computing device cluster is shown. Figure 10 The computing device cluster 100 shown includes multiple computing devices, and the data processing apparatus 200 can be deployed on multiple computing devices in the computing device cluster 100 in a distributed manner. Figure 10 As shown, the computing device cluster 100 includes multiple computing devices 1000 , each computing device 1000 includes a memory 1020 , a processor 1010 , a communication interface 1030 and a bus 1040 , wherein the memory 1020 , the processor 1010 , and the communication interface 1030 are communicatively connected to each other via the bus 1040 .

[0152] The processor 1010 can be a CPU, GPU, ASIC, or one or more integrated circuits. The processor 1010 can also be an integrated circuit chip with signal processing capabilities. During implementation, some functions of the data processing device 200 can be completed through the integrated logic circuit in the hardware of the processor 1010 or instructions in software form. The processor 1010 can also be a DSP, FPGA, general-purpose processor, other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, which can implement or execute some of the methods, steps, and logic block diagrams disclosed in the embodiments of this application. Among them, the general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory 1020. In each computing device 1000, the processor 1010 reads the information in the memory 1020 and can complete some functions of the data processing device 200 in combination with its hardware.

[0153] The memory 1020 can include ROM, RAM, static storage devices, dynamic storage devices, hard disks (such as SSD, HDD), etc. The memory 1020 can store program codes, for example, program codes for implementing some or all of the functions of the data processing device 200. For each computing device 1000, when the program code stored in the memory 1020 is executed by the processor 1010, the processor 1010 executes some of the methods executed by the data processing device 200 based on the communication interface 1030. The memory 1020 can also store data, such as: intermediate data or result data generated during the execution of the processor 1010, such as the above-mentioned data 1, data 2, log files, etc.

[0154] The communication interface 1030 in each computing device 1000 is used for external communication, such as interacting with other computing devices 1000.

[0155] The bus 1040 can be a peripheral component interconnect standard bus or an extended industry standard architecture bus, etc. For ease of representation, Figure 10 the bus 1040 in each computing device 1000 is only represented by a thick line, but it does not mean that there is only one bus or one type of bus.

[0156] A communication path is established between the above-mentioned multiple computing devices 1000 through a communication network to implement the functions of the data processing device 200. Any computing device can be a computing device in a cloud environment (for example, a server) or a computing device in an edge environment.

[0157] In addition, an embodiment of the present application further provides a computer-readable storage medium storing instructions that, when run on one or more computing devices, cause the one or more computing devices to execute the method performed by the data processing device 200 in the above embodiment.

[0158] In addition, an embodiment of the present application further provides a computer program product that, when executed by one or more computing devices, causes the one or more computing devices to execute any of the foregoing data processing methods. The computer program product may be a software installation package. In the case where any of the foregoing data processing methods is required, the computer program product may be downloaded and executed on a computer.

[0159] It should be further noted that the system embodiments described above are merely illustrative. The devices described as separate components may or may not be physically separated, and the components shown as devices may or may not be physical devices, that is, they may be located in one place or distributed to multiple network devices. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the system embodiments provided in the present application, the connection relationships between the devices indicate that they have communication connections, which may be specifically implemented as one or more communication buses or signal lines.

[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general-purpose hardware. Of course, it can also be implemented by means of dedicated hardware, including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, in more cases, software program implementation is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, including several instructions for causing a computer device (which may be a personal computer, a training device, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0161] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0162] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. A data processing method, characterized in that: The method comprises: Obtain a first database statement, where the first database statement is used to perform a first target operation on first data in a main table, where the first target operation includes an update operation or a delete operation, wherein the database includes the main table, and the main table stores data in columns; During the execution of the first database statement, recording the first data at a first position in the main table in a log file, wherein the log file also records the first target operation in the first database statement; The log file is persistently stored.

2. The method according to claim 1, characterized in that The first position is located in a first compression unit corresponding to the main table, and the log file further records a second position of the second data in the first compression unit; The method further comprises: Access the first compression unit according to the first position and the second position recorded in the log file to obtain the updated first data and the second data; updating or deleting the first data persistently stored in the database according to the first target operation recorded in the log file; According to the second target operation recorded in the log file, the second data persistently stored in the database is updated or deleted.

3. The method according to claim 1 or 2, characterized in that: The first target operation includes an update operation, and the first target operation is used to update the first data to third data. The log file also records the third data.

4. The method according to claim 3, characterized in that The first target operation includes an update operation, and the update operation includes a pre-update identifier and a post-update identifier, the pre-update identifier is associated with the first position, and the post-update identifier is associated with the third data.

5. The method according to claim 3 or 4, characterized in that: The method further comprises: When the transaction corresponding to the first database statement is successfully submitted, a mapping relationship between the transaction identifier and a global transaction submission sequence number CSN is established; In the log file, the CSN corresponding to the identifier of the transaction is recorded for the third data.

6. The method according to claim 5, characterized in that The CSN is used to indicate the order of synchronizing the first target operation, the first data, and the third data to the target storage area.

7. The method according to claim 6, characterized in that The log file records the first position, the first target operation, the second position, and the second target operation, the first position and the second position are located in a first compression unit corresponding to the main table, and the method further includes: Access the first compression unit according to the first position and the second position recorded in the log file to obtain the updated first data and the second data; The first target operation, the first data, the second target operation, and the second data are synchronized to the target storage area according to the CSN corresponding to the first data and the CSN corresponding to the second data recorded in the log file.

8. The method according to any one of claims 1 to 7, characterized in that: The first position includes an identifier of a first compression unit where the first data is located, and a row number of the first data in the first compression unit.

9. A data processing device, characterized in that: The device comprises: an acquisition module, configured to acquire a first database statement, wherein the first database statement is used to perform a first target operation on first data in a main table, wherein the first target operation includes an update operation or a delete operation, wherein the database includes the main table, and the main table stores data in columns; an execution module, configured to record the first data at a first position in the main table in a log file during the execution of the first database statement, wherein the log file also records the first target operation in the first database statement; The storage module is used to persistently store the log file.

10. The device according to claim 9, characterized in that The first position is located in a first compression unit corresponding to the main table, and the log file further records a second position of the second data in the first compression unit; The device further comprises a data changing module, wherein the data changing module is used to: Access the first compression unit according to the first position and the second position recorded in the log file to obtain the updated first data and the second data; updating or deleting the first data persistently stored in the database according to the first target operation recorded in the log file; According to the second target operation recorded in the log file, the second data persistently stored in the database is updated or deleted.

11. The device according to claim 9 or 10, characterized in that The first target operation includes an update operation, and the first target operation is used to update the first data to third data. The log file also records the third data.

12. The device according to claim 11, characterized in that The first target operation includes an update operation, and the update operation includes a pre-update identifier and a post-update identifier, the pre-update identifier is associated with the first position, and the post-update identifier is associated with the third data.

13. The device according to claim 11 or 12, characterized in that The execution module is further used for: When the transaction corresponding to the first database statement is successfully submitted, a mapping relationship between the transaction identifier and a global transaction submission sequence number CSN is established; In the log file, the CSN corresponding to the identifier of the transaction is recorded for the third data.

14. The device according to claim 13, characterized in that The CSN is used to indicate the order of synchronizing the first target operation, the first data, and the third data to the target storage area.

15. The device according to claim 14, characterized in that The log file records the first position, the first target operation, the second position, and the second target operation, wherein the first position and the second position are located in a first compression unit corresponding to the main table; The device also includes a synchronization module, which is used to: Access the first compression unit according to the first position and the second position recorded in the log file to obtain the updated first data and the second data; The first target operation, the first data, the second target operation, and the second data are synchronized to the target storage area according to the CSN corresponding to the first data and the CSN corresponding to the second data recorded in the log file.

16. The device according to any one of claims 9 to 15, characterized in that The first position includes an identifier of a first compression unit where the first data is located, and a row number of the first data in the first compression unit.

17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the computing device cluster executes the method according to any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on at least one computing device, enable the at least one computing device to perform the method according to any one of claims 1 to 8.

19. A computer program product comprising instructions which, when executed on at least one computing device, cause the at least one computing device to perform the method as claimed in any one of claims 1 to 8.