Data processing method, data compression method and device
By determining the version status of multiple associated sets and determining the data reading version, the problem of data inconsistency in multiple versions of reporting systems is solved, and storage costs and query delays are reduced through data compression, efficient data processing and storage are achieved.
Patent Information
- Application Number
- CN202011245053.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-10
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-11-10
AI Technical Summary
In a multi-version reporting system, due to the successive update times of tables of different dimensions, data inconsistencies may occur when reading data. In addition, the existing data compression method will generate unnecessary redundant data when the data writing task is stable, wasting storage space and increasing query delay.
Data consistency between multiple associated sets is ensured by determining the state of each version of multiple versions of multiple associated sets and determining the data read version based on the state of each version. At the same time, based on the relationship between the data reading version and different versions, the data of different versions is compressed, and the necessary version data is retained to reduce redundancy.
It realizes the data consistency between different related sets when reading data, and while ensuring the correctness of data reading, it reduces storage costs and reduces query delays.
Smart Images

Figure CN114461651B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and more particularly to a data processing method, a data compression method, and corresponding devices, apparatuses, and computer-readable storage media. Background Art
[0002] Nowadays, whether in commercial, scientific research, or personal applications, it is often necessary to process a large amount of data. Usually, a reporting system is used to store, manage, update, and query this large amount of data. The reporting system can be stored, for example, in a database built on a server. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication means, which is not limited in this application.
[0003] In a reporting system, to facilitate data query, the bottom-layer data can be aggregated according to different dimensions, and a table is created for each dimension, so that the corresponding data can be directly read from the tables of different dimensions without performing real-time calculation during data reading, thereby accelerating the data query speed. However, since the update times of the tables of different dimensions are different, data inconsistency may occur between the tables of different dimensions during data reading. In addition, in a multi-version reporting system, since the number of data versions that the system can store is limited, usually, the maximum number of versions N that each piece of data can store is set. When the number of data versions exceeds N, the data is compressed to clear some versions. The commonly used data compression method is to retain the latest N versions and clear other versions. However, when the data write task execution is stable, this method will generate unnecessary redundant data, waste storage space, and increase query latency. Summary of the Invention
[0004] To solve the above problems, the present disclosure provides a data processing method, a data compression method, and corresponding devices, apparatuses, and computer-readable storage media for an associated set with multiple versions.
[0005] According to one aspect of the present disclosure, there is provided a data processing method for multiple associated sets with multiple versions, including: determining the status of each version among the multiple versions of the multiple associated sets, where the multiple associated sets are multiple mutually associated sets obtained by aggregating and calculating the same data set according to different dimensions, the multiple versions have different creation times, and for each creation time, each associated set among the multiple associated sets corresponds to the same version, and the status of each version indicates the completion of the write operation corresponding to that version; determining a data read version of the multiple associated sets based on the status of each version among the multiple versions, where the data read version is the maximum version that satisfies data consistency among the multiple associated sets; and performing data reading on the multiple associated sets based on the data read version.
[0006] According to an example of the present disclosure, determining the status of each version among the multiple versions of the multiple associated sets includes: determining one or more write values of the multiple associated sets based on input values of at least one data in the data set; setting the status of the version to a first status; performing a write operation on the multiple associated sets using the one or more write values to obtain the multiple associated sets of the version; and updating the status of the version to a second status or a third status based on the completion of the write operation.
[0007] According to an example of the present disclosure, determining one or more write values of the multiple associated sets includes: performing an aggregation calculation on the multiple associated sets using the input values of the at least one data to determine one or more write values of the multiple associated sets.
[0008] According to an example of the present disclosure, updating the status of the version to a second status or a third status based on the completion of the write operation includes: when the write operation is completed, updating the status of the version to the second status; and when the write operation fails, performing a restart write operation corresponding to the one or more write values, and when the restart write operation is completed, updating the status of the version to the third status, where the write operation and the restart write operation correspond to different versions.
[0009] According to an example of the present disclosure, when the status of the version is the third status, the method further includes: repairing the status of the version to the second status based on the status of one or more versions having smaller version numbers than the version and the completion of one or more versions between the version and the version corresponding to the restart write operation.
[0010] According to an example of the present disclosure, when one or more versions having a smaller version number than the version are in the second state or the third state, and one or more versions between the version and the version corresponding to the restart write operation are in the second state or the third state, the state of the version is repaired to the second state.
[0011] According to an example of the present disclosure, determining the data read version of a plurality of associated sets includes: determining the version that is the last consecutive second state among the plurality of versions as the data read version of the plurality of associated sets.
[0012] According to an example of the present disclosure, determining the data read version of a plurality of associated sets includes: determining the version that is the last consecutive second state among the plurality of versions as the first data read version; determining the version with the largest version number among the plurality of versions as the second data read version; and determining the data read version based on the first data read version and the second data read version.
[0013] According to an example of the present disclosure, determining the data read version based on the first data read version and the second data read version includes: determining the data read version as the first data read version when the difference between the version number of the second data read version and the version number of the first data read version is less than a predetermined threshold; and determining the data read version as the second data read version when the difference between the version number of the second data read version and the version number of the first data read version is greater than or equal to the predetermined threshold.
[0014] According to another aspect of the present disclosure, there is provided a data compression method for a plurality of associated sets having a plurality of versions, including: determining the state of each version among the plurality of versions of the plurality of associated sets, where the plurality of associated sets are a plurality of mutually associated sets obtained by aggregating and calculating the same data set according to different dimensions, the plurality of versions have different creation times, and for each creation time, each associated set among the plurality of associated sets corresponds to the same version, and the state of each version indicates the completion status of the write operation corresponding to the version; determining the data read version of the plurality of associated sets based on the state of each version among the plurality of versions, where the data read version is the largest version that satisfies data consistency among the plurality of associated sets; and compressing the data of different versions based on the relationship between the data read version and the different versions of the data of the plurality of associated sets that make up the plurality of versions.
[0015] According to an example of the present disclosure, determining the status of each version among the multiple versions of the associated set includes: determining one or more write values of the multiple associated sets based on input values of at least one data among the multiple data sets; setting the status of the version to a first status; performing a write operation on the multiple associated sets using the one or more write values to obtain multiple associated sets of the version; and updating the status of the version to a second status or a third status based on the completion of the write operation.
[0016] According to an example of the present disclosure, determining a data read version of multiple associated sets includes: determining the last consecutive version among the multiple versions that is in the second status as the data read version.
[0017] According to an example of the present disclosure, compressing data of different versions based on the relationship between the data read version and different versions of data constituting the multiple associated sets includes: retaining data of versions having a larger version number compared to the data read version; and retaining data of the version having the largest version number with a version number less than or equal to the data read version.
[0018] According to another aspect of the present disclosure, there is provided a data processing apparatus for multiple associated sets having multiple versions, including: a status determination unit configured to determine the status of each version among the multiple versions of the multiple associated sets, where the multiple associated sets are multiple mutually associated sets obtained by aggregating and calculating the same data set according to different dimensions, the multiple versions have different creation times, and for each creation time, each associated set among the multiple associated sets corresponds to the same version, and the status of each version indicates the completion of the write operation corresponding to the version; and a data read unit configured to determine a data read version of the multiple associated sets based on the status of each version among the multiple versions, where the data read version is the largest version that satisfies data consistency among the multiple associated sets, and perform data reading on the multiple associated sets based on the data read version.
[0019] According to another aspect of the present disclosure, there is provided a data compression device for multiple associated sets having multiple versions, including: a status determination unit configured to determine the status of each version among the multiple versions of the multiple associated sets, where the multiple associated sets are multiple mutually associated sets obtained by aggregating and calculating the same data set according to different dimensions, the multiple versions have different creation times, and for each creation time, each associated set among the multiple associated sets corresponds to the same version, and the status of each version indicates the completion status of the write operation corresponding to this version; a compression unit configured to determine a data read version of the multiple associated sets based on the status of each version among the multiple versions, where the data read version is the maximum version that satisfies data consistency among the multiple associated sets, and compress the data of different versions based on the relationship between the data read version and different versions of the data constituting the multiple associated sets.
[0020] According to another aspect of the present disclosure, there is also provided a data processing device for an associated set having multiple versions, including: one or more processors; and one or more memories, where computer-readable code is stored in the memories, and when the computer-readable code is run by the one or more processors, the one or more processors are caused to execute the data processing method or data compression method described in each of the above aspects.
[0021] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium having instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the data processing method or data compression method described in each of the above aspects.
[0022] According to another aspect of the present disclosure, there is also provided a computer program product or computer program, where the computer program product or computer program includes computer-readable instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer-readable instructions from the computer-readable storage medium, and when the processor executes the computer-readable instructions, the computer device is caused to execute the data processing method or data compression method described in each of the above aspects.
[0023] By using the data processing method, apparatus, and device according to the various aspects of the present disclosure above, by determining the status of each version among multiple versions of multiple associated sets and determining the data reading version of the multiple associated sets based on the status of each version, it can be ensured that when the data reading version is used to read data from the multiple associated sets, the data read from different associated sets is consistent. Additionally, through the conversion logic of the version from the first state to the third state and from the third state to the second state, the data processing method, apparatus, and device according to the embodiments of the present disclosure can also ensure the correctness and consistency of data reading when the data writing operation restarts. Moreover, by using the data compression method, apparatus, and device for multiple associated sets with multiple versions according to the various aspects of the present disclosure above, by compressing the data of different versions based on the relationship between the data reading version and the different versions of the data constituting the multiple associated sets, it is possible to reduce the storage cost and query latency while ensuring the correctness of data reading. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] By describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. They are used together with the embodiments of the present disclosure to explain the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0025] Figure 1 Shows different schemes for data aggregation in an advertising report system according to an example of the present disclosure;
[0026] Figure 2 Shows the data inconsistency problem of data tables in different dimensions of an advertising report system according to an example of the present disclosure;
[0027] Figure 3 Shows an example scenario of the MVCC mechanism according to an example of the present disclosure;
[0028] Figure 4 Shows the overall architecture of the data processing method for multiple associated sets with multiple versions according to an embodiment of the present disclosure;
[0029] Figure 5 Shows the flowchart of the data processing method 500 for multiple associated sets with multiple versions according to an embodiment of the present disclosure;
[0030] Figure 6 Shows the flowchart of the example data processing method 500 according to an embodiment of the present disclosure;
[0031] Figure 7Shows the input value of at least one data according to an example of an embodiment of the present disclosure;
[0032] Figure 8 Shows the data writing process according to an example of an embodiment of the present disclosure;
[0033] Figures 9A - 9D Shows the conversion of a version from a first state to a second state according to an example of an embodiment of the present disclosure;
[0034] Figure 10 Shows the conversion of a version from a first state to a third state according to an example of an embodiment of the present disclosure;
[0035] Figures 11A - 11B Shows the conversion of a version from a third state to a second state according to an example of an embodiment of the present disclosure;
[0036] Figure 12 Shows the states of each version of a certain data according to an example of an embodiment of the present disclosure;
[0037] Figure 13 Shows the read and write processes of a data processing method according to an example of an embodiment of the present disclosure;
[0038] Figure 14 Shows the data compression process according to an example of the present disclosure;
[0039] Figure 15 Shows the flowchart of a data compression method for multiple associated sets with multiple versions according to an embodiment of the present disclosure;
[0040] Figure 16 Shows an example process of a data compression method according to an example of an embodiment of the present disclosure;
[0041] Figure 17 Shows the structural schematic diagram of a data processing device 1700 according to an embodiment of the present disclosure;
[0042] Figure 18 Shows the structural schematic diagram of a data compression device 1800 according to an embodiment of the present disclosure;
[0043] Figure 19 Shows the schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure. Detailed implementation
[0044] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0045] In a reporting system, the bottom-layer data is unaggregated detailed data. The reporting system can aggregate these detailed data to different extents for output to external customers or for internal operations. To display data in different dimensions, there are generally two aggregation schemes: (1) Create a table with the lowest dimension and write all the data aggregated according to the lowest dimension into this table. When querying high-dimensional data, obtain the high-dimensional data through real-time calculation; (2) Create a table for each dimension, that is, aggregate the same data according to different dimensions and write them into the corresponding dimension tables respectively. The following takes the advertising reporting system as an example to specifically illustrate these two schemes.
[0046] For example, in the advertising reporting system, it is possible to perform a division from high dimension to low dimension according to account, promotion plan, advertisement, and creative. Usually, one account can contain multiple promotion plans, one promotion plan can have multiple advertisements, and one advertisement can have multiple creatives. Figure 1 The differences between the above two schemes are shown by taking the account and advertisement dimensions as examples. Figure 1 Different schemes for data aggregation in the advertising reporting system according to the examples of the present disclosure are shown, where UID represents the account, AID represents the advertisement, and C represents the cost data of the advertisement.
[0047] As Figure 1 shown, in Scheme 1, there is only the advertisement dimension table AID_LEVEL, that is, the cost data in the original log is aggregated according to the advertisement dimension and written into this table. When reading the costs of each advertisement, it can be directly queried from this table. For example, it can be directly queried that the cost of advertisement A1 is 30; but if it is necessary to read the account cost of a higher dimension, it is necessary to aggregate the advertisement costs in this table by account. For example, if querying the cost of account U1, it is necessary to perform real-time calculation on the costs of advertisements A1 and A2 included in account U1, for example, obtain the cost of account U1 as 80 through summation calculation (SUM). It can be seen that although Scheme 1 has lower calculation and storage costs during data aggregation, its calculation cost during data query is very high, resulting in an increase in query latency.
[0048] To solve the problem of data query latency, Solution 2 can be adopted. In Solution 2, a table is created for each of the account dimension and the advertisement dimension. That is, after aggregating the cost data in the original log according to the account dimension and the advertisement dimension respectively, the data is written into the corresponding tables. In this way, whether querying the advertisement cost of a low dimension or the account cost of a high dimension, it can be directly queried from the corresponding advertisement dimension table AID_LEVEL or account dimension table UID_LEVEL without real-time calculation. Solution 2 greatly reduces the data query latency, and for the increased aggregation calculation cost and storage cost, it will not cause trouble for the currently widely used database technologies. For example, the reporting system applying Solution 2 can be stored in a database built on a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0049] However, for Solution 2, due to the different update times of the tables of different dimensions, there may be a problem of data inconsistency between the tables of different dimensions when reading data. Taking the Figure 2 advertisement reporting system as an example for illustration. Figure 2 shows the data inconsistency problem between the data tables of different dimensions of the advertisement reporting system according to the examples of the present disclosure. As Figure 2 shown, at time T1, that is, before the data update occurs, the sum of the cost C of the advertisement dimension table and the account dimension table is both 10. At time T2, first update the advertisement dimension table and update the cost C of advertisement A1 to 5. At this time, the sum of the cost C of the advertisement dimension table is 12, while the cost of the account dimension table has not been updated yet and is still 10, which is inconsistent with the data of the advertisement dimension table. At time T3, update the account dimension table and update the cost C of account U1 to 12. At this time, the data of the account dimension table and the advertisement dimension table is restored to be consistent. During the time interval from T2 to T3, the cost data of the advertisement dimension table and the account dimension table is inconsistent, resulting in inconsistent data read during data query.
[0050] To solve the data inconsistency problem of Solution 2, the present disclosure proposes a data processing method for multiple associated sets with multiple versions. In the method of the present disclosure, a multi-version concurrency control (MVCC) mechanism is introduced. First, a brief introduction to the MVCC mechanism is given below.
[0051] The MVCC mechanism supports multi - version storage of reports and enables concurrent read - write access to data. In MVCC, each written data corresponds to a version number. For example, the version number corresponding to a write operation can be represented by WP (the corresponding write operation may not be completed), and the maximum version number RP that can be currently read to satisfy data consistency is used, that is, starting from the first version, it is the version number of the last "continuously completed" version. In MVCC, before the data write starts, WP is incremented first, and then the corresponding write operation is performed. Only when the write operation corresponding to this WP is completed, RP is advanced. When reading data, the current RP is used to read the data, and one or more versions of data with version numbers less than or equal to RP are returned. Here, the number of versions returned during data reading can be set as needed. In the following, by way of example and not limitation, unless otherwise specified, the example of returning the data of the maximum version with a version number less than or equal to RP is used for illustration. Take Figure 3 as an example to illustrate the MVCC mechanism, Figure 3 shows an example scenario of the MVCC mechanism according to an example of the present disclosure. Figure 3 shows 4 versions W1, W2, W3, W4 of X, where W1, W2, and W4 are completed, and W3 is not completed (shown in Figure 3 with gray shading), then at this time RP is the last continuously completed version W2. Using RP = W2 to query data, the value of X can be read as 2.
[0052] First, the overall architecture of a data processing method for multiple associated sets with multiple versions according to an embodiment of the present disclosure is described in combination with Figure 4 Here, multiple associated sets refer to multiple mutually associated sets obtained by aggregating the same data set according to different dimensions. For example, in the example shown in Figure 1 , after aggregating the original log data according to the advertisement dimension and the account dimension respectively, an advertisement - dimension table and an account - dimension table are obtained. Among them, the advertisement - dimension table and the account - dimension table are mutually associated data sets. Each associated set can have multiple different versions created in chronological order. For example, in the example shown in Figure 2 , at time T1, the advertisement - dimension table and the account - dimension table can have the version number W1, and when updates start at time T2, the version numbers of the advertisement - dimension table and the account - dimension table can be W2.
[0053] As Figure 4As shown, the data processing method according to an embodiment of the present disclosure may include a plurality of write tasks 410 and a plurality of read tasks 420. Among them, the global status table is a global table for storing the status of each WP version. In the present disclosure, the status of WP may include, for example, a start (S) state, a completed (E) state, a restart (R) state, etc., where the restart state R indicates that the current task is overwritten by a subsequent task. In a certain write task 410 among a plurality of associated sets, first, in step 411, set the version number WP corresponding to the write task, and for example, set the status of the WP version to the start state S to mark the start of the WP version, and write the status of WP into the global status table. For example, a new version number WP can be set by incrementing on the initial WP. Then, in step 412, perform the write operation corresponding to the write task 410 to update the plurality of associated sets. Subsequently, in step 413, update the status of the WP according to the completion of the write operation. For example, the status of the WP can be updated to the completed E or restart R, and the updated WP status is stored in the global status table. In a certain read task 420, first, in step 421, read the status of each WP version of the plurality of associated sets from the global status table, and then in step 422, perform a data query according to the read WP status to complete the data reading.
[0054] The following combines Figure 5 to describe a data processing method for multiple associated sets with multiple versions according to an embodiment of the present disclosure. Figure 5 FIG. shows a flowchart of a data processing method 500 for multiple associated sets with multiple versions according to an embodiment of the present disclosure.
[0055] As Figure 5 shown, in step S510, determine the status of each version among the multiple versions of the multiple associated sets. As described above, the multiple versions may be different versions of the multiple associated sets created in different chronological orders. That is, the multiple versions have different creation times, and for each creation time, each associated set among the multiple associated sets corresponds to the same version. The status of each version indicates the completion of the write operation corresponding to the version. The status may include, for example, a first state, a second state, and a third state. Among them, the first state may be, for example, a start state S indicating the start of the write operation of the version, the second state may be, for example, a completed state E indicating that the write operation of the version has been completed, and the third state may be, for example, a restart state R indicating the restart of the write operation of the version. However, the present disclosure is not limited thereto, and the first, second, and third states may also be set to other states as needed.
[0056] In step S520, based on the status of each version among multiple versions, a data reading version of multiple associated sets is determined. The data reading version can be the maximum version that satisfies consistency and can currently be read by multiple associated sets, that is, the maximum version for which the data among multiple associated sets remains consistent. The version number of this data reading version can be represented by RP. According to an example of the embodiments of the present disclosure, the version that is the last consecutive version in the second state (for example, the completed state E) among multiple versions can be determined as the data reading version. Then, in step S530, based on this data reading version, data reading is performed on multiple associated sets, which can ensure that the data read from each associated set is consistent.
[0057] The following combines Figures 6 to 8 to describe in detail step S510 of the data processing method 500 according to the embodiments of the present disclosure. Figure 6 FIG. shows a flowchart of a data processing method 500 according to an example of the embodiments of the present disclosure. Figure 7 FIG. shows input values of at least one data according to an example of the embodiments of the present disclosure. Figure 8 FIG. shows a data writing process according to an example of the embodiments of the present disclosure.
[0058] As Figure 6 shown, in step S511, based on the input values of at least one data in the data set, one or more write values of multiple associated sets are determined. In a certain write task of multiple associated sets, the values of one or more data in the data set can be input. For example, still taking the advertising report system as an example, as Figure 7 shown, at the data time (i.e., the time of data input) D = 10:00 of a certain write task, the newly added cost of advertisement A1 of account U1 is 0.5, and the newly added cost of advertisement A2 is 0.5 are respectively input. At this time, it is necessary to update multiple associated sets of the data set by using the input values of the at least one data, that is, write the data changes into multiple associated sets. For example, in Figure 7 , after the newly added costs of advertisements A1 and A2 are input, the advertisement dimension table and the account dimension table both need to be updated accordingly. The at least one data input value can be used to perform an aggregation calculation on multiple associated sets to determine one or more write values of multiple associated sets.
[0059] The following takes the account dimension table as an example to illustrate the process of determining one or more write values of the associated set. As Figure 8 shown, in the first step, the input data at this data time D can be initially aggregated to obtain an intermediate result Y. For example, for Figure 7The newly added costs of advertisements A1 and A2 are initially aggregated to obtain the newly added cost 1 of account U1 as the intermediate result Y. Assuming the historical data of the account dimension table is X, at this time, in step ②, a snapshot of the historical data X and the intermediate result Y is taken, and the full amount of data for this snapshot is {X, Y}. This full amount of data is stored and can be locked to ensure data uniqueness and order. The time for taking the snapshot is the creation time of the version corresponding to the current write task, or the version number corresponding to the current write task. For example, in this example, the snapshot time is WP = 10:10, and this WP is the version number corresponding to the current write task. In step ③, the full amount of data is aggregated and calculated to obtain the write value for the associated set. For example, the newly added cost 1 of account U1 in the intermediate result Y is aggregated with the cost 10 of account U1 in the historical data to obtain the write value for the account dimension table as 11.
[0060] After determining one or more write values for multiple associated sets, in step S512, the status of the WP version corresponding to the current write task is set to the first status, for example, set to the start status S (such as Figure 8 step ④ in). To mark the start of the write operation. For example, in the above example, after determining that the write value for account U1 in the account dimension table is 11, the status of the WP version is set to the start status S.
[0061] In step S513, using the determined one or more write values, a write operation is performed on multiple associated sets to obtain multiple associated sets for the WP version corresponding to the current write task. Here, only the incremental data can be output to the associated set, that is, only the values of the data that has changed are output to the associated set (such as Figure 8 step ⑤ in). For example, only the changed write value 11 of account U1 is updated to the account dimension table to obtain the account dimension table for this WP version.
[0062] Then, in step S514, based on the completion status of the write operation of the current write task, the status of the corresponding version is updated to the second status or the third status. As mentioned before, the second status can be, for example, the completed status E indicating that the write operation of this version has been completed, and the third status can be, for example, the restart status R indicating that the write operation of this version is restarted. Specifically, when the write operation is completed, the status of this version can be updated to the second status, for example, the completed status E, such as Figure 8 step ⑥ in; while when the write operation fails, a restart write operation corresponding to the determined one or more write values will be executed.
[0063] The following combines Figures 9A - 9D to describe the conversion of the version from the first status to the second status. Figures 9A - 9DShows the conversion of an example version from a first state to a second state according to an embodiment of the present disclosure.
[0064] As Figure 9A shown, a write task is initiated at data time D = 10:00, the determined update content to be written is account U1 = 1, the corresponding snapshot time, i.e., version WP = W1, and the status of W1 is set to start S. At Figure 9A this time, the write operations for both the advertisement dimension table and the account dimension table are completed, so the status of W1 is set to completed E. At this time, the last version with a consecutive completed status is W1, that is, the data read version that multiple associated sets can currently read and meet consistency is W1, i.e., RP = W1. Using RP = W1 to query data from the advertisement dimension table and the account dimension table, the read account U1 is all 1, that is, the data between these two different dimension tables is consistent.
[0065] As Figure 9B shown, a write task is initiated at data time D = 10:10, the determined update content to be written is account U2 = 2, the corresponding snapshot time, i.e., version WP = W2, and the status of W2 is set to start S. In Figure 9B this example, the update of the account dimension table is completed, while the update of the advertisement dimension table is not completed, that is, at this time, the data of different dimension tables of W2 is inconsistent, and the last version with a consecutive completed status is still W1. Therefore, RP is still W1, that is, when using RP to query data, the data of version W1 is read, so as to ensure that the data read from different tables is consistent. Otherwise, for example, assuming that RP = W2 at this time, it will cause the data of U2 to be read, and the data of U2 whose update is not completed is inconsistent in the advertisement dimension table and the account dimension table.
[0066] As Figure 9C shown, a write task is initiated at data time D = 10:20, the determined update content to be written is account U3 = 3, the corresponding snapshot time, i.e., version WP = W3, and the status of W3 is set to start S. In Figure 9C this time, the write operations for both the advertisement dimension table and the account dimension table are completed, so the status of W3 is set to completed E. At this time, the last version with a consecutive completed status is still W1. Therefore, RP is still W1, that is, when using RP to query data, the data of version W1 is read, so as to ensure that the data read from different tables is consistent. Otherwise, for example, assuming that RP = W2 or W3 at this time, it will cause the data of U2 to be read, and the data of U2 whose update is not completed is inconsistent in the advertisement dimension table and the account dimension table.
[0067] Subsequently, at Figure 9DIn this case, the write operation of W2 is completed, and the status of W2 is set to E. At this time, the last version with a consecutive completed status is W3. Therefore, RP = W3. Regardless of whether the data of U1, U2, or U3 is read, the data read from the account dimension table and the advertisement dimension table is consistent.
[0068] By using the data processing method according to the embodiments of the present disclosure, by determining the status of each version in multiple versions of multiple associated sets and determining the data read version of the multiple associated sets based on the status of each version, it can be ensured that when the data of the multiple associated sets is read using the data read version, the data read from different associated sets is consistent.
[0069] In some cases, when the write task needs to be re-executed for some reason. For example, if the write operation of the write task fails due to a short network disconnection, the write task can be restarted to perform a restart write operation corresponding to one or more determined write values. For example, when performing the data processing of the present disclosure using a real-time stream data processing technology such as Spark streaming, when a write task fails, the stream processing will restart a new write task. The restarted write task still processes the same data, so the corresponding data time remains unchanged, but different data versions are generated. In the present disclosure, the write operation and the restart write operation correspond to the same data time, perform a write operation on the same one or more write values, but correspond to different versions. For example, in Figure 8 the example, the data time of the write task is D = 10:00, and the snapshot time, that is, the version number, is WP = 10:10. If the write operation of the write task fails, a restart write operation will be performed. The data time of the restart write operation is still D = 10:00, but corresponds to a new snapshot time, for example, WP = 10:20. When the write operation fails and the restart write operation is completed, the status of the version corresponding to the write operation can be updated to a third status, such as the restart status R. In the present disclosure, when the status of a certain version is the restart status R, it means that the data of this version has been overwritten by a subsequent version.
[0070] The following specifically illustrates the conversion between the first state and the third state of the version by taking Figure 10 as an example. Figure 10 shows the conversion of the version from the first state to the third state in the example according to the embodiments of the present disclosure, that is, the conversion from the start state S to the restart state R.
[0071] The conversion logic from S to R can be described as "the data of this version has been overwritten by a subsequent version". As Figure 10As shown, a write task is started at data time D1. The determined update content to be written is account U1 = 1, and the corresponding snapshot time, i.e., the version, is W1. The status of W1 is set to start S. Due to some reason, the write operation of W1 fails. At this time, the write task D1 is restarted to perform a restart write operation, and the corresponding version is W3. After the write operation of W3 is completed, the status of W3 is set to E. At this time, W3 can query the global status table storing the status of each WP version and learn that W1 has the same data time as this task (which can be called "same D"), and the status of W1 is set to R (such as Figure 8 step ⑦ in
[0072] ). Thus, it can be marked that the data of version W1 has been overwritten by a subsequent version (such as W3).
[0073] By converting the version from the first state to the third state when the write operation fails and the restart write operation is completed, the failed data version can be marked (for example, marked as R). However, since the data read version is the last consecutive version in the second state, in order to be able to read the correct data written by the restart write operation and ensure that the failed data version is not read, a method of restoring the third state to the second state is also required. Figures 11A - 11B According to an example of an embodiment of the present disclosure, when the write operation of a certain version fails and its restart write operation is successful, the status of this version is set to the third state. At this time, based on the status of one or more versions with a smaller version number compared to this version, and the completion status of one or more versions between this version and the version corresponding to the restart write operation, the status of this version is restored to the second state. Specifically, for this version in the third state, when one or more versions with a smaller version number compared to this version are all in the second state or the third state, and one or more versions between this version and the version corresponding to the restart write operation are all in the second state or the third state, the status of this version is restored to the second state. The following combines
[0074] Figures 11A - 11B shows the conversion of the version from the third state to the second state according to an example of an embodiment of the present disclosure, that is, the conversion from the restart state R to the completed state E. The conversion logic from R to E can be expressed as "there exists an RP that satisfies consistency such that the version in R will not be read". Specifically, assume a certain version W i =R, and the version W i with the same data time as W (i.e., same D) is W n =E (n>i), where W i or W nFor example, it can be any version number such as W1, W2, W3, etc. described above. Here, for the sake of clear expression, the digital serial numbers are represented in the form of subscripts. Then, only when there exists an RP≥W that satisfies consistency n will the data of W i be covered by W n such that the data of W i cannot be read. At this time, the status of W i can be repaired to E.
[0075] In Figures 11A - 11B the example, W1 and W4 are in the completed state E, W2 is in the restart state R and has the same D as W4. As Figure 11A shown, when W3 is in the completed state E, at this time W4 is a legal RP, that is, it is assumed that data reading using RP = W4 can satisfy consistency. Thus, W4 can cover the data of W2, and W2 can be repaired from R to E (as in Figure 8 step ⑧). In Figure 11B , W3 is in an incomplete state, such as S or R. At this time, U2 of the account dimension table and the advertisement dimension table in W2 is inconsistent, U2 and U3 of the account dimension table and the advertisement dimension table in W3 are both inconsistent, U3 of the account dimension table and the advertisement dimension table in W4 is inconsistent, and only the data of the account dimension table and the advertisement dimension table in W1 is consistent, that is, only W1 is a legal RP that satisfies consistency, and RP = W1 < W4. Therefore, the status of W2 cannot be set to E.
[0076] In the above example, the conversion logic from R to E "there exists a RP that satisfies consistency such that the version in the R state will not be read" can be formally described for program judgment. Specifically, if the status of a certain version W i is to be converted from R to E, the following all rules should be satisfied:
[0077] · For any W j ≤W i , the status of the W j version ∈ {R, E}, that is, the status of all versions with version numbers less than or equal to W i is R or E;
[0078] · For any W j ≤W i , let the set of versions with the same D as W j be M, and the version with the largest version number in the set M be W k , then the status of all versions between W i and W k is R or E.
[0079] When all the above rules are satisfied, it indicates that there is a consistent RP such that version W of R i will not be read, so that W i can be set to state E.
[0080] The above combination Figure 10 and Figures 11A - 11B describe the method of converting the version state from the first state to the third state (e.g., from S to R), and from the third state to the second state (e.g., from R to E). Through this conversion method, the correctness and consistency of data reading can be ensured when the write task is restarted.
[0081] Back to Figure 5 , in step S520, based on the state of each version among multiple versions, determine the data reading version of multiple associated sets for data reading of multiple associated sets in step S520. As mentioned above, the last consecutive version in the second state among multiple versions can be determined as the data reading version of multiple associated sets. In the ideal case where the number of versions of each piece of data that the system can store has no upper limit, it can be ensured that when using this data reading version for data reading, the data between multiple associated sets is always consistent. However, in actual situations, the number of versions of each piece of data that the system can store is limited. Therefore, when using a certain data reading version for data reading, it is possible that a certain piece of data of this version is not stored in the system, resulting in a system error. The present disclosure provides a method for determining the data reading version, so that in this case, inconsistent data can be preferentially returned, enabling the user to read the data instead of directly reporting an error.
[0082] Specifically, according to an example of an embodiment of the present disclosure, the last consecutive version in the second state among multiple versions is determined as the first data reading version; the version with the largest version number among multiple versions is determined as the second data reading version; and the data reading version is determined based on the first data reading version and the second data reading version. According to an example of an embodiment of the present disclosure, when the difference between the version number of the second data reading version and the version number of the first data reading version is less than a predetermined threshold, the data reading version is determined to be the first data reading version; and when the difference between the version number of the second data reading version and the version number of the first data reading version is greater than or equal to the predetermined threshold, the data reading version is determined to be the second data reading version. For example, the predetermined threshold can be the maximum version number N of each piece of data of multiple associated sets that the system can store, or the predetermined threshold can also be set according to actual storage requirements. The present disclosure does not make specific limitations on this.
[0083] For example, the version number of the first data reading version is called the logical RP (RP logic), the version number of the second data reading version is called the maximum RP (RP max ), the version number of the finally determined data reading version is called the actual RP (RP real ). Assuming that the predetermined threshold is the maximum number of versions N of each piece of data that the system can store for multiple associated sets constituting multiple versions, the data reading version RP can be determined according to the following rules real :
[0084] · When the difference between RP max and RP logic is less than N, RP real = RP logic ;
[0085] · When the difference between RP max and RP logic is greater than or equal to N, RP real = RP max , that is, equal to the current maximum version number.
[0086] The following specifically describes the above method for determining the data reading version in combination with the examples in Figure 12 . Figure 12 shows the status of each version of a certain piece of data according to the example of the embodiment of the present disclosure. As Figure 12 shown, data X has three versions W1, W2, and W3, where W1 is completed, and W2 and W3 are not completed. At this time, RP max is the maximum version W3, RP logic is the last continuously completed version W1, and the difference between RP max and RP logic is 2. If N = 3, that is, the system can store all three versions of data X, then RP real = RP logic = W1. At this time, using RP real = W1 can read data X that meets the consistency. If N = 2, the system can only store two versions of data X, for example, storing versions W2 and W3. At this time, if RP logic = W1 is used, data X cannot be read. If the above rules are used to determine RP real = RP max = W3, using RP real = W3 for data reading ensures that the latest data can be read.
[0087] According to the example of the embodiment of the present disclosure, after determining the data reading version RP real , for example, a structured query language (SQL) can be used to generate with RP realquery statement to perform data query operations. For example, an open-source cluster computing framework application (Apache Phoenix) can be used to construct a query statement as shown in Table 1 below. Querying using the query statement in Table 1 can return data of L versions where the version number is less than or equal to RP real of the data of L versions, for example, L can be set to 1, that is, only return data of 1 version where the version number is less than or equal to RP real of the data of 1 version. It should be understood that the present disclosure is not limited to using Apache Phoenix as an example here for data query, but any suitable method in the art can be used for data query.
[0088] Table 1
[0089]
[0090] The above combination Figures 5 to 12 has described the data processing method 500 according to the embodiments of the present disclosure. To have a clearer understanding of the data processing method of the present disclosure, the following combines Figure 13 the example read and write processes in to systematically elaborate on the data processing method of the present disclosure.
[0091] Figure 13 shows the read and write processes of the data processing method 500 of the example according to the embodiments of the present disclosure. As Figure 13 shown, the data processing method 500 can include multiple write tasks 1310 and multiple read tasks 1320.
[0092] After any write task 1310 is started, first, in step 1311, set the version number WP corresponding to this write task. The WP version can be set to the start state to mark the start of the WP version, and the status of WP is written into the global status table. For example, a new version number WP can be set by incrementing on the initial WP. Then, in step 1312, perform the write operation corresponding to the write task 1310 to update multiple associated sets. Subsequently, in step 1313, update the status of WP according to the completion of the write operation. For example, the status of WP can be updated to completed E or restarted R, and the updated WP status is stored in the global status table. When this WP is completed, in step 1314, it can query the global status table for versions with the same data time (same D) as it, such as versions that failed to be written before, and set the versions with the same D as it to R. In addition, in step S1315, according to the data processing method of the embodiments of the present disclosure, it can also try to repair the versions with the status of R to E by querying the status of each version before the current WP version from the global status table. For the specific repair method, refer to the content described above about Figures 11A - 11B which will not be elaborated here again.
[0093] After any read task 1320 is started, first, in step 1321, the status of each WP version is read from the global status table. In step 1322, according to each WP status, the logical RP is determined. For example, the version number of the last consecutive completed version can be determined as the logical RP. Then, in step 1323, the actual RP is determined, that is, the final data read version. As described above, for example, the current maximum version number can be determined as the maximum RP, and then the actual RP is determined based on the size relationship between the difference between the maximum RP and the logical RP and a predetermined threshold. For the specific method of determining the actual RP, refer to the rules described above and the content about Figure 12 what has been described, which will not be elaborated here. After determining the actual RP, in step 1324, a query statement with the actual RP is generated, and in step 1325, the data query is executed.
[0094] The data processing method according to the embodiments of the present disclosure is described above. By determining the status of each version in multiple versions of multiple associated sets and determining the data read version of the multiple associated sets based on the status of each version, it can be ensured that when the data read version is used to read data from multiple associated sets, the data read from different associated sets is consistent. In addition, through the conversion logic of the version from the first state to the third state and from the third state to the second state, the data processing method according to the embodiments of the present disclosure can also ensure the correctness and consistency of data reading when the data write operation restarts.
[0095] As mentioned before, in practical applications, the number of versions of the data of multiple associated sets that the system can store is limited. Therefore, usually, the maximum number of versions N that each piece of data can store is set. When the number of data versions exceeds N, the data is compressed to clean up some versions. Usually, when compressing the data, only the latest N versions can be retained, and other versions are cleaned up. For example, in Figure 14 the example shown, it is assumed that N is 3. Before compression, data K1 has 4 versions W1 to W4. After compression, only the latest 3 versions W2 to W4 are retained. However, when the task execution situation is stable, that is, when basically each write task can be successfully completed, if N versions are stored for each piece of data, unnecessary redundant data will be generated, wasting storage space and increasing query latency. Therefore, the present disclosure provides a data compression method for multiple associated sets with multiple versions, which can reduce the storage cost and query latency while ensuring the correctness of data reading.
[0096] Next, in combination with Figure 15 describe the data compression method for multiple associated sets with multiple versions according to the embodiments of the present disclosure. Figure 15The flowchart of data compression method 1500 for multiple associated sets with multiple versions according to an embodiment of the present disclosure is shown.
[0097] As Figure 15 shown, in step S1510, the status of each version among multiple versions of multiple associated sets is determined. As described above, the multiple versions may be different versions of multiple associated sets created in different chronological orders. That is, the multiple versions have different creation times, and for each creation time, each associated set among the multiple associated sets corresponds to the same version. The status of each version indicates the completion status of the write operation corresponding to that version. This status may include, for example, a first status, a second status, and a third status. Among them, the first status may be, for example, a start status S indicating the start of the write operation for that version, the second status may be, for example, a completed status E indicating that the write operation for that version has been completed, and the third status may be, for example, a restart status R indicating the restart of the write operation for that version. However, the present disclosure is not limited thereto, and the first, second, and third statuses may also be set to other statuses as needed.
[0098] According to an example of an embodiment of the present disclosure, each version among multiple versions of multiple associated sets may be obtained, for example, through a process as Figure 6 shown. As Figure 6 shown, in step S511, based on the input values of at least one data in the data set, one or more write values of multiple associated sets are determined. In a certain write task of multiple associated sets, the values of one or more data in the input data set may be input. After determining one or more write values of multiple associated sets, in step S512, the status of the WP version corresponding to the current write task is set to the first status, for example, set to the start status S, to mark the start of the write operation. In step S513, using the determined one or more write values, a write operation is performed on multiple associated sets to obtain multiple associated sets corresponding to the WP version of the current write task. Here, only the incremental data may be output to the associated sets, that is, only the values of the data that have changed are output to the associated sets (such as step ⑤ in Figure 8 ). Then, in step S514, based on the completion status of the write operation of the current write task, the status of the corresponding version is updated to the second status or the third status. As described above, the second status may be, for example, a completed status E indicating that the write operation for that version has been completed, and the third status may be, for example, a restart status R indicating the restart of the write operation for that version. Specifically, when the write operation is completed, the status of that version may be updated to the second status, for example, the completed status E, such as step ⑥ in Figure 8 ; and when the write operation fails, a restart write operation corresponding to the determined one or more write values will be executed.
[0099] Specifically, in some cases, when a write task needs to be re-executed for some reason, for example, the write operation of the write task fails due to a short network disconnection, the write task can be restarted to perform a restart write operation corresponding to one or more determined write values. For example, when performing data processing of the present disclosure using a real-time stream data processing technology such as Spark streaming, when a certain write task fails, the stream processing will restart a new write task. The restarted write task still processes the same data, so the corresponding data time remains unchanged, but different data versions are generated. In the present disclosure, the write operation and the restart write operation correspond to the same data time, perform write operations on the same one or more write values, but correspond to different versions. When the write operation fails and the restart write operation is completed, the status of the version corresponding to the write operation can be updated to a third status, such as the restart status R. When the status of a certain version is the restart status R, it means that the data of this version has been overwritten by a subsequent version.
[0100] By converting the version from the first status to the third status when the write operation fails and the restart write operation is completed, the failed data version can be marked (for example, marked as R). However, since the data read version is the last consecutive version in the second status, in order to be able to read the correct data written by the restart write operation and ensure that the failed data version is not read, a method for restoring the third status to the second status is also required.
[0101] According to an example of an embodiment of the present disclosure, when the write operation of a certain version fails and its restart write operation is successful, the status of this version is set to the third status. At this time, based on the status of one or more versions with a smaller version number compared to this version, and the completion status of one or more versions between the version corresponding to the write operation and the version corresponding to the restart write operation, the status of this version can be restored to the second status. Specifically, for this version in the third status, when one or more versions with a smaller version number compared to this version are all in the second status or the third status, and one or more versions between this version and the version corresponding to the restart write operation are all in the second status or the third status, the status of this version is restored to the second status.
[0102] Through the conversion logic of converting the version status from the first status to the third status (for example, from S to R), and from the third status to the second status (for example, from R to E), the correctness and consistency of data reading can be ensured when the write task is restarted. For the specific conversion method, refer to the content combined with Figure 10 and Figures 11A - 11B described above, which will not be elaborated here.
[0103] Back to Figure 15, in step S1520, based on the status of each version among multiple versions, determine the data reading version of multiple associated sets. The data reading version can be the maximum version that satisfies consistency and can be currently read by multiple associated sets, that is, the maximum version for which the data among multiple associated sets remains consistent, and the version number of this data reading version can be represented by the logical RP (RP logic ). According to an example of an embodiment of the present disclosure, the version that is the last consecutive one in the second state (for example, the completed state E) among multiple versions can be determined as the data reading version.
[0104] In step S1530, based on the relationship between the data reading version and different versions of the data constituting multiple associated sets, compress the data of different versions. Specifically, multiple associated sets of multiple versions are composed of data of different versions. For each piece of data of different versions, only the data of the version with a larger version number compared to the data reading version is retained; and, the data of the version with the largest version number less than or equal to the version number of the data reading version is retained. For example, if each piece of data has multiple versions greater than RP logic , then all these versions are retained because these versions may be read subsequently; if each piece of data has multiple versions less than or equal to RP logic , then only the version with the largest version number among these versions is retained. For example, if W i is the largest among these versions, then only W i is retained, so that only W i can be read, and the remaining versions are cleared and thus not read.
[0105] The following is an explanation of the data compression method 1500 according to an embodiment of the present disclosure in conjunction with Figure 16 . Figure 16 shows an example process of the data compression method 1500 according to an example of an embodiment of the present disclosure. As Figure 16 described, there are 4 versions of the associated set, among which W1, W2, and W4 are completed, and W3 is not completed, then RP logic = W2. Before data compression, data K1 has 4 versions W1 to W4, data K2 has 3 versions W1, W3, and W4, and data K3 has two versions W1 and W2. When compressing using the data compression method 1500 according to an embodiment of the present disclosure, for data K1, versions W3 and W4 with version numbers greater than RP logic = W2 are both retained, while for versions W1 and W2 with version numbers less than RP logic = W2, only the largest version W2 is retained; similarly, for data K2, versions W3 and W4 with version numbers greater than RP logic = W2 are retained, and for versions with version numbers less than RP logic= The maximum version W1 of W2; for data K3, retain the version number less than RP logic = The version W1 of W2 and the maximum version W2 in W1 and W2, thereby obtaining the compressed data K1, K2, and K3. In this way, when using RP logic = W2 to read data from the associated set after data compression, both data K1 and K3 can return the data of version W2, while data K2 can return the data of a version less than RP logic The maximum version W1 of the data.
[0106] Using the data compression method for multiple associated sets with multiple versions according to the embodiments of the present disclosure, by compressing data of different versions based on the relationship between the data read version and different versions of the data constituting the multiple associated sets, it is possible to reduce the storage cost and query latency while ensuring the correctness of data reading.
[0107] The following is combined with Figure 17 Describe the data processing device according to the embodiments of the present disclosure. Figure 17 FIG. shows a schematic structural diagram of a data processing device 1700 according to an embodiment of the present disclosure. Since the data processing device 1700 is the same as the details of the data processing method 500 described above in combination with Figure 5 For simplicity, the detailed description of the same content is omitted here. As Figure 17 Shown, the data processing device 1700 includes a status determination unit 1710 and a data reading unit 1720. In addition to these two units, the data processing device 1700 may further include other components. However, since these components are not related to the content of the embodiments of the present application, their illustrations and descriptions are omitted here.
[0108] The status determination unit 1710 is configured to determine the status of each version among multiple versions of multiple associated sets. As described above, the multiple versions may be different versions of multiple associated sets created in different chronological orders. That is, the multiple versions have different creation times, and for each creation time, each associated set among the multiple associated sets corresponds to the same version. The status of each version indicates the completion status of the write operation corresponding to that version. The status may include, for example, a first status, a second status, and a third status. Among them, the first status may be, for example, a start status S indicating the start of the write operation of that version, the second status may be, for example, a completed status E indicating that the write operation of that version has been completed, and the third status may be, for example, a restart status R indicating the restart of the write operation of that version. However, the present disclosure is not limited thereto, and the first, second, and third statuses may also be set to other statuses as needed.
[0109] According to an example of an embodiment of the present disclosure, the status determination unit 1710 is configured to determine the status of each version among multiple versions of multiple associated sets according to the following steps. As Figure 6 shown, in step S511, based on the input values of at least one data in the data set, one or more write values of the multiple associated sets are determined. In a certain write task of the multiple associated sets, the values of one or more data in the input data set may be input. After determining one or more write values of the multiple associated sets, in step S512, the status of the WP version corresponding to the current write task is set to a first status, for example, set to the start status S to mark the start of the write operation. In step S513, using the determined one or more write values, a write operation is performed on the multiple associated sets to obtain multiple associated sets corresponding to the WP version of the current write task. Here, only the incremental data can be output to the associated set, that is, only the values of the data that have changed are output to the associated set (such as Figure 8 step ⑤ in). Then, in step S514, based on the completion of the write operation of the current write task, the status of the corresponding version is updated to a second status or a third status. As described above, the second status may be, for example, the completed status E indicating that the write operation of this version has been completed, and the third status may be, for example, the restart status R indicating that the write operation of this version is restarted. Specifically, when the write operation is completed, the status of this version can be updated to the second status, for example, the completed status E, such as Figure 8 step ⑥ in; and when the write operation fails, a restart write operation corresponding to the determined one or more write values will be executed.
[0110] Specifically, in some cases, when the write task needs to be re-executed for some reason, for example, the write operation of the write task fails due to a short network disconnection, the write task can be restarted to perform a restart write operation corresponding to the determined one or more write values. For example, when applying real-time stream data processing technologies such as Spark streaming to process the data of the present disclosure, when a certain write task fails, the stream processing will restart a new write task, and the restarted write task still processes the same data, so the corresponding data time remains unchanged, but different data versions are generated. In the present disclosure, the write operation and the restart write operation correspond to the same data time, perform a write operation on the same one or more write values, but correspond to different versions. When the write operation fails and the restart write operation is completed, the status of the version corresponding to the write operation can be updated to the third status, for example, the restart status R. In the present disclosure, when the status of a certain version is the restart status R, it means that the data of this version has been overwritten by a subsequent version.
[0111] When the write operation fails and the restarted write operation is completed, the version can be converted from the first state to the third state, and the failed data version can be marked (e.g., marked as R). However, since the data read version is the last version that is continuously in the second state, in order to be able to read the correct data written by the restarted write operation and ensure that the failed data version is not read, a method for restoring the third state to the second state is also required.
[0112] According to an example of an embodiment of the present disclosure, when the write operation of a certain version fails and its restarted write operation is successful, the state of this version is set to the third state. At this time, based on the states of one or more versions with smaller version numbers compared to this version, and the completion status of one or more versions between the version corresponding to the write operation and the version corresponding to the restarted write operation, the state of this version can be restored to the second state. Specifically, for this version in the third state, when one or more versions with smaller version numbers compared to this version are all in the second state or the third state, and one or more versions between this version and the version corresponding to the restarted write operation are all in the second state or the third state, the state of this version is restored to the second state.
[0113] The specific conversion methods by which the state determination unit 1710 converts the version state from the first state to the third state (e.g., from S to R) and from the third state to the second state (e.g., from R to E) can be seen in the content described above in combination with Figure 10 and Figures 11A - 11B what has been described, and will not be elaborated here. Through this conversion method, the correctness and consistency of data reading can be ensured when the write task is restarted.
[0114] The data reading unit 1720 is configured to determine a data reading version of multiple associated sets based on the status of each of multiple versions, and perform data reading on the multiple associated sets based on the data reading version. The data reading version may be the maximum version that satisfies consistency and can be currently read by the multiple associated sets, that is, the maximum version in which the data between the multiple associated sets remains consistent. The version number of the data reading version may be represented by RP. According to an example of an embodiment of the present disclosure, the version that is the last consecutive one in the multiple versions to be in the second state (for example, the completed state E) may be determined as the data reading version. In an ideal case where the number of versions of each piece of data that the system can store has no upper limit, it can be ensured that when data reading is performed using this data reading version, the data between the multiple associated sets is always consistent. However, in an actual situation, the number of versions of each piece of data that the system can store is limited. Therefore, it is possible that when data reading is performed using a certain data reading version, a certain piece of data of this version may not be stored in the system, resulting in a system error. The present disclosure provides a method for determining a data reading version, so that in this case, inconsistent data can be preferentially returned, enabling the user to read the data instead of directly reporting an error.
[0115] Specifically, according to an example of an embodiment of the present disclosure, the data reading unit 1720 is configured to determine the version that is the last consecutive one in the multiple versions to be in the second state as the first data reading version; determine the version with the largest version number in the multiple versions as the second data reading version; and determine the data reading version based on the first data reading version and the second data reading version. According to an example of an embodiment of the present disclosure, in a case where the difference between the version number of the second data reading version and the version number of the first data reading version is less than a predetermined threshold, the data reading version is determined to be the first data reading version; and in a case where the difference between the version number of the second data reading version and the version number of the first data reading version is greater than or equal to the predetermined threshold, the data reading version is determined to be the second data reading version. For example, the predetermined threshold may be the maximum version number N of each piece of data of the multiple associated sets that make up the multiple versions that the system can store, or the predetermined threshold may also be set according to actual storage requirements. The present disclosure does not make a specific limitation thereto.
[0116] For example, the version number of the first data reading version is referred to as logical RP (RP logic ), the version number of the second data reading version is referred to as maximum RP (RP max ), and the version number of the finally determined data reading version is referred to as actual RP (RP real ). Assuming that the predetermined threshold is the maximum version number N of each piece of data of the multiple associated sets that make up the multiple versions that the system can store, the data reading version RP can be determined according to the following rules real :
[0117] · When RP max and RP logic the difference is less than N, RP real = RP logic ;
[0118] · When RP max and RP logic the difference is greater than or equal to N, RP real = RP max , that is, equal to the current maximum version number.
[0119] According to an example of the present disclosure, after determining the data reading version RP real , for example, a query statement with RP real can be generated using the Structured Query Language (SQL) to perform a data query operation. For example, an open-source cluster computing framework application (Apache Phoenix) can be used to construct the query statement as shown in Table 1 above. Querying using the query statement in Table 1 can return data of L versions whose version numbers are less than or equal to RP real . For example, L can be set to 1, that is, only data of 1 version whose version number is less than or equal to RP real is returned. It should be understood that the present disclosure is not limited to using Apache Phoenix as an example here for data query, but any suitable method in the art can be used for data query.
[0120] The above describes a data processing device according to an embodiment of the present disclosure. By determining the status of each version among multiple versions of multiple associated sets and determining the data reading version of the multiple associated sets based on the status of each version, it can be ensured that when data is read from multiple associated sets using this data reading version, the data read from different associated sets is consistent. In addition, through the conversion of the version from the first state to the third state and the conversion from the third state to the second state, the data processing method according to an embodiment of the present disclosure can also ensure the correctness and consistency of data reading when the data writing operation restarts.
[0121] Next, a data compression device according to an embodiment of the present disclosure will be described in conjunction with Figure 18 . Figure 18 FIG. shows a schematic structural diagram of a data compression device 1800 according to an embodiment of the present disclosure. Since the data compression device 1800 has the same details as the data compression method 1500 described above in conjunction with Figure 15 , for simplicity, the detailed description of the same content is omitted here. As Figure 18As shown, the data compression device 1800 includes a status determination unit 1810 and a compression unit 1820. In addition to these two units, the data compression device 1800 may further include other components. However, since these components are not related to the content of the embodiments of the present application, their illustrations and descriptions are omitted here.
[0122] The status determination unit 1810 is configured to determine the status of each version among multiple versions of multiple associated sets. Since the status determination unit 1810 is similar to the details of the status determination unit 1710 described above with reference to Figure 17 For simplicity, the detailed description of the same content is omitted here.
[0123] The compression unit 1820 is configured to determine a data read version of multiple associated sets based on the status of each version among the multiple versions. The data read version may be the maximum version that satisfies consistency and can be currently read from multiple associated sets, that is, the maximum version in which the data among multiple associated sets remains consistent, and the version number of this data read version can be represented by a logical RP (RP logic ). According to an example of the embodiments of the present disclosure, the last consecutive version in the multiple versions in the second state (for example, the completed state E) may be determined as the data read version.
[0124] The compression unit 1820 is further configured to compress data of different versions based on the relationship between the data read version and different versions of the data constituting the multiple associated sets. Specifically, for each piece of data in the multiple associated sets, only the data of the version with a larger version number compared to the data read version is retained; and, the data of the version with the largest version number less than or equal to the version number of the data read version is retained. For example, if there are multiple versions of each piece of data greater than RP logic , then all these versions are retained because these versions may be read later; if there are multiple versions of each piece of data less than or equal to RP logic , then only the version with the largest version number among these versions is retained. For example, if W i is the largest among these versions, then only W i is retained, so that only W i can be read, and the remaining versions are cleared and thus not read.
[0125] By using the data compression device for associated sets with multiple versions according to the embodiments of the present disclosure, by compressing data of different versions based on the relationship between the data read version and different versions of the data constituting the multiple associated sets, it is possible to reduce the storage cost and query latency while ensuring the correctness of data reading.
[0126] In addition, the devices according to the embodiments of the present application (e.g., data processing devices, data compression devices, etc.) can also be implemented by means of Figure 19 the architecture of the exemplary computing device shown. Figure 19 FIG. shows a schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure. As Figure 19 shown, the computing device 1900 may include a bus 1910, one or more CPUs 1920, a read-only memory (ROM) 1930, a random access memory (RAM) 1940, a communication port 1950 connected to a network, an input / output component 1960, a hard disk 1970, etc. The storage devices in the computing device 1900, such as the ROM 1930 or the hard disk 1970, may store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU. The computing device 1900 may also include a user interface 1980. Of course, Figure 19 the architecture shown is only exemplary, and when implementing different devices, one or more components in the shown computing device may be omitted according to actual needs. Figure 19 shown in the computing device.
[0127] The embodiments of the present application may also be implemented as a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium according to the embodiments of the present application. When the computer-readable instructions are run by a processor, the data processing method or data compression method according to the embodiments of the present application described with reference to the above drawings may be executed. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0128] According to an embodiment of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer-readable instructions, and the computer-readable instructions are stored in a computer-readable storage medium. The processor of the computer device may read the computer-readable instructions from the computer-readable storage medium, and the processor executes the computer-readable instructions, so that the computer device executes the data processing method or data compression method described in the above various embodiments.
[0129] Those skilled in the art can understand that the content disclosed in the present application may have various variations and improvements. For example, the various devices or components described above may be implemented by hardware, or may be implemented by software, firmware, or some or all of the combinations of the three.
[0130] In addition, as shown in this application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. The terms "first", "second", and similar terms used in this application do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, words such as "comprising" or "including" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0131] In addition, flowcharts are used in this application to illustrate the operations performed by the systems of the embodiments according to the embodiments of this application. It should be understood that the operations before or below do not necessarily have to be performed precisely in order. On the contrary, various steps can be performed in reverse order or simultaneously. At the same time, other operations can also be superimposed on these processes, or one or more steps can be removed from these processes.
[0132] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense, unless specifically defined as such herein.
[0133] The above has described this application in detail, but for those skilled in the art, it is obvious that this application is not limited to the embodiments described in this specification. This application can be implemented in the form of modifications and changes without departing from the spirit and scope of this application as determined by the claims. Therefore, the description in this specification is for illustrative purposes and does not have any restrictive meaning for this application.
Claims
1. A data processing method for multiple associated sets with multiple versions, comprising: Determining the status of each version among the multiple versions of the multiple associated sets, where the multiple associated sets are multiple mutually associated sets obtained by aggregating and calculating the same data set according to different dimensions, the multiple versions have different creation times, and for each creation time, each associated set among the multiple associated sets corresponds to the same version, and the status of each version indicates the completion of the write operation corresponding to this version, and the status includes a first status where the write operation starts, a second status where the write operation has been completed, or a third status where the write operation restarts; Based on the status of each version among the multiple versions, determining a data read version of the multiple associated sets, where the data read version is the maximum version that satisfies data consistency among the multiple associated sets; And Based on the data read version, performing data reading on the multiple associated sets, where determining the data read version of the multiple associated sets includes: Determining the first data read version as the version among the multiple versions that is continuously the second status last; Determining the second data read version as the version with the largest version number among the multiple versions; and Determining the data read version based on the first data read version and the second data read version.
2. The data processing method according to claim 1, wherein, Determining the status of each version among the multiple versions of the multiple associated sets includes: Based on the input value of at least one data in the data set, determining one or more write values for the multiple associated sets; Setting the status of the version to the first status; Performing a write operation on the multiple associated sets using the one or more write values to obtain the multiple associated sets of this version; Based on the completion of the write operation, updating the status of the version to the second status or the third status.
3. The data processing method according to claim 2, wherein, Determining one or more write values for the multiple associated sets includes: Performing an aggregation calculation on the multiple associated sets using the input value of the at least one data to determine one or more write values for the multiple associated sets.
4. The data processing method according to claim 2, wherein, Based on the completion of the write operation, updating the status of the version to the second status or the third status includes: When the write operation is completed, updating the status of the version to the second status; and When the write operation fails, performing a restart write operation corresponding to the one or more write values, and when the restart write operation is completed, updating the status of the version to the third status, where the write operation and the restart write operation correspond to different versions.
5. The data processing method according to claim 4, wherein, When the status of the version is the third status, the method further includes: Based on the status of one or more versions with smaller version numbers than this version, and the completion of one or more versions between this version and the version corresponding to the restart write operation, repairing the status of this version to the second status.
6. The data processing method according to claim 5, wherein, When one or more versions with a smaller version number than the said version are in the second state or the third state, and one or more versions between the said version and the version corresponding to the restart write operation are in the second state or the third state, repair the state of the said version to the second state.
7. The data processing method according to claim 1, wherein, Determining the data read version based on the first data read version and the second data read version includes: When the difference between the version number of the second data read version and the version number of the first data read version is less than a predetermined threshold, determine that the data read version is the first data read version; and When the difference between the version number of the second data read version and the version number of the first data read version is greater than or equal to the predetermined threshold, determine that the data read version is the second data read version.
8. A data compression method for multiple associated sets with multiple versions, including: Determine the state of each version among the multiple versions of the multiple associated sets, where the multiple associated sets are multiple mutually associated sets obtained by aggregating and calculating the same data set according to different dimensions, the multiple versions have different creation times, and for each creation time, each associated set among the multiple associated sets corresponds to the same version, and the state of each version indicates the completion of the write operation corresponding to this version, and the state includes the first state where the write operation starts, the second state where the write operation has been completed, or the third state where the write operation restarts; Based on the state of each version among the multiple versions, determine the data read version of the multiple associated sets, where the data read version is the largest version that satisfies data consistency among the multiple associated sets; And Based on the relationship between the data read version and different versions of the data constituting the multiple associated sets, compress the data of the different versions, where determining the data read version of the multiple associated sets includes: Determine the first data read version as the last version among the multiple versions that is continuously in the second state; Determine the second data read version as the version with the largest version number among the multiple versions; and Determine the data read version based on the first data read version and the second data read version.
9. The data compression method according to claim 8, wherein, Determining the state of each version among the multiple versions of the associated set includes: Based on the input value of at least one data in the data set, determine one or more write values of the multiple associated sets; Set the state of the said version to the first state; Perform a write operation on the multiple associated sets using the one or more write values to obtain the multiple associated sets of the said version; Based on the completion of the write operation, update the state of the said version to the second state or the third state.
10. The data compression method according to claim 8, wherein, Based on the relationship between the data read version and different versions of the data constituting the multiple associated sets, compressing the data of the different versions includes: Retain the data of the version with a larger version number compared to the data read version; and Retain the data of the version with the maximum version number having a version number less than or equal to the data read version.
11. A data processing device for an associated set with multiple versions, comprising: One or more processors; and One or more memories, wherein computer-readable code is stored in the memories, and when the computer-readable code is run by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1-10.
12. A computer-readable storage medium having instructions stored thereon, which when executed by a processor, cause the processor to execute the method according to any one of claims 1-10.
Citation Information
Patent Citations
In-memory database system providing lockless read and write operations for OLAP and OLTP transactions
CN105868228A
Distributed data reading method and device
CN110196856A