Incremental construction method, construction device and construction system of a cube
By obtaining the start and end times of Cube shards, deleting shards that do not meet the conditions and reconstructing shards, the problem of the inability to reconstruct the merged Cube fragments in the existing technology is solved, and Cube shard storage is realized according to natural cycles, improving Kylin's efficiency.
Patent Information
- Application Number
- CN202211460619.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-11-17
AI Technical Summary
The merged Cube fragment cannot be rebuilt in the prior art, resulting in an increase in cross-film queries and construction errors, which cannot meet the data storage needs in natural cycles.
By obtaining the target start and termination time, determining the boundary conditions of Cube shards, deleting shards that do not meet the conditions, and building a new Cube shard according to the new start and termination time, splitting or merging shards to meet the time interval requirements.
Cube shard storage is implemented regularly according to natural cycles, reducing cross-slice queries, improving Kylin's usage and operation and maintenance efficiency, and solving the problem of being unable to rebuild merged Cube fragments.
Smart Images

Figure CN115934710B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multi-dimensional data analysis. Specifically, it relates to an incremental construction method, a construction device, a computer-readable storage medium, and a construction system for a Cube cube. Background Art
[0002] In the field of big data analysis, constructing a Cube based on the Kylin analytical data warehouse is a commonly used solution. Kylin provides a method for constructing CubeSegments according to corresponding natural cycles for star model / snowflake model data sets stored and growing by natural cycles, which is called incremental Cube construction. At the same time, Kylin will merge some consecutive CubeSegments to improve storage efficiency and reduce cross-segment queries. However, this method has the following two disadvantages:
[0003] Incremental construction is related to the natural cycle, but merging is not related to the natural cycle. Merging is only related to the number, resulting in some irregular CubeSegments being merged, and in some scenarios, cross-segment queries are increased instead. For example, when constructing a Cube incrementally by day, some daily-level CubeSegments at the beginning of this month may be merged into the previous month, and cross-CubeSegment queries will occur when analyzing and querying this month.
[0004] When it is necessary to reconstruct a CubeSegment of a certain natural cycle and it has been merged, a direct error will be reported and construction cannot be carried out. Summary of the Invention
[0005] The main purpose of this application is to provide an incremental construction method, a construction device, a computer-readable storage medium, and a construction system for a Cube cube to solve the problem in the prior art that a merged Cube segment cannot be reconstructed.
[0006] According to one aspect of the embodiments of the present application, an incremental construction method of a Cube cube is provided. The Cube cube is a data structure for storing data in a database. The Cube cube includes a plurality of Cube shards, and the boundaries of the Cube shards are marked with a start time and an end time. One Cube shard is used to store a data set for a time interval. The method includes: obtaining a target start time, a target end time, and the start time and the end time of the Cube shards in the database. The target start time is the start point of the time interval corresponding to the data set to be stored, and the target end time is the end point of the time interval corresponding to the data set to be stored; determining whether the start time of the first target Cube shard meets a first preset condition, and determining whether the end time of the first target Cube shard meets a second preset condition. The first preset condition is that the start time of the first target Cube shard is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube shard is greater than or equal to the target end time. The first target Cube shard is any one of the Cube shards; in the case that the start time of the first target Cube shard meets the first preset condition and the end time of the first target Cube shard meets the second preset condition, deleting the first target Cube shard in the database, and constructing a second target Cube shard according to the start time and the end time of the first target Cube shard to store the data set to be stored.
[0007] Optionally, after deleting the first target Cube shard in the database and constructing a second target Cube shard according to the start time and the end time of the first target Cube shard to store the dataset to be stored, the method further includes: obtaining the start time and the end time of the Cube shard in the database; a first obtaining step of obtaining a first target time and a second target time from the start time and the end time of a third target Cube shard, where the third target Cube shard is any one of the Cube shards, the first target time is the month corresponding to the start time of the third target Cube shard, and the second target time is the month corresponding to the end time of the third target Cube shard; a first determining step of determining that the third target Cube shard needs to be split when the first target time is different from the second target time; a second determining step of determining that the third target Cube shard does not need to be split when the first target time is the same as the second target time; sequentially executing the first obtaining step, the first determining step, and the second determining step at least once until the determination work of all the Cube shards is completed, and obtaining a plurality of Cube shards that need to be split and a plurality of Cube shards that do not need to be split.
[0008] Optionally, after sequentially performing the first obtaining step, the first determining step, and the second determining step at least once until the determination work of all the Cube shards is completed, obtaining a plurality of the Cube shards to be split and a plurality of the Cube shards not to be split, the method further includes: determining a plurality of Cube shard sets among the Cube shards not to be split, where one Cube shard set includes the Cube shards with the same month corresponding to the start time; a second obtaining step of obtaining a target Cube shard set, where the target Cube shard set is any one of the Cube shard sets; a sorting step of sorting the Cube shards in the target Cube shard set in ascending order of the day corresponding to the start time of the Cube shards; a third determining step of determining multiple groups of Cube shards to be merged according to the start time and the end time of the Cube shards in the sorted target Cube shard set, where for any two adjacent Cube shards in each group of Cube shards to be merged, the end time of the previous Cube shard is the same as the start time of the next Cube shard; sequentially performing the second obtaining step, the sorting step, and the third determining step at least once until the determination work of all the Cube shard sets is completed, obtaining multiple groups of Cube shards to be merged; merging all the Cube shards in each group of Cube shards to be merged to obtain a merged Cube shard and storing it in the database, where one group of Cube shards to be merged corresponds to one merged Cube shard.
[0009] Optionally, after the first obtaining step, the first determining step, and the second determining step are sequentially executed at least once until the determination work of all the Cube shards is completed, obtaining a plurality of Cube shards to be split and a plurality of Cube shards not to be split, the method further includes: a third obtaining step of obtaining a third target time, a fourth target time, a fifth target time, and a sixth target time, where the third target time is the month corresponding to the start time of a fourth target Cube shard, the fourth target time is the month corresponding to the end time of the fourth target Cube shard, the fifth target time is the day corresponding to the start time of the fourth target Cube shard, the sixth target time is the day corresponding to the end time of the fourth target Cube shard, and the fourth target Cube shard is any one of the Cube shards to be split; a fourth determining step of determining the start time and the end time for splitting the Cube shard according to the third target time, the fourth target time, the fifth target time, and the sixth target time, where the split Cube shard is the Cube shard after splitting the fourth target Cube shard, the month corresponding to the start time of the split Cube shard is the same as the month corresponding to the end time, and the months corresponding to the start times of any two split Cube shards are different; a splitting step of splitting the fourth target Cube shard according to the start time and the end time of the split Cube shard; and sequentially executing the third obtaining step, the fourth determining step, and the splitting step at least once until the splitting work of all the Cube shards to be split is completed.
[0010] Optionally, splitting the fourth target Cube shard according to the start time and the end time of the split Cube shard includes: deleting the fourth target Cube shard from the database; and constructing the split Cube shard in the database according to the start time and the end time of the split Cube shard.
[0011] Optionally, determining the start time and end time of splitting the Cube shard according to the third target time, the fourth target time, the fifth target time, and the sixth target time includes: when the difference between the third target time and the fourth target time is equal to 1, determining the start time and end time of the two split Cube shards according to the third target time, the fourth target time, the fifth target time, and the sixth target time; when the difference between the third target time and the fourth target time is greater than 1, determining the start time and end time of multiple split Cube shards according to the third target time, the fourth target time, the fifth target time, and the sixth target time.
[0012] Optionally, based on the situation that the start time of the first target Cube shard does not meet the first preset condition or the end time of the first target Cube shard does not meet the second preset condition, after determining whether the start time of the first target Cube shard meets the first preset condition and determining whether the end time of the first target Cube shard meets the second preset condition, the method further includes: constructing the second target Cube shard in the database according to the target start time and the target end time to store the dataset to be stored.
[0013] According to another aspect of the embodiments of the present application, an incremental construction device for a Cube cube is further provided. The Cube cube is a data structure for storing data in a database. The Cube cube includes a plurality of Cube slices. The boundaries of the Cube slices are marked with a start time and an end time. One Cube slice is used to store a data set for a time interval. The device includes: an acquisition unit, configured to acquire a target start time, a target end time, and the start time and the end time of the Cube slices in the database. The target start time is the start point of the time interval corresponding to the data set to be stored, and the target end time is the end point of the time interval corresponding to the data set to be stored; a determination unit, configured to determine whether the start time of a first target Cube slice satisfies a first preset condition and determine whether the end time of the first target Cube slice satisfies a second preset condition. The first preset condition is that the start time of the first target Cube slice is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube slice is greater than or equal to the target end time. The first target Cube slice is any one of the Cube slices; a construction unit, configured to, when the start time of the first target Cube slice satisfies the first preset condition and the end time of the first target Cube slice satisfies the second preset condition, delete the first target Cube slice in the database and construct a second target Cube slice according to the start time and the end time of the first target Cube slice to store the data set to be stored.
[0014] According to still another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program. When the program is executed by a processor, the processor executes any one of the incremental construction methods of the Cube cube.
[0015] According to yet another aspect of the embodiments of the present application, an incremental construction system for a Cube cube is further provided, including: one or more processors, a memory, and one or more programs. The one or more programs are stored in the memory and are configured to be executed by the one or more processors. The one or more programs include those for executing any one of the incremental construction methods of the Cube cube.
[0016] The above-mentioned incremental construction method of the Cube cube, where the Cube cube is a data structure used to store data in a database. The Cube cube includes multiple Cube shards, and the boundaries of the Cube shards are marked with a start time and an end time. One Cube shard is used to store a data set for a time interval. First, obtain the target start time, the target end time, and the start time and the end time of the Cube shards in the database. The target start time is the starting point of the time interval corresponding to the data set to be stored, and the target end time is the ending point of the time interval corresponding to the data set to be stored. Then, determine whether the start time of the first target Cube shard satisfies a first preset condition, and determine whether the end time of the first target Cube shard satisfies a second preset condition. The first preset condition is that the start time of the first target Cube shard is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube shard is greater than or equal to the target end time. The first target Cube shard is any one of the Cube shards. Finally, when the start time of the first target Cube shard satisfies the first preset condition and the end time of the first target Cube shard satisfies the second preset condition, delete the first target Cube shard in the database, and construct a second target Cube shard according to the start time and the end time of the first target Cube shard to store the data set to be stored. This method first determines the time interval corresponding to the data set to be stored and the start and end times of the existing Cube shards in the database. Then, it determines whether there is a Cube shard in the database whose start and end times include the time interval corresponding to the data set to be stored. If so, reconstruct the second target Cube shard with the start and end times of this Cube shard to store the data set to be stored. This method solves the problem in the prior art that it is impossible to reconstruct the merged Cube segments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0018] Figure 1 shows a flowchart of an incremental construction method of a Cube cube according to an embodiment of this application;
[0019] Figure 2 shows a flowchart of an incremental construction method of a Cube cube according to a specific embodiment of this application;
[0020] Figure 3Shows a schematic diagram of an incremental construction device for a Cube cube according to an embodiment of the present application;
[0021] Figure 4 Shows a schematic diagram of an incremental construction system for a Cube cube according to a specific embodiment of the present application. Detailed implementation manners
[0022] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0023] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0024] It should be understood that when an element (such as a layer, film, region, or substrate) is described as being "on" another element, the element can be directly on the other element, or there can also be an intermediate element. Moreover, in the specification and claims, when an element is described as being "connected" to another element, the element can be "directly connected" to the other element, or "connected" to the other element through a third element.
[0025] For the convenience of description, the following explains some nouns or terms related to the embodiments of the present application:
[0026] Kylin: In the present application, Kylin refers to Apache Kylin rather than the Kylin operating system. It is an analytical data warehouse that provides OLAP (online analytical processing) capabilities on top of Hadoop. Its main idea is to "trade space for time". For a star / snowflake model source data set, through pre-computation, all combinations of dimensions are measured and statistically analyzed to form a multi-dimensional data cube.
[0027] Cube cube: A multi-dimensional data cube, a multi-dimensional space data model constructed from dimensions, which contains all the metric basic data to be statistically analyzed, and all aggregation data query operations are performed on it.
[0028] Cube Construction: The computational process from the source dataset to the physical generation of Cube data. In response to changes in the source dataset, Kylin provides two methods: full build and incremental build. In a full build, all data in the source dataset is read and loaded at once for calculation. In an incremental build, it supports source datasets stored partitioned by natural cycles and builds in multiple steps according to natural cycles. Each time, the data in the corresponding natural cycle partition of the source dataset is read and loaded for calculation.
[0029] Cube Shard (CubeSegment): A shard of the multi-dimensional data cube, which is the Cube generated each time in an incremental build scenario. Its boundaries are marked by the start and end times of the natural cycle selected during the incremental build, following the principle of left-closed and right-open, that is, including the data at the start time but not including the data at the end time. Moreover, in the same Cube cube, the boundaries of each Cube shard cannot intersect.
[0030] REST Server: A service module within Kylin that provides RESTful API interfaces. These RESTful API interfaces are called by external applications. This application mainly uses combinations of these interfaces to perform operations such as analyzing Kylin metadata information, building, merging, rebuilding Cubes, and querying the status of build tasks.
[0031] Hive: A data warehouse based on Hadoop that stores Kylin data sources, usually datasets in star / snowflake models. The above-mentioned datasets to be stored are stored in Hive.
[0032] HBase: A columnar database based on Hadoop that stores Kylin's Cube cubes and Kylin's metadata information.
[0033] As mentioned in the background art, in the prior art, it is impossible to rebuild merged Cube segments. To solve the above problems, in a typical embodiment of this application, an incremental construction method, construction device, computer-readable storage medium, and construction system for Cube cubes are provided.
[0034] According to an embodiment of the present application, an incremental construction method for Cube cubes is provided.
[0035] Among them, the above-mentioned Cube cube is a data structure in the database for storing data. The above-mentioned Cube cube includes multiple Cube shards. The boundaries of the above-mentioned Cube shards are marked by start times and end times. One above-mentioned Cube shard is used to store a dataset for a time interval.
[0036] Figure 1It is a flowchart of an incremental construction method of a Cube cube according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0037] Step S101, obtain the target start time, the target end time, and the start time and the end time of the above-mentioned Cube shard in the above-mentioned database. The target start time is the start point of the above-mentioned time interval corresponding to the dataset to be stored, and the target end time is the end point of the above-mentioned time interval corresponding to the dataset to be stored;
[0038] Among them, the target start time and the target end time are respectively the start time and the end time of the Cube shard that needs to be reconstructed for storing the dataset to be stored this time. As Figure 2 shown, obtain the Cube shard list of the database, traverse the Cube shard list, parse the boundaries of each Cube shard, and obtain the start time and the end time of each Cube shard.
[0039] Step S102, determine whether the start time of the first target Cube shard meets the first preset condition, and determine whether the end time of the first target Cube shard meets the second preset condition. The first preset condition is that the start time of the first target Cube shard is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube shard is greater than or equal to the target end time. The first target Cube shard is any one of the above-mentioned Cube shards;
[0040] Among them, as Figure 2 shown, while traversing the Cube shard list, determine whether the start time of each Cube shard in the database is less than or equal to the target start time, and determine whether the end time of each Cube shard is greater than or equal to the target end time, so as to determine whether there is a Cube shard in the existing Cube shards in the database whose start and end times include the start and end times of the Cube shard that needs to be reconstructed this time, that is, determine whether the Cube shard that needs to be reconstructed this time is merged.
[0041] In order to directly construct a Cube shard when the Cube shard that needs to be reconstructed this time is not merged, based on the situation that the start time of the first target Cube shard does not meet the first preset condition or the end time of the first target Cube shard does not meet the second preset condition, in an optional implementation manner, after the above-mentioned step S102, the method further includes:
[0042] Step S201: Construct the second target Cube shard in the above database according to the above target start time and the above target end time to store the above dataset to be stored.
[0043] In the above embodiment, if there is no Cube shard in the existing Cube shards in the database whose start and end times include the start and end times of the Cube shard that needs to be reconstructed this time, that is, it is determined that the Cube shard that needs to be reconstructed this time has not been merged. At this time, construct the second target Cube shard in the database according to the target start time and the target end time to store the dataset to be stored.
[0044] It should be noted that, as Figure 2 shown, when constructing the second target Cube shard, continuously detect the execution status of the construction task and wait for the task to end normally.
[0045] Step S103: When the start time of the above first target Cube shard meets the first preset condition and the end time of the above first target Cube shard meets the second preset condition, delete the above first target Cube shard in the above database, and construct the second target Cube shard according to the start time and the end time of the above first target Cube shard to store the above dataset to be stored.
[0046] Among them, as Figure 2 shown, when traversing the list of Cube shards, if it is determined that there is a Cube shard in the database whose start and end times include the start and end times of the Cube shard that needs to be reconstructed this time, that is, it is determined that the Cube shard that needs to be reconstructed this time has been merged. At this time, jump out of the traversal, and construct the second target Cube shard in the database according to the start time and the end time of the Cube shard whose start and end times include the start and end times of the Cube shard that needs to be reconstructed this time to store the dataset to be stored.
[0047] In order to determine whether the start and end times of the existing Cube shards in the database cross natural cycles, in an optional embodiment, after the above step S103, the method further includes:
[0048] Step S301: Obtain the start time and the end time of the above Cube shard in the above database;
[0049] Step S302, the first acquisition step: Obtain a first target time and a second target time from the start time and the end time of the third target Cube slice. The third target Cube slice is any one of the Cube slices. The first target time is the month corresponding to the start time of the third target Cube slice, and the second target time is the month corresponding to the end time of the third target Cube slice.
[0050] Step S303, the first determination step: When the first target time is different from the second target time, determine that the third target Cube slice needs to be split.
[0051] Step S304, the second determination step: When the first target time is the same as the second target time, determine that the third target Cube slice does not need to be split.
[0052] Step S305: Execute the first acquisition step, the first determination step, and the second determination step at least once in sequence until the determination work of all the Cube slices is completed, obtaining multiple Cube slices that need to be split and multiple Cube slices that do not need to be split.
[0053] In the above embodiment, when determining whether the start and end times of the Cube slices already stored in the database span a natural cycle, as Figure 2 shown, at this time, it is necessary to obtain the Cube slice list again, parse the boundaries of all Cube slices, and determine whether the month corresponding to the start time of each Cube slice is the same as the month corresponding to the end time, so as to determine whether the start and end times of each Cube slice span a natural cycle. The natural cycle is a month. For example, when the start time of the Cube slice is May 1st and the end time is May 22nd, it is determined that the start and end times of the Cube slice do not span months. At this time, it is determined that the Cube slice does not need to be split. When the start time of the Cube slice is May 1st and the end time is June 22nd, it is determined that the start and end times of the Cube slice span a natural cycle. At this time, it is determined that the Cube slice needs to be split.
[0054] It should be noted that the start and end times of the Cube slice follow the principle of left-closed and right-open, that is, the start time is included and the end time is not included. For example, the start and end times of the Cube slice are [May 1st, June 1st). Among them, May 1st is the start and end time of the Cube slice, June 1st is the cut-off time of the Cube slice, and May 31st is the above-mentioned end time of the Cube slice. Therefore, when determining whether the start and end times of the Cube slice span months, it is necessary to compare whether May 1st and May 31st are in the same month.
[0055] In order to merge the Cube shards of the same natural cycle in the database into one Cube shard, in an optional implementation manner, after the above-mentioned step S305, the above method further includes:
[0056] Step S401: In the above-mentioned Cube shards that do not need to be split, determine multiple Cube shard sets. One of the above-mentioned Cube shard sets includes the above-mentioned Cube shards with the same month corresponding to the above-mentioned start time;
[0057] Step S402: A second acquisition step, to acquire a target Cube shard set, where the above-mentioned target Cube shard set is any one of the above-mentioned Cube shard sets;
[0058] Step S403: A sorting step, to sort the above-mentioned Cube shards in the above-mentioned target Cube shard set in ascending order of the day corresponding to the above-mentioned start time of the above-mentioned Cube shards;
[0059] Step S404: A third determination step, according to the above-mentioned start time and the above-mentioned end time of the above-mentioned Cube shards in the sorted above-mentioned target Cube shard set, determine multiple groups of above-mentioned Cube shards to be merged. For any two adjacent above-mentioned Cube shards in each group of above-mentioned Cube shards to be merged, the above-mentioned end time of the previous above-mentioned Cube shard is the same as the above-mentioned start time of the latter above-mentioned Cube shard;
[0060] Step S405: Execute the above-mentioned second acquisition step, the above-mentioned sorting step, and the above-mentioned third determination step at least once in sequence until the determination work of all the above-mentioned Cube shard sets is completed, and obtain multiple groups of above-mentioned Cube shards to be merged;
[0061] Step S406: Merge all the above-mentioned Cube shards in each group of above-mentioned Cube shards to be merged to obtain a merged Cube shard and store it in the above-mentioned database. One group of above-mentioned Cube shards to be merged corresponds to one above-mentioned merged Cube shard.
[0062] In the above embodiments, first, a plurality of Cube shard sets are determined. The Cube shards in one Cube shard set belong to the same natural cycle, that is, the same month. Then, each Cube shard is sorted. After sorting, it is determined whether any two adjacent Cube shards in the Cube shard set are continuous. For example, if the start and end time of the previous Cube shard is [May 1, May 6), and the start and end time of the next Cube shard is [May 6, May 8), it is determined that these two Cube shards can be merged. If the start and end time of the previous Cube shard is [May 1, May 6), and the start and end time of the next Cube shard is [May 8, May 9), it is determined that these two Cube shards cannot be merged. Thus, multiple groups of Cube shards to be merged are determined. As Figure 2 shown, each group of Cube shards to be merged is packaged as a merge task and put into the merge task list. The merge task list is traversed. For each merge task, the merge task is submitted, and the execution status of the merge task is cyclically detected, waiting for the task to end normally until the traversal is completed, so as to form a Cube shard list regularized by natural cycle and improve the efficiency of Kylin use and operation and maintenance.
[0063] In order to split the Cube shards across natural cycles in the database, in an optional embodiment, after the above step S305, the above method further includes:
[0064] Step S501, a third acquisition step, to acquire a third target time, a fourth target time, a fifth target time, and a sixth target time. The third target time is the month corresponding to the start time of the fourth target Cube shard. The fourth target time is the month corresponding to the end time of the fourth target Cube shard. The fifth target time is the day corresponding to the start time of the fourth target Cube shard. The sixth target time is the day corresponding to the end time of the fourth target Cube shard. The fourth target Cube shard is any one of the Cube shards to be split;
[0065] Step S502, a fourth determination step, to determine the start time and end time of the split Cube shards according to the third target time, the fourth target time, the fifth target time, and the sixth target time. The split Cube shards are the Cube shards after splitting the fourth target Cube shard. The month corresponding to the start time of the split Cube shards is the same as the month corresponding to the end time, and the months corresponding to the start times of any two split Cube shards are different;
[0066] Optionally, the present application does not limit the specific process of determining the start time and end time of the split Cube shards according to the above-mentioned third target time, the above-mentioned fourth target time, the above-mentioned fifth target time, and the above-mentioned sixth target time. Any feasible method belongs to the protection scope of the present application.
[0067] In an alternative embodiment, the above step S502 includes:
[0068] Step S5021: When the difference between the above-mentioned third target time and the above-mentioned fourth target time is equal to 1, determine the start time and end time of the two above-mentioned split Cube shards according to the above-mentioned third target time, the above-mentioned fourth target time, the above-mentioned fifth target time, and the above-mentioned sixth target time;
[0069] Step S5022: When the difference between the above-mentioned third target time and the above-mentioned fourth target time is greater than 1, determine the start time and end time of multiple above-mentioned split Cube shards according to the above-mentioned third target time, the above-mentioned fourth target time, the above-mentioned fifth target time, and the above-mentioned sixth target time.
[0070] In the above embodiment, when splitting the Cube shards across natural cycles in the database, if the Cube shard spans two natural cycles, the Cube shard is split into two split Cube shards. For example, if the start and end times of the Cube shard are [May 20th, June 20th), the start and end times of the split Cube shards are determined to be [May 20th, June 1st) and [June 1st, June 20th). If the Cube shard spans more than two natural cycles, the Cube shard is split into multiple split Cube shards. For example, if the start and end times of the Cube shard are [May 20th, July 20th), the start and end times of the split Cube shards are determined to be [May 20th, June 1st), [June 1st, July 1st), and [July 1st, July 20th).
[0071] Step S503, the splitting step, splits the above-mentioned fourth target Cube shard according to the start time and end time of the above-mentioned split Cube shard;
[0072] Optionally, the present application does not limit the specific process of splitting the above-mentioned fourth target Cube shard according to the start time and end time of the above-mentioned split Cube shard. Any feasible method belongs to the protection scope of the present application.
[0073] In an alternative embodiment, the above step S503 includes:
[0074] Step S5031, delete the above-mentioned fourth target Cube shard from the above-mentioned database;
[0075] Step S5032: Construct the split Cube shards in the above database according to the above start time and the above end time of the split Cube shards.
[0076] In the above implementation, the method of splitting the cross-natural-cycle Cube shards into split Cube shards is to first delete the existing cross-natural-cycle Cube shards, and then reconstruct the split Cube shards respectively.
[0077] Step S504: Execute the above third acquisition step, the above fourth determination step, and the above split step at least once until the splitting of all the Cube shards that need to be split is completed.
[0078] In the above implementation, when splitting the cross-natural-cycle Cube shards in the database, the start time and the end time of each Cube shard are determined. The splitting principle is that the split Cube shards obtained after splitting each Cube shard must belong to different natural cycles, that is, different months. For example, Figure 2 as shown, list the split Cube shards after splitting as corresponding reconstruction tasks and put them into the reconstruction task list. For each reconstruction task, first delete the existing cross-natural-cycle Cube shards, then submit the reconstruction subtasks respectively, and loop to detect the execution status of the subtasks. Wait for all subtasks to end normally, then the overall reconstruction task is completed. Until traversal is completed, finally form a list of Cube shards regularized by natural cycle, reducing cross-Cube shard queries and improving the efficiency of Kylin use and operation and maintenance.
[0079] The above-mentioned incremental construction method of the Cube cube, where the Cube cube is a data structure used to store data in a database. The Cube cube includes multiple Cube shards, and the boundaries of the Cube shards are marked with start time and end time. One of the above-mentioned Cube shards is used to store a dataset for a time interval. First, obtain the target start time, target end time, and the start time and end time of the above-mentioned Cube shards in the above-mentioned database. The target start time is the starting point of the time interval corresponding to the dataset to be stored, and the target end time is the ending point of the time interval corresponding to the dataset to be stored. Then, determine whether the start time of the first target Cube shard satisfies a first preset condition, and determine whether the end time of the first target Cube shard satisfies a second preset condition. The first preset condition is that the start time of the first target Cube shard is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube shard is greater than or equal to the target end time. The first target Cube shard is any one of the above-mentioned Cube shards. Finally, when the start time of the first target Cube shard satisfies the first preset condition and the end time of the first target Cube shard satisfies the second preset condition, delete the first target Cube shard in the above-mentioned database, and construct a second target Cube shard according to the start time and end time of the first target Cube shard to store the above-mentioned dataset to be stored. This method first determines the time interval corresponding to the dataset to be stored and the start and end times of the existing Cube shards in the database. Then, it determines whether there is a Cube shard in the database whose start and end times include the time interval corresponding to the dataset to be stored. If so, reconstruct a second target Cube shard with the start and end times of this Cube shard to store the dataset to be stored. This method solves the problem in the prior art that the merged Cube segments cannot be reconstructed.
[0080] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0081] The embodiments of the present application also provide an incremental construction device for a Cube cube. It should be noted that the incremental construction device for a Cube cube in the embodiments of the present application can be used to execute the incremental construction method for a Cube cube provided by the embodiments of the present application. The following introduces the incremental construction device for a Cube cube provided by the embodiments of the present application.
[0082] Figure 3 It is a schematic diagram of an incremental construction device for a Cube cube according to an embodiment of the present application. As Figure 3 shown, the device includes:
[0083] An acquisition unit 10, configured to acquire a target start time, a target end time, and the start time and the end time of the above-mentioned Cube shard in the above-mentioned database, where the target start time is the start point of the above-mentioned time interval corresponding to the dataset to be stored, and the target end time is the end point of the above-mentioned time interval corresponding to the dataset to be stored;
[0084] Among them, the target start time and the target end time are respectively the start time and the end time of the Cube shard that needs to be reconstructed for storing the dataset to be stored this time. As Figure 2 shown, obtain the Cube shard list of the database, traverse the Cube shard list, parse the boundaries of each Cube shard, and obtain the start time and the end time of each Cube shard.
[0085] A determination unit 20, configured to determine whether the start time of the first target Cube shard satisfies a first preset condition, and determine whether the end time of the first target Cube shard satisfies a second preset condition. The first preset condition is that the start time of the first target Cube shard is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube shard is greater than or equal to the target end time. The first target Cube shard is any one of the above-mentioned Cube shards;
[0086] Among them, as Figure 2 shown, while traversing the Cube shard list, determine whether the start time of each Cube shard in the database is less than or equal to the target start time, and determine whether the end time of each Cube shard is greater than or equal to the target end time, so as to determine whether there is a Cube shard in the existing Cube shards in the database whose start and end times include the start and end times of the Cube shard that needs to be reconstructed this time, that is, to determine whether the Cube shard that needs to be reconstructed this time is merged.
[0087] In order to directly construct a Cube shard when the Cube shard that needs to be reconstructed this time is not merged, based on the situation that the start time of the first target Cube shard does not satisfy the first preset condition or the end time of the first target Cube shard does not satisfy the second preset condition, in an alternative embodiment, the device further includes:
[0088] A reconstruction unit is used to construct the second target Cube shard in the above database according to the above target start time and the above target end time to store the above dataset to be stored.
[0089] In the above embodiment, if there is no Cube shard in the existing Cube shards in the database whose start and end times include the start and end times of the Cube shard to be reconstructed this time, it is determined that the Cube shard to be reconstructed this time has not been merged. At this time, a second target Cube shard is constructed in the database according to the target start time and the target end time to store the dataset to be stored.
[0090] It should be noted that, as Figure 2 shown, when constructing the second target Cube shard, the execution status of the construction task is detected cyclically and the task is waited to end normally.
[0091] A construction unit 30 is used to delete the first target Cube shard in the above database when the above start time of the first target Cube shard meets the first preset condition and the above end time of the first target Cube shard meets the second preset condition, and construct a second target Cube shard according to the above start time and the above end time of the first target Cube shard to store the above dataset to be stored.
[0092] Among them, as Figure 2 shown, when traversing the list of Cube shards, if it is determined that there is a Cube shard in the database whose start and end times include the start and end times of the Cube shard to be reconstructed this time, that is, it is determined that the Cube shard to be reconstructed this time has been merged. At this time, the traversal is jumped out, and a second target Cube shard is constructed in the database according to the start time and the end time of the Cube shard whose start and end times include the start and end times of the Cube shard to be reconstructed this time to store the dataset to be stored.
[0093] In order to determine whether the start and end times of the existing Cube shards in the database span natural cycles, in an optional embodiment, the above device further includes:
[0094] A first acquisition unit is used to acquire the above start time and the above end time of the above Cube shard in the above database;
[0095] A second acquisition unit is used to execute the above first acquisition step, and acquire a first target time and a second target time from the above start time and the above end time of the third target Cube shard. The third target Cube shard is any one of the above Cube shards. The first target time is the month corresponding to the above start time of the third target Cube shard, and the second target time is the month corresponding to the above end time of the third target Cube shard;
[0096] A first determination unit, configured to perform the above-mentioned first determination step, and determine that the third target Cube shard needs to be split when the above-mentioned first target time is different from the above-mentioned second target time;
[0097] A second determination unit, configured to perform the above-mentioned second determination step, and determine that the third target Cube shard does not need to be split when the above-mentioned first target time is the same as the above-mentioned second target time;
[0098] A first iteration unit, configured to sequentially perform the above-mentioned first acquisition step, the above-mentioned first determination step, and the above-mentioned second determination step at least once until the determination work of all the above-mentioned Cube shards is completed, and obtain a plurality of the above-mentioned Cube shards that need to be split and a plurality of the above-mentioned Cube shards that do not need to be split.
[0099] In the above-mentioned embodiment, when determining whether the start and end times of the Cube shards already stored in the database span a natural cycle, as Figure 2 shown, at this time, it is necessary to obtain the Cube shard list again, parse the boundaries of all Cube shards, and determine whether the month corresponding to the start time of each Cube shard is the same as the month corresponding to the end time, so as to determine whether the start and end times of each Cube shard span a natural cycle, and the natural cycle is a month. For example, when the start time of the Cube shard is May 1st and the end time is May 22nd, it is determined that the start and end times of the Cube shard do not span months. At this time, it is determined that the Cube shard does not need to be split. When the start time of the Cube shard is May 1st and the end time is June 22nd, it is determined that the start and end times of the Cube shard span a natural cycle. At this time, it is determined that the Cube shard needs to be split.
[0100] It should be noted that the start and end times of the Cube shard follow the principle of left-closed and right-open, that is, the start time is included and the end time is not included. For example, the start and end times of the Cube shard are [May 1st, June 1st), where May 1st is the start and end time of the Cube shard, June 1st is the cut-off time of the Cube shard, and May 31st is the above-mentioned end time of the Cube shard. Therefore, when determining whether the start and end times of each Cube shard span months, it is necessary to compare whether May 1st and May 31st are in the same month.
[0101] In order to merge the Cube shards in the same natural cycle in the database into one Cube shard, in an optional embodiment, the above-mentioned apparatus further includes:
[0102] A third determination unit, configured to determine a plurality of Cube shard sets among the above-mentioned Cube shards that do not need to be split, and one of the above-mentioned Cube shard sets includes the above-mentioned Cube shards with the same month corresponding to the above-mentioned start time;
[0103] A third acquisition unit, configured to perform the second acquisition step above to acquire a target Cube shard set, where the target Cube shard set is any one of the Cube shard sets above;
[0104] A sorting unit, configured to perform the sorting step above, and sort the Cube shards in the target Cube shard set in ascending order of the days corresponding to the start times of the Cube shards above;
[0105] A fourth determination unit, configured to perform the third determination step above, and determine multiple groups of Cube shards to be merged according to the start time and the end time of the Cube shards in the sorted target Cube shard set above, where for any two adjacent Cube shards in each group of Cube shards to be merged, the end time of the previous Cube shard is the same as the start time of the next Cube shard;
[0106] A second iteration unit, configured to perform the second acquisition step, the sorting step, and the third determination step above in sequence at least once until the determination work of all the Cube shard sets above is completed, and obtain multiple groups of Cube shards to be merged;
[0107] A merging unit, configured to merge all the Cube shards in each group of Cube shards to be merged, obtain a merged Cube shard and store it in the database above, and one group of Cube shards to be merged corresponds to one merged Cube shard.
[0108] In the above embodiment, first, multiple Cube shard sets are determined. The Cube shards in one Cube shard set belong to the same natural cycle, that is, belong to the same month. Then, each Cube shard is sorted. After sorting, it is determined whether any two adjacent Cube shards in the Cube shard set are continuous. For example, if the start and end time of the previous Cube shard is [May 1st, May 6th), and the start and end time of the next Cube shard is [May 6th, May 8th), it is determined that these two Cube shards can be merged. If the start and end time of the previous Cube shard is [May 1st, May 6th), and the start and end time of the next Cube shard is [May 8th, May 9th), it is determined that these two Cube shards cannot be merged. Thus, multiple groups of Cube shards to be merged are determined, such as Figure 2As shown, each group of Cube shards to be merged is packaged into a merge task and placed in the merge task list. The merge task list is traversed. For each merge task, the merge task is submitted, and the execution status of the merge task is cyclically detected. Wait for the task to end normally until the traversal is completed, so as to form a list of Cube shards regularized according to the natural cycle and improve the efficiency of Kylin use and operation and maintenance.
[0109] In order to split the Cube shards across natural cycles in the database, in an alternative embodiment, the above device further includes:
[0110] A fourth acquisition unit, configured to execute the above third acquisition step to acquire a third target time, a fourth target time, a fifth target time, and a sixth target time. The third target time is the month corresponding to the start time of the fourth target Cube shard, the fourth target time is the month corresponding to the end time of the fourth target Cube shard, the fifth target time is the day corresponding to the start time of the fourth target Cube shard, and the sixth target time is the day corresponding to the end time of the fourth target Cube shard. The fourth target Cube shard is any one of the Cube shards to be split;
[0111] A fifth determination unit, configured to execute the above fourth determination step to determine the start time and end time of the split Cube shards according to the third target time, the fourth target time, the fifth target time, and the sixth target time. The split Cube shards are the Cube shards after splitting the fourth target Cube shard. The month corresponding to the start time of the split Cube shards is the same as the month corresponding to the end time, and the months corresponding to the start times of any two split Cube shards are different;
[0112] Optionally, the present application does not limit the specific process of determining the start time and end time of the split Cube shards according to the third target time, the fourth target time, the fifth target time, and the sixth target time. Any feasible method belongs to the protection scope of the present application.
[0113] In an alternative embodiment, the above fifth determination unit includes:
[0114] A first determination module: configured to, when the difference between the third target time and the fourth target time is equal to 1, determine the start time and end time of the two split Cube shards according to the third target time, the fourth target time, the fifth target time, and the sixth target time;
[0115] The second determination module: It is used to determine the start time and end time of multiple said split Cube shards according to the said third target time, the said fourth target time, the said fifth target time, and the said sixth target time when the difference between the said third target time and the said fourth target time is greater than 1.
[0116] In the above-mentioned embodiment, when splitting the cross-natural cycle Cube shards in the database, if a Cube shard spans two natural cycles, the Cube shard is split into two split Cube shards. For example, if the start and end times of the Cube shard are [May 20th, June 20th), the start and end times of the split Cube shards are determined to be [May 20th, June 1st) and [June 1st, June 20th). If a Cube shard spans more than two natural cycles, the Cube shard is split into multiple split Cube shards. For example, if the start and end times of the Cube shard are [May 20th, July 20th), the start and end times of the split Cube shards are determined to be [May 20th, June 1st), [June 1st, July 1st), and [July 1st, July 20th).
[0117] The splitting unit: It is used to execute the splitting step and split the said fourth target Cube shard according to the start time and end time of the said split Cube shard;
[0118] Optionally, the present application does not limit the specific process of splitting the said fourth target Cube shard according to the start time and end time of the said split Cube shard, and any feasible method belongs to the protection scope of the present application.
[0119] In an optional embodiment, the said splitting unit includes:
[0120] The deletion module: It is used to delete the said fourth target Cube shard from the said database;
[0121] The reconstruction module: It is used to construct the said split Cube shard in the said database according to the start time and the end time of the said split Cube shard.
[0122] In the above-mentioned embodiment, the method of splitting the cross-natural cycle Cube shard into split Cube shards is to first delete the existing cross-natural cycle Cube shard, and then reconstruct the split Cube shards respectively.
[0123] The third iteration unit: It is used to execute the above-mentioned third acquisition step, the above-mentioned fourth determination step, and the above-mentioned splitting step at least once until the splitting work of all the Cube shards that need to be split is completed.
[0124] In the above embodiments, when splitting the cross-natural-cycle Cube slices in the database, the start time and end time of each Cube slice are determined. The principle of splitting is that the split Cube slices obtained after splitting each Cube slice must belong to different natural cycles, that is, different months. For example, Figure 2 As shown, list the split Cube slices after splitting as corresponding reconstruction tasks and put them into the reconstruction task list. For each reconstruction task, first delete the existing cross-natural-cycle Cube slices, then submit the reconstruction subtasks respectively, and loop to detect the execution status of the subtasks. Wait for all subtasks to end normally, then the entire reconstruction task is completed. Until the traversal is completed, finally form a list of Cube slices regularized by natural cycles, reducing cross-Cube slice queries and improving the efficiency of Kylin use and operation and maintenance.
[0125] The incremental construction device of the above-mentioned Cube cube, where the Cube cube is a data structure for storing data in a database. The Cube cube includes multiple Cube shards, and the boundaries of the Cube shards are marked with a start time and an end time. One of the Cube shards is used to store a data set for a time interval. An acquisition unit is used to acquire a target start time, a target end time, and the start time and the end time of the above-mentioned Cube shards in the above-mentioned database. The target start time is the start point of the time interval corresponding to the data set to be stored, and the target end time is the end point of the time interval corresponding to the data set to be stored. A determination unit is used to determine whether the start time of the first target Cube shard meets a first preset condition and determine whether the end time of the first target Cube shard meets a second preset condition. The first preset condition is that the start time of the first target Cube shard is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube shard is greater than or equal to the target end time. The first target Cube shard is any one of the above-mentioned Cube shards. A construction unit is used to, when the start time of the first target Cube shard meets the first preset condition and the end time of the first target Cube shard meets the second preset condition, delete the first target Cube shard in the above-mentioned database and construct a second target Cube shard according to the start time and the end time of the first target Cube shard to store the above-mentioned data set to be stored. The device first determines the time interval corresponding to the data set to be stored and the start and end times of the existing Cube shards in the database. Then, it determines whether there is a Cube shard in the database whose start and end times include the time interval corresponding to the data set to be stored. If so, it reconstructs the second target Cube shard according to the start and end times of the Cube shard to store the data set to be stored. The device solves the problem in the prior art that it is impossible to reconstruct the merged Cube segments.
[0126] An embodiment of the present application also provides an incremental construction system for a Cube cube, including: one or more processors, a memory, and one or more programs. Among them, the one or more programs are stored in the memory and are configured to be executed by the one or more processors. The one or more programs include those for executing any one of the above-mentioned incremental construction methods for the Cube cube.
[0127] In the above incremental construction system of the Cube cube, it includes: one or more processors, a memory, and one or more programs. Among them, the above one or more programs are stored in the above memory and are configured to be executed by the above one or more processors. The above one or more programs include methods for performing the incremental construction of any of the above Cube cubes. This system first determines the time interval corresponding to the data set to be stored and the start and end times of the existing Cube shards in the database. Then, it determines whether there is a Cube shard in the database whose start and end times include the time interval corresponding to the data set to be stored. If so, it reconstructs the second target Cube shard with the start and end times of this Cube shard to store the data set to be stored. This system solves the problem in the prior art that it is impossible to reconstruct the merged Cube segments.
[0128] To execute the above incremental construction method of the Cube cube, as Figure 4 shown, the above incremental construction system of the Cube cube provides a scheduling tool and an adaptive KylinCube construction and merger application program for executing the above incremental construction method of the Cube cube. Since Kylin does not provide an automatic and timed Cube shard construction function, in the data batch processing process, this application uses the scheduling tool to start the incremental construction Cube shard task regularly. The solution of this application is completely decoupled from the scheduling tool and supports commonly used scheduling tools in the industry, such as Oozie, etc. The adaptive KylinCube construction and merger application program of this application receives parameters from the scheduling tool. The parameters include the name of the Cube shard to be constructed this time, the above target start time, and the above target end time. The above database is an HBase database, and the data set to be stored is stored in the Hive database. This application program is divided into 5 execution modules, namely, the boundary detection module, the construction module, the boundary scanning and natural cycle analysis module, the merger module, and the reconstruction module. The functions of each module are as follows:
[0129] Boundary detection module: Obtain the Cube shard list using the Kylin metadata information stored in HBase, parse the start and end times of each Cube shard, and determine whether the Cube shard to be constructed this time has been merged;
[0130] Building Module: If the boundary detection module determines that the Cube shard to be built this time has not been merged, it directly calls the build API of the RESR Server to submit the build task, and at the same time periodically checks the task execution status in a loop and feeds it back to the upper-level scheduling tool. If the boundary detection module determines that the Cube shard to be built this time has been merged, it adjusts the start time and end time of the Cube shard to be built this time to the start time and end time of the Cube shard that includes the start and end times of the Cube shard that needs to be rebuilt this time, and then calls the rebuild API of the RESR Server to submit the rebuild task;
[0131] It should be noted that the build API is the build interface provided by the RESR Server, and the rebuild API is the rebuild interface provided by the RESR Server.
[0132] Boundary Scanning and Natural Cycle Analysis Module: After the boundary detection module completes the build task, it uses the Kylin metadata information stored in HBase again to obtain the Cube shard list, parses the start and end times of each Cube shard, and determines the Cube shards in the merge task list and the rebuild task list;
[0133] Merge Module: Traverse the merge task list, submit the merge task for each merge task, and periodically check the execution status of the merge task in a loop and wait for the task to end normally until the traversal is completed;
[0134] Rebuild Module: First, delete the existing Cube shards that cross natural cycles, then call the rebuild API of the RESR Server to submit the rebuild subtasks respectively, and periodically check the execution status of the subtasks in a loop and wait for all subtasks to end normally, then the overall rebuild task is completed until the traversal is completed.
[0135] The above incremental building device of the Cube cube includes a processor and a memory. The above obtaining unit, determining unit, building unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement the corresponding functions.
[0136] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem that the existing technology cannot rebuild the merged Cube fragments can be solved.
[0137] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0138] An embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program is executed by a processor, the processor executes the incremental construction method of the above Cube cube.
[0139] An embodiment of the present application provides a device. The device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, at least the following steps are implemented:
[0140] Step S101, obtain a target start time, a target end time, and the start time and the end time of the above Cube shard in the above database. The target start time is the start point of the above time interval corresponding to the data set to be stored, and the target end time is the end point of the above time interval corresponding to the data set to be stored;
[0141] Step S102, determine whether the start time of the first target Cube shard satisfies a first preset condition, and determine whether the end time of the first target Cube shard satisfies a second preset condition. The first preset condition is that the start time of the first target Cube shard is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube shard is greater than or equal to the target end time. The first target Cube shard is any one of the above Cube shards;
[0142] Step S103, when the start time of the first target Cube shard satisfies the first preset condition and the end time of the first target Cube shard satisfies the second preset condition, delete the first target Cube shard in the above database, and construct a second target Cube shard according to the start time and the end time of the first target Cube shard to store the above data set to be stored.
[0143] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0144] The present application also provides a computer program product. When executed on a data processing device, it is suitable for executing a program initialized with at least the following method steps:
[0145] Step S101, obtain a target start time, a target end time, and the start time and the end time of the above Cube shard in the above database. The target start time is the start point of the above time interval corresponding to the data set to be stored, and the target end time is the end point of the above time interval corresponding to the data set to be stored;
[0146] Step S102: Determine whether the start time of the first target Cube slice satisfies the first preset condition, and determine whether the end time of the first target Cube slice satisfies the second preset condition. The first preset condition is that the start time of the first target Cube slice is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube slice is greater than or equal to the target end time. The first target Cube slice is any one of the Cube slices.
[0147] Step S103: When the start time of the first target Cube slice satisfies the first preset condition and the end time of the first target Cube slice satisfies the second preset condition, delete the first target Cube slice in the database, and construct a second target Cube slice according to the start time and the end time of the first target Cube slice to store the dataset to be stored.
[0148] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0149] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the above unit division can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0150] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0151] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0152] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of this application. The aforementioned computer-readable storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0153] From the above description, it can be seen that the above embodiments of this application achieve the following technical effects:
[0154] 1) The incremental construction method of the Cube cube in the present application. The above Cube cube is a data structure for storing data in the database. The above Cube cube includes multiple Cube shards. The boundaries of the above Cube shards are marked with a start time and an end time. One of the above Cube shards is used to store a data set for a time interval. First, obtain the target start time, the target end time, and the start time and the end time of the above Cube shards in the above database. The above target start time is the starting point of the time interval corresponding to the data set to be stored, and the above target end time is the end point of the time interval corresponding to the data set to be stored. Then, determine whether the start time of the first target Cube shard satisfies a first preset condition, and determine whether the end time of the first target Cube shard satisfies a second preset condition. The above first preset condition is that the start time of the first target Cube shard is less than or equal to the above target start time, and the above second preset condition is that the end time of the first target Cube shard is greater than or equal to the above target end time. The above first target Cube shard is any one of the above Cube shards. Finally, when the start time of the first target Cube shard satisfies the first preset condition and the end time of the first target Cube shard satisfies the second preset condition, delete the first target Cube shard in the above database, and construct a second target Cube shard according to the start time and the end time of the first target Cube shard to store the above data set to be stored. This method first determines the time interval corresponding to the data set to be stored and the start and end times of the existing Cube shards in the database. Then, it determines whether there is a Cube shard in the database whose start and end times include the time interval corresponding to the data set to be stored. If so, reconstruct a second target Cube shard with the start and end times of this Cube shard to store the data set to be stored. This method solves the problem in the prior art that the merged Cube fragments cannot be reconstructed.
[0155] 2) The incremental construction device of the Cube cube in the present application. The above-mentioned Cube cube is a data structure used to store data in the database. The above-mentioned Cube cube includes multiple Cube shards. The boundaries of the above-mentioned Cube shards are marked with a start time and an end time. One of the above-mentioned Cube shards is used to store a data set for a time interval. An acquisition unit is used to acquire a target start time, a target end time, and the above-mentioned start time and the above-mentioned end time of the above-mentioned Cube shards in the above-mentioned database. The above-mentioned target start time is the starting point of the time interval corresponding to the data set to be stored, and the above-mentioned target end time is the end point of the time interval corresponding to the data set to be stored. A determination unit is used to determine whether the above-mentioned start time of the first target Cube shard meets a first preset condition, and determine whether the above-mentioned end time of the first target Cube shard meets a second preset condition. The above-mentioned first preset condition is that the above-mentioned start time of the first target Cube shard is less than or equal to the above-mentioned target start time, and the above-mentioned second preset condition is that the above-mentioned end time of the first target Cube shard is greater than or equal to the above-mentioned target end time. The above-mentioned first target Cube shard is any one of the above-mentioned Cube shards. A construction unit is used to, when the above-mentioned start time of the first target Cube shard meets the first preset condition and the above-mentioned end time of the first target Cube shard meets the second preset condition, delete the above-mentioned first target Cube shard in the above-mentioned database, and construct a second target Cube shard according to the above-mentioned start time and the above-mentioned end time of the first target Cube shard to store the above-mentioned data set to be stored. This device first determines the time interval corresponding to the data set to be stored and the start and end times of the existing Cube shards in the database. Then, it determines whether there is a Cube shard in the database whose start and end times include the time interval corresponding to the data set to be stored. If so, it reconstructs the second target Cube shard with the start and end times of this Cube shard to store the data set to be stored. This device solves the problem in the prior art that it is impossible to reconstruct the merged Cube fragments.
[0156] 3) The incremental construction system of the Cube cube of the present application includes: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include methods for performing any of the above-mentioned incremental construction methods of the Cube cube. The system first determines the time interval corresponding to the data set to be stored and the start and end times of the existing Cube shards in the database, and then determines whether there is a Cube shard in the database whose start and end times include the time interval corresponding to the data set to be stored. If so, it reconstructs the second target Cube shard with the start and end times of the Cube shard to store the data set to be stored. The system solves the problem in the prior art that the merged Cube fragments cannot be reconstructed.
[0157] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. An incremental construction method for a Cube cube, where the Cube cube is a data structure used to store data in a database. The Cube cube includes multiple Cube shards, and the boundaries of the Cube shards are marked with a start time and an end time. One Cube shard is used to store a data set for a time interval, and it is characterized in that The method includes: Obtaining a target start time, a target end time, and the start time and the end time of the Cube shards in the database, where the target start time is the start point of the time interval corresponding to the dataset to be stored, and the target end time is the end point of the time interval corresponding to the dataset to be stored; Determining whether the start time of the first target Cube shard meets a first preset condition, and determining whether the end time of the first target Cube shard meets a second preset condition. The first preset condition is that the start time of the first target Cube shard is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube shard is greater than or equal to the target end time. The first target Cube shard is any one of the Cube shards; In the case where the start time of the first target Cube shard meets the first preset condition and the end time of the first target Cube shard meets the second preset condition, deleting the first target Cube shard in the database and constructing a second target Cube shard according to the start time and the end time of the first target Cube shard to store the dataset to be stored; Obtaining the start time and the end time of the Cube shards in the database; A first obtaining step of obtaining a first target time and a second target time from the start time and the end time of a third target Cube shard. The third target Cube shard is any one of the Cube shards. The first target time is the month corresponding to the start time of the third target Cube shard, and the second target time is the month corresponding to the end time of the third target Cube shard; A first determining step of determining that the third target Cube shard needs to be split in the case where the first target time is different from the second target time; A second determining step of determining that the third target Cube shard does not need to be split in the case where the first target time is the same as the second target time; Sequentially executing the first obtaining step, the first determining step, and the second determining step at least once until the determination work of all Cube shards is completed, obtaining a plurality of Cube shards that need to be split and a plurality of Cube shards that do not need to be split.
2. The construction method according to claim 1, characterized in that After sequentially executing the first obtaining step, the first determining step, and the second determining step at least once until the determination work of all Cube shards is completed, obtaining a plurality of Cube shards that need to be split and a plurality of Cube shards that do not need to be split, the method further includes: Determining a plurality of Cube shard sets among the Cube shards that do not need to be split. One Cube shard set includes the Cube shards with the same month corresponding to the start time; Second obtaining step: obtaining a target Cube shard set, where the target Cube shard set is any one of the Cube shard sets; Sorting step: sorting the Cube shards in the target Cube shard set in ascending order of the day corresponding to the start time of the Cube shards; Third determining step: determining multiple groups of Cube shards to be merged according to the start time and the end time of the Cube shards in the sorted target Cube shard set, where for any two adjacent Cube shards in each group of Cube shards to be merged, the end time of the previous Cube shard is the same as the start time of the next Cube shard; Execute the second obtaining step, the sorting step, and the third determining step at least once in sequence until the determination work of all the Cube shard sets is completed, to obtain multiple groups of Cube shards to be merged; Merge all the Cube shards in each group of Cube shards to be merged to obtain a merged Cube shard and store it in the database, where one group of Cube shards to be merged corresponds to one merged Cube shard.
3. The construction method according to claim 1, characterized in that, After executing the first obtaining step, the first determining step, and the second determining step at least once in sequence until the determination work of all Cube shards is completed, obtaining multiple Cube shards that need to be split and multiple Cube shards that do not need to be split, the method further includes: Third obtaining step: obtaining a third target time, a fourth target time, a fifth target time, and a sixth target time, where the third target time is the month corresponding to the start time of a fourth target Cube shard, the fourth target time is the month corresponding to the end time of the fourth target Cube shard, the fifth target time is the day corresponding to the start time of the fourth target Cube shard, the sixth target time is the day corresponding to the end time of the fourth target Cube shard, and the fourth target Cube shard is any one of the Cube shards that need to be split; Fourth determining step: determining the start time and the end time of the split Cube shards according to the third target time, the fourth target time, the fifth target time, and the sixth target time, where the split Cube shards are the Cube shards after splitting the fourth target Cube shard, the month corresponding to the start time of the split Cube shards is the same as the month corresponding to the end time, and the months corresponding to the start times of any two split Cube shards are different; Splitting step: splitting the fourth target Cube shard according to the start time and the end time of the split Cube shards; Execute the third obtaining step, the fourth determining step, and the splitting step at least once in sequence until the splitting work of all the Cube shards that need to be split is completed.
4. The construction method according to claim 3, wherein, Splitting the fourth target Cube slice according to the start time and end time of the split Cube slice includes: Deleting the fourth target Cube slice from the database; Constructing the split Cube slice in the database according to the start time and end time of the split Cube slice.
5. The construction method according to claim 3, characterized in that, Determining the start time and end time of the split Cube slice according to the third target time, the fourth target time, the fifth target time, and the sixth target time includes: When the difference between the third target time and the fourth target time is equal to 1, determining the start time and end time of the two split Cube slices according to the third target time, the fourth target time, the fifth target time, and the sixth target time; When the difference between the third target time and the fourth target time is greater than 1, determining the start time and end time of multiple split Cube slices according to the third target time, the fourth target time, the fifth target time, and the sixth target time.
6. The construction method according to claim 1, characterized in that Based on the situation that the start time of the first target Cube slice does not meet the first preset condition or the end time of the first target Cube slice does not meet the second preset condition, after determining whether the start time of the first target Cube slice meets the first preset condition and determining whether the end time of the first target Cube slice meets the second preset condition, the method further includes: Constructing the second target Cube slice in the database according to the target start time and the target end time to store the dataset to be stored.
7. An incremental construction device for a Cube cube, where the Cube cube is a data structure for storing data in a database. The Cube cube includes multiple Cube slices, and the boundaries of the Cube slices are marked with a start time and an end time. One Cube slice is used to store a data set for a time interval, and it is characterized in that The device includes: An acquisition unit, configured to acquire a target start time, a target end time, and the start time and end time of the Cube slice in the database, where the target start time is the start point of the time interval corresponding to the dataset to be stored, and the target end time is the end point of the time interval corresponding to the dataset to be stored; A determination unit, configured to determine whether the start time of the first target Cube slice meets the first preset condition, and determine whether the end time of the first target Cube slice meets the second preset condition, where the first preset condition is that the start time of the first target Cube slice is less than or equal to the target start time, and the second preset condition is that the end time of the first target Cube slice is greater than or equal to the target end time, and the first target Cube slice is any one of the Cube slices; A construction unit, configured to delete the first target Cube shard in the database and construct a second target Cube shard according to the start time and the end time of the first target Cube shard when the start time of the first target Cube shard meets a first preset condition and the end time of the first target Cube shard meets a second preset condition, so as to store the data set to be stored; A first acquisition unit, configured to acquire the start time and the end time of the Cube shard in the database; A second acquisition unit, configured to execute a first acquisition step to acquire a first target time and a second target time from the start time and the end time of a third target Cube shard, where the third target Cube shard is any one of the Cube shards, the first target time is the month corresponding to the start time of the third target Cube shard, and the second target time is the month corresponding to the end time of the third target Cube shard; A first determination unit, configured to execute a first determination step to determine that the third target Cube shard needs to be split when the first target time is different from the second target time; A second determination unit, configured to execute a second determination step to determine that the third target Cube shard does not need to be split when the first target time is the same as the second target time; A first iteration unit, configured to execute the first acquisition step, the first determination step, and the second determination step at least once in sequence until the determination work of all Cube shards is completed, so as to obtain a plurality of Cube shards that need to be split and a plurality of Cube shards that do not need to be split.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where when the program is executed by a processor, the processor executes the incremental construction method of the Cube cube according to any one of claims 1 to 6.
9. An incremental construction system for a Cube cube, characterized in that, Including: One or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include the incremental construction method of the Cube cube according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-dimensional calculation method and device
CN111026817A
Method for retrieving data object based on spatial-temporal database
US20190266138A1