Methods, apparatuses, and computing devices for processing a dataset

By dividing the time-series dataset into data shards and performing materialization operations, the problem of low efficiency in updating and querying materialized views is solved, achieving efficient and real-time materialized view management and reducing resource waste and redundant calculations.

CN116340322BActive Publication Date: 2026-05-19上海炎凰数据科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
上海炎凰数据科技有限公司
Filing Date
2023-03-24
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency in updating materialized views and querying in real time when processing time-series datasets, resulting in wasted resources and insufficient real-time performance of computation results.

Method used

The time series data is divided into multiple data shards in chronological order, and materialization operations are performed on each shard to create materialized view shards. The materialized view is updated only when a new data shard is created, unnecessary shards are deleted, the modified shards are rebuilt, and the materialized view can be queried precisely for any time range.

Benefits of technology

It enables the creation and updating of materialized views with fewer resources, improves the real-time performance and accuracy of queries, reduces redundant calculations, and saves resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116340322B_ABST
    Figure CN116340322B_ABST
Patent Text Reader

Abstract

A method for processing a dataset including a plurality of time-series data is provided. The method includes: dividing the plurality of time-series data into a plurality of data shards in a time order according to a predetermined rule, each of the plurality of data shards including a portion of the time-series data in the dataset; and performing a materialization operation on a number of data shards of the plurality of data shards respectively to obtain a corresponding number of materialized view shards. The method also provides a corresponding materialized view query method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of dataset processing, and more particularly to methods for processing time-series datasets, apparatus for processing time-series datasets, computing devices, non-transitory computer-readable storage media, and computer program products. Background Technology

[0002] A time-series dataset is a collection of data with time attributes or timestamps. Time-series datasets are widely used across various industries. Typical time-series datasets include real-time CPU (or other resource) utilization of servers, sensor data, etc. The characteristics of time-series datasets are large data volumes, continuous writing, and generally no modification after writing.

[0003] Analyzing time-series datasets requires accessing large amounts of data and performing extensive computations, which is very time-consuming. Furthermore, the analysis of the same time-series dataset may involve redundant computations. Materialized views provide a pre-computation method, saving the results of time-consuming operations for direct reuse during queries, ultimately accelerating the query process. By accessing materialized views, users avoid real-time computation and directly obtain pre-computed results, thus speeding up the overall process. Most relational and non-relational databases on the market have developed materialized view functionality.

[0004] Currently, as time-series datasets continue to increase over time, there is still significant room for improvement in the materialized views used to achieve efficient updates and real-time queries.

[0005] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0006] This disclosure provides a method for processing a dataset, wherein the dataset includes multiple time-series data, and the method includes: dividing the multiple time-series data into multiple data slices according to a predetermined rule in chronological order, each data slice including a portion of the time-series data in the dataset; and performing materialization operations on several data slices in the multiple data slices to obtain corresponding several materialized view slices.

[0007] According to one aspect of this disclosure, an apparatus for processing a dataset is provided, wherein the dataset includes a plurality of data arranged in time, the apparatus comprising: a first module for dividing the dataset into a plurality of data shards according to a predetermined rule, such that each of the plurality of data shards includes a portion of the data in the dataset; and a second module for performing a materialization operation on each of the plurality of data shards to obtain a corresponding plurality of materialized view shards.

[0008] According to another aspect of this disclosure, a computing device is provided, comprising: a processor; and a memory storing instructions thereon, which, when executed by the processor, cause the processor to perform the method as described above.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing instructions, wherein the instructions, when executed by a processor, cause the processor to perform the method described above.

[0010] According to another aspect of this disclosure, a computer program product is provided, including instructions that, when executed by a processor, cause the processor to perform the method as described above.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0013] Figure 1 A schematic diagram illustrates a method for processing time-series datasets by creating materialized views, based on relevant technologies;

[0014] Figure 2 An update based on the relevant technology is shown. Figure 1 A schematic diagram of the method for creating a materialized view in the process;

[0015] Figure 3 A flowchart of a method for processing a time-series dataset according to an example embodiment of this disclosure is shown;

[0016] Figure 4 A schematic diagram illustrating the process of partitioning data according to an exemplary embodiment of this disclosure is shown;

[0017] Figure 5 A schematic diagram illustrating the process of creating a materialized view fragment according to an exemplary embodiment of this disclosure is shown;

[0018] Figure 6 A flowchart is shown of a method for updating a materialized view fragment according to an exemplary embodiment of this disclosure;

[0019] Figure 7 A schematic diagram of a method for updating materialized view fragments according to an exemplary embodiment of this disclosure is shown;

[0020] Figure 8 A flowchart is shown of a method for querying a materialized view according to an example embodiment of this disclosure;

[0021] Figure 9 A schematic diagram of a method for querying a materialized view according to an exemplary embodiment of this disclosure is shown;

[0022] Figure 10 A schematic diagram illustrating a method for dividing a materialized view into multiple time buckets according to an exemplary embodiment of this disclosure; and

[0023] Figure 11 A structural block diagram of an apparatus for processing a dataset according to an exemplary embodiment of the present disclosure is shown. Detailed Implementation

[0024] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present disclosure, including various details of these embodiments to aid understanding; however, these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0025] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0026] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof. The term "based on" should be interpreted as "at least partially based on".

[0027] Time series data often grows infinitely over time. When processing such time series datasets, dividing the data into multiple data fragments according to time sequence can facilitate the implementation of subsequent solutions.

[0028] In related technologies, such as Figure 1 As shown, materialized views 120 can be created for multiple data shards 110 to save pre-computed results for easy subsequent queries. However, the created materialized view 120 is a static snapshot of the data shard materialization operation. That is, as time-series data increases over time, developers need to actively initiate or define periodic tasks to update the materialized view. The process of updating a materialized view in related technologies is as follows... Figure 2 As shown, when a new data shard 210 is imported, updating the materialized view requires recreating the materialized view 220 containing the new data shard 210. This process not only requires recalculating the new data shard 210 but also recalculating all historical data shards 110. Repeatedly calculating all historical data each time leads to unnecessary resource waste. Furthermore, the cached pre-calculated results are static data snapshots, resulting in insufficient real-time availability of calculation results for users.

[0029] The inventors recognized that time-series data rarely changes from past times. Even after sharding the data by time and performing preprocessing calculations on the historical data shards, the results of the preprocessing calculations generally remain unchanged. Therefore, this disclosure fully utilizes these characteristics of time-series data and proposes a method for processing time-series datasets. This method can create and update materialized views of time-series datasets with fewer resources and can accurately and in real-time query materialized views for any time range.

[0030] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0031] 1. About creating materialized views

[0032] Figure 3A flowchart of a method 300 for processing a time-series dataset according to an example embodiment of this disclosure is shown. A time-series dataset refers to a collection of data including multiple time-series data. Time-series data refers to data with time attributes or timestamps. Typical time-series datasets include real-time CPU (or other resource) utilization of a server, sensor-collected information, online store transaction data, etc.

[0033] like Figure 3 As shown, the method 300 for processing a time-series dataset according to an example embodiment of this disclosure includes the following steps:

[0034] In step 310, the multiple time-series data are divided into multiple data shards according to a predetermined rule and in chronological order. Each data shard comprises a portion of the time-series data in the dataset. An example of the process of dividing the data into multiple shards is shown below. Figure 4 As shown. In this example, the data in the original dataset grows over time. The dataset, arranged along a timeline, is divided into multiple data shards 410 according to predetermined rules. The amount of data or the time range of each data shard may not be exactly the same. In some embodiments, the rule for dividing the data shards may be a predetermined time interval, such that the data shards contain data within a predetermined time range. For example, if the time-series dataset is transaction data from an online store, the resulting data shards may contain transaction data from the online store within one hour. In other embodiments, the rule for dividing the data shards may be the size of the amount of data contained in the data shards, such that the data shards contain a predetermined amount of data. For example, if the time-series dataset is temperature information collected by a sensor, the resulting data shards may contain 50 temperature data points collected by the sensor. It should be understood that other predetermined rules can also be selected to divide the data shards according to actual needs.

[0035] In step 320, materialization operations are performed on several data shards from multiple data shards to obtain corresponding materialized view shards. In some examples, the materialization operation is the preprocessing of time-series data in the data shards, including at least one of aggregation, filtering, and field projection of the time-series data. The following section combines... Figure 5 This will illustrate the process of creating materialized view fragments according to embodiments of the present disclosure. For example... Figure 5 As shown, after dividing the original time-series data into multiple data shards 510, a corresponding materialized view shard 520 is created for each data shard 510. When creating a materialized view, if it is necessary to materialize historical data shards that already exist as materialized view shards, then each historical data shard needs to be materialized. Since the historical data shards of the time series remain essentially unchanged, when creating a materialized view of the time-series data, the materialized view shards corresponding to the historical data shards will also remain unchanged and can be directly used without needing to be recreated.

[0036] In some embodiments, users can create materialized views using SQL and specify whether historical data needs to be materialized. The SQL statement can be written using common patterns in the prior art. Example SQL statement for creating a materialized view:

[0037] "CREATE MATERIALIZED VIEW demo_view

[0038] AS

[0039] SELECT CustomerID,ItemID,SUM(ItemCount),AVG(ItemPrice)

[0040] FROM orders GROUP BY CustomerID,ItemID

[0041] WITH NO DATA

[0042] Among them, "WITH NO DATA" can be used to determine whether historical data needs to be materialized.

[0043] It should be understood that users can also implement the methods for processing time-series datasets according to the exemplary embodiments of this disclosure using any language.

[0044] 2. Updates and maintenance of materialized views

[0045] based on Figure 3 The method 300 shown for processing time series datasets is able to update the materialized view of time series data in real time with fewer resources. Figure 6 A flowchart of a method 600 for updating a materialized view fragment according to an exemplary embodiment of this disclosure is shown. The method 600 for updating a materialized view fragment includes steps 610 of adding a new data fragment and its materialized view fragment, steps 620 of deleting an unnecessary data fragment and its materialized view fragment, and steps 630 of modifying the data fragment and its materialized view fragment. The following is in conjunction with… Figure 7 This describes the process of updating materialized view fragments according to embodiments of the present disclosure.

[0046] In step 610, as data continues to be imported, new data fragments and their materialized view fragments are added. Step 610 may specifically include the following operations:

[0047] (1a) In response to the import of new time-series data into the dataset, the new time-series data is divided into at least one new data fragment according to a predetermined rule, each of the at least one new data fragment including a portion of the new time-series data; and

[0048] (1b) Perform materialization operations on at least one new data fragment to obtain at least one corresponding materialized view fragment.

[0049] refer to Figure 7 Because materialized views are created for individual data shards rather than all time-series data, step 610 can create only the materialized view shard 720 for the newly imported data shard 710, provided the original data shard 510 remains unchanged. This eliminates the need to recalculate the materialized view of the previously calculated data shard 510, thus enabling the creation of materialized views of time-series data with fewer resources.

[0050] In step 620, when outdated data no longer needs to be retained, unnecessary data fragments and their materialized view fragments can be deleted. Step 620 may specifically include the following operations:

[0051] (2a) In response to deleting at least one of the plurality of data fragments, delete at least one materialized view fragment corresponding to the at least one data fragment.

[0052] refer to Figure 7 Because the data is sharded and materialized view shards are created for each shard, only the unnecessary data shard 711 and its materialized view shard 721 can be deleted. After deletion, there is no need to re-materialize the retained data, thus saving resources.

[0053] In step 630, when historical data changes, the corresponding materialized view fragment can be recreated to replace the existing materialized view fragment. Step 620 may specifically include the following operations:

[0054] (3a) In response to changing the data in at least one of the plurality of data fragments, materialize the changed at least one data fragment to replace the corresponding at least one materialized view fragment.

[0055] refer to Figure 7 Because data in historical data shard 712 has changed or some data has become unavailable, the previously created corresponding materialized view shard 722 is no longer available. In this case, the materialized view shard can be recreated to replace the previously existing materialized view shard 722. Since it's unnecessary to materialize all data due to individual data changes or unavailability, resources are saved.

[0056] 3. Queries regarding materialized views

[0057] based on Figure 3 The method 300 shown for processing time-series datasets is also capable of accurately and in real-time querying materialized views of any time range. Figure 8A flowchart of a method 800 for querying a materialized view according to an example embodiment of this disclosure is shown. The method 800 for querying a materialized view includes the following steps:

[0058] In step 810, a data query request is received. The data query request specifies the time range of the materialized view to be queried.

[0059] In step 820, based on the time range to be queried, a set of data shards to be queried is determined from the data shards currently existing in the dataset, wherein the set of data shards to be queried covers the time range.

[0060] In step 830, for each query data shard in the query data shard set, the corresponding query materialized view shard is obtained.

[0061] It should be understood that since the time range to be queried may not necessarily correspond to a complete data shard but may only cover a portion of the data in a partial data shard, the corresponding materialized view shard may also not correspond to a complete materialized view shard but only a portion of a materialized view shard. Therefore, query data shards covering the desired time range can exist in three types: the query data shard has no corresponding materialized view shard, the query data shard corresponds to a complete materialized view shard, and the query data shard corresponds to only a portion of a materialized view shard. For these three types of query data shards, the following corresponding operations should be used to obtain the corresponding query materialized view shards:

[0062] (4a) In response to the fact that there is no corresponding materialized view fragment for the query data fragment, materialize the data fragment to obtain the corresponding query materialized view fragment;

[0063] (4b) In response to the query data shard having a corresponding materialized view shard and the corresponding materialized view shard falling entirely within the time range to be queried, the entirety of the corresponding materialized view shard is taken as the corresponding query materialized view shard.

[0064] (4c) In response to the fact that the queried data shard has a corresponding materialized view shard but the corresponding materialized view shard only partially falls within the time range to be queried, the portion of the corresponding materialized view shard that falls within the time range is taken as the corresponding queried materialized view shard.

[0065] The following is for reference. Figure 9 This explains the process of obtaining corresponding materialized view shards based on different types of query data shards.

[0066] like Figure 9As shown, for the desired time range 900, if query data shard 910 is in the importing state and has no corresponding materialized view shard, then materialization is performed on this data shard to obtain the corresponding query materialized view shard P2. If query data shard 911 has a corresponding materialized view shard 921 and this corresponding materialized view shard 921 completely falls within the desired time range 900, then the entirety of this corresponding materialized view shard 921 is used as the corresponding query materialized view shard P1. If query data shard 912 has a corresponding materialized view shard 922, but this corresponding materialized view shard 922 only partially falls within the desired time range 900, then the portion of this corresponding materialized view shard that falls within this time range is used as the corresponding query materialized view shards P3 and P4.

[0067] It should be understood that if a query data shard has a corresponding materialized view shard, but the corresponding materialized view shard only partially falls within the time range to be queried, the original data portion corresponding to the query data shard needs to be recalculated to obtain the corresponding query materialized view shard. To further reduce redundant calculations, in some embodiments, step 830 may also include the following time bucket division operation:

[0068] (4d) Divide the materialized view fragments into multiple time buckets in chronological order at predetermined time intervals. For each time bucket:

[0069] In response to the time bucket falling entirely within the time range of the query, the time bucket is used as a shard of the materialized view for the query; or

[0070] In response to the fact that only a portion of the time bucket falls within the time range to be queried, materialization operations are performed on the original data corresponding to the time bucket to obtain the corresponding materialized view shards for the query.

[0071] Further reference below Figure 10 Explain the process of dividing time into buckets. For example... Figure 10 As shown, the materialized view fragment 1100 is divided into multiple time buckets 1110 in chronological order at predetermined time intervals tb1, tb2, tb3...tbn. The size of the time bucket (i.e., the predetermined time intervals tb1, tb2, tb3...tbn) needs to be customized at creation time. In some embodiments, the predetermined time interval is determined based on the time range of the materialized view to be queried or the density of data import. In some embodiments, the predetermined time interval is an equal time interval.

[0072] Combination Figure 9If the time bucket of the materialized view shard falls entirely within the desired time range of 900, the time bucket can be directly used as the materialized view shard for the query, for example, materialized view shard P4. If the time bucket only partially falls within the desired time range, the original data corresponding to the time bucket is materialized to obtain the corresponding materialized view shard P3.

[0073] Back Figure 8 In step 840, the materialized view fragments corresponding to each query data fragment in the query data fragment set are combined to obtain the materialized view to be queried, for example... Figure 9 The expression P1+P2+P3+P4 is shown in the figure.

[0074] Figure 11 A structural block diagram of an apparatus for processing a dataset according to an exemplary embodiment of the present disclosure is shown. The apparatus 1200 includes a first module 1210 and a second module 1220.

[0075] The first module 1210 is used to divide the dataset into multiple data shards according to predetermined rules, such that each data shard includes a portion of the dataset.

[0076] The second module 1220 is used to perform materialization operations on each of the plurality of data fragments to obtain the corresponding plurality of materialized view fragments.

[0077] It should be understood that Figure 11 The various modules of the device 1200 shown can be connected to the reference. Figure 3 The steps in method 300 described correspond to each other. Therefore, the operations, features, and advantages described above for method 300 also apply to apparatus 1200 and its included modules. For the sake of brevity, some operations, features, and advantages will not be repeated here.

[0078] While specific functions have been discussed above with reference to specific modules, it should be noted that the functions of the modules discussed herein can be divided into multiple modules, and / or at least some functions of multiple modules can be combined into a single module. The specific actions performed by the modules discussed herein include the specific module itself performing the action, or alternatively, the specific module calling or otherwise accessing another component or module that performs the action (or performs the action in conjunction with the specific module). Therefore, a specific module performing an action can include the specific module performing the action itself and / or another module that performs the action, called or otherwise accessed by the specific module.

[0079] It should also be understood that this article can describe various technologies in the general context of software and hardware components or program modules. The above regarding... Figure 11The various modules described can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these modules can be implemented as computer program code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these modules can be implemented as hardware logic / circuit. For example, in some embodiments, one or more of these modules can be implemented together in a system-on-a-chip (SoC). The SoC may include an integrated circuit chip (which includes a processor (e.g., a central processing unit (CPU), microcontroller, microprocessor, digital signal processor (DSP), etc.), memory, one or more communication interfaces, and / or one or more components of other circuitry) and may optionally execute received program code and / or include embedded firmware to perform functions.

[0080] According to embodiments of the present disclosure, a computing device is also provided. The computing device includes a processor; and a memory storing instructions thereon that, when executed by the processor, cause the processor to perform the methods described in any embodiment of the present disclosure.

[0081] According to embodiments of this disclosure, a non-transitory computer-readable storage medium storing instructions is also provided. When executed by a processor, the instructions cause the processor to perform the methods described in any embodiment of this disclosure.

[0082] According to embodiments of the present disclosure, a computer program product is also provided. The computer program product includes instructions that, when executed by a processor, cause the processor to perform the methods described in any embodiment of the present disclosure.

[0083] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0084] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of this disclosure is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A method for processing datasets, wherein, The dataset includes multiple time-series data sets, and the method includes: The multiple time-series data are divided into multiple data slices according to a predetermined rule in chronological order, and each data slice includes a portion of the time-series data in the dataset; Materialization operations are performed on several data fragments among the multiple data fragments to obtain corresponding materialized view fragments; Receive a data query request, wherein the data query request specifies the time range of the materialized view to be queried; Based on the time range, a set of query data shards is determined from the currently existing data shards in the dataset, wherein the set of query data shards covers the time range; For each query data shard in the query data shard set: In response to the absence of a corresponding materialized view shard for the query data shard, a materialization operation is performed on the data shard to obtain the corresponding query materialized view shard; and In response to the query data shard having a corresponding materialized view shard, at least the corresponding portion of that materialized view shard is used as the corresponding query materialized view shard; and The materialized view fragments corresponding to each query data fragment in the query data fragment set are combined to obtain the materialized view to be queried.

2. The method of claim 1, further comprising: In response to the import of new time-series data into the dataset, the new time-series data is divided into at least one new data fragment according to the predetermined rules, and each new data fragment in the at least one new data fragment includes a portion of the time-series data in the new time-series data; as well as Materialization operations are performed on each of the at least one newly added data fragment to obtain at least one corresponding materialized view fragment.

3. The method of claim 1, further comprising: In response to deleting at least one of the plurality of data shards, at least one materialized view shard corresponding to the at least one data shard is deleted.

4. The method of claim 1, further comprising: In response to a change in data in at least one of the plurality of data shards, a materialization operation is performed on the changed at least one data shard to replace the corresponding at least one materialized view shard.

5. The method of claim 1, wherein, The step of using at least a corresponding portion of the corresponding materialized view fragment as the corresponding query materialized view fragment includes: In response to the fact that the corresponding materialized view fragment falls entirely within the time range, the entirety of the corresponding materialized view fragment is used as the corresponding query materialized view fragment; and In response to the fact that the corresponding materialized view fragment falls only partially within the time range, the portion of the corresponding materialized view fragment that falls within the time range is taken as the corresponding queried materialized view fragment.

6. The method of claim 5, wherein, The portion of the corresponding materialized view fragment that falls within the time range is taken as the corresponding query materialized view fragment, including: The corresponding materialized view fragment is divided into multiple time buckets in chronological order at predetermined time intervals; and For each of the multiple time buckets: In response to the time bucket falling entirely within the stated time range, the time bucket is used as a shard of the query materialized view; and In response to the fact that the time bucket only partially falls within the time range, materialize the original data corresponding to the time bucket to obtain the corresponding query materialized view shard.

7. The method of claim 6, wherein, The predetermined time interval is determined based on the time range of the materialized view to be queried or the density of data import.

8. The method of claim 7, wherein, The predetermined time interval is an equal time interval.

9. The method according to any one of claims 1-8, wherein, The predetermined rules include the size of the data contained in the predetermined time interval or data slice.

10. The method according to any one of claims 1-8, wherein, The materialization operations include at least one of the following: aggregation, filtering, field projection, and joining of time-series data.

11. An apparatus for processing a dataset, wherein, The dataset includes multiple datasets arranged chronologically, and the device includes: The first module is used to divide the dataset into multiple data shards according to predetermined rules, such that each data shard includes a portion of the dataset; and The second module is used to perform materialization operations on each of the multiple data fragments to obtain corresponding multiple materialized view fragments. The third module is used to receive data query requests, wherein the data query requests specify the time range of the materialized view to be queried; The fourth module is used to determine a set of query data shards from the currently existing data shards in the dataset based on the time range, wherein the set of query data shards covers the time range; The fifth module is used for each query data shard in the query data shard set: In response to the absence of a corresponding materialized view shard for the query data shard, a materialization operation is performed on the data shard to obtain the corresponding query materialized view shard; and In response to the query data shard having a corresponding materialized view shard, at least the corresponding portion of that materialized view shard is used as the corresponding query materialized view shard; and The sixth module is used to combine the materialized view fragments corresponding to each query data fragment in the query data fragment set to obtain the materialized view to be queried.

12. A computing device, comprising: processor; as well as A memory storing instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 10.

13. A non-transitory computer-readable storage medium storing instructions, wherein, When executed by a processor, the instructions cause the processor to perform the method according to any one of claims 1 to 10.

14. A computer program product comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 10.