Data service system and method

The data service system addresses inefficiencies in data virtualization by integrating time series data from multiple sources using virtual tables and query division, enhancing search processing efficiency by avoiding redundant data acquisition.

JP7731820B2Active Publication Date: 2025-09-01KK TOSHIBA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022022160
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-09-01
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

Existing data virtualization systems face inefficiencies in search processing due to duplicate data management across multiple data sources, leading to prolonged processing times and difficulty in accessing data without knowing the source.

Method used

A data service system that utilizes virtual tables to manage and integrate time series data from multiple sources, including a first data source for real-time data and a second data source for batch-processed data, with boundary management and query division mechanisms to avoid redundant data acquisition.

Benefits of technology

Enables efficient search processing by identifying overlapping data ranges and dividing queries to access data from appropriate sources, reducing redundancy and improving processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007731820000001
    Figure 0007731820000001
  • Figure 0007731820000002
    Figure 0007731820000002
  • Figure 0007731820000003
    Figure 0007731820000003
Patent Text Reader

Abstract

To provide a data service system allowed to realize efficient retrieval processing, and a method.SOLUTION: There is provided according to an embodiment a data service system including first and second data sources, in which the data service system comprises storage means, acquiring means, dividing means and aggregating means. The storage means stores boundary management data indicative of a boundary condition defined based on an overlap relationship of first and second time ranges corresponding to first and second time-series data in first and second databases. The acquiring means acquires a first query including a first time condition. The dividing means divides the first query into second and third queries based on the boundary management data and the first time condition included in the first query. The aggregating means aggregates the first and second time-series data acquired from the first and second data sources based on the second and third queries.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION Embodiments of the present invention relate to a data service system and method. [Background technology]

[0002] In recent years, a technology called data virtualization has become known, which virtually integrates data managed in multiple data sources without duplicating them, providing data that can be used in fields such as business. In this data virtualization, a virtual table corresponding to an external data source is stored in a search engine, allowing integrated search processing of multiple external data sources via the virtual table.

[0003] Here, there are cases where the same data is managed in duplicate in multiple data sources (i.e., some of the data managed in each of the multiple data sources is duplicated), and if the same data is searched (obtained) from each of the multiple data sources, efficient search processing cannot be achieved. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 6371136 Summary of the Invention [Problem to be solved by the invention]

[0005] Therefore, an object of the present invention is to provide a data service system and method that can realize efficient search processing. [Means for solving the problem]

[0006] According to an embodiment, there is provided a data service system that includes a first data source that manages first time series data corresponding to a first time range, and a second data source that manages second time series data corresponding to a second time range that at least partially overlaps with the first time range, and that executes search processing using virtual tables corresponding to the first and second data sources. a third data source; and Storage means, acquisition means, division means, and aggregation means and To be equipped. The third data source manages third time series data in which at least a portion of the second time series data is updated. The storage means stores boundary management data indicating a temporal boundary condition for the first and second time series data, the boundary condition being determined based on an overlapping relationship between the first and second time ranges. and update management data indicating an eighth time range corresponding to the third time series data. The acquiring means acquires a first query for searching data from the first and second data sources, the first query including a first time condition indicating a third time range corresponding to data to be searched based on the first query. The dividing means divides the boundary management data stored in the storage means. Ta and Beauty Update management data; a first time condition included in the acquired first query; and and generating a second query including a second time condition for the first data source based on the first query. 、 a third query including a third time condition for the second data source; and a fourth query including a fourth time condition for the third data source. The aggregation means divides the first time series data acquired from the first data source based on the second query into second time series data acquired from the second data source based on the third query, third time series data obtained from the third data source based on the fourth query; Aggregate the do. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram for explaining a lambda architecture, which is one of the mechanisms for constructing a data service system according to the first embodiment. [Figure 2] FIG. 4 is a diagram illustrating an example of an overlapping relationship between first and second time ranges. [Figure 3] FIG. 1 is a diagram illustrating an overview of a data service system. [Figure 4] FIG. 2 is a diagram showing an example of the hardware configuration of an integrated search engine server. [Figure 5] FIG. 1 is a block diagram showing an example of the functional configuration of a data service system according to a comparative example of the present embodiment. [Figure 6] FIG. 1 is a diagram showing an example of the functional configuration of a data service system according to an embodiment of the present invention. [Figure 7] 10 is a flowchart showing an example of a processing procedure of an integrated search engine server included in the data service system. [Figure 8] 10 is a flowchart showing an example of a processing procedure for a query division process. [Figure 9] FIG. 10 is a diagram showing an example of a target query. [Figure 10] FIG. 4 is a diagram showing an example of the data structure of boundary management data. [Figure 11] FIG. 10 is a diagram showing an example of a time range for each virtual table. [Figure 12] FIG. 10 is a diagram showing an example of a time condition for each virtual table. [Figure 13] FIG. 10 is a diagram showing an example of a query for a time-series DB. [Figure 14] FIG. 10 is a diagram showing an example of a query for object storage. [Figure 15] FIG. 10 is a diagram showing an example of first time-series data acquired based on a query for a time-series DB. [Figure 16] FIG. 10 is a diagram showing an example of second time-series data acquired based on a query for object storage. [Figure 17] FIG. 10 is a diagram showing an example of a result of aggregating first and second time series data. [Figure 18] 5A and 5B are diagrams for explaining an example of improving the efficiency of search processing achieved in the present embodiment. [Figure 19] 5A and 5B are diagrams for explaining an example of improving the efficiency of search processing achieved in the present embodiment. [Figure 20] 5A and 5B are diagrams for explaining an example of improving the efficiency of search processing achieved in the present embodiment. [Figure 21]FIG. 1 is a diagram for explaining an example of a technology related to a data service system according to an embodiment of the present invention. [Figure 22] FIG. 2 is a diagram for explaining a first specific example of the present embodiment. [Figure 23] FIG. 2 is a diagram for explaining a first specific example of the present embodiment. [Figure 24] FIG. 2 is a diagram for explaining a first specific example of the present embodiment. [Figure 25] FIG. 2 is a diagram for explaining a first specific example of the present embodiment. [Figure 26] FIG. 2 is a diagram for explaining a first specific example of the present embodiment. [Figure 27] FIG. 10 is a diagram for explaining a second specific example of the present embodiment. [Figure 28] FIG. 10 is a diagram for explaining a second specific example of the present embodiment. [Figure 29] FIG. 10 is a diagram for explaining a second specific example of the present embodiment. [Figure 30] FIG. 10 is a diagram for explaining a second specific example of the present embodiment. [Figure 31] FIG. 10 is a diagram for explaining a second specific example of the present embodiment. [Figure 32] FIG. 10 is a diagram for explaining a third specific example of the present embodiment. [Figure 33] FIG. 10 is a diagram for explaining a third specific example of the present embodiment. [Figure 34] FIG. 10 is a diagram for explaining a third specific example of the present embodiment. [Figure 35] FIG. 10 is a diagram for explaining a third specific example of the present embodiment. [Figure 36] FIG. 10 is a diagram for explaining a third specific example of the present embodiment. [Figure 37] FIG. 10 is a diagram showing another example of the data structure of boundary management data. [Figure 38] FIG. 10 is a diagram showing yet another example of the data structure of boundary management data. [Figure 39] FIG. 10 is a diagram showing an example of the functional configuration of a data service system according to a second embodiment. [Figure 40] 10 is a flowchart showing an example of a processing procedure for a query division process. [Figure 41] FIG. 4 is a diagram showing an example of the data structure of update management data. [Figure 42] FIG. 10 is a diagram showing an example of a time condition for each virtual table. [Figure 43] FIG. 10 is a diagram showing an example of a query for object storage A. [Figure 44] FIG. 10 is a diagram showing an example of a query for object storage B. [Figure 45] FIG. 10 is a diagram showing an example of second time-series data acquired based on a query for object storage A. [Figure 46] FIG. 10 is a diagram showing an example of third time-series data acquired based on a query for object storage B. [Figure 47] FIG. 10 is a diagram showing an example of a result of aggregating first to third time series data. [Figure 48] FIG. 10 is a diagram for explaining a specific example of the present embodiment. [Figure 49] FIG. 10 is a diagram for explaining a specific example of the present embodiment. [Figure 50] FIG. 10 is a diagram for explaining a specific example of the present embodiment. [Figure 51] FIG. 10 is a diagram for explaining a specific example of the present embodiment. [Figure 52] FIG. 10 is a diagram for explaining a specific example of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, each embodiment will be described with reference to the drawings. (First embodiment) The data service system according to the first embodiment includes, for example, a first data source that manages first time series data corresponding to a first time range and a second data source that manages second time series data corresponding to a second time range that at least partially overlaps with the first time range, and provides a service that provides search results for the time series data by performing an integrated search process using virtual tables corresponding to the first and second data sources. In this embodiment, the data source refers to a database or storage (service) that stores (holds) time series data and allows access to the time series data. In this embodiment, the time series data refers to a collection of data annotated with information related to time.

[0009] Here, with reference to FIG. 1, a brief description will be given of the lambda architecture, which is one of the mechanisms for constructing a data service system according to this embodiment.

[0010] The Lambda architecture is an architecture (data analysis platform) for processing massive amounts of big data, and consists of a first layer that includes a batch layer and a serving layer, and a second layer that includes a speed layer.

[0011] The data (new data) supplied to the first layer is managed in the batch layer, and batch processing is performed on the large amount of data (master dataset) managed in the batch layer. In the serving layer, the results of the batch processing performed in the batch layer can be provided as a batch view.

[0012] On the other hand, when new data is supplied to the second layer, real-time processing of the data is performed in the speed layer. As a result, a real-time view is created in the speed layer and can be provided. Note that the second layer (speed layer) can process data that cannot be processed due to the latency of the batch layer (i.e., data that has not yet been provided as a batch view).

[0013] The above-mentioned batch view and real-time view are acquired (searched) by specifying a query, and the acquired results (batch view and real-time view) can be merged and handled.

[0014] When the above-described lambda architecture is applied to the data service system according to this embodiment, for example, the second layer corresponds to the first data source, and the real-time view (real-time processing result) corresponds to the first time-series data managed in the first data source. Also, for example, the first layer corresponds to the second data source, and the batch view (batch processing result) corresponds to the second time-series data managed in the second data source.

[0015] In this embodiment, the first and second time series data are assumed to be a collection of sensor data (data measured by various sensors) collected at predetermined intervals within a predetermined time range, but may be data other than sensor data as long as it corresponds to the predetermined time range.

[0016] As described above, it is expected that the first time range corresponding to the first time series data (e.g., real-time view) managed in the first data source and the second time range corresponding to the second time series data (e.g., batch view) managed in the second data source will partially overlap.

[0017] Here, an example of the overlapping relationship between the first and second time ranges will be described with reference to Figure 2. The example shown in the upper part of Figure 2 indicates that a first time range 1 (e.g., the temporal range in which the first time series data was collected) corresponding to first time series data managed in a first data source is after January 1, 2021. Also, it indicates that a second time range 2a (e.g., the temporal range in which the second time series data was collected) corresponding to second time series data managed in a second data source is before January 2, 2021.

[0018] In this case, for example, the time series data corresponding to time range 3a up to January 1, 2021 (time series data collected during that time range) is included in both the first and second time series data, and can be said to overlap between the first and second data sources (first and second time series data).

[0019] As shown in the lower part of Figure 2, for example, when the above-mentioned batch processing is newly executed (i.e., a new batch view is added to the second data source as second time series data), and time range 2a corresponding to the second time series data becomes time range 2b, the time range corresponding to the time series data that overlaps between the first and second data sources changes from time range 3a to time range 3b.

[0020] Although not shown in Figure 2, for example, when a new real-time view is added to the first data source as first time series data, and part of the first time series data (e.g., past data corresponding to the time range up to January 1, 2021) is deleted, the time range corresponding to the overlapping time series data between the first and second data sources also changes. Note that deleting past data (i.e., old data) according to the period for which time series data is managed (retained) (or the period itself) is referred to as retention.

[0021] That is, the time range corresponding to the overlapping time-series data between the first and second data sources (i.e., the range of overlapping data) changes dynamically depending on the progress of batch processing (batch view generation processing) or the implementation of retention, etc.

[0022] Here, it is assumed that the first and second time-series data acquired (searched) from the first and second data sources based on a query specified in the data service system are merged and provided.

[0023] A query may include, for example, a time condition indicating a time range corresponding to the data to be searched based on the query. For example, if a query including a time condition indicating a time range from December 31, 2021 to before January 3, 2021, as shown in Figure 2, is specified, time series data that overlaps between the first and second data sources (duplicate data) will be obtained from each of the first and second data sources.

[0024] When the same data is acquired redundantly from multiple data sources (first and second data sources) in this way, it cannot be said that an efficient search process is being performed, and there is a concern that the processing time based on the data will be long. Furthermore, to achieve an efficient search process, it is sufficient to search (acquire) the duplicate data from one of the first and second data sources, but to achieve such a search, it is necessary to identify the time series data (corresponding time ranges) managed in each of the first and second data sources and then specify a query for each of the first and second data sources, which makes it difficult to search (access) time series data equivalently without being aware of the data source to be searched.

[0025] Therefore, the data service system according to this embodiment has a configuration for performing efficient search processing, taking into consideration the fact that the same data exists overlappingly between the first and second data sources.

[0026] An overview of the data service system according to this embodiment will be described with reference to Fig. 3. As shown in Fig. 3, the data service system 10 includes a data management service server 11, a time-series DB server 12, an object storage server 13, an application program 14, and an integrated search engine server 15.

[0027] Here, the data service system 10 is communicatively connected to an edge server 20, and the edge server 20 collects time series data (sensor data group) measured by multiple sensor nodes 30 (sensor group) arranged based on a technology known as IoT (Internet of Things).

[0028] The data management service server 11 distributes the time-series data collected by the edge server 20 to the time-series DB server 12 .

[0029] The time-series DB server 12 stores and manages the time-series data distributed by the data management service server 11 in a database (hereinafter referred to as time-series DB) owned by the time-series DB server 12. The time-series DB corresponds to the first data source described above. The time-series DB server 12 is configured to mainly manage new data with a short retention period and to make access to individual data lightweight.

[0030] The data management service server 11 also has an aggregation processing worker 11a, and distributes the time-series data collected by the edge server 20 to the aggregation processing worker 11a. The aggregation processing worker 11a executes a process of aggregating the time-series data distributed from the data management service server 11 in units of a predetermined period. Specifically, the aggregation processing worker 11a executes a process of aggregating the time-series data collected from a plurality of sensor nodes 30 over, for example, one day in a data format that prioritizes compression efficiency.

[0031] The object storage server 13 stores and manages the time-series data (data in a file format that prioritizes compression efficiency) aggregated by the aggregation processing worker 11a in an object storage included in the object storage server 13. The object storage corresponds to the second data source described above. Since the object storage is configured to store (register) the time-series data aggregated by the aggregation processing worker 11a, the object storage has a large delay until the time-series data is registered compared to the time-series DB described above. Furthermore, since the object storage stores data in a file format obtained by aggregating time-series data corresponding to a predetermined period, access to the time-series data managed in the object storage is not lighter than access to the time-series DB server 12.

[0032] In other words, the data service system 10 can be said to be a system having a plurality of storage systems that manage time-series data according to the freshness of the data.

[0033] As described above, the formats of time-series data managed in the time-series DB and the object storage are different, but in the following description, for convenience, it is assumed that the time-series data is managed in the time-series DB and the object storage. When distinguishing between the time-series data managed in the time-series DB and the object storage, the time-series data managed in the time-series DB server 12 will be referred to as first time-series data, and the time-series data managed in the object storage server 13 will be referred to as second time-series data.

[0034] The application program 14 runs on a terminal device (hereinafter referred to as a user terminal) used by a user who uses a service realized by the data service system 10. The user can specify a query for searching time-series data managed in the time-series DB and object storage by operating the user terminal (the application program 14 running on the user terminal). The application program 14 has a function for implementing processes such as visualization and ad-hoc analysis of time-series data acquired based on a query specified by the user.

[0035] The integrated search engine server 15 realizes a function of executing an integrated search process on the above-mentioned time-series DB and object storage (that is, a plurality of data sources corresponding to a plurality of storage systems).

[0036] The integrated search engine server 15 holds a virtual table for executing integrated search processing on the time-series DB and the object storage. The virtual table is a virtual table defined to refer to data whose entity exists in an external database.

[0037] The virtual tables held in the integrated search engine server 15 include a virtual table for referencing first time-series data managed in the time-series DB (hereinafter referred to as the first virtual table), a virtual table for referencing second time-series data managed in the object storage (hereinafter referred to as the second virtual table), and a virtual table combining (integrating) the first and second virtual tables (hereinafter referred to as the joined virtual table). The joined virtual table (a virtual table corresponding to the time-series DB and the object storage) virtually integrates the time-series data managed in the time-series DB and the object storage, and realizes a function of allowing a user (application program 14) to access the time-series data as if it were managed in the integrated search engine server 15. That is, the joined virtual table makes it possible to search (acquire) time-series data across the first and second virtual tables.

[0038] Fig. 4 shows an example of the hardware configuration of integrated search engine server 15 shown in Fig. 3. As shown in Fig. 4, integrated search engine server 15 includes CPU 101, nonvolatile memory 102, main memory 103, communication device 104, etc.

[0039] The CPU 101 is a processor that controls the operation of each component in the integrated search engine server 15. The CPU 101 executes various programs that are loaded from the nonvolatile memory 102, which is a storage device, to the main memory 103. These programs include an operating system (OS) and programs that allow the integrated search engine server 15 to operate in the data service system 10.

[0040] The communication device 104 is a device configured to perform wired or wireless communication with the time-series DB server 12, the object storage server 13, and a user terminal on which the application program 14 runs.

[0041] Although only a CPU 101, a non-volatile memory 102, a main memory 103, and a communication device 104 are shown in Figure 4, the integrated search engine server 15 may further include other storage devices such as an HDD (Hard Disk Drive) and an SSD (Solid State Drive), or may further include other devices.

[0042] Although the hardware configuration of the integrated search engine server 15 has been described here, it is assumed that other devices constituting the data service system 10 (for example, the time-series DB server 12, etc.) also have similar hardware configurations.

[0043] 5 is a block diagram showing an example of the functional configuration of a data service system 10' according to a comparative example of this embodiment. Here, the time-series DB server 12', the object storage server 13', the application program 14', and the integrated search engine server 15' corresponding to the time-series DB server 12, the object storage server 13, the application program 14, and the integrated search engine server 15 shown in FIG. 3 will be mainly described.

[0044] The time-series DB server 12' includes a time-series DB 12a', a data management unit 12b', and a query execution unit 12c'. The time-series DB 12a' stores the above-described first time-series data. The data management unit 12b' has a function of managing the first time-series data stored in the time-series DB 12a'. The query execution unit 12c' executes a query sent from the integrated search engine server 15, and acquires time-series data based on the query from the time-series DB 12a' via the data management unit 12b'.

[0045] The object storage server 13′ includes an object storage 13a′. The object storage 13a′ stores the second time-series data. Note that the object storage server 13′ does not have functional units equivalent to the data management unit 12b′ and the query execution unit 12c′ of the time-series DB server 12′.

[0046] The application program 14 (or the user terminal on which it runs) includes a query request unit 14a', a communication unit 14b', and a result display unit 14c'.

[0047] The query request unit 14a' creates a query (a query for searching time series data from the joined virtual table) specified by a user using a user terminal (an application program 14' running on the user terminal) and requests a search for time series data based on the query.

[0048] The communication unit 14b' transmits a query to the integrated search engine server 15' in response to a request from the query request unit 14a. The communication unit 14b' also receives a response to the query (i.e., a search result based on the query) from the integrated search engine server 15'.

[0049] The search results received by the communication unit 14b' are displayed by the result display unit 14c' on a display or the like provided in the user terminal.

[0050] The integrated search engine server 15' includes a communication unit 15a', a query analysis unit 15b', a query request unit 15c', a first connector unit 15d', a second connector unit 15e', and a result aggregation unit 15f'.

[0051] The communication unit 15a' receives a query sent by the application program 14' (communication unit 14b').

[0052] The query analysis unit 15b' analyzes the query received by the communication unit 15a' and checks the validity of the query.

[0053] The query request unit 15c' requests the time-series DB server 12 and the object storage server 13, via the first connector unit 15d' and the second connector unit 15e', to search for time-series data based on the query whose validity has been confirmed by the query analysis unit 15b'.

[0054] The first connector unit 15d' includes a query processor and a result acquisition unit. Here, the query specified by the user is, for example, a command statement (SQL statement) written in a language called SQL (Structured Query Language), and the SQL is used in a relational database (RDB). On the other hand, the time-series DB 12a' is, for example, a database other than a relational database called NoSQL (i.e., a database that does not use SQL). In this case, the query processor of the first connector unit 15d' acquires the query from the query request unit 15c' and converts the query into a query that can be processed by the time-series DB server 12' by referring to the first virtual table (a virtual table for referencing first time-series data managed in the time-series DB 12a'). The query converted to be processable by the time-series DB server 12 is transmitted from the first connector unit 15d' to the time-series DB server 12'.

[0055] The result acquisition unit of the first connector unit 15d' acquires first time series data based on the query sent from the first connector unit 15d' to the time-series DB server 12' from the time-series DB 12a' (time-series DB server 12'). The result acquisition unit of the first connector unit 15d' converts the first time series data acquired from the time-series DB 12a' into data in a format that can be handled by the integrated search engine server 15', for example, and outputs the converted first time series data to the result aggregation unit 15f'.

[0056] The second connector unit 15e' includes a query processing unit and a result acquisition unit. The query processing unit of the second connector unit 15e' acquires a query from the query request unit 15c' described above, and downloads file-formatted data from the object storage 13a' (object storage server 13') by referring to a second virtual table (a virtual table for referencing second time-series data managed in the object storage 13a'). The query processing unit of the second connector unit 15e' extracts second time-series data to be searched based on the query from the file-formatted data downloaded from the object storage 13a'.

[0057] The result acquisition unit of the second connector unit 15e' acquires the second time series data extracted by the query processing unit of the second connector unit 15e', and outputs the second time series data to the result aggregation unit 15f'.

[0058] The result aggregation unit 15f' aggregates the first time series data output from the result acquisition unit of the first connector unit 15d' and the second time series data output from the result acquisition unit of the second connector unit 15e', and transmits the aggregation result to the application program 14 (or the user terminal on which the application program 14 is running) as a response to the query (search result based on the query).

[0059] FIG. 6 is a block diagram showing an example of the functional configuration of the data service system 10 according to this embodiment.

[0060] The time-series DB server 12 corresponds to the time-series DB server 12′ shown in Fig. 5. Note that the time-series DB 12a, the data management unit 12b, and the query execution unit 12c included in the time-series DB server 12 are similar to the time-series DB 12a′, the data management unit 12b′, and the query execution unit 12c included in the time-series DB server 12′, and therefore detailed description thereof will be omitted here.

[0061] The object storage server 13 corresponds to the object storage server 13' shown in Fig. 5. Note that the object storage 13a included in the object storage server 13 is similar to the object storage 13a' included in the object storage server 13', and therefore a detailed description thereof will be omitted here.

[0062] Application program 14 corresponds to application program 14' shown in Fig. 5. Note that query request unit 14a, communication unit 14b, and result display unit 14c included in application program 14 are similar to query request unit 14a', communication unit 14b', and result display unit 14c' included in application program 14', and therefore detailed description thereof will be omitted here.

[0063] Integrated search engine server 15 corresponds to integrated search engine server 15' shown in Fig. 5. Note that communication unit 15a, query analysis unit 15b, query request unit 15c, first connector unit 15d, second connector unit 15e, and result aggregation unit 15f included in integrated search engine server 15 are similar to communication unit 15a', query analysis unit 15b', query request unit 15c', first connector unit 15d', second connector unit 15e', and result aggregation unit 15f' included in integrated search engine server 15', and therefore detailed description thereof will be omitted here.

[0064] The integrated search engine server 15 includes a storage unit 15g and a query division unit 15h in addition to the above-mentioned communication unit 15a, query analysis unit 15b, query request unit 15c, first connector unit 15d, second connector unit 15e, and result aggregation unit 15f. In other words, the data service system 10 according to this embodiment differs from the data service system 10' according to the comparative example of this embodiment in that it includes the storage unit 15g and the query division unit 15h.

[0065] Here, assuming that the time range corresponding to the first time series data managed in the time series DB 12a and the time range corresponding to the second time series data managed in the object storage 13a overlap at least partially as described above, the storage unit 15g stores boundary management data indicating the temporal boundary conditions for the first and second time series data that are determined based on the overlapping relationship of the time ranges.

[0066] The query dividing unit 15h acquires a query (a query specified by a user) whose validity has been confirmed by the query analyzing unit 15b. Assuming that the query acquired by the query dividing unit 15h includes the above-described time condition, the query dividing unit 15h divides the query into a query for the time-series DB (i.e., the first virtual table) and a query for the object storage (i.e., the second virtual table) based on the boundary condition indicated by the boundary management data stored in the storage unit 15g and the time condition included in the query. Note that, as will be described in detail later, the time condition included in the query for the time-series DB and the time condition included in the query for the object storage do not overlap.

[0067] When the query specified by the user is divided into a query for the time-series DB and a query for the object storage by the query division unit 15h in this manner, the result aggregation unit 15f aggregates the first time-series data acquired from the time-series DB 12a based on the query for the time-series DB and the second time-series data acquired from the object storage 13a based on the query for the object storage, and transmits the aggregated result to the application program 14 as a search result.

[0068] In addition, the data service system 10' (i.e., the configuration shown in Figure 5) related to the comparative example of this embodiment described above is configured to simply acquire time series data from the time series DB 12a' and the object storage 13a' based on a query specified by a user, so there is a possibility that the same data will be acquired redundantly from both the time series DB 12a' and the object storage 13a'.

[0069] In contrast, in the data service system 10 according to this embodiment (i.e., the configuration shown in FIG. 6), a query specified by a user is divided into a query for the time-series DB and a query for the object storage, each of which includes a time condition indicating a non-overlapping time range based on the boundary management data. This makes it possible to avoid duplicate acquisition of the same data from both the time-series DB 12a and the object storage 13a, as occurs in the data service system 10′ according to the comparative example of this embodiment.

[0070] 6 has been described as including one query request unit 15c in the integrated search engine server 15, but the integrated search engine server 15 may be configured to include, for example, a query request unit (i.e., a query request unit for the time-series DB) that requests the time-series DB server 12 to search for first time-series data based on a query for the time-series DB, and a query request unit (i.e., a query request unit for the object storage) that requests the object storage server 13 to search for second time-series data based on a query for the object storage. That is, the integrated search engine server 15 in this embodiment may be configured so that the first connector unit 15d and the second connector unit 15e can simultaneously execute processes.

[0071] Next, an example of a processing procedure of the integrated search engine server 15 included in the data service system 10 according to this embodiment will be described with reference to the flowchart of FIG.

[0072] First, when a query (hereinafter referred to as a target query) is specified by a user on a user terminal on which the application program 14 runs, the communication unit 15a receives the target query from the application program 14 (step S1). The target query received in step S1 is passed from the communication unit 15a to the query analysis unit 15b.

[0073] Next, the query analysis unit 15b analyzes the target query (step S2). When the process of step S2 is executed, the query analysis unit 15b creates, for example, a parse tree representing the syntax of the target query, and determines (confirms) whether the target query is written in correct grammar based on the parse tree.

[0074] If it is determined that the target query is written in correct grammar, the query dividing unit 15h executes a process of dividing the target query into a query for the time-series DB and a query for the object storage (hereinafter referred to as a query dividing process) (step S3). The details of the query dividing process will be described later.

[0075] When the process of step S3 is executed, the first connector unit 15d acquires a query for the time-series DB via the query request unit 15c, and the second connector unit 15e acquires a query for the object storage via the query request unit 15c. In this case, the first connector unit 15d and the second connector unit 15e execute processes (query processes) on the acquired query for the time-series DB and query for the object storage (step S4).

[0076] In the query processing executed by (the query processing unit of) the first connector unit 15d, for example, a process is executed in which a query for a time-series DB is analyzed, the query is converted into a query that can be processed by the time-series DB server 12, and the converted query is output (transmitted) to the time-series DB server 12. On the other hand, in the query processing executed by (the query processing unit of) the second connector unit 15e, for example, a process is executed in which a query for an object storage is analyzed, and second time-series data (data in a file format) based on the query is downloaded.

[0077] When the processing of step S4 is executed, the first connector unit 15d (the result acquisition unit) acquires first time series data (first time series data based on a query for the time series DB) from the time series DB 12a, and the second connector unit 15e (the result acquisition unit) acquires second time series data (second time series data based on a query for the storage object) from the object storage 13a (step S5).

[0078] The processes in steps S4 and S5 described above are executed for each of the queries for the time-series DB and the object storage, but may be executed repeatedly depending on the amount and range of data to be acquired (searched) based on the queries.

[0079] Next, the result aggregation unit 15f executes a process of aggregating the first and second time series data acquired in step S5 (hereinafter referred to as a result aggregation process) (step S6). Note that the first time series data aggregated in the result aggregation process (i.e., the first time series data acquired from the time series DB 12a) may be data converted into a format that can be handled by the integrated search engine server 15, as described above. Furthermore, the result aggregation process may be executed so as to sequentially aggregate the first time series data acquired from the time series DB 12a and the second time series data acquired from the object storage 13a.

[0080] When the process of step S6 is executed, the results of the result aggregation process are returned to the application program 14 as a response (search result) to the target query.

[0081] Next, an example of the processing procedure of the above-mentioned query division processing (the processing of step S3 shown in FIG. 7) will be described with reference to the flowchart of FIG.

[0082] Here, Fig. 9 shows an example of a target query. In the example shown in Fig. 9, the target query represents a search for time series data corresponding to a time range from 00:00 on January 1, 2021, from the joined virtual table (vTable). Note that the target query shown in Fig. 9 includes "vTable" as the table name of the joined virtual table, and includes "WHERE time>'2021 / 1 / 1 00:00'" (that is, a WHERE condition indicating a time range) as the time condition.

[0083] In this case, the query dividing unit 15h acquires the time condition included in the target query from the target query (Step S11).

[0084] Next, the query dividing unit 15h acquires the boundary management data stored in the storage unit 15g from the storage unit 15g.

[0085] Here, Fig. 10 shows an example of the data structure of boundary management data. In the example shown in Fig. 10, the boundary management data includes a boundary condition list and a table list. The boundary condition list indicates temporal boundary conditions determined based on overlapping relationships of time ranges (hereinafter referred to as time ranges of virtual tables) corresponding to time-series data referenced by virtual tables included in the table list (i.e., time-series data managed in multiple data sources). The table list includes the table names of each virtual table, and in Fig. 10, the first virtual table (the table name) is shown as Table1 and the second virtual table (the table name) is shown as Table2. The same applies to the drawings used in the following description.

[0086] Specifically, the boundary management data shown in Figure 10 indicates the boundary condition determined based on the overlapping relationship between the time range of the first virtual table (Table 1) and the time range of the second virtual table (Table 2) in a format such as ">='2021 / 1 / 2 00:00'". Also, the table list is shown in a format such as "Table 1, Table 2".

[0087] In this case, the boundary condition list ">='2021 / 1 / 2 00:00'" and the table list "Table1, Table2" indicate that if time series data is obtained from time series DB 12a and object storage 13a with 00:00 on January 2, 2021 as the boundary, it is possible to avoid duplicate acquisition of the same data from the time series DB 12a and object storage 13a.

[0088] Returning to FIG. 8 again, the query dividing unit 15h creates a time range for each virtual table based on the boundary management data (boundary conditions indicated by the boundary management data) described above, which can prevent the same data (time series data corresponding to the same time range) from being acquired redundantly from the time series DB 12a and the object storage 13a (step S12).

[0089] In this case, the query division unit 15h creates the time range after 00:00 on January 2, 2021 as the time range of the first virtual table, and creates the time range before 00:00 on January 2, 2021 as the time range of the second virtual table, as shown in Figure 11.

[0090] Next, the query dividing unit 15h creates a time condition for each virtual table (i.e., an entry having a combination of each virtual table and a time condition indicating a time range corresponding to the time series data to be searched from the virtual table) based on the time range indicated by the time condition acquired in step S11 and the time range for each virtual table created in step S12 (step S13).

[0091] 12 shows an example of the time condition for each virtual table created in step S13. The example shown in FIG. 12 shows that a time condition indicating a time range after 00:00 on January 2, 2021 has been created as the time condition for the first virtual table (Table 1). Also, it shows that a time condition indicating a time range from after 00:00 on January 1, 2021 to before 00:00 on January 2, 2021 has been created as the time condition for the second virtual table (Table 2).

[0092] That is, in step S13, a time condition indicating a time range based on the logical product of the time range indicated by the time condition acquired in step S11 and the time range created in step S12 is created for each virtual table. Note that in step S13, time conditions indicating non-overlapping time ranges are created for each virtual table.

[0093] Next, the query dividing unit 15h creates a query for the time series DB (first connector unit 15d) and a query for the object storage (second connector unit 15e) (i.e., a query for each data source) based on the target query and the time conditions for each virtual table created in step S13 (step S14).

[0094] As described above, the first time-series data managed in the time-series DB 12a is referenced using the first virtual table, and the query dividing unit 15h creates a query for the time-series DB (i.e., a query including a time condition for the time-series DB) by replacing the time condition (WHERE condition) included in the target query with the time condition of the first virtual table. Similarly, the second time-series data managed in the object storage 13a is referenced using the second virtual table, and the query dividing unit 15h creates a query for the object storage (i.e., a query including a time condition for the object storage) by replacing the time condition (WHERE condition) included in the target query with the time condition of the second virtual table.

[0095] Note that, although the explanation here is that the time conditions included in the target query are replaced with the time conditions of each virtual table, the table names of the joined virtual tables included in the target query are also replaced with the table names of each virtual table (Table1 and Table2).

[0096] In the present embodiment, the query dividing unit 15h can create a query for the time-series DB and a query for the object storage from the target query (that is, divide the target query into a query for the time-series DB and a query for the object storage) by executing the process shown in Fig. 8 described above. Fig. 13 shows the query for the time-series DB divided from the target query shown in Fig. 9, and Fig. 14 shows the query for the object storage divided from the target query shown in Fig. 9.

[0097] 13, first time series data (data with time-related information) as shown in FIG. 15 is acquired from the time series DB 12a. Furthermore, according to the query for the object storage shown in FIG. 14, second time series data (data with time-related information) as shown in FIG. 16 is acquired from the object storage 13a. Note that it is assumed here that the time series DB 12a stores time series data corresponding to a time range similar to that of the first data source shown in the upper part of FIG. 2 (time series data collected every minute in that time range), and the object storage 13a stores time series data corresponding to a time range similar to that of the second data source shown in the upper part of FIG. 2 (data in a file format obtained by aggregating the time series data collected every minute in that time range by a predetermined period).

[0098] That is, in this embodiment, for example, as shown in the upper part of Figure 2 above, when data (duplicate data) corresponding to the same time range is managed in time series DB 12a (first data source) and object storage 13a (second data source), the duplicate data (i.e., time series data corresponding to the time range from 00:00 on January 1, 2021 to before 00:00 on January 2, 2021) is obtained only from object storage 13a.

[0099] The first time series data acquired from the time series DB 12a and the second time series data acquired from the object storage 13a are aggregated, for example, as shown in FIG.

[0100] As described above, the data service system 10 according to this embodiment includes, for example, a time series DB 12a (first data source) that manages first time series data corresponding to a first time range, and an object storage 13a (second data source) that manages second time series data corresponding to a second time range that at least partially overlaps with the first time range, and executes search processing using a join virtual table corresponding to the time series DB 12a and the object storage 13a. Furthermore, the data service system 10 acquires a target query including a time condition (first time condition) indicating a time range (third time range) corresponding to data searched from the time-series DB 12a and the object storage 13a, and divides the target query into a query (second query) including a time condition for the time-series DB (second time condition) and a query (third query) including a time condition for the object storage (third time condition) based on boundary management data indicating temporal boundary conditions for the time-series data determined based on the overlapping relationship of the time ranges of the time-series data managed in the time-series DB 12a and the object storage 13a and the time condition included in the target query, and then aggregates the first time-series data acquired from the time-series DB 12a based on the query for the time-series DB and the second time-series data acquired from the object storage 13a based on the query for the object storage.

[0101] That is, in this embodiment, by utilizing boundary management data in integrated search processing using virtual tables corresponding to multiple data sources (time series DB 12a and object storage 13a) that manage time series data, it is possible to avoid retrieving the same data from multiple data sources in duplicate and realize efficient search processing.

[0102] Furthermore, in this embodiment, there is no need to specify a query while taking into consideration the time-series data managed in each data source (i.e., there is no need to implement an application program 14 that specifies a complex query), so it is possible to achieve both implementation efficiency of the application program 14 and processing efficiency of the data service system 10.

[0103] 18, the time required to acquire first time series data from the time series DB 12a (first time series data acquisition time) is defined as d1, the time required to acquire second time series data from the object storage 13a (second time series data acquisition time) is defined as d2, and the ratio of the time required to acquire duplicate data (identical data that is duplicated between the first time series data and the second time series data) from the object storage 13a (duplicate data acquisition time) to the time d2 is defined as α. Note that FIG. 19 shows the relationship between a time range (first time range) 1 corresponding to the first time series data managed in the time series DB 12a, a time range (second time range) 2a corresponding to the second time series data managed in the object storage 13a, the first time series data acquisition time d1, the second time series data acquisition time d2, and the ratio α.

[0104] In this case, if duplicate data is acquired from both the time series DB 12a and the object storage 13a, the total time required to acquire the time series data (first and second time series data) from both the time series DB 12a and the object storage 13a is d1+d2, as shown in the upper part of Figure 20.

[0105] On the other hand, if this embodiment is applied in this case, the time required to acquire duplicate data from object storage 13a (i.e., second time series data acquisition time d2*α) can be reduced, so the total time required to acquire time series data (first and second time series data) from each of time series DB 12a and object storage 13a becomes d1+d2*(1-α), as shown in the lower part of Figure 20.

[0106] That is, in the present embodiment, since the time for acquiring data from the time-series DB 12a and the object storage 13a can be shortened, the efficiency of the search process can be improved.

[0107] Here, it is assumed that the process of acquiring the first time-series data from the time-series DB 12a and the process of acquiring the second time-series data from the object storage 13a are executed sequentially (one after another). In this case, the search process can be similarly optimized by configuring to acquire duplicate data from the object storage 13a (that is, not to acquire duplicate data from the time-series DB 12a). On the other hand, when the process of acquiring the first time-series data from the time-series DB 12a and the process of acquiring the second time-series data from the object storage 13a are executed simultaneously in parallel, as shown in FIG. 20, it is preferable to configure not to acquire duplicate data (the second time-series data) from the object storage 13a with a longer time-series data acquisition time (that is, processing time). That is, FIGS. 18 to 20 show examples of optimizing the search process when the process of acquiring time-series data from the time-series DB 12a and the object storage 13a is executed simultaneously in parallel and the longer time-series data acquisition time (processing time) becomes the overall processing time. Note that in FIG. 20, it is assumed that the first time-series data acquisition time d1 is sufficiently smaller than the second time-series data acquisition time d2 (that is, d1≦<d2).

[0108] Also, in the present embodiment, since the amount of data acquired from the time-series DB 12a and the object storage 13a can be reduced, the processing amount in the integrated search engine server 15 can be reduced.

[0109] Furthermore, in this embodiment, a user can specify a query without understanding the time series data (corresponding time range) managed in each of the multiple data sources (time series DB 12a and object storage 13a), and can equivalently search for time series data without being aware of the data source to be searched (i.e., equivalently access multiple data sources).

[0110] Incidentally, sharding (horizontal partitioning), which distributes and manages data in different areas of the same database management system, is known as a technology related to the data service system 10 according to this embodiment. While this sharding does not acquire duplicate data, it requires that a method for managing (storing) data be determined in advance to prevent data duplication, and is clearly distinguishable from the data service system 10 according to this embodiment. Furthermore, while sharding always references data using a single view, compared to lambda architecture, which is one of the mechanisms for building the data service system 10 according to this embodiment, this embodiment (lambda architecture) has the advantage of being able to provide optimal views by dividing the views according to the purpose of the application program 14.

[0111] 21, it is possible to obtain data without duplication by performing a union operation on data (tables) managed in the first and second data sources (the time-series DB 12a and the object storage 13a), for example, but if there is no information on the overlapping range, it is necessary to obtain duplicate data from each of the first and second data sources (that is, the transfer cost of the duplicate data becomes high), which deteriorates search latency. In this embodiment, compared to the case where the processing shown in FIG. 21 is executed, there is an advantage in that it is possible to avoid duplicate acquisition of the same data from the first and second data sources (the time-series DB 12a and the object storage 13a).

[0112] In this embodiment, a time range (fourth time range) for acquiring the first time series data and a time range (fifth time range) for acquiring the second time series data are created based on the boundary conditions indicated by the boundary management data, and the query for the time series DB and the query for the object storage are created to include time conditions indicating time ranges (sixth and seventh time ranges) based on the logical product of the time range indicated by the time condition (WHERE condition of the time column) included in the target query (query for the joined virtual integrated table) and the created time range. In other words, the query for the time series DB and the query for the object storage are created by combining, with an AND condition, the time range indicated by the time condition included in the target query and the time range for each virtual table created from the boundary conditions indicated by the boundary management data.

[0113] According to this configuration, it is possible to appropriately acquire (search) time-series data that satisfies the time condition included in the target query while avoiding duplicate acquisition of the same data from both the time-series DB 12a and the object storage 13a.

[0114] In the present embodiment, the first time series data is acquired from the time-series DB 12a based on a query for a time-series DB split (created) from the target query, and the second time series data is acquired from the object storage 13a based on a query for an object storage split (created) from the target query. However, if there is no time range based on the logical product of the time range indicated by the time condition included in the target query and the time range of the first virtual table, the first time series data is not acquired from the time-series DB 12a. Similarly, if there is no time range based on the logical product of the time range indicated by the time condition included in the target query and the time range of the second virtual table, the second time series data is not acquired from the object storage 13a. Furthermore, in FIG. 10, the boundary management data is described as including a boundary condition list and a table list. However, a query-based search for time series data is not requested for a data source corresponding to a virtual table not included (not set) in the table list (i.e., time series data is not acquired from the data source).

[0115] Specific examples of the operation of the data service system 10 according to this embodiment will be described below. First to third specific examples will be described here.

[0116] First, a first specific example will be described. Here, as shown in FIG. 22, it is assumed that, at time 1 (13:00 on January 31, 2021), first time series data (data for the most recent month) corresponding to the time range from 00:00 on December 31, 2020 to 12:59 on January 31, 2021 is managed in the time series DB 12a. Also, it is assumed that, at time 1, second time series data (data for approximately one year) corresponding to the time range from 00:00 on January 30, 2020 to 12:30 on January 31, 2021 is managed in the object storage 13a. Also, it is assumed that the storage unit 15g included in the integrated search engine server 15 stores boundary management data including the boundary condition list ">='2021 / 1 / 31 12:30'" and the table list "Table1, Table2".

[0117] Here, at the above-mentioned time 1, it is assumed that the user specified "SELECT*FROM vTable WHERE time>='2021 / 1 / 1 00:00'" as a target query (data search expression) for a joined virtual table (vTable) that combines a first virtual table for referencing first time series data and a second virtual table for referencing second time series data. Note that this target query means searching for time series data (first and second time series data) that corresponds to a time range after 00:00 on January 1, 2021 (that is, time series data that meets the time condition indicating after 00:00 on January 1, 2021) among the time series data (first and second time series data) referenced using the joined virtual table (the first and second virtual tables joined).

[0118] In this case, as shown in the upper part of Fig. 23, a time range for each virtual table is created from the boundary condition indicated by the boundary management data shown in Fig. 22. Next, a time condition for each virtual table is created based on the logical product (i.e., overlapping range) of the time range indicated by the time condition included in the target query and the time range for each virtual table, thereby dividing the target query into a query for the time-series DB server and a query for the object storage.

[0119] Specifically, as shown in the lower part of Figure 23, the query for the time series DB includes a time condition indicating the time range "time>='2021 / 1 / 31 12:30'" based on the logical product of the time range "time>='2021 / 1 / 1 00:00'" indicated by the time condition (WHERE condition) included in the target query and the time range "time>='2021 / 1 / 31 12:30'" of the first virtual table shown in the upper part of Figure 23. On the other hand, as shown in the lower part of Figure 23, the query for object storage includes a time condition indicating the time range "time>='2021 / 01 / 01 00:00' AND time<'2021 / 01 / 31 12:30'" based on the logical product of the time range "time>='2021 / 01 / 01 00:00'" indicated by the time condition (WHERE condition) included in the target query and the time range "time<'2021 / 01 / 31 12:30'" of the second virtual table shown in the upper part of Figure 23.

[0120] In this case, first time series data based on a query for the time series DB shown in the lower part of Figure 23 is acquired from the time series DB 12a, and second time series data based on a query for the object storage shown in the lower part of Figure 23 is acquired from the object storage 13a.

[0121] As described above, if the time series data is collected at predetermined intervals, the time series data is periodically added to the time series DB 12a and the object storage 13a, and the time range corresponding to the first time series data managed in the time series DB 12a and the time range corresponding to the second time series data managed in the object storage 13a change according to the addition of the time series data.

[0122] It is assumed that time series data collected from 2 minutes ago to 1 minute ago is sequentially added to the time series DB 12a every minute, for example, as first time series data. Also, it is assumed that file-format data aggregating time series data collected from 45 minutes ago to 30 minutes ago is sequentially added to the object storage 13a every 15 minutes, for example, as second time series data.

[0123] Here, Figure 24 shows the time range corresponding to the first time series data, the time range corresponding to the second time series data, and the boundary management data as of time 2 (13:16 on January 31, 2021), which is later than time 1 mentioned above.

[0124] The time range corresponding to the first and second time series data at time 2 is different from the time range corresponding to the first and second time series data at time 1, but in this case, the boundary management data indicating the temporal boundary conditions for the first and second time series data is updated (changed) in accordance with the change in the time range.

[0125] Specifically, for example, at 13:15 on January 31, 2021, the aggregation processing worker 11a included in the data management service server 11 adds to the object storage 13a the result (one file) of aggregating the time series data collected from 12:30 to 12:45 on January 31, 2021. In this case, the aggregation processing worker 11a updates the boundary management data shown in FIG. 22 (boundary condition list ">='2021 / 1 / 31 12:30'") to the boundary management data shown in FIG. 24 (boundary condition list ">='2021 / 1 / 31 12:45'").

[0126] 22 and 24 show that data summarized in 15-minute increments is written to the object storage 13a, and the boundary condition list (boundary management data) is updated to include values ​​15 minutes later. In this embodiment, it is sufficient that the aggregation processing worker 11a updates the boundary management data, and there is no need to change the application program 14 even if the time ranges corresponding to the first and second time-series data are changed.

[0127] Here, it has been explained that the updating (processing) of the boundary management data in response to the addition of the time series data described above is performed by the aggregation processing worker 11a (data management service server 11), but the updating of the boundary management data may also be performed by the integrated search engine server 15 in response to, for example, a notification from the aggregation processing worker 11a.

[0128] Here, it is assumed that at time 2, the above-mentioned target query "SELECT*FROM vTable WHERE time >='2021 / 01 / 01 00:00'" is specified by the user.

[0129] In this case, as shown in the upper part of Fig. 25, a time range for each virtual table is created from the boundary condition indicated by the boundary management data shown in Fig. 24. Next, a time condition for each virtual table is created based on the logical product (i.e., overlapping range) of the time range indicated by the time condition included in the target query and the time range for each virtual table, whereby the target query is divided into a query for the time-series DB and a query for the object storage shown in the lower part of Fig. 25.

[0130] Note that the query for the time series DB shown in the lower part of Figure 25 includes a time condition indicating the time range "time>='2021 / 1 / 31 12:45'" based on the logical product of the time range "time>='2021 / 1 / 1 00:00'" indicated by the time condition (WHERE condition) included in the target query and the time range "time>='2021 / 1 / 31 12:45'" of the first virtual table shown in the upper part of Figure 25. On the other hand, the query for object storage shown in the lower part of Figure 25 includes a time condition indicating the time range "time>='2021 / 1 / 1 00:00' AND time<'2021 / 01 / 31 12:45'" based on the logical product of the time range "time>='2021 / 1 / 1 00:00'" indicated by the time condition (WHERE condition) included in the target query and the time range "time<'2021 / 1 / 31 12:45'" of the second virtual table shown in the upper part of Figure 25.

[0131] In this case, first time series data based on a query for the time series DB shown in the lower part of Figure 25 is acquired from the time series DB 12a, and second time series data based on a query for the object storage shown in the lower part of Figure 25 is acquired from the object storage 13a.

[0132] Here, Figure 26 shows the relationship between the boundary management data (boundary condition list) at time 1 and time 2 and the data source (hereinafter referred to as the referenced data source) from which the time series data is obtained at time 1 and time 2. Note that Figure 26 is a simple diagram for explaining changes in the referenced data source according to the boundary management data, and the time ranges and the like in Figure 26 are shown in outline. The same applies to the following drawings similar to Figure 26.

[0133] 26 represents the time range corresponding to the first time series data managed in the time series DB 12a at time 1, and time range 202a represents the time range corresponding to the second time series data managed in the object storage 13a at time 1. Also, time range 201b represents the time range corresponding to the first time series data managed in the time series DB 12a at time 2, and time range 202b represents the time range corresponding to the second time series data managed in the object storage 13a at time 2.

[0134] According to Figure 26, at time 1 when the boundary condition list is ">='2021 / 01 / 31 12:30'", the data source (referenced data source) that obtains the time series data corresponding to time range 203a is object storage 13a, and the data source (referenced data source) that obtains the time series data corresponding to time range 204a is time series DB 12a.

[0135] Furthermore, at time 2, where the boundary condition list is ">='2021 / 01 / 31 12:45'", the data source (referenced data source) that obtains the time series data corresponding to time range 203b is object storage 13a, and the data source (referenced data source) that obtains the time series data corresponding to time range 204b is time series DB 12a.

[0136] According to the first specific example described above, by updating the boundary management data (boundary condition list) at the timing when new data is added to the object storage 13a, for example, every 15 minutes, even if the time range corresponding to the time series data managed in the time series DB 12a and the object storage 13a is changed, the query specified by the user on the user terminal (the data search formula executed by the application program 14) can be expressed in the same way, thereby realizing efficient search processing.

[0137] In the first specific example, it is assumed that for identical data (duplicate data in the first and second time series data) managed in the time series DB 12a and the object storage 13a, the second time series data is preferentially acquired. With this configuration, it is possible to construct a data service system 10 that reduces the access load to the time series DB 12a, for example.

[0138] Here, we have described a configuration in which duplicate data is acquired with priority given to object storage 13a (i.e., for duplicate data, the second time series data is acquired with priority), but below we will describe as a second specific example a configuration in which duplicate data is acquired with priority given to time series DB 12a (i.e., for duplicate data, the first time series data is acquired with priority).

[0139] Here, as shown in FIG. 27, it is assumed that, at time 1 (13:00 on January 31, 2021), the time-series DB 12a manages first time-series data (data for the most recent month) corresponding to the time range from 00:00 on December 31, 2020 to 12:59 on January 31, 2021. Also, it is assumed that, at time 1, the object storage 13a manages second time-series data (data for approximately one year) corresponding to the time range from 00:00 on January 30, 2020 to 12:30 on January 31, 2021. Also, it is assumed that the storage unit 15g included in the integrated search engine server 15 stores boundary management data including the boundary condition list ">='2020 / 12 / 31 00:00'" and the table list "Table1, Table2."

[0140] Note that the time ranges corresponding to the first and second time series data shown in Fig. 27 are the same as those in Fig. 22 described above, but the boundary condition lists are different between Fig. 22 (i.e., the first specific example) and Fig. 27 (i.e., the second specific example). That is, it can be said that the boundary condition list in Fig. 22 indicates the boundary conditions for preferentially acquiring duplicate data from the object storage 13a for the time ranges corresponding to the first and second time series data, whereas the boundary condition list in Fig. 27 indicates the boundary conditions for preferentially acquiring duplicate data from the time series DB 12a for the time ranges corresponding to the first and second time series data.

[0141] Here, it is assumed that at the above-mentioned time 1, the user specified "SELECT*FROM vTable WHERE time>='2020 / 12 / 31 00:00'" as the target query. Note that this target query means searching for time series data (first and second time series data) that corresponds to the time range after 00:00 on December 31, 2020 (that is, time series data that satisfies the time condition indicating after 00:00 on December 31, 2020) among the time series data (first and second time series data) referenced using the joined virtual table (first and second virtual tables joined to it).

[0142] In this case, as shown in the upper part of Fig. 28, a time range for each virtual table is created from the boundary condition indicated by the boundary management data shown in Fig. 22. Next, a time condition for each virtual table is created based on the logical product of the time range indicated by the time condition included in the target query and the time range for each virtual table, thereby dividing the target query into a query for the time-series DB and a query for the object storage.

[0143] Specifically, as shown in the lower part of Figure 28, the query for the time-series DB includes a time condition indicating the time range "time >= '2020 / 12 / 31 00:00'" based on the logical product of the time range "time >= '2020 / 12 / 31 00:00'" indicated by the time condition (WHERE condition) included in the target query and the time range "time >= '2020 / 12 / 31 00:00'" of the first virtual table shown in the upper part of Figure 28. Note that in this case, there is no time range based on the logical product of the time range indicated by the time condition included in the target query and the time range "time < '2020 / 12 / 31 00:00'" of the second virtual table shown in the upper part of Figure 28, so a query for object storage is not created.

[0144] In this case, the first time series data based on the query for the time series DB shown in the lower part of FIG. 28 is acquired from the time series DB 12a, and the second time series data is not acquired from the object storage 13a.

[0145] Incidentally, the first time series data managed in the time series DB 12a is periodically deleted by the retention described above, and the time range corresponding to the first time series data managed in the time series DB 12a changes in accordance with the deletion of the first time series data.

[0146] As described above, new data is periodically added to the time-series DB 12a, but here, as the data is added, the oldest one-day data is deleted each day, so that only the most recent month's worth of data is managed in the time-series DB 12a as the first time-series data. In the following description, it is assumed that the first time-series data managed in the time-series DB 12a is deleted, but the second time-series data managed in the object storage 13a may also be periodically deleted in the same manner.

[0147] Here, Figure 29 shows the time range corresponding to the first time series data, the time range corresponding to the second time series data, and the boundary management data as of time 3 (00:01 on February 1, 2021), which is later than time 1 mentioned above.

[0148] The time ranges corresponding to the first and second time series data at time 3 are different from the time ranges corresponding to the first and second time series data at time 1, but in this case, the boundary management data indicating the temporal boundary conditions for the first and second time series data is updated (changed) in accordance with the change in the time ranges. Note that the addition of time series data to the time series DB 12a and the object storage 13a (change in the time range in accordance with the addition) is as described in the first specific example above.

[0149] Specifically, for example, at 00:00 on February 1, 2021, the data management service server 11 (time-series DB maintenance program) deletes the oldest one day's worth of past data (data for December 31, 2021) from the first time-series data managed in the time-series DB 12a. In this case, the data management service server 11 updates the boundary management data shown in FIG. 22 (boundary condition list ">='2021 / 01 / 31 12:30'") to the boundary management data shown in FIG. 29 (boundary condition list ">='2021 / 01 / 01 00:00'").

[0150] That is, according to the examples shown in Figures 27 and 29, when a new date is reached, the time series DB 12a is accessed and a process is executed to delete data older than one month, thereby updating the boundary condition list (boundary management data) to include values ​​one day later.

[0151] Here, it is assumed that at time 3, the above target query "SELECT*FROM vTable WHERE time >='2020 / 12 / 31 00:00'" is specified by the user.

[0152] In this case, as shown in the upper part of Fig. 30, a time range for each virtual table is created from the boundary condition indicated by the boundary management data shown in Fig. 29. Next, a time condition for each virtual table is created based on the logical product of the time range indicated by the time condition included in the target query and the time range for each virtual table, whereby the target query is divided into a query for the time-series DB and a query for the object storage shown in the lower part of Fig. 30.

[0153] Note that the query for the time series DB shown in the lower part of Figure 30 includes a time condition indicating the time range "time>='2021 / 1 / 1 00:00'" based on the logical product of the time range "time>='2020 / 12 / 31 00:00'" indicated by the time condition (WHERE condition) included in the target query and the time range "time>='2021 / 1 / 1 00:00'" of the first virtual table shown in the upper part of Figure 30. On the other hand, the query for object storage shown in the lower part of Figure 30 includes a time condition indicating the time range ``time>=`2020 / 12 / 31 00:00' AND time<`2021 / 01 / 01 00:00''' based on the logical product of the time range ``time>=`2020 / 12 / 31 00:00''' indicated by the time condition (WHERE condition) included in the target query and the time range ``time<`2021 / 1 / 1 00:00''' of the second virtual table shown in the upper part of Figure 30.

[0154] In this case, first time series data based on a query for the time series DB shown in the lower part of Figure 30 is obtained from the time series DB 12a, and second time series data based on a query for the object storage shown in the lower part of Figure 30 is obtained from the object storage 13a.

[0155] Here, Figure 31 shows the relationship between the boundary management data (boundary condition list) at time 1 and time 3 and the data source (referenced data source) from which the time series data is obtained at time 1 and time 2.

[0156] 30 represents the time range corresponding to the first time series data managed in the time series DB 12a at time 1, and time range 202a represents the time range corresponding to the second time series data managed in the object storage 13a at time 1. Also, time range 201c shown in Fig. 30 represents the time range corresponding to the first time series data managed in the time series DB 12a at time 3, and time range 202c represents the time range corresponding to the second time series data managed in the object storage 13a at time 3.

[0157] According to Figure 30, at time 1 when the boundary condition list is ">='2020 / 12 / 31 00:00'", the data source (referenced data source) that obtains the time series data corresponding to time range 205a is object storage 13a, and the data source (referenced data source) that obtains the time series data corresponding to time range 206a is time series DB 12a.

[0158] Furthermore, at time 3, where the boundary condition list is ">='2021 / 01 / 01 00:00'", the data source (referenced data source) that obtains the time series data corresponding to time range 205b is object storage 13a, and the data source (referenced data source) that obtains the time series data corresponding to time range 206b is time series DB 12a.

[0159] According to the second specific example described above, by updating the boundary management data (boundary condition list) at the timing when past data is deleted from the time-series DB 12a, for example, on a daily basis, even if the time range corresponding to the time-series data managed in the time-series DB 12a is changed, the query specified by the user on the user terminal (the data search formula executed by the application program 14) can be expressed in the same way, thereby realizing efficient search processing.

[0160] In the second specific example, it is assumed that when identical data (duplicate data in the first and second time series data) is acquired in the time series DB 12a and the object storage 13a, the first time series data is acquired first. With this configuration, it is possible to build a data service system 10 that emphasizes search response performance, for example, when it is expected that the search latency to the time series DB 12a is low compared to the object storage 13a.

[0161] Furthermore, in the second specific example, a case where the data management service server 11 (time series DB maintenance program) deletes past data is described, but the second specific example may also be applied when the time range corresponding to the time series data changes due to past data being added to the time series DB 12a or object storage 13a.

[0162] In the above-described first and second specific examples, the cases where the multiple data sources are the time-series DB 12a and the object storage 13a have been described, but in the third specific example, the case where the multiple data sources are multiple object storages will be described. Note that in the third specific example, it is assumed that the multiple object storages are included in a single object storage server 13, but they may also be located in different object storage servers 13. In the following description, for convenience, the multiple object storages will be referred to as object storages A and B.

[0163] Here, as shown in FIG. 32, it is assumed that, at time 1 (00:00 on January 31, 2021), object storage A manages second time series data (data for the most recent year) corresponding to the time range from 00:00 on January 31, 2020 to 23:59 on January 30, 2021. Also, it is assumed that, at time 1, object storage B manages second time series data (data for the most recent two years) corresponding to the time range from 00:00 on January 31, 2019 to 23:59 on January 30, 2021. That is, the time range corresponding to the second time series data managed in object storage A (hereinafter referred to as time series data A) and the time range corresponding to the second time series data managed in object storage B (hereinafter referred to as time series data B) partially overlap.

[0164] Furthermore, it is assumed that storage unit 15g included in integrated search engine server 15 stores boundary management data including a boundary condition list ">='2021 / 01 / 31 00:00'" and a table list "TableA, TableB." TableA indicates a virtual table (hereinafter referred to as virtual table A) for referencing time-series data A managed in object storage A, and TableB indicates a virtual table (hereinafter referred to as virtual table B) for referencing time-series data B managed in object storage B.

[0165] In the third specific example, it is assumed that object storages A and B each store data in a file format that aggregates time-series data on a daily basis.

[0166] Furthermore, in object storages A and B, new data (file-format data) for one day is added and past data for one day is deleted every day.

[0167] Furthermore, object storage A is assumed to be, for example, a storage with low latency and a high data storage cost relative to its capacity. On the other hand, object storage B is assumed to be, for example, a storage with high latency and a low data storage cost relative to its capacity. In other words, object storages A and B differ in that they provide different storage services (different service levels).

[0168] Here, it is assumed that at the above-mentioned time 1, the user specified "SELECT*FROM vTable WHERE time>='2020 / 1 / 1 00:00'" as a target query (data search expression) for the joined virtual table (vTable) that joins virtual table A for referencing time series data A and virtual table B for referencing time series data B. Note that this target query means searching for time series data (time series data A and B) that corresponds to the time range after 00:00 on January 1, 2020 (that is, time series data that meets the time condition indicating after 00:00 on January 1, 2020) among the time series data (time series data A and B) referenced using the joined virtual table (virtual tables A and B joined to them).

[0169] In this case, as shown in the upper part of Fig. 33, a time range for each virtual table is created from the boundary conditions indicated by the boundary management data shown in Fig. 31. Next, a time condition for each virtual table is created based on the logical product of the time range indicated by the time condition included in the target query and the time range for each virtual table, thereby dividing the target query into a query for object storage A and a query for object storage B.

[0170] Specifically, as shown in the lower part of Figure 33, the query for object storage A includes a time condition indicating the time range "time>='2020 / 01 / 31 00:00'" based on the logical conjunction of the time range "time>='2020 / 01 / 1 00:00'" indicated by the time condition (WHERE condition) included in the target query and the time range "time>='2020 / 01 / 31 00:00'" of virtual table A shown in the upper part of Figure 33. On the other hand, as shown in the lower part of Figure 33, the query for object storage B includes a time condition indicating the time range "time>='2020 / 1 / 1 00:00'" based on the logical product of the time range indicated by the time condition (WHERE) included in the target query (time>='2020 / 1 / 1 00:00') and the time range "time<'2020 / 1 / 31 00:00'" of virtual table B shown in the upper part of Figure 33.

[0171] In this case, time series data A based on the query for object storage A shown in the lower part of Figure 33 is obtained from object storage A, and time series data B based on the query for object storage B shown in the lower part of Figure 33 is obtained from object storage B.

[0172] Here, assuming that new data is added to object storages A and B every day and past data is deleted as described above, the time range corresponding to the time series data managed in object storages A and B changes due to the addition and deletion of the data.

[0173] Figure 34 shows the time range corresponding to time series data A, the time range corresponding to time series data B, and boundary management data as of time 3 (00:01 on February 1, 2021), which is later than time 1 mentioned above.

[0174] The time range corresponding to time series data A and B at time 3 is different from the time range corresponding to time series data A and B at time 1, but in this case, the boundary management data indicating the temporal boundary conditions for time series data A and B is updated (changed) in accordance with the change in the time range.

[0175] Specifically, for example, at 00:00 on February 1, 2021, the aggregation processing worker 11a included in the data management service server 11 adds the result (one file) of aggregating the time series data collected up to January 31, 2021 to object storages A and B. Furthermore, because object storage A manages data for the most recent year, the data management service server 11 deletes one file corresponding to the result of aggregating the time series data collected up to January 31, 2020 from object storage A. Furthermore, because object storage B manages data for the most recent two years, the data management service server 11 deletes one file corresponding to the result of aggregating the time series data collected up to January 31, 2019 from object storage B.

[0176] In this case, the data management service server 11 (aggregation processing worker 11a) updates the boundary management data in Figure 32 (boundary condition list ">='2020 / 01 / 31 00:00'") to the boundary management data shown in Figure 34 (boundary condition list ">='2020 / 02 / 01 00:00'").

[0177] In other words, the examples shown in Figures 32 and 34 show that the boundary condition list (boundary management data) is updated to include values ​​for one day later by adding and deleting data summarized on a daily basis to object storages A and B.

[0178] Here, it is assumed that at time 3, the above target query "SELECT*FROM vTable WHERE time >='2020 / 1 / 1 00:00'" is specified by the user.

[0179] In this case, as shown in the upper part of Fig. 35, a time range for each virtual table is created from the boundary conditions indicated by the boundary management data shown in Fig. 34. Next, a time condition for each virtual table is created based on the logical product of the time range indicated by the time condition included in the target query and the time range for each virtual table, thereby dividing the target query into a query for object storage A and a query for object storage B shown in the lower part of Fig. 35.

[0180] Note that the query for object storage A shown in the lower part of Figure 35 includes a time condition indicating the time range "time>='2020 / 2 / 1 00:00'" based on the logical product of the time range "time>='2020 / 1 / 1 00:00'" indicated by the time condition (WHERE' condition) included in the target query and the time range "time>='2020 / 2 / 1 00:00'" of virtual table A shown in the upper part of Figure 35. On the other hand, the query for object storage B shown in the lower part of Figure 35 includes a time condition indicating the time range "time>='2020 / 1 / 1 00:00' AND time<'2021 / 2 / 1 00:00'" based on the logical product of the time range "time>='2020 / 1 / 1 00:00'" indicated by the time condition (WHERE' condition) included in the target query and the time range "time<'2020 / 2 / 1 00:00'" of virtual table B shown in the upper part of Figure 35.

[0181] In this case, time series data A based on the query for object storage A shown in the lower part of Figure 35 is obtained from object storage A, and time series data B based on the query for object storage B shown in the lower part of Figure 35 is obtained from object storage B.

[0182] Here, Figure 36 shows the relationship between the boundary management data (boundary condition list) at time 1 and time 3 and the data source (referenced data source) from which the time series data is obtained at time 1 and time 3.

[0183] 36 represents the time range corresponding to time series data A managed in object storage A at time 1, and time range 208a represents the time range corresponding to time series data B managed in object storage B at time 1. Also, time range 207b shown in Fig. 36 represents the time range corresponding to time series data A managed in object storage A at time 3, and time range 208b represents the time range corresponding to time series data B managed in object storage B at time 3.

[0184] According to Figure 36, at time 1 when the boundary condition list is ">='2020 / 01 / 31 00:00'", the data source (referenced data source) that obtains the time series data corresponding to time range 209a is object storage B, and the data source (referenced data source) that obtains the time series data corresponding to time range 210a is object storage A.

[0185] Furthermore, at time 3, where the boundary condition list is ">='2020 / 02 / 01 00:00'", the data source (referenced data source) that obtains the time series data corresponding to time range 209b is object storage B, and the data source (referenced data source) that obtains the time series data corresponding to time range 210b is object storage A.

[0186] According to the third specific example described above, by updating the boundary management data (boundary condition list) when new data is added to object storages A and B on a daily basis and past data is deleted from object storages A and B, even if the time range corresponding to the time series data managed in object storages A and B is changed, the query specified by the user on the user terminal (the data search formula executed by application program 14) can be expressed in the same way, thereby realizing efficient search processing.

[0187] In the third specific example, a data service system 10 that emphasizes search response performance can be constructed by adopting a configuration in which, among multiple object storages, priority is given to retrieving duplicate data from object storage that is expected to have low search latency.

[0188] In this embodiment, the first to third specific examples have been described, but the configurations described in the first to third specific examples are merely examples, and this embodiment may be configured to avoid duplicate retrieval of the same data from multiple data sources by using boundary management data to divide the target query into queries for multiple data sources.

[0189] In this embodiment, the boundary condition list (boundary conditions) included in the boundary management data is described as being expressed in a format such as ">='2021 / 1 / 2 00:00'" as shown in Figure 10, but the boundary condition list may be in other formats.

[0190] Specifically, FIG. 37 shows another example of the data structure of boundary management data (boundary condition list included in boundary management data). FIG. 37 shows a boundary condition list (including boundary management data) using a function that can dynamically change boundary conditions according to the current time. According to the example shown in FIG. 37, the boundary condition list is automatically updated based on the time (date and time) calculated by "current time - 30 minutes." Such a boundary condition list has the advantage that the aggregation processing worker 11a, etc. does not need to perform processing to periodically update the boundary management data (boundary condition list). In the example shown in FIG. 37, the boundary management data (boundary condition list) is updated, for example, every minute, but the boundary management data may also be automatically updated at predetermined intervals, for example, every 15 minutes.

[0191] Furthermore, in this embodiment, the boundary management data has been described as including a boundary condition list and a table list, but the boundary management data may have a data structure such as that shown in FIG.

[0192] Specifically, the boundary management data shown in Figure 38 includes a table, a data period, and a priority, which are associated with each other. The table indicates a virtual table (i.e., a data source). The data period indicates a time range corresponding to the time series data referenced using the associated virtual table. The priority indicates a priority (i.e., a value for determining where duplicate data is obtained) assigned to the data source that manages the time series data referenced using the associated virtual table. Note that the smaller the priority number, the higher the degree of priority.

[0193] 38, the second virtual table (Table2) has a higher priority than the first virtual table (Table1). In this case, the integrated search engine server 15 (query dividing unit 15h) calculates the overlapping time range (i.e., boundary conditions) between the data period associated with table "Table1" and the data period associated with table "Table2," and divides the target query so that the time-series data corresponding to the overlapping time range is searched for by referring to the second virtual table (i.e., obtained from the object storage 13a).

[0194] In this case, in the process of calculating the boundary conditions, the data period associated with the virtual table having a small priority value (i.e., a high priority) is preferentially reflected in the query (the time conditions included in the query) for the data source that manages the time-series data referenced using the virtual table, thereby making it possible to obtain duplicate data from the data source having the high priority.

[0195] In the above-mentioned first specific example, a configuration that prioritizes object storage 13a was described, and in the above-mentioned second specific example, a configuration that prioritizes time-series DB 12a was described. However, by using the boundary management data shown in FIG. 38, it is possible to switch between the configuration described in the first specific example and the configuration described in the second specific example (i.e., the data source that is preferentially used to obtain duplicate data) by changing the priority included in the boundary management data.

[0196] Furthermore, in updating (processing) the boundary management data shown in Figure 38, when the time range corresponding to the time series data managed in each data source is changed, the data period associated with the virtual table for referencing the time series data is updated.

[0197] (Second embodiment) Next, a second embodiment will be described. The outline of the data service system according to this embodiment is the same as that of the first embodiment, and therefore will be described with reference to FIG.

[0198] Fig. 39 is a block diagram showing an example of the functional configuration of the data service system 10 according to this embodiment. In Fig. 39, the same parts as those in Fig. 6 described above are given the same reference numerals, and detailed explanations thereof will be omitted, and the explanation will focus mainly on parts that differ from Fig. 6.

[0199] The data service system 10 according to this embodiment differs from the first embodiment in that the object storage server 13 further includes an object storage 13b in addition to the object storage 13a. Note that the object storage 13a in this embodiment stores the second time-series data, similar to the first embodiment.

[0200] Here, the object storage 13a is configured to manage time series data for a longer period of time than the time series DB 12a, and at least a part of the second time series data managed in the object storage 13a may be updated after being stored in the object storage 13a. Therefore, the object storage 13b stores the time series data (hereinafter referred to as third time series data) after (at least a part of) the second time series data is updated as described above.

[0201] In the first embodiment, the storage unit 15g included in the integrated search engine server 15 stores the boundary management data, but in the present embodiment, the storage unit 15g further stores update management data. The update management data indicates a time range corresponding to the third time-series data stored in the object storage 13b.

[0202] Next, the operation of the integrated search engine server 15 included in the data service system 10 according to this embodiment will be described. The processing procedure of the integrated search engine server 15 according to this embodiment is the same as the processing shown in Fig. 7 described in the first embodiment, but the processing of step S3 shown in Fig. 7 (query division processing) is different from that of the first embodiment.

[0203] An example of the processing procedure for the query division processing in this embodiment will be described below with reference to the flowchart in FIG.

[0204] First, the processes of steps S21 and S22, which correspond to the processes of steps S11 and S12 shown in FIG. 8, are executed.

[0205] In step S21, it is assumed that the time condition "WHERE time>'2021 / 1 / 1 00:00" is acquired from the target query shown in Fig. 9. In addition, in step S22, it is assumed that the time ranges of the first and second virtual tables shown in Fig. 11 are created from the boundary management data.

[0206] Next, the query dividing unit 15h acquires the update management data stored in the storage unit 15g from the storage unit 15g.

[0207] Here, Fig. 41 shows an example of the data structure of the update management data. In the example shown in Fig. 40, the update management data includes (information indicating) an update date and time, a time range, and a table. The update date and time indicates the date and time when (at least a part of) the second time series data stored in the object storage 13a was updated to the third time series data stored in the object storage 13b. The time range indicates the time range corresponding to the third time series data. The table indicates the table name of a virtual table for referencing the third time series data.

[0208] 41, the update management data includes the update date and time "2021 / 1 / 2 00:00." This indicates that the date and time when the second time series data was updated to the third time series data was 00:00 on January 2, 2021.

[0209] The update management data also includes the time range "time>'2021 / 1 / 1 00:00' AND time<='2021 / 1 / 1 00:01'", which indicates that the second time series data corresponding to the time range from 00:00 on January 1, 2021 to 00:01 on January 1, 2021 has been updated to the third time series data.

[0210] Furthermore, the update management data includes a table "Table 3." Note that Table 3 indicates a virtual table (hereinafter referred to as a third virtual table) for referencing third time-series data managed in the object storage 13b.

[0211] 40 again, the query dividing unit 15h creates a time range of the third virtual table from the update management data (step S23). In this case, the time range included in the update management data is created as the time range of the third virtual table.

[0212] Next, the query dividing unit 15h creates a time condition for each virtual table (i.e., an entry having a combination of each virtual table and a time condition indicating a time range corresponding to the time series data searched from the virtual table) based on the time range indicated by the time condition acquired in step S21, the time ranges of the first and second virtual tables created in step S22, and the time range of the third virtual table created in step S23 (step S24).

[0213] 42 shows an example of the time conditions for each virtual table created in step S24. In the example shown in FIG. 42, a time condition indicating a time range after 00:00 on January 2, 2021 is shown as the time condition of the first virtual table (Table 1). In addition, a time condition indicating a time range before 00:00 on January 1, 2021 or a time range from after 00:01 on January 1, 2021 to before 00:00 on January 2, 2021 is shown as the time condition of the second virtual table (Table 2). Furthermore, a time condition indicating a time range from after 00:00 on January 1, 2021 to before 00:01 on January 1, 2021 is shown as the time condition of the third virtual table (Table 3).

[0214] That is, in step S24, a time condition indicating a time range based on the logical product of the time range indicated by the time condition acquired in step S21 and the time ranges created in steps S22 and S23 is created for each virtual table. According to this, the time range corresponding to the updated time series data is excluded from (the time range indicated by) the time condition of the second virtual table, and a time condition indicating this time range is created for the third virtual table. Note that the time conditions for each virtual table created in step S24 do not overlap with each other.

[0215] Next, the query dividing unit 15h creates a query for the time-series DB, a query for the object storage A, and a query for the object storage B (i.e., queries for each data source) based on the target query and the time conditions for each virtual table created in step S24 (step S25). Note that in this embodiment, the query for the object storage A is a query for acquiring the second time-series data from the object storage 13a. Also, the query for the object storage B is a query for acquiring the third time-series data from the object storage 13b.

[0216] As described above, the first time series data managed in the time-series DB 12a is referenced using the first virtual table, and the query dividing unit 15h creates a query for the time-series DB (i.e., a query including a time condition for the time-series DB) by replacing the time condition (WHERE condition) included in the target query with the time condition of the first virtual table. Similarly, the second time series data managed in the object storage 13a is referenced using the second virtual table, and the query dividing unit 15h creates a query for the object storage A (i.e., a query including a time condition for the object storage A) by replacing the time condition (WHERE condition) included in the target query with the time condition of the second virtual table. Furthermore, the third time series data managed in the object storage 13b is referenced using the third virtual table, and the query dividing unit 15h creates a query for the object storage B (i.e., a query including a time condition for the object storage B) by replacing the time condition (WHERE condition) included in the target query with the time condition of the third virtual table.

[0217] Note that, although the explanation here is given assuming that the time conditions included in the target query are replaced with the time conditions of each virtual table, the table names of the joined virtual tables included in the target query are also assumed to be replaced with the table names of each virtual table (Table1, Table2, and Table3).

[0218] In this embodiment, the query dividing unit 15h can create a query for the time-series DB, a query for object storage A, and a query for object storage B from the target query by executing the process of FIG. 40 described above (that is, divide the target query into a query for the time-series DB, a query for object storage A, and a query for object storage B). Note that the query for the time-series DB in this case is the same as the query shown in FIG. 13 described above. Also, FIG. 43 shows a query for object storage A, and FIG. 44 shows a query for object storage B.

[0219] According to the query for the time-series DB shown in Fig. 13, for example, first time-series data as shown in Fig. 15 is acquired from the time-series DB 12a. Furthermore, according to the query for the object storage A shown in Fig. 43, for example, second time-series data as shown in Fig. 45 is acquired from the object storage 13a. Furthermore, according to the query for the object storage B shown in Fig. 44, for example, third time-series data as shown in Fig. 46 is acquired from the object storage 13b.

[0220] That is, in this embodiment, when third time series data (data obtained by updating the second time series data managed in object storage 13a) is managed in object storage 13b, the third time series data (data after update) is acquired in preference to the second time series data (i.e., data before update).

[0221] The first time series data acquired from the time series DB 12a, the second time series data acquired from the object storage 13a, and the third time series data acquired from the object storage 13b are aggregated, for example, as shown in Figure 47, and returned to the application program 14.

[0222] As described above, in this embodiment, the object storage server 13 further includes the object storage 13b (third data source) that manages third time series data obtained by updating at least a portion of the second time series data managed in the object storage 13a (second data source), and the storage unit 15g included in the integrated search engine server 15 further stores update management data indicating a time range (eighth time range) corresponding to the third time series data. In this embodiment, based on the boundary management data and update management data stored in the storage unit 15g and the time condition (first time condition) included in the target query (first query), the target query is divided into a query for the time series DB (second query), a query for object storage A (third query), and a query for object storage B (fourth query).

[0223] In this embodiment, with this configuration, for example, when the second time series data managed in the object storage 13a is updated to the third time series data, it is possible to acquire the third time series data without acquiring the second time series data, thereby enabling efficient search processing.

[0224] A specific example of the operation of the data service system 10 according to this embodiment will be described below. In the following description, the object storages 13a and 13b will be referred to as object storages A and B for convenience.

[0225] Here, as shown in FIG. 48, it is assumed that, at time 1 (13:00 on January 31, 2021), first time series data (data for the most recent month) corresponding to the time range from 00:00 on December 31, 2020 to 12:59 on January 31, 2021 is managed in time series DB 12a. Also, it is assumed that, at time 1, second time series data corresponding to the time range from 00:00 on January 30, 2020 to 12:30 on January 31, 2021 is managed in object storage A. Also, it is assumed that boundary management data including the boundary condition list ">='2021 / 01 / 31 12:30'" and the table list "Table1, Table2" is stored in storage unit 15g included in integrated search engine server 15. It is assumed that at time 1, the third time series data is not stored in the object storage B, and the update management data is not stored in the storage unit 15g.

[0226] Here, it is assumed that at the above-mentioned time 1, the user specified "SELECT*FROM vTable WHERE time>='2020 / 1 / 1 00:00'" as a target query (data search expression) for a joined virtual table (vTable) that joins a first virtual table for referencing first time series data and a second virtual table for referencing second time series data. Note that this target query means searching for time series data (first and second time series data) that corresponds to a time range after 00:00 on January 1, 2020 (that is, time series data that satisfies the time condition indicating after 00:00 on January 1, 2020) among the time series data (first and second time series data) referenced using the joined virtual table (the first and second virtual tables joined).

[0227] In this case, as shown in the upper part of Fig. 49, the time ranges of each of the first and second virtual tables are created from the boundary conditions indicated by the boundary management data shown in Fig. 48. Next, the time conditions of each of the first and second virtual tables are created based on the logical product of the time range indicated by the time condition included in the target query and the time ranges of each of the first and second virtual tables, thereby dividing the target query into a query for the time-series DB and a query for object storage A.

[0228] Specifically, as shown in the lower part of Figure 49, the query for the time series DB includes a time condition indicating the time range "time>='2021 / 01 / 31 12:30'" based on the logical product of the time range "time>='2020 / 1 / 1 00:00'" indicated by the time condition (WHERE' condition) included in the target query and the time range "time>='2021 / 1 / 31 12:30'" of the first virtual table shown in the upper part of Figure 49. On the other hand, as shown in the lower part of Figure 49, the query for object storage A includes a time condition indicating the time range "time>='2021 / 01 / 01 00:00' AND time<'2021 / 01 / 31 12:30'" based on the logical product of the time range "time>='2020 / 1 / 1 00:00'" indicated by the time condition (WHERE' condition) included in the target query and the time range "time<'2021 / 01 / 31 12:30'" of the second virtual table shown in the upper part of Figure 49.

[0229] In this case, first time series data based on a query for the time series DB shown in the lower part of Figure 49 is obtained from time series DB 12a, and second time series data based on a query for object storage A shown in the lower part of Figure 49 is obtained from object storage A.

[0230] Here, when the second time series data managed in the object storage A is updated to the third time series data, the aggregation processing worker 11a included in the data management service server 11 detects the update and adds the third time series data (aggregated file format data) as updated data to the object storage B. Furthermore, the aggregation processing worker 11a adds update management data (an entry in the update management table) to the storage unit 15g, which includes the update date and time when the second time series data was updated to the third time series data, the time range corresponding to the third time series data, and (the table name of) a virtual table for referencing the third time series data managed in the object storage B.

[0231] FIG. 50 shows the time range, boundary management data, and update management data corresponding to the first to third time series data as of time 4 (13:05 on January 31, 2021), which is later than time 1 described above.

[0232] Note that addition and deletion of time-series data to and from the time-series DB and object storage A are not considered here, and the time ranges corresponding to the time-series data managed in the time-series DB and object storage A are assumed to be the same as those in Fig. 48. The boundary management data is also assumed to be the same as those in Fig. 48.

[0233] On the other hand, for example, assume that the second time series data corresponding to the time range from 00:00 on January 31, 2020 to 23:59 on January 31, 2021, among the second time series data managed in object storage A, is updated to third time series data. In this case, the third time series data is added to object storage B, and third time series data (data for January 31) corresponding to the time range from 00:00 on January 31, 2020 to 23:59 on January 31, 2020 is managed in object storage B.

[0234] The third time series data (file format data) is added to object storage B by aggregation processing worker 11a included in data management service server 11, and aggregation processing worker 11a creates update management data in response to the addition of the third time series data and registers (stores) the created update management data in storage unit 15g included in integrated search engine server 15. If the second time series data is updated to the third time series data at time 4 (13:05 on January 31, 2021) as shown in FIG. 50, update management data including the update date and time "2021 / 1 / 31 13:05", the time range ">='2020 / 01 / 31 00:00'" AND "<'2020 / 2 / 1 00:00'"", and table "Table 3" is stored in storage unit 15g.

[0235] That is, according to Figures 48 and 50, when the third time series data (update data) for January 31, 2020 is added to object storage B, update management data (an entry having the time range) indicating the time range corresponding to the third time series data is created in the third virtual table for referencing the third time series data managed in object storage B.

[0236] Here, it is assumed that at time 4, the above target query "SELECT*FROM vTable WHERE time >='2020 / 1 / 1 00:00'" is specified by the user.

[0237] In this case, as shown in the upper part of Fig. 51, a time range for each virtual table is created from the boundary management data and the update management data shown in Fig. 50. Next, a time condition for each virtual table is created based on the logical product of the time range indicated by the time condition included in the target query and the time range for each virtual table, and the target query is divided into a query for the time-series DB, a query for object storage A, and a query for object storage B shown in the lower part of Fig. 51.

[0238] Note that the query for the time series DB shown in the lower part of Figure 51 includes a time condition indicating the time range "time>='2021 / 01 / 31 12:30'" based on the logical product of the time range "time>='2020 / 01 / 01 00:00'" indicated by the time condition (WHERE' condition) included in the target query and the time range "time>='2021 / 01 / 31 12:30'" of the first virtual table shown in the upper part of Figure 51. In addition, the query for object storage A shown in the lower part of Figure 51 includes a time condition indicating the time range "(time>='2020 / 1 / 1 00:00' AND time<'2020 / 1 / 31 00:00') OR (time>='2020 / 2 / 1 00:00' AND time<'2021 / 1 / 31 12:30')" based on the logical product of the time range "time>='2020 / 1 / 1 00:00'" indicated by the time condition (WHERE' condition) included in the target query and the time range of the second virtual table shown in the upper part of Figure 51 "time<'2020 / 1 / 31 00:00' OR (time>='2020 / 2 / 1 00:00' AND time<'2021 / 1 / 31 12:30')". Furthermore, the query for object storage B shown in the lower part of Figure 51 includes a time condition indicating the time range "time>='2020 / 1 / 31 00:00' AND time<'2020 / 2 / 1 00:00'" based on the logical product of the time range "time>='2020 / 1 / 1 00:00'" indicated by the time condition (WHERE' condition) included in the target query and the time range of the third virtual table shown in the upper part of Figure 51, "time>='2020 / 1 / 31 00:00' AND time<'2020 / 2 / 1 00:00'".

[0239] In this case, first time series data based on a query for the time series DB shown in the lower part of Figure 51 is obtained from time series DB 12a, second time series data based on a query for object storage A shown in the lower part of Figure 51 is obtained from object storage A, and third time series data based on a query for object storage B shown in the lower part of Figure 51 is obtained from object storage B.

[0240] Here, Figure 52 shows the relationship between the boundary management data and update management data at time 1 and time 4 and the data source (referenced data source) from which the time series data at time 1 and time 4 is obtained.

[0241] A time range 211 shown in FIG. 52 represents a time range corresponding to the first time series data managed in the time series DB at time 1 and time 4, and a time range 212 represents a time range corresponding to the second time series data managed in object storage A at time 1 and time 4. Furthermore, a time range 213 shown in FIG. 52 represents a time range corresponding to the third time series data managed in object storage B at time 4. Note that, at time 1, the third time series data has not been added (stored) in object storage B.

[0242] According to Figure 52, at time 1 when no update management data is stored (the third time series data has not been added to object storage B), the data source (referenced data source) that acquires the time series data corresponding to time range 214a is object storage A, and the data source (referenced data source) that acquires the time series data corresponding to time range 215a is the time series DB.

[0243] Also, at time 4 when the update management data was stored (the third time series data was added to object storage B), the data source (referenced data source) for acquiring the time series data corresponding to time range 214b was object storage A, and the data source (referenced data source) for acquiring the time series data corresponding to time range 215b was the time series DB. Furthermore, at time 4, the data source (referenced data source) for acquiring the time series data corresponding to time range 216 was object storage B.

[0244] According to the specific example of this embodiment described above, for example, when an update of the second time series data is detected by the aggregation processing worker 11a, the third time series data is added to the object storage B, and the update management data (an entry having a time range corresponding to the third time series data) is added to the integrated search engine server 15 (storage unit 15g). This allows the query (data search formula executed by the application program 14) specified by the user on the user terminal to obtain only the updated data while keeping the same expression, thereby realizing efficient search processing.

[0245] Although one update management data (entry) has been described in this embodiment, the update management data is added each time the second time series data is updated (i.e., the third time series data is added to object storage B). When the same second time series data has been updated multiple times (i.e., when there are multiple update management data with overlapping time ranges), the process of dividing the target query can be executed based on the update management data with the most recent update date and time. With this configuration, it is possible to realize a configuration that always acquires the latest time series data (updated data) even when the second time series data has been updated multiple times.

[0246] Furthermore, although the update management data has been described as including the update date and time, the update date and time may be used as a system column of the joined virtual table. In this embodiment, when the second time series data managed in object storage B is updated to the third time series data, the third time series data is preferentially acquired over the second time series data (i.e., object storage B is prioritized). However, by using the system column (update date and time) of the joined virtual table as a condition in the target query, it becomes possible to acquire the second time series data before the update from object storage A (search for past data that does not include the update data), for example, even after the second time series data has been updated to the third time series data.

[0247] Specifically, by adding a condition indicating the update date and time, such as "update_time<'2021 / 1 / 1 00:00", to the target query in the specific example above, it becomes possible to obtain the same time series data (i.e., search results) as at time 1, even at time 4, without referencing object storage B. In this case, "update_time" represents the column name of the system column.

[0248] According to at least one of the embodiments described above, it is possible to provide a data service system and method that can realize efficient search processing.

[0249] The first embodiment described above may be applied to the second embodiment as appropriate. Specifically, in the second embodiment, when the same data is managed in the time-series DB 12a and the object storage 13a (object storage A) (data is duplicated), the object storage 13a is given priority (i.e., the second time-series data is preferentially acquired), but by applying the second specific example described in the first embodiment described above to the second embodiment, the time-series DB 12a may be given priority in the second embodiment.

[0250] In addition, in the second embodiment, the multiple data sources are described as a time-series DB, object storage 13a, and object storage 13b. However, by applying the third specific example described in the first embodiment to the second embodiment, the multiple data sources in the second embodiment may be three object storages.

[0251] Furthermore, in the second embodiment, the boundary management data shown in the above-described FIGS. 37 and 38 may be adopted.

[0252] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention described in the claims and their equivalents. [Explanation of symbols]

[0253] 10...data service system, 11...data management service server, 11a...aggregation processing worker, 12...time series DB server, 12a...time series DB (first data source), 12b...data management unit, 12c...query execution unit, 13...object storage server, 13a, 3b...object storage (second and third data sources), 14...application program, 14a...query request unit, 14b...communication unit, 14c...result display unit, 15...integrated search engine server, 15a...communication unit, 15b...query analysis unit, 15c...query request unit, 15d...first connector unit, 15e...second connector unit, 15f...result aggregation unit, 15g...storage unit, 20...edge server, 30...sensor node, 101...CPU, 102...non-volatile memory, 103...main memory, 104...communication device.

Claims

1. A data service system includes a first data source that manages first time series data corresponding to a first time range, and a second data source that manages second time series data corresponding to a second time range that at least partially overlaps with the first time range, and executes a search process using virtual tables corresponding to the first and second data sources, a third data source that manages third time series data in which at least a part of the second time series data has been updated; a storage means for storing boundary management data indicating a temporal boundary condition for the first and second time series data determined based on an overlapping relationship between the first and second time ranges, and update management data indicating an eighth time range corresponding to the third time series data; an acquiring means for acquiring a first query for retrieving data from the first and second data sources, the first query including a first time condition indicating a third time range corresponding to data retrieved based on the first query; a dividing means for dividing the first query into a second query including a second time condition for the first data source, a third query including a third time condition for the second data source, and a fourth query including a fourth time condition for the third data source, based on the boundary management data and update management data stored in the storage means and a first time condition included in the acquired first query; aggregating means for aggregating first time series data acquired from the first data source based on the second query, second time series data acquired from the second data source based on the third query, and third time series data acquired from the third data source based on the fourth query; A data service system comprising:

2. the dividing means creates a fourth time range for acquiring the first time series data and a fifth time range for acquiring the second time series data based on the boundary condition indicated by the boundary management data; the second query includes a second time condition indicating a sixth time range based on a logical product of a third time range indicated by a first time condition included in the first query and the created fourth time range; The third query includes a third time condition indicating a seventh time range based on a logical product of a third time range indicated by a first time condition included in the first query and the created fifth time range.

2. The data service system according to claim 1.

3. if a sixth time range based on a logical product of a third time range indicated by a first time condition included in the first query and the created fourth time range does not exist, the first time series data is not acquired from the first data source; If a seventh time range based on a logical product of a third time range indicated by a first time condition included in the first query and the created fifth time range does not exist, the second time series data is not acquired from the second data source.

3. The data service system according to claim 2.

4. A data service system according to any one of claims 1 to 3, wherein the boundary management data is updated in response to changes in a first time range corresponding to first time series data managed in the first data source and a second time range corresponding to second time series data managed in the second data source.

5. the boundary management data includes a first time range corresponding to the first time series data, a second time range corresponding to the second time series data, and priorities assigned to the first and second data sources; The dividing means divides the first query into the second and third queries so that time-series data whose time ranges overlap between the first and second data sources is acquired from a data source with a high priority included in the boundary management data. The data service system according to any one of claims 1 to 4.

6. the storage means stores a plurality of update management data; each of the plurality of update management data includes an update date and time when at least a part of the second time series data was updated to the third time series data; 2. The data service system according to claim 1, wherein the dividing unit divides the first query based on the update management data that includes the most recent update date and time among the plurality of update management data stored in the storage unit.

7. the update management data includes an update date and time when at least a part of the second time series data was updated to the third time series data; The update date and time included in the update management data is used as a column of the virtual table, The first query includes a condition indicating the update date and time.

2. The data service system according to claim 1.

8. A method executed by a data service system including a first data source that manages first time series data corresponding to a first time range and a second data source that manages second time series data corresponding to a second time range that at least partially overlaps with the first time range, and that executes a search process using virtual tables corresponding to the first and second data sources, storing, in a storage means, boundary management data indicating a temporal boundary condition for the first and second time series data determined based on an overlapping relationship between the first and second time ranges and update management data indicating an eighth time range corresponding to third time series data in which at least a part of the second time series data managed in a third data source included in the data service system is updated; obtaining a first query for retrieving data from the first and second data sources, the first query including a first time condition indicating a third time range corresponding to data retrieved based on the first query; dividing the first query into a second query including a second time condition for the first data source, a third query including a third time condition for the second data source, and a fourth query including a fourth time condition for the third data source, based on the boundary management data and update management data stored in the storage means and a first time condition included in the acquired first query; aggregating first time series data acquired from the first data source based on the second query, second time series data acquired from the second data source based on the third query, and third time series data acquired from the third data source based on the fourth query; A method comprising:

Citation Information

Patent Citations

  • Pizza pie and toast

    JP1988071136A

  • Apparatus and method for processing information, program, and storage medium

    JP2010267075A

  • Data virtualization server, method for processing query in data virtualization server, and query processing program

    JP2016009425A

  • Method, program and system for automatic discovery of relationship between fields in environment where different types of data sources coexist

    JP2017188137A

  • Model generation device, abnormality occurrence prediction device, abnormality occurrence prediction model generation method and abnormality occurrence prediction method

    JP2020166407A