Data processing device, data processing method, and program product

Through the mapping and inverted structure of D5A file format and virtual table data, the problem of inheritance in indexes in virtual databases is solved, and quick sorting, retrieval and summary are realized, which is suitable for interactive operations of large-scale distributed data.

CN120584343APending Publication Date: 2025-09-02古庄 晋二
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480009105.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-27
Filing Date
2024-01-11
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing virtual database technology cannot effectively inherit the sorting, retrieval and summary index of actual databases, resulting in long operation time and large storage requirements.

Method used

By using D5A file format and virtual table data, the index inheritance is achieved using mapping and inverted structures, and the order, search and summary operations are automatically performed in virtual table data.

Benefits of technology

It shortens operation time, reduces storage requirements, and achieves fast sorting, retrieval and summary results, suitable for interactive operations of large-scale distributed data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120584343A_ABST
    Figure CN120584343A_ABST
Patent Text Reader

Abstract

A data processing device according to one aspect of the present invention comprises: a data operation unit that performs a sorting operation, a search operation, or an aggregation operation using a second data structure as an index on virtual table format data comprising virtual columns having the second data structure, and performs a sorting operation, a search operation, or an aggregation operation on the virtual table format data; the second data structure is obtained by converting a first data structure included in one or more columns included in the one or more pieces of table-format data using a predetermined map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data processing device, a data processing method and a program product. Background Art

[0002] In recent years, the development of various sensor devices and observation equipment has made it possible to obtain tabular data that stores large amounts of data representing sensor and observation results (so-called big data). Consequently, there is a growing need to extract multiple columns from one or more tabular data sets and generate virtual tabular data (hereinafter referred to as virtual tabular data) based on their intended use. One method for achieving these needs is to utilize a technology known as virtual databases or data virtualization (e.g., see Non-Patent Document 1). This technology executes subqueries on a distributed database when a query is received from a user.

[0003] Non-Patent Document 1: "Data Hub vs. Data Lake vs. Data Virtualization - MarkLogic (Data Center vs. Data Lake vs. Data Virtualization - MarkLogic)", Internet <URL: https: / / jp.marklogic.com / product / comparisons / data-hub-vs-data-lake / > Summary of the Invention

[0004] <Problems to be Solved by the Invention>

[0005] However, in a virtual database or a technology called data virtualization, the indexes used to implement sorting, retrieval, and aggregation of the actual database cannot be inherited by the virtual database.

[0006] In view of the above points, the present invention provides a technology that can inherit the indexes for sorting, searching, and summarizing actual table format data to virtual table format data.

[0007] <Methods for solving the problem>

[0008] According to one aspect of the present invention, a data processing device includes a data operation unit that performs a sorting operation, a search operation, or an aggregation operation on virtual table format data composed of virtual columns having a second data structure, using the second data structure as an index. The second data structure is obtained by converting a first data structure of one or more columns included in one or more table format data using a prescribed mapping.

[0009] <Effects of the Invention>

[0010] Provided is a technology that can inherit indexes used to sort, retrieve, and summarize actual table-based data to virtual table-based data. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a diagram for explaining an example of virtual table format data constructed hierarchically starting from D5A holding a value.

[0012] Figure 2 This is a diagram for explaining an example (one of) of map A.

[0013] Figure 3 This is a diagram for explaining an example (part 2) of mapping A.

[0014] Figure 4 This is a diagram for explaining an example of D5A.

[0015] Figure 5 This is a diagram for explaining an example of allocation.

[0016] Figure 6 This is a diagram for explaining an example of an allocation method based on enumeration-type mapping.

[0017] Figure 7 This is a diagram for explaining an example of an allocation method based on a linear function type mapping.

[0018] Figure 8 This is a diagram showing an example of a source column and a virtual column.

[0019] Figure 9 This is a diagram showing an example of a source column, an inverted structure, and a virtual inverted structure.

[0020] Figure 10 This is a diagram for explaining an example of allocation mapping of two-layer virtual table format data.

[0021] Figure 11 This is a diagram for explaining an example of a virtual inverted structure of two-layer virtual table format data.

[0022] Figure 12 This is a diagram showing an example of the overall configuration of a system including a data processing device according to this embodiment.

[0023] Figure 13 This is a diagram showing an example of the hardware configuration of the data processing device according to this embodiment.

[0024] Figure 14 This is a flowchart showing an example of the flow of virtual table format data generation processing.

[0025] Figure 15 This is a flowchart showing an example of the flow of sorting processing in virtual table format data.

[0026] Figure 16This is a flowchart showing an example of the flow of search processing in virtual table format data.

[0027] Figure 17 This is a flowchart showing an example of the flow of summary processing in virtual table format data. DETAILED DESCRIPTION

[0028] The following describes one embodiment of the present invention. In the following embodiment, necessary explanations and definitions are first provided, followed by a definition of data called D5A, which represents tabular data. Next, a method for generating new virtual tabular data from D5A or other virtual tabular data is described, as well as a method for inheriting the index of (actual) tabular data and sorting, searching, and aggregating the virtual tabular data. Furthermore, a method for inheriting the index even when the virtual tabular data is hierarchically constructed is described. Finally, a data processing device 10 capable of generating, sorting, searching, and aggregating virtual tabular data is described.

[0029] <Preface>

[0030] Archived data, such as IoT (Internet of Things) data, various observational data, and log data, accumulates daily. In most cases, tabular data is combined and added to the archive at regular intervals (e.g., daily or monthly). Tabular data added to the archive can be treated as read-only. This archived tabular data can be large and is often distributed across LANs (Local Area Networks) and the internet.

[0031] Typically, attempting to utilize the above-mentioned archived tabular data requires the following two steps, both of which take a long time and often have the problem of consuming a large storage area.

[0032] The first step is to generate new tabular data. This is done by combining multiple tabular data sets, joining them, or extracting only the necessary columns. This process can be very time-consuming if the original tabular data is large or distributed across a wide area network such as the Internet. Furthermore, if the new tabular data is large, a large storage area is required to accommodate it.

[0033] The second step involves sorting, searching, or summarizing. Newly generated tabular data is not indexed, so sorting, searching, and summarizing take time. Furthermore, large amounts of new tabular data, such as large sorting results, large search results with a large number of hits, and large summaries, require significant storage space.

[0034] With the huge growth of archived data and the increasing demand for distributed archived data usage today, the problems of the above two steps are becoming increasingly serious.

[0035] Therefore, the present application proposes a technology that shortens the time required for the operation of the above two steps to the time required to establish interactive operation, and only requires a small amount of necessary storage area. The tabular data to be operated can be, for example, 1 million records or 100,000 columns. In addition, combinations of UNIONs and JOINs of the tabular data to be operated can be accumulated in layers. In addition, the tabular data to be operated can be distributed on a LAN or on multiple HTTP (Hypertext Transfer Protocol) servers on the Internet.

[0036] This is possible because a network system that maps archived data is constructed using a file format called D5A, which contains tabular data with indexes on all columns to accelerate sorting, retrieval, and aggregation; virtual tabular data that inherits values ​​directly or indirectly from the D5A file; and virtual indexes on the virtual tabular data, automatically established using a data structure directly or indirectly inherited from the indexes used by the D5A file. The virtual tabular data and virtual indexes are immediately usable simply by being directly or indirectly linked to the D5A file, and they consume only a small amount of storage space. Furthermore, even the largest sorting, retrieval, and aggregation results obtained using the virtual indexes require only a small recording area.

[0037] This archive data mapping network using D5A makes it possible to use new archive data, which is described as "users creating virtual tabular data for each purpose and interactively using archive data distributed across the network." For example, within each organization, archive data distributed across each department within the organization can be used for various purposes, and archive data such as IoT (Internet of Things) data from around the world can be combined or extracted on the Internet and utilized in various ways.

[0038] Here, the terms used in this specification will be sorted out. First, D5A refers to a file format for tabular data, which has an index that accelerates the sorting, retrieval and aggregation of all columns. Secondly, virtual tabular data is tabular data that inherits values ​​from a D5A file (a file in D5A format) or other virtual tabular data. It should be noted that the tabular data of the inheritance source is called source tabular data. In addition, the columns on D5A are called D5A columns, and the indexes that accelerate the sorting, retrieval and aggregation of D5A columns are called D5A indexes. Similarly, virtual tabular data are called virtual columns and virtual indexes. Similarly, source tabular data are called source columns and source indexes. In addition, the data structure used by the D5A index is called an inverted structure, the data structure used by the virtual index is called a virtual inverted structure, and the data structure used by the source index is called a source inverted structure. It should be noted that virtual tabular data can be called "mapped tabular data" or the like, and similarly, virtual columns can be called "mapped columns" or the like.

[0039] <Solution to the problem>

[0040] As mentioned above, attempting to utilize archived tabular data presents the following problem: Generally, two steps are required: Step 1 and Step 2. Each step is time-consuming and often consumes a large amount of storage space. Step 1 involves extracting only the necessary columns through UNION or JOIN to generate new tabular data, while Step 2 involves sorting, searching, and summarizing.

[0041] The reason why step 1 takes time is because it takes time to read, compare and store the values. Therefore, the following method can be considered instead: there is tabular data as the source tabular data, and a mapping defined by a correspondence table and rules that stipulate which unit is mapped to which unit of the newly generated tabular data is used to generate new tabular data through the mapping. In the case of archived data, it is not uncommon to be able to define the mapping in a concise and inductive manner. In this case, the new tabular data does not need to have values, which is a great advantage especially when the new tabular data is large. The new tabular data can be displayed in a short time after the mapping definition is loaded, and the required storage area can be a small block for saving the mapping definition. Because the new tabular data does not save values, but inherits values ​​from the original tabular data, it is virtual tabular data. The problem of step 1 can be solved with such virtual tabular data.

[0042] In step 2, the newly generated tabular data columns lack indexes, and writing sorting, retrieval, and summary results often requires time. These two factors contribute to the time consumption. However, all columns in the virtual tabular data automatically include virtual indexes that accelerate column sorting, retrieval, and summary. Furthermore, even sorting, retrieval, and summary results with large virtual indexes can be stored using only a small amount of storage space, thus shortening the writing time. These two factors address the time consumption issue. Furthermore, by using only a small amount of storage space to accommodate the results of the virtual indexes, the storage space issue is also addressed. Thus, the issue in step 2 can be resolved using virtual indexes.

[0043] The D5A file expressed in a storage format of table format data called D5A holds values ​​serving as sources of virtual table format data and a data structure serving as a source of a data structure used as a virtual index.

[0044] D5A: Storage format for tabular data as a source of values ​​and indices

[0045] Virtual tabular data inherits values ​​from one or more source tabular data. The source tabular data can be either another virtual tabular data or tabular data called a D5A file. In the former case, the virtual tabular data further inherits values ​​from another source tabular data, ultimately arriving at the D5A file. Therefore, virtual tabular data can be considered to be hierarchically structured by directly or indirectly inheriting values ​​from D5A files.

[0046] Virtual indexes use virtual inverted structures to accelerate sorting, retrieval, and aggregation. These virtual inverted structures are automatically created by inheriting from one or more source inverted structures. A source inverted structure can be either a virtual inverted structure on other virtual tabular data or an inverted structure on a D5A file. In the former case, the virtual inverted structure further inherits from other source inverted structures, ultimately arriving at the inverted structure on the D5A file. Therefore, virtual inverted structures can be considered to be hierarchically constructed, directly or indirectly inheriting from the inverted structure on the D5A file.

[0047] D5A is a storage format for tabular data that provides values ​​to virtual tabular data and inverted structures to the virtual inverted structures, and has a D5A index that holds values ​​and uses inverted structures in all its columns.

[0048] Implementation of a Network System for Mapping Archived Data

[0049] A simple hypothetical example of a network system representing a map of archived data and will refer to Figure 1This section explains how to use D5A, virtual table format data, and virtual indexes described so far. Figure 1 This shows the process of combining the weather observation data from Sunday to Saturday for each area of ​​Tokyo, Osaka, and Nagoya into one. This process is done in two stages. First, the 7 daily data for each area are combined into weekly data, and then the weekly data are concatenated to form a single tabular data. Here, Figure 1 The 21 daily data for one week in the three areas on the left are stored in the 55A file. Figure 1 The three tabular data in each area in the center are virtual tabular data. Each of these virtual tabular data uses Figure 1 The seven D5A files for one week on the left are used as source tabular data. The necessary columns are extracted from them, and then a UNION operation is performed. Figure 1 The virtual table data on the right is used Figure 1 The three virtual tabular data of each area in the center are used as source tabular data and are concatenated.

[0050] like Figure 1 As shown, the network of archive data mappings is constructed hierarchically from virtual table-like data starting with the D5A file. As described above, the virtual table-like data automatically incorporates virtual indexes, enabling quick sorting, retrieval, and aggregation of any column. Furthermore, the storage area required to store sorting, retrieval, and aggregation results is minimal. This virtual table-like data can be generated and used by users, potentially promoting widespread use of archive data.

[0051]

[0052] D5A, virtual table data, and virtual indexes are all described by combining mappings from a continuous interval of natural numbers starting at 0 to a specific value (which is also a natural number in most cases). Because the subscript (index) of this mapping can be represented by a one-dimensional array starting at 0, it is called an A-map. By using the A-map, the correspondence relationship can be viewed holistically, resulting in the ability to derive algorithms that utilize the properties of sets or groups. The advantages of the A-map are described below.

[0053] The first advantage of A-map is that A-maps can be combined in various ways to generate new A-maps. As a simple example, by combining A-map S representing the result column of the search with A-map C representing column A A Combined, we can get the A mapping that represents the search results of column A, and by combining the A mapping C that represents column B B By combining them, we can obtain the A mapping that represents the search results of column B. In this case, S only needs to be the A mapping that represents the result column, and C A and CB Any A-map that represents a column can be used. Furthermore, it can exist on a local storage device or on a network. Such a combination of A-maps can be considered an algebraic combination.

[0054] Furthermore, the combination result of A mapping and A mapping does not necessarily have to be written into the storage area, but can also be a virtual A mapping. A virtual A mapping refers to an A mapping that uses one or more A mappings as a combination source and can extract any i-th part from the combination source knowing the size of the whole. Since a virtual A mapping is also an A mapping, further virtual A mappings can be generated hierarchically by combining virtual A mappings. The column of virtual table format data is a type of virtual A mapping and can be generated hierarchically. An index is a mechanism implemented by a data structure for indexing and an algorithm using the data structure. A virtual index is an index that uses one or more virtual A mappings as a data structure for indexing and can be generated hierarchically.

[0055] The second advantage of the A-map is that it can decompose the A-map to generate a new A-map. By decomposing the A-map, a new A-map with functions that the A-map before decomposition does not have can be generated and used. One of the particularly effective decompositions is the LP decomposition (decomposition into A-map L and A-map P) of the A-map M that determines the correspondence (mapping) between cells between tabular data. When M is decomposed into L and P, elements can be efficiently retrieved through a bisection search in L, and the inverse mapping can be obtained through P. At this time, M only needs to be an A-map that defines the mapping, L automatically becomes an A-map that can be binary searched, and P automatically becomes an A-map with inverse elements. Since the decomposition of the A-map produces a new A-map, it can be considered an algebraic decomposition.

[0056] On the other hand, since the A map is a one-dimensional array, it takes time to insert and delete elements in a large A map. However, in the case of archived data, updates are rarely performed, so this disadvantage is not a problem.

[0057] By using such an A-map, it is possible to design a D5A, a storage format for tabular data that includes indexes in any column, enabling high-speed sorting, retrieval, and aggregation. Furthermore, D5A and A-maps can be used to define virtual tabular data. This virtual tabular data automatically includes virtual indexes using one or more A-maps as the data structure for the index. This allows for the realization of a network system for mapping archived data that can be used by combining archived data distributed across a network in various ways.

[0058] Therefore, below, we first define the A-map and its representation. Next, we introduce the index operator as an operator for combining A-maps. Next, we describe the decomposition of the A-map, and explain the particularly important SN decomposition, LP decomposition, and spectral decomposition. Finally, we classify A-maps according to four aspects.

[0059] Definition of A-mapping

[0060] Consider a one-dimensional array with an index starting at 0 and size N. This one-dimensional array can be viewed as a mapping whose domain is a continuous interval of natural numbers from 0 to N-1 and whose range is a discrete interval of at most N possible values. This is called an A-map. If records are assigned consecutive record numbers starting at 0, then the columns of tabular data can also be viewed as an A-map with the record number as the domain.

[0061] For example, Figure 2 As shown in the figure, consider a one-dimensional array that stores "4" as the 0th element, "0" as the 1st element, "6" as the 2nd element, and "3" as the 3rd element. Since this one-dimensional array maps 0 to "4", 1 to "0", 2 to "6", and 3 to "3", it can be considered an A-map with a domain of {0, 1, 2, 3} and a range of {0, 3, 4, 6}.

[0062] In addition, for example Figure 3 As shown, consider a one-dimensional array that stores "Bob" as the 0th element, "Alice" as the 1st element, "Cathy" as the 2nd element, and "Bob" as the 3rd element. Since this one-dimensional array corresponds to 0 for "Bob", 1 for "Alice", 2 for "Cathy", and 3 for "Bob", it can be considered an A-map with a domain of {0, 1, 2, 3} and a range of {Alice, Bob, Cathy}.

[0063] It should be noted that although the A-map can be expressed as a one-dimensional array, when discussing aspects of this mapping, it is referred to as a map, while when discussing operations on this mapping, it is referred to as an array. Furthermore, when considering it as a mapping, the terms domain and range are used, while when it is an array of data in column and table form, the terms record number and value are used. However, whenever possible, the A-map is expressed as an array.

[0064] Representation of A-mapping

[0065] The following describes the representation of A-map including the representation of general permutations.

[0066] Notation 1. When defining an A map by enumerating elements, it is written as (a0, a1, ..., a n-1 ).

[0067] Representation 2. When the domain n of A is explicitly represented, it is denoted as A (n) According to this notation, for a column, for example, R is the total number of records and C is (R) wait.

[0068] Representation 3. When it is known that the range of A mapping A is on natural numbers and the maximum value of its range does not exceed n-1, in order to clearly represent the range, it is recorded as A (n) .

[0069] Representation 4. A maps A (n) The i-th element of (n) [i]. Therefore, A (n) ≡(A (n) [0],A (n) [1],...,A (n) [n-1]).

[0070] Representation 5. The concatenation of (i0,i1,...),(j0,j1,...),... is expressed as (i0,i1,...)+(j0,j1,...)+...

[0071] Representation 6. Let the mapping of λ on the value range of A map to i0,i1,... from the definition domain be denoted as λ:(i0,i1,...). For example, Figure 3 The example shown in can be denoted as Alice:(1), Bob:(0,3), Cathy:(2).

[0072] Notation 7. When λ0:(i0,i1,...),λ1:(j0,j1,...),... are connected together, the symbol + is used to express it as λ0:(i0,i1,...)+λ1:(j0,j1,...)+... However, in this case, the order is λ0<λ1<... For example, Figure 3 The example shown in can be written as Alice:(1)+Bob:(0,3)+Cathy:(2).

[0073] Index Operators for Combining A Maps

[0074] The following defines an operator "·" called the Index Operator, which combines A-maps to produce a new A-map.

[0075] A (n) ·B (m) (n) =(A (n) [B (m) (n)[0]],...,A (n) [B (m) (n) [m-1]]) Formula (1)

[0076] The above index operator has the properties shown in the following properties 1 to 3.

[0077] Property 1: The associative rule holds. That is, (A (n) ·B (m) )·C (k) =A (n) ·(B (m) ·C (k) ) was established.

[0078] Property 2: The size of the permutation of the combined result is the size of the permutation on the right.

[0079] Property 3: The data type of the permutation representing the result of the combination is the data type of the permutation on the left.

[0080] Here, examples are used to determine the binding rules. (3) =(Alice, Bob, Cathy), B (4) =(1,0,2,1),C (3) =(0,2,1). At this time, (A (3) ·B (4) )·C (3) =(Bob, Alice, Cathy, Bob)·(0,2,1)=(Bob, Cathy, Alice). On the other hand, A (3) ·(B (4) ·C (3) )=(Alice,Bob,Cathy)·(1,2,0)=(Bob,Cathy,Alice), the two are consistent.

[0081] According to Properties 1 and 2 above, the number of possible values ​​in the A-map obtained by combining A-maps using an index operator is equal to or smaller than the size of the smallest A-map among the A-maps combined by the index operator. This is a metric similar to the rank in a matrix.

[0082] Decomposition of A-Mapping

[0083] When using the index operator to combine A-maps, only one A-map is obtained. On the other hand, a single A-map can be decomposed into various A-maps. For example, (Bob, Alice, Cathy, Bob) can be decomposed into (Bob, Alice, Cathy, Bob) = (Alice, Bob, Cathy) (1, 0, 2, 1) or (Bob, Alice, Cathy, Bob) = (Bob, Alice, Cathy) (0, 1, 2, 0).

[0084] In this case, a unique decomposition method is particularly valuable, either for accelerating sorting and retrieval or for obtaining an inverse mapping. Therefore, the following describes the unique decomposition methods, namely, SN decomposition, LP decomposition, and spectral decomposition.

[0085] SN decomposition

[0086] In a method of decomposing an A-map into two, the left A-map of the two obtained A-maps only retains the values ​​that appear in the left A-map, and there is only one way to keep these values ​​in a unique and ascending order. This decomposition is called SN decomposition.

[0087] For example, (Bob, Alice, Cathy, Bob) = (Alice, Bob, Cathy)·(1, 0, 2, 1) is an example of SN decomposition.

[0088] The first item on the right, unique and in ascending order, is called the Sorted Value List (SVL). In the example above, SVL = (Alice, Bob, Cathy). The second item on the right can be thought of as replacing the elements on the left with the storage locations in the SVL. This is called the Natural Numbered Column (NNC). In the example above, NNC = (1, 0, 2, 1).

[0089] Such SN decomposition is generally defined as follows.

[0090] C (R) =SVL (K) ·NNC (R) (K) ;

[0091]

[0092] In formula (2), the upper part represents column C (R) Decomposition into SVL (K) and NNC (R) (K) , the lower part represents SVL (K)The elements of NNC are arranged in ascending order. (R) (K) As can be seen from its label, it is an A map of size R with natural numbers 0,...,K-1 as elements.

[0093] By using SN decomposition, an A-map NNC containing natural numbers can be extracted from an A-map that does not necessarily contain natural numbers. Using the extracted NNC, an efficient algorithm using the properties of natural numbers, such as counting sort, can be used.

[0094] (SN decomposition algorithm)

[0095] Here, an example of the SN decomposition algorithm is described using (Bob, Alice, Cathy, Bob) as an example. In the SN decomposition algorithm, the following steps 1-1 to 1-5 are performed.

[0096] Step 1-1: Generate ((Bob,0),(Alice,1),(Cathy,2),(Bob,3)) by adding the position in the A map to each value of (Bob,Alice,Cathy,Bob).

[0097] Step 1-2: Sort the value and position groups while evaluating the size relationship to obtain ((Alice, 1), (Bob, 0), (Bob, 3), (Cathy, 2)).

[0098] Step 1-3: Assign a value number starting from 0 to each value in ascending order to obtain the sorting result ((Alice=0,1),(Bob=1,0),(Bob=1,3),(Cathy=2,2)).

[0099] Note that in the example above, the same value, Bob, appears twice, and each occurrence is assigned the value number 1. Furthermore, the sorted result has the structure ((value = value number, position in map A), ...). For example, the first element in the sorted result (Alice = 0, 1) has the value Alice, the value number 0, and the position in map A.

[0100] Step 1-4: Use the above sorting results to secure a storage area for SVL (area size 3) and write it as "SVL[value number] = value". That is, write it as SVL[0] = Alice, SVL[1] = Bob, and SVL[2] = Cathy. This completes SVL = (Alice, Bob, Cathy).

[0101] Step 1-5: Use the above sorting results to secure a storage area for NNC (area size 4) and write it as "NNC[position in A map] = value number." That is, write NNC[1] = 0, NNC[0] = 1, NNC[3] = 1, and NNC[2] = 2. This completes NNC = (1, 0, 2, 1).

[0102] Based on the above, we can obtain (Bob, Alice, Cathy, Bob) = (Alice, Bob, Cathy)·(1, 0, 2, 1).

[0103] LP decomposition

[0104] LP decomposition is a special case of SN decomposition, which decomposes a mapping M (a non-repeating natural number) into two elements. For example, (5,2,7,3) = (2,3,5,7)·(2,0,3,1) is an example of LP decomposition.

[0105] The first term on the right side, obtained through LP decomposition, consists of natural numbers, unique elements, and in ascending order. This term is called L (selection). In the above example, L = (2, 3, 5, 7). Meanwhile, the second term on the right side can be considered to be formed by permuting the elements on the left side with their respective positions in L. This term is called P (permutation). In the above example, P = (2, 0, 3, 1). In LP decomposition, M, L, and P are all the same size.

[0106] If LP decomposition is performed, the presence and position of elements in M ​​can be determined using L, and the inverse element can be determined using P. It should be noted that LP decomposition can be implemented using the same algorithm as SN decomposition.

[0107] Spectral decomposition

[0108] The above representation 7 represents the position (index) of each value on the A-map where that value appears as a new A-map. This is called spectral decomposition of the A-map. Spectral decomposition allows the A-map to be viewed from its range to its domain, enabling various algorithms to be implemented. The following example shows how spectral decomposition of columns as A-maps facilitates column sorting, retrieval, and aggregation.

[0109] (Execution of sorting)

[0110] exist Figure 3In the example shown, for sorting, by removing the value from the spectral decomposition of the column Alice:(1)+Bob:(0,3)+Cathy:(2) and setting it to (1)+(0,3)+(2), and performing the connection specified in Representation 5, the sorted result column (1,0,3,2) can be obtained. It should be noted that the sorting result can be obtained by combining the column C as follows (R) =(Bob,Alice,Cathy,Bob) and sort the result columns to obtain.

[0111] Sorting result = C (R) Sorted result column = (Bob, Alice, Cathy, Bob) (1, 0, 3, 2) = (Alice, Bob, Bob, Cathy)

[0112] (Execution of Search)

[0113] exist Figure 3 In the example shown, for the search of Bob, the part specifying Bob can be decomposed from the spectrum of the column Alice:(1)+Bob:(0,3)+Cathy:(2) by binary search, and Bob:(0,3) can be extracted to obtain the search result column (0,3). It should be noted that the column C can be decomposed as follows (R) =(Bob, Alice, Cathy, Bob) is combined with the search result column to obtain the search results.

[0114] Search results = C (R) Search result column = (Bob, Alice, Cathy, Bob) (0, 3) = (Bob, Bob)

[0115] (Execution of Aggregation)

[0116] exist Figure 3 In the example shown, the summary takes the values ​​and their occurrence times from the spectral decomposition of the column Alice:(1)+Bob:(0,3)+Cathy:(2), and the summary result = Alice: 1 time, Bob: 2 times, Cathy: 1 time can be obtained.

[0117] (General formula for spectral decomposition)

[0118] Here, the general rewritten spectrum is decomposed as follows.

[0119] [Number 1]

[0120]

[0121] In formula (3), R is the total number of records, λ is the key value, w is the number of postings numbers belonging to the key value λ, and X is the permutation of the postings numbers.

[0122] For example, when written in the form of formula (3) Figure 3 When the sorting results of the one-dimensional arrangement are shown, the situation is as follows.

[0123] R=4, λ0=Alice, λ1=Bob, λ2=Cathy

[0124] w0=1, w1=2, w2=1

[0125] X 0(1) (4) =(1),X 1(2) (4) =(0,3),X 2(1) (4) =(2)

[0126] (Algorithm for spectral decomposition)

[0127] Here, an example of an algorithm for performing spectral decomposition on the A-map is shown, taking (Bob, Alice, Cathy, Bob) as an example. This algorithm consists of the following steps 2-1 to 2-3.

[0128] Step 2-1: Add the position in the A map to each value of (Bob, Alice, Cathy, Bob) to generate ((Bob, 0), (Alice, 1), (Cathy, 2), (Bob, 3)).

[0129] Step 2-2: Use the value and position groups to sort while evaluating the size relationship, and obtain ((Alice, 1), (Bob, 0), (Bob, 3), (Cathy, 2)).

[0130] Step 2-3: After summarizing the positions within the A map for each identical value, the spectral decomposition Alice:(1)+Bob:(0,3)+Cathy:(2) is completed.

[0131] Classification of A-Mappings

[0132] A-maps can be classified based on four criteria: associativity, searchability, whether they represent sets, and ease of inversion. These categories will be helpful for understanding the following explanations, so we'll explain them here first.

[0133] ·Associability

[0134] A mappings with natural numbers as elements can be on both sides of the index operator. A mappings without natural numbers as elements can only be on the left side of the index operator.

[0135] Retrieval

[0136] An A-map with elements arranged in ascending order can be efficiently searched using binary search to determine whether a specific element exists and, if so, where it is located. This type of A-map is called an ascending order. For an A-map with unique, ascending elements, if a search for a certain value yields a match, there will only be one such match. This type of A-map is called a unique ascending order.

[0137] Whether it represents a set

[0138] If the elements of an A map are not repeated, then it can be considered to represent a set. However, this specification also adds the premise that the elements are natural numbers. In this specification, a set is a result column and is an A map with natural numbers (record numbers) as elements.

[0139] Is it easy to find the inverse?

[0140] For a map A with natural numbers as its elements, its size is N. If A includes all elements with values ​​from 0 to N-1, then the inverse of the index operator can be easily found. This is called a symmetric permutation. The definition of a symmetric permutation P is as follows.

[0141]

[0142] Symmetrical permutations generate groups with respect to the index operator. Therefore, there exists an identity element and an inverse element. The inverse element of P is used to identify a cell of the source tabular data from a cell of the virtual tabular data.

[0143] Unit element: E = (0, 1, ..., N-1)

[0144] Inverse element:

[0145] <D5A>

[0146] Let R represent the number of records in tabular data, and K represent the types of values ​​contained in a column. Tabular data refers to data consisting of R records, identified sequentially by numbers from 0 to R-1 from the top, and one or more columns, identified sequentially by numbers or names from the left. Each record and column defines a single value, and each column has a single data type (e.g., integer, floating-point number, string, etc.). Furthermore, indexing is established when, given the size W of the sorting, searching, or summarizing results for the target column, the i-th row of the sorting, searching, or summarizing results can be read in approximately O(log(R)) or less when i∈0,1,...,W-1 is specified.

[0147] The following first describes the internal structure of D5A, then explains how to extract values, and finally explains the D5A index in conjunction with related matters.

[0148] The Internal Structure of D5A

[0149] Will refer to Figure 4 Describe the internal structure of the D5A. Figure 4 Including column C (4) = (Bob, Alice, Cathy, Bob) in CSV (Comma Separated Value) format and its D5A format.

[0150] D5A has a structure formed by juxtaposing two structures described below.

[0151] The first structure is used to obtain the value of the column and is composed of Figure 4 The portion surrounded by the dotted line in . This structure is composed of (a) storing the column as A mapping as is, (b) storing the SVL and NNC obtained by performing SN decomposition on the column, and (c) storing either of the two. In the case of (a) storing as is, the performance of reading values ​​is improved, but the algorithm using NNC cannot be used. In the case of (b) storing the SVL and NNC, the algorithm using NNC can be used, but the performance of reading values ​​is reduced. In the case of (c) storing both, the size of the D5A file increases. In this embodiment, the case of (b) will be explained.

[0152] The second structure is a structure for D5A index, which is composed of Figure 4 The dotted line encloses the portion in [ ]. This structure, consisting of SVL, ACM, and INV, is called an inverted structure. The inverted structure is obtained by rewriting the spectral decomposition of the D5A column described above into the form of an A map. The details of the inverted structure will be described later.

[0153] If an inverted structure (ie, the second structure) is obtained, it forms a D5A column together with the first structure. If the D5A columns are combined, D5A is formed.

[0154] Structure for getting column values

[0155] By (R) Perform SN decomposition to generate a structure for obtaining the column value. For example, Figure 4 In the example shown, C (4) =(Bob,Alice,Cathy,Bob)=(Alice,Bob,Cathy)·(1,0,2,1). Therefore, SVL (3) =(Alice, Bob, Cathy), NNC (4) (3)=(1,0,2,1).

[0156] On the other hand, the column is obtained by the following formula using SVL and NNC.

[0157] C (R) =SVL (K) ·NNC (R) (K) Formula (4)

[0158] Therefore, the i-th row of the column is obtained by the following formula.

[0159] value=C (R) [i]=SVL (K) ·NNC (R) (K) [i] Formula (5)

[0160] For example, in Figure 4 In the example shown, the first row of the column is (Alice, Bob, Cathy)·(1,0,2,1)[1]=Alice.

[0161] In addition, the following information 1 to information 4 can also be obtained from SVL.

[0162] Information 1: From SVL (K) Find the number of types of values ​​in the column K. For example, Figure 4 In the example shown, K is 3.

[0163] Information 2: Starting from the smallest value, the i-th value can be read as SVL[i].

[0164] Information 3: Whether a value v exists in the column is known through a binary search using SVL. Furthermore, if the value v exists, the order of the value v is known.

[0165] Information 4: SVL is the summary dimension. It should be noted that sorting is required in the past to generate the summary dimension.

[0166] Inverted Structure

[0167] Next, the method of generating an inverted structure as a structure of a D5A index is described. Although it has been described that sorting, retrieval, and aggregation can be performed at high speed through spectral decomposition columns, there are two difficulties in spectral decomposition. The first difficulty is that since it is not completed through A mapping, it cannot be accessed using an index operator (the first problem point). The second difficulty is that when trying to obtain the i-th item in the sorted result column, it is necessary to retrieve the i-th item while adding w0, w1,... (the second problem point). The structure obtained by converting the spectral decomposition into three A mappings is an inverted structure, which can solve these two difficulties. The following describes its definition and generation method.

[0168] Definition of inverted structures and their generation based on spectral decomposition

[0169] The definition of the inverted structure is as follows.

[0170] [Number 2]

[0171]

[0172] Here, SVL (K) 、ACM (K) (R+1) 、INV (R) (R) They are obtained through the following steps 3-1 to 3-3, respectively.

[0173] Step 3-1: SVL (K) is obtained by decomposing the value part of the spectrum of formula (3) into λ0,λ1,...,λ K-1 It is obtained by arranging it into A mapping as it is. In addition, compared with the SVL obtained by formula (4), (K) same.

[0174] Step 3-2: ACM (K) (R+1) yes

[0175] [Number 3]

[0176]

[0177] is the value λ of formula (3) i The number of occurrences w i This is done by accumulating and mapping to A. This step takes O(K) time.

[0178] Step 3-3: INV (R) (R) yes

[0179] [Number 4]

[0180]

[0181] is to convert [number 5] of formula (3)

[0182]

[0183] Connect them to form an A map. Here, x i,j is λ i The j-th record number from the top among the record numbers that appear.

[0184] [Number 6]

[0185]

[0186] is corresponding to λ0<λ1<...<λ K-1 , so INV is the permutation of the inverted record numbers.

[0187] For example, in Figure 4 In the example shown, the inverted structure is <(Alice, Bob, Cathy):(1,3,4),(1,0,3,2)>.

[0188] Spectral decomposition is not an A-mapping. If you want to get the i-th column of the sorted result, you need w i The inverted structure derived from spectral decomposition is a structure composed of three A maps, which can solve the first problem. Since the i-th item in the sorted result column can be obtained through INV[i], it can also solve the second problem.

[0189] Methods for obtaining spectral decomposition from inverted structures

[0190] Spectral decomposition can be obtained from the inverted structures within D5A by the following method.

[0191] First, in the spectral decomposition of formula (3), λ i By λ i =SVL ( K) [i] is obtained. In addition, if it is predetermined that ACM[-1]≡0, then

[0192] [Number 7]

[0193]

[0194] Get it below.

[0195] [Number 8]

[0196]

[0197] According to the above two, the i-th term of the spectral decomposition is calculated below.

[0198] [Number 9]

[0199]

[0200] For example, in Figure 4 In the case of the example shown, the first term of the spectral decomposition is obtained as follows.

[0201] [Number 10]

[0202]

[0203] Information obtained from the inverted structure

[0204] When the inverted structure is used, in addition to the information obtained from SVL (the above-mentioned information 1 to information 4), the following information 5 to information 8 can also be obtained.

[0205] Information 5: The number of occurrences of the i-th value starting from the smaller value can be found below.

[0206] count=ACM (K) [i]-ACM (K) [i-1] Formula (8)

[0207] For example, in Figure 4 In the example shown, the number of occurrences of the first value starting from the smaller value is ACM (K) [1]-ACM (K) [0]=3-1=2.

[0208] Information 6: The i-th record number of the sorting result can be read below.

[0209] The i-th record number of the sorting result = INV[i] Formula (9)

[0210] For example, in Figure 4 In the example shown, the first record number of the sort result can be read as INV[1]=0(Bob).

[0211] Information 7: Record numbers having i-th values ​​from small to large can be read out as the following arrangement.

[0212] (INV (R) (R) [ACM (K) [i-1]],...,INV (R) (R) [ACM (K) [i]-1])

[0213] For example, in Figure 4 In the case of the example shown, the record number having the first value (Bob) from smallest to largest can be read as (INV[1], ..., INV[2]) = (0, 3).

[0214] Information 8: Record numbers with values ​​v0 to v1 can be read below.

[0215] Let i0 be the smallest i0 that satisfies v0≤SVL[i0], and i1 be the largest i1 that satisfies SVL[i1]≤i1, (INV (R) (R) [ACM (K)[i0-1]],...,INV (R) (R) [ACM (K) [i1]-1]) Formula (10)

[0216] For example, in Figure 4 In the example shown, since the record numbers with values ​​Bob to Cathy are i0=1 and i1=2, they can be read as (INV[ACM[1-1]],...,INV[ACM[2]-1])=(INV[1],...,INV[4-1])=(0,3,2).

[0217] Ability to generate indexes using inverted structures

[0218] As mentioned above, the requirement for an index is that, given the size W of the sorted, searched, or summarized results, the i-th row of the sorted, searched, or summarized results can be read in approximately O(log(R)) or less, given i∈0,1,...,W-1. Below, we demonstrate that an index can be generated using an inverted structure.

[0219] (Sort)

[0220] The size of the sorted result is known to be R. The i-th element of the sorted result can be read out in O(1) using formula (9). Therefore, the inverted structure meets the requirements for generating an index for sorting.

[0221] (Search)

[0222] Explain the case of formula (10). It is found that the size of the search result is ACM (K) [i1]-ACM (K) [i0-1]. For example, in Figure 4 In the example shown, the record numbers with values ​​Bob to Cathy are i0=1, i1=2, so ACM (K) [2]-ACM (K) [1-1]=4-1=3.

[0223] Can be achieved through (INV (R) (R) [ACM (K) [i0-1]],...,INV (R) (R) [ACM (K) [i1]-1])[i] reads the i-th element of the search result in O(1). For example, Figure 4 In the example shown, when searching by values ​​Bob to Cathy, since i0=1 and i1=2, the search can be performed by (INV (R) (R) [ACM(K) [1-1]],...,INV (R) (R) [ACM (K) [2]-1])[i]=(0,3,2)[i] is read.

[0224] Therefore, the inverted structure meets the requirements for generating an index for retrieval.

[0225] (Summary)

[0226] The size of the summary result is known to be from information 1 to K. The i-th value of the summary result is given by information 2, which is obtained in O(1). The number of occurrences of the i-th value of the summary result is given by formula (8), which is obtained in O(1). Therefore, the inverted structure meets the requirements for generating an index for the summary.

[0227] For example, in Figure 4 In the case of the example shown, when i=1, the size of the summary result is 3, the first value of the summary result is Bob, and the number of occurrences of the first value of the summary result is ACM[1]-ACM[1-1]=3-1=2.

[0228] As described above, it can be said that a D5A index can be generated that uses an inverted structure to accelerate sorting, retrieval, and aggregation.

[0229] Less storage space is required to maintain sorted, retrieved, and summarized results through D5A indexes

[0230] (Sorting results)

[0231] The sorting results are already stored in INV, so no new storage area is required.

[0232] (Search results)

[0233] The case of formula (10) is explained. The search results are arranged by (INV (R) (R) [ACM (K) [i0-1]],...,INV (R) (R) [ACM (K) [i1]-1]). In this case, INV and ACM have already been generated, and the storage area required to store the search results is only the storage area of ​​i0 and i1, which is irrelevant to the number of hits.

[0234] (Summary results)

[0235] The i-th value of the summary result is given by the above information 2, and no new storage area is required. The number of occurrences of the i-th value of the summary result is given by formula (8), and no new storage area is required.

[0236] As described above, D5A provides values ​​for virtual table data and indexes for virtual indexes. Using the SVL and NNC obtained by performing SN decomposition on a column, not only can values ​​be provided, but also the four types of information (Information 1 to Information 4) described above can be obtained. Furthermore, it is shown that D5A indexes can be generated using inverted structures (i.e., inverted structures can be used as indexing data structures for sorting, searching, and summarizing). D5A indexes not only accelerate sorting, searching, and summarizing, but also require only a small storage area to store the results of sorting, searching, and summarizing.

[0237] <Virtual table format data>

[0238] For example, when executing a JOIN or UNION in an RDB (Relational Database), the source tabular data is read, compared, the data to be written is generated, its storage location is determined, and it is written to a newly secured storage area. This then requires time and storage space. This is the first step in attempting to utilize archived tabular data. The larger the newly generated tabular data, the more serious this problem becomes. Generating it takes a long time and requires a large storage area. Moreover, only a small portion of the generated tabular data is typically used.

[0239] Virtual table form data solves the above problem by generating a mechanism of necessary parts when necessary. That is, in terms of virtual table form data, when a certain unit needs to be displayed, the unit of the source table form data corresponding to the unit at that time point is referenced and displayed. However, to do this, it is necessary to pre-define a mapping to define the correspondence between the unit of the virtual table form data and the units of one or more source table form data. The mapping can be a mapping from virtual table form data to multiple source table form data, or vice versa, it can be a mapping from multiple source table form data to virtual table form data, and the two are essentially the same. However, for the former, the mapping target is a set of two pieces of information: which unit of which source table form data, while for the latter, only one piece of information: which unit of the virtual table form data is required, which is easy to define. Therefore, the latter will be explained below. It should be noted that both require inverse mapping. For the former, the virtual index requires inverse mapping, and for the latter, when referencing the source table form data to obtain a value, the virtual table form data requires inverse mapping.

[0240] The mapping of which cells in each source tabular data item correspond to which cells in the virtual tabular data item is called an allocation map. However, the actual allocation map defines the mapping from each source column (i.e., a column in the source tabular data item) to a virtual column (i.e., a column in the virtual tabular data item). Each source column is divided into one or more intervals, and for each interval, the cells in that interval are associated with the cells in the virtual column using a correspondence table or rule. This association for each interval is called an interval map. The allocation map is the combination of all interval maps.

[0241] 《Generation steps of allocation mapping》

[0242] Will refer to Figure 5 The generation steps of the allocation map will be described. Note that, in the following text, rs represents the record number on the source table format data, and rv represents the record number on the virtual table format data.

[0243] The columns on the virtual table format data, that is, the virtual columns, are composed of allocation maps from one or more source columns. Figure 5 In the example shown, the virtual column is allocated from source column 0 and source column 1.

[0244] have Figure 5 The blank D5A for source column 1 in the lower left corner is a D5A programmatically generated automatically when the virtual column is generated. The blank D5A is used to assign blank values ​​to unassigned cells in the virtual column so that the virtual column satisfies the fully injective condition described later. When the fully injective condition is met, a virtual index is automatically included in the virtual tabular data. On the other hand, in some cases, sorting, searching, and summarizing the virtual tabular data is unnecessary. For example, if you want to use the entire result of a UNION or JOIN as is. Of course, in this case, the virtual tabular data can be defined and used without considering the fully injective condition.

[0245] Each source column is divided into one or more source intervals. Figure 5 In the example shown, source column 0 is divided into interval 1 and interval 2. There are two types of mappings from each source interval to a virtual column (i.e., interval mapping). The first is an enumeration-type mapping defined by A mapping M, where the elements are natural numbers and there is no overlap. The second is a linear function-type mapping defined by a linear function. The former, i.e., the enumeration-type mapping, can define any allocation mapping, but consumes a large amount of storage area. The latter, i.e., the function-type mapping, is a mapping obtained by replacing the M of the enumeration-type mapping with a linear function, which consumes only a small amount of storage area and can define a huge interval mapping, but it only corresponds to regular allocations. When interval mappings are defined in all source intervals, the allocation mapping is completed.

[0246] If any of the above interval mappings are defined, their inverse mappings and the conditions for their existence are also defined. When displaying data in virtual table format, the source interval where the inverse mapping exists is first determined based on the record number in the virtual column. The inverse mapping of the interval mapping is then used to determine the record number in the source column, and the value obtained from that record number is displayed.

[0247] Below, after explaining the interval mapping based on enumeration type mapping and the interval mapping based on linear function type mapping, refer to Figure 5 This section explains how to display data in a virtual table format.

[0248] Interval Mapping and Its Inverse Mapping Based on Enumeration Mapping

[0249] Figure 6 This example shows an allocation based on an enumeration map. Figure 6 In the table, Rs represents the size of the source column, rs represents the record number of the source column, Rv represents the size of the virtual column, rv represents the record number of the virtual column, Q represents the interval on the source column, q represents the starting position of interval Q, u represents the length of interval Q, V represents the interval on the virtual column corresponding to interval Q, v represents the starting position of interval V, and w represents the length of interval V.

[0250] At this time, the interval mapping F:Q→V based on the enumeration type mapping is expressed as follows.

[0251] rv=M (u) (Rv) [rs-q] Formula (11)

[0252] In addition, its inverse mapping F -1 :V→Q is expressed as follows.

[0253] M (u) (Rv) =L (u) (Rv) ·P (u) (u) ;

[0254] L (u) (Rv) [j]=rv;

[0255] rs=P -1 (u) (u) [j]+q formula (12)

[0256] Enumeration mapping

[0257] First, we will explain the formula (11) that defines the interval mapping based on the enumeration type mapping. Figure 6In , the numbers written next to the source column and virtual column are record numbers. M can be defined by writing the record numbers on the corresponding virtual columns in sequence from the beginning of the source interval. Figure 6 In the example shown, it is (7, 5, 9, 1). Here, the starting position q of the source interval Q is 2.

[0258] The source interval (interval Q) starts at rs = 2 and ends at rs = 5. When equation (11) is applied for each rs, the result is as follows.

[0259] When rs = 2, rv = (7, 5, 9, 1) [2-2] = 7

[0260] When rs = 3, rv = (7, 5, 9, 1) [3-2] = 5

[0261] When rs = 4, rv = (7, 5, 9, 1) [4-2] = 9

[0262] When rs = 5, rv = (7, 5, 9, 1) [5-2] = 1

[0263] Inverse mapping of enumeration mapping

[0264] Next, the formula (12) defining the inverse mapping of the enumeration type mapping is described.

[0265] M is a mapping of A with non-repeating natural numbers as elements, which can be decomposed by LP. The top part of formula (12) represents this LP decomposition. Through LP decomposition, we get L = (1, 5, 7, 9) and P = (2, 1, 3, 0). L is a unique ascending order, and the presence of an element (if an element exists, its position) can be found by binary search. P is a symmetric order with an inverse element P. -1 =(3,1,0,2).

[0266] The middle part of formula (12) shows the method of finding the number of occurrences of rv in L, where j is its occurrence position. It should be noted that if j is not found, rv means that it has not received the mapping from the source interval Q and there is no inverse mapping.

[0267] The lower part of formula (12) shows that P -1 =(3,1,0,2), q=2 and the above j to find rs.

[0268] Based on the above, try to calculate rs when rv = 1, 5, 7, 9.

[0269] When rv=1, from (1,5,7,9)[j]=1,j=0; rs=(3,1,0,2)[j]+2=5

[0270] When rv=5, from (1,5,7,9)[j]=5,j=1;rs=(3,1,0,2)[j]+2=3

[0271] When rv=7, from (1,5,7,9)[j]=7,j=2; rs=(3,1,0,2)[j]+2=2

[0272] When rv=9, from (1,5,7,9)[j]=9, j=3; rs=(3,1,0,2)[j]+2=4

[0273] Interval Mapping and Its Inverse Mapping Based on Linear Function Type Mapping

[0274] Figure 7 This shows an example of allocation based on a primary function mapping. Figure 7 In the table, Rs represents the size of the source column, rs represents the record number of the source column, Rv represents the size of the virtual column, rv represents the record number of the virtual column, Q represents the interval on the source column, q represents the starting position of interval Q, u represents the length of interval Q, V represents the interval on the virtual column corresponding to interval Q, v represents the starting position of interval V, and w represents the length of interval V.

[0275] At this time, the interval mapping F:Q→V based on the linear function type mapping is as follows.

[0276] rv=a×rs+b; a, b are integers, a≠0 Formula (13)

[0277] In addition, its inverse mapping F -1 :V→Q is expressed as follows.

[0278] rs=(rv-b) / a Formula (14)

[0279] When v≤rv<v+w is not satisfied and rs is not an integer, rv is not the value range of F.

[0280] ·One-time function mapping

[0281] First, the formula (13) defining the linear function type mapping is explained. Figure 7 In the example, the numbers next to the source column and virtual column are record numbers. a is the interval on the virtual column. Figure 7 In the example shown, a=3 and b=-5.

[0282] The source interval (interval Q) starts at rs = 2 and ends at rs = 5. When each rs is applied to equation (13), the following results are obtained.

[0283] When rs=2, rv=3×rs-5=1

[0284] When rs=3, rv=3×rs-5=4

[0285] When rs=4, rv=3×rs-5=7

[0286] When rs=5, rv=3×rs-5=10

[0287] Inverse mapping of a linear functional mapping

[0288] Next, the formula (14) defining the inverse mapping of the linear function type mapping will be described: As described above, a=3, b=-5.

[0289] First, we need to determine v ≤ rv < v + w. Since v = 1 and w = 10, 1 ≤ rv < 1 + 10 must be true. If this condition is satisfied, use the upper part of equation (14) to find rs. At this point, as shown in the lower part of equation (14), if rs is not an integer, then rv means that the mapping from the source interval Q is not accepted, and there is no inverse mapping.

[0290] When rv=1, rs=(1-(-5)) / 3=2

[0291] When rv=4, rs=(4-(-5)) / 3=3

[0292] When rv=7, rs=(7-(-5)) / 3=4

[0293] When rv=10, rs=(10-(-5)) / 3=5

[0294] Display of data in virtual table format

[0295] Refer again Figure 5 This section explains how to display data in a virtual table format. Source interval 1 in source column 0 is an interval mapping based on a linear function type mapping, source interval 2 in source column 0 is an interval mapping based on an enumeration type mapping, and source interval 1 in source column 1 is an interval mapping based on a linear function type mapping. The definitions of each interval mapping are as follows.

[0296] The interval mapping in the source interval 1 of the source column 0 is defined by equation (13) as F:rv=rs×2+1.

[0297] The interval mapping in source interval 2 of source column 0 is defined by equation (11) as F:rv=(2,7,0)[rs-3]. Substituting rs=3,4,5 as shown below confirms that it is correct.

[0298] When rs = 3, rv = (2, 7, 0) [3-3] = 2

[0299] When rs = 4, rv = (2, 7, 0) [4-3] = 7

[0300] When rs = 5, rv = (2, 7, 0) [5-3] = 0

[0301] According to equation (13), the interval mapping in the source interval 1 of the source column 1 of the blank D5A is defined as F: rv = rs × 2 + 4.

[0302] When the above interval mappings F are defined, their inverse mappings F -1 Automatically determined as follows.

[0303] According to formula (14), the inverse mapping of source column 0 in source interval 1 is F -1 :rs=(rv-1) / 2.

[0304] According to formula (12), the inverse mapping of source column 0 in source interval 2 is F -1 :rs=(2,0,1)[j]+3;rv=(0,2,7)[j]. First, when rv=0, a binary search is performed on (0,2,7) and the storage location j=0 of rv is obtained, so rs=(2,0,1)[0]+3=5. Similarly, when rv=2, a binary search is performed on (0,2,7) and the storage location j=1 of rv is obtained, so rs=(2,0,1)[1]+3=3. Similarly, when rv=7, a binary search is performed on (0,2,7) and the storage location j=2 of rv is obtained, so rs=(2,0,1)[2]+3=4.

[0305] According to formula (13), the inverse mapping of source column 1 of blank D5A in source interval 1 is F -1 :rs=(rv-4) / 2.

[0306] To sum up, when rv=0,1,2,3,5,7, we get rs=5,0,3,1,2,4 in source column 0, and when rv=4,6, we get rs=0,1 in source column 1.

[0307] "How to Structure Virtual Table Format Data"

[0308] In order to realize the functions described above, the virtual table format data is composed of the following information.

[0309] 1. Number of records

[0310] 2. Virtual column name and data type

[0311] 3. One or more URLs or paths to source tabular data

[0312] 4. Definition of source columns and source intervals in the above source table format data

[0313] 5. Definition of interval mapping and its inverse mapping for each source interval

[0314] Except for 5, no large storage area is required in the above. Furthermore, the need for a large storage area in 5 is limited to cases where the source range is large and the mapping is enumerated. Therefore, virtual table format data can be compactly constructed in many cases.

[0315] Detecting allocation conflicts from different intervals

[0316] Here, a method of checking whether there is a conflict in allocation destinations based on allocation mapping will be described. Hereinafter, it is assumed that two allocation mappings are F0:Q0→V0 and F1:Q1→V1.

[0317] First, it is obvious that the two allocation maps F0 and F1 do not conflict when any one of the following (conditions 1-1) to (conditions 1-3) is satisfied.

[0318] (Condition 1-1) V0 and V1 are not on the same virtual column

[0319] (Condition 1-2) When v0+w0-1<v1

[0320] (Condition 1-3) When v1+w1-1<v0

[0321] It should be noted that v0 is the starting position of V0, w0 is the length of V0, v1 is the starting position of V1, and w1 is the length of V1.

[0322] If it is not the obvious case mentioned above, it can be checked whether there is a conflict between the two allocation maps F0 and F1 through the following (case 1) to (case 3).

[0323] (Case 1) F0 and F1 are both enumeration mappings

[0324] Assume L0 = (v0, ..., v0 + w0 - 1) and L1 = (v1, ..., v1 + w1 - 1). The common range of L0 and L1 is newly set to L0 and L1. Note that w0 is the length of V0, and w1 is the length of V1.

[0325] Then, for i=0, j=0, the following steps 4-1 to 4-3 are executed.

[0326] Step 4-1: If L0[i]=L1[j], end as "there is a conflict."

[0327] Step 4-2: If L0[i] < L1[j], return to step 4-1 as i←i+1. However, as a result of i←i+1, if i exceeds the maximum value of the index representing the element of L0, the process ends as "no conflict".

[0328] Step 4-3: If L0[i]>L1[j], return to step 4-1 as j←j+1. However, as a result of j←j+1, if j exceeds the maximum value of the index representing the element of L1, the process ends with "no conflict".

[0329] (Case 2) F0 and F1 are both linear function mappings

[0330] The common part of V0 and V1 is selected and F0 and F1 are re-described by the following equation with this common part as the value range. In this case, a0 and a1 are adjusted to be positive.

[0331] y0=a0×x0+b0; x0=0,1,2,...,N0-1

[0332] y1=a1×x1+b1; x1=0,1,2,...,N1-1

[0333] Here, a0 and a1 are integers greater than or equal to 1, and b0 and b1 are integers.

[0334] Then, let x0=0 and x1=0, and execute the following steps 5-1 to 5-3.

[0335] Step 5-1: If y0=y1, end as "there is a conflict".

[0336] Step 5-2: If y0 < y1, then x0 ← x0 + MaxInt(1, (y1 - y0) div a0) is executed, and the process returns to Step 5-1. However, if x0 > MinInt(N0 - 1, LCM(a0, a1) div a0) as a result of x0 ← x0 + MaxInt(1, (y1 - y0) div a0), the process ends as "no conflict."

[0337] Step 5-3: If y0 > y1, then x1←x1+MaxInt(1,(y0-y1)div a1) is executed and the process returns to step 5-1. However, if x1 > MinInt(N1-1,LCM(a0,a1)div a1) as a result of x1←x1+MaxInt(1,(y0-y1)div a1), the process ends as "no conflict."

[0338] Among them, MaxInt represents a function that obtains the maximum integer, MinInt represents a function that obtains the minimum integer, div represents integer division (discarding the remainder), and LCM represents a function that returns the least common multiple.

[0339] (Case 3) When either F0 or F1 is an enumeration type mapping and the other is a linear function type mapping

[0340] Below, let F0 be a linear function mapping and F1 be an enumeration mapping. Select the common part of V0 and V1 and rewrite F0 and F1 using the following equation with this common part as the range. In this case, adjust a0 to make it positive.

[0341] y0(n)=a0×n+b0; x0=0,1,2,...,N-1

[0342] y1(m)=(v1,...,v1+w1-1)[m]

[0343] Among them, a0>0. In addition, (v1,...,v1+w 1-1 ) is the ascending order of size M.

[0344] Then, let n=0 and m=0, and execute the following steps 6-1 to 6-3.

[0345] Step 6-1: If y0(n)=y1(m), end as "there is a conflict."

[0346] Step 6-2: If y0(n)>y1(m), return to step 6-1 as m←m+1. However, as a result of m←m+1, if m=M, the process ends as "no conflict".

[0347] Step 6-3: If y0(n) < y1(m), return to step 6-1 as n←(y1(m) - y0(n)) div a0. However, as a result of n←(y1(m) - y0(n)) div a0, when n≥N, the process ends as "no conflict".

[0348] <Virtual Index>

[0349] We have described a method for generating virtual table-like data that inherits values ​​directly or indirectly from D5A. Below, we explain why a virtual inverted structure automatically establishes on virtual table-like data by directly or indirectly inheriting an inverted structure from D5A. When a virtual inverted structure is established, a virtual index for the data structure that uses it as an index also automatically establishes. Virtual indexes use only virtual inverted structures instead of inverted structures, and because the algorithm is the same as that of D5A indexes, they have the same functions and characteristics as D5A indexes. This means that column sorting, retrieval, and aggregation can be accelerated, and even large sorting, retrieval, and aggregation results can be stored using only a small storage area.

[0350] Similar to inverted structures, virtual inverted structures have virtual SVL, virtual ACM, and virtual INV. Virtual SVL, virtual ACM, and virtual INV are virtual permutations. Virtual permutations are mechanisms that function similarly to permutations. If the size W is known in advance and i∈0, 1, ..., W-1 is specified, the i-th element can be retrieved in approximately O(log(R)) or less, eliminating the need to store values ​​as required by virtual tabular data. If the allocation mapping defined for virtual tabular data satisfies the following fully injective conditions, virtual SVL, virtual ACM, and virtual INV automatically hold true.

[0351] The conditions for total injection

[0352] 1. (Mapping is successful) All cells in the source column are mapped to cells in the virtual column.

[0353] 2. (It is injective) No cell on any virtual column is mapped to from cells in more than two source columns.

[0354] 3. (It is fully reflective) All cells on the virtual column are mapped from cells in the source column.

[0355] 4. (Source sequence is fully injective) When the source sequence is D5A, it can be considered to be fully injective. When the source sequence is a virtual sequence, if the source sequence satisfies 1 to 3 above, it can be considered to be fully injective.

[0356] First, we explain how a virtual inverted structure can be constructed from multiple inverted structures. We also confirm that using the virtual indexes of this virtual inverted structure allows for high-speed sorting, retrieval, and aggregation, and that the sorting, retrieval, and aggregation results can be stored by reserving only a small amount of new storage space. We then confirm that even with virtual inverted structures and inverted structures, further virtual inverted structures can be generated hierarchically, and that even with their virtual indexes, high-speed sorting, retrieval, and aggregation can be performed, and that the sorting, retrieval, and aggregation results can be stored by reserving only a small amount of new storage space. Finally, we confirm that using spectral decomposition makes it easy to understand the behavior of even hierarchically generated virtual inverted structures.

[0357] Method for constructing virtual inverted structures

[0358] As an example, in Figure 8 In the source column #0:C 0(4) and source column #1:C 1(4) Define virtual column C V(8) In order to distinguish C 0(4) and C 1(4)The elements of are represented by subscripts A0, B0, C0, B1, C1, and D1, along with the source column number, with values ​​such as B0 = B1 and C0 = C1. In this case, since the allocation targets from source column #0 are the 5th, 7th, 2nd, and 3rd rows of the virtual column, the allocation map F0 = (5, 7, 2, 3). Similarly, since the allocation targets from source column #1 are the 4th, 0th, 6th, and 1st rows of the virtual column, the allocation map F1 = (4, 0, 6, 1).

[0359] The inverted structure of source column #0, the inverted structure of source column #1 and the virtual inverted structure are as follows Figure 9 Here, the SVL, ACM, and INV in the inverted structure of source column #0 are represented as SVL0, ACM0, and INV0, respectively. The SVL, ACM, and INV in the inverted structure of source column #1 are represented as SVL1, ACM1, and INV1, respectively. In addition, the SVL, ACM, and INV in the virtual inverted structure of the virtual column are represented as SVL V 、ACM V 、INV V .

[0360] INV0 and INV1 both store inverted record numbers. The record numbers on the source column are mapped to the record numbers on the virtual column through allocation mapping. Therefore, when reading INV0 and INV1, they must be replaced with INV0' and INV1', respectively. INV0' and INV1' can be calculated using the following formula (15).

[0361] INV0′=F0·INV0=(5,7,2,3)·(3,0,2,1)=(3,5,2,7) Formula (15) −1

[0362] INV1′=F1·INV1=(4,0,6,1)·(1,2,0,3)=(0,6,4,1) Formula (15)-2

[0363] Below, Figure 9 For example, according to the virtual arrangement SVL V 、ACM V 、INV V The order of describing the construction of the virtual inverted structure is as follows.

[0364] Virtual Array SVL V The composition method

[0365] First, explain how to find SVL V The size of the value is determined by adding the order of the source column to the values ​​B and C that appear in both SVL0 and SVL1. That is, B0 of SVL0 is smaller than B1 of SVL1.V The size of is 6, that is, the sum of the size of SVL0 = 3 and the size of SVL1 = 3, which is uniquely determined.

[0366] Next, we explain how to find SVL. V The i-th element of SVL V [i] method. Select an appropriate value v from SVL0 or SVL1. However, this v complies with the above-mentioned rule that determines the size relationship by considering the order of the source column. Next, find the sum j of the number of values ​​smaller than v in SVL0 and SVL1. If j < i, select v' that is larger than v and repeat the above operation. The same is done when j > i (however, select v' that is smaller than v). If j = i, then v at this time is SVL V [i].

[0367] The above SVL V [i]'s method can be performed in about O(log(K)). Therefore, since we know SVL V The size of , and its i-th element can be taken out in less than O(log(R)), so it can be considered that there is a virtual permutation SVL V .

[0368] The example when i=0 is shown. First, select B1 of SVL1 as v. There are two values ​​less than v in SVL0, and there are 0 values ​​less than v in SVL1. Therefore, j=2+0=2, which is greater than i=0. Therefore, next, try to select A0 of SVL0 as v. Then, j=0, j=i. Therefore, SVL V [0]=A0.

[0369] The example when i=4 is shown. First, select B0 of SVL0 as v. There is a value smaller than v in SVL0, and there is 0 in SVL1. Therefore, j=1+0=1, which is smaller than i=4. Therefore, next, try to select C1 of SVL1 as v. Then, j=3+1=4, j=i. Therefore, SVL V [4]=C1.

[0370] Through the above, the virtual array SVL can be generated V =(A0,B0,B1,C0,C1,D1).

[0371] Virtual Arrangement ACM V The composition method

[0372] First, as SVL V Find the size of ACM V size.

[0373] Next, explain the ACM V The i-th element of ACMV [i] Method. First, according to the “virtual array SVL V The method described in "Construction method" is used to obtain v, that is, v = SVL V [i]. Next, find the maximum i0 that satisfies SVL0[i0]≤v. If i0 is not found, then i0=-1. Note that ACM[-1]≡0.

[0374] Similarly, find the maximum i1 that satisfies SVL1[i1]≤v. Then, ACM V [i] As ACM V [i]=ACM0[i0]+ACM1[i1].

[0375] Obtain the above ACM V The method [i] can be executed in about O(log(K)). Therefore, knowing that ACM V The size of , and its i-th element can be taken out in less than O(log(R)), so it can be considered that there is a virtual permutation ACM V .

[0376] The example when i=3 is shown. First, select B1 of SVL1 as v. There are two values ​​less than v in SVL0, and 0 in SVL1. Therefore, j=2, which is less than i=3. Therefore, next, select C0 of SVL0 as v. Then, j=2+1=3, SVL V [3] = C0.

[0377] Next, if we find the largest i0 that satisfies SVL0[i0]≤v, we find i0=2. Next, if we find the largest i1 that satisfies SVL1[i1]≤v, we find i1=0. Therefore, ACM V [3] As ACM V [3]=ACM0[2]+ACM1[0]=4+2=6.

[0378] According to the above, it is possible to generate a virtual array ACM V =(1,3,5,6,7,8).

[0379] Virtual Arrangement INV V The composition method

[0380] First, know INV V The size of the source columns is R, where R is the sum of the sizes of the source columns.

[0381] Next, we explain how to obtain the ACM V [j-1]≤i≤(ACM V [j]-1) to obtain INVV [i] The method of .j is a V Determined INV V The number of the interval above.

[0382] ACM V The size is K V .K V is the sum of the sizes of ACM0 and ACM1. First, from 0,1,...,K V -1, select the appropriate j, through the "virtual arrangement ACM V The interval ACM is calculated using the method described in "Construction method of V [j-1]~ACM V [j]-1.

[0383] If i < ACM V [j-1], then reselect the interval by making j have a smaller value. On the other hand, if i>ACM V [j] - 1, then reselect the interval by making j have a larger value. Repeat this step to determine j. If j is determined, then pre-calculate offset = i - ACM V [j-1].

[0384] Next, we find v = SVL V [j]. Then, find the value v and its source column. Using this v, find k that makes SVL[k]=v through SVL on the source column. Thus, as INV V [i]=INV'[ACM[k-1]+offset] and find INV V [i] Here, INV', ACM are INV', ACM on the source column.

[0385] The above is to find INV V The method [i] can be performed in about O(log(K)). Then, since we know that INV V The size of , and can efficiently extract its i-th element, so it can be considered that there is a virtual arrangement INV V .

[0386] The example when i=4 is shown. First try to select j=3. The lower limit of the interval is ACM V [j-1]=ACM V [2]=5, the upper limit of the interval is ACM V [j]-1=ACM V [3]-1=5, it can be seen that the selected interval is too large. Therefore, try to select j=2 again. The lower limit of the interval is ACM V [j-1]=ACM V[1]=3, the upper limit of the interval is ACM V [j]-1=ACM V [2] -1 = 4. Therefore, we know that i = 4 belongs to the interval of j = 2. offset = i - ACM V [j-1]=4-ACM V [1] = 1. Then, we know that v = SVL V [j]=SVL V [2] = B1. Next, we can see that k is 0 when SVL[k] = B1 on source column #1. Thus, we get INV V [4]=INV'[ACM[k-1]+offset]=INV'[1]=6.

[0387] Through the above, it is possible to generate a virtual array INV V =(3,5,2,0,6,7,4,1).

[0388] Sorting using virtual indexes

[0389] Try to use virtual indexes using virtual inverted structures to sort. Figure 8 and Figure 9 NNC is not described in , so it is used instead Figure 8 C v(8) The sorting results are as follows.

[0390] C v(8) ·INV V =(B1,D1,B0,A0,C1,B0,B1,C0)·(3,5,2,0,6,7,4,1)=(A0,B0,B0,B1,B1,C0,C1,D1)

[0391] From the above results, we can see that the sorting is indeed done. In addition, we can also see that even if there are the same "B" and "C" in the sorting results, C 0(4) The "B" and "C" also appear in C 1(4) Before "B" and "C." This reflects the rule that when values ​​are the same, the smaller the column number, the smaller the value. Also, like the D5A index, no new storage area is required to store the sort results.

[0392] Search using virtual indexes

[0393] Try to use the virtual index using the virtual inverted structure to search. Figure 8 、 Figure 9 NNC is not listed in the table, so it is used instead. Figure 8 C v(8) . In addition, suppose that a search is performed under the conditions "B" to "C".

[0394] Since i0=1、i1=4 in formula (10), so (INV (R) (R) [ACM (K) [i0-1]],...,INV (R) (R) [ACM (K) [i1]-1])=(5,2,0,6,7,4). Therefore, the search results are as follows.

[0395] C v(8) ·(5,2,0,6,7,4)=(B1,D1,B0,A0,C1,B0,B1,C0)·(5,2,0,6,7,4)=(B0,B0,B1,B1,C0,C1)

[0396] In addition, similar to the D5A index, only i0 and i1 need to be saved to save the search results, so only a small amount of storage space is required.

[0397] Aggregation using virtual indexes

[0398] Summarize using a virtual index using a virtual inverted structure. The i-th value of the summary result is given by the above information 2, and the number of occurrences of the i-th value of the summary result is given by formula (8). If the summary is performed based on this, the following table 1 will be shown.

[0399] [Table 1]

[0400] i value frequency Remark 0 <![CDATA[A0]]> 1 From source column #0 1 <![CDATA[B0]]> 2 From source column #0 2 <![CDATA[B1]]> 2 From source column #1 3 <![CDATA[C0]]> 1 From source column #0 4 <![CDATA[G1]]> 1 From source column #1 5 <![CDATA[D1]]> 1 From source column #1

[0401] Here, even for identical values, if the D5A indexes providing those values ​​differ, separate summary results are generated. Therefore, users who prefer this approach remain unchanged, while users who prefer not to do so need to sum the number of occurrences of identical values. Since the maximum sum does not exceed the number of D5As, processing time is minimal. Furthermore, similar to the D5A index, no new storage space is required to store the summary results.

[0402] As mentioned above, we find the existence of virtual SVL, virtual ACM, and virtual INV. Therefore, we can say that there is a virtual inverted structure. The inverted structure is the data structure used for D5A indexing, and for the same reason, the virtual inverted structure can be considered the data structure used for virtual indexing.

[0403] Furthermore, the virtual inverted structure is composed of references to the inverted structure of the source column. However, when the source column is a virtual column, the inverted structure is a virtual inverted structure. Therefore, in this case, the virtual inverted structure is constructed hierarchically.

[0404] Virtual index of virtual table format data generated hierarchically

[0405] Although it has been described that virtual inverted structures can be generated hierarchically, this is verified below through an example. Figure 10 Shows the use of D5A file C 0(3) and C 1(2) Generate virtual table format data C V0(5) And use the virtual table data C V0(5) and D5A file C 2(3) Generate further virtual table form data C V1(8) The following is an example of an allocation mapping definition.

[0406] F(C 0(3) →C V0(5) )=(2,3,0)

[0407] F(C 1(2) →C V0(5) )=(4,1)

[0408] F(C V0(5) →C V1(8) )=(2,3,7,1,5)

[0409] F(C 2(3) →C V1(8) )=(0,6,4)

[0410] Figure 11 Indicates that this time Figure 10 How is the inverted structure of C? 0(3) 、C 1(2) 、C V0(5) 、C 2(3) The inverted structure or virtual inverted structure of is shown in the figure. As in formula (15), each INV is replaced by INV' as follows.

[0411] INV0'=F(C 0(3) →C V0(5) )·INV0=(2,3,0)·(1,0,2)=(3,2,0)

[0412] INV1'=F(C 1(2) →C V0(5) )·INV1=(4,1)·(0,1)=(4,1)

[0413] INV V0 '=F(C V0(5) →C V1(8) )·INV V0 =(2,3,7,1,5)·(3,4,2,0,1)=(1,5,7,2,3)

[0414] INV2′=F(C 2(3) →C V1(8) )·INV2=(0,6,4)·(2,1,0)=(4,6,0)

[0415] C V0(5) 、C V1(8) The virtual inverted structure can be constructed using the SVL, ACM, and INV' described above through the steps described above.

[0416] Here, an attempt is made to sort, retrieve, and summarize hierarchically generated virtual table-form data using a virtual index.

[0417] Sorting using virtual indexes of hierarchically formed virtual table-like data

[0418] Try to sort using virtual indexes using a hierarchically formed virtual inverted structure. Figure 10 and Figure 11 NNC is not described in , so it is used instead Figure 10 C in V1(8) The sorting results are as follows.

[0419] C V1(8) ·INV V1 =(D2,A0,B0,C1,A2,A1,B2,B0)·(1,5,4,7,2,6,3,0)=(A0,A1,A2,B0,B0,B2,C1,D2)

[0420] From the above results, we can see that the sorting is indeed done. In addition, we can also see that even if there are the same "A", "B" and "C" in the sorting results, C 0(3) "A" and "B" also appear in C 1(2) and C 2(3) Before "A" and "B." This reflects the rule that when values ​​are the same, the smaller the column number, the smaller the value. Also, like the D5A index, no new storage area is required to store the sort results.

[0421] Search using virtual indexes of hierarchically formed virtual table format data

[0422] Try to use the virtual index that uses the virtual inverted structure formed hierarchically to search. Figure 10 、 Figure 11 NNC is not listed in the table, so it is used instead. Figure 10 C V1(8) . In addition, suppose that a search is performed under the conditions "B" to "C".

[0423] Since i0=3、i1=5 in formula (8), so (INV (R) (R) [ACM (K) [i0-1],...,INV (R) (R) [ACM (K) [i1]-1])=(7,2,6,3). Therefore, the search results are as follows.

[0424] C V1(8) ·(7,2,6,3)=(D2,A0,B0,C1,A2,A1,B2,B0)·(7,2,6,3)=(B0,B0,B2,C1)

[0425] In addition, similar to the D5A index, only i0 and i1 need to be saved to save the search results, so only a small amount of storage space is required.

[0426] Aggregation of virtual indexes using hierarchically formed virtual table format data

[0427] Aggregation is performed using virtual indices using a hierarchically formed virtual inverted structure. The i-th value of the aggregated result is given by the information 2 above, and the number of occurrences of the i-th value of the aggregated result is given by equation (8). If the aggregation is performed based on this, the following table 2 is shown.

[0428] [Table 2]

[0429] i value frequency Remark 0 <![CDATA[A0]]> 1 <![CDATA[From C 0(3) > 1 <![CDATA[A1]]> 1 <![CDATA[From C 1(2) > 2 <![CDATA[A2]]> 1 <![CDATA[From C 2(3) > 3 <![CDATA[B0]]> 2 <![CDATA[From C 0(3) > 4 <![CDATA[B2]]> 1 <![CDATA[From C 2(3) > 5 <![CDATA[C1]]> 1 <![CDATA[From C 1(2) > 6 <![CDATA[D2]]> 1 <![CDATA[From C 2(3) >

[0430] Here, even for identical values, if the D5A providing that value is different, separate summaries are generated. Therefore, users who prefer this approach remain unchanged, while users who prefer not to do so need to sum the number of occurrences of the same value. Even in this case, since the maximum number of occurrences does not exceed the number of D5As, processing time is minimal. Furthermore, similar to the D5A index, no new storage area is required to store the summation results.

[0431] Understanding Virtual Indexing via Spectral Decomposition

[0432] If we explain it from the perspective of spectral decomposition Figure 11 C in 0(3) , C 1(2) →C V0(5) and C V0(5) , C 2(3) →C V1(8) For ease of understanding, all the following formulas are recorded using allocation mapping with INV replaced by INV'.

[0433] C 0(3) and C1(2) It can be expressed using spectral decomposition as follows.

[0434] C 0(3) =A0:(3)+B0:(2,0)

[0435] C 1(2) =A1:(4)+C1(1)

[0436] Therefore, C V0(5) It can be expressed as follows.

[0437] C V0(5) =C 0(3) +C 1(2) =A0:(3)+A1:(4)+B0:(2,0)+C1(1)

[0438] In addition, C V0(5) and C 2(3) It is expressed as follows through spectral decomposition.

[0439] C V0(5) =A0:(1)+A1:(5)+B0:(7,2)+C1(3)

[0440] C 2(3) =A2: (4)+B2: (6)+D2: (0)

[0441] Therefore, C V1(8) It can be expressed as follows.

[0442] C V1(8) =C V0(5) +C 2(3) =A0: (1)+A1: (5)+A2: (4)+B0: (7, 2)+B2: (6)+C1(3)+D2:(0)

[0443] Therefore, the sorting results are listed as (1, 5, 4, 7, 2, 6, 3, 0). Furthermore, the search results for "B" through "C" are listed as (7, 2, 6, 3). This can be summarized in the same way as in Table 3 above. This shows that calculations based on spectral decomposition are easier to understand and superior to manual calculations.

[0444] <Overall Configuration Example of a System Including the Data Processing Device 10>

[0445] Figure 12 FIG. 1 shows an example of the overall configuration of a system including the data processing device 10 according to this embodiment. Figure 12 As shown, the data processing device 10 according to this embodiment is communicably connected to database servers distributed over a network 20 such as the Internet. In these database servers, at least one of one or more D5A or virtual table format data is stored.

[0446] The data processing device 10 of this embodiment includes a virtual table format data generating unit 101, a sorting unit 102, a search unit 103, a compilation unit 104, and a storage unit 105. It should be noted that the virtual table format data generating unit 101, the sorting unit 102, the search unit 103, and the compilation unit 104 are implemented by, for example, a processor such as a CPU (Central Processing Unit) executing one or more programs installed in the data processing device 10. The storage unit 105 is implemented by, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), a flash memory, or the like.

[0447] The virtual table format data generation unit 101 generates virtual table format data having the information 1 to 5 described in the "Virtual table format data configuration method" above. Furthermore, when defining the interval map 5 described in the "Virtual table format data configuration method" above, the virtual table format data generation unit 101 uses the method described in "Detecting Conflicts in Allocations from Different Intervals" to determine whether the allocation targets based on the interval map (allocation map) conflict.

[0448] The sorting unit 102 performs sorting using the virtual index of the virtual inverted structure of the virtual table format data. For example, the sorting unit 102 uses the virtual column of the virtual table format data to be sorted and the virtual INV contained in the virtual inverted structure of the virtual column to calculate the sorting result through index operation.

[0449] After receiving the search conditions, the search unit 103 performs a search using a virtual index of the virtual inverted structure of the virtual table format data. For example, the search unit 103 calculates equations (8) and (10) using a virtual column of the virtual table format data, the search conditions for the virtual column, and the virtual INV and virtual ACM contained in the virtual inverted structure of the virtual column. The search unit 103 then calculates the search result by performing an index operation on the virtual column and the result of the calculation of equations (8) and (10).

[0450] The aggregation unit 104 aggregates using the virtual index of the virtual inverted structure of the virtual table format data. For example, the aggregation unit 104 aggregates based on the information 2 and equation (8) using the virtual inverted structure of the virtual column of the aggregation target of the virtual table format data.

[0451] The storage unit 105 stores various data (for example, virtual table format data, D5A, sorting results, search results, summary results, etc.).

[0452] <Hardware Configuration Example of Data Processing Device 10>

[0453] Figure 13 FIG. 1 shows an example of a hardware configuration of the data processing device 10 according to this embodiment. Figure 13 As shown, the data processing device 10 according to this embodiment includes an input device 201, a display device 202, an external I / F 203, a communication I / F 204, a RAM (Random Access Memory) 205, a ROM (Read Only Memory) 206, an auxiliary storage device 207, and a processor 208. These hardware components are communicably connected via a bus 209.

[0454] The input device 201 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 202 is, for example, a display, a display panel, etc. In addition, the data processing device 10 may not include at least one of the input device 201 and the display device 202, for example.

[0455] The external I / F 203 is an interface with an external device such as a recording medium 203a. Examples of the recording medium 203a include a CD (Compact Disc), a DVD (Digital Versatile Disc), an SD memory card (Secure Digital Memory Card), a USB (Universal Serial Bus) memory card, and the like.

[0456] Communication I / F 204 is an interface for connecting data processing device 10 to network 20. RAM 205 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. ROM 206 is a nonvolatile semiconductor memory (storage device) that retains programs and data even when the power is turned off. Auxiliary storage device 207 is a nonvolatile storage device such as an HDD, SSD, or flash memory. Processor 208 is a CPU or other computing device.

[0457] It should be noted that Figure 13 The hardware configuration shown is merely an example, and the data processing device 10 may have other hardware configurations. For example, the data processing device 10 may include multiple auxiliary storage devices 207 or multiple processors 208, may not include some of the hardware shown, or may include various hardware other than the hardware shown.

[0458] <Flow of virtual table format data generation processing>

[0459] Below, refer to Figure 14 The flow of the virtual table format data generation process will be described.

[0460] Step S101: First, the virtual table format data generation unit 101 receives a specification of the number of records, virtual column names, and data types for the virtual table format data (the information in steps 1 and 2 described above in the "Virtual Table Format Data Configuration Method"). Note that the number of records, virtual column names, and data types for the virtual table format data are specified by the user, for example.

[0461] Step S102: Next, the virtual tabular format data generation unit 101 receives the designation of the URL or path of one or more source tabular format data (the information in step 3 described in the "Virtual Tabular Format Data Construction Method"). Note that, for example, the user may designate the URL or path of one or more source tabular format data. Alternatively, the user may designate the URL or path of one or more columns contained in one or more source tabular format data.

[0462] Step S103: Next, the virtual tabular format data generation unit 101 receives the definitions of the source columns and source intervals in each source tabular format data item (the information in step 4 described in the "Virtual Tabular Format Data Construction Method" above). Note that the definitions of the source columns and source intervals in each source tabular format data item are, for example, specified by the user.

[0463] Step S104: Next, the virtual table format data generation unit 101 receives the definition of the interval mapping for each source interval (the information in step 5 described in the "Virtual table format data construction method") and calculates the definition of the inverse mapping. It should be noted that the definition of the interval mapping for each source interval and its inverse mapping are specified by the user, for example. Furthermore, when the interval mapping is defined, its inverse mapping is automatically determined using equations (12) and (14) and stored in the storage unit 105.

[0464] Step S105 : Finally, the virtual table format data generating unit 101 stores the information received in the above-mentioned steps S101 to S104 in the storage unit 105 as virtual table format data.

[0465] <Flow of sorting processing in virtual table format data>

[0466] Below, refer to Figure 15 The flow of sorting processing in virtual table format data will be described.

[0467] Step S201 : The sorting unit 102 performs sorting using a virtual index of a virtual inverted structure of a virtual column to be sorted in a virtual column using virtual table format data.

[0468] Step S202: Then, the sorting unit 102 stores the sorting result in the above step S201 in the storage unit 105. However, the sorting unit 102 may not store the sorting result.

[0469] <Flow of search processing in virtual table format data>

[0470] Below, refer to Figure 16 The flow of search processing in virtual table format data will be described.

[0471] Step S301: First, the search unit 103 receives search conditions.

[0472] Step S302 : Next, the search unit 103 performs a search using a virtual index of a virtual inverted structure of a virtual column to be searched, among virtual columns of the virtual table format data.

[0473] Step S303 : Then, the search unit 103 stores the search result in the above step S302 in the storage unit 105 .

[0474] <Flow of summary processing in virtual table format data>

[0475] Below, refer to Figure 17 The flow of summary processing in virtual table format data will be described.

[0476] Step S401 : The aggregation unit 104 performs aggregation using a virtual index of a virtual inverted structure of a virtual column to be aggregated in a virtual column of virtual table format data.

[0477] Step S402: Then, the aggregation unit 104 stores the aggregation result in the above-mentioned step S303 in the storage unit 105. However, the aggregation unit 104 may not store the aggregation result.

[0478] Summary

[0479] As described above, in the data processing device 10 according to this embodiment, new virtual tabular data can be defined by using D5A or previously defined virtual tabular data as an allocation map for source tabular data. Furthermore, this virtual tabular data can be quickly sorted, retrieved, and aggregated using virtual indexes.

[0480] Therefore, for example, various virtual tabular data can be generated hierarchically starting from the D5A of the original tabular data, and the sorting, retrieval, and aggregation of this virtual tabular data can be performed at high speed. In addition, users can also generate and publish new virtual tabular data from the already published D5A and virtual tabular data according to their usage purposes.

[0481] Supplementary information

[0482] Next, as a supplement, the reason why the existing technology cannot inherit the index but the present embodiment can inherit the index is explained.

[0483] Indexing is achieved through a set of data structures for indexing and algorithms that use the data structures to accelerate processing. The inverted structure can be regarded as a data structure for indexing that is used to accelerate sorting, retrieval, and aggregation. Based on the inverted structure, a virtual inverted structure can be generated for the column obtained by rearranging the elements of the column. Moreover, a virtual inverted structure can be generated based on multiple inverted structures. In addition, further virtual inverted structures can be generated by layering multiple virtual inverted structures. In this way, virtual inverted structures can be generated hierarchically, or even if the elements of the column are rearranged, a virtual inverted structure can be generated. The algorithm for indexing using the inverted structure can also be applied to the virtual inverted structure, in which case the index is a virtual index. Existing indexes do not have a data structure for indexing that can be combined in multiple or hierarchical ways like an inverted structure, so inheritance is not possible when combining or layering.

[0484] The present invention is not limited to the specifically disclosed embodiments above, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.

[0485] This application is based on basic applications No. 2023-010716 and No. 2023-010717 filed in Japan on January 27, 2023, the entire contents of which are hereby incorporated by reference.

[0486] Description of Reference Numerals

[0487] 10 Data processing device

[0488] 20 Network

[0489] 101 Virtual table format data generation unit

[0490] 102 Sorting Department

[0491] 103 Search Department

[0492] 104 Summary Department

[0493] 105 Storage

[0494] 201 Input Device

[0495] 202 display device

[0496] 203 External I / F

[0497] 203a Recording medium

[0498] 204 Communication I / F

[0499] 205 RAM

[0500] 206 ROM

[0501] 207 Auxiliary storage device

[0502] 208 processor

[0503] 209 bus

Claims

1. A data processing device, comprising: A data operation unit performs a sorting operation, a search operation, or a summary operation on virtual tabular data composed of virtual columns having a second data structure obtained by converting a first data structure of one or more columns included in one or more tabular data using a predetermined mapping, using the second data structure as an index.

2. The data processing apparatus according to claim 1, wherein: The one or more tabular format data include virtual tabular format data composed of virtual columns having the second data structure.

3. The data processing apparatus according to claim 1, wherein: The one or more tabular data include D5A consisting of SVL, INV, and ACM, and the virtual tabular data, wherein the SVL is an arrangement that maintains the values ​​of the column in ascending order, the INV is an arrangement that maintains the posting record numbers of the values ​​of the column, and the ACM is an arrangement that maintains the posting record numbers of the values ​​of the column.

4. The data processing apparatus according to claim 3, wherein: The one or more tabular data further include the columns.

5. The data processing apparatus according to claim 3, wherein: The one or more tabular data further include NNC, which is an array of natural numbers obtained by replacing the values ​​of the column with positions on the SVL.

6. The data processing apparatus according to any one of claims 3 to 5, wherein: The data operation unit performs the sorting operation by sequentially acquiring values ​​of the sorting target column using INV included in the second data structure associated with the sorting target column of the virtual table format data.

7. The data processing apparatus according to any one of claims 3 to 5, wherein: The data operation unit performs the retrieval operation by using INV and ACM contained in the second data structure related to the column of the retrieval object of the virtual table format data, and the values ​​v0 and v1 given as retrieval conditions, based on the minimum i0 satisfying v0≤SVL[i0] and the maximum i1 satisfying SVL[i1]≤i1, to obtain (INV[ACM[i0-1]],...,INV[ACM[i1]-1]) as the retrieval result.

8. The data processing apparatus according to any one of claims 3 to 5, wherein: The data operation unit performs the aggregation operation by calculating the aggregation result of the i-th element of the SVL of the virtual table format data using ACM[i]-ACM[i-1] using the SVL and ACM included in the second data structure related to the aggregation target column of the virtual table format data.

9. A data processing method, which is executed by a computer: The data operation step is for performing a sorting operation, a search operation, or an aggregation operation on virtual table format data composed of virtual columns having a second data structure, using the second data structure as an index, wherein the second data structure is obtained by converting the first data structure of one or more columns included in one or more table format data using a prescribed mapping.

10. A program product causing a computer to execute: The data operation step is for performing a sorting operation, a search operation, or an aggregation operation on virtual table format data composed of virtual columns having a second data structure, using the second data structure as an index, wherein the second data structure is obtained by converting the first data structure of one or more columns included in one or more table format data using a prescribed mapping.

11. A data processing device comprising: an accepting unit that accepts designation of the number of records of virtual tabular format data to be generated, a virtual column representing a column of the virtual tabular format data, one or more tabular format data, and one or more columns included in the one or more tabular format data; and The allocation mapping definition unit generates the virtual tabular format data by allocating a second data structure obtained by converting a first data structure of the one or more columns included in the one or more tabular format data using a predetermined mapping to the virtual column.

12. The data processing apparatus according to claim 11, wherein: The method further includes a display unit configured to display the values ​​of the one or more columns on the virtual column of the virtual table format data.

13. The data processing apparatus according to claim 11 or 12, wherein: The mapping includes: a linear function type mapping, in which the correspondence between the first data structure and the second data structure is represented by a linear function; and an enumeration type mapping, in which the correspondence between the first data structure and the second data structure is represented by an enumeration of the correspondence.

14. The data processing apparatus according to claim 13, wherein: include: The conflict determination unit determines whether or not second data structures obtained by converting different first data structures using different mappings conflict with each other.

15. The data processing apparatus according to claim 14, wherein: The conflict determination unit Assume that any two of the different mappings are F0 and F1, If the value range of the mapping F0 and the value range of the mapping F1 are not on the same virtual column, it is determined that the second data structure of the conversion target of the mapping F0 does not conflict with the second data structure of the conversion target of the mapping F1. Assuming that the index of the starting position of the interval representing the value range of the image F0 is v0, the length of the interval representing the domain of the image F0 is w0, and the index of the starting position of the interval representing the value range of the image F1 is v1, if v0+w0-1<v1, then it is determined that the second data structure of the conversion target of the image F0 does not conflict with the second data structure of the conversion target of the mapping F1. Assuming that the length of the interval representing the domain of the mapping F1 is w1, if v1+w1-1<v0, it is determined that the second data structure of the conversion target of the mapping F0 does not conflict with the second data structure of the conversion target of the mapping F1.

16. The data processing apparatus according to claim 15, wherein: The conflict determination unit In the case where both the mapping F0 and the mapping F1 are enumeration type mappings, let the length of the interval representing the value range of the mapping F0 be w0, the length of the interval representing the value range of the mapping F1 be w1, L0 = (v0, ..., v0 + w0 - 1), L1 = (v1, ..., v1 + w1 - 1), and find the common ranges L0' and L1' of L0 and L1. When L0'[i] = L1'[j] is not satisfied for each i and j, it is determined that the second data structure of the conversion target of the mapping F0 does not conflict with the second data structure of the conversion target of the mapping F1. In the case where both the mapping F0 and the mapping F1 are linear function type mappings, the mapping F0 is re-described using y0=a0×x0+b0; x0=0, 1, 2, ..., N-1, where the common portion of the value range of the mapping F0 and the value range of the mapping F1 is used as the value range, and the mapping F1 is re-described using y1=a1×x1+b1; x0=0, 1, 2, ..., N-1, where the common portion is used as the value range. When x0 and x1 do not satisfy y0=y1 for x0≤MinInt(N0-1,LCM(a0,a1)div a0) and x1≤MinInt(N1-1,LCM(a0,a1)div a1), it is determined that the second data structure of the conversion target of the mapping F0 does not conflict with the second data structure of the conversion target of the mapping F1. When the mapping F0 is a linear function type mapping and the mapping F1 is an enumeration type mapping, the mapping F0 is re-described using y0(n) = a0×n+b0; x0 = 0, 1, 2, ..., N-1 with the common part of the value range of the mapping F0 and the value range of the mapping F1 as the value range, and the mapping F1 is re-described using y1(m) = (v1, ..., v1+w1-1)[m] with the common part as the value range. When y0(n) = y1(m) is not satisfied for each n and m, it is determined that the second data structure of the conversion target of the mapping F0 does not conflict with the second data structure of the conversion target of the mapping F1.

17. The data processing apparatus according to claim 16, wherein: The first data structure and the second data structure are data structures including at least INV indicating an array having posting numbers as elements.

18. A data processing method, wherein: Executed by a computer: an accepting step of accepting designations of the number of records of virtual tabular format data to be generated, virtual columns representing columns of the virtual tabular format data, one or more tabular format data, and one or more columns included in the one or more tabular format data; and The allocation mapping definition step generates the virtual tabular data by allocating a second data structure obtained by converting the first data structure of the one or more columns included in the one or more tabular data using a predetermined mapping to the virtual column.

19. A program product causing a computer to execute the following steps: an accepting step of accepting designations of the number of records of virtual tabular format data to be generated, virtual columns representing columns of the virtual tabular format data, one or more tabular format data, and one or more columns included in the one or more tabular format data; and The allocation mapping definition step generates the virtual tabular data by allocating a second data structure obtained by converting the first data structure of the one or more columns included in the one or more tabular data using a predetermined mapping to the virtual column.