Code conversion device, code conversion method, and program

The code conversion device and method optimize feature generation processing by detecting and replacing duplicate key columns with group names, thereby reducing calculation time in machine learning operations.

JP7708216B2Active Publication Date: 2025-07-15NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023570534
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-07-15
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

Existing methods for feature generation processing in machine learning, such as Target Encoding, suffer from increased calculation time due to duplicate processing when executing grouping operations with overlapping key columns in two-dimensional array data.

Method used

A code conversion device and method that detects duplicate key columns, generates group keys, and replaces them with duplicate group names to eliminate duplicate processing, thereby accelerating the grouping operation.

Benefits of technology

The solution significantly reduces the calculation time required for feature generation processing by eliminating duplicate operations in grouping operations using multiple key columns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708216000001
    Figure 0007708216000001
  • Figure 0007708216000002
    Figure 0007708216000002
  • Figure 0007708216000003
    Figure 0007708216000003
Patent Text Reader

Abstract

This code conversion device performs the following: first, combines a plurality of key columns which are included in two-dimensional array data, and detects a first function code for executing a grouping calculation for each combined key column; next, detects a key column name which represents the name of the key column from the first function code; next, generates a group key by using the key column name in each unit of two-dimensional array data for the first function code; next, detects a key column name which is duplicated in the plurality of group keys; generates a duplicate group comprising the duplicate key column names; adds a second function code for executing the grouping calculation and a third function code for adding a duplicate group name which represents the name of the duplicate group to the two-dimensional array data as a key column, by using the duplicate group name which represents the name of the duplicate group and the key column name which constitutes the duplicate group; and next, replaces the key column name used in the duplicate group and included in the first function code with the duplicate group name.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical field relates to a code conversion device for converting into codes, a code conversion method, and further, a program for realizing these. to the mud Related.

Background Art

[0002] Preprocessing for generating learning data used in machine learning includes feature generation processing. Also, it is known that the feature generation processing takes time.

[0003] Therefore, it is desired to shorten the time required for the feature generation processing. The reason why the feature generation processing takes time is that a plurality of columns included in two-dimensional array data are used as key columns, and a grouping operation is executed for each combination of the key columns. That is, when there are overlapping columns among the key columns, overlapping processing is executed.

[0004] As related art, Patent Documents 1 and 2 disclose a technique of changing combinations of a plurality of key columns included in two-dimensional array data and executing an operation using the key columns for each combination.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, the techniques of Patent Documents 1 and 2 do not convert the code of the grouping operation used in the feature generation processing and the like into a code for speeding up (shortening the operation time).

[0007] As one aspect, there is provided a code conversion device, a code conversion method, and a program for accelerating (shortening the calculation time) a grouping operation using a plurality of key columns included in two-dimensional array data. program The object is to achieve this. **Means for Solving the Problem**

[0008] To achieve the above object, a code conversion device in one aspect includes: a function detection unit that detects, from code for causing a computer to execute, which is stored in a storage device in advance, a first function code that combines a plurality of key columns included in two-dimensional array data and executes a grouping operation for each combined key column; for each of the detected first function codes, a group key generation unit that detects a key column name representing the name of the key column from the first function code and generates a group key using the key column name for each two-dimensional array data of the first function code; a duplicate group generation unit that detects key column names that are duplicated among the plurality of group keys and generates a duplicate group composed of the detected duplicated key column names; a function code addition unit that adds a second function code for executing the grouping operation and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to the two-dimensional array data, using the duplicate group name representing the name of the duplicate group and the key column names constituting the duplicate group; a key column replacement unit that replaces the key column names used in the duplicate group included in the first function code with the duplicate group name; and is characterized by having the above.

[0009] Also, to achieve the above object, a code conversion method in one aspect includes the steps of: a computer A function detection step of detecting a first function code that combines a plurality of key columns included in two-dimensional array data and executes a grouping operation for each combined key column from the code stored in a storage device in advance for causing a computer to execute; For each of the detected first function codes, a key column name representing the name of the key column is detected from the first function code, and a group key is generated using the key column name for each of the two-dimensional array data of the first function code, a group key generation step; A duplicate group generation step of detecting key column names that are duplicated among the plurality of group keys and generating a duplicate group composed of the detected duplicate key column names; Using the duplicate group name representing the name of the duplicate group and the key column names constituting the duplicate group, a second function code for executing the grouping operation and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to the two-dimensional array data are added, a function code addition step; A key column replacement step of replacing the key column names used in the duplicate group included in the first function code with the duplicate group name; characterized by having.

[0010] Furthermore, to achieve the above object, in one aspect, the program the mud is , On the computer, A function detection step of detecting a first function code that combines a plurality of key columns included in two-dimensional array data and executes a grouping operation for each combined key column from the code stored in a storage device in advance for causing a computer to execute; For each of the detected first function codes, a key column name representing the name of the key column is detected from the first function code, and a group key is generated using the key column name for each of the two-dimensional array data of the first function code, a group key generation step; A duplicate group generation step of detecting key column names that overlap among the plurality of group keys and generating a duplicate group composed of the detected overlapping key column names; A function code addition step of adding a second function code for executing the grouping operation using a duplicate group name representing the name of the duplicate group and the key column names constituting the duplicate group, and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to the two-dimensional array data; A key column replacement step of replacing the key column names used in the duplicate group included in the first function code with the duplicate group name; to be executed to do characterized by the above.

Advantages of the Invention

[0011] As one aspect, it is possible to speed up (shorten the calculation time) the grouping operation using a plurality of key columns included in the two-dimensional array data.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

DETAILED DESCRIPTION OF THE INVENTION

[0013] First, an overview will be described to facilitate understanding of the embodiments described hereinafter. In the preprocessing for generating learning data used in machine learning, there is feature quantity generation processing. Feature quantity generation processing includes, for example, Target Encoding (or Target Mean Encoding (Likelihood Encoding)) for quantifying (feature quantity conversion) categorical variables. Target Encoding is a process of aggregating the target variable for each categorical variable and quantifying it with the aggregated value (for example, average value, variance value, etc.).

[0014] FIG. 1 is a diagram for explaining Target Encoding. When using Table 1 as shown in FIG. 1 as the input for machine learning, the data in the "Category" column of Table 1 is not numerical, so it cannot be used directly as the input for machine learning.

[0015] Therefore, using Target Encoding, the data in the "Category" column of Table 1 as shown in FIG. 1 is converted into a numerical value obtained by aggregating the target variable, such as the data shown in the "Category Tgt-Mean" column of Table 3.

[0016] In that case, first, for the data in the "Category" column of Table 1, set each of the categorical variables A, B, C, and D to information that has no meaning in itself, such as integer values, like the data shown in the "Category ID" column of Table 2. In the example of Figure 1, 1 is set for categorical variable A, 2 is set for categorical variable B, 3 is set for categorical variable C, and 4 is set for categorical variable D.

[0017] Next, using the data shown in the "Category ID" column of Table 2, calculate the average value for each categorical variable, like the data shown in the "Category Tgt-Mean" column of Table 3. In the example of Figure 1, categorical variable A is quantified as 0.50 (= (1 + 0) / 2), categorical variable B is quantified as 0.33 (= (1 + 0 + 0) / 3), categorical variable C is quantified as 0.75 (= (1 + 0 + 1 + 1) / 4), and categorical variable D is quantified as 1.00 (= (1) / 1).

[0018] Next, using Figure 2, an example of target encoding with combinations of multiple categorical variables as well as a single categorical variable will be described. Figure 2 is a diagram for explaining target encoding when extended to multiple categorical variables.

[0019] In the example of Figure 2, target encoding is performed using 4 out of the categorical variables "CategoryA", "CategoryB", "CategoryC", "CategoryD", and "CategoryE" shown in Table 4. Note that in the example of Figure 2, the data in each column is omitted for reasons of convenience.

[0020] In the example of Figure 2, target encoding using the categorical variables "CategoryA", "CategoryB", "CategoryC", "CategoryD" and target encoding using the categorical variables "CategoryB", "CategoryC", "CategoryD", "CategoryE" are executed.

[0021] As a result, the categorical variables "CategoryABCD Tgt-Mean" and "CategoryBCDE Tgt-Mean" of Table 5 shown in Figure 2 are generated.

[0022] Target Encoding using a table processing library will be described. Figure 3 is a diagram for explaining the code of Target Encoding. The code shown in Figure 3 is an example of code using "groupby" and "transform" of pandas, which is a table processing library in the Python language.

[0023] Code 6 in Figure 3 is the code for Target Encoding using one categorical variable described in Figure 1. Code 7 in Figure 3 is the code for Target Encoding using multiple categorical variables described in Figure 2.

[0024] "groupby" used in Codes 6 and 7 is a function (or method) for grouping. "transform" is a function (or method) for rewriting data using the acquired statistical information (e.g., mean value, variance value, etc.).

[0025] "Category", "CatA", "CatB", "CatC", "CatD", and "CatE" described in Codes 6 and 7 represent the columns "Category", "CategoryA", "CategoryB", "CategoryC", "CategoryD", and "CategoryE" shown in Figures 1 and 2. "Target" represents "Target" shown in Figures 1 and 2. "Category_TgtMean", "ABCD_TgtMean", and "BCDE_TgtMean" represent "Category Tgt-Mean", "CategoryABCD Tgt-Mean", and "CategoryBCDE Tgt-Mean" shown in Figures 1 and 2.

[0026] The processes executed by Codes 6 and 7 include a process of generating groups and a process of calculating aggregated values for each group. In the case of Code 6, the process of generating groups generates the following groups GRP0, GRP1, GRP2, and GRP3 for each categorical variable.

[0027] Note that the numerical values representing the elements included in the following groups GRP0 to GRP3 are represented using the line numbers shown in FIG. 1.

[0028] GRP0: 0, 1 (Group of CategoryA) GRP1: 2, 3, 4 (Group of CategoryB) GRP2: 5, 6, 7, 8 (Group of CategoryC) GRP3: 9 (Group of CategoryD)

[0029] Furthermore, in the case of Code 6, by calculating the aggregated values for each group, the following average values for each group are calculated.

[0030] GRP0: Average value of 0, 1 (0.50) (A of Category Tgt-Mean) GRP1: Average value of 2, 3, 4 (0.33) (B of Category Tgt-Mean) GRP2: Average value of 5, 6, 7, 8 (0.75) (C of Category Tgt-Mean) GRP3: Average value of 9 (1.00) (D of Category Tgt-Mean)

[0031] However, when "groupby" using multiple columns (key columns) is executed multiple times by changing the combination of key columns, if there are overlapping columns between the key columns, duplicate processing (similar and wasteful processing) will be executed.

[0032] Specifically, when "groupby" is executed twice with two combinations of categorical variables, namely, the categorical variables "CategoryA", "CategoryB", "CategoryC", "CategoryD" and the categorical variables "CategoryB", "CategoryC", "CategoryD", "CategoryE" as shown in code 7, the categorical variables "CategoryB", "CategoryC", "CategoryD" overlap, resulting in duplicate processing (similar unnecessary processing).

[0033] Therefore, the calculation speed of the feature generation process becomes slower (the calculation time increases) only for the time spent on duplicate processing. Furthermore, the amount of calculation increases as the number of key columns increases.

[0034] Through such a process, the inventor found the problem of speeding up the calculation speed of the feature generation process (shortening the calculation time) in the method as described above, and derived means to solve such problems accordingly.

[0035] That is, the inventor derived means to convert the code used to perform the grouping operation using multiple key columns included in the two-dimensional array data (table) into code that can speed up the calculation speed (shorten the calculation time). As a result, the calculation speed of the feature generation process can be increased (the calculation time can be shortened).

[0036] Hereinafter, embodiments will be described with reference to the drawings. In the drawings described below, elements having the same function or corresponding functions are denoted by the same reference numerals, and repeated descriptions may be omitted.

[0037] (Embodiment) The configuration of the code conversion device 10 in the embodiment will be described with reference to FIG. 4. FIG. 4 is a diagram for explaining an example of the code conversion device.

[0038] [Device Configuration] The code conversion device 10 shown in FIG. 4 is a device that generates a code for accelerating the calculation speed (shortening the calculation time) based on a code used to execute a grouping operation using a plurality of key columns included in two-dimensional array data (table).

[0039] For example, when the code to be executed by a computer includes a description of a function that executes a grouping operation using a plurality of key columns included in two-dimensional array data, the code conversion device 10 generates a code for accelerating the calculation speed (shortening the calculation time).

[0040] The code conversion device 10 includes a function code detection unit 11, a group key generation unit 12, a duplicate group generation unit 13, a function code addition unit 14, and a key column replacement unit 15.

[0041] The function code detection unit 11 detects a first function code that combines a plurality of key columns included in two-dimensional array data (table) from the code to be executed by a computer and executes a grouping operation for each combined key column.

[0042] Specifically, the function code detection unit 11 detects, for example, "groupby" (the first function code) from the code to be executed by a computer.

[0043] FIG. 5 is a diagram for explaining an example of code conversion. The function code detection unit 11 detects, for example, code 7 shown in FIG. 5 (code 7 in FIG. 3) from the code to be executed by a computer.

[0044] For each detected first function code, the group key generation unit 12 detects a key column name representing the name of the key column from the first function code, and generates a group key (group key list) using the key column name for each two-dimensional array data (table) of the first function code.

[0045] For example, in the upper part of Code 7, the group key generation unit 12 detects the key column names "CatA", "CatB", "CatC", and "CatD" included in "groupby", and uses the detected key column names "CatA", "CatB", "CatC", and "CatD" to generate a group key (0:CatA, CatB, CatC, CatD).

[0046] Also, for example, in the lower part of Code 7, the group key generation unit 12 detects the key column names "CatB", "CatC", "CatD", and "CatE" from the table, and uses the detected key column names "CatB", "CatC", "CatD", and "CatE" to generate a group key (1:CatB, CatC, CatD, CatE).

[0047] The duplicate group generation unit 13 detects key column names that are duplicated among a plurality of group keys (group key list), and generates a duplicate group composed of the detected duplicate key column names.

[0048] For example, the duplicate group generation unit 13 uses a duplicate group name (t0) representing the name of the duplicate group and the duplicate key column names ("CatB", "CatC", "CatD") to generate a duplicate group (t0:CatB, CatC, CatD).

[0049] The function code addition unit 14 adds a second function code for executing a grouping operation using a duplicate group name representing the name of the duplicate group and the key column names constituting the duplicate group, and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to two-dimensional array data (table).

[0050] For example, the function code addition unit 14 adds Code 8 to the upper part of Code 7 as shown in FIG. 5. That is, a second function code (the upper code of Code 8) and a third function code (the lower code of Code 8) are added to the upper part of Code 7.

[0051] Note that the code below the code of 8 is a function (or method) not available in pandas, which is a table processing library in the Python language.

[0052] The key column replacement unit 15 replaces the key column name used in the duplicate group included in the first function code with the duplicate group name.

[0053] Specifically, the key column replacement unit 15 replaces the key column name used in the duplicate group included in "groupby" (the first function code) with the duplicate group name of the duplicate group composed of the duplicate key columns. The key column replacement unit 15 is converted as shown in, for example, code 9 in FIG. 5.

[0054] In this way, in the embodiment, the code used to execute the grouping operation using a plurality of key columns included in the two-dimensional array data (table) is converted into code that can speed up the operation speed (shorten the operation time). As a result, the operation speed of the feature quantity generation process can be increased (the operation time can be shortened).

[0055] [System Configuration] Using FIG. 6, the configuration of the code conversion device 10 in the embodiment will be described more specifically. FIG. 6 is a diagram for explaining an example of a system having the code conversion device 10. In the example of FIG. 6, the system 100 includes a code conversion device 10 and a storage device 20.

[0056] The code conversion device 10 is an information processing device such as a programmable device such as a CPU (Central Processing Unit), an FPGA (Field-Programmable Gate Array), a GPU (Graphics Processing Unit), or a circuit, a server computer, a personal computer, a mobile terminal, etc. equipped with any one or more of them.

[0057] The memory device 20 stores the code (pre-conversion code) that can be executed by a computer used to generate learning data. Also, the memory device 20 stores the code (post-conversion code) that can increase the calculation speed (shorten the calculation time).

[0058] The code conversion device will be specifically described. The specific process of code conversion will be described using the code that uses "groupby" and "transform" of pandas, which is a table processing library in the Python language. However, the language for writing the code is not limited to the Python language.

[0059] Figures 7, 8, 9, 10, and 11 are diagrams for explaining code conversion. The case where the function code detection unit 11 detects the code 71 shown in Figure 7 from the code to be executed by the computer will be described. In the following Figures 7 to 11, for the sake of easy understanding of the explanation, descriptions such as [Aggregation] "transform" for "group" are omitted.

[0060] Next, when the group key generation unit 12 detects the code 71 shown in Figure 7, it generates a group key list 72 composed of a plurality of group keys shown in Figure 7.

[0061] Next, the duplicate group generation unit 13 detects two key columns that overlap from the group key list 72 in Figure 7, and further obtains the number of the two overlapping key columns detected. As a result, as shown in 81 of Figure 8, the number of each of the two overlapping key columns is obtained.

[0062] Next, the duplicate group generation unit 13 compares the number of the two overlapping key columns and selects the two overlapping key columns with the maximum number. In the example of 81 in Figure 8, since the number of the two overlapping key columns "A" and "B" is the maximum value (6), the two overlapping key columns "A" and "B" are selected.

[0063] Next, the duplicate group generation unit 13 generates a duplicate group 82 (t0: ['A', 'B']) shown in FIG. 8 using the selected duplicate two-column key columns "A" and "B".

[0064] Furthermore, the duplicate group generation unit 13 replaces the duplicate two-column key columns "A" and "B" in the group key list 72 of FIG. 7 with "t0" representing the name of the duplicate group, and generates a new group key list 83 shown in FIG. 8.

[0065] Next, the duplicate group generation unit 13 detects the duplicate two-column key columns from the group key list 83 of FIG. 8, and further obtains the number of the detected duplicate two-column key columns. As a result, as shown in 91 of FIG. 9, the number of each of the duplicate two-column key columns is obtained.

[0066] Next, the duplicate group generation unit 13 compares the number of the duplicate two-column key columns, and selects the duplicate two-column key column with the maximum number. In the example of 91 in FIG. 9, since the number of each of the duplicate two-column key columns "t0", "D", "t0", "E", "t0", "F", and "E", "F" is 4, which is the maximum value, one of these duplicate two-column key columns, "t0", "E", is selected. However, as the duplicate two-column key column to be selected, any one of the duplicate two-column key columns "t0", "D", "t0", "F", and "E", "F" may be selected.

[0067] Next, the duplicate group generation unit 13 generates a duplicate group 92 shown in FIG. 9 (add t1: ['t0', 'E'] to the duplicate group 82) using the selected duplicate two-column key columns "t0" and "E".

[0068] Furthermore, the duplicate group generation unit 13 replaces the duplicate two-column key columns "t0" and "E" in the group key list 83 of FIG. 8 with "t1" representing the name of the duplicate group, and generates a new group key list 93 shown in FIG. 9.

[0069] Next, the duplicate group generation unit 13 detects two key columns that overlap from the group key list 93 in FIG. 9, and further obtains the number of the two overlapping key columns detected. As a result, as shown in 101 of FIG. 10, the number of each of the two overlapping key columns is obtained.

[0070] Next, the duplicate group generation unit 13 compares the number of the two overlapping key columns and selects the two overlapping key columns with the maximum number. In the example of 101 in FIG. 10, since there are only two overlapping key columns 't1' and 'F' (the maximum value of the two overlapping key columns is 4), the two overlapping key columns 't1' and 'F' will be selected.

[0071] Next, the duplicate group generation unit 13 uses the selected two overlapping key columns 't1' and 'F' to generate the duplicate group 102 shown in FIG. 10 (add t2:['t1','F'] to the duplicate group 92). Further, the duplicate group generation unit 13 replaces the two overlapping key columns 't1' and 'F' in the group key list 93 in FIG. 9 with 't2' representing the name of the duplicate group, Figure 10 and generates the new group key list 103 shown in.

[0072] In this way, the duplicate group generation unit 13 removes the duplication of the key column names in a plurality of group keys.

[0073] Next, after the duplicate group generation unit 13 generates the group key list 103 without overlapping key columns, the duplicate group 102 shown in FIG. 10 is converted into the duplicate group shown in FIG. 11 104 to.

[0074] That is, among 't0:['A','B']', 't1:['t0','E']', and 't2:['t1','F']' in the duplicate group 102 shown in FIG. 10, 't1:['t0','E']' not included in the group key list 103 is removed, expanded to 't2:['t1','F']', and 't2:['t0','E','F']' shown in FIG. 11 is obtained.

[0075] In the group key list 103 corresponding to "groupby" of the group key list 72 (original) from "0" to "6", "t0" and "t2" are used, but "t1" is not used. That is, "t1" was created for duplicate elimination and will not be a key for "groupby" in the group key list 72, so it is deleted. However, since it is used by "t2", it is deleted after expanding to "t2".

[0076] Next, the function code addition section 14 105 generates a second function code and a third function code based on the group key list, and adds the second function code and the third function code as shown in 106 of FIG. 11.

[0077] That is, add the second function code "grp_t0 = table.groupby(['A','B'])" corresponding to "t0" and the third function code "table['t0'] = grp_t0.getid()". Also, add the second function code "grp_t2 = table.groupby(['t0','E','F'])" corresponding to "t2" and the third function code "table['t2'] = grp_t2.getid()".

[0078] Next, the key column replacement section 15 replaces "A", "B", "A", "B", "E", "F" included in the first function code "groupby" of the code 71 shown in FIG. 7 with the duplicate group names "t0", "t2" based on the group key list 105, and obtains the code shown in 106 of FIG. 11 (replacement of the key column name of the first function code).

[0079] In this way, in the embodiment, the code used to execute the grouping operation using a plurality of key columns included in the two-dimensional array data (table) is converted into code that can accelerate the operation speed (shorten the operation time). As a result, the operation speed of the feature quantity generation process can be accelerated (the operation time can be shortened).

[0080] [Device Operation] Next, the operation of the code conversion device in the embodiment will be described with reference to FIG. 12. FIG. 12 is a diagram for explaining an example of the operation of the code conversion device. In the following description, the figures will be referred to as appropriate. In the embodiment, by operating the code conversion device, the code conversion method is implemented. Therefore, the description of the code conversion method in the embodiment will be replaced by the following description of the operation of the code conversion device.

[0081] As shown in FIG. 12, first, the function code detection unit 11 acquires the code for causing the computer to execute, which is stored in the storage device in advance (step A1). Next, the function code detection unit 11 combines a plurality of key columns included in the two-dimensional array data (table) from the acquired code, and detects the first function code for executing the grouping operation for each combined key column (step A2).

[0082] Next, the group key generation unit 12 detects the key column name representing the name of the key column from the detected first function code for each first function code, and generates a group key (group key list) using the key column name for each two-dimensional array data (table) of the first function code (step A3).

[0083] Next, the duplicate group generation unit 13 detects the key column names that are duplicated in the plurality of group keys (group key list), and generates a duplicate group composed of the detected duplicate key column names (step A4).

[0084] Specifically, in step A4, the duplicate group generation unit 13 first combines two key column names included in the group key for each group key, sets the combination with the most combinations as the duplicate group, and further replaces the two key column names used in the duplicate group included in the group key with the duplicate group name. Next, the duplicate group generation unit 13 removes the duplication of the key column names in the plurality of group keys (group key list).

[0085] Next, the function code addition unit 14 adds a second function code for executing a grouping operation using a duplicate group name representing the name of the duplicate group and a key column name constituting the duplicate group, and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to two-dimensional array data (table) (step A5).

[0086] The key column replacement unit 15 replaces the key column name used in the duplicate group included in the first function code with the duplicate group name (step A6).

[0087] [Advantages of the Embodiment] According to the embodiment as described above, the code used for executing a grouping operation using a plurality of key columns included in two-dimensional array data (table) can be converted into a code capable of accelerating the operation speed (shortening the operation time). As a result, the operation speed of the feature quantity generation process can be accelerated (the operation time can be shortened).

[0088] [Program] The program in the embodiment may be a program that causes a computer to execute steps A1 to A6 shown in FIG. 12. By installing and executing this program on a computer, the code conversion device and the code conversion method in the embodiment can be realized. In this case, the processor of the computer functions as the function code detection unit 11, the group key generation unit 12, the duplicate group generation unit 13, the function code addition unit 14, and the key column replacement unit 15, and performs processing.

[0089] Also, the program in the embodiment may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the function code detection unit 11, the group key generation unit 12, the duplicate group generation unit 13, the function code addition unit 14, and the key column replacement unit 15.

[0090] [Physical Configuration] Here, a computer that realizes a code conversion device by executing a program in an embodiment will be described with reference to FIG. 13. FIG. 13 is a diagram for explaining an example of a computer that realizes a code conversion device in an embodiment.

[0091] As shown in FIG. 13, the computer 110 includes a CPU (Central Processing Unit) 111, a main memory 112, a storage device 113, an input interface 114, a display controller 115, a data reader / writer 116, and a communication interface 117. These components are connected to be able to communicate with each other via a bus 121. Note that in addition to or instead of the CPU 111, the computer 110 may include a GPU or an FPGA.

[0092] The CPU 111 expands a program (code) in an embodiment stored in the storage device 113 into the main memory 112 and executes these in a predetermined order to perform various operations. The main memory 112 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory). Also, the program in the embodiment is provided in a state stored in a computer-readable recording medium 120. Note that the program in the embodiment may be distributed on the Internet connected via the communication interface 117. Note that the recording medium 120 is a non-volatile recording medium.

[0093] In addition, specific examples of the storage device 113 include a hard disk drive and a semiconductor storage device such as a flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and a mouse. The display controller 115 is connected to a display device 119 and controls the display on the display device 119.

[0094] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, and performs reading of programs from the recording medium 120 and writing of processing results in the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.

[0095] Specific examples of the recording medium 120 include general-purpose semiconductor memory devices such as CF (Compact Flash (registered trademark)) and SD (Secure Digital), magnetic recording media such as Flexible Disks, or optical recording media such as CD-ROM (Compact Disk Read Only Memory).

[0096] Note that the code conversion device 10 in the embodiment can also be realized by using hardware corresponding to each part, rather than a computer in which a program is installed. Furthermore, part of the code conversion device 10 may be realized by a program and the remaining part may be realized by hardware.

[0097] [Appendix] Regarding the above embodiments, the following appendix is further disclosed. Some or all of the above-described embodiments can be expressed by (Appendix 1) to (Appendix 9) described below, but are not limited to the following description.

[0098] (Appendix 1) A function detection unit that combines a plurality of key columns included in two-dimensional array data from code for causing a computer pre-stored in a storage device to execute, and detects first function code for executing a grouping operation for each combined key column; For each of the detected first function codes, a key column name representing the name of the key column is detected from the first function code, and a group key is generated using the key column name for each two-dimensional array data of the first function code. A group key generation unit; Detect key column names that are duplicated among the plurality of group keys, and generate a duplicate group composed of the detected duplicated key column names; a duplicate group generation unit A second function code for causing the grouping operation to be executed using a duplicate group name representing the name of the duplicate group and the key column names constituting the duplicate group, and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to the two-dimensional array data; a function code addition unit A key column replacement unit that replaces the key column names used in the duplicate group included in the first function code with the duplicate group name A code conversion device having the above

[0099] (Appendix 2) The code conversion device according to Appendix 1, wherein For each of the group keys, the duplicate group generation unit combines two of the key column names included in the group key, sets the combination with the most occurrences as the duplicate group, and further replaces the two key column names used in the duplicate group included in the group key with the duplicate group name A code conversion device

[0100] (Appendix 3) The code conversion device according to Appendix 2, wherein The duplicate group generation unit removes the duplication of the key column names among the plurality of group keys A code conversion device

[0101] (Appendix 4) A computer A function detection step of detecting a first function code for combining a plurality of key columns included in two-dimensional array data and executing a grouping operation for each combined key column from code for causing a computer to execute, which is stored in a storage device in advance For each of the detected first function codes, detect a key column name representing the name of the key column from the first function code, and for each of the two-dimensional array data of the first function code, generate a group key using the key column name, a group key generation step; Detect key column names that are duplicated among the plurality of group keys, and generate a duplicate group composed of the detected duplicated key column names, a duplicate group generation step; Using the duplicate group name representing the name of the duplicate group and the key column names constituting the duplicate group, add a second function code for executing the grouping operation and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to the two-dimensional array data, a function code addition step; Replace the key column names used in the duplicate group included in the first function code with the duplicate group name, a key column replacement step; A code conversion method for executing the above.

[0102] (Appendix 5) A code conversion method according to Appendix 4, In the duplicate group generation step, for each group key, combine the two key column names included in the group key, and set the combination with the most occurrences as the duplicate group. Further, replace the two key column names used in the duplicate group included in the group key with the duplicate group name; A code conversion method.

[0103] (Appendix 6) A code conversion method according to Appendix 5, In the duplicate group generation step, remove the duplication of the key column names among the plurality of group keys A code conversion method.

[0104] (Appendix 7) On a computer, A function detection step of combining a plurality of key sequences included in two-dimensional array data from code stored in a storage device in advance for causing a computer to execute, and detecting first function code for executing a grouping operation for each of the combined key sequences; For each of the detected first function codes, detecting a key sequence name representing the name of the key sequence from the first function code, and generating a group key using the key sequence name for each of the two-dimensional array data of the first function code; a group key generation step; A duplicate group generation step of detecting key sequence names that are duplicated among a plurality of the group keys, and generating a duplicate group composed of the detected duplicate key sequence names; Using the duplicate group name representing the name of the duplicate group and the key sequence name constituting the duplicate group, adding a second function code for executing the grouping operation and a third function code for adding the duplicate group name representing the name of the duplicate group as a key sequence to the two-dimensional array data; a function code addition step; A key sequence replacement step of replacing the key sequence name used in the duplicate group included in the first function code with the duplicate group name; A program including an instruction to execute mud

[0105] (Appendix 8) As described in Appendix 7 program where In the duplicate group generation step, for each of the group keys, combining two of the key sequence names included in the group key, setting the combination with the most occurrences as the duplicate group, and further replacing the two key sequence names used in the duplicate group included in the group key with the duplicate group name; program 。

[0106] (Appendix 9) As described in Appendix 8 program where The duplicate group generation step removes the duplication of the key column names among the plurality of the group keys. program 。

[0107] Although the invention has been described with reference to the embodiments above, the invention is not limited to the above-described embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the invention within the scope of the invention.

Industrial Applicability

[0108] According to the above description, it is possible to speed up (shorten the calculation time) the grouping operation using a plurality of key columns included in two-dimensional array data (table). Further, it is useful in fields where a grouping operation using a plurality of key columns included in two-dimensional array data (table) is required.

Explanation of Signs

[0109] 10 Code conversion device 11 Function code detection unit 12 Group key generation unit 13 Duplicate group generation unit 14 Function code addition unit 15 Key column replacement unit 20 Storage device 100 System 110 Computer 111 CPU 112 Main memory 113 Storage device 114 Input interface 115 Display controller 116 Data reader / writer 117 Communication interface 118 Input device 119 Display device 120 Recording medium 121 Bus

Claims

1. Function detection means for detecting first function code that combines a plurality of key sequences included in two-dimensional array data from code for causing a computer to execute, which is stored in a storage device in advance, and executes a grouping operation for each of the combined key sequences; Group key generation means for detecting, for each of the detected first function codes, a key sequence name representing the name of the key sequence from the first function code, and generating a group key using the key sequence name for each of the two-dimensional array data of the first function code; Duplicate group generation means for detecting key sequence names that are duplicated among a plurality of the group keys and generating a duplicate group composed of the detected duplicated key sequence names; Function code addition means for adding a second function code for executing the grouping operation using the duplicate group name representing the name of the duplicate group and the key sequence name constituting the duplicate group, and a third function code for adding the duplicate group name representing the name of the duplicate group as a key sequence to the two-dimensional array data; Key sequence replacement means for replacing the key sequence name used in the duplicate group, which is included in the first function code, with the duplicate group name; A code conversion device having the above.

2. The code conversion device according to claim 1, wherein the duplicate group generation means combines, for each of the group keys, two of the key sequence names included in the group key, sets the combination having the most combinations as the duplicate group, and further replaces the two key sequence names used in the duplicate group included in the group key with the duplicate group name. Code conversion device.

3. The code conversion device according to claim 2, wherein the duplicate group generation means removes duplication of the key sequence names among a plurality of the group keys. Code conversion device.

4. A computer detects first function code that combines a plurality of key sequences included in two-dimensional array data from code for causing a computer to execute, which is stored in a storage device in advance, and executes a grouping operation for each of the combined key sequences; for each of the detected first function codes, detects a key sequence name representing the name of the key sequence from the first function code, and generates a group key using the key sequence name for each of the two-dimensional array data of the first function code; Detect key column names that are duplicated among the plurality of group keys, and generate a duplicate group composed of the detected duplicated key column names. Add a second function code for executing the grouping operation using the duplicate group name representing the name of the duplicate group and the key column names constituting the duplicate group, and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to the two-dimensional array data. Replace the key column names used in the duplicate group included in the first function code with the duplicate group name. Code conversion method.

5. The code conversion method according to claim 4, In the generation of the duplicate group, for each group key, combine two key column names included in the group key, and set the combination with the most occurrences as the duplicate group. Further, replace the two key column names used in the duplicate group included in the group key with the duplicate group name. Code conversion method.

6. The code conversion method according to claim 5, In the generation of the duplicate group, remove the duplication of the key column names among the plurality of group keys. Code conversion method.

7. On a computer, Detect a first function code for combining a plurality of key columns included in two-dimensional array data and executing a grouping operation for each combined key column from the code stored in a storage device in advance for execution by the computer. For each detected first function code, detect a key column name representing the name of the key column from the first function code, and generate a group key using the key column name for each two-dimensional array data of the first function code. Detect key column names that are duplicated among the plurality of group keys, and generate a duplicate group composed of the detected duplicated key column names. Add a second function code for executing the grouping operation using the duplicate group name representing the name of the duplicate group and the key column names constituting the duplicate group, and a third function code for adding the duplicate group name representing the name of the duplicate group as a key column to the two-dimensional array data. Cause the key column names used in the duplicate group included in the first function code to be replaced with the duplicate group name. A program including instructions.

8. The program according to claim 7, in the generation of the duplicate group, for each of the group keys, combine the two key column names included in the group key, and set the combination with the most occurrences as the duplicate group. Further, replace the two key column names used in the duplicate group included in the group key with the duplicate group name. Program.

9. The program according to claim 8, in the generation of the duplicate group, remove the duplication of the key column names in a plurality of the group keys Program.

Citation Information

Patent Citations

  • Efficient large-scale filtering and / or sorting for queries on column-based data-encoded structures

    JP2012504825A

  • Analysis method, analyzer and analysis program

    JP2014228974A

  • Evaluation of rollups with distinct aggregates by using sequence of sorts and partitioning by measures

    US6775682B1

  • Data stream processing parallelization program, and data stream processing parallelization system

    WO2014188500A1

  • Elimination of query fragment duplication in complex database queries

    WO2020131243A2