Data model optimization system, data model optimization method, and data model optimization program
The data model optimization system addresses the limitation of existing techniques by calculating call counts and similarity to generate a data model optimized for application processing, enhancing communication efficiency and suitability for database table operations.
Patent Information
- Application Number
- JP2024574614
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-06-15
AI Technical Summary
Existing techniques, such as those described in Patent Document 1, do not provide a method for regenerating a conversion expression based on the evaluation of a conversion expression in JSON format, which limits the ability to generate a data model suitable for applications considering the synthesis and division of database tables.
The data model optimization system calculates call counts for individual columns and column sets, similarity between column names, and generates a data model based on these metrics to optimize data processing for applications, ensuring the data model is suitable for the application's data acquisition scenario.
This approach enables the generation of a data model that is optimized for application processing, improving communication efficiency by considering the merging and division of database tables, and ensuring data is processed in a structure suitable for the application.
Smart Images

Figure 0007682411000001 
Figure 0007682411000002 
Figure 0007682411000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for generating a data model for an application.
Background Art
[0002] Patent Document 1 proposes a technique for generating new JSON-formatted data using data described in JSON format and conversion expressions described in JSON format.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Patent Document 1 does not propose a technique for regenerating a conversion expression based on the evaluation of a conversion expression described in JSON format. Therefore, it is not possible to generate a data model suitable for an application in consideration of the synthesis and division of database tables.
[0005] An object of the present disclosure is to be able to generate a data model suitable for an application.
Means for Solving the Problems
[0006] The data model optimization system of the present disclosure Based on database configuration information indicating the configuration of a plurality of tables in a database and a data acquisition scenario in which data acquired from the database is specified, for each column in the plurality of tables, when data is acquired from the database according to the data acquisition scenario, a single call count calculation unit that calculates the number of times the column is called a set call count calculation unit that calculates, for each column set in the plurality of tables, a call count at which the column set is called at the same timing when data is acquired from the database according to the data acquisition scenario, based on the database configuration information and the data acquisition scenario; a similarity calculation unit that calculates a similarity between names of columns for each column group in the plurality of tables based on the database configuration information; a data model generation unit that generates a data model for expressing data acquired according to the data acquisition scenario in a structure suitable for processing by an application that uses the data, based on the number of calls for each column, the number of calls for each column set, and the similarity for each column set; Equipped with. Effect of the Invention
[0007] According to the present disclosure, it is possible to generate a data model suitable for an application. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a configuration diagram of a data model optimization system 100 according to a first embodiment. [Diagram 2] FIG. 1 is a configuration diagram of a data processing system 200 according to a first embodiment. [Diagram 3] FIG. 2 is a functional configuration diagram of a data processing system 200 according to the first embodiment. [Figure 4] 3 is a flowchart of a data model optimization method according to the first embodiment. [Diagram 5] 3 is a flowchart of a data model optimization method according to the first embodiment. [Figure 6] FIG. 2 shows an example of the configuration of a database 212 according to the first embodiment. [Figure 7] FIG. 13 shows an example of a data acquisition scenario D02 according to the first embodiment. [Figure 8] FIG. 2 is a diagram showing an example of a first edition model D20 according to the first embodiment. [Figure 9] FIG. 2 is a diagram showing an example of a data conversion image according to the first embodiment. [Figure 10] FIG. 2 shows an example of a data model D31 according to the first embodiment. [Figure 11] FIG. 2 shows an example of a data model D31 according to the first embodiment. [Figure 12] FIG. 2 shows an example of a data model D31 according to the first embodiment. [Figure 13] FIG. 2 shows an example of a functional configuration of a data processing system 200 according to the first embodiment. [Figure 14] FIG. 11 is a functional configuration diagram of a data processing system 200 according to a second embodiment. [Figure 15] 11 is a flowchart of a data model optimization method according to the second embodiment. [Figure 16] 11 is a flowchart of a data model optimization method according to the second embodiment. [Figure 17] 11 is a flowchart of a data model optimization method according to the second embodiment. [Figure 18] FIG. 1 is a hardware configuration diagram of a data model optimization system 100 according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] In the embodiments and drawings, the same or corresponding elements are denoted by the same reference numerals. Descriptions of elements denoted by the same reference numerals as those of the elements described above will be omitted or simplified as appropriate. Arrows in the drawings mainly indicate data flow or processing flow.
[0010] Embodiment 1 The data model optimization system 100 will be described with reference to FIGS.
[0011] ***Configuration Description*** The configuration of a data model optimization system 100 will be described with reference to FIG. The data model optimization system 100 is a computer that includes hardware such as a processor 101, a memory 102, an auxiliary storage device 103, a communication device 104, and an input / output interface 105. These hardware components are connected to each other via signal lines. The data model optimization system 100 may be configured with multiple computers instead of one computer (device).
[0012] The processor 101 is an IC that performs arithmetic processing and controls other hardware. For example, the processor 101 is a CPU. IC is an abbreviation for Integrated Circuit. CPU is an abbreviation for Central Processing Unit.
[0013] The memory 102 is a volatile or non-volatile storage device. The memory 102 is also called a primary storage device or a main memory. For example, the memory 102 is a RAM. Data stored in the memory 102 is saved in the secondary storage device 103 as necessary. RAM is an abbreviation for Random Access Memory.
[0014] The auxiliary storage device 103 is a non-volatile storage device. For example, the auxiliary storage device 103 is a ROM, a HDD, a flash memory, or a combination of these. Data stored in the auxiliary storage device 103 is loaded into the memory 102 as needed. ROM is an abbreviation for Read Only Memory. HDD is an abbreviation for Hard Disk Drive.
[0015] The communication device 104 is a receiver and a transmitter. For example, the communication device 104 is a communication chip or a NIC. The communication of the data model optimization system 100 is performed using the communication device 104. NIC is an abbreviation for Network Interface Card.
[0016] The input / output interface 105 is a port to which an input device and an output device are connected. For example, the input / output interface 105 is a USB terminal, the input device is a keyboard and a mouse, and the output device is a display. Input and output of the data model optimization system 100 is performed using the input / output interface 105. USB is an abbreviation for Universal Serial Bus.
[0017] The data model optimization system 100 includes components such as a calculation unit 110, an evaluation unit 120, and a generation unit 130. These components are realized in software. The calculation unit 110 includes elements such as a single call count calculation unit 111 , a pair call count calculation unit 112 , and a similarity calculation unit 113 . The evaluator 120 comprises elements such as a data model evaluator 121 . The generation unit 130 includes elements such as a data model comparison unit 131 and a data model generation unit 132 .
[0018] The auxiliary storage device 103 stores a data model optimization program for causing a computer to function as the calculation unit 110, the evaluation unit 120, and the generation unit 130. The data model optimization program is loaded into the memory 102 and executed by the processor 101. The auxiliary storage device 103 further stores an OS. At least a part of the OS is loaded into the memory 102 and executed by the processor 101. The processor 101 executes the data model optimization program while executing the OS. OS is an abbreviation for Operating System.
[0019] The input and output data of the data model optimization program are stored in the storage unit 190 . The memory 102 functions as the storage unit 190. However, a storage device such as the auxiliary storage device 103, a register in the processor 101, or a cache memory in the processor 101 may function as the storage unit 190 instead of the memory 102 or together with the memory 102.
[0020] The data model optimization program can be recorded (stored) in a computer-readable manner on a non-volatile recording medium such as an optical disk or a flash memory.
[0021] The configuration of the data processing system 200 will be described with reference to FIG. The data processing system 200 is a computer system that utilizes the data model optimization system 100 . The data processing system 200 includes a data platform 210 and an application unit 221 . The data platform 210 is a computer system, and includes the data model optimization system 100 , a data conversion unit 211 , and a database 212 . The data conversion unit 211 is an element that executes data conversion software, and is realized by a processing circuit (for example, a processor) of a computer. The data conversion software causes the computer to function as the data conversion unit 211. The application unit 221 is an element that executes an application program for data processing, and is realized by a processing circuit of a computer. The application program causes the computer to function as the application unit 221.
[0022] The data model optimization system 100 is introduced into the communication system between the application unit 221 and the data platform 210 . The data model optimization system 100 generates a data model D1 and transmits the data model D1 to the data conversion unit 211. The data conversion unit 211 receives the data model D1. The application unit 221 transmits the data request D2 to the data platform 210. The data conversion unit 211 receives the data request D2. The data conversion unit 211 obtains the data D4 requested in the data request D2 from the database 212 by making an inquiry D3 to the database 212. The data conversion unit 211 converts the data D4 into data D5 based on the data model D1, and transmits the data D5 to the application unit 221. The application unit 221 receives the data D5 and performs data processing using the data D5. The data model D1 represents only the data necessary for the application from the vast amount of source data (DB) in a structure (format and grouping) suitable for application processing. For example, data model D1 converts data D4 into data D5 as follows. Data model D1 represents a "person" whose internal structure includes name, age, and gender. Then, separate data D4, "Suzuki," "26," and "Female," are obtained from database 212. In this case, data model D1 stores "Suzuki" in the "Name" field of the internal structure, stores "26" in the "Age" field of the internal structure, and stores "Female" in the "Gender" field of the internal structure. As a result, data D4 is converted into data D5 that represents "a person named Suzuki." The application uses this group of data D5.
[0023] FIG. 3 shows the functional configuration of the data model optimization system 100. The functions of each element of the data model optimization system 100 and the data input and output between the elements will be described later.
[0024] ***Explanation of Operation*** The operation procedure of the data model optimization system 100 corresponds to a data model optimization method, and also corresponds to a processing procedure by a data model optimization program.
[0025] Based on FIGS. 4 and 5, a data model optimization method will be described. In step S110, the calculation unit 110 calculates the number of calls per column, the number of calls per column set, and the similarity per column set based on the database configuration information D01 and the data acquisition scenario D02.
[0026] The database configuration information D01 is data indicating the configurations of a plurality of tables in the database 212. The database configuration information D01 is acquired from the database 212, for example.
[0027] FIG. 6 shows an example of the configurations of a plurality of tables in the database 212. The database 212 has a first table, a second table, and a third table. Time-series data is registered in each table. The first table has columns such as "ID", "person's name", "person's age", "person's location", and "observation time". The second table has columns such as "ID", "person's name", "person's gender", "person's heart rate", and "observation time". The third table has columns such as "ID", "robot model number", "robot location", "remaining battery level of the robot", and "observation time". The database configuration information D01 indicates such a configuration of the database 212.
[0028] Returning to FIG. 4, the description of step S110 will be continued. The data acquisition scenario D02 is data in which the data to be acquired from the database 212 is specified. The data acquisition scenario D02 is acquired from the application unit 221, for example.
[0029] FIG. 7 shows an example of the data acquisition scenario D02. The first scenario indicates that the data of each column specified in the usage field column is acquired in the order specified in the timing column. Data in two or more columns specified in the usage field column at the same timing is acquired at the same timing. "The same timing" may be read as "simultaneously". The second scenario shows that data in a plurality of columns specified in the usage field column is acquired in a batch.
[0030] Returning to FIG. 4, the details of step S110 will be described. The single call count calculation unit 111 calculates the call count for each column in a plurality of tables shown in the database configuration information D01 based on the data acquisition scenario D02. The calculated call count is the number of times a column is called when data is acquired from the database 212 according to the data acquisition scenario D02.
[0031] Based on the first scenario in FIG. 7, an example of calculating the call count for each column will be described. In the first scenario, the column "person's age" is specified in the usage field column at the first timing and the eleventh timing, respectively. Therefore, when the column "person's age" is not specified in the usage field column after the fifteenth timing, the call count of the column "person's age" is 2. In the case of the second scenario, the call count of each column specified in the usage field column is 1, and the call count of the other columns is 0.
[0032] Returning to FIG. 4, the description of step S110 will be continued. The group call count calculation unit 112 calculates the call count for each column group in a plurality of tables shown in the database configuration information D01 based on the data acquisition scenario D02. The calculated call count is the number of times a column group is called at the same timing when data is acquired from the database 212 according to the data acquisition scenario D02. A column group consists of two or more columns. For example, each of all combinations in all columns of all tables in the database 212 is a column group.
[0033] An example of calculating the number of calls for each column group will be described based on the first scenario in FIG. In the first scenario, the pair of columns "person's age" and "person's heart rate" is specified in the used fields section of each of the first timing and the eleventh timing. Therefore, if the pair of columns "person's age" and "person's heart rate" is not specified in the used fields column from the 15th timing onwards, the number of calls for the pair of columns "person's age" and "person's heart rate" is 2. In the case of the second scenario, the number of calls for each column group specified in the field used field column is 1, and the number of calls for the other column groups is 0.
[0034] Returning to FIG. 4, the description of step S110 will continue. The similarity calculation unit 113 calculates the similarity for each column group in the multiple tables indicated in the database configuration information D01. The calculated similarity is the similarity between the names of the columns included in the column group.
[0035] An example of calculation of the similarity for each column set will be described with reference to FIG. If the character strings of the names of the columns are exactly the same, the similarity of the column pair is calculated by multiplying the reference value by the number of columns. The number of columns means the number of columns included in the column pair. The column "Person's Name" in the first table and the column "Person's Name" in the second table have the string "Person's Name" in each name that is a perfect match. Therefore, if the reference value is 10, the value "20" calculated by multiplying the reference value "10" by the number of columns "2" is the similarity between the pair of column "person's name" in the first table and column "person's name" in the second table. If the character strings in the names of the columns do not match completely, the similarity between the column pairs is calculated by multiplying the number of common words by the number of columns. The number of common words means the number of words that are common in the character strings in the names of the columns. The column "person position" in the first table and the column "robot position" in the third table have one word in common in their name strings: "position." Therefore, the value "2" calculated by multiplying the number of common words "1" by the number of columns "2" is the similarity between the pair of column "person position" in the first table and column "robot position" in the third table.
[0036] Returning to FIG. 4, the description of step S110 will continue. The single call count calculation unit 111 stores the call count for each column in the storage unit 190. The pair call count calculation unit 112 stores the call count for each column pair in the storage unit 190. The similarity calculation unit 113 stores the similarity for each column pair in the storage unit 190. The data indicating the number of calls for each column, the number of calls for each column set, and the similarity for each column set is referred to as calculation information D11.
[0037] In step S120, the data model evaluation unit 121 evaluates the data model D21 based on the data acquisition scenario D02 to calculate an evaluation value of the data model D21. The data model D21 is data that indicates rules for expressing data acquired according to the data acquisition scenario D02 in a structure suitable for processing by an application that uses the data.
[0038] In the first step S120, the first edition model D20 is evaluated. The first version model D20 is a data model D21 that is generated in advance. For example, the first version model D20 is input to the data model optimization system 100, and the data model evaluation unit 121 receives the input first version model D20.
[0039] FIG. 8 shows an example of the first edition model D20. The first edition model D20 indicates rules for expressing data acquired from each table in the database 212 in a structure suitable for processing by an application that uses the data. Data model x is the original model D20 for the first table. The data model y is the original model D20 for the second table.
[0040] FIG. 9 shows an example of data whose structure has been converted according to the first edition model D20. The data transformation image x represents data whose structure has been transformed according to the data model x. The data transformation image y represents data whose structure has been transformed in accordance with the data model y.
[0041] Returning to FIG. 4, the description of step S120 will continue. In step S120 from the second time onwards, the data model D31 generated in step S140 is evaluated as the data model D21. The data model D31 is stored in the storage unit 190 and is read out from the storage unit 190.
[0042] The evaluation value (score) of the data model D21 is a value obtained by evaluating the data model D21 based on the evaluation axis. The evaluation axis means a criterion, rule, condition, or the like for evaluation. An example of an evaluation axis is the number of queries to the database 212. The number of queries corresponds to the number of accesses to a table. The evaluation axis may be the data communication volume or the number of data models. The evaluation axis may be a combination of factors such as the number of inquiries, the data communication volume, and the number of data models. The evaluation axis may be related to communication performance. In addition, other evaluation axes may be used.
[0043] In the first embodiment, the smaller the evaluation value, the higher the evaluation of the data model D21, and the larger the evaluation value, the lower the evaluation of the data model D21.
[0044] The evaluation value of the data model is calculated as follows. The data model evaluation unit 121 simulates the behavior of the data processing system 200 (particularly, at least one of the data conversion unit 211 and the application unit 221) when the data model to be evaluated is used. Then, the data model evaluation unit 121 calculates an evaluation value of the data model based on the result of the simulation.
[0045] An example of calculation of the evaluation value will be described. In this example, the evaluation axis is the number of queries to the database 212. The database 212 has the table shown in Fig. 6. Furthermore, the data acquisition scenario D02 is the first scenario in Fig. 7, and the data model D21 to be evaluated is the first edition model D20 in Fig. 8. First, the data model evaluation unit 121 calculates an evaluation value for each timing indicated in the first scenario. At the first timing, data on "person's age" and "person's heart rate" are acquired. "Person's age" is a column of the first table, and "person's heart rate" is a column of the second table. Therefore, when the first edition model D20 is used, an inquiry is made to the first table of the database 212 and an inquiry is made to the second table of the database 212. In other words, the number of inquiries made to the database 212 is two. Therefore, the evaluation value for the first timing is two. Then, the data model evaluation unit 121 sums up the evaluation values for each timing, and the calculated sum becomes the evaluation value of the first edition model D20.
[0046] Continuing with the explanation of step S120, the evaluation value calculated for the data model D21 is referred to as an evaluation value D23. The data model evaluation unit 121 stores the evaluation information D22 in the storage unit 190. The evaluation information D22 indicates an evaluation value D23 of the data model D21 in association with the identifier of the data model D21. The data model evaluation unit 121 stores the data model D21 in the storage unit 190.
[0047] In step S131, the data model comparison unit 131 compares the latest evaluation value D23 with the reference value D24 to determine whether the evaluation represented by the latest evaluation value D23 is higher than the evaluation represented by the reference value D24.
[0048] Specifically, the data model comparison unit 131 determines whether the latest evaluation value D23 is smaller than the reference value D24. When the latest evaluation value D23 is smaller than the reference value D24, the evaluation represented by the latest evaluation value D23 is higher than the evaluation represented by the reference value D24. When the latest evaluation value D23 is greater than the reference value D24, the evaluation represented by the latest evaluation value D23 is lower than the evaluation represented by the reference value D24.
[0049] The latest evaluation value D23 is the evaluation value D23 of the latest data model D21 among the generated data models D21. In other words, the latest evaluation value D23 is the evaluation value D23 calculated in the immediately preceding step S120. The latest evaluation value D23 is read out from the storage unit 190.
[0050] The reference value D24 is an evaluation value that represents the highest evaluation among one or more evaluation values D23 for one or more data models D21 generated before the latest data model D21. In other words, the reference value D24 is the minimum evaluation value among one or more evaluation values D23 for one or more data models D21 generated before the latest data model D21. The reference value D24 used in the first step S131 is an initial value (e.g., the maximum value). The reference value D24 is stored in the storage unit 190 and read out from the storage unit 190.
[0051] If the latest evaluation value D23 is smaller than the reference value D24, that is, if the evaluation represented by the latest evaluation value D23 is higher than the evaluation represented by the reference value D24, the process proceeds to step S132. If the current evaluation value D23 is equal to or greater than the reference value D24, that is, if the evaluation represented by the latest evaluation value D23 is lower than the evaluation represented by the reference value D24, the process proceeds to step S151.
[0052] In step S132, the data model comparison unit 131 updates the reference value D24 to the latest evaluation value D23.
[0053] In step S133, the data model comparison unit 131 resets the unupdated count to zero. The non-update count is the number of times that the reference value D24 has not been updated. The non-update count is stored in the storage unit 190.
[0054] In step S140, the data model generation unit 132 generates a data model D31 based on the calculation information D11. The data model D31 is generated so that the evaluation value D23 of the data model D31 is small. In other words, the data model D31 is generated so that the evaluation represented by the evaluation value D23 of the data model D31 is high.
[0055] The data model D31 is generated by at least one of the following methods (1) to (3).
[0056] (1) The data model D31 is generated as follows. First, the data model generation unit 132 selects a column set with a large number of calls based on the number of calls for each column set. A set count threshold is used for the selection. The set count threshold is stored in the storage unit 190. The selected column set is called a target column set. Specifically, the data model generation unit 132 compares the number of calls for each column set with a set count threshold, and selects a column set whose number of calls is equal to or greater than the set count threshold as a target column set. Then, the data model generation unit 132 generates a data model D31 for expressing data of each column of only the target column group among the column groups in the multiple tables indicated in the database configuration information D01, in a structure suitable for processing by an application that uses the data.
[0057] An example of a data model D31 generated by the method (1) is shown in Fig. 10. The data model D31 in Fig. 10 will be described below. When the first scenario in FIG. 7 is used, the pair of columns "person's age" and "person's heart rate" is called at the first timing and the eleventh timing. That is, the number of times that the pair of columns "person's age" and "person's heart rate" is called is 2. When the pair count threshold is 2 or less, the pair of columns "person's age" and "person's heart rate" is a target column pair. In the database 212 in Fig. 6, "person's age" is a column in the first table, and "person's heart rate" is a column in the second table. When the first edition model D20 in Fig. 8 is used, conversion is performed for each table, so that the evaluation value D23 based on the number of inquiries or the amount of data communication, etc. becomes high. Therefore, in order to lower the evaluation value D23, the data model D31 in FIG. 10 is generated. Data model D31 in FIG. 10 is a data model D31 for expressing each piece of data in the column "person's age" and the column "person's heart rate" in a structure suitable for processing by an application that uses the data. In addition, since the columns "ID" and "Observation Time" are required items when reading data, the data model D31 indicates names for the columns "ID" and "Observation Time".
[0058] (2) The data model D31 is generated as follows: First, the data model generation unit 132 selects a column pair with high similarity from among combinations of columns with high call counts based on the call count for each column and the similarity for each column pair. A single call count threshold and a similarity threshold are used for the selection. The single call count threshold and the similarity threshold are stored in the storage unit 190. The selected column pair is called a target column pair. Specifically, the data model generation unit 132 compares the number of calls for each column with a single-time threshold, and selects each column whose number of calls is equal to or greater than the single-time threshold as a target column. Then, the data model generation unit 132 compares the similarity of each column pair, which is a combination of the target columns, with a similarity threshold, and selects, from among the column pairs that are combinations of the target columns, a column pair whose similarity is equal to or greater than the similarity threshold as a target column pair. Then, the data model generation unit 132 generates a data model D31 for expressing data of each column of only the target column group among the column groups in the multiple tables indicated in the database configuration information D01, in a structure suitable for processing by an application that uses the data.
[0059] An example of a data model D31 generated by the method (2) is shown in Fig. 11. The data model D31 in Fig. 11 will be described below. When the first scenario in FIG. 7 is used, the column "Person Location" is called at the second and twelfth timings. Also, the column "Robot Location" is called at the third and thirteenth timings. In other words, the number of times each of the columns "Person Location" and "Robot Location" is called is two. When the single-time threshold is two or less, each of the columns "Person Location" and "Robot Location" is a target column. 6, "person position" is a column of the first table, and "robot position" is a column of the third table. When the first edition model D20 is used, conversion is performed for each table, so that the evaluation value D23 based on the number of inquiries or the amount of data communication, etc. becomes high. The column pair “human position” and the column pair “robot position” have a high similarity because they share the word “position.” Therefore, the similarity between the column pair “human position” and the column pair “robot position” is equal to or greater than the similarity threshold. Therefore, in order to lower the evaluation value D23, the data model D31 in FIG. 11 is generated. Data model D31 in FIG. 11 is a data model D31 for expressing each of the data in the column “person location” and the column “robot location” in “location,” which is a structure suitable for processing by an application that uses the data. When there are multiple common words in the character strings of the names of columns, the data model D31 indicates conversion of the multiple common words.
[0060] (3) The data model D31 is generated as follows: First, the data model generation unit 132 selects a column pair with low similarity from among combinations of columns with many call counts, based on the call count for each column and the similarity for each column pair. A single call count threshold and a similarity threshold are used for the selection. The single call count threshold and the similarity threshold are stored in the storage unit 190. The selected column pair is called a target column pair. Specifically, the data model generation unit 132 compares the number of calls for each column with a single-time threshold, and selects each column whose number of calls is equal to or greater than the single-time threshold as a target column.The data model generation unit 132 then compares the similarity of each column pair, which is a combination of the target columns, with a similarity threshold, and selects a column pair, which is a combination of the target columns, whose similarity is less than the similarity threshold as a target column pair. Then, the data model generation unit 132 generates a data model D31 for expressing data of each column of only the target column group among the column groups in the multiple tables indicated in the database configuration information D01, in separate structures suitable for processing by an application that uses the data.
[0061] An example of a data model D31 generated by the method (3) is shown in Fig. 12. The data model D31 in Fig. 12 will be described below. When the first scenario in FIG. 7 is used, the column "Person's Name" is called at the fourth and fourteenth times. Also, the column "Robot's Model Number" is called at the fifth and fifteenth times. In other words, the number of times each of the columns "Person's Name" and "Robot's Model Number" is called is two. When the single-time threshold is two or less, each of the columns "Person's Name" and "Robot's Model Number" is a target column. 6, "person's name" is a column in the first and second tables, and "robot model number" is a column in the third table. When the first edition model D20 is used, conversion is performed for each table, so that the evaluation value D23 based on the number of inquiries or the amount of data communication, etc. becomes high. The pair of columns "person's name" and "robot's model number" has no common words, so the similarity between the pair of columns is low. Therefore, the similarity between the pair of columns "person's name" and "robot's model number" is below the similarity threshold. Therefore, in order to lower the evaluation value D23, the data model D31 in FIG. 12 is generated. Data model D31 in FIG. 12 is a data model D31 for expressing each of the data in the column “person’s name” and the column “robot model number” in separate (independent) structures suitable for processing by an application that uses the data.
[0062] The explanation of step S140 will be continued. The data model generating unit 132 changes each of the pair count threshold, the single count threshold, and the similarity threshold. For example, thresholds such as the pair count threshold, the single count threshold, and the similarity threshold are changed using machine learning as follows. The data model generation unit 132 changes the threshold to an appropriate value by using the trained model. The trained model is generated in advance and stored in the storage unit 190. The trained model is generated by a training device, which may be, for example, a device separate from the data model optimization system 100. The learning device generates a trained model by learning training data using, for example, a convolutional neural network (CNN). The training data indicates a relationship between a threshold value and an evaluation value of a data model. For example, the training data indicates a relationship between a threshold value used in another data model optimization system and an evaluation value of a data model generated by another data model optimization system.
[0063] Returning to FIG. 4, the description of step S140 will continue. The data model generation unit 132 passes the data model D31 to the data model evaluation unit 121. In addition, the data model generation unit 132 stores the data model D31 in the storage unit 190. After step S140, the process proceeds to step S120.
[0064] Proceeding to FIG. 5, the description continues from step S151. In step S151, the data model comparison unit 131 adds 1 to the unupdated count to update it.
[0065] In step S152, the data model comparison unit 131 compares the number of not-updated times with the not-updated threshold to determine whether the number of not-updated times has reached the not-updated threshold. The non-update threshold is a threshold for the number of non-updates, and is stored in advance in the storage unit 190.
[0066] If the number of unupdated times reaches the unupdated threshold, the process proceeds to step S153. If the unupdated count has not reached the unupdated threshold, the process proceeds to step S140.
[0067] In step S153, the data model comparison unit 131 outputs the data model D1. The data model D1 is a data model D21 that corresponds to the reference value D24.
[0068] The data model D1 is output as follows: First, the data model comparison unit 131 selects the evaluation information D22 that indicates the same evaluation value D23 as the reference value D24, and obtains a data model identifier from the selected evaluation information D22. Next, the data model comparison unit 131 acquires from the storage unit 190 the data model D21 identified by the acquired data model identifier. Then, the data model comparison unit 131 outputs the acquired data model D21 as the data model D1. The output data model D1 is input to the data conversion unit 211. After step S153, the process ends.
[0069] ***Advantages of the First Embodiment*** The first embodiment aims to optimize communication efficiency by simulating the behavior of a communication system for an application and generating a data model suitable for each application, taking into consideration the merging and division of database tables. The data model optimization system 100 repeats the generation of a data model based on the number of column invocations and the column similarity, and the evaluation of the data model based on the evaluation axis.
[0070] The data model optimization system 100 includes a calculation unit 110, an evaluation unit 120, and a generation unit . The calculation unit 110 calculates the number of calls of a single column, the number of calls of a combination of columns, and column similarity based on the database configuration and the data acquisition scenario of the application. The evaluation unit 120 tally up evaluation values based on the evaluation axes for cases in which data is acquired by an application using the data model, based on the data model and a data acquisition scenario of the application. The generation unit 130 compares the evaluation value based on the evaluation result with the minimum value of the data model evaluation result after the system is started. If the evaluation value exceeds the minimum value of the data model evaluation result after the system is started, the generation unit 130 generates a data model based on the evaluated data model, the number of column calls, and the similarity so that the evaluation value is smaller. This makes it possible to generate a data model suitable for each application by taking into account the merging and division of database tables, thereby optimizing communication efficiency.
[0071] The generating unit 130 generates a data model based on the column call count, the similarity, and the evaluated data model as follows. The generation unit 130 generates a data model that converts only columns in a combination that has a large number of simultaneous calls. The generating unit 130 generates a data model that converts only columns that have a high number of independent calls and high column similarity. The generating unit 130 generates a data model that converts a column having a high number of independent calls and low column similarity independently. This makes it possible to generate a data model suitable for each application, taking into account the merging and division of database tables.
[0072] The calculation unit 110 calculates column similarity from character strings in the names of columns in tables of a database, based on a database configuration and a data acquisition scenario of an application. This makes it possible to generate a data model that can convert columns with high column similarity into the same data model when generating the data model.
[0073] The generation unit 130 compares the data model evaluation value with the minimum value of the data model evaluation result after the system is started. Then, in any of the following cases, the generation unit 130 generates a new data model based on the column call count, the similarity, and the evaluated data model so as to reduce the data model evaluation value. The generating unit 130 generates a new data model when the data model evaluation value falls below the minimum value of the data model evaluation result after the system is started. The generating unit 130 generates a new data model when the data model evaluation value exceeds the minimum value of the data model evaluation result after the system is started but the number of times the minimum value of the data model evaluation result has not been updated since the system is started does not reach a threshold value. This makes it possible to generate a data model that is more communication efficient and more suitable for the application.
[0074] The generation unit 130 compares the data model evaluation value with the minimum value of the data model evaluation result after the system is started. When the data model evaluation value exceeds the minimum value of the data model evaluation result after the system is started and the number of times the minimum value of the data model evaluation result has not been updated since the system is started reaches a threshold, the generation unit 130 outputs a data model. The output data model is a data model whose data model evaluation value is the minimum value of the data model evaluation result since the system is started. This makes it possible to output an optimal data model, thereby optimizing communication efficiency.
[0075] ***Example of the first embodiment*** FIG. 13 shows an example of the functional configuration of the data processing system 200. The input and output data of the data model optimization system 100 may be stored in the network storage 230 instead of or in addition to the storage unit 190 . The network storage 230 is a storage unit provided outside the data model optimization system 100, and is composed of one or more storage devices. The data model optimization system 100 communicates with the network storage 230 to store data in the network storage 230 and to retrieve data from the network storage 230 .
[0076] ***Addition to the first embodiment*** When data from various fields is handled, such as in smart cities, a data platform is built to collect and manage the data. For data integration across disciplines, a software platform is used that enables data to be handled using a common model between applications. Such a software infrastructure converts data according to a data model defined in the interface portion of the data platform and provides the converted data to the application. Therefore, the efficiency of communication between the application and the data platform depends on the data model. The first embodiment is a technology for functions implemented within a data platform that handles various data, such as a smart city. Application developers may not know the database configuration of the data platform, so the application's data retrieval scenario only indicates which fields of data are used and the order in which the data is used.
[0077] Embodiment 2 The mode of outputting the data model D1 for which a target evaluation has been obtained will be described below with reference to Figs. 14 to 17, mainly with respect to the points that differ from the first embodiment.
[0078] ***Configuration Description*** The configuration of the data processing system 200 will be described with reference to FIG. The configuration of the data processing system 200 is the same as that in the first embodiment. However, in the data processing method, a target value D03 is used. The target value D03 is a value that indicates a target evaluation, and is set in the setting file 191.
[0079] ***Explanation of Operation*** The data model optimization method will be described with reference to FIGS. In step S210, the calculation unit 110 calculates the number of calls for each column, the number of calls for each column set, and the similarity for each column set, based on the database configuration information D01 and the data acquisition scenario D02. Step S210 is the same as step S110 in the first embodiment.
[0080] In step S220, the data model evaluation unit 121 evaluates the data model D21 based on the data acquisition scenario D02 to calculate an evaluation value of the data model D21. Step S220 is the same as step S120 in the first embodiment.
[0081] In step S231, the data model comparison unit 131 compares the latest evaluation value D23 with the target value D03 to determine whether the evaluation represented by the latest evaluation value D23 is higher than the evaluation represented by the target value D03.
[0082] Specifically, the data model comparison unit 131 determines whether the latest evaluation value D23 is smaller than the target value D03. When the latest evaluation value D23 is smaller than the target value D03, the evaluation represented by the latest evaluation value D23 is higher than the evaluation represented by the target value D03. When the latest evaluation value D23 is greater than the target value D03, the evaluation represented by the latest evaluation value D23 is lower than the evaluation represented by the target value D03.
[0083] The target value D03 is obtained from the setting file 191. The setting file 191 is stored in advance in the storage unit 190, for example.
[0084] If the latest evaluation value D23 is smaller than the target value D03, that is, if the evaluation represented by the latest evaluation value D23 is higher than the evaluation represented by the target value D03, the process proceeds to step S232. If the current evaluation value D23 is equal to or greater than the target value D03, that is, if the evaluation represented by the latest evaluation value D23 is lower than the evaluation represented by the target value D03, the process proceeds to step S241.
[0085] In step S232, the data model comparison unit 131 outputs the data model D1. The data model D1 is the data model D21 that corresponds to the latest evaluation value D23. In other words, the data model D1 is the latest data model D21.
[0086] The data model D1 is output as follows: First, the data model comparison unit 131 selects the evaluation information D22 indicating the latest evaluation value D23, and obtains a data model identifier from the selected evaluation information D22. Next, the data model comparison unit 131 acquires from the storage unit 190 the data model D21 identified by the acquired data model identifier. Then, the data model comparison unit 131 outputs the acquired data model D21 as the data model D1. The output data model D1 is input to the data conversion unit 211. After step S232, the process ends.
[0087] The process in step S231 in which the latest evaluation value D23 is equal to or greater than the target value D03 is the same as the process in the steps following step S131 in the first embodiment. That is, steps S241 to S243 are the same as steps S131 to S133 in the first embodiment. Moreover, step S250 is the same as step S140 in the first embodiment. Moreover, steps S261 to S263 are the same as steps S151 to S153 in the first embodiment.
[0088] ***Effects of the second embodiment*** According to the second embodiment, a data model that satisfies the target values can be generated using the evaluation target value setting file.
[0089] The generation unit 130 compares the data model evaluation value, the target value in the target value setting file, and the minimum value of the data model evaluation result after system startup. Then, in any of the following cases, the generation unit 130 generates a new data model based on the column call count, the similarity, and the evaluated data model so that the data model evaluation value becomes smaller. The generating unit 130 generates a new data model when the data model evaluation value exceeds the target value based on the target value setting file but falls below the minimum value of the data model evaluation result after the system is started. The generation unit 130 generates a new data model when the data model evaluation value exceeds the target value in the target value setting file, the data model evaluation value exceeds the minimum value of the data model evaluation result after system startup, and the number of times the minimum value of the data model evaluation result has not been updated after system startup has not reached a threshold value. This makes it possible to generate a data model that has higher communication efficiency and meets the target values.
[0090] The generation unit 130 compares the data model evaluation value, the target value in the target value setting file, and the minimum value of the data model evaluation result after the system is started. Then, in any of the following cases, the generation unit 130 outputs a data model whose data model evaluation value is lower than the target value in the target value setting file, or a data model whose data model evaluation value is the minimum value of the data model evaluation result after the system is started. The generating unit 130 outputs the data model when the data model evaluation value is below the target value in the target value setting file. The generation unit 130 outputs a data model when the data model evaluation value exceeds the target value in the target value setting file, the data model evaluation value exceeds the minimum value of the data model evaluation result after system startup, and the number of times the minimum value of the data model evaluation result has not been updated after system startup has reached a threshold value. This makes it possible to output a data model that satisfies the target value, or a data model that does not satisfy the target value but has high communication efficiency.
[0091] ***Supplementary explanation of implementation form*** The hardware configuration of the data model optimization system 100 will be described with reference to FIG. The data model optimization system 100 includes a processing circuit 109 . The processing circuit 109 is hardware that realizes the calculation unit 110, the evaluation unit 120, and the generation unit 130. The processing circuitry 109 may be dedicated hardware, or may be a processor 101 that executes a program stored in the memory 102 .
[0092] When the processing circuitry 109 is dedicated hardware, the processing circuitry 109 may be, for example, a single circuit, a multiple circuit, a programmed processor, a parallel programmed processor, an ASIC, an FPGA, or a combination thereof. ASIC is an abbreviation for Application Specific Integrated Circuit. FPGA is an abbreviation for Field Programmable Gate Array.
[0093] The data model optimization system 100 may include multiple processing circuits replacing the processing circuit 109 .
[0094] In the processing circuit 109, some functions may be realized by dedicated hardware, and the remaining functions may be realized by software or firmware.
[0095] Thus, the functionality of the data model optimization system 100 may be implemented in hardware, software, firmware, or a combination thereof.
[0096] Each embodiment is an example of a preferred embodiment, and is not intended to limit the technical scope of the present disclosure. Each embodiment may be implemented in part or in combination with other embodiments. The procedures described using flow charts, etc. may be modified as appropriate.
[0097] The "section" of each element of the data model optimization system 100 may be read as "processing", "step", "circuit", or "circuitry".
Explanation of Signs
[0098] 100 data model optimization system, 101 processor, 102 memory, 103 auxiliary storage device, 104 communication device, 105 input / output interface, 109 processing circuit, 110 calculation unit, 111 single call count calculation unit, 112 group call count calculation unit, 113 similarity calculation unit, 120 evaluation unit, 121 data model evaluation unit, 130 generation unit, 131 data model comparison unit, 132 data model generation unit, 190 storage unit, 191 setting file, 200 data processing system, 210 data platform, 211 data conversion unit, 212 database, 221 application unit, 230 network storage, D1 data model, D2 data request, D3 inquiry, D01 database configuration information, D02 data acquisition scenario, D03 target value, D11 calculation information, D20 first version model, D21 data model, D22 evaluation information, D23 evaluation value, D24 reference value, D31 data model.
Claims
1. a single call count calculation unit that calculates, for each column in the plurality of tables, a call count for the column when data is acquired from the database in accordance with the data acquisition scenario, based on database configuration information indicating a configuration of a plurality of tables in the database and a data acquisition scenario that specifies data to be acquired from the database; a set call count calculation unit that calculates, for each column set in the plurality of tables, a call count at which the column set is called at the same timing when data is acquired from the database according to the data acquisition scenario, based on the database configuration information and the data acquisition scenario; a similarity calculation unit that calculates a similarity between names of columns for each column group in the plurality of tables based on the database configuration information; a data model generation unit that generates a data model for expressing data acquired according to the data acquisition scenario in a structure suitable for processing by an application that uses the data, based on the number of calls for each column, the number of calls for each column set, and the similarity for each column set; A data model optimization system comprising:
2. The data model optimization system includes a data model evaluation unit, the data model evaluation unit evaluates the generated data model based on the data acquisition scenario every time a data model is generated, and calculates an evaluation value of the generated data model; The data model generation unit generates the new data model so that the evaluation represented by the evaluation value of the new data model is high. The data model optimization system of claim 1 .
3. The data model optimization system includes a data model comparison unit, the data model comparison unit compares a latest evaluation value, which is an evaluation value of a latest data model among the generated data models, with a reference value, which is an evaluation value representing a highest evaluation among one or more evaluation values for one or more data models generated before the latest data model, and determines whether the evaluation represented by the latest evaluation value is higher than the evaluation represented by the reference value; The data model generation unit generates the new data model when the evaluation represented by the latest evaluation value is higher than the evaluation represented by the reference value. The data model optimization system of claim 2 .
4. the data model comparison unit updates the reference value to the latest evaluation value when the evaluation represented by the latest evaluation value is higher than the evaluation represented by the reference value, and does not update the reference value to the latest evaluation value when the evaluation represented by the latest evaluation value is lower than the evaluation represented by the reference value; The data model generation unit generates the new data model when the evaluation represented by the latest evaluation value is lower than the evaluation represented by the reference value, but a non-update count, which is a number of times the reference value has not been updated, does not reach a non-update threshold. The data model optimization system of claim 3 .
5. The data model comparison unit outputs a data model corresponding to the reference value when the number of times that the non-update has occurred reaches the non-update threshold. The data model optimization system of claim 4.
6. The data model optimization system includes a data model comparison unit, The data model comparison unit compares a latest evaluation value, which is an evaluation value of a latest data model among the generated data models, with a target value, determines whether the evaluation represented by the latest evaluation value is higher than the evaluation represented by the target value, and outputs the latest data model when the evaluation represented by the latest evaluation value is higher than the evaluation represented by the target value. The data model optimization system of claim 2 .
7. The data model generation unit generates the new data model when the evaluation represented by the latest evaluation value is lower than the evaluation represented by the target value. The data model optimization system of claim 6.
8. the data model comparison unit compares the latest evaluation value with a reference value that is an evaluation value that represents a highest evaluation among one or more evaluation values for one or more data models generated before the latest data model, and determines whether the evaluation represented by the latest evaluation value is higher than the evaluation represented by the reference value; The data model generation unit generates the new data model when the evaluation represented by the latest evaluation value is higher than the evaluation represented by the reference value among cases in which the evaluation represented by the latest evaluation value is lower than the evaluation represented by the target value. The data model optimization system of claim 7.
9. the data model comparison unit updates the reference value to the latest evaluation value when the evaluation represented by the latest evaluation value is higher than the evaluation represented by the reference value, and does not update the reference value to the latest evaluation value when the evaluation represented by the latest evaluation value is lower than the evaluation represented by the reference value; The data model generation unit generates the new data model when the evaluation represented by the latest evaluation value is lower than the evaluation represented by the target value, and when the evaluation represented by the latest evaluation value is lower than the evaluation represented by the reference value but a non-update count, which is a number of times the reference value has not been updated, does not reach a non-update threshold. The data model optimization system of claim 8.
10. The data model comparison unit outputs a data model corresponding to the reference value when the evaluation represented by the latest evaluation value is lower than the evaluation represented by the target value but the number of times of non-update reaches the non-update threshold. The data model optimization system of claim 9.
11. The data model generation unit selects a column set that is called frequently based on the number of calls for each column set as a target column set, and generates a data model for expressing data of each column of the target column set only among the column sets in the plurality of tables in a structure suitable for processing of an application that uses the data. The data model optimization system according to any one of claims 1 to 10.
12. The data model generation unit selects, as a target column set, a column set having a high similarity from among combinations of columns having a high number of calls based on the number of calls for each column and the similarity for each column set, and generates a data model for expressing data of each column of the target column set only for the target column set among the column sets in the plurality of tables in a structure suitable for processing of an application that uses the data. The data model optimization system according to any one of claims 1 to 10.
13. The data model generation unit selects, as a target column set, a column set having a low similarity from among combinations of columns having a high call count based on the number of calls for each column and the similarity for each column set, and generates a data model for expressing data of each column of the target column set only for the target column set among the column sets in the plurality of tables in separate structures suitable for processing of an application that uses the data. The data model optimization system according to any one of claims 1 to 10.
14. based on database configuration information indicating the configuration of a plurality of tables in a database and a data acquisition scenario specifying data to be acquired from the database, calculates, for each column in the plurality of tables, the number of times that the column will be called when data is acquired from the database in accordance with the data acquisition scenario; calculating, for each column group in the plurality of tables, a number of invocations that the column group will be invoked at the same timing when data is acquired from the database according to the data acquisition scenario, based on the database configuration information and the data acquisition scenario; Calculating a similarity between names of columns for each column group in the plurality of tables based on the database configuration information; A data model is generated for expressing data acquired according to the data acquisition scenario in a structure suitable for processing by an application that uses the data, based on the number of calls for each column, the number of calls for each column set, and the similarity for each column set. Data model optimization methods.
15. a single call count calculation process for calculating, for each column in the plurality of tables, the number of calls that the column will be called when data is acquired from the database according to the data acquisition scenario, based on database configuration information indicating the configuration of a plurality of tables in the database and a data acquisition scenario that specifies data to be acquired from the database; a set call count calculation process for calculating, for each column set in the plurality of tables, a call count at which the column set is called at the same timing when data is acquired from the database according to the data acquisition scenario, based on the database configuration information and the data acquisition scenario; a similarity calculation process for calculating a similarity between names of columns for each column group in the plurality of tables based on the database configuration information; a data model generation process for generating a data model for expressing data acquired according to the data acquisition scenario in a structure suitable for processing by an application that uses the data, based on the number of calls for each column, the number of calls for each column set, and the similarity for each column set; A data model optimization program for running the above on a computer.
Citation Information
Patent Citations
Synonymous column detecting device and synonymous column detecting method
JP2011232879A
Virtual database system management device, management method, and management program
JP2015179449A
Correlation rule analysis device and correlation rule analysis method
JP2016014944A
Analysis support method, analysis support server and storage media
JP2019109676A
Techniques to determine relationships of items in web-based content
US10922374B1