Data Management System and Data Management Method

The data management system improves search efficiency and convenience for semi-structured data by grouping keys and optimizing queries, addressing the inefficiencies in existing relational database management systems.

JP7706353B2Active Publication Date: 2025-07-11HITACHI LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021205565
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-07-11
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

Managing semi-structured data in relational databases leads to increased construction and operation costs, decreased search convenience, and efficiency due to the configuration of tables and columns, especially with increasing table and column numbers.

Method used

A data management system that divides semi-structured data keys into key groups, generates key group tables with specified search keys, and optimizes search queries using these tables to improve efficiency.

Benefits of technology

Enhances search convenience and efficiency for semi-structured data by reducing unnecessary table and column counts through strategic key grouping and query optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007706353000001
    Figure 0007706353000001
  • Figure 0007706353000002
    Figure 0007706353000002
  • Figure 0007706353000003
    Figure 0007706353000003
Patent Text Reader

Abstract

To provide a data management system capable of improving search convenience and search efficiency for semi-structured data.SOLUTION: A data management system 1 manages JSON data. The data management system 1 includes one or more processors and a JSON table 24 that stores JSON data. For each of the processors, multiple keys included in multiple data units of JSON data are divided into multiple key groups based on the relationship between keys in multiple data units and a search key for JSON data is identified. For each of the multiple key groups, when a search key is included in the key groups, a key group table that includes a search key column and a data unit column is generated.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for managing semi-structured data.

Background Art

[0002] For example, the production record data of a product may have different measurement parameters for each process, and even in the same process, the measurement parameters may be changed over time due to changes in the production process or operation. Therefore, it may be managed by semi-structured data such as JSON (JavaScript Object Notation: JavaScript is a registered trademark) format data (referred to as JSON data).

[0003] Since the search efficiency for JSON data is low, a technique for storing JSON data in a relational database table is known.

[0004] For example, as a technique for storing XML data in a relational database, a technique is known in which XML data is stored as it is in a single column, and a side table with XML keys as columns is defined according to the access definition of the XML column (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] When managing semi-structured data in the tables of a relational database, depending on the configuration of the table and its columns, the cost required for table construction and operation may increase significantly, resulting in a decrease in search convenience and search efficiency, or there may be a risk that the search efficiency will not improve. For example, when the number of tables increases or the number of columns in a table increases, the cost required for table construction and operation increases, search convenience decreases, and search efficiency decreases.

[0007] The present invention has been made in view of the above circumstances, and an object thereof is to provide a technology capable of improving search convenience and search efficiency for semi-structured data.

Means for Solving the Problems

[0008] To achieve the above object, a data management system according to one aspect is a data management system for managing semi-structured data, the data management system including one or more processors and a semi-structured data storage unit for storing the semi-structured data, and the processor divides a plurality of keys included in a plurality of data units of the semi-structured data into a plurality of key groups based on the relationship between the keys in the plurality of data units, specifies a search key for the semi-structured data, and for each of the plurality of key groups, when the search key is included in the key group, generates a key group table including a column of the search key and a column indicating the data unit.

Effects of the Invention

[0009] According to the present invention, it is possible to improve search convenience and search efficiency for semi-structured data.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

[0011] The embodiments will be described with reference to the drawings. Note that the embodiments described below do not limit the invention according to the claims, and not all of the elements and combinations thereof described in the embodiments are essential for the solution means of the invention.

[0012] In the following description, information may be described using expressions such as "AAA table" and "AAA log". However, the information may be represented in any data structure. That is, in order to indicate that the information is independent of the data structure, "AAA table" and "AAA log" can be referred to as "AAA information".

[0013] FIG. 1 is an overall configuration diagram of a data management system according to the first embodiment.

[0014] The data management system 1 includes a key group table generation server 10, a database management server 20, a data search server 30, a data insertion server 40, a client PC 50, and a network 60 that communicably connects these devices.

[0015] The network 60 is, for example, a wired LAN (Local Area Network), a wireless LAN, a WAN (Wide Area Network), or the like.

[0016] The key group table generation server 10 is composed of, for example, a computer such as a PC (Personal Computer) or a server. The key group table generation server 10 performs a process of generating a key group table. The key group table generation server 10 includes a JSON key grouping unit 11, a search key extraction unit 12, a key group table generation unit 13, a key group table generation interface 14, JSON key group information 15, search key aggregation information 16, a JSON key update log 17, a search key log 18, and a key-data mapping table 19. Details of each configuration will be described later.

[0017] The database management server 20 is composed of, for example, a computer such as a PC or a server. The database management server 20 performs a process of managing a database. The database management server 20 includes a database management system 21, a data conversion / transfer unit 22, a search query rewriting unit 23, a JSON table 24, a search query log 25, and a key group table group 26. The database management system 21 manages a database composed of tables, and has, for example, a function of creating an index or the like for searching a table and performing a search using the index. Details of each configuration will be described later.

[0018] The data search server 30 is composed of, for example, a computer such as a PC or a server. The data search server 30 performs a search process on the database managed by the database management server 20 and performs various processes based on the search results. The data search server 30 has a data search unit 31. Details of the data search unit 31 will be described later.

[0019] The data insertion server 40 is composed of, for example, a computer such as a PC or a server. The data insertion server 40 performs an insertion process of inserting the manufacturing result data obtained at a factory (not shown) etc. as data in JSON format (JSON data), which is an example of semi-structured data, into the database managed by the database management server 20. The data insertion server 40 has a data insertion unit 41. Details of the data insertion unit 41 will be described later.

[0020] The client PC 50 is composed of, for example, a computer such as a PC. The client PC 50 is used by a user who generates a key group table, and for example, receives a generation instruction (generation request) of the key group table from the user by the key group table generation interface 14 and transmits it to the key group table generation server 10.

[0021] Next, the JSON table 24 will be described.

[0022] FIG. 2 is a configuration diagram of the JSON table according to the first embodiment.

[0023] The JSON table 24 stores an entry for each JSON data (data unit) inserted from the data insertion server 40. An entry of the JSON table 24 stores fields of an ID 24a and a JSON 24b. In the ID 24a, identification information (ID) for identifying the JSON data corresponding to the entry is stored. In the JSON 24b, the JSON data corresponding to the entry is stored.

[0024] Next, the search query log 25 will be described.

[0025] FIG. 3 is a configuration diagram of the search query log according to the first embodiment.

[0026] The search query log 25 stores one or more search queries 251 issued by the data search unit 31. The search query is described by, for example, a SELECT statement and has a SELECT clause that specifies keys to be selected from a table, a FROM clause that describes the table to be searched, and a WHERE clause that describes conditions for one or more search keys.

[0027] Next, the key group table group 26 will be described.

[0028] FIG. 4 is a configuration diagram of the key group table group according to the first embodiment.

[0029] After the key group tables are generated, the key group table group 26 stores one or more key group tables. In the example of FIG. 4, the key group table group 26 includes, as key group tables, a group 1 table (group_1_Table) 261, a group 2 table (group_2_Table) 262, and a group 3 table (group_3_Table) 263.

[0030] The group 1 table 261 is a table corresponding to the keys of the JSON data divided into the key group of group 1, and each entry includes fields of an ID 261a, a parent ID 261b, a Time 261c, and a Process 261d.

[0031] The ID 261a stores the identifier (ID) of the entry. The parent ID 261b stores the identifier (ID: parent ID) of the JSON data having the key value of the entry. The Time 261c stores the value of Time of the JSON data corresponding to the entry. The Process 261d stores the value of Process of the JSON data corresponding to the entry.

[0032] The group 2 table 262 is a table corresponding to the keys of the JSON data divided into the key group of group 2, and each entry includes fields of an ID 262a, a parent ID 262b, and an Acc 262c.

[0033] In ID262a, the identifier (ID) of the entry is stored. In the parent ID262b, the identifier (ID: parent ID) of the JSON data having the key value of the entry is stored. In Acc262c, the Acc value of the JSON data corresponding to the entry is stored.

[0034] The group 3 table 263 is a table corresponding to the keys of the JSON data divided into key groups of group 3, and each entry includes fields of ID263a, parent ID263b, Temp263c, and Velocity263d.

[0035] In ID263a, the identifier (ID) of the entry is stored. In the parent ID263b, the identifier (ID: parent ID) of the JSON data having the key value of the entry is stored. In Temp263c, the Temp value of the JSON data corresponding to the entry is stored. In Velocity263d, the Velocity value of the JSON data corresponding to the entry is stored.

[0036] Next, the JSON key group information 15 will be described.

[0037] FIG. 5 is a configuration diagram of the JSON key group information according to the first embodiment.

[0038] The JSON key group information 15 manages the keys of each group (JSON key group, key group) when each key in the JSON data is grouped. The JSON key group information 15 stores entries for each JSON key group. The entry of the JSON key group information 15 includes fields of JSON key group 15a and JSON key 15b. In the JSON key group 15a, the group name of the JSON key group corresponding to the entry is stored. In the JSON key 15b, the key (JSON key) of the JSON data belonging to the JSON key group corresponding to the entry is stored.

[0039] Next, the JSON key update log 17 will be described.

[0040] FIG. 6 is a configuration diagram of the JSON key update log according to the first embodiment.

[0041] The JSON key update log 17 manages the update of keys in the JSON key group. The JSON key update log 17 includes an entry for each key to be updated. An entry in the JSON key update log 17 includes fields of a date and time 17a, a JSON key group 17b, a JSON key 17c, and an update content 17d.

[0042] The date and time 17a stores the date and time when the key corresponding to the entry was updated. The JSON key group 17b stores the group name of the JSON key group to which the key corresponding to the entry belongs. The JSON key 17c stores the updated JSON key corresponding to the entry. The update content 17d stores the content of the update corresponding to the entry.

[0043] Next, the search key aggregation information 16 will be described.

[0044] FIG. 7 is a configuration diagram of the search key aggregation information according to the first embodiment.

[0045] The search key aggregation information 16 manages the search keys in the JSON key group. The search key aggregation information 16 stores an entry for each JSON key group. An entry in the search key aggregation information 16 includes fields of a JSON key group 16a and a search key 16b.

[0046] The JSON key group 16a stores the group name of the JSON key group corresponding to the entry. The search key 16b stores the search key in the JSON key group corresponding to the entry. If there is no search key in the JSON key group corresponding to the entry, nothing is stored in the search key 16b. If there are multiple search keys, multiple search keys are stored.

[0047] Next, the search key log 18 will be described.

[0048] FIG. 8 is a configuration diagram of the search key log according to the first embodiment.

[0049] The search key log 18 manages information on the search keys used in searches. The search key log 18 stores entries corresponding to searches. An entry in the search key log 18 includes fields of a date and time 18a, a JSON key group 18b, a search key 18c, and a search time 18d.

[0050] The date and time 18a stores the date and time when the search corresponding to the entry was performed. The JSON key group 18b stores the group name of the JSON key group to which the search key in the search corresponding to the entry belongs. The search key 18c stores the search key in the search corresponding to the entry. The search time 18d stores the time required for the search (search time) corresponding to the entry. Note that when there are a plurality of entries corresponding to a single search in the search time 18d, the time required for a single search is stored.

[0051] For example, the entries in the first and second lines are entries corresponding to a single search, indicating that the search was performed at 13:00 on August 31, 2021, and the search was performed using Time, Process in Group_1 and Acc in Group_2 as search keys, and it took 10 seconds for the search.

[0052] Next, the key group table generation interface 14 will be described.

[0053] FIG. 9 is a diagram showing a configuration example of the key group table generation interface according to the first embodiment.

[0054] The key group table generation interface 14 is an interface for a user to give an instruction to create a key group table, and the user can specify a JSON key to be a column of the key group table to be created. According to the key group table generation interface 14, the user can generate a key group table having desired keys as columns. In the example of FIG. 9, it shows an instruction to generate a key group table having keys of Time, Process, S1, S2, and S4 among the keys in the JSON data as columns.

[0055] Next, the key-data mapping table 19 will be described.

[0056] FIG. 10 is a configuration diagram of the key-data mapping table according to the first embodiment.

[0057] The key-data mapping table 19 is a table for managing JSON keys included in JSON data, and a key group is determined by this key-data mapping table 19. The row 191 of the key-data mapping table 19 stores identification information (ID) indicating the JSON data stored in the JSON table 24. The column 192 stores the keys stored in the JSON data. In the cell corresponding to the ID of the JSON data in the row 191 and the key in the column 192, information indicating whether or not the key is included in the JSON data indicated by the ID is stored. In the present embodiment, when the key is included in the JSON data indicated by the ID, a black circle is stored. According to the key-data mapping table 19, the keys included in each JSON data can be grasped, and the keys in a plurality of JSON data can be divided into several groups (JSON key groups).

[0058] Next, an example of a computer constituting the key group table generation server 10, the database management server 20, the data search server 30, the data insertion server 40, and the client PC 50 will be described.

[0059] FIG. 11 is a hardware configuration diagram of a computer according to the first embodiment.

[0060] The computer 100 includes a communication interface (I / F) 101, a CPU 102 as an example of a processor, an input device 103, a storage device 104, a memory 105, and a display device 106. The CPU 102, the input device 103, the storage device 104, the memory 105, and the display device 106 are connected via a bus 107. Note that in a computer that constitutes any one of the key group table generation server 10, the database management server 20, the data search server 30, and the data insertion server 40, the input device 103 and the display device 106 may not be provided.

[0061] The communication I / F 101 is an interface such as a wired LAN card or a wireless LAN card, and communicates with other computers via a network 60.

[0062] The CPU 102 executes various processes according to programs stored in the memory 105 and / or the storage device 104. In the case of the computer 100 that constitutes the key group table generation server 10, when the CPU 102 executes a program, it constitutes a JSON key grouping unit 11, a search key extraction unit 12, and a key group table generation unit 13. Also, in the case of the computer 100 that constitutes the database management server 20, when the CPU 102 executes a program, it constitutes a database management system 21, a data conversion / transfer unit 22, and a search query rewriting unit 23. Also, in the case of the computer 100 that constitutes the data search server 30, when the CPU 102 executes a program, it constitutes a data search unit 31. Also, in the case of the computer 100 that constitutes the data insertion server 40, when the CPU 102 executes a program, it constitutes a data insertion unit 41.

[0063] The memory 105 is, for example, a RAM (RANDOM ACCESS MEMORY), and stores programs executed by the CPU 102 and necessary information.

[0064] The memory device 104 is, for example, a hard disk, a flash memory, etc., and stores programs executed by the CPU 102 and data used by the CPU 102. For example, in the case of the computer 100 constituting the key group table generation server 10, the memory device 104 stores JSON key group information 15, search key aggregation information 16, JSON key update log 17, search key log 18, and key-data mapping table 19. Also, in the case of the computer 100 constituting the database management server 20, the memory device 104 stores a JSON table 24, a search query log 25, and a key group table group 26. In this case, the memory device 104 corresponds to a semi-structured data storage unit.

[0065] The input device 103 is, for example, a mouse, a keyboard, etc., and receives input of information by the user. The display device 106 is, for example, a display, and displays and outputs a user interface including various information.

[0066] Next, the processing by the data management system 1 according to the first embodiment will be described.

[0067] First, the JSON data insertion process will be described.

[0068] FIG. 12 is a diagram for explaining the JSON data insertion process according to the first embodiment.

[0069] The data insertion unit 41 of the data insertion server 40 uses the production result data acquired at a factory or the like (not shown) as JSON data, and transmits an insertion query for inserting the JSON data into the JSON table 24 to the database management system 21 of the database management server 20. Here, the insertion query includes the JSON data. When the database management system 21 receives the insertion query, it adds a new entry to the JSON table 24 and stores the JSON data included in the insertion query in the JSON 24b of the entry.

[0070] Next, the data search process will be described.

[0071] FIG. 13 is a diagram for explaining the data search process according to the first embodiment.

[0072] The data search unit 31 of the data search server 30 transmits a search query for searching for desired data from the JSON table 24 to the search query rewriting unit 23 of the database management server 20.

[0073] When the search query rewriting unit 23 receives a search query, it stores the received search query in the search query log 25. The search query rewriting unit 23 acquires information (table information) of the key group table in the key group table group 26. Here, the table information includes information for identifying the key group table existing in the key group table group 26 and information on the keys (which are columns) managed by the key group table.

[0074] Based on the acquired table information, the search query rewriting unit 23 rewrites the search query into a search query (rewritten search query) including a partial query for the key group table that performs a search using the key group table and the JSON table 24, and a partial query for the semi-structured data storage unit for searching the semi-structured data in the semi-structured data storage unit. When no key group table has been generated in the key group table group 26 or when the search query has no relation to the key group table, the search query rewriting unit 23 uses the received search query as the rewritten search query as it is.

[0075] The search query rewriting unit 23 passes the rewritten search query to the database management system 21. The database management system 21 executes a search process on the JSON table 24 or the JSON table 24 and the key group tables in the key group table group 26 based on the rewritten search query, and returns the search result data (searched data) to the search query rewriting unit 23.

[0076] The search query rewriting unit 23 transmits the returned search data to the data search unit 31 that is the source of the search query.

[0077] By this process, search data corresponding to the search query is returned to the data search unit 31.

[0078] Next, the rewriting of the search query by the search query rewriting unit 23 will be described.

[0079] FIG. 14 is a diagram for explaining the rewriting of the search query according to the first embodiment. FIG. 14 shows an example in which the search query 32 is rewritten to the rewritten search query 33 when the key group table group 26 is in the state shown in FIG. 4.

[0080] The search query 32 is a search query for obtaining the values of the keys of Time, Process, Acc, Rot, Volt, and Volt2 from the JSON table 24. This search query 32 is a search query for the JSON table 24, and the data search unit 31 that transmits the search query 32 does not need to create a search query in consideration of the configuration of the key group table in the key group table group 26.

[0081] The search query rewriting unit 23 obtains table information from the key group table group 26 and recognizes the configuration of the key group table in the key group table group 26. In the present embodiment, the search query rewriting unit 23 includes a group 1 table 261, a group 2 table 262, and a group 3 table 263 having the configuration shown in FIG. 4, and can recognize the keys that are the columns of the respective tables.

[0082] Here, the search query rewriting unit 23 identifies whether each key included in the search query 32 exists in any of the key group tables. For the values of the keys that exist in the key group tables, they are obtained from the columns of the key group tables, and for the values of the keys that do not exist in any key group, they are obtained from the JSON table 24, and thus rewrites the target of the SELECT. Further, the search query rewriting unit 23 rewrites the content of the FROM clause to each table that obtains the key values, and as the WHERE clause, sets the part where the ID of the JSON table and the parent ID of each group table match so that the JSON data of the JSON table 24 and the key values of each key group table are the same, and completes the rewritten search query 33. Here, the part that designates the search by the search key for the key group table in the rewritten search query 33 corresponds to the partial query for the key group table, and the part that designates the search of the JSON data of the JSON table 24 corresponds to the partial query for the semi-structured data storage unit.

[0083] By this process, the search query 32 is rewritten like the rewritten search query 33.

[0084] Next, the data conversion and transfer process to the key group table will be described.

[0085] FIG. 15 is a diagram for explaining the data conversion and transfer process to the key group table according to the first embodiment.

[0086] The data conversion and transfer unit 22 performs data conversion and transfer processing on the entries in the JSON table 24 that have not yet been transferred to the key group tables of the key group table group 26.

[0087] The data conversion and transfer unit 22 retrieves the target entry from the JSON table 24, and identifies the keys that exist as columns in the key group tables of the key group table group 26 from among the JSON data included in the JSON 24b of the entry. Next, the data conversion and transfer unit 22 adds a new entry to the key group table having the column of the identified key, and stores the value of the key in that entry. Here, the data conversion and transfer unit 22 stores the ID of the target entry of the JSON table 24 in the parent ID in the added entry.

[0088] By this process, the JSON data stored in the JSON table 24 can be appropriately transferred to the key group tables of the key group table group 26.

[0089] Next, the generation process of the key group table will be described.

[0090] FIG. 16 is a diagram for explaining the generation process of the key group table according to the first embodiment.

[0091] The JSON key grouping unit 11 refers to the entries of the JSON table 24 and creates a key-data mapping table 19.

[0092] Next, the JSON key grouping unit 11 refers to the key-data mapping table 19 and divides (groups) the JSON keys into groups of related keys. For example, in the example of the key-data mapping table 19 in FIG. 10, the JSON key grouping unit 11 divides the keys into group 1, which is a group of keys common to each JSON data, and groups 2, 3, and 4, which are groups of keys common to the JSON data of each Process. Next, the JSON key grouping unit 11 registers or updates the JSON key group information 15 according to the grouping, and registers the content in the JSON key update log 17.

[0093] On the one hand, the search key extraction unit 12 refers to the search query log 25, extracts search keys from the search query, refers to the JSON key group information 15, identifies the JSON key group to which the search keys belong, registers the relationship between the search keys and the JSON key group in the search key aggregation information 16, and registers the log of the search keys related to this search query in the search key log 18.

[0094] Next, the key group table generation unit 13 generates a key group table based on the information of the JSON key group information 15, the JSON key update log 17, the search key aggregation information 16, and the search key log 18, and stores it in the key group table group 26. When a key group table generation instruction is received from the user through the key group table generation interface 14, the key group table generation unit 13 generates a key group table based on the key group table generation instruction and stores it in the key group table group 26.

[0095] Next, the JSON key grouping process for dividing the keys of the JSON data into JSON key groups will be described.

[0096] FIG. 17 is a flowchart of the JSON key grouping process according to the first embodiment.

[0097] The JSON key grouping process is executed, for example, at regular intervals. First, the JSON key grouping unit 11 of the key group table generation server 10 determines whether there is new JSON data in the JSON table 24 (step S11).

[0098] As a result, if there is no new JSON data (step S11: No), the JSON key grouping section 11 ends the process. On the other hand, if there is new JSON data (step S11: Yes), the JSON key grouping section 11 acquires the new JSON data, analyzes the JSON data to extract JSON keys, and updates the key-data mapping table 19 (step S12). Specifically, the JSON key grouping section 11 adds a column corresponding to the JSON data, and registers a symbol indicating the existence of the JSON key in the row corresponding to the JSON key included in the JSON data. Note that when the JSON data includes a JSON key that does not exist in the row of the key-data mapping table 19, the JSON key grouping section 11 adds a row corresponding to this JSON key.

[0099] Next, the JSON key grouping section 11 executes the process of loop 1 (steps S13 to S16) with one of the JSON keys in the key-data mapping table 19 as the processing target. Here, in the description of this process, the JSON key to be processed is referred to as the target key.

[0100] First, the JSON key grouping section 11 detects the similarity between the set of JSON data having the target key and the set of JSON data having other keys. If the similarity is high (equal to or higher than a predetermined threshold value), the target key and the other keys are defined as the same JSON key group (step S13). Note that each key defined as a JSON key group is excluded from the processing target of the subsequent loop 1.

[0101] Next, the JSON key grouping section 11 determines whether a new JSON key group has been defined in step S13 (step S14). If no new JSON key group has been defined (step S14: No), the process of loop 1 for the target key ends, and the process of loop 1 for the next processing target is executed.

[0102] On the other hand, when a new JSON key group is defined (step S14: Yes), the JSON key grouping unit 11 updates the JSON key group information 15 to register the new JSON key group (step S15). Next, the JSON key grouping unit 11 writes the updated content of the JSON key group information 15 to the JSON key update log 17 (step S16), ends the process of loop 1 for the target key, and executes the process of loop 1 for the next processing target.

[0103] If all the keys of the JSON keys in the key-data mapping table 19 are divided into any JSON key group in the process of loop 1, the JSON key grouping unit 11 exits loop 1 and ends the JSON key grouping process.

[0104] According to this JSON key grouping process, the JSON keys of the JSON data can be appropriately divided into JSON key groups.

[0105] Next, the search key extraction process will be described.

[0106] FIG. 18 is a flowchart of the search key extraction process according to the first embodiment.

[0107] The search key extraction process is executed, for example, at regular intervals. First, the search key extraction unit 12 of the key group table generation server 10 determines whether there is a new search query for the JSON table 24 in the search query log 25 (step S21).

[0108] As a result, when there is no new search query (step S21: No), the search key extraction unit 12 ends the search key extraction process. On the other hand, when there is a new search query (step S21: Yes), the search key extraction unit 12 acquires the new search query, analyzes the search query, extracts the search key used for narrowing down the search conditions (step S22), and writes the log related to the extracted search key to the search key log 18 (step S23).

[0109] Next, the search key extraction unit 12 determines whether the extracted search key is a search key that is used for the first time in the search, specifically, whether it is a search key not stored in the search key aggregation information 16 (step S24). If the search key is not a newly used one (step S24: No), the search key extraction process ends.

[0110] On the other hand, if the search key is a newly used one (step S24: Yes), the search key extraction unit 12 updates the search key aggregation information 16 by registering the extracted search key in the entry corresponding to the JSON key group to which this search key belongs (step S25), and then ends the process.

[0111] According to this search key extraction process, the search keys used in the JSON key group can be appropriately registered in the search key aggregation information 16.

[0112] Next, the key group table generation process will be described.

[0113] FIG. 19 is a flowchart of the key group table generation process according to the first embodiment.

[0114] The key group table generation process is repeatedly executed, for example, when the key group table generation server 10 is activated. The key group table generation unit 13 of the key group table generation server 10 executes the process of loop 2 (steps S31 to S37) for each JSON key group in the JSON key group information 15. Here, the JSON key group to be processed is referred to as the target key group.

[0115] First, the key group table generation unit 13 determines whether there is a request to generate a key group table from the user through the key group table generation interface 14 (step S31). As a result, if there is a request to generate a key group table (step S31: Yes), the key group table generation unit 13 generates a key group table including the key specified by the user included in the generation request as a column (step S37), and proceeds to the end of loop 2. Note that for the generated key group table, the key value for the JSON data of the JSON table 24 will be transferred by the data conversion / transfer unit 22.

[0116] On the other hand, if there is no request to generate a key group table (step S31: No), the key group table generation unit 13 refers to the search key aggregation information 16 and determines whether the key of the target key group exists in the search key (step S32).

[0117] As a result, if it is determined that the key of the target key group does not exist in the search key (step S32: No), it means that the target key group is not used for search, and since there is no need to create a key group table, the key group table generation unit 13 ends the process of loop 2 for the target key group without generating a key group table for the target key group. Thereby, it is possible to appropriately prevent the generation of a useless key group table that is not used for search.

[0118] On the other hand, if it is determined that the key of the target key group exists in the search key (step S32: Yes), the key group table generation unit 13 refers to the search key log 18 and determines whether the search time using the search key of the target key group is a certain amount or more, or whether it tends to deteriorate (step S33).

[0119] As a result, when it is determined that the search time using the search key for the target key group is not longer than a certain period and does not tend to deteriorate (step S33: No), the current state is acceptable and there is no need to generate a key group table for the target key group. Therefore, the key group table generation unit 13 ends the processing of loop 2 for the target key group without generating a key group table for the target key group.

[0120] On the other hand, when the search time using the search key for the target key group is longer than a certain period or tends to deteriorate (step S33: Yes), the key group table generation unit 13 determines whether the search time can be shortened by a certain amount or more by searching the key group table using the search key rather than searching the JSON table 24 (step S34). Specifically, the key group table generation unit 13 calculates the search cost when the JSON table 24 is assumed to be a relational table without an index (JSON table search cost) and the search cost in the key group table with an index on the key (key group table search cost), and determines whether the difference between the JSON table search cost and the key group table search cost is a certain amount or more.

[0121] As a result, when it is determined that the search time cannot be shortened by a certain amount or more by searching the key group table rather than searching the JSON table 24 (step S34: No), there is no need to generate a key group table. Therefore, the key group table generation unit 13 ends the processing of loop 2 for the target key group without generating a key group table for the target key group.

[0122] On the other hand, when it is determined that the search time can be shortened by a certain amount or more by searching the key group table rather than searching the JSON table 24 (step S34: Yes), the key group table generation unit 13 refers to the JSON key update log 17 and the search key log 18 and determines whether the addition of the JSON key and the search key for the target key group is at a frequency of a certain amount or less recently (step S35).

[0123] As a result, when it is determined that the addition of the JSON key and the search key of the target key group is not at a frequency equal to or lower than a certain level recently (step S35: No), even if a key group table is generated, there may be a need for columns of keys that are not in the table, which means that the generated key group table may become useless. Therefore, the key group table generation unit 13 ends the process of loop 2 for the target key group without generating a key group table for the target key group.

[0124] On the other hand, when it is determined that the addition of the JSON key and the search key of the target key group is at a frequency equal to or lower than a certain level recently (step S35: Yes), the key group table generation unit 13 generates a key group table including the search key of the target key group as a column (step S36) and ends the process of loop 2 for the target key group. In this embodiment, the key group table does not include columns for keys that are not the search key of the target key group, and the number of columns in the table can be reduced. Note that for the generated key group table, the key values for the JSON data in the JSON table 24 are transferred by the data conversion / transfer unit 22.

[0125] The key group table generation unit 13 executes the process of loop 2 (steps S31 to S37) with the JSON key group not targeted in the JSON key group information 15 as the next target. When all JSON key groups in the JSON key group information 15 have been targeted, it exits loop 2 and ends the key group table generation process.

[0126] According to the above-described embodiment, by determining whether to generate a table in units of key groups and creating a key group table, the number of generated tables can be suppressed. Further, since the search key in the key group is included as a column and keys other than the search key are not used as columns, the number of columns in the table can be suppressed. Thereby, the search convenience and search efficiency for semi-structured data can be improved.

[0127] Next, a data management system according to the second embodiment will be described.

[0128] FIG. 20 is an overall configuration diagram of a data management system according to the second embodiment. In FIG. 20, components similar to those of the data management system 1 according to the first embodiment are denoted by the same reference numerals.

[0129] The data management system 1A includes a database integrated management server 70, a JSON database management server 80, and a relational database management server 90 instead of the database management server 20, and realizes the functions of the database management server 20 by the functions of the database integrated management server 70, the JSON database management server 80, and the relational database management server 90.

[0130] The database integrated management server 70, the JSON database management server 80, and the relational database management server 90 are connected to the network 60.

[0131] The database integrated management server 70 is configured by a computer such as a PC or a server (for example, computer 100), for example. The database integrated management server 70 performs a process of integrally managing the JSON database management server 80 and the relational database management server 90. The database integrated management server 70 includes a search query rewriting unit 71 and a search query log 25. Details of the search query rewriting unit 71 will be described later.

[0132] The JSON database management server 80 is composed of, for example, a computer such as a PC or a server (e.g., computer 100). The JSON database management server 80 manages JSON data using the JSON table 24. The JSON database management server 80 includes a JSON database management system 81 and the JSON table 24. Details of the JSON database management system 81 will be described later.

[0133] The relational database management server 90 is composed of, for example, a computer such as a PC or a server (e.g., computer 100). The relational database management server 90 manages data using the key group table group 26. The relational database management server 90 includes a relational database management system 91 and the key group table group 26. Details of the relational database management system 91 will be described later.

[0134] First, the insertion process of JSON data will be described.

[0135] FIG. 21 is a diagram for explaining the insertion process of JSON data according to the second embodiment.

[0136] The data insertion unit 41 of the data insertion server 40 converts the production result data obtained at a factory (not shown) etc. into JSON data, and sends an insertion query for inserting the JSON data into the JSON table 24 to the JSON database management system 81 of the JSON database management server 80. Here, the insertion query includes the JSON data. When receiving the insertion query, the JSON database management system 81 adds a new entry to the JSON table 24 and stores the JSON data included in the insertion query in the JSON24b of the entry.

[0137] Next, the data search process will be described.

[0138] FIG. 22 is a diagram for explaining the data search process according to the second embodiment.

[0139] The data search unit 31 of the data search server 30 sends a search query for retrieving desired data from the JSON table 24 to the search query rewriting unit 71 of the database integrated management server 70.

[0140] When receiving the search query, the search query rewriting unit 71 stores the received search query in the search query log 25. The search query rewriting unit 71 acquires information (table information) of key group tables in the key group table group 26 from the relational database management system 91. Here, the table information includes information for identifying key group tables existing in the key group table group 26 and information on keys (which are columns) managed by the key group tables.

[0141] Based on the acquired table information, the search query rewriting unit 71 rewrites the search query into a search query (rewritten search query) including a partial query for the key group table for performing a search using the key group table and the JSON table 24, and a partial query for the semi-structured data storage unit for searching the semi-structured data in the semi-structured data storage unit. When no key group table has been generated in the key group table group 26 or when the search query is not related to the key group table, the search query rewriting unit 71 sets the received search query as the rewritten search query as it is.

[0142] The search query rewriting unit 71 passes the rewritten search query to the relational database management system 91. The relational database management system 91 performs a search on the key group table group 26 based on the search content for the key group tables in the key group table group 26 of the rewritten search query. Further, the relational database management system 91 generates a search query (sub-query) indicating the search content for the JSON table 24 based on the search result for the key group table group 26 and the rewritten search query, and transmits the sub-query to the JSON database management system 81 of the JSON database management server 80. The JSON database management system 81 receives the sub-query, searches the JSON table 24 according to the sub-query, and returns the data of the search result (search-in-progress data) to the relational database management system 91.

[0143] The relational database management system 91 combines the search result for the key group table group 26 and the search-in-progress data received from the JSON database management system 81 to create the data of the search result (search data) for the rewritten search query, and returns the search data to the search query rewriting unit 71.

[0144] The search query rewriting unit 71 receives the search data from the relational database management system 91 and transmits the search data to the data search unit 31 which is the source of the search query. Thereby, the data search unit 31 can perform various processes using the search data.

[0145] Note that the present invention is not limited to the above-described embodiments, and can be appropriately modified and implemented without departing from the spirit of the present invention.

[0146] For example, in the above embodiment, in step S35 of the key group table generation process, it was determined whether the JSON keys and search keys of the target key group had occurred at a frequency equal to or lower than a certain level recently. However, the present invention is not limited to this. For example, it may be determined whether only the search keys of the target key group have occurred at a frequency equal to or lower than a certain level recently.

[0147] Also, in the above embodiment, it was determined whether to generate a key group table by making all the determinations in steps S33, S34, and S35. However, the present invention is not limited to this. For example, it may be determined whether to generate a key group table by omitting at least one of steps S33, S34, and S35.

[0148] Also, the data management system of the above embodiment is not limited to the above configuration. For example, at least a plurality of the plurality of servers may be realized by one computer. Also, the division of functions of each server is not limited to the above, and it is sufficient if the functions are realized by a processor of any computer included in the data management system.

[0149] Also, in the above embodiment, an example using JSON data was shown. However, the present invention is not limited to this. For example, other semi-structured data such as XML data may be used.

[0150] Also, in the above embodiment, part or all of the processing performed by the functional unit may be performed by a hardware circuit. Also, the program constituting the functional unit in the above embodiment may be installed from a program source. The program source may be a program distribution server or a storage medium (for example, a portable storage medium).

Explanation of Reference Numerals

[0151] 1,1A… Data management system, 10… Key group table generation server, 11… JSON key grouping section, 12… Search key extraction section, 13… Key group table generation section, 20… Database management server, 21… Database management system, 22… Data conversion / transfer section, 23… Search query rewriting section, 30… Data search server, 31… Data search section, 40… Data insertion server, 41… Data insertion section, 50… Client PC, 60… Network, 70… Database integrated management server, 71… Search query rewriting section, 80… JSON database management server, 90… Relational database management server

Claims

1. A data management system for managing semi-structured data, wherein the data management system includes one or more processors and a semi-structured data storage unit for storing the semi-structured data, and the processor divides a plurality of keys included in a plurality of data units of the semi-structured data into a plurality of key groups based on the relationships between the keys in the plurality of data units, identifies a search key for the semi-structured data, and for each of the plurality of key groups, when the search key is included in the key group, generates a key group table including a column of the search key and a column indicating the data unit data management system.

2. The processor sets the first key and the second key to be in the same key group when the similarity between a set of data units having the first key and a set of data units having a second key other than the first key is equal to or greater than a predetermined value The data management system according to claim 1.

3. The processor does not generate a key group table for the key group when the search time by the search key included in the key group is not equal to or greater than a predetermined value and there is no tendency for the search time to deteriorate The data management system according to claim 1.

4. The processor does not generate a key group table for the key group when the reduction in search time by searching the key group table having the search key included in the key group as a column is not expected to be equal to or greater than a predetermined value compared to searching the semi-structured data in the semi-structured data storage unit by the search key included in the key group The data management system according to claim 1.

5. The processor repeatedly performs the process of adding the key included in the new data unit of the semi-structured data to the key group and adding a new search key for the semi-structured data, and does not generate a key group table for the key group when the addition of keys and search keys to the key group is not less than a certain frequency The data management system according to claim 1.

6. The processor receives a designation of a key for generating a key group table from a user, and generates a key group table having the received key as a column The data management system according to claim 1.

7. The processor Receive a search query for the semi-structured data in the semi-structured data storage unit, Convert the search query into a rewritten search query including a partial query for the key group table related to the search keys stored in the key group table and a partial query for the semi-structured data storage unit for searching the semi-structured data in the semi-structured data storage unit, Perform a search on the key group table using the partial query for the key group table, perform a search on the semi-structured data storage unit based on the partial query for the semi-structured data storage unit, and collectively return the search results to the source of the search query The data management system according to claim 1.

8. A data management method by a data management system for managing semi-structured data, The data management system includes a semi-structured data storage unit for storing the semi-structured data, Divide a plurality of keys included in a plurality of data units of the semi-structured data into a plurality of key groups based on the relationship between the keys in the plurality of data units, Specify a search key for the semi-structured data, For each of the plurality of key groups, when the search key is included in the key group, generate a key group table including a column of the search key and a column indicating the data unit Data management method.

Citation Information

Patent Citations

  • Structured document management device, structured document sub management device, program and structured document managing method

    JP2007265248A

  • Structured document management system and program

    JP2008052662A

  • Method and device for XML data storage, query rewrite, visualization, mapping, and reference

    JP2011181106A

  • XML document processing

    US20020123993A1

  • Schema-less access to stored data

    US20150088924A1