Data processing system, data processing method, and data processing program

The data processing system addresses the challenge of balancing data creator flexibility and user efficiency by generating semi-structured data from tabular data, enabling efficient data utilization and search functionality.

JP2025080081AInactive Publication Date: 2025-05-23DAIKIN INDUSTRIES LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023193085
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing data processing systems face challenges in allowing data creators to freely select spreadsheet data formats while ensuring data users can efficiently extract necessary information, and vice versa.

Method used

A data processing system that includes a processor to obtain and store tabular data, generate semi-structured data by defining element relationships, and output this data in a format that can be efficiently utilized by users.

Benefits of technology

The system enables efficient data utilization by allowing data creators to maintain flexibility in data format while ensuring data users can easily extract and search the necessary information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080081000001_ABST
    Figure 2025080081000001_ABST
Patent Text Reader

Abstract

To solve a problem in which a user of data created by a creator may have difficulty in extracting necessary information from the data, the creator being allowed to freely select a format of the spreadsheet data.SOLUTION: A data processing system 10 includes a first processor 22, a first storage device 24, a second input device 36, and a second output device 38. The first processor 22 is configured to: (a) acquire an input file including tabular data to be stored in a first storage device 24; (b) output an interface to the second output device 38, for inputting first information that defines element relationships between a plurality of elements in the tabular data; (c) generate semi-structured data including the element relationships defined by the input first information from the tabular data included in the input file stored in the first storage device 24; and (d) generate an output file including the generated semi-structured data.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a data processing system, a data processing method, and a data processing program. [Background technology]

[0002] As disclosed in Patent Document 1 (JP 2017-146923 A), a device is known that converts data in a spreadsheet format having a predefined format into data in a semi-structured data format. Summary of the Invention [Problem to be solved by the invention]

[0003] If the creator of data is allowed to freely select the format of spreadsheet-style data, it may be difficult for users of the created data to extract necessary information from the data. On the other hand, if the format of spreadsheet-style data is predefined for the convenience of data users, it may be difficult for data creators to smoothly carry out their data creation work. [Means for solving the problem]

[0004] The data processing system according to a first aspect includes a processor, a storage device, an input device, and an output device. The processor performs the following processes. · Obtain an input file containing tabular data. The acquired input file is stored in a storage device. Outputting, to an output device, an interface for inputting, via an input device, first information defining element relationships between a plurality of elements in the tabular data. From tabular data included in an input file stored in a storage device, semi-structured data including element relationships defined by first information input via an input device is generated. Generate an output file containing the generated semi-structured data.

[0005] The data processing system according to the first aspect can realize efficient utilization of data by data users.

[0006] A data processing system according to a second aspect is the data processing system according to the first aspect, wherein the processor further performs the following processes. Generate semi-structured data from tabular data contained in multiple input files. Generate one output file containing multiple generated semi-structured data.

[0007] The data processing system according to the second aspect can realize efficient utilization of data by data users.

[0008] A data processing system according to a third aspect is the data processing system according to the first or second aspect, in which the processor generates an output file further including metadata of tabular data corresponding to the semi-structured data included in the output file, The metadata includes at least one of the name of the input file, the registration date of the input file, the update date of the input file, the creation date of the input file, and the last update date of the input file.

[0009] The data processing system according to the third aspect can realize efficient utilization of data by data users.

[0010] A data processing system according to a fourth aspect is the data processing system according to any one of the first to third aspects, in which the processor further performs the following processes. · Store the generated semi-structured data in a storage device. An interface for inputting a first search condition related to element relationships included in the semi-structured data via an input device is output to an output device. Semi-structured data that satisfies a first search condition input via an input device is extracted from the semi-structured data stored in the storage device. Generate an output file that contains the semi-structured data that meets the first search criteria.

[0011] The data processing system according to the fourth aspect can improve the efficiency of searching tabular data in any format.

[0012] A data processing system of a fifth aspect is any one of the data processing systems of the first to fourth aspects, in which the processor further generates an output file containing the generated semi-structured data as data represented in a two-dimensional table.

[0013] The data processing system according to the fifth aspect can realize efficient utilization of data by data users.

[0014] A data processing system of a sixth aspect is any one of the data processing systems of the first to fifth aspects, wherein the processor further generates an output file further including tabular data corresponding to the semi-structured data included in the output file.

[0015] The data processing system according to the sixth aspect can realize efficient utilization of data by data users.

[0016] A data processing system according to a seventh aspect is the data processing system according to any one of the first to sixth aspects, in which the element relationship is at least one of the following relationships: The first relationship is defined as a key-value pair. · A second relationship defined as multiple pairs that are connected in series. · The third relation is defined as a table containing pairs of elements arranged in a matrix. · The fourth relationship, defined as two or more elements connected in a given order.

[0017] A key is an element. A value is an element or a table. A table includes a first region consisting of a plurality of elements arranged in a matrix. Each element of the first region is a value.

[0018] A data processing system according to an eighth aspect is the data processing system according to any one of the first to seventh aspects, in which the tabular data is a spreadsheet and the elements are cells of the spreadsheet.

[0019] The data processing system according to the eighth aspect can improve the degree of freedom for the data creator in inputting data.

[0020] A data processing system according to a ninth aspect is the data processing system according to any one of the first to eighth aspects, in which the semi-structured data is described in JSON.

[0021] The data processing system according to the ninth aspect can realize efficient utilization of data by data users.

[0022] A data processing method according to a tenth aspect is used in a data processing system including a processor, a storage device, an input device, and an output device, and executes the following steps. Obtaining an input file containing tabular data. Storing the acquired input file in a storage device. A step of outputting, to an output device, an interface for inputting, via the input device, first information defining element relationships between a plurality of elements in the tabular data. A step of generating semi-structured data including element relationships defined by first information input via the input device from tabular data included in an input file stored in a storage device. Generating an output file containing the generated semi-structured data.

[0023] A data processing program according to an eleventh aspect is used in a data processing system including a processor, a storage device, an input device, and an output device, and realizes the following functions. Ability to take input files containing tabular data. The ability to store acquired input files in a storage device. A function for outputting, to an output device, an interface for inputting, via an input device, first information defining element relationships between a plurality of elements in the tabular data. A function for generating semi-structured data including element relationships defined by first information input via an input device from tabular data included in an input file stored in a storage device. The ability to generate an output file containing the generated semi-structured data. [Brief description of the drawings]

[0024] [Figure 1] 1 is an overall configuration diagram of a data processing system 10 according to an embodiment. [Diagram 2] FIG. 2 is a block diagram of a first processor 22 and a second processor 32 according to the embodiment. [Diagram 3] 2 is a schematic diagram of a main window W according to an embodiment. [Figure 4] 11 is a screen transition diagram of the main screen P2 in the embodiment. FIG. [Diagram 5] 13 is an example of a filer screen D1 according to the embodiment. [Figure 6] 13 is an example of a perspective screen D2 according to an embodiment. [Figure 7] FIG. 2 is a diagram illustrating a cell parser and a header connector of an embodiment. [Figure 8] FIG. 2 is a diagram illustrating a cell parser and a sequence connector according to an embodiment. [Figure 9] FIG. 2 is a diagram illustrating a table parser according to an embodiment. [Figure 10] 7 is a diagram illustrating a state after a perspective object has been placed on the perspective screen D2 shown in FIG. 6 by a perspective operation. [Figure 11] 13 is an example of a spreadsheet after performing a header connection, which is one of the parsing operations of the embodiment. [Figure 12] 13 is an example of a spreadsheet after execution of continuous heading connection, which is one of the parsing operations of the embodiment. [Figure 13]13 is an example of a spreadsheet after table parsing, which is one of the parsing operations of the embodiment, is executed. [Figure 14] 13 is an example of a spreadsheet after execution of an order connection, which is one of the parsing operations of the embodiment. [Figure 15] 13 illustrates an example of a search screen D3 according to the embodiment. [Figure 16] 13 illustrates an example of a search condition panel D3a in an item-specified search according to the embodiment. [Figure 17] 13 is an example of an element relationship that is a target of an item-specified search according to the embodiment. [Figure 18] 13 is an example of an input pattern in an item-specified search according to the embodiment. [Figure 19] 13 is an example of an element relationship that is a target of an item-specified search according to the embodiment. [Figure 20] 13 is an example of an input pattern in an item-specified search according to the embodiment. [Figure 21] 13 illustrates an example of a search condition panel D3a for a procedure search according to the embodiment. [Figure 22] FIG. 11 is a diagram illustrating the function of a first check box B11 according to the embodiment. [Figure 23] 13 illustrates an example of a search condition panel D3a for a metadata search and a full-text search according to the embodiment. [Figure 24] 13 is an example of semi-structured data displayed in a preview area D2a according to an embodiment. [Diagram 25] 13 is an example of semi-structured data displayed in a preview area D2a according to an embodiment. [Figure 26] FIG. 4 is a state transition diagram of an input file according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] --Embodiment-- (1) Overall Configuration of Data Processing System 10 1, a data processing system 10 includes a server 20 and a client 30. The server 20 and the client 30 are communicatively connected to each other via a network 40. The number of servers 20 may be one or more. The number of clients 30 may be one or more. In the following description, the number of servers 20 and clients 30 is each assumed to be one.

[0026] The server 20 is a computer including a first processor 22, a first memory 23, a first storage device 24, a first input device 26, and a first output device 28. The server 20 is, for example, a workstation or a personal computer on which a server operating system is installed. The first processor 22 is, for example, a CPU or a GPU. The first memory 23 is a main storage device that the first processor 22 can directly access. The first storage device 24 is an auxiliary storage device such as a hard disk drive (HDD) or a solid state drive (SSD). The first input device 26 is, for example, a keyboard or a mouse. The first output device 28 is, for example, a display.

[0027] The client 30 is a computer including a second processor 32, a second memory 33, a second storage device 34, a second input device 36, and a second output device 38. The client 30 is, for example, a personal computer on which a client operating system is installed, or a mobile terminal such as a tablet. The second processor 32 is, for example, a CPU or a GPU. The second memory 33 is a main storage device that the second processor 32 can directly access. The second storage device 34 is an auxiliary storage device such as an HDD or SSD. The second input device 36 is, for example, a keyboard, a mouse, or a touch panel. The second output device 38 is, for example, a display or a touch panel.

[0028] The network 40 is made up of various hardware components that enable the server 20 and the client 30 to communicate with each other using a predetermined protocol. The network 40 includes a network adapter, a LAN cable, a router, a switching hub, an optical cable, and the like.

[0029] The first storage device 24 of the server 20 stores computer software programs for executing various processes and data processed by the programs. The first processor 22 of the server 20 loads the programs and data stored in the first storage device 24 into the first memory 23 and processes the data by executing the programs. The first processor 22 has a functional unit that provides a predetermined function realized by executing one or more programs. The second processor 32, the second memory 33, and the second storage device 34 of the client 30 have the same functions as the first storage device 24, the first memory 23, and the first storage device 24, respectively. The first processor 22 and the second processor 32 each have one or more functional units. FIG. 2 shows the functional units of the first processor 22 and the second processor 32.

[0030] (2) Functions of Data Processing System 10 The data processing system 10 is a system that can be used by users belonging to a specific organization, such as an in-house network system. Specifically, the data processing system 10 is used in an organization to which both those who create data and those who use data belong. In this case, the users of the data processing system 10 are mainly classified into creators A and searchers S. When the data processing system 10 is used in a company's research and development site, the creator A is, for example, a researcher who creates an experiment plan and an experiment result report. The searcher S is a data scientist who analyzes the data created and accumulated by the creator A to extract useful information.

[0031] The data processing system 10 mainly has a filer function, a parse function, a search function, and an output function. These functions are realized by the first processor 22 and the second processor 32 executing predetermined programs. The filer function and the parse function are normally used by the creator A. The search function is normally used by the creator A and the searcher S. The output function is normally used by the searcher S.

[0032] A user of the data processing system 10 starts a predetermined data processing application installed in the client 30. After starting the data processing application, the user operates the client 30 using the second input device 36 and the second output device 38 to realize each function of the data processing system 10. When the user starts the data processing application, a main window W shown in FIG. 3 is displayed on the second output device 38. The main window W includes a side panel P1 and a main screen P2. A filer button B1 and a search button B2 are displayed on the side panel P1. Any one of a filer screen D1, a purse screen D2, and a search screen D3 is displayed on the main display area R0 of the main screen P2. A filer tab T1 is displayed on the top of the main screen P2. A purse tab T2 is further displayed on the top of the main screen P2 by a predetermined operation. When the filer tab T1 is selected, the filer screen D1 is displayed on the main screen P2. When the purse tab T2 is selected, the purse screen D2 is displayed on the main screen P2. The file tab T1 displays the name of the current folder, and the parse tab T2 displays the name of the file being parsed (described later).

[0033] FIG. 4 shows a screen transition diagram of the main screen P2. When the user starts the data processing application, the filer screen D1 is displayed with the filer tab T1 selected. When the parse tab T2 is displayed and the parse tab T2 is selected, the parse screen D2 is displayed. When the filer tab T1 is selected in this state, the filer screen D1 is displayed. When the search button B2 is clicked with the filer screen D1 displayed, the search screen D3 is displayed. When the filer button B1 is clicked with the search screen D3 displayed, the filer screen D1 is displayed. When the parse screen D2 is displayed, the search button B2 is disabled. In this embodiment, the action of clicking an interface such as a button displayed on the screen may be replaced with an action of tapping the interface, or another action of selecting the interface and starting a predetermined process.

[0034] When implementing the filer function, the user causes the filer screen D1 to be displayed. When implementing the parse function, the user causes the parse screen D2 to be displayed. When implementing the search function or output function, the user causes the search screen D3 to be displayed. Next, details of each function of data processing system 10 will be described.

[0035] (2-1) Filer function The filer function is a function for registering, managing, and viewing data created by creator A on a folder or file basis. The data created by creator A is tabular data. Tabular data is, for example, a spreadsheet created by spreadsheet software such as Excel (registered trademark). Tabular data has multiple elements arranged two-dimensionally. When the tabular data is a spreadsheet, an element is a cell having row and column identifiers that indicate a position on the spreadsheet. Tabular data has one or more spreadsheets.

[0036] To realize the filer function, the second processor 32 has, as functional units, a creation unit 91 and a management unit 92. The creation unit 91 is executed, for example, by spreadsheet software installed in the client 30. The management unit 92 is executed, for example, by file management software installed in the client 30.

[0037] The creation unit 91 creates tabular data and stores it in the second storage device 34. The creation unit 91 allows creator A to operate the client 30 to create an input file including tabular data. Creator A stores the created input file in the second storage device 34. Also, the creation unit 91 allows creator A to operate the client 30 to edit the tabular data included in the input file. Creator A may obtain an input file from an external storage device or an external network and store it in the second storage device 34. The input file is, for example, an Excel (registered trademark) file.

[0038] The management unit 92 manages the input files stored in the second storage device 34 under a folder tree. Creator A operates the client 30 using the management unit 92 to move and copy input files between folders, delete input files, and the like. The management unit 92 also manages metadata of the input files. The metadata of the input files is, for example, the last update date and time of the input file, the person who last updated it, the file size, and access rights.

[0039] The management unit 92 displays a filer screen D1 for the creator A to manage input files on the second output device 38. An example of the filer screen D1 is shown in Fig. 5. The filer screen D1 includes a folder tree panel D1a and a file list panel D1b.

[0040] The folder tree panel D1a visually displays a hierarchical structure of folders stored in a predetermined storage area of ​​the second storage device 34. The user operates folders in the folder tree panel D1a. Folder operations include, for example, creating, moving, deleting, copying, and selecting folders.

[0041] The file list panel D1b displays a list of input files when the folder selected in the folder tree panel D1a includes input files. In FIG. 5, the file list panel D1b displays a list of input files included in "Folder 2-2" selected in the folder tree panel D1a, and "File 5" is selected as an input file in the file list panel D1b. The file list panel D1b may display metadata of the input file in addition to the name of the input file. The user operates the input file in the file list panel D1b. The input file operations include, for example, creating, deleting, copying, and selecting the input file. The user can move the input file displayed in the file list panel D1b to any folder by dragging and dropping the input file to the folder displayed in the folder tree panel D1a.

[0042] The filer screen D1 may further display a "Parse" button C11, a "New Folder" button C12, and a "Download" button C13, as shown in Fig. 5. The "Parse" button C11 has a function of starting a parse operation, described later, for an input file currently selected in the file list panel D1b. The "New Folder" button C12 has a function of creating a new folder under a folder currently selected in the folder tree panel D1a. The "Download" button C13 has a function of acquiring an input file currently selected in the file list panel D1b, which is an input file uploaded to the server 20, as described later.

[0043] (2-2) Perspective function The parsing function is a function that generates semi-structured data from tabular data and centrally stores the generated semi-structured data. Semi-structured data is data whose data structure is defined to some extent among unstructured data whose data structure is not clearly defined. In this embodiment, the semi-structured data is text data written in JSON (JavaScript (registered trademark) Object Notation).

[0044] To realize the parsing function, the first processor 22 has, as functional units, an acquisition unit 41, a first storage unit 42, a first output unit 43, a first generation unit 44, and a second storage unit 45. These functional units are realized by software installed in the server 20.

[0045] The acquisition unit 41 acquires tabular data for generating semi-structured data by the parsing function. To this end, the creator A operates the client 30 to upload an input file including the tabular data from the client 30 to the server 20 in advance. Specifically, the creator A selects an input file to be uploaded on the filer screen D1, opens a context menu in this state, and selects an item "upload", thereby starting uploading of the input file. Alternatively, the creator A may start uploading of the input file by dragging and dropping the input file to be uploaded into a predetermined area of ​​the filer screen D1. When uploading of the input file starts, the input file is transmitted to the server 20 via the network 40. The acquisition unit 41 acquires the input file transmitted from the client 30 to the server 20.

[0046] The first storage unit 42 stores the tabular data acquired by the acquisition unit 41 in the first storage device 24. When the input file is transmitted to the server 20 by uploading the input file, the first storage unit 42 stores the input file transmitted to the server 20 in the first storage device 24.

[0047] The first output unit 43 outputs to the second output device 38 an interface for inputting, via the second input device 36, first information that defines the relationships between multiple elements in the tabular data stored in the first storage device 24. Creator A operates the client 30 to perform a parsing operation to input the first information. Hereinafter, the relationships between multiple elements defined by the first information will be referred to as element relationships. Specific examples of the parsing operation and details of the first information and element relationships will be described later.

[0048] The first output unit 43 displays a parse screen D2 for creator A to perform a parse operation on the second output device 38 as an interface for inputting the first information. Specifically, creator A selects an input file to be parsed on the filer screen D1, and opens a context menu and selects an item "Open parse screen" or selects the "Parse" button C11 at the bottom of the filer screen D1 to display the parse screen D2 on the second output device 38. Creator A performs a parse operation on the tabular data included in the selected input file on the parse screen D2.

[0049] The first generating unit 44 generates semi-structured data including element relationships defined by the first information input via the second input device 36 from the tabular data stored in the first storage device 24. After completing the parsing operation on the tabular data, creator A clicks the register button B3 displayed on the parse screen D2, and an input file including tabular data reflecting the first information input by the parse operation is registered. When the input file is registered, the first generating unit 44 automatically generates semi-structured data including element relationships based on the first information input by the parse operation of the tabular data.

[0050] The second storage unit 45 stores the semi-structured data generated by the first generation unit 44 in the first storage device 24. The second storage unit 45 centrally accumulates the semi-structured data generated from the tabular data included in the registered input file in a database stored in the first storage device 24. The second storage unit 45 stores the tabular data in the first storage device 24 in association with the corresponding semi-structured data.

[0051] Next, a specific example of the parse operation and details of the first information and element relationships will be described. In the following, it is assumed that the tabular data is a spreadsheet, and the first information defines the element relationships, which are relationships between multiple cells in the spreadsheet. The first information input by the parse operation relates to a parse object for associating the relationships between multiple cells. The parse operation is an operation in which creator A arranges parse objects on a spreadsheet displayed on the parse screen D2. Information regarding the arrangement of all parse objects arranged on a certain spreadsheet is called a parse layout. Creator A performs a parse operation to change the parse layout by adding, moving, copying, deleting, etc., parse objects on the spreadsheet.

[0052] Fig. 6 shows an example of a perspective screen D2 before a perspective operation is performed. A spreadsheet SS on which no perspective objects are arranged is displayed on the perspective screen D2 in Fig. 6. This spreadsheet SS represents an experiment plan created by a creator A. The format of the spreadsheet may be freely determined by each creator A.

[0053] (2-2-1) Perspective Object There are four types of parse objects: cell parser, table parser, header connector, and order connector. Figures 7 to 9 show examples of these parse objects arranged on a spreadsheet. Figure 10 shows an example of a parse object arranged on the parse screen D2 of Figure 6.

[0054] (2-2-1-1) Cell Parser A cell parser specifies the location of a single cell on a spreadsheet and is shown in Figures 7 and 8 as a thick line surrounding a single cell.

[0055] A cell parser is not placed alone on a spreadsheet. A cell parser is placed connected to a header connector or an order connector. If all header connectors and order connectors connected to a cell parser are deleted, the cell parser is automatically deleted.

[0056] (2-2-1-2) Table Parser The table parser specifies the area in which multiple cells are arranged in a rectangular shape (matrix shape) on the spreadsheet. In other words, the table parser specifies the table range that is considered to be a table. In Figure 9, the table parser is shown as a thick line surrounding multiple cells arranged in a rectangular shape.

[0057] The table range specified by the table parser has a header area R1, an index area R2, and a value area R3, as shown in FIG. 9. The header area R1, the index area R2, and the value area R3 are each a rectangular area. The header area R1 is located above the value area R3. The number of columns in the header area R1 is the same as the number of columns in the value area R3. The index area R2 is located to the left of the value area R3. The number of rows in the index area R2 is the same as the number of rows in the value area R3. The header area R1 may have multiple rows. The index area R2 may have multiple columns.

[0058] 9, header area R1 is a rectangular area of ​​cells C2 to E3. Index area R2 is a rectangular area of ​​cells B4 to B8. Index area R2 is a rectangular area of ​​cells C4 to E8.

[0059] The table range may not have either the header area R1 or the index area R2. Some cells in the header area R1 and the index area R2 may be combined. In FIG. 9, the cells in the hatched area in the upper left of the table range are not included in the element relationship.

[0060] (2-2-1-3) Header Connector A header connector connects two perspective objects on a spreadsheet with a header relationship. In Figure 7, header connectors are shown as solid arrows.

[0061] The header connector connects a first cell parser to a second cell parser different from the first cell parser. When the header connector starts at the first cell parser and ends at the second cell parser, data stored in a cell specified by the first cell parser becomes a header for referencing data stored in a cell specified by the second cell parser.

[0062] Also, a header connector connects a cell parser and a table parser. A header connector starts from a cell parser and ends at a table parser. In this case, data stored in a cell specified by a cell parser becomes a header for referencing data stored in a cell specified by a table parser. A header connector cannot start from a table parser.

[0063] As shown in Fig. 7, three or more parse objects may be connected in series by multiple header connectors. In this case, the parse objects other than the terminal parse object are cell parsers and are not table parsers.

[0064] A heading connector cannot start at one perspective object and end at multiple perspective objects. A heading connector cannot start at multiple perspective objects and end at a single perspective object.

[0065] Two perspective objects connected by a header connector may be adjacent or separate on a spreadsheet. If two perspective objects connected by a header connector are adjacent, the header connector is displayed on the border separating the two perspective objects, as in cells B8 and C8 in Figure 7.

[0066] (2-2-1-4) Order connector An order connector connects multiple cell parsers on a spreadsheet in a predefined order. In Figure 8, order connectors are represented by dashed arrows.

[0067] Cell parsers connected by an order connector are connected to only two cell parsers, except for the start cell parser and the end cell parser, which are connected to only one cell parser.

[0068] An order connector cannot be connected such that one cell parser is the starting point and multiple cell parsers are the ending point. An order connector cannot be connected such that multiple cell parsers are the starting point and one cell parser is the ending point.

[0069] Two cell parsers connected by an order connector may be adjacent or separate on a spreadsheet. When two cell parsers connected by an order connector are adjacent, the order connector is displayed on the border separating the two cell parsers, as shown in cells E2 to E5 in FIG.

[0070] (2-2-2) Perth Menu There are four types of parse operations: heading connection, consecutive heading connection, table parse, and order connection. In Figs. 6 and 10, four buttons C1 to C4 corresponding to each parse operation are displayed at the top of the parse screen D2. The "Key-Value" button C1 corresponds to heading connection. The "Multiple cell" button C2 corresponds to consecutive heading connection. The "Table" button C3 corresponds to table parse. The "Order" button C4 corresponds to order connection.

[0071] As shown in Figs. 6 and 10, the perspective screen D2 may further display a "Save" button C5, an "Undo" button C6, and a "Redo" button C7. The "Save" button C5 has a function of saving the perspective operation as first information. The "Undo" button C6 has a function of canceling the immediately previous perspective operation. The "Redo" button C7 has a function of redoing the most recently canceled perspective operation.

[0072] (2-2-2-1) Heading connection A header connection is a parser operation that connects two parse objects with a header connector. When creator A clicks the "Key-Value" button C1, the parse screen D2 transitions to a header connection mode in which header connection can be executed as a parse operation.

[0073] In the header connection mode, creator A performs a parser operation to specify an element relationship defined as a pair of a key and a value. Specifically, creator A first specifies a cell corresponding to the key. Next, creator A specifies a cell or table parser corresponding to the value. When a cell corresponding to the value is specified, a header connector that connects a specified cell corresponding to the key and a specified cell corresponding to the value in a header relationship is automatically inserted in the spreadsheet. When a table parser corresponding to the value is specified, a header connector that connects a specified cell corresponding to the key and a specified table parser corresponding to the value in a header relationship is automatically inserted in the spreadsheet. As shown in FIG. 11, the header connector is represented as an arrow that starts from a cell corresponding to the key and ends at a cell or table parser corresponding to the value. Creator A cannot specify a table parser corresponding to the key. When a cell corresponding to the key or value is specified, a cell parser corresponding to the specified cell is automatically inserted.

[0074] Creator A specifies the cell or table parser that corresponds to the key or value by two clicks or one drag. When specifying by a click, the first clicked cell is specified as the cell that corresponds to the key, and the second clicked cell or table parser is specified as the cell or table parser that corresponds to the value, respectively.

[0075] When specifying by dragging, the cell over which the mouse is placed when dragging starts is specified as the cell corresponding to the key. While dragging, the cell or table parser over which the mouse is placed is highlighted. When dragging ends, the highlighted cell or table parser is specified as the cell or table parser corresponding to the value.

[0076] 11, creator A first clicks cell A1 and then clicks cells B1 to E1, or creator A drags from cell A1 to cells B1 to E1.

[0077] The parsing operation in the header connection mode defines first information. The first information includes information regarding the position of a cell corresponding to a key and the position of a cell or table parser corresponding to a value. The position of a cell is the row and column number of the cell. The position of a table parser is, for example, the positions of the top left cell and the bottom right cell of the table range of the table parser, and the positions of the top left cell and the bottom right cell of the value area R3 of the table parser.

[0078] (2-2-2-2) Continuous heading connection A continuous header connection is a parser operation that connects multiple parse objects consecutively with header connectors. When creator A clicks the "Multiple cell" button C2, the parse screen D2 transitions to a continuous header connection mode in which continuous header connection can be executed as a parse operation.

[0079] In the continuous header connection mode, creator A performs a parser operation to specify an element relationship defined as a plurality of pairs connected in series. Next, a specific parser operation in the case of continuous header connection of n different cells will be described. First, creator A specifies the first cell as the starting point. Next, creator A specifies the 2nd to n-1th cells in a predetermined order. Finally, creator A specifies the nth cell as the end point. When creator A specifies the tth cell (t is 2 to n), a header connector that connects the t-1th cell and the tth cell in a header relationship is automatically inserted on the spreadsheet. As shown in FIG. 12, the header connector is represented as an arrow starting from the t-1th cell and ending at the tth cell. Creator A can specify a table parser instead of the nth cell. Creator A cannot specify a table parser instead of the 1st to n-1th cells. When a cell is specified, a cell parser corresponding to the specified cell is automatically inserted.

[0080] Creator A specifies n cells or n-1 cells and one table parser connected by consecutive headers by multiple click actions or one drag action. In the case of specification by click action, the 1st to n-1th clicked cells are specified as the 1st to n-1th cells, respectively. After the 1st to n-1th cells are specified, the cell or table parser that is double-clicked is specified as the nth cell or table parser, respectively. Alternatively, when the 1st to n-1th cells are specified, and a cell or table parser is clicked, a predetermined operation may be further performed to specify the last clicked cell or table parser as the nth cell or table parser, respectively.

[0081] When specifying by dragging, the cell over which the mouse is placed when the dragging starts is specified as the 1st cell. The cell or table parser over which the mouse is placed when the dragging ends is specified as the nth cell or table parser. After the dragging ends, the cells located between the specified 1st cell and the specified nth cell, or between the specified 1st cell and the specified table parser, are automatically specified as the 2nd to n-1th cells. In this case, the 1st to nth cells must all be in the same row or the same column. Also, the table parser must have a cell in the same row or the same column as the 1st to n-1th cells.

[0082] When creating the perspective object shown in FIG. 12, creator A clicks cells A4, B4, D4, and E4 in this order.

[0083] The parsing operation in the continuous header connection mode defines first information. The first information includes information on the position of each cell in the continuous header connection or the position of the table parser in the continuous header connection. The first information further includes information on the order in which the cells in the continuous header connection are connected.

[0084] (2-2-2-3) Table Perth The table parse is a parser operation that places a table parser in a specified range on a spreadsheet. When creator A clicks the "Table" button C3, the parse screen D2 transitions to a table parse mode in which table parse can be executed as a parse operation.

[0085] In the table parsing mode, creator A performs a parser operation to specify the table range of the table parser and the value region R3 included in the table range. Specifically, creator A first specifies the table range, which is the region occupied by the entire table parser. Next, creator A specifies the value region R3, which is a part of the specified table range. After that, as shown in FIG. 13, the header region R1 and the index region R2 are automatically specified, and the table parser is automatically arranged. The table parser includes an element relationship defined as a pair of a key and a value. The element relationship included in the table parser is defined as a pair of a key stored in a cell of the header region R1 and a value included in each cell of the value region R3 in the same column as the header region R1. In addition, the element relationship included in the table parser is defined as a pair of a key stored in a cell of the index region R2 and a value included in each cell of the value region R3 in the same row as the index region R2.

[0086] Creator A specifies the table range of the table parser by dragging or clicking. When specifying by dragging, a rectangular range including the start cell that the mouse is over when the drag begins, and the end cell that the mouse is over when the drag ends is specified as the table range. While dragging, the rectangular range including the start cell and the cell that the mouse is over is highlighted. The start cell is the top left cell of the specified table range. The end cell is the bottom right cell of the specified table range. When specifying by clicking, the first cell clicked is specified as the start cell, and the second cell clicked is specified as the end cell.

[0087] Creator A specifies a value region R3 that is part of the specified table range by dragging or clicking in the same way as the table range. Creator A cannot specify an area that includes cells that are not included in the table range as value region R3.

[0088] When creating the parse object shown in Fig. 13, creator A first drags from cell B7 to cell F10, and then drags from cell B8 to cell F10. The table parser shown in Fig. 13 has a header region R1 and a value region R3, but does not have an index region R2.

[0089] The parsing operation in the table parsing mode defines first information, which includes information about the positions of the start and end cells of the table range of the table parser and the positions of the start and end cells of the value range R3 included in the table range.

[0090] (2-2-2-4) Sequence connection An order connection is a parser operation that connects multiple parse objects with an order connector. When creator A clicks the "Order" button C4, the parse screen D2 transitions to an order connection mode in which an order connection can be executed as a parse operation.

[0091] In the order connection mode, creator A performs a parser operation to specify an element relationship defined as a plurality of cells connected in a predetermined order. Next, a specific parser operation in the case of sequentially connecting n different cells from each other will be described. First, creator A specifies the first cell as the starting point. Next, creator A specifies the 2nd to n-1th cells in a predetermined order. Finally, creator A specifies the nth cell as the end point. When creator A specifies the tth cell (t is 2 to n), an order connector that connects the t-1th cell and the tth cell is automatically inserted on the spreadsheet. The order connector is represented as an arrow starting from the t-1th cell and ending at the tth cell, as shown in FIG. 14. Creator A can specify a table parser instead of the nth cell. Creator A cannot specify a table parser instead of the 1st to n-1th cells. When a cell is specified, a cell parser corresponding to the specified cell is automatically inserted.

[0092] Creator A designates n cells or n-1 cells and one table parser to be connected in order by multiple click actions or one drag action. In the case of designation by click action, the 1st to n-1th clicked cells are designated as the 1st to n-1th cells, respectively. After the 1st to n-1th cells are designated, the cell or table parser that is double-clicked is designated as the nth cell or table parser, respectively. Alternatively, after the 1st to n-1th cells are designated, when a cell or table parser is clicked, a predetermined operation may be further performed to designate the clicked cell or table parser as the nth cell or table parser, respectively.

[0093] When specifying by dragging, the cell over which the mouse is placed when the dragging starts is specified as the 1st cell. The cell or table parser over which the mouse is placed when the dragging ends is specified as the nth cell or table parser. After the dragging ends, the cells located between the specified 1st cell and the specified nth cell, or between the specified 1st cell and the specified table parser, are automatically specified as the 2nd to n-1th cells. In this case, the 1st to nth cells must all be in the same row or the same column. Also, the table parser must have a cell in the same row or the same column as the 1st to n-1th cells.

[0094] When creating the perspective object shown in FIG. 14, creator A clicks cells G2, G4, G6, G8 and G10 in this order.

[0095] The parsing operation in the sequential connection mode defines first information. The first information includes information on the position of each sequentially connected cell or the position of the sequentially connected table parser. The first information further includes information on the order in which the sequentially connected cells are connected.

[0096] (2-3) Search function The search function is a function for searching for input files that satisfy predetermined conditions among the input files uploaded to the server 20. The search function includes a function for extracting semi-structured data that satisfies predetermined conditions from semi-structured data that is generated from spreadsheets and centrally stored in the database of the first storage device 24. The search function further includes a function for extracting spreadsheets that satisfy predetermined conditions from spreadsheets that are associated with the semi-structured data and stored in the first storage device 24.

[0097] To realize the search function, the first processor 22 has, as functional units, a fifth output unit 51, a first extraction unit 52, a second extraction unit 53, a third extraction unit 54, a third storage unit 55, and a sixth output unit 56. These functional units are realized by software installed in the server 20.

[0098] The fifth output unit 51 outputs an interface for inputting search conditions related to the input file via the second input device 36 to the second output device 38. The searcher S operates the client 30 to perform a search operation for inputting search conditions related to the input file.

[0099] The search conditions related to the input file include a first search condition, a second search condition, and a third search condition. The first search condition relates to element relationships included in the semi-structured data. The second search condition relates to metadata of the input file. The third search condition relates to elements included in the input file. The first to third search conditions will be described in detail later.

[0100] The fifth output unit 51 displays a search screen D3 for the searcher S to perform a search operation on the second output device 38 as an interface for inputting the first to third search conditions. The searcher S can display the search screen D3 on the second output device 38 by clicking a search button B2 while the filer screen D1 is displayed on the second output device 38.

[0101] As shown in Fig. 15, the search screen D3 includes a search condition panel D3a, a search result panel D3b, and a preview panel D3c. The search condition panel D3a displays an interface for inputting the first to third search conditions, a search execution button B4, and a search folder text box E1. The search result panel D3b displays a list of input files extracted based on the first to third search conditions. The preview panel D3c displays semi-structured data corresponding to the input file selected in the search result panel D3b. The semi-structured data displayed in the preview panel D3c is, for example, text data described in JSON.

[0102] The first extraction unit 52 extracts semi-structured data that satisfies a first search condition input via the second input device 36 from the semi-structured data stored in the first storage device 24. The second extraction unit 53 extracts input files that satisfy a second search condition input via the second input device 36 from the input files stored in the first storage device 24. The third extraction unit 54 extracts input files that satisfy a third search condition input via the second input device 36 from the input files stored in the first storage device 24.

[0103] The third storage unit 55 stores the search results in the first storage device 24 based on the extraction results of the semi-structured data or input files by the first extraction unit 52, the second extraction unit 53, and the third extraction unit 54. The search results include information about the input files corresponding to the semi-structured data that satisfies the first search condition. The search results include information about the input files that satisfy the second search condition. The search results include information about the input files that satisfy the third search condition. The information about the input files is, for example, the name and metadata of the input files. The third storage unit 55 stores the search results related to the input files that satisfy the input search conditions among the first to third search conditions in the first storage device 24.

[0104] The sixth output unit 56 outputs the search results stored in the first storage device 24 to the second output device 38.

[0105] When the searcher S inputs at least one of the first to third search conditions in the search condition panel D3a of the search screen D3 and clicks the search execution button B4, the first extraction unit 52, the second extraction unit 53, and the third extraction unit 54 start the process of extracting the above-mentioned input files. After that, the third storage unit 55 stores the search results in the first storage device 24. The sixth output unit 56 displays a list of input files that satisfy the input search conditions in the search result panel D3b based on the search results. If no search conditions are input in the search condition panel D3a, the sixth output unit 56 displays all input files having spreadsheets from which semi-structured data has been generated in the search result panel D3b.

[0106] The searcher S can also input the path of the folder he / she wants to search in the search folder text box E1 of the search condition panel D3a. When the searcher S clicks the search execution button B4 with a path input in the search folder text box E1, only the input files stored in the folder corresponding to the input path will be searched. If the search folder text box E1 is empty, all input files stored in the first storage device 24 will be searched.

[0107] By using the search function, the searcher S can perform four types of searches, which will be described below: item-specific search, procedure search, metadata search, and full-text search. The searcher S may perform only one of these four types of searches, or may perform a combination of two or more types of searches.

[0108] (2-3-1) Search by item In an item-specified search, a searcher S inputs a first search condition related to element relationships included in the semi-structured data into a search condition panel D3a. In an item-specified search, the first search condition relates to element relationships defined as key-value pairs in the semi-structured data. Specifically, the first search condition specifies a key-value pair corresponding to two cell parsers connected by a header connector in a spreadsheet after a parsing operation that corresponds to the semi-structured data, or a key-value pair included in a table parser.

[0109] As shown in Fig. 16, the search condition panel D3a displays one or more pairs of an item text box E2 and a keyword text box E3. The searcher S inputs a character string to be searched for in both the item text box E2 and the keyword text box E3 of each pair as a first search condition. The searcher S inputs an "item" which is a character string corresponding to the key to be searched for in the item text box E2, and inputs a "keyword" which is a character string corresponding to the value to be searched for in the keyword text box E3. Wildcards can be used in the input values ​​of the item and keyword.

[0110] When the searcher S clicks the pull-down button B5 to the right of the item text box E2, a list of recently searched items or a list of pre-set items is displayed. The searcher S may select an item to be searched from the displayed list. When the searcher S clicks the clear button B6 to the right of the keyword text box E3, the keyword entered in the keyword text box E3 is deleted. When the searcher S clicks the first add button B7, a pair of the item text box E2 and the keyword text box E3 is added to the bottom. When the searcher S clicks the first delete button B8, the pair of the item text box E2 and the keyword text box E3 displayed at the bottom is deleted.

[0111] The first extraction unit 52 extracts, from the semi-structured data stored in the first storage device 24, semi-structured data including element relationships defined as pairs of a key that matches an item entered in the item text box E2 and a value that matches a keyword entered in the keyword text box E3.

[0112] Next, an explanation will be given of input patterns of items and keywords, which are conditions for items and keywords that can be input in the item text box E2 and the keyword text box E3 in an item-specified search.

[0113] Fig. 18 shows an example of an input pattern of items and keywords when the semi-structured data has the element relationship as shown in Fig. 17. In Fig. 17, multiple cell parsers are connected by header connectors. The items and keywords in the table in Fig. 18 correspond to the data stored in the cell parser in Fig. 17.

[0114] FIG. 20 shows an example of an input pattern of items and keywords when semi-structured data has the element relationship as shown in FIG. 19. In FIG. 19, one cell parser and one table parser are connected by a header connector. The items in the table in FIG. 20 correspond to the data stored in the header area R1 and index area R2 of the table parser in FIG. 19. The keywords in the table in FIG. 20 correspond to the data stored in the value area R3 of the table parser in FIG. 19.

[0115] The searcher S can input items and keywords of an input pattern for which the "Searchability" column is set to "Yes" in the tables of Figures 18 and 20 into the item text box E2 and keyword text box E3, respectively. However, the searcher S cannot input items and keywords of an input pattern for which the "Searchability" column is set to "No" into the item text box E2 and keyword text box E3, respectively.

[0116] Examples of parameters for a specific item search in which the search function extracts semi-structured data generated from a spreadsheet having the parse layout shown in FIG. 10 are as follows: - Item "Theme", keyword "Single crystal sample preparation" - Item "Experiment date and time", keyword "202? / 09 / *" (wildcard used) - Item "Space group", keyword "l4 / mcm" Item "No.", keyword "140"

[0117] (2-3-2) Procedure search In the procedure search, the searcher S inputs a first search condition related to an element relationship included in the semi-structured data into the search condition panel D3a. In the procedure search, the first search condition relates to an element relationship defined as a plurality of elements connected in a predetermined order in the semi-structured data. Specifically, the first search condition specifies the values ​​of cells corresponding to a plurality of cell parsers connected by order connectors in a spreadsheet after a parsing operation that corresponds to the semi-structured data, and the order of the values ​​of the cells.

[0118] On the search condition panel D3a, a plurality of step text boxes E4 are displayed in a lined up state as shown in Fig. 21. A character string to be searched is input in a predetermined order into each of the plurality of step text boxes E4. The searcher S inputs a character string to be searched in a predetermined order into each step text box E4 as a first search condition.

[0119] The first extraction unit 52 extracts, from the semi-structured data stored in the first storage device 24, semi-structured data that includes element relationships in which cells containing strings that match values ​​entered in multiple procedure text boxes E4 are connected in the same order as the order shown in the multiple procedure text boxes E4.

[0120] When the searcher S clicks the second add button B9, one step text box E4 is added to the bottom. When the searcher S clicks the second delete button B10, the step text box E4 displayed at the bottom is deleted. When there are two step text boxes E4, the second delete button B10 is disabled.

[0121] A first check box B11 is displayed on the search condition panel D3a. The searcher S can switch the first check box B11 on / off by clicking the first check box B11. By default, the first check box B11 is off.

[0122] When the first check box B11 is off, the first extraction unit 52 extracts only semi-structured data including element relationships in which the number of cells including strings matching the values ​​input in the multiple procedure text boxes E4 is the same. When the first check box B11 is on, the first extraction unit 52 also extracts semi-structured data including element relationships in which the number of cells including strings matching the values ​​input in the multiple procedure text boxes E4 is not the same. In other words, when the first check box B11 is on, even if any cell exists between two cells including strings matching the values ​​input in the two procedure text boxes E4 in the semi-structured data, the semi-structured data is subject to extraction as long as the order of the two cells matches the order of input in the procedure text boxes E4.

[0123] 22, which illustrates the function of the first check box B11, shows (a) input examples of multiple procedure text boxes E4, (b) a first example of an element relationship included in the semi-structured data, and (c) a second example of an element relationship included in the semi-structured data. When the first check box B11 is in an on state, the first extraction unit 52 extracts both the first example and the second example based on the input example. When the first check box B11 is in an off state, the first extraction unit 52 extracts only the first example based on the input example.

[0124] When the first check box B11 is off, an example of a procedure search parameter for extracting semi-structured data generated from a spreadsheet having the perspective layout shown in FIG. 10 using a search function is as follows: Step 1: Weighing, Step 2: Crushing Step 1: Grinding, Step 2: Mixing, Step 3: Vacuum sealing

[0125] When the first check box B11 is checked, an example of a procedure search parameter for extracting semi-structured data generated from a spreadsheet having the perspective layout shown in FIG. 10 using a search function is as follows: Step 1: Weighing, Step 2: Mixing Step 1: Crushing, Step 2: Vacuum sealing, Step 3: Firing

[0126] (2-3-3) Metadata search In the metadata search, the searcher S inputs a second search condition related to the metadata of the input file into the search condition panel D3a. The metadata of the input file includes at least one of the name of the input file, the name of the registrant of the input file, the registration date of the input file, the creation date of the input file, and the last update date of the input file. At least one of the start date and the end date is specified for the registration date of the input file, the creation date of the input file, and the last update date of the input file.

[0127] As shown in FIG. 23, the search condition panel D3a displays a file name text box E5, a registrant name text box E6, a date type selection list box E7, a start date designation date picker E8, and an end date designation date picker E9 for inputting the second search condition. The name of the input file is input in the file name text box E5. The name of the registrant of the input file is input in the registrant name text box E6. When the date type selection list box E7 is clicked, a list box for selecting one of the registration date of the input file, the creation date of the input file, and the last update date of the input file is displayed. When the start date designation date picker E8 is clicked, a date picker for designating the start date of the date selected in the date type selection list box E7 is displayed. When the end date designation date picker E9 is clicked, a date picker for designating the end date of the date selected in the date type selection list box E7 is displayed.

[0128] The second extraction section 53 extracts, from the input files stored in the first storage device 24, input files that satisfy the second search conditions inputted in the search condition panel D3a.

[0129] (2-3-4) Full-text search In a full-text search, the searcher S inputs a third search condition related to an element included in the input file into the search condition panel D3a. As shown in FIG. 23, the search condition panel D3a displays a full-text search text box E10 for inputting the third search condition. A character string that will be a condition for the full-text search is input into the full-text search text box E10. This character string may include at least one of an AND condition and an OR condition. Furthermore, a portion enclosed in double quotation marks in this character string is regarded as a phrase to be searched.

[0130] The search target of the full-text search is only the input file when the full-text search is used in combination with the item-specified search, procedure search, and metadata search. The search target of the full-text search is all files stored in the first storage device 24 when only the full-text search is used. When the search target of the full-text search is only the input file, the third search condition relates to, for example, an element contained in the input file. In this case, the third extraction unit 54 extracts input files that satisfy the third search condition entered in the search condition panel D3a from the input files stored in the first storage device 24.

[0131] The search condition panel D3a further displays a search start button B12. When the searcher S inputs the first to third search conditions in the search condition panel D3a and then clicks the search start button B12, a list of input files that contain semi-structured data that satisfies the first search condition and also satisfies the second and third search conditions is displayed in the search result panel D3b.

[0132] (2-4) Output function The output function is a function for outputting the semi-structured data generated from the spreadsheet to a file.

[0133] In order to realize the output function, the first processor 22 has, as functional units, a second generating unit 57 and a fourth storing unit 58. These functional units are realized by software installed in the server 20.

[0134] The second generating unit 57 generates an output file including the generated semi-structured data as text data. The second generating unit 57 outputs, for example, a file including text data described in JSON as the output file.

[0135] The searcher S selects an input file for which an output file is to be generated from the list of input files displayed in the search result panel D3b on the search screen D3. When the searcher S then clicks the output button B13 displayed on the search screen D3, the second generation unit 57 generates an output file from the semi-structured data generated from the spreadsheet included in the selected input file.

[0136] On the search screen D3, the searcher S selects an input file from the list of input files displayed on the search result panel D3b in a state where no search conditions are entered in the search condition panel D3a. In this case, the searcher S can select any input file from all input files having spreadsheets from which semi-structured data has been generated, and generate an output file.

[0137] Furthermore, the searcher S can input search conditions in the search condition panel D3a on the search screen D3 and click the search start button B12 to display a list of input files that satisfy the input search conditions in the search result panel D3b. In this case, the searcher S can select any input file from the input files displayed in the search result panel D3b and generate an output file.

[0138] The fourth storage unit 58 stores the output file generated by the second generation unit 57 in the first storage device 24.

[0139] The searcher S can utilize the output file stored in the fourth storage unit 58 by using it as input data for other applications installed in the client 30 or by saving it on an external storage medium.

[0140] (3) Other Functions of Data Processing System 10 (3-1) Other functions of the perspective function In order to realize the parsing function, the first processor 22 may further have, as functional units, a second output unit 46, a third output unit 47, a fifth storage unit 48, and a fourth output unit 49. These functional units are realized by software installed in the server 20.

[0141] The second output unit 46 outputs the semi-structured data stored in the first storage device 24 to the second output device 38. In this case, the second output unit 46 may output the spreadsheet stored in the first storage device 24 to the second output device 38 together with the semi-structured data generated from the spreadsheet and stored in the first storage device 24.

[0142] Specifically, while a spreadsheet parsing operation is being performed on the parse screen D2, the second output unit 46 displays a preview of semi-structured data generated from the spreadsheet that is the target of the parse operation in a partial area of ​​the parse screen D2. Every time an input file including a spreadsheet is registered by the first generating unit 44 and the second storing unit 45, semi-structured data is automatically generated and stored in the first storage device 24. Therefore, every time creator A registers an input file after the parse operation, the second output unit 46 updates the preview of the semi-structured data generated from the spreadsheet included in the input file.

[0143] 24 and 25, a preview of the semi-structured data is displayed, for example, in a preview area D2a occupying a part of the perspective screen D2 on the right side of the perspective screen D2. In this state, the area of ​​the perspective screen D2 excluding the preview area D2a is an area in which a spreadsheet on which a perspective operation is performed is displayed. Creator A can display the preview area D2a on the perspective screen D2 or delete it from the perspective screen D2 by, for example, selecting a "preview display" menu on the perspective screen D2.

[0144] As shown in FIG. 24, the semi-structured data displayed in the preview area D2a is listed for each type of parse object. When the parse object is a heading connector or an order connector, the position on the spreadsheet of the starting parse object, the data stored in the starting parse object, and the data stored in the ending parse object are displayed for each heading connector or order connector. When the parse object is a table parser, the positions of the starting cell and the ending cell of the table range of the table parser, and the positions of the starting cell and the ending cell of the value area R3 included in the table range are displayed for each table parser.

[0145] As shown in FIG. 25, the semi-structured data displayed in the preview area D2a may be displayed as text data described in JSON. In this case, in the preview area D2a, when the list display tab T3 is selected, the semi-structured data shown in FIG. 24 may be displayed, and when the JSON tab T4 is selected, the semi-structured data shown in FIG. 25 may be displayed.

[0146] The third output unit 47 outputs the element relationship defined by the first information input via the second input device 36 to the second output device 38. Specifically, each time the creator A performs a parsing operation, the third output unit 47 updates the parse layout displayed on the parse screen D2.

[0147] The fifth storage unit 48 stores the first information input via the second input device 36 in the first storage device 24. The fourth output unit 49 outputs an interface for changing the element relationship defined by the first information stored in the first storage device 24 via the second input device 36 to the second output device 38. Specifically, the fifth storage unit 48 stores the first information input by the creator A performing a parsing operation in the first storage device 24 to save the parse layout. The fourth output unit 49 reads out the saved parse layout onto the parse screen D2 so that the creator A can change the parse layout.

[0148] The fourth output unit 49 may further have a function of applying the read perspective layout to any spreadsheet displayed on the perspective screen D2. In this case, the creator A can place all the perspective objects included in the read perspective layout in the spreadsheet displayed on the perspective screen D2.

[0149] (3-2) Other output functions In order to realize the output function, the first processor 22 may further have, as functional units, a third generating unit 59, a fourth generating unit 60, a fifth generating unit 61, and a sixth generating unit 62. These functional units are realized by software installed in the server 20.

[0150] The third generation unit 59 generates one output file including a plurality of semi-structured data generated from the spreadsheets included in the plurality of input files. In other words, the third generation unit 59 generates one output file by collecting a plurality of text data described in JSON, for example.

[0151] The fourth generating unit 60 generates an output file that further includes metadata of the input file that includes a spreadsheet corresponding to the semi-structured data included in the output file. In other words, the fourth generating unit 60 generates an output file that includes both the metadata of the input file and the semi-structured data generated from the spreadsheet included in the input file.

[0152] The fifth generating unit 61 generates an output file including the semi-structured data generated from the spreadsheet included in the input file as data represented in a two-dimensional table. The semi-structured data included in the output file is, for example, text data written in CSV (comma-separated values). In this case, the table generated from the text data written in CSV is displayed in the preview panel D3c of the search screen D3.

[0153] The sixth generation unit 62 generates an output file that further includes a spreadsheet corresponding to the semi-structured data included in the output file. In other words, the sixth generation unit 62 generates an output file that includes both the spreadsheet included in the input file and the semi-structured data generated from the spreadsheet.

[0154] (3-3) Input file status management function The first processor 22 may further have a function of managing the state of an input file stored in the first storage device 24. In this case, the first processor 22, for example, limits a part of a search function for the input file depending on the state of the input file.

[0155] The state of an input file stored in the first storage device 24 will be described with reference to the transition diagram shown in FIG. 26. An input file can be in any one of the following states: "unregistered", "unregistered and being parsed", "registered and being parsed", and "registered". Depending on the state of the input file, the search function and preview display function are restricted. The state of each input file may be displayed as metadata on the filer screen D1.

[0156] When creator A uploads an input file to server 20, the status of the input file becomes "unregistered." When creator A selects an input file in the "unregistered" status on filer screen D1 and selects parse tab T2 to open parse screen D2, the status of the input file becomes "unregistered and parsing." When creator A then performs a parse operation on parse screen D2 and registers the input file after the parse operation is completed, the status of the input file becomes "registered and parsing."

[0157] If the input file's status is "unregistered and being parsed" and creator A selects filer tab T1 and closes parse screen D2, the input file's status will become "unregistered." If the input file's status is "registered and being parsed" and creator A selects filer tab T1 and closes parse screen D2, the input file's status will become "registered." If creator A selects an input file with a "registered" status on filer screen D1, selects parse tab T2 and opens parse screen D2, the input file's status will become "registered and being parsed."

[0158] When the status of an input file is "unregistered" or "unregistered and being parsed," searcher S can only execute full-text search among the search functions for that input file. When the status of an input file is "registered and being parsed" or "registered," searcher S can execute all search functions for that input file (item-specified search, procedure search, metadata search, and full-text search). In addition, when the status of an input file is "registered and being parsed" or "registered," searcher S can execute the preview display function of the input file on the filer screen D1 and the search screen D3.

[0159] When the state of the input file is "unregistered and being parsed" or "registered and being parsed", users other than the creator A of the input file can open the parse screen D2 to view the parse layout and copy the parse objects. However, users other than the creator A of the input file cannot open the parse screen D2 to perform a parse operation to change the parse layout.

[0160] (4) Features (4-1) The data processing system 10's tabular data parsing function enables it to generate semi-structured data from tabular data of any format with simple operations, thereby improving the freedom of data input for data creators and enabling data users to make efficient use of data.

[0161] The data processing system 10 has a function of displaying a preview of the semi-structured data on the second output device 38, enabling data users to utilize the data efficiently.

[0162] The data processing system 10, by using a function for displaying the perspective layout on the second output device 38, can improve the efficiency of the parsing work performed by the data creator.

[0163] The data processing system 10 can save a perspective layout and later recall and edit it, thereby improving the efficiency of parsing work by the data creator.

[0164] The data processing system 10 has a function of displaying tabular data together with semi-structured data on the second output device 38, thereby enabling data users to utilize data efficiently.

[0165] The data processing system 10 uses a spreadsheet as tabular data, thereby improving the degree of freedom for the data creator in inputting data.

[0166] The data processing system 10 uses text data described in JSON as semi-structured data, thereby enabling data users to make efficient use of data.

[0167] (4-2) The data processing system 10 uses the semi-structured data search function to search for relationships between multiple elements defined by the semi-structured data, thereby improving the efficiency of searching tabular data in any format.

[0168] The data processing system 10 can improve the efficiency of searching tabular data in any format by using the metadata search function of the input file.

[0169] The data processing system 10, by using a function for displaying a preview of search results on the second output device 38, can realize efficient utilization of data by data users.

[0170] The data processing system 10 enables data users to efficiently utilize data through item-specified searches and procedure searches of semi-structured data, as well as full-text search functions.

[0171] (4-3) The data processing system 10, by using the semi-structured data output function, enables data users to utilize data efficiently.

[0172] The data processing system 10 can realize efficient use of data by data users by using a function for consolidating semi-structured data generated from tabular data contained in multiple files into one file and outputting the file.

[0173] The data processing system 10 can realize efficient utilization of data by data users by using a function for outputting semi-structured data together with metadata of the input file.

[0174] The data processing system 10 can realize efficient use of data by data users by providing a function for outputting two-dimensional table data of semi-structured data as a file together with text data of semi-structured data.

[0175] The data processing system 10 can realize efficient utilization of data by data users by using a function for outputting semi-structured data together with an input file.

[0176] --Variations-- (1) Variation A In the embodiment, the first generating unit 44 generates semi-structured data based on first information that defines element relationships, which are relationships between a plurality of cells in a spreadsheet. The first generating unit 44 may generate semi-structured data based on first information that defines element relationships of elements other than cells of a spreadsheet. In other words, the first generating unit 44 may generate semi-structured data by treating elements other than cells of a spreadsheet in the same way as cells.

[0177] Elements that the first generating unit 44 can handle in the same way as a cell are, for example, image data, text boxes, and check boxes.

[0178] Image data is an image that can be embedded in a spreadsheet regardless of the cell position. Image data is, for example, a png file. In the parsing screen D2, creator A can handle image data in the same way as the cell parser. For example, creator A can connect image data with a header connector or an order connector, or embed image data in a cell in the value area R3 of the table parser.

[0179] A text box is text data that can be embedded in a spreadsheet regardless of the cell position. In the parsing screen D2, creator A can handle text boxes in the same way as a cell parser. For example, creator A can connect text boxes with a header connector or an order connector, or embed a text box in a cell in the value area R3 of a table parser.

[0180] A check box is a user interface that can take two states: a checked state and an unchecked state. A check box is placed in a cell and determines the state of data included in an adjacent cell. For example, when a check box is checked on the perspective screen D2, the first output unit 43 enables the creator A to treat the data included in the adjacent cell as a cell. On the other hand, when a check box is unchecked on the perspective screen D2, the first output unit 43 prevents the creator A from treating the data included in the adjacent cell as a cell.

[0181] The elements that the first generating unit 44 can handle in the same way as a cell are not limited, and may be graphic data, video data, and other user interfaces that are generally used in applications.

[0182] (2) Variation B In the embodiment, the sequential connection is a parser operation that sequentially connects a plurality of parse objects from a start point to an end point. However, in the sequential connection, a plurality of parse objects may be connected so as to branch or merge. Specifically, the first output unit 43 may enable the creator A to connect a plurality of cell parsers so that one cell parser is a start point and a plurality of cell parsers are end points on the parse screen D2. Similarly, the first output unit 43 may enable the creator A to connect a plurality of cell parsers so that a plurality of cell parsers is a start point and a plurality of cell parsers is an end point on the parse screen D2.

[0183] (3) Variation C To realize the search function, the first processor 22 may have a function of saving the first to third search conditions. For example, the third storage unit 55 may have a function of saving the first to third search conditions input by the searcher S on the search screen D3. In this case, the first processor 22 further has a function of reading out the first to third search conditions saved in the third storage unit 55 and displaying them on the search screen D3 when the searcher S performs a new search, without inputting the first to third search conditions.

[0184] (4) Variation D The data processing system 10 mainly has a filer function, a parsing function, a search function, and an output function, and may have other functions such as a login function, an API function, and a management function.

[0185] The login function is a function for authenticating users (creator A and searcher S) when they start using the client 30. The login function is realized, for example, by the management unit 92. The management unit 92 performs authentication using, for example, an ID and password assigned to each user. The management unit 92 displays the filer screen D1 on the second output device 38 only if the user authentication is successful. The user logs out when they finish using the client 30.

[0186] The API function is, for example, a function for transmitting information about input files that satisfy first to third search conditions from the client 30 to the server 20 to the client 30. The information about the input files is, for example, the name or ID of the input files.

[0187] The management function is a function for managing users and user groups. For example, the management function is creating a group, adding and deleting members of a group, and adding and deleting users. The management function may also be a function for limiting the number of logins, changing passwords, and outputting various logs.

[0188] Although the embodiments of the present disclosure have been described above, it will be understood that various changes in form and details can be made without departing from the spirit and scope of the present disclosure described in the claims. [Explanation of symbols]

[0189] 10: Data processing system 20: Server 22: First processor (processor) 24: 1st storage device (storage device) 26: First input device 28: First output device 30: Client 32: Second processor 34:Second storage device 36: Second input device (input device) 38: Second output device (output device) 40: Network [Prior art documents] [Patent documents]

[0190] [Patent Document 1] JP 2017-146923 A

Claims

1. A data processing system comprising a processor (22), a storage device (24), an input device (36), and an output device (38), The processor, Take an input file containing tabular data, storing the acquired input file in the storage device; outputting, to the output device, an interface for inputting, via the input device, first information defining relationships between a plurality of elements in the tabular data; generating semi-structured data including the relationship defined by the first information input via the input device from the tabular data included in the input file stored in the storage device; generating an output file containing the generated semi-structured data; A data processing system (10).

2. The processor, generating the semi-structured data from the tabular data included in each of the plurality of input files; generating one output file including the generated plurality of semi-structured data; 2. The data processing system of claim 1.

3. the processor generates an output file further comprising metadata for the tabular data corresponding to the semi-structured data included in the output file; The metadata includes: The name of the input file, The registration date of the input file; The update date of the input file; The creation date of the input file; and The last modification date of said input file; At least one of 3. The data processing system according to claim 1.

4. The processor, storing the generated semi-structured data in the storage device; outputting, to the output device, an interface for inputting, via the input device, a first search condition related to the relationship included in the semi-structured data; extracting the semi-structured data that satisfies the first search condition inputted via the input device from the semi-structured data stored in the storage device; generating the output file including the semi-structured data that satisfies the first search criteria; 3. The data processing system according to claim 1.

5. The processor generates the output file including the generated semi-structured data as data represented in a two-dimensional table.

3. The data processing system according to claim 1.

6. the processor generates an output file further comprising tabular data corresponding to the semi-structured data contained in the output file.

3. The data processing system according to claim 1.

7. The relationship is: A first relation defined as a key-value pair, a second relationship defined as a plurality of said pairs connected in series; a third relationship defined as a table including the pairs, the third relationship being composed of a plurality of the elements arranged in a matrix; and a fourth relationship defined as a plurality of said elements connected in a predetermined order; At least one of The key is the element, the value is the element or the table, the table includes a first region including a plurality of the elements arranged in a matrix; each said element of said first region is said value; 3. The data processing system according to claim 1.

8. the tabular data is a spreadsheet, the elements are cells of the spreadsheet; 3. The data processing system according to claim 1.

9. The semi-structured data is written in JavaScript Object Notation (JSON), 3. The data processing system according to claim 1.

10. A data processing method for use in a data processing system including a processor (22), a storage device (24), an input device (36), and an output device (38), comprising: The processor, obtaining an input file containing tabular data; storing the acquired input file in the storage device; outputting, to the output device, an interface for inputting, via the input device, first information defining relationships between a plurality of elements in the tabular data; generating semi-structured data including the relationship defined by the first information input via the input device from the tabular data included in the input file stored in the storage device; generating an output file containing the generated semi-structured data; Implementing data processing methods.

11. A data processing program for use in a data processing system including a processor (22), a storage device (24), an input device (36), and an output device (38), comprising: The processor, The ability to take an input file containing tabular data; A function of storing the acquired input file in the storage device; a function of outputting, to the output device, an interface for inputting, via the input device, first information defining relationships between a plurality of elements in the tabular data; a function of generating semi-structured data including the relationship defined by the first information input via the input device from the tabular data included in the input file stored in the storage device; generating an output file including the generated semi-structured data; A data processing program that realizes the above.

Citation Information

Patent Citations

  • Specification input / output device, specification input / output system, and specification input / output method

    JP2017146923A