Data search system, data search method, and data search program
The data search system addresses the challenge of identifying relevant data by calculating field adaptability through metadata and process usage analysis, ensuring accurate data selection based on user needs.
Patent Information
- Application Number
- JP2022134320
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-25
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2042-08-25
AI Technical Summary
Existing data search systems struggle to provide appropriate information about searched registered data, as they fail to determine whether the data meets user needs based on data reliability alone.
A data search system that calculates field adaptability by analyzing metadata and process usage information to identify suitable registered data, using a processor to search for data corresponding to specified fields and evaluate its adaptability.
The system provides appropriate information about searched registered data by determining its relevance and suitability for user needs, enhancing data selection accuracy.
Smart Images

Figure 0007819059000001 
Figure 0007819059000002 
Figure 0007819059000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique that can be used when searching for desired data from among multiple types of data. [Background technology]
[0002] Systems are being conceived and implemented to promote cross-sectoral business development and industrialization using various data. However, data comes in a wide variety of accuracy levels, making it extremely difficult to select data that meets the requirements of cross-sectoral applications.
[0003] Known data-related technologies include, for example, a technology that extracts valid values of similar data from registered data and metadata of valid data from similar data, taking into account market trends over a specified period of time (see, for example, Patent Document 1).
[0004] In addition, a technique is known for calculating the reliability of data to be used from the scores of data registrants and data viewers and the reliability of the data group used for data processing as a quantitative evaluation index for the degree of reliability of data to be used in the data distribution market (see, for example, Patent Document 2). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 2021-96792 [Patent Document 2] Patent Publication No. 2021-114077 Summary of the Invention [Problem to be solved by the invention]
[0006] For example, the technology disclosed in Patent Document 1 is said to extract valid data, but it is not easy to determine whether the data is what the user needs. Also, the technology disclosed in Patent Document 2 indicates data reliability, but it is not easy to determine whether the data is what the user needs based on this data reliability alone.
[0007] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a technique that can provide appropriate information about searched registered data. [Means for solving the problem]
[0008] In order to achieve the above-mentioned object, one aspect of a data search system is a data search system that searches for appropriate registered data from among a plurality of registered data, the data search system having a processor and a memory unit, the memory unit including the plurality of registered data, metadata including information about the fields of the plurality of registered data and the processes performed to obtain the registered data, and process usage information regarding the history of the use of one or more processes performed to obtain registered data for the fields of the plurality of registered data, the processor accepts specification of the field of registered data to be searched, searches for registered data corresponding to the field, and calculates a field adaptability for the registered data obtained by the search, based on the metadata of the registered data obtained by the search and the process usage information, which indicates the degree of adaptability in the field of the process performed to obtain the registered data obtained by the search. [Effects of the Invention]
[0009] According to the present invention, it is possible to provide appropriate information about the searched registration data. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a functional configuration diagram of a data search system according to an embodiment. [Figure 2] FIG. 2 is a hardware configuration diagram of a data search system according to an embodiment. [Figure 3] FIG. 3 is a diagram showing the configuration of a registered data metadata table according to an embodiment. [Figure 4] FIG. 4 is a configuration diagram of an accuracy type management table according to an embodiment. [Figure 5] FIG. 5 is a diagram illustrating the configuration of an accuracy calculation method management table according to an embodiment. [Figure 6] FIG. 6 is a diagram illustrating the configuration of a process management table according to an embodiment. [Figure 7] FIG. 7 is a diagram illustrating a configuration of a process field management table according to an embodiment. [Figure 8] FIG. 8 is a diagram showing a data field input screen according to one embodiment. [Figure 9] FIG. 9 is a diagram showing a search result screen according to one embodiment. [Figure 10] FIG. 10 is a flowchart of a process management process according to an embodiment. [Figure 11] FIG. 11 is a flowchart of a field adaptability calculation process according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] The following description of the embodiments will be given with reference to the drawings. Note that the embodiments described below do not limit the scope of the invention as claimed, and not all of the elements and combinations thereof described in the embodiments are necessarily essential to the solution of the invention.
[0012] In the following explanation, information may be described using the expression "AAA table", but the information may be expressed in any data structure. In other words, to show that the information does not depend on the data structure, the "AAA table" can be called "AAA information".
[0013] FIG. 1 is a functional configuration diagram of a data search system according to an embodiment.
[0014] The data search system 1 includes a data search unit 10, an evaluation index management unit 20, an accuracy calculation unit 30, a data accuracy unit 40, a field adaptability calculation unit 50, a data processing unit 60, an output unit 70, and a registered data management unit 80.
[0015] The data search unit 10 receives a search request including information specifying the field to be searched from the user via a data field input screen 200 (see FIG. 8), and performs a metadata search or a full-text search on the registered data 81 based on the received search request. Specifically, the data search unit 10 identifies a data ID group that satisfies the search request received from the user (note that there may be only one data ID or no data ID depending on the search content), and passes the data ID group to the evaluation index management unit 20 as a calculation request for accuracy and field adaptability. In addition, the data search unit 10 passes the evaluation index results for the data ID group from the evaluation index management unit 20 to the output unit 70.
[0016] The evaluation index management unit 20 passes the data ID group received from the data search unit 10 to the accuracy calculation unit 30 and field adaptability calculation unit 50, obtains the accuracy of the registered data corresponding to each of the data ID group from the accuracy calculation unit 30, and obtains field adaptability indicating the degree of adaptability in the field of the registered data process corresponding to each of the data ID group from the field adaptability calculation unit 50. The evaluation index management unit 20 converts the obtained accuracy and field adaptability into evaluation index values, sorts the data IDs in descending order of evaluation index value, and returns them to the data search unit 10.
[0017] The accuracy calculation unit 30 receives a data ID group from the evaluation index management unit 20, and acquires information on the accuracy of the target to be calculated by referring to the metadata of each piece of data corresponding to the data ID group. Next, the accuracy calculation unit 30 performs an accuracy calculation process to calculate the accuracy based on the acquired accuracy information.
[0018] The data accuracy unit 40 includes an accuracy calculation method management unit 41 , an accuracy type management table 42 , and an accuracy calculation method management table 43 .
[0019] The accuracy calculation method management unit 41 manages the calculation function of each accuracy by managing an accuracy type management table 42 and an accuracy calculation method management table 43. The accuracy calculation method management unit 41 may change the accuracy and calculation method to be used according to an application request.
[0020] The field adaptability calculation unit 50 receives a data ID group from the evaluation index management unit 20, performs a field adaptability calculation process to calculate the field adaptability of the registered data corresponding to the data ID group, and returns the processing result to the evaluation index management unit 20.
[0021] The data processing unit 60 includes a process management unit 61 , a process management table 62 , one or more process field management tables 63 , and one or more application data management tables 64 .
[0022] The process management unit 61 manages the processes and model formulas used when creating registered data, and also performs process management processing to manage process usage information related to the usage history of processes for each field of registered data. The usage data management table 64 stores the data ID of the registered data in which each process was used.
[0023] The output unit 70 displays and outputs a search result screen 300 (see FIG. 9) including the search results obtained by the data search unit 10.
[0024] The registered data management unit 80 stores a plurality of registered data 81 (registered data units) and a registered data metadata table 82 corresponding to each registered data (registered data unit). The registered data may be in file units such as a CSV file, for example.
[0025] FIG. 2 is a hardware configuration diagram of a data search system according to an embodiment.
[0026] The data search system 1 is configured by a computer such as a PC (Personal Computer), and includes a communication interface (I / F) 101, a CPU (Central Processing Unit) 102 as an example of a processor, an input device 103, a storage device 104 as an example of a storage unit, a memory 105 as an example of a storage unit, and a display device 106. The communication I / F 101, CPU 102, input device 103, storage device 104, memory 105, and display device 106 are connected via a bus 107.
[0027] The communication I / F 101 is an interface such as a wired LAN card or a wireless LAN card, and communicates with other devices via a network.
[0028] The CPU 102 executes various processes in accordance with a program stored in the memory 105 and / or the storage device 104. In this embodiment, the CPU 102 executes a program (data search program) to configure a data search unit 10, an evaluation index management unit 20, an accuracy calculation unit 30, a data accuracy unit 40, a field adaptability calculation unit 50, a data processing unit 60, and an output unit 70.
[0029] The memory 105 is, for example, a RAM (RANDOM ACCESS MEMORY), and stores programs executed by the CPU 102 and necessary information. The storage device 104 is, for example, a hard disk or flash memory, and stores programs executed by the CPU 102 and data used by the CPU 102. In this embodiment, the memory 105 stores an accuracy type management table 42, an accuracy calculation method management table 43, a process management table 62, a process field management table 63, and an application data management table 64, and the storage device 14 stores registered data 81 and a registered data metadata table 82.
[0030] The input device 103 is, for example, a mouse, a keyboard, etc., and accepts information input by a user. The display device 106 is, for example, a display, and displays and outputs a user interface including various types of information.
[0031] Next, the registered data metadata table 82 will be described.
[0032] FIG. 3 is a diagram showing the configuration of a registered data metadata table according to an embodiment.
[0033] The registered data metadata table 82 is provided for each registered data 81. The registered data metadata table 82 has records with No. 82a, Item 82b, and Content 82c. No. 82a stores the number of each record for the registered data 81. Item 82b stores information items for the registered data 81. The items include a data ID indicating identification information of the registered data, one or more fields (Field 1, Field 2) to which the registered data belongs, one or more process types (Process 1, 2, 3, 4) used to create the registered data from the basic data, a model formula used in the process (Process 1 Model Formula, Process 2 Model Formula, ...), process usage data indicating a table in the usage data management table 64 that manages the data IDs of the registered data used in the process (Process 1 Usage Data, Process 2 Usage Data, ...), and one or more types related to the accuracy of the registered data 81 (in the example of FIG. 3, the number of references and reliability). Content 82c stores content corresponding to the item. For a portion where there is no content corresponding to an item, information indicating no data (for example, "-") is stored.
[0034] Next, the accuracy type management table 42 will be described.
[0035] FIG. 4 is a configuration diagram of an accuracy type management table according to an embodiment.
[0036] The accuracy type management table 42 is a table for managing the accuracy of registered data, and stores entries (e.g., records) for each type of accuracy (accuracy type). An entry in the accuracy type management table 42 includes fields for an accuracy ID 42a and an accuracy type 42b. The accuracy ID 42a stores the ID of the accuracy type corresponding to the entry. The accuracy type 42b stores the name of the accuracy type corresponding to the entry (accuracy type name).
[0037] Next, the accuracy calculation method management table 43 will be described.
[0038] FIG. 5 is a diagram illustrating the configuration of an accuracy calculation method management table according to an embodiment.
[0039] The accuracy calculation method management table 43 manages the calculation method for each accuracy type of registered data. The accuracy calculation method management table 43 stores an entry for each accuracy type. An entry of the accuracy calculation method management table 43 includes fields for an accuracy ID 43a, an accuracy type 43b, and a calculation method 43c.
[0040] The accuracy ID 43a stores the ID of the accuracy type corresponding to the entry. The accuracy type 43b stores the name of the accuracy type corresponding to the entry (accuracy type name). The calculation method 43c stores information specifying a method for calculating the accuracy for the accuracy type corresponding to the entry. Information specifying the method for calculating the accuracy includes information indicating an API (Application Programming Interface) for calculating the accuracy, the item name of the table in which the accuracy is stored, etc.
[0041] Next, the process management table 62 will be described.
[0042] FIG. 6 is a diagram illustrating the configuration of a process management table according to an embodiment.
[0043] The process management table 62 is a table for managing process types (process types), and includes an entry for each process type. The entry of the process management table 62 includes fields for a process ID 62a, a process type 62b, and a model formula type 62c.
[0044] The process ID 62a stores the ID of the process corresponding to the entry. The process type 62b stores the process type corresponding to the entry. The model formula type 62c stores the type of model formula used in the process type corresponding to the entry.
[0045] Next, the process field management table 63 will be described.
[0046] FIG. 7 is a diagram illustrating a configuration of a process field management table according to an embodiment.
[0047] The process field management table 63 is provided for each field of registered data and manages process usage information related to the usage history of processes used in creating the registered data for that field. The process field management table 63 includes an entry for each process type. An entry in the process field management table 63 stores fields for ID 63a, process type 63b, model formula type 63c, process usage rate 63d, number of processes used 63e, and usage data management table ID 63f.
[0048] The ID 63a stores the ID of the process type corresponding to the entry. The process type 63b stores the name of the process type corresponding to the entry. The model formula type 63c stores the model formula type used for the process type corresponding to the entry. The process utilization rate 63d stores the process utilization rate, which is the rate at which processes of the process type corresponding to the entry are used in the registered data of the field corresponding to the process field management table 63. The process utilization count 63e stores the process utilization count, which is the number of processes of the process type corresponding to the entry are used in the registered data of the field corresponding to the process field management table 63. Here, the process utilization count corresponds to the number of registered data in which processes of the process type corresponding to the entry are used in the registered data of the field corresponding to the process field management table 63. The utilization data management table ID 63f stores the table ID of the utilization data management table 64 that manages the data IDs of the registered data in which the process corresponding to the entry is used.
[0049] Next, the data field input screen 200 displayed and output by the data search unit 10 will be described.
[0050] FIG. 8 is a diagram showing a data field input screen according to one embodiment.
[0051] The data field input screen 200 includes a field list display area 201 , a search field input area 202 , and a search button 203 .
[0052] The field list display area 201 is an area in which a list of fields for registered data is displayed in a selectable manner. The selectable fields may be displayed hierarchically in the field list display area 201. When a field is selected in the field list display area 201, the selected field name may be input into the search field input area 202, or the data search unit 10 may be requested to perform a search for the selected field. The search field input area 202 is an area in which a user inputs characters for the field to be searched. When a field is selected in the field list display area 201, the selected field name may be input into the search field input area 202. The search button 203 is a button that receives an instruction to search for appropriate registered data in the field input into the search field input area 202. When the search button 203 is pressed, the data search unit 10 starts a search process for the field input into the search field input area 202.
[0053] Next, the search result screen 300 displayed and output by the output unit 70 will be described.
[0054] FIG. 9 is a diagram showing a search result screen according to one embodiment.
[0055] The search result screen 300 includes a field list display area 301 and a search result display area 302. The field list display area 301 is similar to the field list display area 201.
[0056] The search result display area 302 includes a search field input and search instruction area 303, a display order condition selection field 304, a display order selection field 305, and a registered data display area 310. One or more registered data display areas 310 are displayed when the target registered data is found by the search. Note that if only one piece of the target registered data is found, information indicating that there is no comparison target may be displayed in the search result display area 302.
[0057] The search field input and search instruction area 303 is an area that displays the searched field name and also accepts input of a new field name to be searched and a search instruction.
[0058] The display order condition selection field 304 is an area for selecting the conditions for the display order of registered data obtained by a search. The display order conditions include accuracy (average of multiple accuracies), any accuracy type within the accuracy, field suitability, etc. When a condition is selected in the display order condition selection field 304, the registered data display area 310 is displayed in the order according to the selected condition.
[0059] The display order selection field 305 is an area for selecting the order (e.g., ascending order, descending order) in which the registered data obtained by the search is to be displayed under the display order conditions. When the order is selected using the display order selection field 305, the registered data display areas 310 are displayed in the order selected under the selected conditions.
[0060] The registered data display area 310 is an area that displays information about registered data found by a search, and includes a registered data name display field 311, an accuracy display field 312, and one or more fields and field adaptability display fields 313.
[0061] The registered data name display field 311 displays the name of the registered data corresponding to the registered data display area 310. The accuracy display field 312 displays various accuracies of the registered data corresponding to the registered data display area 310. The field and field adaptability display field 313 displays information about the field of the registered data corresponding to the registered data display area 310 and the field adaptability of that field. In this embodiment, information about the field adaptability is displayed by placing an index button 314 at a position corresponding to the field adaptability of the registered data on a slider connecting "General," which indicates a field adaptability of 100 percent, and "Special," which indicates a field adaptability of 0 percent. In this embodiment, when the index button 314 is pressed, for example, a breakdown of the process (process type, process model formula, etc.) used to create the registered data is displayed. This search result screen 300 allows the user to easily understand the field adaptability in the field of the process used for the registered data obtained by the search, and can use this information as information to decide which registered data to use.
[0062] Next, the processing operations in the data search system 1 according to the embodiment will be described.
[0063] First, the process management process performed by the process management unit 61 will be described.
[0064] FIG. 10 is a flowchart of a process management process according to an embodiment.
[0065] First, the process management unit 61 checks whether new registration data has been registered in the registration data management unit 80, and if there is new registration data, acquires the registration data metadata table 82 corresponding to the registration data (S11). Note that whether new registration data has been registered in the registration data management unit 80 may be checked using a pre-prepared API.
[0066] Next, the process management unit 61 extracts all fields to which the registered data corresponds, all processes used, and model formulas in the processes from the acquired registered data metadata table 82 (S12).
[0067] Next, the process management unit 61 performs the processing (S13 to S16) of loop 1 with each extracted field as the processing target. Here, in the explanation of this processing, the field to be processed is referred to as the target field.
[0068] In the processing of loop 1, first, the process management unit 61 acquires the process field management table 63 of the target field (S13).
[0069] Next, the process management unit 61 executes the processing of steps S14 to S16 for each process extracted in step S12. That is, the process management unit 61 determines whether or not the extracted process and model formula exist in the process field management table 63 (S14), and if the process and model formula exist (S14: Yes), it updates the process utilization rate of the corresponding entry in the process field management table 63 and adds the data ID of the registered data to the utilization data management table 64 whose table ID is registered in the utilization data management table ID 63f (S15). On the other hand, if the process and model formula do not exist (S14: No), the process management unit 61 adds a new entry corresponding to the process, calculates and stores the process utilization rate, registers the ID of the new utilization data management table 64 in the utilization data management table ID 63f, creates the utilization data management table 64, and adds the data ID of the registered data (S16).
[0070] Here, the process utilization rate can be calculated as follows. First, the process management unit 61 calculates the number of processes in use (new) = the number of processes in use (old) + 1. If a corresponding entry already exists in the process field management table 63, the process utilization rate (old) is the number of processes in use for that entry, and if no entry exists, the process utilization rate (old) is 0. Next, the process management unit 61 calculates the number of data items in the field (new) = the number of data items in the field (old) + 1. The number of data items in the field (old) can be calculated using the same method as before. Next, the process management unit 61 calculates the process utilization rate (new) by calculating the number of processes in use (new) ÷ the number of data items in the field (new).
[0071] After executing the processing of Loop 1 for the target field, the process management unit 61 performs the processing of Loop 1 for the unprocessed fields within the extracted fields as new target fields, and when the processing of Loop 1 has been performed for all extracted fields, it exits Loop 1.
[0072] Next, the process management unit 61 acquires the process management table 62 (S17), and executes the processing of steps S18 to S19 for each process extracted in step S12. Specifically, the process management unit 61 determines whether the extracted process and model formula exist in the process management table 62 (S18), and if the extracted process and model formula do not exist (S18: No), it adds an entry to the process management table 62 and registers the extracted process and model formula in this entry (S19).
[0073] Next, the process management unit 61 executes the processes of steps S18 to S19 for all the processes extracted in step S12, and then ends the process management process.
[0074] According to this process management processing, when new registration data is added, the process management table 62 and the process field management table 63 can be appropriately updated to the latest status.
[0075] Next, the registered data search process will be described.
[0076] First, the data search unit 10 receives a search request from the user via the data field input screen 200, and based on the received search request, identifies a group of data IDs that satisfy the search request, and passes the group of data IDs to the evaluation index management unit 20 as a request to calculate accuracy and field adaptability.
[0077] The evaluation index management unit 20 passes the data ID group received from the data search unit 10 to the accuracy calculation unit 30 and the field adaptability calculation unit 50 .
[0078] The accuracy calculation unit 30 receives a data ID group from the evaluation index management unit 20, and acquires information on the accuracy of the target to be calculated by referring to the metadata of each piece of data corresponding to the data ID group. Next, the accuracy calculation unit 30 performs accuracy calculation processing based on the acquired accuracy information. Specifically, the accuracy calculation unit 30 acquires the accuracy calculation method by referring to the accuracy type management table 42 and the accuracy calculation method management table 43, calculates the accuracy based on the calculation method, and returns the processing result to the evaluation index management unit 20.
[0079] On the other hand, the field adaptability calculation unit 50 receives a data ID group from the evaluation index management unit 20, performs a field adaptability calculation process (see Figure 11) to calculate the field adaptability for the registered data corresponding to the data ID group, and returns the processing result to the evaluation index management unit 20.
[0080] The evaluation index management unit 20 acquires the accuracy of the registered data corresponding to each of the data ID groups from the accuracy calculation unit 30, and acquires the field adaptability of the registered data corresponding to each of the data ID groups from the field adaptability calculation unit 50.
[0081] Next, the evaluation index management unit 20 converts the obtained accuracy and field adaptability into an evaluation index value.
[0082] Here, the conversion process into the evaluation index value is performed, for example, as follows.
[0083] The evaluation index management unit 20 identifies the maximum (MAX) and minimum (MIN) accuracy values from the accuracy of the acquired data ID group, sets the maximum accuracy value as an evaluation index value of 100 points, the minimum accuracy value as an evaluation index value of 0 points, and for accuracy between the maximum and minimum values, sets the evaluation index value as (accuracy - minimum value) ÷ maximum value × 100. This allows the highest accuracy to be set as 100 points, and the lowest accuracy to be set as 0 points.
[0084] Furthermore, the evaluation index management unit 20 identifies the maximum value (MAX) and minimum value (MIN) of field adaptability from the field adaptability of the acquired data ID group, and sets the maximum value of field adaptability as an evaluation index value of 100 points, the minimum value as an evaluation index value of 0 points, and for field adaptability between the maximum and minimum values, sets the evaluation index value as (field adaptability - minimum value) ÷ maximum value × 100. This allows the most general case to be assigned 100 points, and the most specific (least general) case to be assigned 0 points. Note that if all of the registered data in the same field obtained by the search have the same evaluation index value, the evaluation index management unit 20 may set all evaluation index values to 100 points, indicating generality.
[0085] Next, the evaluation index management unit 20 sorts the data IDs of the registered data in descending order of evaluation index value and returns them to the data search unit 10. Here, the evaluation index management unit 20 sorts the registered data in descending order of the evaluation index value of accuracy, and may sort the registered data in descending order of the evaluation index value of field adaptability if there is registered data with the same evaluation index value. Here, if there are multiple types of evaluation index values for accuracy, the data may be sorted according to the evaluation index value of the type of accuracy that takes priority in sorting, or according to the average value of the multiple types of evaluation index values.
[0086] Next, the data search unit 10 obtains the search results of the sorted data ID groups from the evaluation index management unit 20 and passes them to the output unit 70.
[0087] Next, the output unit 70 displays and outputs a search result screen 300 (see FIG. 9) including the search results obtained by the data search unit 10.
[0088] Next, the field adaptability calculation process will be described.
[0089] FIG. 11 is a flowchart of a field adaptability calculation process according to an embodiment.
[0090] The field adaptability calculation unit 50 receives a group of data IDs from the evaluation index management unit 20 (S21), and performs the processing of loop 2 (S22 to S26) for each data ID as a processing target. Here, in the description of this processing, the data ID to be processed is referred to as a target data ID.
[0091] In the processing of loop 2, first, the field adaptability calculation unit 50 acquires the registered data metadata table 82 of the registered data 81 of the target data ID, and identifies all fields of this registered data (S22).
[0092] Next, the field adaptability calculation unit 50 performs the processing (S23 to S26) of loop 3 with each of the identified fields as the processing target. Here, in the description of this processing, the processing target field will be referred to as the target field.
[0093] In the processing of loop 3, first, the field adaptability calculation unit 50 acquires the process field management table 63 corresponding to the target field (S23), and identifies all processes of the registered data from the registered data metadata table 82 (S24). Here, the field adaptability calculation unit 50 identifies information on all processes (sets of process type, process model formula, and process usage data) in the registered data metadata table 82.
[0094] Next, the field adaptability calculation unit 50 performs the processing of loop 4 (S25) with each of the identified processes as the processing target. Here, in the description of this processing, the processing target process will be referred to as the target process.
[0095] In the processing of loop 4, the field adaptability calculation unit 50 extracts the utilization rate of the target process (process utilization rate) from the process field management table 63 (S25).
[0096] After executing the processing of Loop 4 for the target process, the field adaptability calculation unit 50 performs the processing of Loop 4 for the unprocessed process in the identified process as a new target process, and when the processing of Loop 4 has been performed for all the identified processes, it exits Loop 4 and proceeds to step S26.
[0097] In step S26, the field adaptability calculation unit 50 calculates the field adaptability of the target field based on the information of the utilization rate extracted from the process field management table 63 by the following method (first method).
[0098] The field adaptability calculation unit 50 calculates the average number of processes (average number of processes) used for registered data in the same field as the target field. Specifically, the field adaptability calculation unit 50 calculates the average number of processes by summing up the process utilization rates of all entries in the process field management table 63 for the target field and dividing the sum by the number of registered data 81 in the target field. In this embodiment, the average number of processes is an integer value obtained by rounding the calculated value to one decimal place. Here, the number of registered data 81 in the target field is, for example, the number of data IDs remaining after obtaining data IDs from the utilization data management table 64 corresponding to the table ID of the utilization data management table ID 63f in all entries in the process field management table 63 for the target field and eliminating duplicates.
[0099] Next, the field adaptability calculation unit 50 calculates the field adaptability as shown below so that if the number of processes of the registered data of the corresponding data ID (the number of processes stored in the registered data metadata table) is the same as the average number of processes of the registered data in the target field, the field adaptability is large, and if the number of processes is different from the average number of processes (i.e., if the number of processes is more or less than the average number of processes), the field adaptability is small.
[0100] Specifically, the domain adaptability calculation unit 50 sorts the processes of the registered data 81 of the target data ID in descending order of process utilization rate. In the subsequent processing, the domain adaptability calculation unit 50 refers to the records in the sorted result in descending order.
[0101] Next, if the number of processes in the registered data 81 of the target data ID is the same as the average number of processes, the field adaptability calculation unit 50 determines the value obtained by dividing the sum of the process utilization rates of each process by the number of processes in the registered data 81 of the target data ID as the field adaptability; if the number of processes in the registered data 81 of the target data ID is greater than the average number of processes, the field adaptability calculation unit 50 determines the following as the field adaptability: (the sum of the process utilization rates of the processes in the registered data 81 of the target data ID up to the average number of processes + the sum of the process utilization rates of the processes exceeding the average number of processes in the registered data 81 of the target data ID × a predetermined value (a value less than 1, for example, 0.5)) ÷ the number of processes in the registered data 81 of the target data ID; and if the number of processes in the registered data 81 of the target data ID is less than the average number of processes, the field adaptability calculation unit 50 determines the following: (the sum of the process utilization rates of the processes in the registered data 81 of the target data ID + a predetermined value (a value less than 1, for example, 0.5) × (average number of processes - number of processes in the registered data 81 of the target data ID)) ÷ the average number of processes.
[0102] Next, after performing the processing of Loop 3 on the target field, the field adaptability calculation unit 50 performs the processing of Loop 3 on the unprocessed fields within the identified field as new target fields, and when the processing of Loop 3 has been performed on all identified fields, it exits Loop 3.
[0103] Next, after executing the processing of Loop 2 on the target data ID, the field adaptability calculation unit 50 performs the processing of Loop 2 using the unprocessed data ID in the received data ID as the new target data ID, and when the processing of Loop 2 has been performed on all the received data IDs, it exits Loop 2 and terminates the field adaptability calculation processing.
[0104] Next, a modified example of the data search system according to this embodiment will be described.
[0105] The process of calculating the field adaptability in the field adaptability calculation unit may be as follows: For convenience, the respective functional units will be described using the same reference numerals as those shown in FIG.
[0106] The field adaptability calculation unit 50 calculates the field adaptability of the target field based on the information of the utilization rate extracted from the process field management table 63 using the following method (second method).
[0107] The field adaptability calculation unit 50 extracts entries from the process field management table 63 that have the maximum process utilization rate for each process type and whose process utilization rate is equal to or greater than a predetermined value (for example, 50 percent), and sorts them in descending order of process utilization rate. Here, one or more processes remaining after sorting are called sorted processes. Note that if there are multiple entries with the maximum process utilization rate for the same process type, only one of the entries is selected.
[0108] Next, the domain adaptability calculation unit 50 extracts processes that are the same as the sorted processes from the processes of the registered data of the corresponding data ID. Next, if the extracted processes include all or part of the sorted processes, the domain adaptability calculation unit 50 calculates a provisional domain adaptability by (number of extracted processes ÷ number of sorted processes) × 100. If the sorted processes include processes other than the extracted processes, the domain adaptability is calculated by dividing the provisional domain adaptability by the total number of processes in the registered data of the corresponding data ID. If the sorted processes include no processes other than the extracted processes, the provisional domain adaptability is calculated as the domain adaptability. On the other hand, if the extracted processes do not include all or part of the sorted processes, the domain adaptability is set to 0.
[0109] Furthermore, the field adaptability calculation unit 50 may calculate the field adaptability of the target field by the following method (third method). Specifically, when the field adaptability calculated by the first method is A and the field adaptability calculated by the second method is B, the field adaptability calculation unit 50 may determine the field adaptability as (A+B) / 2.
[0110] The present invention is not limited to the above-described embodiment, and can be modified appropriately without departing from the spirit of the present invention.
[0111] For example, in the above embodiment, the accuracy and field adaptability of the registered data obtained by the search were calculated, and the registered data was rearranged based on the accuracy and field adaptability, but it is also possible to calculate only the field adaptability and rearrange the registered data based on the field adaptability.
[0112] Furthermore, in the above embodiment, an example was shown in which the data search system 1 was configured with a single computer, but the data search system 1 may also be configured with multiple computers, and external storage may be used as the registered data management unit 80.
[0113] In addition, in the above-described embodiments, some or all of the processing performed by the CPU may be performed by a hardware circuit. Also, the programs in the above-described embodiments may be installed from a program source. The program source may be a program distribution server or a recording medium (e.g., a portable recording medium). [Explanation of symbols]
[0114] 1...data search system, 10...data search unit, 20...evaluation index management unit, 30...accuracy calculation unit, 40...data accuracy unit, 50...field adaptability calculation unit, 60...data processing unit, 70...output unit, 80...registered data management unit, 102...CPU, 104...storage device, 105...memory
Claims
1. A data search system for searching for appropriate registered data from among a plurality of registered data, A processor and a storage unit are included. The storage unit includes the plurality of registration data, metadata including information about fields of the plurality of registration data and processes performed to obtain the registration data, and process usage information regarding a history of usage of one or more processes performed to obtain registration data for each field of the registration data for the plurality of registration data, The processor: Accepts the specification of the field of registered data to be searched, Searching for registered data corresponding to the field; For the registered data obtained by the search, a field adaptability indicating the degree of adaptability in the field of the process performed to obtain the registered data obtained by the search is calculated based on the metadata of the registered data obtained by the search and the process usage information. Data retrieval system.
2. The processor: Outputting information about the registered data obtained by the search and the field adaptability calculated for the registered data.
2. The data retrieval system according to claim 1.
3. The processor: Calculating the accuracy of the registered data obtained by the search; The registered data obtained by the search is output in order based on the accuracy and the field adaptability.
3. The data search system according to claim 2.
4. The processor: The registered data is output in an order based on the accuracy, and for registered data with the same accuracy, the registered data is output in an order based on the field adaptability.
4. The data search system according to claim 3.
5. The process is identified by the type of the process and the type of model formula used in the process.
2. The data retrieval system according to claim 1.
6. the process usage information includes a process usage count, which is the number of times the process is used, for the plurality of registered data; The processor: An average number of processes, which is the average number of processes used in the field, is identified, and the field adaptability is calculated based on the number of processes in the registered data obtained by the search and the average number of processes.
2. The data retrieval system according to claim 1.
7. The processor: The field adaptability is calculated so that the field adaptability is high when the average number of processes is the same as the number of processes in the registered data obtained by the search, and is low when the average number of processes is different from the number of processes in the registered data obtained by the search.
7. The data search system according to claim 6.
8. the process usage information includes a process usage count, which is the number of times the process is used for the plurality of registered data; The processor: The field adaptability is calculated based on the number of process uses corresponding to the process of the registered data obtained by the search.
2. The data retrieval system according to claim 1.
9. The processor: The field adaptability is calculated based on the process type of the registered data obtained by the search and the process types of the plurality of registered data in the corresponding field.
2. The data retrieval system according to claim 1.
10. The processor: If there are multiple fields corresponding to the registered data obtained by the search, calculate the field adaptability for each of the multiple fields.
2. The data retrieval system according to claim 1.
11. The processor: Creating the process utilization information based on the metadata 2. The data retrieval system according to claim 1.
12. A data search method for a data search system that searches for appropriate registered data from among a plurality of registered data, comprising: The data search system includes: storing the plurality of registration data, metadata including information regarding fields of the plurality of registration data and processes performed to obtain the registration data, and process usage information regarding a history of usage of one or more processes performed to obtain registration data for the fields of registration data for the plurality of registration data; The data search system includes: Accepts the specification of the field of registered data to be searched, Searching for registered data corresponding to the field; For the registered data obtained by the search, a field adaptability indicating the degree of adaptability in the field of the process performed to obtain the registered data obtained by the search is calculated based on the metadata of the registered data obtained by the search and the process usage information. Data retrieval methods.
13. A data search program for causing a computer to search for appropriate registered data from among a plurality of registered data, The computer, Accept the specification of the field of registered data to be searched, Searching for registered data corresponding to the field; For the registered data obtained by the search, a field adaptability indicating the degree of adaptability in the field of the process performed to obtain the registered data obtained by the search is calculated based on metadata including information on the fields of the registered data of the plurality of registered data obtained by the search and the process performed to obtain the registered data, and process usage information on the fields of the registered data of the plurality of registered data and the history of usage of one or more processes performed to obtain the registered data in the field. Data retrieval program.
Citation Information
Patent Citations
Method and device for document retrieval and computer- readable recording medium with recorded program making computer actualize method for document retrieving
JP2002032401A
Category selection apparatus, advertisement distribution system, category selection method and program
JP2019003254A
Information processing device and method, and program
JP2021056560A
Valid data information extraction system and valid data information extraction
JP2021096792A
Data reliability calculation device, data reliability calculation method and data reliability calculation program
JP2021114077A