Computer system and data management method
The system efficiently manages pre- and post-processed data by using separate storage areas and archiving less frequently accessed data, addressing the need for cost-effective and accessible data storage.
Patent Information
- Application Number
- JP2023205244
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-17
AI Technical Summary
Existing systems require pre-processing of raw data before application usage and need efficient storage solutions for both pre-processed and post-processed data.
A computer system with separate areas for pre-processed and post-processed data, along with archive areas for archiving less frequently accessed data, and arithmetic units that selectively archive and re-generate data based on access patterns and resource costs.
This solution enables efficient storage and retrieval of data by archiving less frequently used pre- and post-processed data, reducing storage costs and improving data accessibility when needed.
Smart Images

Figure 2025090171000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the management of pre - processed data and post - processed data.
Background Art
[0002] There is Information Lifecycle Management (ILM) as a cost - reduction method when storing data with low access frequency for a long period. For example, in Patent Document 1, there is a system and method for operating a large - volume data platform, which includes steps of receiving discrete client data in a data analysis platform; storing the client data in a real - time storage system; storing the client data in an archive storage system distributed in column - based form; storing the client data in a distributed storage system accessible to a network including these steps; receiving a data query request through an inquiry interface; and selectively interfacing with the client data from the real - time storage system and the archive storage system according to the inquiry. A system and method including these steps are disclosed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is a system that requires pre - processing of raw data before passing it to an application and needs to store data so that the post - processed data can be passed to the application. In this system, a technology for efficiently storing data is desired.
Means for Solving the Problems
[0005] The computer system includes a pre - processed data area for storing pre - processed data, a post - processed data area for storing post - processed data generated by processing the pre - processed data, a pre - processed data archive area for archiving the pre - processed data moved from the pre - processed data area, a post - processed data archive area for archiving the post - processed data moved from the post - processed data area, and one or more arithmetic units. The one or more arithmetic units select one or more pre - processed data that are the sources of one or more post - processed data to be archived, select first pre - processed data that meet preset conditions from the one or more pre - processed data, move the first pre - processed data from the pre - processed data area to the pre - processed data archive area, delete, without archiving, the first post - processed data that is generated using the first pre - processed data and meets the preset conditions from the post - processed data area. The condition for selecting the first pre - processed data includes that the first pre - processed data is the source of multiple types of post - processed data. The one or more arithmetic units generate the first post - processed data using the first pre - processed data in the pre - processed data archive area in response to an access request to the first post - processed data from an application, and return it to the application.
Advantages of the Invention
[0006] Data can be stored efficiently.
Brief Description of the Drawings
[0007]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Modes for Carrying Out the Invention
[0008] In the following, for convenience, when necessary, it will be described by dividing it into a plurality of sections or examples. However, unless otherwise specified, they are not unrelated to each other. One is a variation, detail, supplementary explanation, etc. of part or all of the other. Also, in the following, when referring to the number of elements, etc. (including the number, numerical value, quantity, range, etc.), unless otherwise specified and in cases where it is clearly limited to a specific number in principle, it is not limited to that specific number, and it may be more than or less than the specific number.
[0009] The system or device in this specification may be a physical computer system (one or more physical computers), or a system constructed on a group of computing resources (a plurality of computing resources) such as a cloud platform. The computer system or group of computing resources may include one or more interface devices (including, for example, communication devices and input / output devices), one or more storage devices (including, for example, memory (main memory) and auxiliary storage devices), and one or more arithmetic devices.
[0010] When a function is realized by a program being executed by an arithmetic unit, since the defined processing is performed while appropriately using a storage device and / or an interface device, etc., the function may be at least part of one or more arithmetic units. The processing described with the function as the subject may be the processing performed by a system including one or more arithmetic units.
[0011] The program may be installed from a program source. The program source may be, for example, a program distribution computer or a computer-readable storage medium (for example, a computer-readable non-transitory storage medium). The description of each function is an example, and a plurality of functions may be grouped into one function, or one function may be divided into a plurality of functions.
Example
[0012] FIG. 1A shows the overall logical configuration of the system according to this example. The system according to this example includes an input 101, a load data area 102 (pre-processed data area), an online data processor 103, a published data area 104 (post-processed data area), a load data archive area 121 (pre-processed data archive area), an on-demand data processor 131, a published data archive area 141 (post-processed data archive area), an archive creator 108, an information life cycle management 105, a request receiver 106, applications 107 and 171.
[0013] The input 101 is the data receiving unit in this embodiment. In this embodiment, it is assumed that all inputs are digital data and messages, and the input 101 only receives data and does not process the data. The message indicates the value at a specific time. For example, it indicates the time information and the value at the time indicated by the time information. However, in another embodiment where the input is analog data, data conversion such as analog-to-digital conversion is performed so that the essential meaning remains unchanged. Also, even if the input is digital data, if it is data other than a message, such as stream data that only indicates values continuous on the time axis, the digital data is converted into a message. Note that data other than messages may also be handled.
[0014] The raw data area 102 is a storage area that stores the data received by the input 101 as it is without processing. In this embodiment, the raw data received by the input 101 is stored in different tables for each source. Note that the storage format of the input raw data is not particularly limited.
[0015] The online data processor 103 processes the data stored in the tables in the raw data area 102 at specified frequencies or data sizes. For example, the raw data is processed for use by the application 107 or 171. A plurality of types of processed data processed in different ways can be generated from one type of raw data, and one type of processed data can be generated from a plurality of types of raw data.
[0016] The published data area 104 is a storage area that stores the data processed by the online data processor 103. In the following, it is assumed that different processed data are stored in different tables. Note that the storage format of the processed data is not particularly limited.
[0017] The archive creator 108 moves the data stored in the raw data area 102 or the published data area 104 that has a low access frequency from the applications 107 and 171 to the archive areas 121 or 141.
[0018] The archive area 121 for load data is a storage area for archiving tables with low access frequencies among the tables stored in the load data area 102.
[0019] The archive area 141 for published data is a storage area for archiving tables with low access frequencies among the tables stored in the published data area 104.
[0020] The on-demand data processor 131 reprocesses old data stored in the archive area 121 for load data when accessed from the application 107 or 171.
[0021] The request receiver 106 receives access from the application 107 or 171 and returns the requested data to the application 107 or 171. Specifically, the request receiver 106 acquires the requested data from the published data area 104 or the archive area 141 for published data, or processes the data acquired from the archive area 121 for load data using the on-demand data processor 131.
[0022] The management module 151 of the information lifecycle management 105 issues an instruction to the archive creator 108 to move specific data from the load data area 102 or the published data area 104 to the archive area 121 for load data or the archive area 141 for published data. Also, it passes the data infrastructure settings of the on-demand data processor 131 to the request receiver 106.
[0023] The access frequency table 152 aggregates the access frequencies to each table in the published data area 104 and the archive area 141 for published data.
[0024] The data lineage table 153 manages the correspondence relationships among the tables included in the load data area 102 or the archive area for load data 121 corresponding to the tables included in the published data area 104 and the archive area for published data 141, respectively, and the names of the data processing programs used for their processing, that is, the processed tables, the tables before processing, and their data processing programs.
[0025] The data conversion infrastructure setting table 154 manages the amount of resources required for the operation of each data processing program. The infrastructure cost table 155 manages the unit prices of the resources required for data areas, archive areas, and the execution of data processing programs.
[0026] The logical configuration of the system shown in FIG. 1A can be implemented in a computer system, and the computer system can include one or more computers. Also, each logical component can be implemented on one or more computers. For example, applications 107 and 171 and input 101 can be implemented on different computers, and other components can be implemented on computers different from them. The computers communicate with each other via a network. In other examples, each element of the system may be implemented on a different computer. Thus, each component of the system can be realized by a computer system including one or more arithmetic units and one or more storage devices.
[0027] FIG. 1B shows an example of the hardware configuration of a computer that can be used in the embodiments of this specification. The computer includes a CPU (arithmetic unit) 211 that executes various programs, a memory (main storage device) 202 that stores various programs, and an auxiliary storage device 213 that stores various data. The CPU 211 can include one or more cores, and the memory 212 is, for example, a RAM that includes a volatile storage area. The auxiliary storage device 213 is, for example, an HDD (Hard Disk Drive) or a flash memory, etc., and can provide a non-volatile storage area.
[0028] The computer further includes an output device 214 for presenting information to the user, an input device 215 for inputting instructions, images, etc. by the user, and a communication device 206 for communicating with other devices. These are interconnected by a bus 207.
[0029] The functional units of the system shown in FIG. 1A can be implemented, for example, by the CPU 211 operating according to a program. The CPU 211 reads and executes various programs from the memory 212 as necessary. The memory 212 stores programs. Each program is, for example, loaded from the auxiliary storage device 213 into the memory 212 and executed by the CPU 211. Note that at least a part of the functional units may be constituted by logic circuits.
[0030] The auxiliary storage device 213 stores data referenced or managed by various programs. In the configuration example shown in FIG. 1A, the data storage areas 102, 121, 104, 141 can be provided by one or more auxiliary storage devices 213, respectively. For example, the bit cost of the storage device providing the load data archive area 121 or the published data archive area 141 is lower than the bit cost of the storage device providing the load data area 102 or the published data area 104. This enables an efficient system configuration.
[0031] Data may be stored in different ways between the load data area 102 and the load data archive area 121. For example, the load data area 102 may store uncompressed data, and the load data archive area 121 may store compressed data. The bit cost can be reduced by compression. The same explanation also applies to the published data area 104 and the published data archive area 141.
[0032] For example, the archive area 121 for load data and the archive area 141 for publish data are each provided by one or more HDDs. And the load data area 102 and the publish data area 104 are each provided by one or more SSDs. Although HDDs are inferior to SSDs in terms of access performance, they can store more data at a lower cost. In this way, different types of storage drives can be used depending on the data area. Note that the types of storage drives in the two archive areas 121 and 141 may be different, and the types of storage drives in the storage areas of the load data area 102 and the publish data area 104 may also be different.
[0033] The output device 214 is composed of devices such as a display, a printer, and a speaker. The input device 215 is composed of devices such as a keyboard, a mouse, and a microphone. The output device 214 presents the input result from the user and also presents the processing result. An instruction from the user is input by the input device 215. The communication device 206 receives data transmitted from other devices connected via a network, for example, and also transmits the processing result to other devices. Note that some devices may be omitted.
[0034] FIG. 2 shows a configuration example of the access frequency table 152. The access frequency table 152 has a converted table ID column 201, an access count column 202 for period 1, and an access count column 203 for period 2.
[0035] Period 1 and period 2 refer to the period including the current of the tumbling window and the period one older. Also, in this embodiment, the window size is set to 1 hour. For example, when the current time is 7:13 am on June 29, 2023, period 1 is the period from 7 am to 8 am on June 29, 2023, and period 2 is the period from 6 am to 7 am on June 29, 2023.
[0036] The converted table ID column 201 holds the IDs of the tables included in the published data area 104 and the archive area 141 for published data. In this embodiment, the ID uses a format in which the hierarchy (pub representing the published data area or raw representing the loaded data area), database name, table name, month, and year are concatenated. However, depending on the type of database used in the data area, other namespace management layers such as a schema may be used.
[0037] The access count column 202 for period 1 holds the total number of accesses during the period including the current tumbling window to the table. The access count column 203 for period 2 holds the total number of accesses during the one-old period of the tumbling window to the table.
[0038] Note that in this embodiment, in a later step 603, as in row 204, tables with zero consecutive access counts for two periods are targeted for archiving by the information lifecycle management 105. Also, instead of managing each table, it is possible to manage each part of the table. Thus, the unit of management is not limited.
[0039] FIG. 3 shows a configuration example of the data lineage table 153. The data lineage table 153 has a converted table ID column 301, a pre-conversion table ID column 302, and a data processing program column 303.
[0040] The converted table ID column 301 holds the IDs of the tables included in the published data area 104 and the archive area 141 for published data. In this embodiment, the ID uses a format in which the hierarchy (pub representing the published data area or raw representing the loaded data area), database name, table name, month, and year are concatenated. However, depending on the type of database used in the data area, other namespace management layers such as a schema may be used.
[0041] In addition, in this embodiment, wild card notation "*" is allowed for the month and year included in the ID. The notation "pub_pdm.vehicle_*_*" includes "pub_pdm.vehicle_apr_2023" and "pub_pdm.vehicle_may_2022". Note that the relationship between the year and month of the data before conversion and the year and month of the data after conversion is set in advance.
[0042] The ID column 302 of the table before conversion holds the IDs of the tables included in the load data area 102 and the archive area 121 for load data. In this embodiment, the ID uses a format in which the hierarchy ( "pub" representing the published data area or "raw" representing the load data area), database name, table name, month, and year are concatenated. However, depending on the type of database used in the data area, other name space management layers such as a schema may be used.
[0043] In addition, in this embodiment, wild card "*" notation is allowed for the month and year included in the ID. That is, the notation "raw_general.vehicle_type_*_*" includes "raw_general.vehicle_type_apr_2023" and "raw_general.vehicle_type_may_2022". Note that the relationship between the year and month of the data before conversion and the year and month of the data after conversion is set in advance.
[0044] The data processing program column 303 holds the ID of the data processing program used for data processing from the table before conversion to the table after conversion. One table after conversion can be generated from one or more tables before conversion. Also, one table before conversion can be used for the generation of one or more tables after conversion.
[0045] Fig. 4 shows a configuration example of the data conversion infrastructure setting table 154. The data conversion infrastructure setting table 154 has a data processing program column 401, a CPU column 402, a main memory column 403, a secondary storage column 404, and a batch size column 405.
[0046] The data processing program sequence 401 holds the IDs of the data processing programs used for data processing from the pre-conversion table to the post-conversion table.
[0047] The CPU sequence 402 holds the number of CPU cores to be allocated when executing the data processing program on the online data processing machine 103.
[0048] The main memory sequence 403 holds the amount of main memory to be allocated when executing the data processing program on the online data processing machine 103.
[0049] The secondary storage sequence 404 holds the amount of secondary storage to be allocated when executing the data processing program on the online data processing machine 103.
[0050] The batch size sequence 405 holds the batch size to be allocated when executing the data processing program on the online data processing machine 103. The batch size indicates the amount of data to be processed together. For example, "1 hour" indicates generating batch data by aggregating data for one hour. Batch data is processed together in one operation.
[0051] When executing the data processing program on the on-demand data processing machine 131, by using these parameters, it is possible to allocate the minimum necessary resources for reliable execution.
[0052] Fig. 5 shows a configuration example of the infrastructure cost table 155. The infrastructure cost table 155 has an infrastructure resource column 501 and a unit price column 502. The infrastructure resource column 501 holds the types of infrastructure resources such as data areas and CPUs. The data area means the areas of the raw data area 102 and the published data area 104. The archive area means the raw data archive area 121 and the published data archive area 141.
[0053] Note that the unit prices of the load data area 102 and the published data area 104 may be different, and the unit prices of the load data area 102 and the published data area 104 may be different. As in this example, the unit price of the archive area is lower than that of the data area. The unit price column 502 holds the unit price (price) of the corresponding infrastructure resource type. In the example shown in FIG. 5, the unit price per hour is shown. For example, the unit price of the data area is $0.01 per GB per second. The unit price may be, for example, the unit price of the service when using a cloud service, or a value calculated from the price when using a purchased device. Note that the unit price may be expressed in other units according to the design.
[0054] FIG. 6 shows a flowchart of the archive process 601 of the management module 151. The archive process 601 is executed periodically, for example, at one-hour intervals.
[0055] After the start of the process, in step 602, the management module 151 initializes the post-archive candidate conversion table ID table 801 and the pre-archive candidate conversion table ID table 901. FIG. 8 shows a configuration example of the post-archive candidate conversion table ID table 801. The post-archive candidate conversion table ID table 801 has a post-conversion table ID column 802. FIG. 9 shows a configuration example of the pre-archive candidate conversion table ID table 901. The pre-archive candidate conversion table ID table 901 has a pre-conversion table ID column 902 and an appearance count column 903.
[0056] Subsequently, in step 603, the management module 151 acquires the post-conversion table IDs whose access counts are less than a predetermined threshold for a predetermined number of consecutive periods from the access frequency table 152 and adds them to the post-archive candidate conversion table ID table 801. As a result, only the post-conversion tables with few recent access counts among the post-conversion tables are set as candidates for archiving. This step selects, in the published data area 104, candidates (archive candidates) for moving the processed (converted) data with few accesses to the archive area 141 for published data.
[0057] The above example uses the number of accesses in the immediately preceding predetermined period as an indicator showing the access frequency. An example of the predetermined period is 2, and an example of the threshold value is 1. That is, the converted table ID with the number of accesses being 0 for two consecutive periods is obtained. Note that the indicator showing the access frequency and the determination condition may be different from the above conditions. For example, the condition may be that the period in which the number of accesses is less than the threshold value in the most recent predetermined number of periods exceeds a predetermined ratio, or the average value of the number of accesses in the most recent predetermined number of periods is less than the threshold value. Also, the threshold value may vary depending on the period. Appropriate archive candidates can be selected based on past accesses.
[0058] Specifically, the converted table ID is obtained by the procedure shown in the following pseudo-code. SELECT Converted Table ID FROM Access Frequency Table WHERE Number of Accesses in Period 1 = 0 AND Number of Accesses in Period 2 = 0
[0059] Subsequently, in step 604, the management module 151 refers to the converted table ID column 301 and the pre-conversion table ID column 302 of the data lineage table 153, obtains the number of occurrences for each pre-conversion table ID corresponding to the converted table ID included in the archive candidate converted table ID table 801, and adds it to the archive candidate pre-conversion table ID table 901. Thereby, only the pre-conversion table corresponding to the converted table with a small number of recent accesses in the pre-conversion table is set as an archive candidate. Specifically, the number of occurrences for each pre-conversion table ID is obtained by the procedure shown in the following pseudo-code.
[0060] SELECT Data Lineage Table.Pre-conversion Table ID, COUNT(Data Lineage Table.Pre-conversion Table ID) FROM Data Lineage Table, Archive Candidate Pre-conversion Table ID Table WHERE Data Lineage Table.Converted Table ID = Archive Candidate Pre-conversion Table ID Table.Converted Table ID GROUP BY Data Hierarchy Table. Pre - conversion Table ID;
[0061] Subsequently, in step 605, the management module 151 refers to the post - conversion table ID column 301 and the pre - conversion table ID column 302 of the data hierarchy table 153, obtains the pre - conversion table ID corresponding to the post - conversion table ID included in the archive candidate post - conversion table ID table 801, and adds it to the archive candidate data hierarchy table 1001. FIG. 10 shows a configuration example of the archive candidate data hierarchy table 1001. The archive candidate data hierarchy table 1001 has a pre - conversion table ID column 1002 and a post - conversion table ID column 1003. Specifically, the pre - conversion table ID corresponding to the post - conversion table ID is obtained by the procedure shown in the following pseudo - code.
[0062] SELECT Data Hierarchy Table.Pre - conversion Table ID, Data Hierarchy Table.Post - conversion Table ID FROM Data Hierarchy Table, Archive Candidate Pre - conversion Table ID Table WHERE Data Hierarchy Table.Post - conversion Table ID = Archive Candidate Pre - conversion Table ID Table.Post - conversion Table ID
[0063] Subsequently, the management module 151 repeatedly executes the determination process 607 from the beginning of the archive candidate pre - conversion table ID table in steps 606 to 608. The details of the determination process 607 will be described with reference to FIG. 7.
[0064] When the management module 151 finishes step 608, it ends this flow.
[0065] Fig. 7 shows a flowchart of the determination process 607 in the archive process 601 of the management module 151. The determination process 607 determines the data to be archived from the converted data (published data) of the archive candidate and its pre-conversion data (loaded data). As described above, the determination process 607 is sequentially executed for each row starting from the first row of the archive candidate pre-conversion table ID list 901. In another embodiment, exception processing is performed such that, regardless of cost comparison, tables that should leave the pre-conversion or post-conversion tables are exceptionally excluded from the target of this process.
[0066] First, in step 702, the management module 151 determines whether the value in the appearance count column 903 of the current row in the archive candidate pre-conversion table ID list 901 is a predetermined value, for example, 2 or less. If the determination result is Yes (Y), the determination process 607 ends. If the determination result is No (N), the process proceeds to step 703. This allows the selection and archiving of the pre-conversion tables that are the basis of many post-conversion tables, enabling efficient data storage. Note that this step may be omitted.
[0067] Subsequently, in step 703, the management module 151 obtains the size (storage area usage amount) of the pre-conversion table corresponding to the value in the pre-conversion table ID column 902 of the current row from the loaded data area 102. Here, the size of metadata related to the table, such as indexes, is also included in the table size.
[0068] Subsequently, in step 704, in addition to the size of the pre-conversion table obtained in step 703, the management module 151 obtains the unit price of the archive area from the infrastructure cost table 155 and calculates the maintenance cost when archiving the pre-conversion table. In an embodiment where data compression is performed during archiving, the average compression rate of the data may be obtained in advance from past operation results and multiplied by the compression rate when calculating the maintenance cost.
[0069] Subsequently, in step 705, the management module 151 obtains all the converted table IDs related to the pre-conversion table ID of the row from the archive candidate data lineage table 1001 from the converted table ID column 1003.
[0070] Subsequently, in step 706, the management module 151 obtains the sizes (storage area usage amounts) of all the converted tables corresponding to all the converted table IDs obtained in step 705 from the publish data area 104. Here, the size of metadata related to the table such as the index is also regarded as the size of the table.
[0071] Subsequently, in step 707, in addition to the sizes of the converted tables obtained in step 706, the management module 151 obtains the unit price of the archive area from the infrastructure cost table 155 and calculates the maintenance cost of all the converted tables. In an embodiment where data compression is performed during archiving, the average compression rate of the data may be obtained in advance from past operation results and multiplied by the compression rate when calculating the maintenance cost here.
[0072] Subsequently, in step 708, the management module 151 obtains all the above-mentioned converted table IDs and the data processing program name corresponding to the pre-conversion table ID from the data lineage table 153.
[0073] Subsequently, in step 709, in addition to the data processing program name obtained in step 708, the management module 151 obtains the set values of the CPU, main memory, and secondary memory corresponding to the data processing program from the data conversion infrastructure setting table 154.
[0074] Subsequently, in step 710, in addition to the set values of the CPU, main memory, and secondary memory corresponding to the data processing program obtained in step 709, the management module 151 obtains the unit prices of the CPU, main memory, and secondary memory from the infrastructure cost table 155 and calculates the recreation cost of all the converted tables.
[0075] Subsequently, in step 711, the management module 151 compares the sum of the maintenance cost when archiving the pre-conversion table obtained in step 704 and the recreation cost of all the post-conversion tables obtained in step 709 with the maintenance cost when archiving the post-conversion table obtained in step 707. If the latter cost is lower, in step 712, the management module 151 archives the post-conversion table. On the other hand, if the former cost is lower, in step 713, the management module 151 archives the pre-conversion table and ends the process.
[0076] Through the above processing, efficient data storage becomes possible. Note that it is also possible to compare only one of the maintenance cost in the storage area of the pre-processed data and the recreation cost of the post-conversion table with the maintenance cost of the post-processed data. Without considering the maintenance cost, it is also possible to compare the recreation cost with a predetermined value. Alternatively, without considering the cost or based on a perspective different from the cost, the table to be archived may be selected from the load data or the publish data.
[0077] When the pre-conversion table (load data) is archived, the pre-conversion table is deleted from the load data area 102. When all (one or more) of the pre-conversion tables used for generating the post-conversion table are archived, the post-conversion table is deleted from the publish data area 104 without being archived.
[0078] The load data area 102 sequentially stores new load data. Therefore, old load data can be overwritten in the load data area 102. When one converted table is generated from a plurality of pre-conversion tables and some of the pre-conversion tables are not archived, the converted table may be archived, that is, deleted from the publish data area 104 and stored in the archive area 141 for publish data. This enables efficient data storage. In other examples, the converted table may be maintained in the publish data area 104. When the load data is stored anywhere, the converted table may be deleted from the publish data area 104 without being archived because it is a table with few recent accesses.
[0079] For example, in an embodiment of further reducing the usage amount of the storage system, in step 713, the converted table obtained by converting the pre-conversion table is deleted. The converted table obtained by converting the pre-conversion table is obtained from the data pedigree table 153. In another example, in step 713, if there is a converted table with a high access frequency among the converted tables obtained by converting the pre-conversion table, the converted table is instead archived. The converted table obtained by converting the pre-conversion table is obtained from the data pedigree table 153. Also, those with a high access frequency refer to, for example, the case where the values in column 202 or column 203 of the access frequency table 152 are included in the top 10.
[0080] FIG. 11 shows a flowchart of the process when the request receiver 106 receives a request from the application 107 or 171.
[0081] In step 1102, the request receiver 106 searches for the table requested from the application 107 or 171 in the publish data area 104 or the archive area 141 for publish data.
[0082] Next, in step 1103, the request receiver 106 determines whether it was found in step 1102. If found in the determination step 1103, in step 1131, the request receiver 106 reads the table from the publish data area 104 and the archive area 141 for publish data.
[0083] If not found in the determination step 1103, in step 1104, the request receiver 106 obtains the pre-conversion table ID of the table and the data processing program from the data lineage table 153.
[0084] Next, in step 1105, the request receiver 106 obtains the set values of the CPU, main memory, secondary memory, and batch size corresponding to the data processing program from the data conversion infrastructure setting table 154.
[0085] Next, in step 1106, the request receiver 106 sets the CPU, main memory, and secondary memory in the on-demand data processor 131 and starts up the data processing program.
[0086] Next, in step 1107, the request receiver 106 searches for the table requested from the application 107 or 171 in the archive area 121 for load data.
[0087] Next, in step 1108, the request receiver 106 reads data in the batch size from the archive area 121 for load data and processes it with the on-demand data processor 131.
[0088] Next, in step 1109, the request receiver 106 returns the result to the application 107 or 171 and ends the process.
[0089] In the above example, the on-demand data processor 131 executes processing with the same amount of resources as the online data processor 103. Thereby, requests from the application can be efficiently and reliably satisfied. Note that the amount of resources allocated to the on-demand data processor 131 may be different from the amount of resources allocated to the online data processor 103. For example, the amount of resources allocated to the on-demand data processor 131 may be less than the amount of resources allocated to the online data processor 103 within a threshold value.
[0090] FIG. 12 shows a setting screen 1201 of the information lifecycle management 105. The information lifecycle management setting screen 1201 has setting menus 1211, 1212, and 1213 corresponding to the data lineage table 153, the data conversion infrastructure setting table 154, and the infrastructure cost table 155, respectively.
[0091] For example, when the user selects the setting menu 1211, the data lineage table 153 is displayed on a display device within the system. The user can update the information in the data lineage table 153. Similarly, when the setting menu 1212 or 1213 is selected, the data conversion infrastructure setting table 154 or the infrastructure cost table 155 is displayed, and the information can be updated by the user.
Example
[0092] In the second embodiment, a method for streamlining user settings of the data lineage table 153 and the data conversion infrastructure setting table 154 is disclosed using a development environment (such as an Integrated Development Environment).
[0093] Figure 13 shows the overall configuration of the embodiment. In this embodiment, a development environment 1301 is added to the configuration of Embodiment 1. The development environment 1301 used for developing an online data processing program for implementing the online data processor 103 is also connected to the data lineage table 153 and the data conversion infrastructure setting table 154. The development environment 1301 automatically creates the data lineage table 153 and the data conversion infrastructure setting table 154 from the data input by the user for developing the online data processing program.
[0094] Note that the present invention is not limited to the above-described embodiments, and various modifications are included. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations is possible. Also, each of the above configurations, functions, processing units, etc. may be realized in hardware by designing a part or all of them, for example, with an integrated circuit. Also, each of the above configurations, functions, etc. may be realized in software by a processor interpreting and executing a program for realizing each function. Information such as a program, table, file, etc. for realizing each function can be placed in a memory, a recording device such as a hard disk, SSD, or a recording medium such as an IC card, SD card. Also, control lines and information lines are shown as those considered necessary for explanation, and not all control lines and information lines are necessarily shown on the product. In reality, it may be considered that almost all configurations are interconnected.
Explanation of Reference Numerals
[0095] 101 Input 102 Load Data Area 103 Online Data Processor 104 Publish Data Area 121 Archive area for raw data 131 On-demand data processor 141 Archive area for published data 108 Archive creator 105 Information lifecycle management 106 Request receiver 107, 171 Application 152 Access frequency table 153 Data lineage table 154 Data conversion infrastructure setting table 155 Infrastructure cost table
Claims
1. A pre - processing data area for storing pre - processing data, A post - processing data area for storing post - processing data generated by processing the pre - processing data, A pre - processing data archive area for archiving the pre - processing data moved from the pre - processing data area, A post - processing data archive area for archiving the post - processing data moved from the post - processing data area, One or more computing devices, and includes, The one or more computing devices, Select one or more pre - processing data that are the sources of one or more post - processing data of archive candidates, Select first pre - processing data that meets a preset condition from the one or more pre - processing data, and move it from the pre - processing data area to the pre - processing data archive area, Delete first post - processing data that is generated using the first pre - processing data and meets a preset condition from the post - processing data area without archiving, The condition for selecting the first pre - processing data includes that the first pre - processing data is the source of multiple types of post - processing data, The one or more computing devices generate the first post - processing data using the first pre - processing data in the pre - processing data archive area in response to an access request to the first post - processing data from an application, and return it to the application, a computer system.
2. The computer system according to claim 1, wherein, A preset index indicating the access frequency of each of the post - processing data of the archive candidates is less than a threshold value, a computer system.
3. The computer system according to claim 1, wherein, It includes cost management information for managing the cost in the use of the computer system, The one or more computing devices, Refer to the cost management information to determine the cost of maintaining the first pre - processed data in the pre - processed data archive area. Refer to the cost management information to determine the cost of maintaining the processed data generated using the first pre - processed data in the processed data archive area. A computer system, wherein the condition for selecting the first pre - processed data includes that the cost of maintaining the first pre - processed data is less than the cost of maintaining the processed data.
4. The computer system according to claim 1, including cost management information for managing the costs in the use of the computer system, wherein the one or more arithmetic units refer to the cost management information to determine the cost of regenerating the processed data generated using the first pre - processed data by using the first pre - processed data. refer to the cost management information to determine the cost of maintaining the processed data generated using the first pre - processed data in the processed data archive area. A computer system, wherein the condition for selecting the first pre - processed data includes that the regeneration cost is less than the cost of maintaining the processed data.
5. The computer system according to claim 4, wherein the one or more arithmetic units refer to the cost management information to determine the cost of maintaining the first pre - processed data in the pre - processed data archive area. A computer system, wherein the condition for selecting the first pre - processed data includes that the sum of the cost of maintaining the first pre - processed data and the regeneration cost is less than the cost of maintaining the processed data.
6. The computer system according to claim 1, The computer system, wherein the one or more arithmetic units regenerate the first processed data from the first pre-processed data in the pre-processed data archive area with reference to the setting of the resource allocation amount for generating the first processed data from the first pre-processed data in the pre-processed data area. **Claim 7** The computer system according to claim 1, wherein the pre-processed data used for generating the first processed data is only the first pre-processed data. **Claim 8** The computer system according to claim 1, wherein the second processed data is generated using a plurality of pre-processed data, and when only a part of the plurality of pre-processed data is stored in the pre-processed data archive area, the one or more arithmetic units move the second processed data from the processed data area to the processed data archive area. **Claim 9** A data management method in a system, wherein the system includes a pre-processed data area for storing pre-processed data, a processed data area for storing processed data generated by processing the pre-processed data, a processed data archive area for archiving the processed data moved from the processed data area, and a pre-processed data archive area for archiving the pre-processed data moved from the pre-processed data area, and the method includes the system selecting one or more pre-processed data that are the generation sources of one or more archive candidate processed data, selecting first pre-processed data that satisfies a preset condition from the one or more pre-processed data, and moving the first pre-processed data from the pre-processed data area to the pre-processed data archive area. Delete the first post - processing data generated using the first pre - processing data and satisfying preset conditions from the post - processing data area without archiving it. The conditions for selecting the first pre - processing data include that the first pre - processing data is the source of multiple types of post - processing data. The method is a data management method in which the system generates the first post - processing data using the first pre - processing data in the pre - processing data archive area in response to an access request from an application to the first post - processing data and returns it to the application.
Citation Information
Patent Citations
System and method for operating a big-data platform
WO2013070873A1