Computer systems, methods for tracking the history of data, and programs
The integration of a data history management system with database and infrastructure systems addresses the challenge of tracking data storage across separate systems, ensuring comprehensive data location management and policy compliance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2026-04-01
AI Technical Summary
In systems where database management systems and infrastructure systems manage data separately, tracking the storage location of data becomes difficult due to independent operations like copying and backing up, leading to a fragmented understanding of data whereabouts.
A computer system that integrates a data history management system with database and infrastructure systems to track data existence periods, storage areas, and replication operations, using identification information and configuration history to comprehensively locate data across multiple storage areas.
Enables comprehensive tracking of data storage locations, facilitating better data management and compliance with storage policies by identifying and managing data across separate database and infrastructure systems.
Smart Images

Figure 0007839021000001 
Figure 0007839021000002 
Figure 0007839021000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for tracking the history of data.
Background Art
[0002] From the viewpoints of protecting personal information and strengthening security measures, strict management of confidential data has become important. In the management of confidential data, it is necessary to grasp the storage location of the confidential data. In contrast, the technique described in Patent Document 1 is known.
[0003] Patent Document 1 describes that "in a computer system including a metadata management server and a history management server that manages history, each of the servers is connected to a client computer that holds the file, and the metadata management server detects an event in which the metadata of the file held by the client computer is changed, requests the history management server to search for the history related to the detected event, and based on the history search result transmitted from the history management server, extracts a derived second file derived from a first file whose metadata has been changed and a third file that is the source from which the first file is derived, and among the extracted second file and third file, identifies the file for which the metadata should be updated, and generates a metadata update request for updating the metadata of the identified file."
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] By using the technology described in Patent Document 1, files and derived files can be identified.
[0006] In recent system configurations, the database management system that manages the data and the infrastructure system that manages the storage area where the data is stored exist separately. Therefore, the infrastructure system performs tasks such as copying and backing up the storage area independently of the data management system. The infrastructure system manages the history of the storage area configuration, but does not manage the data stored in the storage area.
[0007] Therefore, simply combining the operation history of the database management system and the operation history of the underlying system makes it difficult to understand the location of all data.
[0008] This invention provides a technology for comprehensively tracking the storage locations of data in a system that manages data using a database management system and an underlying system. [Means for solving the problem]
[0009] A representative example of the invention disclosed in this application is as follows: A computer system comprising at least one computer having a processor, a storage device connected to the processor, and a network interface connected to the processor, connected to at least one database management system for managing data, and at least one infrastructure system for managing data storage areas, wherein the computer system receives a tracking request including identification information of target data, calculates the data existence period during which the target data was managed by the at least one database management system based on information about the data managed by the at least one database management system, identifies a first storage area that was provided to the at least one database management system and where the target data was stored during the data existence period, based on configuration history information regarding the configuration of the storage area in the at least one infrastructure system,The first storage area is registered in the tracking list, Based on the data existence period and the history of replication operations of the storage area of at least one base system, a tracking process is performed to track the second storage area where the target data is stored by replication operations, starting from the first storage area, and information regarding the first and second storage areas is output as the storage area where the target data is stored. The computer system, in the tracking process, selects one target storage area from the tracking list, calculates the existence period of the target storage area, calculates the replication operation execution period based on the history of replication operations related to the target storage area of at least one base system, identifies the second storage area based on the data existence period, the existence period of the target storage area, and the replication operation execution period, and registers it in the tracking list. . [Effects of the Invention]
[0010] According to the present invention, in a system that manages data using a database management system and an underlying system, the storage location of the data can be comprehensively tracked. Problems, configurations, and effects other than those described above will be clarified by the following description of embodiments. [Brief explanation of the drawing]
[0011] [Figure 1] This figure shows an example of the system configuration in Example 1. This figure shows an example of the computer configuration in Example 1. [Figure 2] This figure shows an example of the information included in the database management system information of Example 1. [Figure 3] This figure shows an example of the information included in the basic system information of Example 1. [Figure 4] This figure shows an example of the information included in the basic system information of Example 1. [Figure 5] This figure shows an example of the information included in the basic system information of Example 1. [Figure 6] This figure shows an example of the information included in the basic system information of Example 1. [Figure 7] This figure shows an example of the information included in the basic system information of Example 1. [Figure 8] This figure shows an example of the information included in the basic system information of Example 1. [Figure 9] This figure shows an example of the information included in the basic system information of Example 1. [Figure 10A]This is a diagram showing an example of the information included in the tracking policy information of Example 1. [Figure 10B] This is a diagram showing an example of the information included in the tracking policy information of Example 1. [Figure 10C] This is a diagram showing an example of the information included in the tracking policy information of Example 1. [Figure 11] This is a flowchart for explaining an example of the data tracking process executed by the data history management system of Example 1. [Figure 12] This is a diagram showing an example of the tracking list generated by the data history management system of Example 1. [Figure 13A] This is a flowchart for explaining an example of the related volume tracking process executed by the data history management system of Example 1. [Figure 13B] This is a flowchart for explaining an example of the related volume tracking process executed by the data history management system of Example 1.
[0014] The designations "First," "Second," "Third," etc., used in this specification are for the purpose of identifying constituent elements and do not necessarily limit their number or order.
[0015] The positions, sizes, shapes, and ranges of each component shown in the drawings, etc., may not represent the actual positions, sizes, shapes, and ranges, etc., in order to facilitate understanding of the invention. Therefore, the present invention is not limited to the positions, sizes, shapes, and ranges, etc., disclosed in the drawings, etc. [Examples]
[0016] Figure 1 shows an example of the system configuration of Example 1. Figure 2 shows an example of the computer configuration of Example 1.
[0017] The system includes a data history management system 100, a database management system 101, and an infrastructure system 102. Each system is connected to the others via a network 105 such as a LAN (Local Area Network) or WAN (Wide Area Network). The connection method of the network 105 may be either wired or wireless.
[0018] Note that there may be two or more database management systems 101. Also, there may be two or more base systems 102.
[0019] The database management system 101 is a system that manages a database that stores user data. The database management system 101 consists of a computer having a processor, main memory, and a network interface. The database management system 101 has a database management unit 130 that manages the database. The database management unit 130 stores data in a storage area (volume 141) provided by the base system 102. Although the database management system 101 holds various control information, it is omitted here as it is not directly related to the invention.
[0020] Furthermore, the present invention is not limited to the type and size of data handled by the database management system 101. The data may include, for example, block data, tables, and files. Additionally, some data from tables and files may be tracked.
[0021] The base system 102 is a system that provides volume 141 to the database management system 101. The following configurations are possible for the base system 102: (Configuration 1) A system consisting of a computer having a processor, main memory, secondary memory (drives), and a network interface. (Configuration 2) A system consisting of a computer having a processor, main memory, and a network interface, as well as a drive box equipped with a group of drives. The base system 102 in configuration (2) is a so-called storage system.
[0022] The base system 102 includes a base management unit 140 that controls the creation, duplication, and deletion of volume 141, as well as the allocation of volume 141 to the database management system 101. Although the base system 102 holds various control information, it is omitted here as it is not directly related to the invention.
[0023] The infrastructure management unit 140 generates a drive group that constitutes a RAID (Redundant Arrays of Inexpensive Disks) from multiple drives, and generates a volume 141 from the drive group.
[0024] The data history management system 100 is a system that manages the history of data handled by the database management system 101. Data History Management System 100 For example, it consists of a computer 200 as shown in Figure 2.
[0025] Computer 200 includes a processor 201, main memory 202, secondary memory 203, and a network interface 204. Each hardware element is connected to the others via a bus 205. Computer 200 may also have input devices such as a keyboard, mouse, and touch panel, as well as output devices such as a display.
[0026] The processor 201 executes the program stored in the main memory 202. The processor 201 in Embodiment 1 functions as an information acquisition unit 110 and a tracking unit 111 by executing processing according to the program. Note that, for each functional unit of the base system 102, multiple functional units may be combined into a single functional unit, or a single functional unit may be divided into multiple functional units according to its function.
[0027] The main memory 202 is a memory unit that stores the program executed by the processor 201 and the information used by the program. In Embodiment 1, the main memory 202 stores the program that implements the information acquisition unit 110 and the tracking unit 111, respectively, and also stores the database management system information 120, the base system information 121, and the tracking policy information 122.
[0028] The secondary storage device 203 is an HDD (Hard Disk Drive) or SSD (Solid State Drive), etc., and permanently stores a large amount of data. Programs and information stored in the main memory 202 may also be stored in the secondary storage device 203. In this case, the processor 201 reads the programs and information from the secondary storage device 203 and loads them into the main memory 202.
[0029] The network interface 204 communicates with external devices via the network.
[0030] Next, we will describe the details of the database management system information 120, the underlying system information 121, and the tracking policy information 122.
[0031] The information acquisition unit 110 acquires information about data managed by the database management system 101 and information about operations on that data from the database management system 101, and stores it in the database management system information 120. The information acquisition unit 110 also acquires information about volume 141 managed by the base system 102 and information about operations on volume 141 from the base system 102, and stores it in the base system information 121.
[0032] The information acquisition unit 110 may directly collect information about data and information about operations on data from the database management system 101, or it may collect it indirectly from another system. For example, it may collect information from a system that manages multiple data and databases together, such as a data catalog system. In addition, information about data operations may be collected from operation logs, or it may be information about operations inferred from changes in the state and configuration of data that are managed periodically. Similarly, information about volume 141 and information about operations on volume 141 may be collected indirectly from a system other than the base system 102. In addition, information about operations inferred from the state and configuration of the base system may be collected.
[0033] This implementation can be used with on-premises systems, public clouds, hybrid clouds, and multi-cloud environments.
[0034] Figure 3 shows an example of the information contained in the database management system information 120 of Example 1. Figures 4, 5, 6, 7, 8, and 9 show an example of the information contained in the base system information 121 of Example 1. Figures 10A, 10B, and 10C show an example of the information contained in the tracking policy information 122 of Example 1.
[0035] Database management system information 120 is information used to manage information about the data handled by the database management system 101.
[0036] The database management system information 120 includes, for example, table 300. Table 300 is information for managing the data handled by the database management system 101 so far, and stores entries that include data UUID 301, data ID 302, DB ID 303, creation date and time 304, and deletion date and time 305. There is one entry for each combination of data and the time the data was handled. Note that the fields included in the entries are not limited to those mentioned above. It may not include any of the fields mentioned above, or it may include other fields.
[0037] Data UUID301 is a field that stores identification information to uniquely identify data handled by the database management system 101. Data ID302 is a field that stores identification information to identify data by the database management system 101. DB ID303 is a field that stores identification information of the database management system 101 that managed the data. Creation date and time304 is a field that stores the date and time the data was created by the database management system 101. Deletion date and time305 is a field that stores the date and time the data was deleted by the database management system 101.
[0038] The data corresponding to entries where the deletion date 305 is blank indicates that the data is still being managed by the database management system 101.
[0039] It is also possible to manage regular data deletion and data deletion in a format that is difficult to recover (wiping) separately. In this case, the entry can include a field to store the time of the wipe.
[0040] The base system information 121 is information for managing the configuration of volume 141 and control logs for volume 141. The base system information 121 includes, for example, tables 400, 500, 600, 700, 800, and 900.
[0041] The table 400 shown in Figure 4 contains information for managing volume 141 that the infrastructure system 102 has managed to date, and stores entries including volume UUID 401, volume ID 402, infrastructure ID 403, creation date and time 404, and deletion date and time 405. There is one entry for each combination of volume 141 and the time when volume 141 was managed. Note that the fields included in the entries are not limited to those described above. It is not necessary to include any of the fields described above, and it may also include other fields.
[0042] Volume UUID 401 is a field that stores identification information to uniquely identify volume 141 that the underlying system 102 has managed to date. Volume ID 402 is a field that stores identification information for the underlying system 102 to identify volume 141. Underground ID 403 is a field that stores identification information for managing volume 141. Creation date and time 404 is a field that stores the date and time when volume 141 was created by the underlying system 102. Deletion date and time 405 is a field that stores the date and time when volume 141 was deleted by the underlying system 102.
[0043] The deletion timestamp 405 indicates that volume 141, which corresponds to the blank entry, is still managed by the underlying system 102.
[0044] Alternatively, the deletion of a regular volume 141 and the deletion of data in a format that is difficult to recover (wiping) can be managed separately. In this case, the entry can include a field to store the time of the wipe.
[0045] The table 500 shown in Figure 5 contains information for managing the correspondence between volume 141 and the drive group that provides the storage area constituting volume 141. It stores entries that include volume UUID 501, drive group ID 502, drive ID 503, base ID 504, creation date and time 505, and deletion date and time 506. There is one entry for volume 141. Note that the fields included in the entries are not limited to those described above. It may not include any of the fields described above, or it may include other fields.
[0046] Volume UUID 501 is the same field as Volume UUID 401. Drive Group ID 502 is a field that stores identification information for the drive group that provides the storage area constituting Volume 141. Drive ID 503 is a field that stores identification information for the drives included in the drive group. Drive ID 503 stores the identification information for all drives included in the drive group in list format. Base ID 504 is a field that stores identification information for the base system 102 that manages the drive group. Creation Date and Time 505 is a field that stores the date and time when the drive group was created by the base system 102. Deletion Date and Time 506 is a field that stores the date and time when the drive group was deleted by the base system 102.
[0047] The table 600 shown in Figure 6 contains information for managing virtual volumes and stores entries that include the volume UUID 601, external volume UUID 602, start date and time 603, and end date and time 604. There is one entry for each virtual volume. Note that the fields included in the entry are not limited to those mentioned above. It may not include any of the fields mentioned above, or it may include other fields.
[0048] A virtual volume is a volume that one infrastructure system 102 provides as its own managed volume 141, which is managed by another infrastructure system 102. When an I / O request is received for the virtual volume, data is sent and received between the infrastructure systems 102.
[0049] Volume UUID 601 and External Volume UUID 602 are the same fields as Volume UUID 401. However, Volume UUID 601 stores the Volume UUID of the virtual volume, and External Volume UUID 602 stores the UUID of Volume 141, which is the actual virtual volume. Start Date and Time 603 is the field that stores the date and time when the provision of the virtual volume began. End Date and Time 604 is the field that stores the date and time when the provision of the virtual volume ended.
[0050] The table 700 shown in Figure 7 contains information for managing the allocation of volume 141 to the database management system 101, and stores entries including DB ID 701, volume UUID 702, start date and time 703, and end date and time 704. There is one entry for each combination of database management system 101 and volume 141. Note that the fields included in the entries are not limited to those described above. It may not include any of the fields described above, or it may include other fields.
[0051] DB ID 701 is the same field as DB ID 303. Volume UUID 702 is the same field as Volume UUID 401. Start Date & Time 703 is the field that stores the date and time when the provision of volume 141 to the database management system 101 began. End Date & Time 704 is the field that stores the date and time when the provision of volume 141 to the database management system 101 ended.
[0052] The table 800 shown in Figure 8 contains information for managing the log of replication operations of volume 141 in the base system 102, and stores entries that include volume UUID (Source) 801, volume UUID (Destination) 802, operation type 803, start date and time 804, and end date and time 805. There is one entry for each replication operation. Note that the fields included in the entries are not limited to those described above. It may not include any of the fields described above, or it may include other fields.
[0053] Volume UUID(Source)801 and Volume UUID(Destination)802 are the same fields as Volume UUID401. However, Volume UUID(Source)801 stores the Volume UUID of the source volume 141, and Volume UUID(Destination)802 stores the UUID of the destination volume 141. Operation type 803 is a field that stores the type of replication operation. Operation type 803 stores one of the following: "Temporary replication," which represents the replication of data at any point in time; "Snapshot," which represents the replication of the state of volume 141 at any point in time; or "Permanent replication," which represents synchronous replication such as mirroring and replication. Note that the operation types mentioned above are just examples and are not limited to them. Start date and time 804 is a field that stores the date and time when the replication process for volume 141 started. End date and time 805 is a field that stores the date and time when the replication process for volume 141 ended.
[0054] The table 900 shown in Figure 9 contains information for managing the log of the migration operation of volume 141 in the base system 102, and stores entries that include the volume UUID 901, drive group ID (Source) 902, drive group ID (Destination) 903, start date and time 904, and end date and time 905. There is one entry for each migration operation. Note that the fields included in the entries are not limited to those described above. It is not necessary to include any of the fields described above, and it may also include other fields.
[0055] Volume UUID 901 is the same field as Volume UUID 401. Drive Group ID (Source) 902 and Drive Group ID (Destination) 903 are the same fields as Drive Group ID 502. However, Drive Group ID (Source) 902 stores the identification information of the source drive group, and Drive Group ID (Destination) 903 stores the identification information of the destination drive group. Start Date and Time 904 is the field that stores the date and time when the movement of Volume 141 started. End Date and Time 905 is the field that stores the date and time when the movement of Volume 141 finished.
[0056] As explained above, tables 400, 500, 600, and 700 contain historical information regarding the configuration of volume 141 in the base system 102, while tables 800 and 900 contain historical information regarding various operations performed on volume 141 in the base system 102.
[0057] Tracking policy information 122 is information for managing policies to track volume 141 and drives, which are the data storage locations. Tracking policy information 122 includes, for example, tables 1000, 1010, and 1020.
[0058] The table 1000 shown in Figure 10A is provided to the base system 102 and has a replication relationship with volume 141 (first-tier volume 141) that stores the data to be tracked. It also contains information for managing the policy for searching for volume 141 (second-tier volume 141) where the data is stored. The table 1000 stores entries that include operation type 1001, storage conditions 1002, and tracking conditions 1003. There is one entry for each type of replication operation. Note that the fields included in an entry are not limited to those described above. It may not include any of the fields described above, or it may include other fields.
[0059] Operation type 1001 is a field that stores the type of replication operation. Storage condition 1002 is a field that stores conditions for determining whether the source volume 141 has a replication relationship with the source volume 141 and whether the data to be tracked is stored there. Tracking condition 1003 is a field that stores conditions for determining whether the source volume 141 has a replication relationship with the source volume 141 and whether the volume 141 is the volume to be tracked.
[0060] Tables 1010 and 1020, shown in Figures 10B and 10C, contain information for managing policies for searching for volumes 141 that have a replication relationship with volumes 141 in the second or higher hierarchical levels and where the data to be tracked is stored.
[0061] Tables 1010 and 1020 are switched depending on the type of replication operation at each tier. Table 1010 is used when the replication operation type for volume 141 at each tier related to volume 141 under evaluation is constant replication. Table 1020 is used when the replication operation type for volume 141 at at least one tier related to volume 141 under evaluation is not constant replication.
[0062] Operation type 1011 and operation type 1021 are the same fields as operation type 1001. Storage condition 1012 and storage condition 1022 are the same fields as storage condition 1002. Tracking condition 1013 and tracking condition 1023 are the same fields as tracking condition 1003.
[0063] Storage conditions 1002, 1012, 1022 and tracking conditions 1003, 1013, 1023 The conditions set are defined based on the data's existence period, the availability period of volume 141, and the execution period of the replication process. The data's existence period is calculated from the creation date 304 and deletion date 305. The volume 141's existence period is calculated from the creation date 404 and deletion date 405. Replication process The execution period is calculated from the start date and time 804 and the end date and time 805.
[0064] Based on conditions using the data's existence period, the availability period of volume 141, and the execution period of the replication process, it is possible to comprehensively track the volumes where data is stored. Furthermore, by considering the replication operation tree and switching the conditions used, it is possible to comprehensively track volumes according to the characteristics of each type of replication operation.
[0065] The above describes the information held by the data history management system 100. Next, we will describe the data tracking process performed by the data history management system 100. Figure 11 is a flowchart illustrating an example of the data tracking process performed by the data history management system 100 in Example 1. Figure 12 is a diagram showing an example of a tracking list generated by the data history management system 100 in Example 1.
[0066] The tracking unit 111 receives a tracking request from the user that includes identification information (data UUID) of the data to be tracked (target data) (step S101). The tracking unit 111 may also refer to table 300 to display the data UUID of the traversable data.
[0067] The tracking unit 111 calculates the existence period of the target data based on the database management system information 120 (step S102).
[0068] Specifically, the tracking unit 111 refers to the table 300 and searches for an entry in which the data UUID of the target data is stored in data UUID 301. The tracking unit 111 calculates the data's existence period based on the creation date and time 304 and deletion date and time 305 of the found entry. If the deletion date and time 305 is blank, the tracking unit 111 calculates the data's existence period based on the creation date and time 304 and the current date and time.
[0069] The tracking unit 111 searches for the volume 141 (first-tier volume 141) that was provided to the database management system 101 and in which the target data was stored, based on the data's existence period and the underlying system information 121 (step S103).
[0070] Specifically, the tracking unit 111 refers to table 700 and searches for an entry in which DB ID 701 is set to the value of DB ID 303 of the entry identified in step S102. The tracking unit 111 calculates a determination period from the creation date and time 404 and deletion date and time 405 of the found entry. If the determination period includes the existence period of the data, the tracking unit 111 identifies the volume 141 of the corresponding entry as the first-tier volume 141.
[0071] The tracking unit 111 registers the first layer volume 141 in the tracking list 1200 (step S104).
[0072] The tracking list 1200 searches for entries containing hierarchy 1201, volume UUID (Destination) 1202, volume UUID (Source) 1203, operation type 1204, tracking flag 1205, and storage flag 1206. There is one entry for each target volume 141. Note that the fields included in an entry are not limited to those described above. It may not include any of the fields described above, or it may include other fields.
[0073] The hierarchy 1201 is a field that stores the hierarchy of the replication relationship starting from the first-tier volume 141. Here, the hierarchy of the first-tier volume 141 is set to 1. Volume UUID(Destination)1202 and Volume UUID(Source)1203 are the same fields as Volume UUID401. However, Volume UUID(Destination)1202 stores the UUID of the searched volume 141, and Volume UUID(Source)1203 stores the UUID of the source volume 141 from which the searched volume 141 was replicated. Operation type 1204 is a field that stores the type of replication operation performed between the two volumes 141.
[0074] The tracking flag 1205 is a field that stores a flag used to determine whether or not to start the search from volume 141, which corresponds to volume UUID (Destination) 1202. If the search is to start from volume 141, the tracking flag 1205 is set to "1". If the search is not to start from volume 141, the tracking flag 1205 is set to "0".
[0075] The storage flag 1206 is a field that stores a flag indicating whether or not the target data is stored in volume 141, which corresponds to volume UUID (Destination) 1202. If the target data is stored in volume 141, the storage flag 1206 is set to "1", and if the target data is not stored in volume 141, the storage flag 1206 is set to "0".
[0076] In step S104, the tracking unit 111 adds an entry to the tracking list 1200, sets the hierarchy 1201 to "1", and sets the volume UUID of the first-tier volume 141 to volume UUID(Destination) 1202. The tracking unit 111 also sets the tracking flag 1205 to "1" and the storage flag 1206 to "1". Volume UUID(Source) 1203 and operation type 1204 are left blank.
[0077] The tracking unit 111 performs related volume tracking processing to track the volume 141 (related volume 141) that has a replication relationship with the first layer volume 141 and stores the target data (step S105). The related volume tracking processing will be described later.
[0078] The tracking unit 111 outputs the tracking results including the first layer volume 141 and related volumes 141 (step S106), and terminates the data tracking process.
[0079] Figure 13 is a flowchart illustrating an example of the related volume tracking process performed by the data history management system 100 of Example 1.
[0080] The tracking unit 111 sets the variable g, which represents the hierarchy, to an initial value of "1" (step S201).
[0081] The tracking unit 111 selects one entry (target volume 141) from among the entries in the tracking list 1200 whose value at hierarchy 1201 matches the value of variable g, and whose tracking flag 1205 is set to "1" (step S202). The volume 141 corresponding to the volume UUID (Destination) 1202 of the entry becomes the target volume 141.
[0082] The tracking unit 111 calculates the lifetime of the target volume 141 (step S203).
[0083] Specifically, the tracking unit 111 refers to table 400 and searches for an entry in volume UUID 401 where the volume UUID of the target volume is stored. The tracking unit 111 calculates the existence period of volume 141 based on the creation date and time 404 and deletion date and time 405 of the searched entry. If the deletion date and time 405 is blank, the tracking unit 111 calculates the existence period of the volume based on the creation date and time 404 and the current date and time.
[0084] The tracking unit 111 searches for a replicated volume 141 that uses the target volume 141 as the source, based on the underlying system information 121 (step S204).
[0085] Specifically, the tracking unit 111 refers to table 800 and searches for an entry where the volume UUID (Source) 801 is set to the value of the target volume 141's volume UUID (Destination) 1202. If the replicated volume 141 does not exist, the tracking unit 111 proceeds to step S211.
[0086] Furthermore, the tracking unit 111 adds an entry to the tracking list 1200 and sets the value of variable g plus 1 to the hierarchy 1201. The tracking unit 111 sets the volume UUID of the replica volume 141 to the volume UUID (Destination) 1202 of the added entry and sets the volume UUID of the target volume 141 to the volume UUID (Source) 1203. In addition, the tracking unit 111 sets the operation type 1204 to ,search The value of the operation type 803 for the added entry is set. The tracking unit 111 also sets the tracking flag 1205 and storage flag 1206 of the added entry to "0".
[0087] The tracking unit 111 selects one copy volume 141 from among the searched copy volumes 141 (step S205).
[0088] The tracking unit 111 determines whether the selected replica volume 141 satisfies the storage conditions (step S206). Specifically, the following processes are performed.
[0089] (S206-1) The tracking unit 111 calculates the execution period of the replication process based on the start date and time 804 and end date and time 805 of the entry retrieved in step S204.
[0090] (S206-2) The tracking unit 111 obtains the storage conditions to be used from the tracking policy information 122.
[0091] If variable g is "1", the tracking unit 111 refers to table 1000 and searches for an entry in which the value of operation type 803 of the searched entry is set for operation type 1001.
[0092] If variable g is not "1", the tracking unit 111 refers to the tracking list 1200 to track the replication relationships of the target volume 141. If all operation types are "steady replication", the tracking unit 111 refers to table 1010 to search for an entry in which operation type 1011 is set to the value of operation type 803 of the searched entry. If at least one operation type is not "steady replication", the tracking unit 111 refers to table 1020 to search for an entry in which operation type 1021 is set to the value of operation type 803 of the searched entry.
[0093] (S206-3) The tracking unit 111 determines whether the acquired storage conditions are met based on the data existence period, the existence period of volume 141, and the execution period of the replication process.
[0094] The above is a description of the process in step S206.
[0095] If the selected replica volume 141 does not meet the storage conditions, the tracking unit 111 proceeds to step S208.
[0096] If the selected replica volume 141 satisfies the storage conditions, the tracking unit 111 sets the storage flag 1206 to "1" (step S207), and then proceeds to step S208.
[0097] The tracking unit 111 determines whether the selected copy volume 141 satisfies the tracking conditions (step S208).
[0098] Specifically, the tracking unit 111 searches for entries from the tracking policy information 122 using the same procedure as in S206-2 and obtains tracking conditions. The tracking unit 111 determines whether the tracking conditions are met based on the data existence period, the volume existence period 141, and the execution period of the replication process.
[0099] If the selected replica volume 141 does not meet the tracking conditions, the tracking unit 111 proceeds to step S210.
[0100] If the selected replica volume 141 satisfies the tracking conditions, the tracking unit 111 sets the tracking flag 1205 to "1" (step S209), and then proceeds to step S210.
[0101] In step S210, the tracking unit 111 determines whether processing has been completed for all the duplicate volumes 141 that were searched in step S204.
[0102] If processing is not complete for all replicated volumes 141, the tracking unit 111 returns to step S205 and performs the same processing.
[0103] Once processing is complete for all replicated volumes 141, the tracking unit 111 determines whether processing has been completed for all volumes 141 in hierarchy g (step S211).
[0104] If processing is not complete for all volumes 141 in hierarchy g, the tracking unit 111 returns to step S202 and performs the same processing.
[0105] When processing is complete for all volumes 141 in hierarchy g, the tracking unit 111 sets the variable g to a value obtained by adding 1 to the variable g (step S212).
[0106] The tracking unit 111 determines whether an entry for hierarchy g exists in the tracking list 1200 (step S213). That is, it determines whether an entry exists in hierarchy 1201 with a value set for variable g.
[0107] If an entry for hierarchy g exists in the tracking list 1200, the tracking unit 111 returns to step S202 and performs the same process.
[0108] If there is no entry for hierarchy g in the tracking list 1200, the tracking unit 111 terminates the related volume tracking process.
[0109] Figure 14 shows an example of the tracking results output by the data history management system 100 of Example 1.
[0110] The data history management system 100 can display tracking results as shown in Figure 14, based on the tracking list 1200, database management system information 120, and underlying system information 121. In the tracking results, volume 141 with storage flag 1206 set to "1" is displayed. The dotted box represents a deleted object. The data history management system 100 may perform the display, or it may output the display information to another device or system.
[0111] According to Example 1, the data history management system 100 can comprehensively track the volume 141 in which data is stored in a system that manages data using a database management system 101 and an infrastructure system 102.
[0112] The data history management system 100 may also accept operations such as data deletion, wiping, and moving, along with the tracking results.
[0113] The data history management system 100 maintains information regarding the data placement policy and the placement information of the underlying system 102 that provides volume 141. Based on the tracking results and the placement policy, it may determine whether volume 141, where the data is stored, violates the policy. For example, it may determine whether the underlying system 102 providing volume 141 is located in a specific country. It may also determine whether the underlying system 102 providing volume 141 is located in a public cloud. Furthermore, if the policy is violated, the data history management system 100 may perform operations such as deleting, wiping, and moving the data. This enables appropriate storage and management of the data.
[0114] Furthermore, the history of operations on the data may include a flag indicating whether or not it is subject to tracking. In this case, it is not necessary to set tracking conditions in the tracking policy information 122.
[0115] Furthermore, if the data is a file, you may want to manage the history of operations on the file's metadata. For example, you can track operations starting from the date and time when information indicating that the metadata is sensitive data was added.
[0116] Furthermore, if the base system 102 is object storage, the data will be stored packet You may want to track it. [Examples]
[0117] In Example 2, the treatment of the existence period of volume 141 differs from that in Example 1. Below, Example 2 will be described, focusing on the differences from Example 1.
[0118] The system configuration of Example 2 is the same as that of Example 1. The functional configurations of the data history management system 100, the database management system 101, and the base system 102 in Example 2 are the same as those in Example 1. Of the information held by the data history management system 100 in Example 2, table 300 is different. The data structure of the other information is the same as in Example 1.
[0119] The entries in Table 300 of Example 2 include the wipe date and time. In Example 2, data tracking is performed assuming that data remains until it is wiped. Data wiping may be performed on a data-by-data basis by the database management system 101, or it may be performed on a volume basis by the underlying system 102. If wiping is performed on a volume basis, the wipe date and time are recorded for all data (including deleted data) stored in volume 141.
[0120] The processing flow executed by the data history management system 100 in Example 2 is the same as in Example 1. However, in Example 2, in step S102, the data existence period is calculated based on the creation date and time 304 and the wipe date and time.
[0121] According to Example 2, volume 141, in which data is stored in a recoverable state, can also be tracked. [Examples]
[0122] The data history management system 100 in Example 3 also tracks the drives that provide storage space to the volume. The following describes Example 3, focusing on the differences from Example 1.
[0123] The system configuration of Example 3 is the same as that of Example 1. The functional configurations of the data history management system 100, the database management system 101, and the base system 102 in Example 3 are the same as those in Example 1. The information held by the data history management system 100 in Example 3 is the same as that in Example 1.
[0124] In the data tracking process of Example 3, the drive tracking process is executed after the process in step S105. Figure 15 is a flowchart illustrating an example of the drive tracking process performed by the data history management system 100 of Example 3. Figure 16 is a diagram showing an example of the drive group list generated by the data history management system 100 of Example 3.
[0125] The tracking unit 111 selects one volume 141 from the tracking list 1200 (step S301).
[0126] The tracking unit 111 refers to table 900 based on the data's existence period and retrieves the migration operation history of the selected volume 141 (step S302). Specifically, the following processes are performed.
[0127] (S302-1) The tracking unit 111 searches for an entry in volume UUID 901 that stores the UUID of the selected volume 141. If no entry exists, the process in step S302 is skipped.
[0128] (S302-2) The tracking unit 111 determines whether the selected volume 141 is the volume 141 of the first layer.
[0129] (S302-3) If the selected volume 141 is the first-tier volume 141, the tracking unit 111 selects from the identified entries an entry in which the start date and time of the migration operation is later than the start date and time of the data's existence period, and the end date and time of the migration operation is earlier than the end date and time of the data's existence period.
[0130] (S302-4) If the selected volume 141 is not the first-tier volume 141, the tracking unit 111 refers to the tracking list 1200 to track the replication relationships of the selected volume 141. If all operation types are "steady replication", the tracking unit 111 selects an entry in the same procedure as in S302-3. If at least one operation type is not "steady replication", the tracking unit 111 selects an entry from the identified entries in which the start date and time of the migration operation is later than the start date and time of the data's existence period.
[0131] The above is a description of the process in step S302.
[0132] The tracking unit 111 acquires the history of external connections to volume 141 (step S303). Specifically, the following processes are performed.
[0133] (S303-1) The tracking unit 111 searches for an entry in volume UUID 601 that stores the UUID of the selected volume 141. If no entry exists, the process in step S303 is skipped.
[0134] (S303-2) The tracking unit 111 obtains the migration operation history of volume 141 corresponding to the external volume UUID 602 of the searched entry. The method for obtaining the migration operation history is the same as in step S302.
[0135] The tracking unit 111 identifies the drive group associated with the selected volume 141 and registers it in the drive group list 1600 (step S304). Subsequently, the tracking unit 111 determines whether processing has been completed for all volumes 141 in the tracking list 1200 (step S305). If processing has not been completed for all volumes 141 in the tracking list 1200, the tracking unit 111 returns to step S301. If processing has been completed for all volumes 141 in the tracking list 1200, the tracking unit 111 terminates the drive tracking process.
[0136] The drive group list 1600 searches for entries containing volume UUID 1601 and list 1602. There is one entry for each volume 141. Note that the fields included in the entry are not limited to those mentioned above. It may not include any of the fields mentioned above, or it may include other fields.
[0137] Volume UUID 1601 is the same field as Volume UUID 401. List 1602 is a field that stores the identification information of the drive group associated with Volume 141. List 1602 stores the identification information of one or more drive groups.
[0138] The drive group where the selected volume 141 resides can be identified based on table 500. The drive group to which the selected volume 141 will be migrated can be identified based on the drive group ID (Destination) 903 in the migration operation history. In addition, the drive group where the externally connected volume 141 of the selected volume 141 resides can be identified based on table 500 and the migration operation history (table 900).
[0139] The tracking unit 111 adds an entry to the drive group list 1600 and sets the volume UUID of the selected volume 141 to volume UUID 1601. The tracking unit 111 also stores the identification information of the identified drive group in list 1602.
[0140] Figure 17 shows an example of the tracking results output by the data history management system 100 of Example 3.
[0141] The data history management system 100 displays the tracking results as shown in Figure 17, based on the tracking list 1200, drive group list 1600, database management system information 120, and underlying system information 121. Dotted boxes represent deleted objects. Dotted links indicate that the correspondence with volume 141 has been resolved.
[0142] The data history management system 100 may also be configured to display only drive groups.
[0143] Furthermore, the data history management system 100 may be configured to display the base system 102, including the drive group, as shown in Figure 17.
[0144] According to Example 3, the data history management system 100 can further track the drive groups in which the data is actually stored.
[0145] Furthermore, by using the history of volume 141, which was used as a cache during operations such as replication and migration, volume 141 can also be tracked. In addition to the operation history of the database management system 101 and the base system 102, the operation history of applications used by the user may also be acquired. This allows for the handling of replication by applications as well.
[0146] Although this embodiment describes a system that tracks the history of data, the present invention can also be applied to a system that tracks the history of containers.
[0147] It should be noted that the present invention is not limited to the embodiments described above, and various modifications are included. Furthermore, for example, the embodiments described above are detailed explanations of the configuration in order to clearly illustrate the present invention, and are not necessarily limited to those having all the configurations described. In addition, some of the configurations in each embodiment can be added to, deleted from, or replaced with other configurations.
[0148] Furthermore, each of the above-mentioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, in whole or in part, for example, by designing them as integrated circuits. The present invention can also be implemented by software program code that realizes the functions of the embodiment. In this case, a storage medium on which the program code is recorded is provided to a computer, and the processor of that computer reads the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the embodiment described above, and the program code itself and the storage medium on which it is stored constitute the present invention. Examples of storage media used to supply such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, SSDs (Solid State Drives), optical disks, magneto-optical disks, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, and the like.
[0149] Furthermore, the program code that implements the functions described in this embodiment can be implemented in a wide range of programming or scripting languages, such as assembler, C / C++, Perl, Shell, PHP, Python, and Java (registered trademark).
[0150] Furthermore, the program code for the software that implements the functions of the embodiment may be distributed via a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the computer's processor may read and execute the program code stored in the storage means or storage medium.
[0151] In the above-described embodiment, the control lines and information lines shown are those deemed necessary for explanation and do not necessarily represent all control lines and information lines in the actual product. All components may be interconnected. [Explanation of Symbols]
[0152] 100 Data History Management System 101 Database Management Systems 102 Infrastructure Systems 105 Network 110 Information Acquisition Department 111 Tracking Department 120 Database Management System Information 121 Basic System Information 122 Tracking Policy Information 130 Database Management Department 140 Infrastructure Management Department 141 Volume 200 calculator 201 Processor 202 Main storage 203 Secondary storage device 204 Network Interfaces Bus 205 1200 Tracking List 1600 Drive Group List
Claims
1. A computer system, A computer comprising at least one computer having a processor, a storage device connected to the processor, and a network interface connected to the processor, It connects to at least one database management system for managing data, and at least one infrastructure system for managing data storage areas. The aforementioned computer system, We accept tracking requests that include identification information for the target data. Based on information regarding the data managed by the at least one database management system, the data existence period during which the target data was managed by the at least one database management system is calculated. Based on the configuration history information relating to the configuration of the storage area in the at least one underlying system, a first storage area is identified that was provided to the at least one database management system and in which the target data was stored during the data existence period, and the first storage area is registered in the tracking list. Based on the data existence period and the history of replication operations of the storage area of the at least one base system, a tracking process is performed to track the second storage area where the target data is stored by replication operations, starting from the first storage area. Regarding the storage area in which the target data is stored, information concerning the first storage area and the second storage area is output. In the tracking process, the aforementioned computer system Select one target storage area from the aforementioned tracking list, The existence period of the aforementioned target storage area is calculated, The replication operation execution period is calculated based on the replication operation history related to the target storage area of the at least one base system. A computer system characterized by identifying the second storage area and registering it in the tracking list based on the data existence period, the existence period of the target storage area, and the replication operation execution period.
2. The computer system according to Claim 1, The aforementioned computer system, It holds tracking policy information for managing the policy for identifying the second storage area, Based on the history of replication operations related to the target storage area of the at least one underlying system, the storage area to which the target storage area has been replicated is identified. A computer system characterized by determining whether the identified storage area satisfies the policy.
3. The computer system according to Claim 2, The aforementioned tracking policy information includes a policy set for each type of replication operation. The computer system is characterized in that the policy is defined using at least two of the data existence period, the existence period of the target storage area, and the replication operation execution period.
4. The computer system according to Claim 1, A computer system characterized in that the data existence period of the target data is calculated based on the date and time when the target data was generated by the at least one database management system and the date and time when the target data was completely deleted by the at least one database management system.
5. The computer system according to Claim 1, Based on configuration history information relating to the configuration of the storage area in the at least one base system, a group of drives that provide storage areas to the first storage area and the second storage area is identified. A computer system characterized by outputting information regarding the identified group of drives.
6. The computer system according to claim 5, Based on the configuration history information relating to the configuration of the storage area in at least one of the aforementioned base systems, the base system equipped with the identified drive group is identified, A computer system characterized by outputting information relating to the identified underlying system.
7. The computer system according to Claim 1, A computer system characterized by performing at least one of the following operations on the first storage area and the second storage area: deletion of target data and duplication of target data, in accordance with instructions from the user.
8. The computer system according to Claim 1, It maintains storage policy information for managing the policy of the storage location of the aforementioned data, A computer system characterized by determining whether the first storage area and the second storage area satisfy the policy for the location of the data, and outputting the determination result.
9. A method for tracking the history of data performed by a computer system, The aforementioned computer system, A computer comprising at least one computer having a processor, a storage device connected to the processor, and a network interface connected to the processor, It connects to at least one database management system for managing data, and at least one infrastructure system for managing data storage areas. The method for tracing the origin of the aforementioned data is: The first step involves at least one computer receiving a tracking request that includes identification information for target data, A second step in which the at least one computer calculates the data existence period during which the target data was managed by the at least one database management system, based on information about the data managed by the at least one database management system. A third step in which the at least one computer identifies a first storage area that is provided to the at least one database management system and stores the target data during the data existence period, based on configuration history information relating to the configuration of the storage area in the at least one base system, A fourth step in which at least one computer performs a tracking process that tracks the second storage area where the target data is stored by a replication operation, starting from the first storage area, based on the data existence period and the history of replication operations of the storage area of the at least one underlying system, The fifth step includes the at least one computer outputting information about the first storage area and the second storage area as the storage area in which the target data is stored, The third step includes the step of at least one computer registering the first storage area in a tracking list, The aforementioned tracking process is, The sixth step involves at least one computer selecting one target storage area from the tracking list, The seventh step involves at least one computer calculating the lifetime of the target storage area, The eighth step involves the at least one computer calculating the replication operation execution period based on the history of replication operations related to the target storage area of the at least one base system, A method for tracking the origin of data, characterized in that at least one computer identifies the second storage area and registers it in the tracking list based on the data existence period, the existence period of the target storage area, and the replication operation execution period.
10. A method for tracing the history of data according to Claim 9, The aforementioned computer system maintains tracking policy information for managing policies for identifying the second storage area, Step 9 above is, The steps include: the at least one computer identifying the storage area to which the target storage area has been replicated based on the history of replication operations related to the target storage area of the at least one base system; A method for tracing the origin of data, characterized in that at least one computer determines whether the identified storage area satisfies the policy.
11. A method for tracing the origin of data according to claim 10, The aforementioned tracking policy information includes a policy set for each type of replication operation. A method for tracking the origin of data, characterized in that the policy is defined using at least two of the data existence period, the existence period of the target storage area, and the replication operation execution period.
12. A program to be executed by a computer, The aforementioned computer is The system comprises a processor, a storage device connected to the processor, and a network interface connected to the processor. It connects to at least one database management system for managing data, and at least one infrastructure system for managing data storage areas. The aforementioned program, The first step involves accepting a tracking request that includes identification information for the target data, A second step of calculating the data existence period during which the target data was managed by the at least one database management system, based on information about the data managed by the at least one database management system, A third step of identifying a first storage area that, during the data existence period, is provided to the at least one database management system and stores the target data, based on configuration history information relating to the configuration of the storage area in the at least one infrastructure system; A fourth step involves performing a tracking process that, based on the data existence period and the history of replication operations of the storage area of at least one infrastructure system, tracks the second storage area where the target data is stored by replication operations, starting from the first storage area. A fifth step of outputting information regarding the first storage area and the second storage area as the storage area in which the target data is stored, The computer is made to execute the above, The third step includes registering the first storage area in the tracking list, The aforementioned tracking process is, A procedure for selecting one target storage area from the aforementioned tracking list, A procedure for calculating the existence period of the target storage area, A procedure for calculating the replication operation execution period based on the replication operation history related to the target storage area of at least one of the underlying systems, A program characterized by including a step of identifying the second storage area and registering it in the tracking list based on the data existence period, the existence period of the target storage area, and the replication operation execution period.
Citation Information
Patent Citations
History management system and history management method
JP2004258733A
Method and apparatus to manage object based tier
JP2011170833A
Computer system and metadata management server
JP2011238165A
File history recording system, file history management device, and file history recording method
WO2012164648A1