Computer system, and method executed by computer system

The computer system addresses the inefficiencies in data cache management by using a storage controller to manage a virtual volume and control a cache volume based on operation mode, resulting in reduced data capacity, transfer, and management costs.

JP2025076590APending Publication Date: 2025-05-16HITACHI VANTARA LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023188232
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing data cache techniques struggle to identify and manage data effectively, leading to increased capacity and transfer of cached data, as well as higher management costs and resource consumption.

Method used

A computer system with a storage controller and storage device that provides a virtual volume for data management, constituting a cache volume based on storage area, and controlling it based on operation mode determined by the operation mode determining unit.

Benefits of technology

This approach reduces the capacity of cached data, minimizes data transfer to cache, and lowers the cost of managing cached data, thereby reducing resource and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025076590000001_ABST
    Figure 2025076590000001_ABST
Patent Text Reader

Abstract

To reduce capacity of cache data, reduce a transfer amount, and reduce a management cost.SOLUTION: A computer system 100 comprises a storage control unit 103 and a storage device. The storage control unit 103 receives notification of information about an operation mode 108 determined by an operation mode determination unit 101 on the basis of processing executed by a process execution unit 102. The storage control unit 103 executes access to an actual volume 106, when object data is not stored in a cache volume 105, when receiving an access request 109 through a virtual volume 104 from the process execution unit 102. The storage control unit 103 controls the cache volume 105 on the basis of the information about the operation mode 108.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a control technique for a storage system (for example, a hybrid cloud-linked storage system) constructed between systems interconnected via a network. [Background technology]

[0002] Digital transformation is now widely recognized as a competitive axis for businesses, and there has been a growing movement to utilize the data held by companies to create new value. Data analysis, which is the basis for data utilization, often uses cloud services (e.g., public clouds). However, the data to be analyzed is generally stored in each company's on-premise environment (or private cloud). As a result, there is an increasing need to analyze data to be analyzed that resides in on-premise environments (or private clouds) using cloud services (e.g., public clouds).

[0003] Patent Document 1 is one of the prior arts relating to analyzing data to be analyzed that exists in an on-premise environment (or a private cloud) using a cloud service (for example, a public cloud). Patent Document 1 discloses a prior art that holds a copy of data that exists in the on-premise environment (or a private cloud) in a cloud environment (public cloud) on a page-by-page basis in order to efficiently access data that exists in the on-premise environment (or a private cloud) from the cloud environment (public cloud).

[0004] In addition to the above-mentioned case where data to be analyzed that resides in an on-premises environment (or private cloud) is analyzed using a cloud service (e.g., a public cloud), there is a more general need for systems that are interconnected via a network to be able to efficiently access data that resides in other systems. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] JP 2019-149077 A Summary of the Invention [Problem to be solved by the invention]

[0006] In the technology for holding a copy of data (so-called technology for caching data), including the prior art disclosed in Patent Document 1, in many cases, the selection of data to be held in the cache is performed based on the past access history of the data. Therefore, it is often not easy to identify data that will be needed now or in the future. In addition, it is often not easy to identify when the data will be needed. If the above-mentioned data identification and timing identification are not performed appropriately, data that does not need to be cached because it is unlikely to be reused or the elapsed time until reuse is long may be cached, which may lead to an increase in the amount of cached data and an increase in the amount of data transferred. In addition, if data that is not accessed after a certain point in time remains in the cache, the cost of managing the data may increase. The above may lead to excessive consumption of resources and power.

[0007] In view of the above, one object of the present disclosure may be to realize a reduction in the amount of data to be cached, a reduction in the amount of data transferred for caching, or a reduction in the cost of managing the data to be cached. [Means for solving the problem]

[0008] In order to achieve at least one of the above objects, the present disclosure may have the following features, for example. One aspect of the present disclosure is a computer system including a storage control unit and a storage device. The storage control unit provides a virtual volume to a processing unit that executes an application. The storage control unit manages data input to and output from a real volume via the virtual volume. The storage control unit configures a cache volume based on a storage area of ​​the storage device. The storage control unit receives notification of information regarding the operation mode determined by the operation mode determination unit based on the processing executed by the processing execution unit. When the storage control unit receives an access request from the processing execution unit via a virtual volume, if the target data of the access request is not stored in the cache volume, the storage control unit executes access to the real volume for the target data. The storage control unit controls the cache volume based on the notified information regarding the operation mode. Effect of the Invention

[0009] As the present disclosure has the above features, the present disclosure can realize a reduction in the amount of data to be cached, a reduction in the amount of data transferred for caching, or a reduction in the cost of managing the data to be cached.

[0010] A method or program for realizing the same processing as that realized by the above system can also obtain the same operational effects as the above system. Other features that the present disclosure may have and the effects corresponding to these features will be disclosed in this specification, the claims, or the drawings. [Brief description of the drawings]

[0011] [Figure 1] 1 illustrates a basic functional configuration of an embodiment of the present disclosure. [Diagram 2] 2 shows a detailed functional configuration of an embodiment of the present disclosure. [Diagram 3] 1 shows a first example of processing phases and operation modes. [Figure 4] Indicates staging. [Diagram 5]Indicates destaging. [Figure 6] 2 illustrates a second example of processing phases and modes of operation. [Figure 7] 1 illustrates an overall configuration of a hardware aspect of an embodiment of the present disclosure. [Figure 8] 1 illustrates an example of a computer architecture for implementing a storage node. [Figure 9] 1 shows an example of a computer architecture. [Figure 10] 1 shows a flowchart in a workflow management unit. [Figure 11] 1 shows a first example of workflow information. [Figure 12] 13 shows a second example of workflow information. [Figure 13] 13 shows a flowchart of an operation mode determination unit. [Figure 14] The first example of an attribute table (for a service) is shown below. [Figure 15] A second example of an attribute table (for a volume) is shown. [Figure 16] 13 shows a flowchart in a storage management unit. [Figure 17] 13 shows a flowchart of a service execution unit. [Figure 18] 13 shows a flowchart of access control in a storage control unit. [Figure 19] Here is an example of where the page containing the data is located: [Figure 20] The conversion table is shown below. [Figure 21] 13 shows a flowchart of purge mode control in a storage control unit. [Figure 22] 13 shows a modified example of implementation in a virtual computer environment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the embodiments described below do not limit the disclosure according to the claims, and all of the elements and combinations thereof described in the embodiments are not necessarily essential to the solution of the present disclosure. The present disclosure is not limited to the embodiments, and any application example that matches the idea of ​​the present disclosure is included in the technical scope of the present disclosure. The following description and drawings are examples for explaining the present disclosure, and are appropriately omitted and simplified for clarity of explanation. Unless otherwise limited, each component may be singular or plural. In order to facilitate understanding of the present disclosure, the position, size, shape, range, etc. of each component shown in the drawings may not represent the actual position, size, shape, range, etc. For this reason, the present disclosure is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings. Each of the systems, servers, nodes, devices, or functional units of the present disclosure may be integrated in terms of hardware, or may be divided into multiple parts that work together to perform their roles. Several systems, servers, nodes, devices, or functional units may be integrated in terms of hardware. Each of the systems, servers, nodes, devices, or functional units may be realized by causing a computer to execute software (programs) (as in FIG. 9). Some of the functions of the systems, servers, nodes, devices, or functional units may be realized by hardware (e.g., hardwired logic or FPGA), and the remaining functions may be realized by executing software (programs). All of the functions of the systems, servers, nodes, devices, or functional units may be realized by hardware. Some or all of the steps shown in the flowcharts described in this disclosure may be realized by hardware. One or more systems, servers, nodes, devices, or functional units of the present disclosure may be realized from one or more hardware resources. For this purpose, each of the systems, servers, nodes, devices, or functional units of the present disclosure may be realized virtually. For example, a virtual computer or container technique may be used. The software (program) of the present disclosure may be included in the concept that generally includes software that cooperates with hardware resources to construct a specific information processing system or its operating method according to the intended use. In other words, the program of the present disclosure is not limited to a specific type or form of program. In addition, the program may be initially recorded in a compressed format. The same reference numbers are used in multiple drawings. In the drawings showing flow charts, rectangular boxes indicate processing steps, and hexagonal boxes indicate conditional branching steps. In the drawings showing flow charts, "step" is abbreviated as "S". Also, in the drawings, database is abbreviated as "DB".

[0013] 1. Basic functional configuration of an embodiment of the present disclosure (FIG. 1) Fig. 1 shows a basic functional configuration of an embodiment of the present disclosure. Note that not all of the functional configurations shown in Fig. 1 are essential to the present disclosure. In addition, functional configurations other than those shown in Fig. 1 may be added. In FIG. 1, a computer system 100 (hereinafter, "computer system 100" will be simply referred to as "system 100") is capable of communicating with another system 190 (another system) via a network 180. The network 180 may be, for example, a WAN. The other system 190 has a real volume 106. In FIG. 1, four real volumes 106 are shown, but the number of real volumes 106 may be any number. The system 100 can implement a cache volume 105. Here, the cache volume 105 is used to cache data held by the real volume 106 held by the other system 190. In FIG. 1, two cache volumes 105 are shown, but the number of cache volumes 105 may be any number. The system 100 handles access (access request 109) to data managed in the virtual volume 104. Specifically, when an access entity (processing execution unit 102 in FIG. 1) in the system 100 accesses (access request 109) data present in a real volume 106 of another system 190 or data present in a cache volume 105 of the system 100, the virtual volume 104 may be the access target. Explaining with the example of FIG. 1, for example, when the processing execution unit 102 in the system 100 accesses data present in a real volume 106-1 in the other system 190 or a cache volume 105-1 in the system 100, the processing execution unit 102 may access the virtual volume 104-1. Similarly, for example, when the processing execution unit 102 accesses data present in a real volume 106-2, the processing execution unit 102 may access the virtual volume 104-2. As shown in FIG. 1, for each virtual volume 104, it may be determined whether data existing in the real volume 106 corresponding to the virtual volume 104 is copied (cached) to the cache volume 105. In the example of FIG. 1, data existing in the real volume 106-1 and the real volume 106-4 may be copied (cached) to the cache volume 105-1 and the cache volume 105-4, respectively. On the other hand, data existing in the real volume 106-2 and the real volume 106-3 is not copied (cached) to the cache volume 105 implemented in the system 100. For each real volume 106 (virtual volume 104), whether or not to use caching by the cache volume 105 may be determined based on the operation mode 108. Here, when it is permitted for a certain virtual volume 104 to copy (cache) data existing in the real volume 106 corresponding to the virtual volume 104 to the cache volume 105, the certain virtual volume 104 may be called a "virtual volume with cache". In FIG. 1, for data present in real volume 106-1 or real volume 106-4, and copies of that data (cached data) exist in cache volume 105-1 or cache volume 105-4, system 100 can access that data quickly (compared to accessing real volume 106-1 or real volume 106-4).

[0014] As shown in FIG. 1, a system 100 has an operation mode determination unit 101, a process execution unit 102, and a storage control unit 103 as functional units. The operation mode determination unit 101 determines the operation mode 108 based on how data is used in the process ("process 107 to be executed" in FIG. 1) executed by the process execution unit 102. The operation mode determination unit 101 directly or indirectly transmits (notifies) information about the operation mode 108 to the storage control unit 103. The storage control unit 103 accepts the information about the operation mode 108. In conjunction with the execution of the process 107, the process execution unit 102 accesses data managed in the virtual volume 104 (access request 109). The storage control unit 103 provides a virtual volume 104 to a process execution unit 102 that executes an application. The storage control unit 103 manages data input / output to / from a real volume 106 via the virtual volume 104. The storage control unit 103 configures a cache volume 105 based on a storage area of ​​a storage device (or a recording device) that the system 100 has. The storage control unit 103 controls access (access request 109) (for example, by the process execution unit 102) to data managed in the virtual volume 104. When the storage control unit 103 receives an access request 109 from the process execution unit 102 via the virtual volume 104, if the target data of the access request is not stored in the cache volume 105, the storage control unit 103 executes access to the real volume 106 for the target data. The storage control unit 103 controls the handling of the cache volume 105 based on information about an operation mode 108 corresponding to a process 107 executed by the process execution unit 102 . For example, when the operation mode 108 corresponding to the process 107 is an operation mode that uses a cache (cache mode or on mode (ON mode) described below), when the process execution unit 102 performs a read access or write access to data managed in the virtual volume 104, the storage control unit 103 may perform control to cache the data present in the real volume 106 corresponding to the virtual volume 104 in the cache volume 105 as necessary (control to leave the data read from the real volume 106 in the cache volume 105). On the other hand, when the operation mode 108 corresponding to the process 107 is an operation mode that does not use a cache (remote mode or off mode (OFF mode) described below), when a read access or write access is made from the process execution unit 102 to data managed in the virtual volume 104, the storage control unit 103 may perform control not to cache the data present in the real volume 106 corresponding to the virtual volume 104 in the cache volume 105 (control not to leave the data read from the real volume 106 in the cache volume 105). Furthermore, when an operation mode (a purge mode described later) is specified in which data existing in the cache volume 105 is written back (destaged) to the real volume 106 and then erased from the cache volume 105, the storage control unit may perform destaging and erasing of the data in the cache volume 105 (control to prevent data stored in the cache volume 105 from remaining in the cache volume 105). Details of the operation mode will be described later.

[0015] Since the system 100 in the present disclosure has the functional configuration described above, it can have the effects shown in the above-mentioned [Effects of the Invention].

[0016] 2. Detailed Functional Configuration of the Embodiment of the Present Disclosure (FIG. 2) Fig. 2 shows a detailed functional configuration of an embodiment of the present disclosure. Note that not all of the functional configurations shown in Fig. 2 are essential to the present disclosure. In addition, functional configurations other than those shown in Fig. 2 may be added. Items already explained using Fig. 1 may not be repeatedly explained below. 2, the system 100 that makes an access via the virtual volume 104 may be a public cloud, and another system 190 that has the real volume 106 may be an on-premise environment or a private cloud. In this way, the present disclosure can be useful in allowing the public cloud to perform data analysis on data to be analyzed that is held in the on-premise environment or the private cloud. The system 100 may include an application server 201, a storage node 203, and a management server 205. The operation mode determination unit 101 and the process execution unit 102 may be functional units realized in the application server 201. The cache volume 105 may be included in the storage node 203. The storage control unit 103 may be a functional unit realized in the storage node 203. The management server 205 may include a storage management unit 251 as a functional unit. The other system 190 may include a storage node 290 in the other system. The real volume 106 may be included in the storage node 290 in the other system. The storage node 290 in the other system may include a storage control unit 293 in the other system as a functional unit. As described above, the system, server node, volume, and functional units are appropriately separated and managed, and thus the embodiment of the present disclosure can realize the functions of the present disclosure while avoiding unexpected mutual interference between the system, server node, volume, and functional units. At least some of the virtual volumes 104 handled by the system 100 may be included in a virtual lake in which a lake for handling unstructured data to be analyzed is virtualized, or in a virtual mart in which a mart for handling structured data to be analyzed is virtualized. FIG. 2 shows an example in which a virtual volume 104-3 and a virtual volume 104-4 are included in a virtual mart 204. In this way, by including the virtual volume 104 in a larger management unit (virtual lake or virtual mart), it is also possible to reduce the management cost of the attributes of the virtual volume 104. If the management cost of the attributes of the virtual volume 104 can be reduced, the consumption of resources and power in the system 100 etc. can also be reduced.

[0017] The process execution unit 102 may have, as functional units, a workflow management unit 212 and a service execution unit 213. For example, the workflow management unit 212 may be realized as a functional unit by the application server 201 executing a workflow management program. Also, the service execution unit 213 may be realized as a functional unit by the application server 201 executing a service execution program. The workflow management unit 212 manages the workflow based on the workflow information 207 in which a workflow indicating the process 107 is described, and also manages the service execution unit 213. The management performed by the workflow management unit 212 may include management related to the execution of various services (various service programs) by the service execution unit 213, management of parameters used in the various services, and management of data to be analyzed that is handled in the various services. The workflow information 207 may include, for each phase of a process included in the process 107, information that identifies a service corresponding to that phase, or information that identifies data input / output in that phase. The workflow information 207 may also include, for each phase of a process, information on an explicit operation mode that explicitly specifies the operation mode 108. The workflow information 207 may also be expressed as a graph in accordance with the form of the workflow. The form of the workflow information 207 will be described later with reference to Figs. 11 and 12. The service execution unit 213 executes a service included in the process 107. For example, the service execution unit 213 may execute a service corresponding to each phase of the process included in the process 107. In conjunction with the execution of the service, the service execution unit 213 may access (access request 109) data managed in the virtual volume 104 (data existing in the real volume 106 or the cache volume 105). When implementing the service execution unit 213, a service program for each service type may be executed under control based on the above-mentioned service execution program. Alternatively, a service program for each service type may be executed without control based on the service execution program (under direct control based on the workflow management program). If the processing execution unit 102 is as described above, the processing execution unit 102 can execute the processing 107 based on the workflow described in the workflow information 207. Furthermore, the processing execution unit 102 can execute a service for each phase of the processing included in the processing 107 based on the description in the workflow information 207, and can access data input / output in the phase (access request 109).

[0018] The operation mode determination unit 101 may be realized as a functional unit by, for example, the application server 201 executing an operation mode determination program. The operation mode determination unit 101 may determine the context (how data is used in the process indicated by the data flow) of the workflow indicating the process 107 based on attributes related to the process 107 executed by the process execution unit 102 (of which, the service execution unit 213), obtain a context determination result 211, and then determine the operation mode 108 corresponding to the context (context determination result 211). Then, the operation mode determination unit 101 may notify the storage management unit 251 in the management server 205 of the operation mode information 208 indicating the operation mode 108. As described above, (instead of handling the cache volume 105 based only on the past access history), the operation mode determination unit 101 determines how data is used in the process 107 as the context of the workflow indicating the process 107. Therefore, when the process 107 is executed, the operation mode determination unit 101 can determine the operation mode 108 related to the handling of the cache volume 105 so as to match the data that needs to be accessed now or in the future and the timing at which the data is needed. To determine the operation mode 108 , the operation mode determination unit 101 may use the workflow information 207 in which a workflow indicating the process 107 is described, and information included in the attribute table 214 . As described above, the workflow information 207 may include, for each phase of a process included in the process 107, information identifying a service corresponding to that phase, information identifying data input / output in that phase, or information on an explicit operation mode that explicitly designates the operation mode 108. The operation mode determination unit 101 may determine the context for each phase of a process based on the above information included in the workflow information 207. On the other hand, the attribute table 214 may have information on the type of operation mode 108 related to the handling of the cache volume 105 according to the type of attribute related to the process 107. As shown in FIG. 14 described later, the attribute table 214 may have information on the operation mode 108 according to the type of service executed by the process execution unit 102 (of which, the service execution unit 213). Or, as shown in FIG. 15 described later, the attribute table 214 may have information on the operation mode 108 according to the type of volume, etc. accessed by the process execution unit 102 (of which, the service execution unit 213). The attribute table 214 provides the operation mode determination unit 101 with a guideline (policy) for determining the operation mode 108. The operation mode determination unit 101 may compare, for each phase of a process included in the process 107, information in the workflow information 207 that identifies a service corresponding to the phase or information that identifies data (volume, etc.) input / output in the phase with information on the operation mode 108 corresponding to the type of service or the type of volume, etc., in the attribute table 214. The operation mode determination unit 101 may determine the operation mode 108 based on the comparison result. Alternatively, when the workflow information 207 includes information on an explicit operation mode that explicitly specifies the operation mode 108 for a certain phase of a process, the operation mode determination unit 101 may preferentially use the information on the explicit operation mode to determine the operation mode 108 in the certain phase. In this way, when the operation mode determination unit 101 determines the operation mode 108, any of the attributes of the service, the attributes of the data (volume, etc.) input / output, or the explicit specification of the operation mode may be utilized for each phase of a process included in the process 107, so that the determination of the operation mode 108 for each phase can be flexibly performed. As described above, the operation mode determination unit 101 determines the control of the virtual volume 104 (and the cache volume 105) in association with the workflow. Therefore, the system 100 can appropriately allocate data, while ensuring the performance of accessing data and reducing the costs associated with maintaining and managing the data. Reducing the costs associated with maintaining and managing data can also reduce the consumption of resources and power in the system 100, etc.

[0019] The storage management unit 251 may be realized as a functional unit by, for example, the management server 205 executing a storage management program. The storage management unit 251 is notified of the operation mode information 208 by the operation mode determination unit 101. Upon receiving the notification, the storage management unit 251 transmits (notifies) an operation mode instruction 252, which is an instruction to apply the operation mode 108 indicated by the operation mode information 208, to the storage control unit 103 that controls access to the cache volume 105 to which the operation mode 108 should be applied (and the virtual volume 104 that uses the cache volume 105). In this way, in the embodiment of the present disclosure, a method may be adopted in which the storage management unit 251 in the management server 205 receives information (operation mode information 208) of the operation mode 108 determined by the operation mode determination unit 101, and then the storage management unit 251 sends (transmits, notifies) the operation mode instruction 252 to the storage control unit 103. By adopting this method, even if the system 100 implements multiple application servers 201 or multiple storage nodes 203, the management server 205 (storage management unit 251 therein) can centrally manage the storage within the system 100. A modified example is also possible in which the operation mode determination unit 101 directly notifies the storage control unit 103 of the information on the operation mode 108 (operation mode information 208).

[0020] The storage control unit 103 may be realized as a functional unit, for example, by the storage node 203 executing a storage control program. When there is an access (access request 109) from the process execution unit 102 (the service execution unit 213 therein) to data managed in the virtual volume 104, the storage control unit 103 may first check whether the data to be accessed exists in the cache volume 105. To make this check, the storage control unit 103 may refer to the information in the conversion table 231. The conversion table 231 may have information regarding the location of each piece of data managed in the virtual volume 104 in the real volume 106, as well as information regarding the presence or absence of the data in the cache volume 105, and (if present) information regarding the location of the data (a copy of the data) in the cache volume 105. By using the conversion table 231, the storage control unit 103 can determine whether it is appropriate to access the cache volume 105 or the real volume 106 for each piece of data managed in the virtual volume 104. The aspects of the conversion table 231 will be described later with reference to Figs. 19 and 20.

[0021] When there is an access (access request 109) from the processing execution unit 102 (the service execution unit 213) to data managed in the virtual volume 104, if the data does not exist in the cache volume 105, how that data is handled depends on the operation mode 108 indicated by the operation mode instruction 252. When the data targeted by the access request 109 does not exist in the cache volume 105, there are roughly two methods for handling the case. In the first method, when the data targeted by the access request 109 does not exist in the cache volume 105, the storage control unit 103 performs a read access or write access to the data in the real volume 106, but does not stage the data (from the real volume 106) in the cache volume 105 (transferring and storing a copy of the data). The "remote mode" and "off mode" described below correspond to this first method. For example, if this first method is applied to data that is unlikely to be reused or data that will take a long time to be reused, the amount of data that is staged (transferred and stored) from the real volume 106 to the cache volume 105 can be reduced while suppressing the disadvantage caused by the accessed data not being cached. In the second method, when the data targeted by the access request 109 does not exist in the cache volume 105, the storage control unit 103 stages the data (from the real volume 106) to the cache volume 105 (transferring and storing a copy of the data) and then performs a read access or write access to the data that now exists in the cache volume 105. The "cache mode" and "on mode" described below correspond to this second method. For example, if this second method is applied to data that has a high reusability and will be reused soon, it becomes possible to obtain the advantage of caching data that corresponds to the amount of data staged (transferred and stored) from the real volume 106 to the cache volume 105.

[0022] There may also be a case where the operation mode instruction 252 specifies an operation mode (a purge mode, described later) for destaging the data in the cache volume 105 to the real volume 106 (writing back the data) and then erasing the data in the cache volume 105. In this case, the storage control unit 103 may perform destaging and erasing the data in the cache volume 105 regardless of the access (access request 109) from the processing execution unit 102 (the service execution unit 213 therein) to the data managed in the virtual volume 104. If the above-mentioned purge mode control is performed on data that is not accessed in the system 100 after a certain point in time (after a certain phase), the possibility that data that no longer needs to be stored in the cache volume 105 will remain in the cache volume 105 is reduced, and it is expected that the data management cost will be reduced. If the data management cost can be reduced, the consumption of resources and power related to the handling of the cache volume 105 can also be reduced.

[0023] The other intra-system storage control unit 293 may be realized as a functional unit, for example, by executing another intra-system storage control program by the other intra-system storage node 290. The other intra-system storage control unit 293 controls access to the real volume 106. When a storage system that traverses systems interconnected by a network is constructed by a storage-related part in the system 100 and a storage-related part in another system 190, the storage control unit 293 in the other system and the storage control unit 103 may cooperate in some way. For example, an operation mode instruction 252 issued by a storage management unit 251 included in the system 100 may be transmitted not only to the storage control unit 103 but also to the storage control unit 293 in the other system. In this case, the storage control unit 103 and the storage control unit 293 in the other system may cooperate to operate in order to make the control of the virtual volume 104 and the control of the real volume 106 and the cache volume 105 corresponding to the virtual volume 104 conform to the operation mode 108 indicated in the operation mode instruction 252. Furthermore, in the case of a modified example in which it is sufficient for the storage control unit 103 to execute the control for realizing the operation mode 108 indicated in the operation mode instruction 252, and the other in-system storage control unit 293 simply responds to read access or write access to data existing in the real volume 106, the operation mode instruction 252 does not need to be sent to the other in-system storage control unit 293.

[0024] 1. First Example of Processing Phases and Operation Modes (Figures 3, 4, and 5) 3 shows a first example of a workflow showing a process 107, a process phase included in the process 107, and an operation mode corresponding to the phase, which may be handled in an embodiment of the present disclosure. As shown in Fig. 3, an operation mode 108 that is a "remote mode", a "cache mode", or a "purge mode" may be set for a process phase (process) or a virtual volume 104 associated with a process phase (process). In the example of Fig. 3, workflow 307 (workflow of data analysis) showing process 107 shows, in the order of execution, a search reference process 372, a pre-processing process 374, an analysis process 376, and a post-processing process 378. Each of the processes shown in workflow 307 corresponds to each of the phases of the process. Workflow 307 may include points 371, 373, 375, and 377 between these processes. Each of the points may indicate the control content of the transition between processes (phases) (start or end of the process (phase)) and the control content of the repetition of the process (phase). The search reference process 372 is a process for searching and referencing data to be analyzed. The search reference process 372 may be a process for searching a data lake (for managing unstructured data as data to be analyzed) to investigate what kind of data exists as data to be analyzed (what kind of data is likely to be used in the analysis process 376) prior to the analysis process 376, for example. Alternatively, the search reference process 372 may perform processing for collecting or adding meta information for identifying a data group to be analyzed to be used in the pre-processing process 374 or the analysis process 376. (Note that the search reference process 372 may be performed by another system 190 (on-premise environment or private cloud) instead of the system 100 (public cloud). Also, if the data to be analyzed to be used in the analysis process 376, etc., and the contents of the processing to be performed on the data are known in advance, the search reference process 372 may not be executed.) The pre-processing process 374 is a process that performs processing in preparation for the analysis process 376. The pre-processing process 374 may be a process that performs pre-processing such as cleansing the data to be analyzed (processing to detect corrupted or inaccurate data and correct or delete the corrupted or inaccurate data, etc.) or shaping the data to be analyzed. The analysis process 376 is a process for performing data analysis and data verification. The analysis process 376 may be, for example, a process for performing some kind of aggregation processing. Alternatively, the analysis process 376 may be, for example, a process for training a learning model by performing machine learning using the data to be analyzed (for example, training model parameters). The post-processing process 378 is a process for performing post-processing based on the processing result by the preceding analysis process 376. The post-processing process 378 may be, for example, a process for transmitting information (e.g., trained model parameters) on the learning model trained in the analysis process 376 from the system 100 (public cloud) to another system 190 (on-premise environment or private cloud) and storing the information in the other system 190 (on-premise environment or private cloud). Alternatively, the post-processing process 378 may be a process in which the learning model is utilized. (Note that, in the case where the other system 190 (on-premise environment or private cloud) utilizes the trained learning model, the process in which the other system 190 utilizes the learning model may be separated from the workflow 307.)

[0025] The workflow management unit 212 refers to the workflow information 207 (workflow 307) in which each of the processes (phases) described above is described as each of the workflow entities. The workflow information 207 (workflow 307) may be directly referred to by the operation mode determination unit 101. In this case, it can be said that the operation mode determination unit 101 analyzes the workflow 307 by itself. Alternatively, the operation mode determination unit 101 may be provided with information (workflow entities) included in the workflow information 207 (workflow 307) from the workflow management unit 212. In this case, for each workflow entity individually provided by the workflow management unit 212, the operation mode determination unit 101 determines the operation mode 108 for the phase of the process corresponding to the workflow entity. The roles played by the workflow management unit 212, service execution unit 213, operation mode determination unit 101, attribute table 214, storage management unit 251, storage control unit 103, and conversion table 231 have already been explained using FIGS.

[0026] 3, in the processing phase in which the search reference process 372 is executed, the virtual volume 104a for managing the data accessed based on the search reference process 372 may be set to be handled in "remote mode". For example, the "remote mode" may be set at point 371. When there is a read access or write access to the data managed in the virtual volume 104a for which the "remote mode" is set, the above-mentioned first method is applied. The data handled by the search reference process 372 is not newly imported into the cache volume 105. The data accessed between the search reference process 372 and the subsequent processes does not necessarily match. Also, a long time may elapse between the completion of execution of the search reference process 372 and the start of execution of the subsequent processes. Due to these circumstances, the effect of caching between the search reference process 372 and the subsequent processes tends to be low. For this reason, it is often appropriate to apply the "remote mode" to the search reference process 372.

[0027] In the example of FIG. 3, in the processing phase in which the pre-processing process 374 and the analysis process 376 are executed, the virtual volume 104b for managing the data accessed based on the pre-processing process 374 and the analysis process 376 may be set to be handled in the "cache mode". For example, the "cache mode" may be set at point 373, and the "cache mode" may be continued at point 375. When there is a read access or write access to the data managed in the virtual volume 104b in which the "cache mode" is set, the above-mentioned second method is applied. The data handled by the pre-processing process 374 and the analysis process 376 may be taken into the cache volume 105b. In other words, when the data handled by the pre-processing process 374 and the analysis process 376 does not exist in the cache volume 105b (when a miss is generated), the data may be staged (transferred and stored) from the real volume 106b to the cache volume 105b. Much of the data handled in the pre-processing process 374 tends to be handled in the analysis process 376 as well. In addition, the processing results of the analysis process 376 tend to be subject to post-processing in the post-processing process 378. Due to these circumstances, it is expected that the cache effect will be high in the pre-processing process 374 and the analysis process 376. Therefore, it is often appropriate to apply the "cache mode" to the pre-processing process 374 and the analysis process 376.

[0028] FIG. 4 explains the staging performed in the "cache mode." The diagram on the left side of Fig. 4 shows a case where, when the service execution unit 213 makes an access request 409 to data managed in the virtual volume 104, the data does not exist in the cache volume 105 (miss hit). (Note that, if "remote mode" is set for the virtual volume 104, read access or write access to the data that is the subject of the access request 409 is made directly to the real volume 106.) 4 shows how page 401 including the data that is the subject of access request 409 is staged from real volume 106 to cache volume 105 in response to the fact that ("cache mode" is set for virtual volume 104) and the data that is the subject of access request 409 is not present in cache volume 105. This staging 403 causes a copy of page 401 (page copy 402) to be held in cache volume 105. The diagram on the right side of Fig. 4 shows how, after staging 403, a read access or write access (404 in Fig. 4) to the data that is the target of the access request 409 is performed in the cache volume 105. If the data that is the target of the read access or write access exists in the cache volume 105, the data can be accessed quickly (compared to the case where the real volume 106 is directly accessed).

[0029] In the example of FIG. 3, in the processing phase in which the post-processing process 378 is executed, the virtual volume 104c (which may be the same as the virtual volume 104b) corresponding to the post-processing process 378 may be set to be handled by the "purge mode". For example, the "purge mode" may be set at point 377. When the "purge mode" is set, a series of processes related to the purge mode are performed at the beginning of the post-processing process 378 to which the "purge mode" is applied. After the data present in the cache volume 105c is destaged to the real volume 106c (the data is written back), the data that was present in the cache volume 105c is erased. Furthermore, the virtual volume 104c or the cache volume 105c may be blocked. After the processing phase in which the post-processing process 378 is executed, the analysis results (e.g., the trained learning model (trained model parameters)) by the analysis process 376 may be used exclusively by other systems 190 (on-premise environments or private clouds) and may not be necessary for the system 100. In such a case, if the above-mentioned purge mode control is performed, the possibility that data that no longer needs to be stored in the cache volume 105c will continue to remain in the cache volume 105 is reduced (or eliminated), and it is expected that the data management cost (data distribution cost) will be reduced. If the data management cost (data distribution cost) can be reduced, the consumption of resources and power in the system 100, etc. can also be reduced.

[0030] FIG. 5 illustrates the destaging performed in the "purge mode." 5 shows a case before destaging is executed and when the service execution unit 213 makes an access request to data managed in the virtual volume 104, the data exists (hits) in the cache volume 105. Before destaging is executed, if the page copy 402 including the data that is the subject of the access request is hit in the cache volume 105, a read access or write access (404 in FIG. 5) to the data is made to the cache volume 105. The central diagram in FIG. 5 shows a state in which the page copy 402 existing in the cache volume 105 is destaged from the cache volume 105 to the real volume 106 as a result of the "purge mode" being set for the virtual volume 104. Among the data existing in the cache volume 105, at least data (dirty data) that has not yet been reflected in the real volume 106 (among the data that needs to be saved) is transferred from the cache volume 105 to the real volume 106, and the dirty data is held in the real volume 106. Note that the operation of transferring data and holding it in the real volume 106 in this destaging 503 may be performed only for the dirty data, may be performed for the page copy 402 including the dirty data, or may be performed for all valid page copies 402 existing in the cache volume 105. Note that in order to reduce the risk of losing dirty data before the destaging 503 is performed, the cache volume 105 may be made redundant (for example, RAID). 5, after destaging 503 is completed, all of the page copies 402 that were present in the cache volume 105 may be erased. Also, in the system 100, the virtual volume 104 or cache volume 105 that was the subject of destaging 503 may be treated as blocked. The diagram on the right side of Fig. 5 shows how the service execution unit 213 accesses the data managed in the virtual volume 104 that has been destaged again after destaging 503 and after all of the page copies 402 that were in the cache volume 105 have been erased (in many cases in connection with a separate workflow). At this point, the data managed in the virtual volume 104 exists only in the real volume 106, so (if "remote mode" is applied to the virtual volume 104) read access or write access (504 in Fig. 5) to the data that is the subject of the access request is made directly to the real volume 106.

[0031] As shown in FIG. 3, by setting the operation mode 108 for each processing phase, the virtual volume 104 is accessed in "cache mode" only when the pre-processing process 374 and the analysis process 376 are executed, and after the analysis process 376 is completed, the data handled in the processing indicated by the workflow 307 does not remain in the system 100 (public cloud). Only when the pre-processing process 374 and the analysis process 376 are executed, (a copy of) data in another system 190 (on-premise environment or private cloud) is placed in the system 100 (public cloud), so that it is expected that the data management cost (data distribution cost) will be reduced, the capacity (size) of data managed in the cache volume 105 used in the system 100 (public cloud) will be reduced, and the amount of data transferred between the system 100 and the other system 190 will be reduced. If the above various reductions can be realized, the consumption of resources and power in the system 100 etc. can also be reduced.

[0032] 2. Second Example of Processing Phases and Operation Modes (Figure 6) 6 shows a second example of a workflow showing process 107, process phases included in process 107, and operation modes corresponding to the phases, which may be handled in an embodiment of the present disclosure. As shown in Fig. 6, a process phase (service) or a virtual volume 104 associated with a process phase (service) may be set to operation mode 108, which is "OFF mode" for read access, "ON mode" for read access, "OFF mode" for write access, "ON mode" or "purge mode" for write access. In the example of Fig. 6, a workflow 607 (workflow for data analysis) showing process 107 shows, in the order of execution, a catalog service 672, an ETL service 674, an analysis service 676, and a model output service 678. Each service corresponds to each phase of the process. The workflow 607 may include points 671, 673, 675, and 677 between these services. Each point may indicate the control content of the transition between services (between phases) (start or end of the service (phase)), and the control content of the repetition of the service (phase). The catalog service 672 may be a type of service that searches for and references data to be analyzed. The catalog service 672 may, for example, search for meta-information stored in a catalog volume. Here, the meta-information is used to search for data to be analyzed. The ETL service 674 may be a type of service that performs preprocessing (e.g., data shaping) on ​​data to be analyzed. The ETL service 674 performs at least one of extracting, transforming, and loading on the data to be analyzed. The analysis service 676 may be a type of service that performs data analysis using the data to be analyzed that has been shaped by the ETL service 674. The analysis service 676 may be, for example, a process that trains a learning model (e.g., trains model parameters) by performing machine learning using the data to be analyzed. The model output service 678 may be a service that outputs the analysis results obtained by the analysis service 676 from the system 100 (public cloud) to another system 190 (on-premise environment or private cloud). The model output service 678 may be, for example, a process that transmits information about the learning model trained by the analysis service 676 (e.g., trained model parameters) from the system 100 (public cloud) to the other system 190 (on-premise environment or private cloud) and stores the information in the other system 190 (on-premise environment or private cloud).

[0033] The workflow management unit 212 refers to the workflow information 207 (workflow 607) in which each of the processes (phases) described above is described as each of the workflow entities. The workflow information 207 (workflow 607) may be directly referred to by the operation mode determination unit 101. In this case, it can be said that the operation mode determination unit 101 analyzes the workflow 607 by itself. Alternatively, the operation mode determination unit 101 may be provided with information (workflow entities) included in the workflow information 207 (workflow 607) from the workflow management unit 212. In this case, for each workflow entity individually provided by the workflow management unit 212, the operation mode determination unit 101 determines the operation mode 108 for the phase of the process corresponding to the workflow entity. The roles played by the workflow management unit 212, service execution unit 213, operation mode determination unit 101, attribute table 214, storage management unit 251, storage control unit 103, and conversion table 231 have already been explained using FIGS.

[0034] In the example of FIG. 6, in the processing phase in which the catalog service 672 is executed, the virtual volume 104d for managing data accessed based on the catalog service 672 may be set to be handled in "OFF mode" for both read access and write access. For example, at point 671, "OFF mode" may be set for both read access and write access. When there is a read access or write access to data managed in the virtual volume 104d in which "OFF mode" is set for both read access and write access, the above-mentioned first method is applied. The data handled by the catalog service 672 is not newly imported into the cache volume 105. The technical reasons and background for the catalog service 672 being in “OFF mode” for both read and write access may be similar to the technical reasons and background for the search reference process 372 in FIG. 3 being in “remote mode.”

[0035] In the example of FIG. 6, in the phase of the process in which the ETL service 674 is executed, the virtual volume 104e for managing data that is read-accessed based on the ETL service 674 may be set to be handled in "off mode (OFF mode)", while the virtual volume 104f for managing data that is write-accessed based on the ETL service 674 may be set to be handled in "on mode (ON mode)". Alternatively, when read access and write access based on the ETL service 674 are performed to the same virtual volume 104, it may be set to be handled in "off mode (OFF mode)" when the virtual volume 104 is read-accessed based on the ETL service 674, and it may be set to be handled in "on mode (ON mode)" when the virtual volume 104 is write-accessed based on the ETL service 674. For example, at the point 673, "off mode (OFF mode)" for read access and "on mode (ON mode)" for write access may be set. During the read access that results in "off mode (OFF mode)", the first method described above is applied. During a write access in the "ON mode", the above-mentioned second method is applied. A page including data written based on the ETL service 674 can be newly imported into the cache volume 105f. In the ETL service 674, data to be analyzed (before reformatting) is read, data reformatting is performed, and reformatted data to be analyzed is obtained. The reformatted data to be analyzed is used in the analysis service 676, while the unformulated data to be analyzed tends not to be used in the analysis service 676. Due to these circumstances, it is expected that the cache effect will be high for data written based on the ETL service 674, while the cache effect will be low for data read based on the ETL service 674. For this reason, it is often appropriate to apply "on mode" to write access and "off mode" to read access.

[0036] In the example of Fig. 6, in the processing phase in which the analysis service 676 is executed, the virtual volume 104g for managing data accessed based on the analysis service 676 may be set to be handled in "ON mode" for both read access and write access. For example, at point 675, "ON mode" may be set for both read access and write access. When there is a read access or write access to data managed in the virtual volume 104g in which "ON mode" is set for both read access and write access, the above-mentioned second method is applied. The technical reasons and background for the analysis service 676 being in “ON mode” for both read access and write access may be similar to the technical reasons and background for the analysis process 376 in FIG. 3 being in “cache mode.”

[0037] In the example of FIG. 6, in the processing phase in which the model output service 678 is executed, the virtual volume 104h (which may be the same as the virtual volume 104g) corresponding to the model output service 678 may be set to be handled by the "purge mode". For example, the "purge mode" may be set at point 677. When the "purge mode" is set, a series of processes related to the purge mode described above is performed at the beginning of the model output service 678 to which the "purge mode" is applied. That is, the data present in the cache volume 105h is destaged to the real volume 106h (the data is written back), and then the data present in the cache volume 105h is erased. Furthermore, the virtual volume 104h or the cache volume 105h may be blocked. The technical reasons and background for the model output service 678 to be in the “purge mode” may be similar to the technical reasons and background for the post-processing process 378 in FIG.

[0038] 3. Overall configuration of the embodiment of the present disclosure (FIG. 7) 7 shows an overall configuration of the hardware aspect of an embodiment of the present disclosure. (Note that, as will be described later as Modification (3), some or all of the servers, nodes, switches, etc. shown in FIG. 7 may be implemented virtually.) In Fig. 7, the system 100 (public cloud) and another system 190 (on-premise environment or private cloud) are interconnected via a network 180 (e.g., a WAN). In the overall configuration shown in Fig. 7, it may be said that a hybrid cloud environment is constructed. For example, a compute node inter-network switch 702 (NSC) in the system 100 (public cloud) and a compute node inter-network switch 702 (NSC) in the other system 190 (on-premise environment or private cloud) may be interconnected via the network 180 (e.g., a WAN). Either system 100 (public cloud) or other system 190 (on-premise environment or private cloud) may have one or more compute nodes 701, one or more compute inter-node network switches 702 (NSC), one or more storage nodes 203 or 290 (SN), and one or more storage inter-node network switches 703 (NSS). A compute node 701 (CN) in the system 100 (public cloud) may play the role of the application server 201 (roles of the operation mode determination unit 101, the workflow management unit 212, and the service execution unit 213). As the compute node 701 (CN) as the application server 201 executes various programs, a request for access to data held by the storage node 203 or 290 (SN) may be made. A storage cluster may be formed by a plurality of storage nodes 203 (SN) in the system 100 (public cloud). Similarly, a storage cluster may be formed by a plurality of storage nodes 290 (SN) in another system 190 (on-premise environment or private cloud). In a hybrid cloud environment, these storage clusters may be connected by a network 180 (e.g., a WAN). The compute node-to-node network switch 702 (NSC) in the system 100 (public cloud) mediates communications between the compute nodes 701 (CN) and between the compute node 701 (CN) and the storage node 203 (SN). The compute node-to-node network switch 702 (NSC) in the system 190 (on-premise environment or private cloud) mediates communications between the compute nodes 701 (CN) and between the compute node 701 (CN) and the storage node 290 (SN). An inter-storage node network switch 703 (NSS) in the system 100 (public cloud) mediates communication between storage nodes 203 (SN). An inter-storage node network switch 703 (NSS) in another system 190 (on-premise environment or private cloud) mediates communication between storage nodes 290 (SN). For management of the system 100 (and further, other systems 190), the system 100 may have a management server 205 (control S.) or a system environment external server 704 (out S.). The management server 205 (control S.) or the system environment external server 704 (out S.) may be capable of communicating with some or all of the compute nodes 701, the inter-compute node network switch 702 (NSC), the storage nodes 203 (SN), and the inter-storage node network switch 703 (NSS) included in the system 100 via a management network switch 706 (NScontrol) or a management network. When the compute node 701 or the storage node 203 (SN) communicates with the management server 205 (control S.) or the system environment external server 704 (out S.) via the management network switch 706 (NScontrol), the communication may be performed via a management port 825 (MNP) that the compute node 701 or the storage node 203 (SN) has (shown in FIG. 8 described below). In order to ensure availability of the system 100 (public cloud) or another system 190 (on-premise environment or private cloud), some or all of the servers, nodes, switches, and paths may be made redundant. However, in FIG. 7, for ease of viewing, some of the redundant configurations (e.g., part of the management network) are omitted. In addition, a terminal of a user of the cloud or the like may be capable of communicating with the system 100 or another system 190 via a network (WAN).

[0039] 4. Computer Architecture In the embodiment of the present disclosure, the computer architecture for realizing each of the application server 201 (compute node 701), the storage node 203 or 290, and the management server 205 may be any as long as it satisfies the requirements of the functional units and storage (volume) to be implemented. In the following, first, an example of a computer architecture for a storage node is shown, and then computer architectures for various servers and nodes are shown.

[0040] 4.1. Architecture for Storage Nodes (Figure 8) FIG. 8 illustrates an example of a computer architecture for implementing storage node 203 (or 293). 8, the storage node 203 has a storage controller 833 and a drive 803 (DRIVE). The storage controller 833 corresponds to the storage control unit 103. The drive 803 (DRIVE) may hold data managed in a cache volume 105. (In a storage node 290 in another system 190, the drive 803 (DRIVE) may hold data managed in a real volume 106.) The storage controller 833 and the drive 803 (DRIVE) are capable of communicating with each other. The storage controller 833 may have one or more processors 801 (P), one or more memories 802 (M), one or more front-end ports 821 (FEP), one or more front-end port switches 822 (FEP SW), one or more back-end port switches 823 (BEP SW), one or more back-end ports 824 (BEP), and one or more management ports 825 (MNP). An instruction (e.g., an access request) from the application server 201 (compute node 701) is transmitted to the processor 801 (P) via a front-end port 821 (FEP) and a front-end port switch 822 (FEP SW). Then, resources such as the processor 801 (P) and memory 802 (M) are used to execute a storage control program, etc., and the drive 803 (DRIVE) is accessed via a back-end port switch 823 (BEP SW) and a back-end port 824 (BEP). The storage node 203 (SN) provides functions such as data input / output processing, data encryption, data compression, snapshot creation, and data virtualization according to instructions from the application server 201 (compute node 701). Intercommunication between the management server 205 and the storage nodes 203 may occur via a management port 825 (MNP).

[0041] FIG. 8 shows a computer architecture for realizing the storage node 203, but it is also possible to realize various servers and nodes other than the storage node 203 (e.g., application server 201 (compute node 701), management server 205) and various functional units using a computer architecture similar to that of FIG.

[0042] 4.2. Other examples of architectures for various servers and nodes (Figure 9) FIG. 9 illustrates another example of a computer architecture for implementing various server node functions in accordance with an embodiment of the present disclosure. In order to realize the system 100 (which may be another system 190; the same applies below in the description of FIG. 9 ) and a server node functional unit included in the system 100, an information processing device 901 (e.g., a CPU or processor; which may be one or more microprocessors), a storage device 902 (e.g., a memory), a non-volatile recording medium or recording device 903 (e.g., a non-volatile memory (e.g., a flash memory), a non-volatile disk device), an external recording medium drive 904 (e.g., a disk drive), an input device 906 (e.g., a mouse, a keyboard, an imaging device, a sensor, a touch panel, a pointing device), a display or output device 907 (e.g., a display, a printer, a speaker), a communication device 908 (e.g., a communication device for wired communication, a communication device for wireless communication; which may be a network interface device (NIC) that controls communication with other systems, servers, nodes, or devices according to a predetermined protocol), an external input / output port 909, and a part or all of a reading device 910 may be interconnected by an interconnection unit 911 (e.g., a bus, a crossbar switch). The non-volatile recording medium or recording device 903 may record one or more of a program 920a (e.g., a program for realizing the functional configuration according to the present disclosure, e.g., an operation mode determination program, a workflow management program, a service execution program, various service programs, a storage management program, or a storage control program), various databases 921, and various information 922. Instead of recording information such as 920a, 921, or 922 in the non-volatile recording medium or recording device 903, the system 100 or a server node functional unit included in the system may acquire (access) part or all of the information such as 920a, 921, or 922 from outside the system shown in FIG. The external recording medium drive 904 can connect to an external recording medium 905 (for example, a portable recording disk (such as a DVD), an IC card, an SD card, a non-volatile memory (for example, a flash memory), or a portable hard disk). In addition, a part or all of the information as shown in 920a, 921, or 922 above may be transferred and stored from this external recording medium 905 to a non-volatile recording medium or recording device 903 or a storage device 902. The external recording medium 905 may be used to record programs and data handled in the system 100 or a server node functional unit included in the system. The external recording medium drive 904 and the external recording medium 905 may be connected to the system 100 or a server node functional unit included in the system 100 via a network. Some or all of the information such as that shown in 920a, 921 or 922 above may be provided via the input device 906, the communication device 908 or the external input / output port 909 and recorded or stored in a non-volatile recording medium or recording device 903 or the storage device 902.

[0043] In order for the architecture of FIG. 9 to function as the system 100, the server node functional unit included in the system 100, or a part of each unit (to execute one or a series of steps), a part or all of the above 920a, 921, or 922 (for example, program 920a) may be loaded into the storage device 902 (for example, from a non-volatile recording medium or recording device 903). The loaded program is indicated by 920b in FIG. 9. Then, the information processing device 901 may execute the program 920b (using various information present in the non-volatile recording medium or recording device 903, etc., as necessary). By executing the program 920b, the function of the system 100, the server node functional unit included in the system 100, or a part of each unit is realized (one or a series of steps are executed). At this time, various buffers 923 temporarily formed in the storage device 902 may also be used.

[0044] 5. Controls performed in the embodiment of the present disclosure In the following, the control and the like performed in the embodiment of the present disclosure will be described in approximately chronological order. Note that not all of the steps shown in the following description are essential. In addition, steps other than those shown in the following description may also be executed.

[0045] 5.1. Workflow Information (Figure 10, Figure 11, Figure 12) FIG. 10 shows the steps executed by the workflow manager 212 . In step 1001 of FIG. 10, the workflow management unit 212 reads the workflow information 207 in which the workflow of the process 107 is described.

[0046] 11 shows a first example of the workflow information 207. The workflow information 207a (or workflow information 1100) shown in FIG. 11 corresponds to the workflow 607 shown in FIG. The workflow information 1100 may include point information 1101, 1103, 1105, and 1107, and phase information 1102, 1104, 1106, and 1108. (The point information and phase information may be called workflow entities.) Each of the point information 1101, 1103, 1105, and 1107 is information describing each of the points 671, 673, 675, and 677 in Fig. 6. Each of the phase information 1102, 1104, 1106, and 1108 is information describing each of the phases of the catalog service 672, the ETL service 674, the analysis service 676, and the model output service 678 in Fig. 6. Each piece of point information may be information describing the control details of the transition between phases (start or end of a phase) and the control details of the repetition of a phase. Each piece of phase information may have information indicating the type of service program (service program name) or the type of I / O data (I / O data name) to be executed in the corresponding phase. The I / O data name may be the name of the volume (e.g., virtual volume) in which the I / O data is managed. For example, phase information 1102 corresponding to the phase of catalog service 672 may have one or both of catalog service program name 1121 and catalog I / O data (volume) name 1122. Phase information including the above information is useful when determining the operation mode or requesting the execution of a service. 11, any of the phase information may have explicit operation mode information that explicitly specifies the operation mode 108 to be applied in the corresponding phase. As described above, the operation mode 108 may be an operation mode associated with the corresponding phase, or an operation mode related to the handling of the virtual volume 104 (or cache volume 105) in which the data accessed in that phase is managed.

[0047] Fig. 12 shows a second example of the workflow information 207. The workflow information 207b (or workflow information 1200) shown in Fig. 12 corresponds to a workflow different from the workflow 307 shown in Fig. 3 and the workflow 607 shown in Fig. 6. 12, the workflow information 1200 may include a two-way branch from point information 1203 to phase information 1204 and phase information 1208, may include a two-way merging from phase information 1204 and phase information 1208 to point information 1205, and may include a route for repetition from point information 1207 to phase information 1202. In this manner, the workflow information 1200 may be described according to the aspect of the workflow. Based on the contents of the point information 1203, only one of the phase indicated by the phase information 1204 and the phase indicated by the phase information 1208 may be executed, or both of these two phases may be executed in parallel. 12, explicit operation mode information is included in phase information 1202 and phase information 1208. In this manner, in a phase corresponding to phase information having explicit operation mode information, the explicit operation mode indicated by the explicit operation mode information may be preferentially set as operation mode 108 in that phase.

[0048] 5.2. Operation mode determination (Fig. 10, Fig. 13, Fig. 14, Fig. 15) Returning to the explanation of Fig. 10, in step 1002 in Fig. 10, the workflow management unit 212 sets the first phase of the processing in the workflow described in the workflow information 207 read in step 1001 as the "current phase of the processing". For example, in the example of Fig. 11, the phase indicated by the phase information 1102 (the phase of the catalog service 672) and in the example of Fig. 12, the phase indicated by the phase information 1202 are the "first phase of the processing". 10, the workflow management unit 212 requests the operation mode determination unit 101 to determine the operation mode 108 corresponding to the "current phase of processing". When making this request, the workflow management unit 212 may transmit to the operation mode determination unit 101 phase information corresponding to the "current phase of processing" from among the information included in the workflow information 207. (Alternatively, the operation mode determination unit 101 itself may directly read the workflow information 207 and use the phase information corresponding to the "current phase of processing".)

[0049] Figure 13 shows steps executed by the operation mode determination unit 101. Each of the steps included in Figure 13 may be considered to be one of the operation mode determination steps. 13, the operation mode determination unit 101 determines whether or not it has been requested by the workflow management unit 212 to determine the operation mode 108 corresponding to the "current phase of processing." Step 1301 may be repeated as long as the determination result of step 1301 is negative. On the other hand, when the above-mentioned step 1003 is executed, the determination result of step 1301 becomes positive, and control transitions to step 1302.

[0050] In step 1302 of FIG. 13, the operation mode determination unit 101 determines whether or not the operation mode 108 corresponding to the "current phase of processing" is explicitly specified in the workflow information 207. In the example of FIG. 11 and the example of FIG. 12, the phase information 1202 and the phase information 1208 each have the first explicit operation mode information 1223 and the third explicit operation mode information 1283, respectively, while the other phase information does not have the explicit operation mode information. If the phase information corresponding to the "current phase of processing" has the explicit operation mode information, the determination result of step 1302 becomes positive, and the control is transitioned to step 1303. If the phase information corresponding to the "current phase of processing" does not have the explicit operation mode information, the determination result of step 1302 becomes negative, and the control is transitioned to step 1304. In step 1303 of FIG. 13, the operation mode determination unit 101 sets the operation mode 108 associated with the "current phase of processing" (or the operation mode 108 related to the handling of the virtual volume 104 (or the cache volume 105) in which the data accessed in the "current phase of processing" is managed) as the operation mode indicated by the explicit operation mode information possessed by the phase information corresponding to the "current phase of processing" in the workflow information 207. The operation mode indicated by the explicit operation mode information may be any one of the "remote mode", "cache mode", and "purge mode" as in the example of FIG. 3, or may be an "off mode (OFF mode)" or "on mode (ON mode)" (which can be set separately for each of read access and write access), or a "purge mode" (which is set regardless of access) as in the example of FIG. 6. After step 1303, the control transitions to step 1307.

[0051] (Because no operating mode is explicitly specified for the "current phase of processing"), in step 1304 of FIG. 13, the operating mode determination unit 101 determines whether or not an operating mode 108 is defined in the attribute table 214 for the type of service to be executed in the "current phase of processing" or the type of volume (virtual volume) in which the data accessed in the "current phase of processing" is managed.

[0052] FIG. 14 shows a first example of the attribute table 214. The attribute table 214a (or the attribute table 1400) shown in FIG. 14 may show a correspondence between the type of service (e.g., microservice) executed in a phase, the operation mode at the time of read access accompanying the execution of the service (microservice), and the operation mode at the time of write access accompanying the execution of the service (microservice). For example, an "OFF mode" at the time of read access and an "ON mode" at the time of write access are associated with an ETL service. The attribute table 214a (or the attribute table 1400) may also show a correspondence between the type of service (e.g., microservice) executed in a phase and a "purge mode". For example, a "purge mode" that is unrelated to read access or write access is associated with a model output service. In addition, the "first type of analysis service" shown in Figure 14 may be, for example, a service that performs aggregation processing without machine learning, while the "second type of analysis service" may be a service that trains a learning model (model parameters) by performing machine learning. By setting the operation mode 108 indicated in the attribute table 1400 for various services (microservices), it becomes possible to handle virtual volumes and cache volumes in a manner appropriate to the way data is used (context) in each service (microservice). In step 1304 of FIG. 13, if the type of service (microservice) executed in the “current phase of processing” is registered in the attribute table 214a (or attribute table 1400), the operation mode determination unit 101 may determine that the determination result of step 1304 is positive.

[0053] FIG. 15 shows a second example of the attribute table 214. The attribute table 214b (or attribute table 1500) shown in FIG. 15 may show the correspondence between the type of volume (logical volume) in which data accessed in a phase is managed, the operation mode at the time of read access to the volume (logical volume), and the operation mode at the time of write access to the volume (logical volume). In the attribute table 1500, an "OFF mode" at both read access and write access is associated with a catalog volume that manages meta information for searching for data to be analyzed. An "OFF mode" at the time of read access and an "ON mode" at the time of write access are associated with a virtual lake that includes a volume in which non-structured data of the data to be analyzed is accumulated. An "ON mode" at both read access and write access is associated with a virtual mart that includes a volume in which structured data of the data to be analyzed is accumulated. The WORK volume, which is a temporary workspace used by the analysis service, is assigned the "ON mode" for both read and write access. The MODEL volume, which stores trained learning models (model parameters), is assigned the "PURGE mode" regardless of read or write access. By setting the operating mode 108 indicated in the attribute table 1500 for various volumes, and for virtual lakes and virtual marts that contain volumes, it becomes possible to handle virtual volumes and cache volumes in a manner appropriate to the usage (context) of the data managed by each volume. In step 1304 of FIG. 13, if the type of volume (logical volume) in which the data accessed in the “current phase of processing” is managed is registered in attribute table 214b (or attribute table 1500), the operation mode determination unit 101 may determine that the result of the determination in step 1304 is positive.

[0054] If the determination result in step 1304 is positive, the control is transferred to step 1305. If the determination result in step 1304 is negative, the control is transferred to step 1306. In step 1305 of FIG. 13, the operation mode determination unit 101 sets the operation mode 108 associated with the "current phase of processing" (or the operation mode 108 related to the handling of the virtual volume 104 (or the cache volume 105) in which the data accessed in the "current phase of processing" is managed) based on the attribute table 214 (for example, the attribute table 1400 of FIG. 14 or the attribute table 1500 of FIG. 15). Note that, if the type of the service executed in the "current phase of processing" is registered in the attribute table 1400 of FIG. 14 and the type of the volume (virtual volume) in which the data accessed in the "current phase of processing" is managed is registered in the attribute table 1500 of FIG. 15, the operation mode determination unit 101 may prioritize the operation mode indicated by any of the attribute tables, may determine the operation mode 108 according to some rule, or may treat it as an error. After step 1305, the control transitions to step 1307.

[0055] 13, the operation mode determination unit 101 may appropriately set the operation mode 108 associated with the "current phase of processing" (or the operation mode 108 related to the handling of the virtual volume 104 (or cache volume 105) in which the data accessed in the "current phase of processing" is managed) to one of the candidate operation modes. For example, the operation mode determination unit 101 may set the operation mode 108 in the "current phase of processing" to the "cache mode" or "on mode (ON mode)" for both read access and write access by default.

[0056] 13, the operation mode determination unit 101 notifies the storage management unit 251 in the management server 205 of the operation mode information 208 indicating the operation mode 108 determined in any one of steps 1303, 1305, and 1306. After step 1307, control is returned to step 1301.

[0057] 5.3. Operation mode instruction (Fig. 16) FIG. 16 shows the steps executed by the storage manager 251 . 16, the storage management unit 251 judges whether the operation mode information 208 indicating the operation mode 108 corresponding to the "current phase of processing" has been notified from the operation mode judgment unit 101. Step 1601 may be repeated as long as the judgment result of step 1601 is negative. On the other hand, when the above-mentioned step 1307 is executed, the judgment result of step 1601 becomes positive, and control transitions to step 1602. In step 1602 of FIG. 16, the storage management unit 251 transmits an operation mode instruction 252 indicating the operation mode 108 to be applied to one or more storage control units 103 that manage the virtual volume 104 (or cache volume 105) to which the operation mode 108 indicated by the operation mode information 208 is applied. Here, as also shown in FIG. 2, the storage management unit 251 may also transmit the operation mode instruction 252 to another intra-system storage control unit 293 that manages the real volume 106 for the virtual volume 104 to which the operation mode 108 indicated by the operation mode information 208 is applied, among other intra-system storage control units 293 in the other system 190. (Note that, if the control performed by the other intra-system storage control unit 293 is limited to input / output of a page requested for access by the storage control unit 103, the operation mode instruction 252 may not be transmitted to the other intra-system storage control unit 293.) With the above steps, the operation mode 108 for the "current phase of processing" is now ready to be applied to the virtual volume 104 (or cache volume 105) accessed by the "current phase of processing".

[0058] 4. Service Execution (Fig. 10, Fig. 17) Returning to the explanation of FIG. 10, in step 1004 in FIG. 10, the workflow management unit 212 requests the service execution unit 213 to execute the "current phase of processing" (execution related to the workflow entity). When making this request, the workflow management unit 212 may transmit to the service execution unit 213 phase information corresponding to the "current phase of processing" from among the information included in the workflow information 207 (for example, information such as the type of service (microservice) executed in the "current phase of processing" and the volume (virtual volume) in which the accessed data is managed). (Alternatively, the service execution unit 213 itself may directly read the workflow information 207 and use the phase information corresponding to the "current phase of processing".)

[0059] FIG. 17 shows the steps executed by the service execution unit 213. 17, the service execution unit 213 judges whether or not it has been requested by the workflow management unit 212 to execute the "current phase of processing" (execution related to a workflow entity). Step 1701 may be repeated as long as the judgment result of step 1701 is negative. On the other hand, when the above-mentioned step 1004 is executed, the judgment result of step 1701 becomes positive, and control transitions to step 1702. In step 1702 (processing execution step) of FIG. 17, the service execution unit 213 executes the "current phase of processing". At this time, the service execution unit 213 may refer to, for example, information on the type of service (microservice) (transmitted from the workflow management unit 212 or read by the service execution unit 213 itself) and information on the volume (virtual volume) in which the accessed data is managed. The service execution unit 213 may, for example, execute a service program corresponding to the specified type of service, and perform read access or write access to the specified volume (virtual volume) or the like. (Alternatively, the service execution unit 213 (service execution program) may not be distinguished from various service programs, and the service execution unit 213 may simply execute the service program corresponding to the specified type of service without the service execution program.) In conjunction with the execution of the “current phase of processing” in step 1702 of FIG. 17 (execution related to the workflow entity), the service execution unit 213 may, as necessary, perform read or write access to the data managed in the virtual volume 104 from the storage control unit 103. In addition, in the case where "purge mode" is set for the virtual volume 104 when executing the "current phase of processing" in step 1702 of Figure 17 (execution related to the workflow entity), it is possible to wait for the completion of destaging from the cache volume 105 to the real volume 106 performed by the storage control unit 103. 17, the service execution unit 213 notifies the workflow management unit 212 that the execution of the "current phase of processing" (the service (microservice) executed in the "current phase of processing") has been completed. After step 1703, control is returned to step 1701.

[0060] 5.4.1. Control of Read and Write Access (Fig. 18, Fig. 19, Fig. 20) Fig. 18 shows steps executed by the storage control unit 103 during read access or write access. Each of the steps included in Fig. 18 may be considered to be one of the storage control steps. In Fig. 18, circles containing numbers are connected to each other at the transition between steps in the flowchart. 18, the storage control unit 103 judges whether or not a read access or a write access to data managed in the virtual volume 104 has arrived from the service execution unit 213, which is associated with the execution of the "current phase of processing" (execution related to the workflow entity). Step 1801 may be repeated as long as the judgment result of step 1801 is negative. On the other hand, when a read access or a write access is executed in the above-mentioned step 1702, the judgment result of step 1801 becomes positive, and control transitions to step 1802. 18, the storage control unit 103 checks the location of data that is the target of the read access or write access received from the service execution unit 213. At this time, the storage control unit 103 refers to the conversion table 231.

[0061] Fig. 19 shows an example of the location of a page including data that is the target of a read access or a write access, and Fig. 20 shows a conversion table 231 (or a conversion table 2000) corresponding to the example of Fig. 19. In the example of FIG. 19, the virtual volume 104 has at least five pages, "VA", "VB", "VC", "VD", and "VE", as virtual volume pages 1903, which are virtual pages. The real volume 106 has at least five pages, "RA", "RB", "RC", "RD", and "RE", as real volume pages 1901, which are real pages. Here, "VA" and "RA" correspond to each other, "VB" and "RB" correspond to each other, "VC" and "RC" correspond to each other, "VD" and "RD" correspond to each other, and "VE" and "RE" correspond to each other. In the example of FIG. 19, for "VA", "VC", and "VD" of the virtual volume pages 1903, "CA", "CC", and "CD", which are corresponding cache volume pages 1902 (page copies), exist in the cache volume 105. On the other hand, in the example of FIG. 19, for “VB” and “VE” of the virtual volume pages 1903, the corresponding cache volume pages 1902 (page copies) do not exist in the cache volume 105. The size of the pages handled above may be determined as appropriate. By adjusting the page size, the balance between access responsiveness and throughput can be adjusted. A translation table 2000 corresponding to the example of Fig. 19 may be as shown in Fig. 20. The translation table 2000 shown in Fig. 20 shows the correspondence between the virtual volume page 1903 and the real volume page 1901. Furthermore, if a cache volume page 1902 exists for the virtual volume page 1903, the translation table 2000 also shows the correspondence between the virtual volume page 1903 and the cache volume page 1902. In step 1802 of FIG. 18, the storage control unit 103 confirms that the virtual volume page 1903 including the data to be accessed by read access or tight access is registered in the conversion table 2000. Then, the storage control unit 103 confirms whether or not a cache volume page 1902 is associated with the virtual volume page 1903 to be accessed in the conversion table 2000. In the example of FIG. 19 and FIG. 20, if the virtual volume page 1903 including the data to be accessed is "VC", the cache volume page 1902 exists as "CC". On the other hand, if the virtual volume page 1903 including the data to be accessed is "VE", the corresponding cache volume page 1902 does not exist, and the actual page becomes "RE" as the real volume page 1901. In step 1803 of FIG. 18, the storage control unit 103 judges whether or not a cache volume page 1902 corresponding to a virtual volume page 1903 including data that is the target of read access or write access exists in the cache volume 105 (cache hit or cache miss). In other words, the storage control unit 103 judges whether or not a page copy (cache volume page 1902) of a real volume page 1901 including data that is the target of read access or write access exists in the cache volume 105. If the judgment result in step 1803 is positive (cache hit), control is transferred to step 1804. If the judgment result in step 1803 is negative (cache miss), control is transferred to step 1805. 18, the storage controller 103 executes a read access or a write access to the cache volume page 1902 in the cache volume 105 which has caused a cache hit. After step 1804, the control is returned to step 1801.

[0062] 18, the storage control unit 103 determines whether the type of the access request that resulted in the cache miss is a read access or a write access. If the cache miss is a read access, control transitions to step 1806. If the cache miss is a write access, control transitions to step 1807. In step 1806 of FIG. 18 (upon the determination of a cache miss for the read access), the storage control unit 103 checks the type of the operation mode 108 for the read access for the "current phase of the process" itself or for the virtual volume 104 accessed by the "current phase of the process". The operation mode 108 may be the operation mode indicated by the operation mode instruction 252 transmitted from the storage management unit 251 to the storage control unit 103 in step 1602. If the operation mode 108 is the "cache mode" described in FIG. 3 or the "ON mode" for the read access described in FIG. 6, control is transitioned to step 1808. If the operation mode 108 is the "remote mode" described in FIG. 3 or the "OFF mode" for the read access described in FIG. 6, control is transitioned to step 1810. In step 1807 of FIG. 18 (upon determination of a cache miss for the write access), the storage control unit 103 checks the type of operation mode 108 for the write access for the "current phase of processing" itself or for the virtual volume 104 accessed by the "current phase of processing". If the operation mode 108 is the "cache mode" described in FIG. 3 or the "ON mode" for the write access described in FIG. 6, control is transitioned to step 1808. If the operation mode 108 is the "remote mode" described in FIG. 3 or the "OFF mode" for the write access described in FIG. 6, control is transitioned to step 1810.

[0063] In step 1808 of FIG. 18 (due to the operating mode in which a page including miss-hit data is cached), the storage control unit 103 stages a real volume page 1901 including data subject to read access or write access from the real volume 106 to the cache volume 105. The staged page copy becomes a cache volume page 1902. This staging is as already shown in the central diagram of FIG. 4. (The so-called write allocate method is used to deal with a write miss. Note that there are various variations in the way that a write miss can be dealt with. Such variations may be applied to the present disclosure.) (Upon completion of staging) in step 1809 of Figure 18, the storage control unit 103 performs read access or write access to the staged cache volume page 1902. The state of read access or write access after staging is as already shown in the diagram on the right side of Figure 4. After step 1809, control is returned to step 1801.

[0064] 18, the storage control unit 103 performs a read access or a write access to the real volume page 1901 including the data that is the target of the read access or the write access (because this is an operating mode in which the page including the mishit data is not cached). At this time, staging from the real volume 106 to the cache volume 105 is not performed. After step 1809, control is returned to step 1801.

[0065] 5.4.2. Purge mode control (Fig. 21) Fig. 21 shows steps executed by the storage control unit 103 when the "purge mode" is specified as the operation mode 108. Each of the steps included in Fig. 21 may be considered to be one of the storage control steps. In step 2101 of FIG. 21, the storage control unit 103 judges whether or not an instruction to specify "purge mode" for the "current phase of processing" (workflow entity) itself or for the virtual volume 104 (or cache volume 105) associated with the "current phase of processing" has arrived as an operation mode instruction 252 from the storage management unit 251. Step 2101 may be repeated while the judgment result of step 2101 is negative. On the other hand, when the storage management unit 251 transmits the operation mode instruction 252 instructing the "purge mode" to the storage control unit 103 in the above-mentioned step 1602, the judgment result of step 2101 becomes positive, and control transitions to step 2102. In step 2102 of FIG. 21, the storage control unit 103 sets the first page in the virtual volume 104 that is the target of destaging in the "purge mode" as the "current page" in the purge mode control in FIG. In step 2103 of FIG. 21, the storage control unit 103 accesses the conversion table 231 (or the conversion table 2000) to obtain information related to the "current page." In step 2104 of FIG. 21, the storage control unit 103 determines whether or not a cache volume page 1902 corresponding to the "current page" exists in the cache volume 105 (whether or not there is a cache hit) based on the information contained in the conversion table 231 (or conversion table 2000). If the cache volume page 1902 corresponding to the "current page" exists (cache hit), control transitions to step 2105. If the cache volume page 1902 corresponding to the "current page" does not exist (cache miss), control transitions to step 2108.

[0066] In step 2105 of FIG. 21 (in the case of a cache hit for the "current page"), the storage control unit 103 may determine whether or not a write access has been performed on the cache volume page 1902 corresponding to the "current page" (whether or not the cache volume page 1902 is dirty data) after the cache volume page 1902 corresponding to the "current page" has been staged from the real volume page 1901. In order to enable this determination, a dirty flag indicating the presence or absence of a write access for each page may be provided in the conversion table 231 (or the conversion table 2000). If the determination result in step 2105 is positive (if the data is dirty), control may transition to step 2106. If the determination result in step 2105 is negative (if the data is clean data and not dirty data), control may transition to step 2107. 21, in step 2106, the storage control unit 103 destages the cache volume page 1902 corresponding to the "current page" to the real volume page 1901. The destaging process has already been shown in the central diagram of FIG. 21, the storage control unit 103 may destage the cache volume page 1902 corresponding to the "current page" to the real volume page 1901, without distinguishing whether the data is dirty or not. In this case, if the determination result of step 2104 is positive (if there is a cache hit), step 2105 may not be executed and control may transition to step 2106. If such control is performed, the conversion table 231 (or conversion table 2000) does not need a dirty flag indicating the presence or absence of a write access for each page. In step 2107 of FIG. 21 (after destaging of the “current page” as necessary), the storage controller 103 erases the cache volume page 1902 corresponding to the “current page” from the cache volume 105 .

[0067] In step 2108 of Figure 21, the storage control unit 103 judges whether or not the "current page" is the last page in the virtual volume 104 that is to be destaging by the "purge mode". If the judgment result in step 2108 is positive, control is transferred to step 2110. If the judgment result in step 2108 is negative, control is transferred to step 2109. In step 2109 of Fig. 21, the storage control unit 103 sets the page next to the "current page" in the virtual volume 104 that is the target of destaging in the "purge mode" as the new "current page". In other words, the page that is the target of destaging is moved to the next page. After step 2109, control is returned to step 2103. In step 2110 of FIG. 21 (after processing by purge mode control has been completed for all pages included in the virtual volume 104 that is the target of destaging by the "purge mode"), the storage control unit 103 may block the virtual volume 104 that is the target of destaging by the "purge mode". Alternatively, the storage control unit 103 may not block the virtual volume 104 that is the target of destaging by the "purge mode". Also, the storage control unit 103 may block only the cache volume 105 without blocking the virtual volume 104. After step 2110, control is returned to step 2101.

[0068] 5.5 Remaining control of workflow management (Figure 10) Returning to the explanation of Fig. 10, in step 1005 in Fig. 10, the workflow management unit 212 judges whether or not there is a notice of completion of the "current phase of processing" from the service execution unit 213 (based on step 1703). As long as the judgment result of step 1005 is negative, step 1005 may be repeated. When the judgment result of step 1005 becomes positive, control transitions to 1006. 10, the workflow management unit 212 judges whether or not the entire processing related to the workflow described in the workflow information 207 is completed. If the judgment result in step 1006 is positive, the control of the workflow information 207 by the workflow management unit 212 is completed. If the judgment result in step 1006 is negative, the control transitions to step 1007. 10, the workflow management unit 212 sets the next phase to be executed as the new "current phase of processing". At this time, the point information contained in the workflow information 207 may be used. After step 1007, the control returns to step 1003.

[0069] 6.Other (variations) The present disclosure is not limited to the above-described embodiment, but includes various modifications. A part of the configuration or processing of the embodiment may be replaced with the configuration or processing of another conceivable embodiment. The configuration or processing of the embodiment may be added to the configuration or processing of another conceivable embodiment. For example, the present disclosure may have the following modified embodiments.

[0070] (1) Hybrid cloud combination In the above embodiment, the system 100 that performs access via a virtual volume is a public cloud, and the other system 190 having a real volume is an on-premise environment or a private cloud. However, the present disclosure can be widely applied to storage systems constructed between systems interconnected by a network. For example, the system 100 that performs access via a virtual volume may be an on-premise environment or a private cloud, and the other system 190 having a real volume may be a public cloud. Also, the system 100 that performs access via a virtual volume may be a first on-premise environment or a private cloud, and the other system 190 having a real volume may be a second on-premise environment or a private cloud. Furthermore, the system 100 that performs access via a virtual volume may be a first public cloud, and the other system 190 having a real volume may be a second public cloud. As described above, the present disclosure can be widely used.

[0071] (2) Variations in cache control In the above embodiment, the operation modes 108 other than the "purge mode" stipulate the processing to be performed when the page including the data to be accessed does not exist in the cache volume 105 (cache miss). When the page including the data to be accessed exists in the cache volume 105 (cache hit), read access or write access is performed to the hit cache volume page 1902 regardless of the operation mode 108. In a modified example, even in the case of a cache hit, the contents of control may be stipulated by the setting of the operation mode 108. For example, when "remote mode" or "off mode (OFF mode)" for read access is set, it is not necessary to perform a cache hit determination using the conversion table 231 at the time of read access. Alternatively, in a similar case, if a cache hit occurs at the time of read access, the hit cache volume page 1902 may be deleted, and then a read access to the real volume page 1901 may be performed. Various variations in cache control are possible, and such variations may be applied to the present disclosure. The above modification makes it possible to realize flexible cache control in accordance with the way data is used (context) in the processing performed by the system 100.

[0072] (3) Virtual computer environment (Figure 22) Various server nodes, functional units, etc. may be implemented virtually. 22 shows a modified example in which various servers, nodes, functional units, etc. constituting the system 100 are virtually implemented. The system 100 may have one or more physical nodes 2203 and one or more other resources 2204 as resources. Here, the physical nodes 2203 may have the same configuration as each other, or may have different configurations from each other. Similarly, the other resources 2204 may have the same configuration as each other, or may have different configurations from each other. The system 100 may include a virtual computer manager 2202. The virtual computer manager 2202 (VMM or hypervisor) may be realized by any of the computer resources included in the system 100 executing a program for implementing a virtual computer environment 2201 (partition). Through control by the virtual computer manager 2202, one or more physical nodes 2203 and one or more other resources 2204 of the system 100 can be treated as virtualized resources that can be provided in a virtual computer environment 2201 (partition). Therefore, resources in the virtual computer environment 2201 (partition) can be allocated to various server nodes, functional units, etc. (for example, each of the application servers 201 (compute nodes 701), each of the storage nodes 203, and the management server 205), and various server nodes, functional units, etc. can be realized. Here, for example, resources of a plurality of physical nodes 2203 can be allocated to one application server 201. Also, resources of a common physical node 2203 can be divided and allocated to a plurality of application servers 201. The same applies to the storage nodes 203, the management server 205, the functional units, etc. By using a virtual computer environment, it is possible to realize a flexible system configuration because it is possible to realize various virtual servers, nodes, functional units, etc. that are different from the configuration of resources that the system 100 physically has. Note that other systems 190 can also use a virtual computer environment.

[0073] (4) Batch judgment of operation mode In the above embodiment, the operation mode 108 for the "current phase of processing" is determined by the steps executed by the operation mode determination unit 101 shown in FIG. In a modified example, the operation mode determination unit 101 may analyze all workflow entities such as point information and phase information described in the workflow information 207 in a lump sum, and determine the operation mode 108 for each processing phase in a lump sum. The above modification can reduce the cooperation between the workflow management section 212 and the operation mode determination section 101, and therefore can increase the degree of freedom in the timing of operation mode determination.

[0074] (5) Physical volumes not connected via a network In the embodiment described above, the network 180 is interposed between the system 100 and the real volume 106. In the embodiment described above, in a storage system in which the network 180 is interposed, it is possible to speed up access to data stored in the real volume 106, while also reducing the capacity, transfer amount, or management cost of the cached data that accompanies caching of the data. In a modified embodiment, the network 180 may not be interposed between the system 100 and the real volume 106. For example, the system 100 may have the real volume 106. Even in the above-mentioned modified example, if there is a sufficient difference in the data transfer capacity of the storage device (or recording device) for implementing the cache volume 105 and the data transfer capacity of the storage device (or recording device) for implementing the real volume 106, as viewed from the perspective of the process execution unit 102 (the service execution unit 213 therein), which is the entity that accesses the data, the modified example can have the same effects as the embodiment described above. In the above-mentioned modified example, for example, the storage device for implementing the cache volume 105 may be a memory such as SRAM or DRAM, and the storage device (or recording device) for implementing the real volume 106 may be a hard disk drive (HDD) or flash memory.

[0075] (6) Separation of application server and storage server In the embodiment described above, the system 100 includes both the application server 201 including the operation mode determination unit 101 and the process execution unit 102, and the storage node 203 including the cache volume 105 and the storage control unit 103. By configuring the system 100 in this way, data access (access request 109) via the virtual volume 104 and the operation mode instruction 252 can be performed quickly within the system 100. In a modified example, a computer system having an application server 201 including an operation mode determination unit 101 and a process execution unit 102 may be a system separate from the system 100 having a storage node 203 including a cache volume 105 and a storage control unit 103. In this modified example, data access (access request 109) via a virtual volume 104 and an operation mode instruction 252 are transmitted between different computer systems. As long as the delay due to transmission between different computer systems is tolerable, such a modified example is also possible. As described above, the present disclosure is applicable regardless of the arrangement of the operation mode determination unit 101, the process execution unit 102, the cache volume 105, and the storage control unit 103.

[0076] (7) Operation mode information - Variations of operation mode instructions In the embodiment described above, the content of the information transmitted (notified) from the operation mode determination unit 101 to the storage control unit 103 (via the storage management unit 251 as necessary) was the operation mode 108. Then, the storage control unit 103 performed control related to the cache volume 105 suitable for the transmitted (notified) operation mode 108. In a modified example, either the operation mode determination unit 101 or the storage management unit 251 may identify control contents for the cache volume 105 suitable for the operation mode 108. Either the operation mode determination unit 101 or the storage management unit 251 may directly or indirectly transmit (notify) information indicating the identified control contents to the storage control unit 103. In the above modification, the storage control unit 103 does not need to determine the control content for the cache volume 105 suitable for the operation mode 108, so that the hardware or software resources for the storage control unit 103 can be reduced.

[0077] (8) Control for cache volumes without purge mode In the embodiment described above, the purge mode is included as one of the operation modes 108 for control related to the cache volume 105. In a modified example, the purge mode does not have to be one of the operation modes 108 for control related to the cache volume 105. For example, when data targeted by the access request 109 misses in the cache volume 105, only an operation mode in which the miss-hit data (and the page including the data) is staged from the real volume 106 to the cache volume 105 (an operation mode in which data read from the real volume 106 is left in the cache volume 105) and an operation mode in which no staging is performed (an operation mode in which data read from the real volume 106 is not left in the cache volume 105) may exist as candidates for the operation mode 108. Even with the above-described modified example, it is possible to reduce the amount of data to be cached, reduce the amount of data transfer required for caching, or reduce the cost of managing the cached data.

[0078] The technical matters described in each of the embodiments of the present disclosure and the modified examples of the embodiments described above can be combined as appropriate as long as no technical contradiction occurs.

Claims

1. A computer system including a storage control unit and a storage device, The storage control unit providing a virtual volume to a processing execution unit that executes an application; managing data input / output to / from a real volume via the virtual volume; configuring a cache volume based on the storage area of ​​the storage device; receiving notification of information regarding an operation mode determined by an operation mode determination unit based on the processing executed by the processing execution unit; when receiving an access request via the virtual volume from the processing execution unit, if the target data of the access request is not stored in the cache volume, access is executed for the target data in the real volume; The computer system controls the cache volume based on the notified information about the operation mode.

2. 2. The computer system of claim 1, A computer system comprising the process execution unit and the operation mode determination unit.

3. 2. The computer system of claim 1, A computer system, wherein the real volume exists in another system connected to the computer system.

4. 3. A computer system according to claim 2, comprising: The operation mode determination unit determines a context of a workflow indicating the process based on attributes related to the process executed by the process execution unit, and then determines the operation mode corresponding to the context.

5. 3. A computer system according to claim 2, comprising: The operation mode determination unit determines the operation mode for each phase of the process included in a workflow indicating the process.

6. 3. A computer system according to claim 2, comprising: A computer system in which the operating mode determination unit determines the operating mode based on an operating mode specified in workflow information describing a workflow indicating the processing, an attribute of a service executed by the processing execution unit, or an attribute of a virtual volume accessed by the processing execution unit.

7. 2. The computer system of claim 1, The storage control unit, when receiving an access request from the processing execution unit, controls whether or not to leave data read from the real volume in the cache volume based on information relating to the notified operation mode.

8. 8. A computer system according to claim 7, comprising: the storage control unit deletes the data stored in the cache volume in response to notification of information relating to the predetermined operation mode; When deleting the data, the computer system stores dirty data that is not stored in the real volume in the real volume.

9. 8. A computer system according to claim 7, comprising: The storage control unit In an operation mode in which the processing execution unit executes a catalog service, data read from the real volume for read access and write access is not left in a cache volume, In an operation mode in which the processing execution unit executes an ETL service, data read from the real volume in response to a read access is not left in a cache volume, and data read from the real volume in response to a write access is left in the cache volume; In an operation mode in which the processing execution unit executes an analysis service, data read from the real volume in relation to read access and write access is left in a cache volume; A computer system in which, in an operation mode in which the process execution unit executes a model output service, data stored in the cache volume is not left in the cache volume.

10. 8. A computer system according to claim 7, comprising: The storage control unit In an operation mode in which the processing execution unit accesses a catalog volume, data read from the real volume for read access and write access is not left in a cache volume, In an operation mode in which the volume accessed by the processing execution unit is included in a virtual lake, data read from the real volume in response to a read access is not left in a cache volume, and data read from the real volume in response to a write access is left in the cache volume; In an operation mode in which the volume accessed by the processing execution unit is included in a virtual mart, data read from the real volume for read access and write access is left in a cache volume; In an operation mode in which the processing execution unit accesses a WORK volume, data read from the real volume for read access and write access is left in a cache volume; A computer system in which, in an operation mode in which the volume in which the process execution unit is involved is a model volume, data stored in the cache volume is not left in the cache volume.

11. 3. A computer system according to claim 2, comprising: The process execution unit includes a workflow management unit and a service execution unit, the workflow management unit manages the workflow and the service execution unit based on workflow information in which a workflow indicating the processing is described; The service execution unit executes the service included in the process.

12. 12. A computer system according to claim 11, comprising: A computer system, wherein the workflow information includes, for each phase of the processing, information for identifying a service corresponding to that phase, or information for identifying data to be input / output in that phase.

13. A method executed by a computer system having a storage control unit and a storage device, comprising: the storage control unit provides a virtual volume to a process execution unit that executes an application, manages data input / output to / from a real volume via the virtual volume, and configures a cache volume based on a storage area of ​​the storage device; The method comprises: a step of the storage control unit receiving a notification of information regarding an operation mode determined by an operation mode determination unit based on the process executed by the process execution unit; when the storage control unit receives an access request via the virtual volume from the processing execution unit, if target data of the access request is not stored in the cache volume, executing access to the real volume for the target data; The method includes a step of the storage control unit controlling the cache volume based on information related to the notified operation mode.

Citation Information

Patent Citations

  • Computing system and storage device

    JP2019149077A