Delivery of Event Notifications from a Distributed File System

The event log buffer synchronizes with recovery log buffers to ensure reliable delivery of event notifications in clustered file systems, addressing the issue of unreliable event information due to node crashes by flushing to disk storage, enabling continuous monitoring and compliance.

JP7703024B2Active Publication Date: 2025-07-04INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023524138
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-05
Filing Date
2021-10-20
Publication Date
2025-07-04
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

In distributed computing environments with clustered file systems, event notifications for file system operations are unreliable due to potential node crashes or interruptions, leading to incomplete or lost event information, which is critical for monitoring and compliance in sectors like finance and healthcare.

Method used

Implementing an event log buffer in each node that synchronizes with a recovery log buffer to flush event information to disk storage before the recovery log, ensuring consistent and reliable delivery of event notifications even in failure conditions.

Benefits of technology

Ensures reliable and uninterrupted delivery of event notifications by maintaining event information in persistent storage, allowing for playback and recovery in case of system interruptions, thus supporting continuous monitoring and compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007703024000001
    Figure 0007703024000001
  • Figure 0007703024000002
    Figure 0007703024000002
  • Figure 0007703024000003
    Figure 0007703024000003
Patent Text Reader

Abstract

In a method for fault tolerance in distribution of event information within a file system cluster, one or more processors identify event information associated with file system activity performed by nodes of the cluster. The one or more processors add the event information to an event log buffer in memory. The one or more processors receive a first log sequence number (LSN) associated with flushing recovery information from the recovery log buffer. The one or more processors identify event information in the event log buffer having a log sequence number less than or equal to the first log sequence number, and in response to determining that the event information includes a log sequence number less than or equal to the first log sequence number, the one or more processors flush the corresponding event information from the event log buffer to disk storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of distributed file system operations, and more particularly to reliable persistence of event notifications for file system operations.

Background Art

[0002] Companies that operate in multiple regions and often use clustered file systems hosted in a cloud environment are increasingly focusing on data security and monitoring data access and other file activity events. To protect business-critical data from internal threats and external security breaches, many companies have included auditing features for their file systems.

[0003] Financial institutions and healthcare businesses have operational and legal needs to identify access and editing activities for ledgers, accounts and transactions, medical records, and personnel files. Some consulting services rely on secure file system access and activity operations to support client needs and maintain competitiveness in the market. In a distributed computing environment where a clustered file system is used, central information may be included in the recovery of file system activities prior to node reboot crashes or hangs, or compliance rules may be required.

[0004] Distributed file systems often include multiple nodes that perform operation activities on the system's files, objects, or records. Nodes may include message queues that are distributed across multiple nodes and support the collection of file activities or events performed on files by each node. The publication of event activities to the message queue is performed by a "producer" component that is executed within a daemon on the node and applies a policy that sets the files and event activities of interest.

Summary of the Invention

[0005] Aspects of the present invention provide a method, computer program product, and system for fault tolerance in the distribution of event information of nodes within a computer node cluster. The method provides one or more processors for identifying event information associated with file system activities, executed by nodes of a computer node cluster. The one or more processors add the event information to an event log buffer in memory. The one or more processors receive a first log sequence number that matches a recovery log sequence number corresponding to the flushing of recovery information from a recovery log buffer. In response to identifying event information included in the event log buffer having a log sequence number less than or equal to the first log sequence number and determining that the event information included in the event log buffer corresponds to a log sequence number less than or equal to the first log sequence number, the one or more processors flush the event information having the corresponding log sequence number from the event log buffer to disk storage.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

[0007] In an embodiment of the present invention, it is recognized that a particular application running on a distributed computing environment including a clustered file system is configured to generate and deliver event notifications associated with file event activities. Also, in an embodiment of the present invention, an application that monitors a file system is recognized to require a consistent and reliable status and activity associated with the file system and the component data of the file system. The file system monitoring application relies on event notifications that are delivered to the application as input for changes in the status, access, and data content of the file system, and event information may be lost due to node crashes, restarts, or other delivery problems, which may require significant intervention to correct. The resiliency of event notifications for the monitored file system is an important reliability issue for such applications. The storage system may be designated for storing files and documents, or may be configured to store items as objects.

[0008] Embodiments of the present invention provide reliable delivery of event notifications for storage management operations and activities regardless of the content type being stored. For the sake of brevity, the discussion of the embodiments refers to "files", "file systems", and "file system clusters" without being limited to "objects" or other units of content.

[0009] Embodiments of the present invention provide a method, a computer program product, and a computer system for reliable delivery of event notifications in a clustered file system. In embodiments of the present invention, an event log buffer is included in each operation node of the clustered file system and receives event metadata of file events executed by each node. In some embodiments of the present invention, the event log buffer has a corresponding event log that receives (i.e., flushes) event information of file accesses and activities processed on each node from the event log buffer and provides persistent storage (i.e., disk storage) of the event information.

[0010] In embodiments of the present invention, the event log buffer is synchronized with a recovery log buffer such that the event log buffer flushes event information stored in memory to disk before flushing the corresponding data of the recovery log buffer to a recovery log on disk. In some embodiments, in response to an interruption in event notification delivery to a notification sink, event information is reproduced from the event log stored on disk, ensuring complete delivery of file system event information.

[0011] In some embodiments of the present invention, a node of a computer cluster that operates a clustered file system includes a running daemon having an event producer that generates event notifications corresponding to file system activities executed on the node. Events of the clustered file system include, but are not limited to, activities such as access (and access ID), creation, open, close, removal, edit change, ownership change, security level change, relocation, etc. The producer of the node sends event information as metadata to an event log buffer and a sink. In one embodiment, the producer of the node sends event metadata to a message queue. In embodiments of the present invention, a message queue and a sequential log on disk are recognized as embodiments of a sink. The message queue is a stream processing software platform that provides handling of high throughput real-time data feeds. The message queue functions as a repository that collects event information notifications from each distributed node of the file system cluster, and thus the information can be used by interested applications. In one embodiment, the producer of the node writes event metadata to a sequential log file on persistent storage. The sink provides an access point for a file system monitoring application to receive event notifications of file system activities and metadata of file or object status, changes, and access information.

[0012] In addition to the data that is sent to the recovery log buffer and stored in the recovery log on disk, the event information includes additional metadata associated with each file system event. In an embodiment of the present invention, when a node failure occurs, it is recognized that the event information necessary to generate an event notification is not included in the recovery log, which is designed to perform the recovery of the operations of system components. The event log includes metadata such as, but not limited to, i-node numbers, part names, storage pools, user IDs, owner IDs, file attributes, file event activities, and event timestamps.

[0013] Embodiments of the present invention provide an event log buffer that flushes event information to persistent disk storage prior to flushing the recovery information contained in the recovery log buffer, ensuring that event notifications are reliably maintained and delivered even in the event of a crash, restart, or other system interruption. The event log provides event information under failure conditions where normal event activity information is not delivered to the sink.

[0014] Next, the present invention will be described in detail with reference to the drawings. FIG. 1 is a functional block diagram showing a distributed data processing environment, generally designated 100, according to an embodiment of the present invention. FIG. 1 provides only an example of one implementation and is not intended to impose any limitations on environments in which different embodiments may be implemented. Those skilled in the art can make many modifications to the depicted environment without departing from the scope of the present invention as described by the claims.

[0015] The distributed data processing environment 100 includes a plurality of server nodes of a distributed file system cluster, represented as intervening server nodes (not shown), including the first server node 110, the second server node 120, and up to the Nth server node 130. The distributed data processing environment 100 also includes a sink 140, a storage system application 160, and a storage system monitoring application 170, all interconnected via a network 150. The network 150 can be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a virtual local area network (VLAN), or any combination that can include a wired connection, a wireless connection, or an optical connection. Generally, the network 150 can be any combination of connections and protocols that supports communication and data transmission between intervening server nodes, including the first server node 110, the second server node 120, and up to the Nth server node 130, the sink 140, the storage system application 160, and the storage system monitoring application 170.

[0016] The intervening nodes (not shown) up to the first server node 110, the second server node 120, and the Nth server node 130 are nodes of a distributed file system cluster. The first server node 110 includes a copy of the event log program 300, a daemon 115, a producer 117, an event log 113, and a recovery log 119. The second server node 120 includes a copy of the event log program 300, a daemon 125, a producer 127, an event log 123, and a recovery log 129. The Nth server node 130 includes a copy of the event log program 300, a daemon 135, a producer 137, an event log 133, and a recovery log 139. In some embodiments of the present invention, the event log program 300 may be remotely hosted from the first server node 110, the second server node 120, and the Nth server node 130 to provide programmable instructions to each server node.

[0017] In some embodiments, the first server node 110, the second server node 120, and the Nth server node 130 are computing devices that operate in a distributed environment and are configured to perform file system operations. In some embodiments, the first server node 110, the second server node 120, and the Nth server node 130 can be blade servers, web servers, laptop computers, desktop computers, stand-alone mobile computing devices, smartphones, tablet computers, or other electronic devices or computing systems capable of receiving, transmitting, and processing data. In other embodiments, the first server node 110, the second server node 120, and the Nth server node 130 can be computing devices that interact with applications and services hosted and operating in a cloud computing environment. In another embodiment, the first server node 110, the second server node 120, and the Nth server node 130 can be netbook computers, personal digital assistants (PDAs), or other programmable electronic devices capable of generating and sending inputs to the event log program 300 and receiving programming instructions from the event log program 300. Alternatively, in some embodiments, the first server node 110, the second server node 120, and the Nth server node 130 may be communicatively connected to the remotely operating event log program 300. The first server node 110, the second server node 120, and the Nth server node 130 may include internal and external hardware components, which are depicted in more detail in FIG. 4.

[0018] Daemons 115, 125, and 135 are computer programs that run as background processes. Daemons 115, 125, and 135 each enable producers 117, 127, and 137 to identify file system event activities occurring at their respective nodes and generate event information associated with the activities of specific files or objects.

[0019] Producers 117, 127, and 137 are entities that operate inside daemons 115, 125, and 135 respectively. Producers 117, 127, and 137 identify file system event activities on their respective nodes, collect metadata for the events, and publish the events and metadata to a message queue. In an embodiment of the present invention, producers 117, 127, and 137 send event metadata to the event log buffers of their respective nodes, which is described in more detail in the discussion of FIG. 2.

[0020] Recovery logs 119, 129, and 139 are collections of log files within the first server node 110, the second server node 120, and the Nth server node 130 respectively, and include activities and status information that enable recovery of their respective nodes from crashes, restarts, or other processing interruptions. Recovery logs 119, 129, and 139 exist in the persistent storage of the first server node 110, the second server node 120, and the Nth server node 130 respectively. Recovery logs 119, 129, and 139 receive recovery information when the corresponding recovery log buffers of their respective server nodes reach or near capacity limits, and flush the recovery information from the respective recovery log buffers in memory to the corresponding recovery logs on disk.

[0021] Event Log 113, Event Log 123, and Event Log 133 are log files of event activities and related metadata in the first server node 110, the second server node 120, and the Nth server node 130. Event Logs 113, 123, and 133 exist in persistent storage and provide consistent event information in case of interruption or malfunction of message queue delivery of event information to a specified sink. Event Logs 113, 123, and 133 receive event information when the corresponding event log buffers of their respective server nodes reach or approach the capacity limit, and flush the event information from each event log buffer in memory to the corresponding event log on disk.

[0022] Event Log Program 300 is depicted as operating on each of the first server node 110, the second server node 120, and the Nth server node 130. In some embodiments of the present invention, Event Log Program 300 is a component of a distributed application that manages data of a clustered file system and is scalable to operate on multiple nodes operating within a file system cluster. The file system monitoring application uses file system event notifications to determine the state and status of the operations of creating and accessing files or objects, directories, and files / objects of file system components. In an application that monitors a file system, there may be a problem with recoverability where the delivery of event notifications to a sink from which the application accesses the event notifications is interrupted. An application that monitors a file system may not be able to tolerate delivery failures, and embodiments of the present invention provide a reliable source of event notifications for playback and delivery to an event notification sink in case of a failure condition of a node in a file system cluster.

[0023] In some embodiments, the event log program 300 determines an event activity and instructs each node's producer to send event information including event metadata to the node's event log buffer. In some embodiments, the event log program 300 attaches the event metadata of a particular event to a JSON file stored in the event log buffer. Each instance of event information receives a log sequence number (LSN) that corresponds to the same log sequence number of recovery information associated with the event stored in the node's recovery log buffer. For a given event of a file system activity, such as a "write" activity to a file, the LSN is assigned to the recovery information associated with the write event and sent to the recovery log buffer, and the same LSN is assigned to the event information of the write event sent to the event log buffer.

[0024] In some embodiments of the present invention, the event log program 300 receives an LSN corresponding to a set of recovery information in the recovery log buffer. The set of recovery information includes all recovery information having an LSN less than or equal to the LSN received by the event log program 300. For example, the recovery log buffer is approaching the capacity limit of the recovery information, and it becomes necessary to flush the recovery information from the buffer to the recovery log in the persistent disk storage. Flushing means the process of writing the contents in the volatile memory (i.e., the buffer) to the persistent storage on the disk. The recovery log buffer identifies the LSN corresponding to the event previously executed at the node (e.g., LSN = 1000). The recovery log buffer also includes multiple LSNs executed at the node before the event corresponding to LSN = 1000. The recovery log buffer identifies the maximum sequence number to be flushed to the disk storage, along with all other recovery information having an LSN less than the identified LSN = 1000 (i.e., LSN = 999, LSN = 998,... etc.), thus clearing the recovery log buffer and enabling the acceptance of subsequent recovery information from the file system activities executed on the node.

[0025] Before executing the flushing of recovery information to disk storage, the LSN identified by the recovery log buffer is received by the event log program 300, and the event log program 300 identifies event information in the event log buffer having an LSN less than or equal to the LSN identified by the recovery log buffer. For example, the event log program 300 receives an LSN = 1000 corresponding to the recovery information in the recovery log buffer that is intended to be flushed to the persistent disk storage in the recovery log. The event log program 300 identifies event information associated with events having an LSN less than or equal to LSN = 1000. The event log program 300 flushes the event information of the events with LSN less than or equal to 1000 in the event log buffer to the event log in the persistent storage on the disk. The event log program 300 communicates the completion of the flushing of the event information to the disk to the recovery log buffer, and then the recovery log buffer executes the flushing of the recovery information from the buffer to the recovery log in the persistent storage on the disk.

[0026] In some embodiments, the event log program 300 determines that the event log buffer is at or near capacity and starts flushing event information to the event log on the disk without starting from the LSN identified by the recovery log buffer. In some embodiments, the event log program 300 determines that the event log buffer does not contain event information having an LSN less than or equal to the LSN identified by the recovery log buffer, in which case the event log program 300 does not execute the flushing of the event information and instead communicates completion to the recovery log buffer to start the flushing of the recovery log buffer.

[0027] Sink 140 is a software application that receives event notification delivery via a message queue from respective producers 117, 127, 137 of the first server node 110, the second server node 120, and the Nth server node 130. Sink 140 provides an interface for a storage system monitoring application to access the event notification, determine the state and status of components of the file system cluster, and perform monitoring activities of the file system cluster.

[0028] The storage system application 160 includes one or more applications that provide file system storage management and tuning services for a file system cluster including the first server node 110, the second server node 120, intervening server nodes (not shown), and the Nth server node 130. In some embodiments, the event log program 300 operates as a module component of the storage system application 160.

[0029] The storage system monitoring application 170 includes the collection of client - side applications that access event notifications from the sink 140 and monitor the respective states, statuses, files, and directory activities of the file system. The storage system monitoring application 170 utilizes the uninterrupted delivery of event notifications to maintain consistent monitoring of each file system component and content, and uses event metadata to identify the states, statuses, and activities associated with the cataloging operations applied to the clustered file system. The storage system monitoring application 170 uses the delivered event notifications, which include activities such as system status, number of files, number of directories, detection of files to archive, etc., to identify users accessing files, identify content attributes (confidential files), determine heavy access, and provide services across the distributed storage environment.

[0030] Figure 2 is a functional block diagram showing an example of flushing an event - log buffer to an event log in persistent disk storage according to an embodiment of the present invention. Figure 2 includes a daemon 215, a recovery - log buffer 220, a recovery log 225, an event - log buffer 230, an event log 235, a flush LSN 240, and a completion response 250. The daemon 215 includes a producer 217 and a policy 219. As described above with respect to daemons 115, 125, and 135, the daemon 215 runs as a background application at a node of the file - system cluster. The producer 217 identifies file - system event activities on the node, collects event metadata, and publishes the events and metadata to a message queue. The producer 217 transmits event information, including metadata associated with events executed by the node, to the event - log buffer 230. In some embodiments, the producer 217 transmits recovery information to the recovery - log buffer 220.

[0031] Policy 219 includes rules and conditions related to the access and recording of event information related to the file system. In some embodiments, Policy 219 identifies events of files and activities included in the notifications published by Producer 217.

[0032] Recovery log buffer 220 collects recovery information associated with the event activities of the file system executed on each node. Recovery log buffer 220 exists in volatile memory and writes the recovery information about the file system held in recovery log buffer 220 to recovery log 225 existing in persistent memory on the disk. Recovery log buffer 220 receives information for recovering from processing interruptions such as crashes, node or component failures, or restarts.

[0033] Recovery log buffer 220 determines whether the buffer is full or almost full and sends flash LSN 240 to event log buffer 230. Event log program 300 uses flash LSN 240 to identify the event information included in event log buffer 230 corresponding to the recovery information in recovery log buffer 220 and flushes it to event log 235 located in persistent storage on the disk. Flash LSN 240 enables event log program 300 to maintain event log buffer 230 in synchronization with recovery log buffer 220. Following the flushing of the event information corresponding to flash LSN 240 or LSNs less than flash LSN 240 from event log buffer 230 to event log 235 existing in persistent storage on the disk, recovery log buffer 220 receives completion response 250 from event log program 300.

[0034] The event log buffer 230 receives event information from the producer 217 associated with each event activity executed on the node. The event log buffer 230 attaches the event information of each event activity to a JSON file. The flash LSN 240 is received by the event log program 300 and reflects a set of recovery information that is flushed from the recovery log buffer 220 when the recovery log buffer is full. The set of recovery information includes recovery information having a corresponding LSN below the flash LSN 240. The event log program 300 uses the received flash LSN 240 to determine whether the event log buffer 230 contains event information corresponding to the set of recovery information identified by the flash LSN 240. If the event log buffer 230 contains event information corresponding to an LSN below the flash LSN 240, the event log program 300 starts flushing the identified event information from the event log buffer 230 to the event log 235 on persistent storage on disk. Flushing of the event log buffer event information prior to flushing of the recovery log buffer 220 ensures that the event information on persistent storage includes event information that is consistent with the file system state to be reproduced for event notification delivery in case of an interruption of the file system event notification delivery due to a failure of the message queue and sink.

[0035] Flash LSN 240 includes a log sequence number corresponding to recovery information for a file system activity designated to be flushed from the recovery log buffer 220 to the recovery log 225. Flushing the recovery log buffer 220 currently also includes recovery information corresponding to LSNs less than the Flash LSN 240 that are currently in the recovery log buffer 220. The event log program 300 can use the Flash LSN 240 to identify event information in the event log buffer 230 and flush it to the event log 235. Flushing the event information includes event information having a corresponding LSN less than the Flash LSN 240 included in the event log buffer 230.

[0036] The completion response 250 is a transmission from the event log program 300 that confirms that the flushing of the event log buffer 230 is complete and indicates that the flushing of the recovery log buffer 220 can begin. In an embodiment of the present invention, the event log buffer 230 flushes event information to the event log 235 based on event information having an LSN less than or equal to the Flash LSN 240 prior to flushing the recovery information of the recovery log buffer 220 to the recovery log 225. In some embodiments, the event log buffer 230 may flush event information to the event log 235 due to the event log buffer 230 being near or at capacity. In an embodiment where the event log program 300 determines that the event log buffer 230 does not include event information having a corresponding LSN less than or equal to the Flash LSN 240, the event log program 300 transmits a completion response to the recovery log buffer 220 and starts flushing the recovery information to the recovery log 225.

[0037] FIG. 3 is a flowchart showing the operation steps of an event log program operating in the distributed data processing environment of FIG. 1 according to an embodiment of the present invention. The event log program 300 receives a log sequence number reflecting a set of recovery information designated to be flushed from the recovery log buffer of a node to the recovery log in persistent storage. The designated flushing includes recovery information having an LSN less than or equal to the received log sequence number. The event log program 300 determines whether the event log buffer includes event information having a corresponding LSN that is less than or equal to the LSN received from the recovery log buffer.

[0038] The event log program 300 identifies the file system event activity of a node of the file system cluster (step 310). The event log program 300 determines the start of a file access activity executed by the node. The event log program 300 identifies the type of activity included as event information and the metadata associated with the event activity. For example, the first server node 110 receives an instruction to execute a write operation to a file in the file system. The event log program 300 of the first server node 110 detects a file event activity and prepares to receive metadata corresponding to the event activity executed on the node.

[0039] The event log program 300 adds event information associated with the event activity of the node to the event log buffer (step 320). The event log program 300 receives the event information of the event activity, and this event information includes metadata associated with the activity and the file in which the event activity is executed. The event log program 300 receives event information uniquely associated with a specific event activity, forms a JSON message, and attaches the message to the event log buffer existing in the volatile memory. The event log buffer collects the event information of multiple event activities on the node and is provided to the recovery log buffer also existing in the volatile memory for the addition and adjustment of recovery information related to a specific event activity.

[0040] For example, when receiving event information including specific metadata associated with the write operation to the file of the file system executed on the first server node 110, the event log program 300 attaches the event information to the JSON file of the event log buffer. The event information includes metadata specifically associated with the write operation executed on the file of the file system to which the write operation is directed. The event information receives a log sequence number (LSN) that uniquely identifies the metadata associated with a specific write event activity. The write event activity is uniquely associated with a specific write activity and also generates a corresponding set of recovery information that is sent to the recovery log buffer existing in the memory. The event information contains significantly more details of the metadata associated with the write event activity to a specific file, while the recovery information contains only the information for supporting failure recovery at the node. The recovery log buffer receives the event log buffer received for the event information of the same write event activity and the log sequence number (LSN) for the same received recovery information.

[0041] The event log program 300 receives a first log sequence number (LSN) associated with the flushing of the recovery log buffer (step 330). The term "first" log sequence number is used to distinguish a particular LSN from other LSNs and does not necessarily correspond to the first LSN such as 0001. The recovery log buffer continues to receive recovery information associated with event activity on the node. The event log program 300 receives an LSN when the recovery log buffer reaches or approaches the capacity level at which the recovery log buffer adds recovery information to the recovery log buffer. The LSN received by the event log program 300 corresponds to the high or highest sequence number LSN of the recovery information currently held in the recovery log buffer to be flushed to disk, and the flushing of the recovery information includes all LSNs below the first LSN. The event log program 300 receives the same first LSN specified for the recovery log buffer flush in order to establish synchronization between the recovery log buffer and the event log buffer.

[0042] For example, the recovery log buffer 220 (FIG. 2) determines that the recovery log buffer 220 is approaching capacity and that it is necessary to flush the recovery information held in the recovery log buffer 220 to the recovery log 225 on disk. The recovery log buffer 220 sends an LSN (flush LSN 240) to the event log program 300, which is used to flush the event log buffer 230. The flush LSN 240 sent corresponds to the same LSN having the high or highest sequence number of the recovery information currently held in the recovery log buffer 220 that is intended to be flushed. The flushing of the recovery information includes all LSNs below the flush LSN 240, and the event log program 300 uses the same flush LSN 240 to identify the event information to be flushed to the event log 235 on disk.

[0043] The event log program 300 identifies the LSN of the event information included in the event log buffer (step 340). The event log program 300 uses the LSN of the event information included in the event log buffer to establish the range of LSNs currently held in the event log buffer. The LSN of the unique instance of the event information added to the event log buffer is used to identify the set of event information corresponding to the individual event activities that are flushed to the event log existing on the disk.

[0044] The event log program 300 determines whether the event log buffer contains an LSN that is less than or equal to the first LSN (determination step 350). If the event log program 300 determines that the received LSN that is less than or equal to the first LSN does not currently exist in the event log buffer (step 350, "NO" branch), the event log program 300 sends a completion confirmation message and starts flushing the recovery information of the recovery log buffer to the recovery log existing on the disk (step 370).

[0045] For example, the event log program 300 receives a first LSN = 1000 indicating the flushing of event information having an LSN corresponding to LSN = 1000 or less. The event log program 300 determines that all the LSNs in the event log buffer exceed LSN = 1000 from the flushing actions associated with the event information that is close to the capacity of the event log buffer 230 and that occurred independently of the adjustment with the recovery log buffer 220. The event log program 300 sends a completion response 250 (Figure 2) and starts flushing the recovery log buffer 220.

[0046] When the event log program 300 sends a completion confirmation to start flushing the recovery log buffer, it terminates but remains active for detecting additional event activities executed on the node. In some embodiments of the present invention, the event log program 300 detects a failure condition that interrupts the delivery of event notifications and starts playing back event notifications from the event log stored in persistent storage from the last confirmed event information LSN. By providing for the playback of securely stored event notifications, the reliability of the delivery of event notifications to an internal sink, an external sink, or both is ensured.

[0047] If the event log program 300 determines that the event log buffer contains event information having an LSN less than or equal to the received first LSN (step 350, "YES" branch), the event log program 300 flushes the event information having an LSN less than or equal to the first LSN to the event log on the disk (step 360). The event log program 300 starts flushing a set of event information corresponding to each separate event activity executed on the node to the event log existing in persistent storage on the disk.

[0048] Upon executing the flush of the event information specified by the received first LSN or less, the event log program 300 sends a completion confirmation to start flushing the recovery log buffer (step 370) and proceeds as described above. In embodiments of the present invention, it is recognized that the recovery information and the event information flushed to the disk are adjusted by the specified LSN range defined by the event LSNs less than or equal to the first LSN. Further, in an embodiment, by flushing the event log buffer before flushing the recovery log buffer, a reliable and protected source of event notifications and corresponding metadata is ensured and is recognized as being available for playback upon detection of a failure condition or interruption of the delivery of event notifications from the node.

[0049] When the event log program 300 starts flushing the recovery log buffer and sends a completion confirmation, it ends, but remains active for detecting additional event activities executed on the node. In some embodiments of the present invention, the event log program 300 detects a failure condition that interrupts the delivery of event notifications, and starts playing back event notifications from the last confirmed event information LSN from the event log existing in the persistent storage. By providing playback of securely stored event notifications, the reliability in the delivery of event notifications to an internal sink or an external sink, or both, is guaranteed.

[0050] Figure 4 shows a block diagram of components of a computing system that includes a computing device 405, is configured to include or be operatively connected to the components depicted in FIG. 1, and has the ability to operatively execute the event log program 300 of FIG. 3, according to an embodiment of the present invention.

[0051] The computing device 405 includes components and functional capabilities similar to those of the first server node 110, the second server node 120, and the Nth server node 130 (FIG. 1) according to an exemplary embodiment of the present invention. It should be understood that FIG. 4 provides only an example of one implementation and does not imply any limitation with respect to the environments in which different embodiments may be implemented. Many modifications can be made to the depicted environment.

[0052] The computing device 405 includes a communication fabric 402 that provides communication among a computer processor 404, a memory 406, a persistent storage 408, a communication unit 410, and an input / output (I / O) interface 412. The communication fabric 402 can be implemented in any architecture designed to transfer data or control information, or both, among processors (such as microprocessors, communication devices, and network processors, etc.), system memory, peripheral devices, and any other hardware components within the system. For example, the communication fabric 402 can be implemented with one or more buses.

[0053] The memory 406, the cache memory 416, and the persistent storage 408 are computer-readable storage media. In this embodiment, the memory 406 includes a random access memory (RAM) 414. Generally, the memory 406 can include any suitable volatile or non-volatile computer-readable storage media.

[0054] In one embodiment, the event log program 300 is stored in the persistent storage 408 for execution by one or more of the respective computer processors 404 via one or more memories of the memory 406. In this embodiment, the persistent storage 408 includes a magnetic hard disk drive. Alternatively, or in addition to the magnetic hard disk drive, the persistent storage 408 can include a solid state hard drive, a semiconductor storage device, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or any other computer-readable storage media capable of storing program instructions or digital information. In one embodiment of the present invention, the recovery log 225 and the event log 235 are included in the persistent storage 408.

[0055] The media used by the persistent storage 408 can also be removable. For example, a removable hard drive can be used for the persistent storage 408. As another example, optical disks, magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer to another computer-readable storage medium that is part of the persistent storage 408 are included.

[0056] In these examples, the communication unit 410 provides communication with other data processing systems or devices, including the resources of the distributed data processing environment 100. In these examples, the communication unit 410 includes one or more network interface cards. The communication unit 410 may provide communication by using either or both physical and wireless communication links. The event log program 300 may be downloaded to the persistent storage 408 via the communication unit 410.

[0057] The I / O interface 412 enables the input and output of data with other devices that can be connected to the computing system 400. For example, the I / O interface 412 can provide connections to external devices 418 such as a keyboard, keypad, touch screen, or other suitable input device, or combinations thereof. The external devices 418 can also include portable computer-readable storage media such as thumb drives, portable optical or magnetic disks, and memory cards. Software and data used to implement embodiments of the present invention, such as the event log program 300, can be stored on such portable computer-readable storage media and loaded into the persistent storage 408 via the I / O interface 412. The I / O interface 412 also connects to a display 420.

[0058] The display 420 provides a mechanism for displaying data to a user and can be, for example, a computer monitor.

[0059] The programs described in this specification are identified, in certain embodiments of the invention, based on the applications for which they are implemented. However, the nomenclature of any particular program herein is used merely for convenience, and thus the invention should not be understood to be limited to use only in any particular application that is identified or implied by such nomenclature, or both.

[0060] The present invention can be a system, method, or computer program product, or a combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to implement aspects of the present invention.

[0061] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy (R) disks, punch cards, or mechanically encoded devices such as raised structures within grooves having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0062] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.

[0063] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. Execution of the computer-readable program instructions may be performed entirely on the user's computer as a stand-alone software package, partly on the user's computer, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer is connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can utilize the state information of the computer-readable program instructions to execute the computer-readable program instructions to personalize the electronic circuit in order to carry out aspects of the present invention.

[0064] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be realized by computer-readable program instructions.

[0065] Computer-readable program instructions provided to a processor of a computer or other programmable data processing apparatus can create means for causing the instructions executed via the processor of the computer or other programmable data processing apparatus to implement the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both, thereby generating a machine. These computer-readable program instructions may also be stored in a computer-readable storage medium having instructions stored therein that, when included in a product, cause a computer, a programmable data processing apparatus, or other devices that function in a particular manner, or combinations thereof, to implement the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both.

[0066] Computer-readable program instructions may also be loaded onto a computer, other programmable apparatus, or other device to generate a computer-implemented process for causing the instructions executed on the computer, other programmable apparatus, or other device to implement the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both, and thereby causing a series of operational steps to be performed on the computer, other programmable apparatus, or other device.

[0067] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be accomplished as one step, executed simultaneously, substantially simultaneously, in a partially or wholly temporally overlapping manner, or the blocks may be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagrams or flowchart diagrams, or both, and combinations of blocks in the block diagrams or flowchart diagrams, or both, can be implemented by a special purpose hardware-based system that performs the specified functions or operations, or implements a combination of special purpose hardware and computer instructions.

Claims

1. A method for fault tolerance in the distribution of event information within a node cluster of a computer, comprising: identifying, by one or more processors, event information associated with file system activity executed by a node of a node cluster of a computer; adding, by the one or more processors, the event information to an event log buffer in memory; receiving, by the one or more processors, a first log sequence number that matches a recovery log sequence number corresponding to flushing of recovery information from a recovery log buffer; identifying, by the one or more processors, the event information included in the event log buffer having a log sequence number less than or equal to the first log sequence number; in response to determining that the event information included in the event log buffer corresponds to a log sequence number less than or equal to the first log sequence number, flushing, by the one or more processors, the event information having the corresponding log sequence number from the event log buffer to disk storage A method comprising the above steps.

2. In response to determining that the event information included in the event log buffer does not correspond to a log sequence number less than or equal to the first log sequence number, further comprising: transmitting, by the one or more processors, an acknowledgement response to the recovery log buffer and starting flushing of the recovery information included in the recovery log buffer without performing flushing of the event information from the event log buffer. The method according to claim 1.

3. Further comprising: in response to detecting, by the one or more processors, a failure state of the node of the node cluster of the computer, transmitting, to a sink, the event information of at least one file system activity executed on the node flushed to disk storage. The method according to claim 1 or 2.

4. The method according to any one of claims 1 to 3, wherein the event information includes metadata corresponding to specific file system activities executed on the nodes of the node cluster of the computer.

5. The method according to any one of claims 1 to 4, wherein each node of the node cluster of the computer includes a producer and a message queue as a sink, and the sink is an event notification access repository external to the node cluster of the computer.

6. The method according to any one of claims 1 to 4, wherein each node of the node cluster of the computer includes a producer and an event log file on persistent storage as a sink, and the sink is an event notification access repository external to the node cluster of the computer.

7. The method according to any one of claims 1 to 6, wherein each event activity executed on the nodes of the node cluster of the computer receives a log sequence number that uniquely corresponds to the event information associated with the respective event activity.

8. Flushing the event information Flushing, by the one or more processors, the event information having a corresponding log sequence number from the event log buffer to an event log on disk storage Flushing, by the one or more processors, the recovery information to a recovery log on disk storage after the flushing of the event information having a corresponding log sequence number to the event log on disk storage The method according to any one of claims 1 to 7, further comprising

9. A computer program for fault tolerance in the delivery of event information of nodes within a computer node cluster, comprising: Program instructions for identifying event information associated with file system activities executed by nodes of a node cluster of a computer; Program instructions for adding the event information to an event log buffer in memory; Program instructions for receiving a first log sequence number that matches the recovery log sequence number, corresponding to flushing recovery information from the recovery log buffer, Program instructions for identifying the event information included in the event log buffer that has a log sequence number less than or equal to the first log sequence number, Program instructions for flushing the event information having the corresponding log sequence number from the event log buffer to disk storage in response to determining that the event information included in the event log buffer corresponds to a log sequence number less than or equal to the first log sequence number A computer program including.

10. Program instructions for transmitting an acknowledgment response to the recovery log buffer in response to determining that the event information included in the event log buffer does not correspond to a log sequence number less than or equal to the first log sequence number, and starting flushing of the recovery information included in the recovery log buffer without performing flushing of the event information from the event log buffer The computer program according to claim 9, further including.

11. Program instructions for transmitting the event information flushed to disk storage to a sink in response to detecting a failure state of the node in the node cluster of the computer The computer program according to claim 9, further including.

12. The computer program according to claim 9 or 10, wherein the event information includes metadata corresponding to specific file system activities executed on the node in the node cluster of the computer.

13. The computer program according to any one of claims 9 to 11, wherein each node in the node cluster of the computer includes a producer and a message queue as a sink.

14. The computer program according to any one of claims 9 to 11, wherein each node in the node cluster of the computer includes an event log file on persistent storage as a sink.

15. The computer program according to any one of claims 9 to 14, wherein each event activity executed on the nodes of the node cluster of the computer receives an LSN that uniquely corresponds to the event information associated with each event activity.

16. The event information having the corresponding log sequence number equal to or less than the first log sequence number is flushed from the event log buffer to an event log existing in persistent storage, and the event information flushed to the event log corresponds to the recovery information subsequently flushed from the recovery log buffer to a recovery log existing in persistent storage. The computer program according to any one of claims 9 to 15.

17. A computer system for fault tolerance in the distribution of event information of nodes within a computer node cluster, One or more computer processors, One or more computer-readable storage media, Program instructions stored in the one or more computer-readable storage media, and Including, the program instructions are Program instructions for identifying event information associated with file system activities executed by nodes of a node cluster of a computer, Program instructions for adding the event information to an event log buffer in memory, Program instructions for receiving a first log sequence number that matches a recovery log sequence number corresponding to flushing of recovery information from a recovery log buffer, Program instructions for identifying the event information included in the event log buffer having a log sequence number equal to or less than the first log sequence number, In response to determining that the event information included in the event log buffer corresponds to a log sequence number equal to or less than the first log sequence number, flushing the event information having the corresponding log sequence number from the event log buffer to disk storage. Program instructions and A computer system including.

18. Program instructions stored in the computer-readable storage medium for execution by at least one of the one or more processors, comprising: In response to determining that the event information included in the event log buffer does not correspond to a log sequence number less than or equal to the first log sequence number, sending an acknowledgement response to the recovery log buffer and starting to flush the recovery information included in the recovery log buffer without flushing the event information from the event log buffer, the program instructions The computer system according to claim 17, further comprising. **Claim 19** The computer system according to claim 17 or 18, wherein the event information includes metadata corresponding to specific file system activities executed on the node of the computer node cluster. **Claim 20** The computer system according to any one of claims 17 to 19, wherein the log sequence number (LSN) for the event log buffer, for each event activity executed on the node of the computer node cluster, matches the LSN associated with the recovery information in the recovery log buffer of the node for the same respective event activity. **Claim 21** The computer system according to any one of claims 17 to 20, wherein the event information having the corresponding log sequence number less than or equal to the first log sequence number is flushed from the event log buffer to an event log existing in persistent storage, and the event information flushed to the event log corresponds to the recovery information subsequently flushed from the recovery log buffer to a recovery log existing in persistent storage. **Claim 22** A method for reproducing undelivered event information due to a failure state of a node within a computer node cluster, comprising: Determining, by one or more processors, a failure state of a first node of a computer node cluster of a storage system Accessing, by the one or more processors, event information of an event executed on the first node of the computer node cluster of the storage system, the event information being stored in an event log existing in persistent storage and uniquely identified by a log sequence number (LSN) corresponding to an event activity executed on the node, wherein an event notification includes delivery of the event information of the event activity, the accessing; Identifying, by the one or more processors, an LSN of a last event notification delivered to a sink of the storage system, the sink being a receiving repository of a plurality of event notifications of the storage system, the identifying; Identifying, by the one or more processors, one or more event notifications having an LSN greater than the LSN of the last event notification delivered to the sink; Replaying, by the one or more processors, the one or more event notifications having an LSN greater than the LSN of the last event notification delivered to the sink from the event log to the sink of the storage system A method comprising.

23. The method according to claim 22, wherein the sink is an external event notification access repository of the computer node cluster.

24. A computer system for replaying undelivered event information due to a failure state of a node within a computer node cluster, comprising: One or more computer processors; One or more computer-readable storage media; Program instructions stored on the one or more computer-readable storage media, and wherein the program instructions include: Program instructions for determining a failure state of a first node of a computer node cluster of a storage system; Program instructions for accessing event information of an event executed on the first node of the computer node cluster of the storage system, which is stored in an event log existing in persistent storage, wherein the event information is uniquely identified by a log sequence number (LSN) corresponding to an event activity executed on the node, and an event notification includes delivery of the event information of the event activity, and the program instructions; Program instructions for identifying the LSN of the last event notification delivered to a sink of the storage system, wherein the sink is a reception repository for a plurality of event notifications of the storage system, and the program instructions; Program instructions for identifying one or more event notifications having an LSN greater than the LSN of the last event notification delivered to the sink; Program instructions for reproducing the one or more event notifications having an LSN greater than the LSN of the last event notification delivered to the sink from the event log to the sink of the storage system; A computer system comprising the above.

25. The computer system according to claim 24, wherein the sink is an external event notification access repository of the computer node cluster.

Citation Information

Patent Citations

  • Method of transaction processing and system

    JP1993006297A

  • fault tolerant computer system

    JP2002522845A

  • Index update pipeline

    JP2016524750A