Event message management in hyperconverged infrastructure environments
By using heartbeat messages to manage event message delivery status in a hyperconverged infrastructure environment, the response latency problem caused by event message traffic overload is resolved, thereby improving system performance and service quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DELL PROD LP
- Filing Date
- 2021-12-23
- Publication Date
- 2026-05-29
AI Technical Summary
In hyperconverged infrastructure environments, excessive event message traffic can cause delays in management resource response, impacting performance and quality of service.
Event messages are managed by sending heartbeat messages through a central controller. Nodes maintain event message transmission strategies based on heartbeat status, including active, restricted, pending, and invalid states, to achieve timely processing and storage of event messages.
Effectively manage event message traffic, ensure coordination and QoS levels between nodes, reduce latency, and improve system performance and service quality.
Smart Images

Figure CN116339902B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to system administration, and more specifically to the management of messages in a virtualized environment. Background Technology
[0002] As the value and uses of information continue to grow, individuals and businesses are seeking alternative ways to process and store information. One option available to users is an information processing system. Information processing systems typically process, compile, store, and / or communicate information or data for business, personal, or other purposes, allowing users to leverage the value of this information. Because technologies and information processing needs and requirements vary among different users or applications, information processing systems can also vary in terms of what information is processed, how it is processed, how much information is processed, stored, or communicated, and how quickly and efficiently it can be processed, stored, or communicated. Variations in information processing systems allow them to be general-purpose or configured for specific users or purposes (such as financial transaction processing, flight booking, corporate data storage, or global communications). Furthermore, information processing systems can include various hardware and software components configured to process, store, and communicate information, and may include one or more computer systems, data storage systems, and networking systems.
[0003] Standard hardware, including (as a non-limiting example) x86-based servers, is increasingly used in hyperconverged infrastructure (HCI) environments. For the purposes of this disclosure, HCI can be characterized as an information technology (IT) paradigm in which computing, storage, networking, and management functions are all implemented in virtualized nodes.
[0004] In an HCI environment, management resources can monitor the operational status of each node, at least in part, based on event messages sent by the nodes in response to various events or conditions. An HCI environment may encompass hundreds or thousands of servers, resulting in a potentially large volume of event messages from numerous sources. If the event message traffic approaches or exceeds the management resources' capacity to process each message with little or no noticeable latency, the management resources' response may slow down, and the environment's performance and / or quality of service parameters may be negatively impacted. Summary of the Invention
[0005] Based on the teachings disclosed herein, an information processing system and method for managing event messages addresses common problems associated with event message processing in distributed systems, wherein a central controller is configured to send heartbeat messages (also referred to herein as heartbeats) to multiple nodes indicating the central controller's message processing capabilities. Each node is configured to receive heartbeats from the central controller and maintain the event message delivery state of the central controller based on the heartbeats. When a node detects a reportable event, the node determines a reporting policy corresponding to the event message delivery state of the central controller and takes an event message action according to the reporting policy. The event message action may include sending the event message without delay or storing the event message for subsequent transmission. In at least some embodiments, the multiple nodes include each node managed by the central controller, and each heartbeat includes one or more UDP-compliant datagram packets multicast by the central controller to the managed nodes.
[0006] Some implementations implement a finite set of heartbeat types and a finite set of event message transmission states, wherein the type of each heartbeat is selected from the set of heartbeat types and each event message transmission state is selected from the set of event message transmission states. In at least one implementation, the heartbeat types include a normal heartbeat type and the event message transmission states include an active state, wherein each node is configured to assign the active state to the central controller in response to receiving a normal heartbeat from the central controller. Furthermore, a reporting policy corresponding to the active state may require each node or enable each node to report new events to the central controller immediately or without significant delay.
[0007] The heartbeat type may further include a flow control heartbeat type, wherein the central controller is configured to send a flow control heartbeat in response to detecting a message processing capability below a threshold. In embodiments that include and / or support flow control heartbeat types, the set of event message delivery states may include a restricted state associated with the flow control heartbeat type, and each node may transition the central controller's event message delivery state to the restricted state in response to receiving a flow control heartbeat from the central controller. A reporting policy corresponding to the restricted state may impose a minimum interval between event messages on one or more of the nodes, wherein the minimum interval may be explicitly indicated within the flow control heartbeat or otherwise included as part of the flow control heartbeat. The flow control heartbeat may include an indication of which nodes the heartbeat is intended for.
[0008] The heartbeat type may also include a paused heartbeat, and the central controller is configured to send a pause message before restarting the central controller. In some embodiments, the event message delivery state may include a pending state, and each node may be configured to transition the event message delivery state to the pending state in response to receiving a paused heartbeat, wherein the pending state prevents multiple nodes from sending reporting messages. In at least some of these embodiments, each node receiving a paused heartbeat records the identifier of the last reported message, and subsequently stores new event messages without reporting them to the control controller until the central controller transitions out of the pending state, such as by sending a normal heartbeat.
[0009] The heartbeat type may include a recovery heartbeat, and a node may be configured to change its event message delivery status from a restricted or pending state to an active state in response to receiving the recovery heartbeat. Any node with an event message delivery status of pending can respond to receiving the recovery heartbeat by sending a stored message that occurs after the last recorded message to the central controller.
[0010] The technical advantages of this disclosure will be apparent to those skilled in the art from the accompanying drawings, specification, and claims included herein. The objectives and advantages of the embodiments will be realized and obtained, at least by means of the elements, features, and combinations particularly pointed out in the claims.
[0011] It should be understood that the foregoing general description and the following detailed description are illustrative and explanatory, and do not limit the claims set forth in this disclosure. Attached Figure Description
[0012] A more complete understanding of the embodiments and advantages of the invention can be obtained from the following description taken in conjunction with the accompanying drawings, in which the same reference numerals indicate the same features, and in the drawings:
[0013] Figure 1 A block diagram of the HCI platform is shown;
[0014] Figure 2 A block diagram of an HCI node is shown;
[0015] Figure 3 A block diagram showing the resources for handling event messages;
[0016] Figure 4 This shows the event message transmission status and state transitions associated with the heartbeat;
[0017] Figure 5 A flowchart illustrating the event message management method is shown; and
[0018] Figure 6 A block diagram of an exemplary information processing system is shown. Detailed Implementation
[0019] By reference Figures 1 to 6 For the best understanding of the exemplary embodiments and their advantages, the same numbers are used to indicate the same and corresponding parts unless otherwise explicitly indicated.
[0020] For the purposes of this disclosure, an information processing system may include any tool or set of tools operable to calculate, classify, process, transmit, receive, retrieve, initiate, exchange, store, display, exhibit, detect, record, reproduce, dispose of, or utilize information, intelligence, or data of any form for commercial, scientific, control, entertainment, or other purposes. For example, an information processing system may be a personal computer, a personal digital assistant (PDA), a consumer electronic device, a network storage device, or any other suitable device, and may vary in size, shape, performance, functionality, and price. An information processing system may include memory, one or more processing resources (such as a central processing unit (“CPU”), a microcontroller, or hardware or software control logic. Additional components of an information processing system may include one or more storage devices, one or more communication ports for communicating with external devices, and various input / output (“I / O”) devices (such as a keyboard, mouse, and video display). An information processing system may also include one or more buses operable to transmit communication between various hardware components.
[0021] Additionally, the information processing system may include firmware for controlling and / or communicating with, for example, hard disk drives, network circuitry, memory devices, I / O devices, and other peripheral devices. For example, a management program and / or other components may include firmware. As used in this disclosure, firmware includes software embedded in information processing system components for performing predefined tasks. Firmware is typically stored in non-volatile memory or memory in which stored data is not lost upon power failure. In some embodiments, firmware associated with an information processing system component is stored in non-volatile memory accessible to one or more information processing system components. In similar or alternative embodiments, firmware associated with an information processing system component is stored in non-volatile memory dedicated to and including as a part of that component.
[0022] For the purposes of this disclosure, computer-readable media may include any tool or set of tools capable of retaining data and / or instructions for a period of time. Computer-readable media may include, but is not limited to: storage media such as direct access storage devices (e.g., hard disk drives or floppy disks), sequential access storage devices (e.g., magnetic tape drives), optical discs, CD-ROMs, DVDs, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and / or flash memory; and communication media such as wires, optical fibers, microwaves, radio waves, and other electromagnetic and / or optical carrier waves; and / or any combination of the foregoing.
[0023] For the purposes of this disclosure, information processing resources can be broadly defined as any component system, apparatus or device of an information processing system, including but not limited to processors, service processors, basic input / output systems (BIOS), buses, memory, I / O devices and / or interfaces, storage resources, network interfaces, motherboards and / or any other components and / or elements of the information processing system.
[0024] In the following description, details are illustrated by way of example to facilitate discussion of the disclosed subject matter. However, it will be apparent to those skilled in the art that the disclosed embodiments are exemplary and do not exhaustively represent all possible embodiments.
[0025] Throughout this disclosure, the hyphenated form of the reference numerals refers to a specific instance of an element, while the non-hyphenated form of the reference numerals refers to an element in general. Thus, for example, "device 12-1" refers to an instance of a device category, which can be collectively referred to as "device 12," and any one of them can be generally referred to as "device 12."
[0026] As used herein, when two or more elements are referred to as “coupled” to each other, this term indicates that such two or more elements are in electronic communication, mechanical connection, including thermal and fluid connection, thermal connection or mechanical connection, whether indirect or direct, with or without intermediate elements.
[0027] Before describing the public features for monitoring and managing event messages in a distributed computing environment, an exemplary HCI platform suitable for implementing these features is provided. Referring now to the accompanying drawings, Figure 1 and Figure 2 An exemplary information processing system 100 is shown. Figure 1 and Figure 2 The information processing system 100 shown includes a platform 101, which is communicatively coupled to a platform administrator 102. Figure 1The platform 101 shown is an HCI platform in which computing, storage, and networking resources are virtualized to provide software-defined information technology (IT) infrastructure. Administrator 102 can be any computing system with functions for overseeing the operation and maintenance of the hardware, software, and / or firmware components of the HCI platform 101. Platform administrator 102 can interact with the HCI platform 101 through requests to an application programming interface (API) (not explicitly depicted) and responses from the API. In such implementations, requests may relate to event messaging monitoring and event messaging state management as described below. Figure 1 The HCI platform 101 shown can be implemented as a data center and / or cloud computing resource or within the data center and / or cloud computing resource, characterized by software-defined integration and virtualization of various information processing resources, including but not limited to servers, storage, network resources, management resources, etc.
[0028] Figure 1 The HCI platform 101 shown includes one or more HCI clusters 106-1 to 106-N, which are communicatively coupled to each other and to the Platform Resource Monitor (PRM) 114. Figure 1 Each HCI cluster 106 shown encompasses a set of HCI nodes 110-1 to 110-M configured to share information processing resources. In some implementations, resource sharing may require virtualizing resources in each HCI node 110 to create a logical pool of those resources, which can then be provided as needed across all HCI nodes 110 in the HCI cluster 106. For example, when considering storage resources, one or more physical devices (e.g., hard disk drives (HDDs), solid-state drives (SSDs), etc.) representing local storage resources on each HCI node 110 can be virtualized to form a clustered distributed file system (DFS) 112. In at least some of these implementations, the clustered DFS 112 corresponds to a logical pool of storage capacity formed by some or all of the storage within the HCI cluster 106.
[0029] HCI cluster 106 and one or more HCI nodes 110 within the cluster may represent or correspond to one or more of the entire application or multiple microservices implementing the application. As an example, HCI cluster 106 may be dedicated to a specific microservice, where multiple HCI nodes 110 provide redundancy and support high availability. In another example, the HCI nodes 110 within HCI cluster 106 may include one or more nodes corresponding to each microservice associated with a specific application.
[0030] Figure 1The HCI cluster 106-1 shown also includes a cluster network device (CND) 108, which facilitates communication and / or information exchange between HCI nodes 110 of HCI cluster 106-1 and other clusters 106, PRM 114, and / or one or more external entities, including (as an example) a platform administrator 102. In at least some embodiments, the CND 108 is implemented as a physical device, examples of which include, but are not limited to, network switches, network routers, network gateways, bridges, or any combination thereof.
[0031] PRM 114 can be implemented using one or more servers, each of which may correspond to a physical server in a data center, a cloud-based virtual server, or a combination thereof. PRM 114 may be communicatively coupled to all HCI nodes 110 of all HCI clusters 106 across HCI platform 101 and to platform administrator 102. PRM 114 may include a Resource Utilization Monitoring (RUM) service or feature with the ability to monitor Resource Utilization Parameters (RUP) associated with HCI platform 101.
[0032] Figure 2 An exemplary HCI node 110 according to the disclosed subject matter is shown. The HCI node 110, implemented with physical appliances (e.g., servers (not shown)), enables a hyperconverged architecture, thereby integrating virtualization, compute, storage, and networking resources into a single solution. The HCI node 110 may include a resource utilization agent (RUA) 202, which is communicatively coupled to network resources 204, compute resources 206, and node controller 216. Figure 2 The node controller 216 shown is coupled to a hypervisor 208 that supports one or more virtual machines (VMs) 210-1 to 210-L, each of which is shown as having an operating system (OS) 214 and one or more applications 212. The node controller 216 shown is also coupled to a storage component (e.g., a Small Computer System Interface (SCSI) controller) including zero or more optional storage controllers 220, and storage resources (including storage components) 222.
[0033] In some implementations, the task of RUA 202 is to monitor the utilization of virtualization, compute, storage, and / or network resources on HCI node 110. Therefore, node RUA 202 may include the following functions: monitoring the utilization of network resources 204 to obtain network resource utilization parameters (RUP); monitoring the utilization of compute resources 206 to obtain compute RUP; monitoring the utilization of virtual machines 210 to obtain virtualization RUP; and monitoring the utilization of storage resources 222 to obtain storage RUP. RUA 202 may periodically provide some or all of the RUPs to the Environmental Resource Monitor (ERM) 226 through pull and / or push mechanisms.
[0034] The focus now shifts to public characteristics used for monitoring and managing event messages in distributed computing environments. Figure 3 Showing with Figure 1 The HCI platform 101 shown is used in conjunction with an exemplary event messaging resource 300. Figure 3 The event message transmission resource 300 shown includes one or more central controllers 301, two of which are in Figure 3 The event message transmission resource 300 shown is represented as central controller 1 (301-1) and central controller 2 (301-2). The event message transmission resource 300 also includes multiple nodes 310, two of which are located in… Figure 3 The nodes are shown as node 1 (310-1) and node K (310-K).
[0035] Figure 3 Each node 310 shown can correspond to Figure 1 An example of HCI node 110 is shown. In at least some embodiments, node 310 communicating with any particular central controller 301 may include all nodes that have established a management trust relationship with the particular central controller. A group of nodes 310 that have established a management trust relationship with the central controller 301 may be referred to as managed nodes, and the term management domain may be used herein to collectively refer to all such managed nodes. Therefore, Figure 3 The diagram shows a management domain 303 consisting of managed nodes of the central controller 1 (301-1), i.e., node 310 managed by the central controller 1 (301-1), such as the central controller 301-1.
[0036] Figure 3 Each central controller 301 depicted includes a Quality of Service (QoS) control resource 302 and an event listener 304, while each node 310 includes an Event Message (EM) controller 311 and a storage device 320 for storing event message states 321 corresponding to each central controller 301. Each central controller 301 may be implemented as follows: Figure 1The components of the HCI-based information processing system 100 shown are implemented therein. As an example, the central controller 301 can be implemented in the platform administrator 102 ( Figure 1 Platform Resource Monitor 114 Figure 1 ), Environmental Resources Monitor 226 ( Figure 2 It can be implemented within another suitable physical or virtual system, device, or resource.
[0037] In at least one implementation, QoS control resource 302 is configured to generate heartbeats 330 and broadcast the heartbeats 330 to each node 310 managed by central controller 301. Central controller 301 may generate heartbeats 330 to convey the central controller's event message handling capabilities. For illustration, QoS control resource 302 causes a first type of heartbeat to be generated when event message handling capabilities are relatively high, a second type of heartbeat to be generated when event message handling capabilities are relatively low, and generates heartbeats including the following... Figure 4 The heartbeat type described has zero or more other types of heartbeats.
[0038] Heartbeat listeners 312 in each event message controller 311 receive and process heartbeats 330 from one or more aspects of the management node 310 and from each central controller 301. In at least some implementations, each event message controller 311 maintains a set of event message states 321 for each central controller 301. These event message states 321 are... Figure 3 The event message state 321 of the central controller determines or influences how node 310 generates event messages 340 and sends them to the applicable central controller to report node events 316 that occur from time to time during node operation. For example, reportable events may include any changes to the configuration of node 310, including any changes to the hardware, software, and / or firmware of node 310 regarding any computing, storage, network, and / or management resources. In this way, each central controller 301 and the nodes 310 within the management domain 303 of central controller 301 coordinate the transmission of event messages at least in part based on the event message handling capabilities of central controller 301. The benefits of the event message delivery handling and management described herein include the ability to differentiate event message delivery policies among various nodes, thereby facilitating the ability to support different QoS levels for different nodes. Another benefit is the ability to detect and respond to changes in event message delivery traffic that will inevitably occur during operation.
[0039] Now for reference Figure 4 The state transition diagram shows that it can be generated by Figure 3 The event messaging resource 300 shown employs an exemplary event messaging strategy 400. Figure 4 The event message delivery strategy 400 shown is based on an implementation that employs four event message states maintained by each node, where each event message state may correspond to the message processing capability of the central controller, and four heartbeat types generated and sent by the central controller to signal the message processing capability of the central controller and to switch the event message delivery state in the applicable nodes according to the message delivery strategy 400.
[0040] Figure 4 The event message states shown in the state transition diagram include active state 401, pending state 402, restricted state 403, and invalid state 404. Figure 4 The heartbeats shown include normal heartbeat 411, paused heartbeat 412, flow-controlled heartbeat 413, and resumed heartbeat 414. Furthermore, Figure 4 The heartbeat timeout condition 415 is shown. A heartbeat timeout condition can occur whenever the time interval since the last heartbeat was generated by the central controller exceeds a specified timeout value. It will be readily understood by those skilled in the art that the use of four event message states and four heartbeat types is an implementation-specific design choice, and that other implementations may employ more, fewer, and / or different event message states and more, fewer, and / or different heartbeat types.
[0041] In at least one implementation, the central controller 301 may issue a flow control (F) heartbeat 413 whenever the event message processing capability of the central controller 301 exceeds or falls below a specified threshold. The message processing capability of the central controller 301 may be measured based on messages / second, maximum latency, or a combination of these and / or other parameters. The flow control heartbeat 413 may include an indication of a minimum interval parameter, wherein the value of the minimum interval parameter may indicate the minimum time interval required between consecutive messages sent from any given node. The flow control heartbeat 413 may also include or otherwise indicate a range parameter, which indicates one or more specific nodes 310 to which a restricted state applies. In this context, range may refer to the node 310 to which a restricted event message delivery state applies. The range characteristic of the flow control heartbeat may facilitate a differentiated level of QoS among nodes 310. As an example, a prioritized node 310 may be excluded from the range of the flow control heartbeat to allow the prioritized node 310 to remain active. Meanwhile, other nodes may transition to a restricted state 403, where event message reporting is subject to the previously referenced minimum interval.
[0042] Figure 4The transition from an active state 401 to a pending state 402 in response to a paused heartbeat 412 is also illustrated. A paused heartbeat may be generated by the central controller 301 in the event of a reset, system boot, or similar event anticipated prior to a planned interruption of the central controller, to implement configuration changes or perform some other administrative or maintenance task. In these embodiments, the pending state 402 may correspond to an event message reporting policy, in which the applicable node stores, rather than sends, a new message corresponding to a reportable event and occurring after the paused heartbeat is processed from node event 316, and the event message delivery state remains pending for a period of time. Figure 4 In the illustrated implementation, the central controller may remain in a pending state until: a normal heartbeat is received, in which case the central controller event message state may transition to active state 401; or a flow control heartbeat is received, in which case the event message transmission state transitions to restricted state 403. In at least some implementations, node 310 may respond to a paused heartbeat 412 by recording the identifier of the last event message processed and / or sent by the node. In some implementations, the event message is assigned a unique value that monotonically increases over time or is otherwise associated with said unique value. A paused heartbeat may also include a next heartbeat parameter indicating an estimate of when the central controller will return to an operational state. Each node 310 may use the value of the next heartbeat parameter to determine when to resume monitoring and processing of the heartbeat.
[0043] Figure 4 The state transition diagram illustrates an implementation in which the event message transmission state of the central controller 301 transitions to an invalid state 404 whenever node 310 fails to detect a heartbeat signal from the central controller for a duration exceeding a specified threshold (referred to herein as the timeout value).
[0044] Now for reference Figure 5 The flowchart illustrates a method for use in distributed computing environments (such as...) Figure 1 Method 500 for managing and monitoring event messages in the HCI environment shown. Method 500 in Figure 5 As shown in the diagram, the left side represents actions performed by the central controller, while the right side represents actions performed by one or more managed nodes.
[0045] The method shown begins with the central controller broadcasting an initial heartbeat (operation 502) to all managed nodes. In at least some implementations, the initial heartbeat transitions each of the managed nodes to the active event message state 401. Figure 4 The normal heartbeat (as shown in the diagram). The initial heartbeat and all subsequent heartbeats can be broadcast simultaneously to all managed nodes, for example, via UDP multicast.
[0046] Following the broadcast of the initial heartbeat, the central controller monitors (operation 504) its ability to load event messages and / or process pending event messages. In at least one implementation, the central controller may distinguish at least two event message handling capabilities, including a normal event message handling capability in which event messages are not processed immediately or without significant delay. In some implementations, the normal event message handling capability is determined based on QoS parameters that may indicate the maximum latency associated with event message handling.
[0047] Based on the event message handling capability determined by the central controller in operation 504, the central controller may send (operation 506) an appropriate heartbeat based on or influenced by the determined event message transmission handling capability.
[0048] like Figure 5 As shown on the right, each managed node can respond to receiving an initial heartbeat from the central controller by initializing the event message state of the central controller to an active state (operation 520). The managed node can then monitor (operation 522) any new heartbeats from the central controller. Upon receiving a new heartbeat, each managed node can update (operation 524) the event message state of the central controller based on the current event message state and heartbeat type, as described above regarding... Figure 3 and Figure 4 The discussion continues. When the managed node subsequently detects (operation 530) a reportable event, the managed node determines (operation 532) the event action based on the event message status of the central controller and the corresponding event message policy, such as... Figure 3 and Figure 4 As shown and as described above.
[0049] Any or all of the HCI components shown or described herein (including virtualization components and resources) can be used in... Figure 6 The information processing system 600 shown is instantiated. The information processing system shown includes one or more general-purpose processors or central processing units (CPUs) 601, which are communicatively coupled to memory resources 610 and input / output hubs 620, with various I / O resources and / or components communicatively coupled to the input / output hubs. Figure 6 The I / O resources explicitly described include a network interface 640, commonly referred to as a NIC (Network Interface Card), storage resources 630, and other I / O devices, components, or resources, including, as non-limiting examples, a keyboard, mouse, monitor, printer, speaker, microphone, etc. (Not included in...) Figure 6The document explicitly describes that some implementations of the information processing system 600 (including some server implementations) may include a baseboard management controller, which, among other features and services, provides out-of-band management resources that can be coupled to a management device. Similarly, although not explicitly stated... Figure 6 While explicitly described, at least some notebook, laptop, and / or tablet computer implementations of the information processing system 600 may include an embedded controller (EC) that provides some management functions, which may include at least some functions, features, or services provided by a baseboard management controller in some server implementations.
[0050] This disclosure covers all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments herein that will be understood by those skilled in the art. Similarly, where appropriate, the appended claims cover all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments herein that will be understood by those skilled in the art. Furthermore, references in the appended claims to a device or system or a component of a device or system adapted, arranged, capable, configured, enabled, operable, or operable to perform a particular function cover that device, system, or component, whether or not it or the particular function is activated, turned on, or unlocked, provided that the device, system, or component is so adapted, arranged, capable, configured, enabled, operable, or operable.
[0051] All examples and conditional language described herein are intended to assist the reader in understanding this disclosure and the educational purposes provided by the inventors to advance the concepts in the art, and should be construed as not being limited to such specific examples and conditions. Although embodiments of this disclosure have been described in detail, it should be understood that various changes, substitutions, and modifications may be made to embodiments of this disclosure without departing from the spirit and scope thereof.
Claims
1. An information processing system management method, wherein the method includes: The central controller is configured to send heartbeats to multiple nodes, wherein the heartbeats indicate the message processing capabilities of the central controller. as well as Configure each of the plurality of nodes as follows: Receive heartbeats from the central controller and maintain the event message transmission status of the central controller based on the heartbeats; as well as In response to the occurrence of a reportable event, a reporting policy corresponding to the event message transmission status is determined, and a message indicating the reportable event is sent to the central controller according to the reporting policy.
2. The method of claim 1, wherein: Each heartbeat consists of one or more packets conforming to the Uniform Datagram Protocol (UDP); The plurality of nodes includes a plurality of managed nodes, wherein each managed node includes a node managed by the central controller; as well as Configuring the central controller to send the heartbeat includes configuring the central controller to multicast one or more UDP-compliant packets to the plurality of managed nodes.
3. The method of claim 1, wherein: The type of each of the heartbeats is selected from a set of heartbeat types; The event message transmission status is selected from a set of event message transmission statuses; The set of heartbeat types includes a normal heartbeat, and the set of event message transmission states includes an active state, wherein the plurality of nodes are configured to assign the active state in response to receiving a normal heartbeat; as well as The reporting strategy corresponding to the activity state enables the multiple nodes to report new events without delay.
4. The method of claim 3, wherein: The set of heartbeat types includes flow control heartbeats, and the central controller is configured to send a flow control heartbeat in response to detecting a message processing capability below a threshold. The set of event message transmission states includes a restricted state, and the plurality of nodes are configured to change the event message transmission state to the restricted state in response to receiving a flow control heartbeat; The reporting strategy corresponding to the restricted state imposes a minimum interval between event messages.
5. The method of claim 4, wherein the flow-controlled heartbeat includes an indication of the minimum interval.
6. The method of claim 4, wherein the flow control heartbeat includes an indication of the node to which the flow control heartbeat is applied.
7. The method of claim 4, wherein: The set of heartbeat types includes paused heartbeats, and the central controller is configured to send heartbeats before the central controller restarts; The set of event message transmission states includes a pending state, and the plurality of nodes are configured to change the event message transmission state to the pending state in response to receiving a pause heartbeat; as well as The reporting policy corresponding to the pending state prevents the multiple nodes from sending report messages.
8. The method of claim 7, wherein each node receiving the paused heartbeat records an identifier of the last reported message, and stores new messages while the central controller remains in the pending state.
9. The method of claim 8, wherein The set of heartbeat types includes a restored heartbeat; and The plurality of nodes are configured to change the event message transmission status from the restricted state or the pending state to the active state in response to receiving a recovery heartbeat.
10. The method of claim 9, wherein, The plurality of nodes whose event message transmission status is pending are configured to respond to receiving the recovery heartbeat by sending a stored message that occurred after the last reported message to the central controller.
11. An information processing system, comprising: processor; A non-transitory storage device, communicatively coupled to the processor, and including processor-executable instructions, which, when executed, cause the information processing system to perform management operations, the management operations including: The central controller is configured to send heartbeats to multiple nodes, wherein the heartbeats indicate the message processing capabilities of the central controller; and Configure each of the plurality of nodes as follows: Receive heartbeats from the central controller and maintain the event message transmission status of the central controller based on the heartbeats; In response to the occurrence of a reportable event, a reporting policy corresponding to the event message transmission status is determined, and a message indicating the reportable event is sent to the central controller according to the reporting policy.
12. The information processing system as described in claim 11, wherein: Each heartbeat consists of one or more packets conforming to the Uniform Datagram Protocol (UDP); The plurality of nodes includes a plurality of managed nodes, wherein each managed node includes a node managed by the central controller; Configuring the central controller to send the heartbeat includes configuring the central controller to multicast one or more UDP-compliant packets to the plurality of managed nodes.
13. The information processing system as described in claim 11, wherein: The type of each of the heartbeats is selected from a set of heartbeat types; The event message transmission status is selected from a set of event message transmission statuses; The set of heartbeat types includes a normal heartbeat, and the set of event message transmission states includes an active state, wherein the plurality of nodes are configured to assign the active state in response to receiving a normal heartbeat; as well as The reporting strategy corresponding to the activity state enables the multiple nodes to report new events without delay.
14. The information processing system as described in claim 13, wherein: The set of heartbeat types includes flow control heartbeats, and the central controller is configured to send a flow control heartbeat in response to detecting a message processing capability below a threshold. The set of event message transmission states includes a restricted state, and the plurality of nodes are configured to change the event message transmission state to the restricted state in response to receiving a flow control heartbeat; The reporting strategy corresponding to the restricted state imposes a minimum interval between event messages.
15. The information processing system of claim 14, wherein the flow control heartbeat includes an indication of the minimum interval.
16. The information processing system of claim 14, wherein the flow control heartbeat includes an indication of the node to which the flow control heartbeat is applied.
17. The information processing system as described in claim 14, wherein... The set of heartbeat types includes paused heartbeats, and the central controller is configured to send heartbeats before the central controller restarts; The set of event message transmission states includes a pending state, and the plurality of nodes are configured to change the event message transmission state to the pending state in response to receiving a pause heartbeat; and The reporting policy corresponding to the pending state prevents the multiple nodes from sending report messages.
18. The information processing system of claim 17, wherein each node receiving the paused heartbeat records an identifier of the last reported message, and stores new messages while the central controller remains in the pending state.
19. The information processing system of claim 18, wherein... The set of heartbeat types includes a restored heartbeat; and The plurality of nodes are configured to change the event message transmission status from the restricted state or the pending state to the active state in response to receiving a recovery heartbeat.
20. The information processing system as described in claim 19, wherein, The plurality of nodes whose event message transmission status is pending are configured to respond to receiving the recovery heartbeat by sending a stored message that occurred after the last reported message to the central controller.