Alarm event processing method, transmission system, processing device, medium and product

By dividing volatile and nonvolatile storage areas in the management controller, combining dual heartbeat packet detection and hashing algorithm to process alarm events, the problem of alarm event loss caused by network interruption is solved, efficient cache and recovery transmission of alarm events is achieved, and fault diagnosis capabilities of remote servers are improved.

CN120342836AActive Publication Date: 2025-07-18JINAN INSPUR DATA TECH CO LTD

Patent Information

Application Number
CN202510797154.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

When the network connection is interrupted, the alarm events generated by the management controller may be lost, resulting in a reduced fault diagnosis capability of the remote server.

Method used

The volatile storage area and the non-volatile storage area are divided in the management controller, and the untransmitted alarm events are stored, and they are transmitted to the remote server when the network is restored. The network status is detected through the dual heartbeat packet detection mechanism, and the hash algorithm is used to generate an alarm event fingerprint to ensure that data is not lost and repeated transmission.

Benefits of technology

It realizes permanent storage and timely recovery of alarm events during network interruption, improves the fault diagnosis capability of remote servers, reduces the risk of data loss and repeated transmission, and improves the reliability and efficiency of network connections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342836A_ABST
    Figure CN120342836A_ABST
Patent Text Reader

Abstract

The invention discloses an alarm event processing method, a transmission system, a processing device, a medium and a product, and relates to the technical field of servers. A management controller is divided into a volatile storage area and a non-volatile storage area which belong to the management controller in advance, so that an efficient caching strategy of a mixed storage area is realized, and the loss of an alarm event when a network is interrupted is completely eradicated. And when the network is recovered, obtaining the target alarm event from the memory area, and transmitting the target alarm event to a remote server. In the process, a returned alarm event fingerprint list needs to be received, so that repeated transmission is avoided through recording of a unique fingerprint identifier. If the target alarm event does not exist in the sending queue, the target alarm event needs to be sent again, so that missing sending of the target alarm event is prevented, recovery transmission of the alarm event under the condition of network recovery is achieved, it is guaranteed that the remote server can completely receive the alarm event, and the fault diagnosis capacity of the remote server is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of servers, and in particular to a method for processing alarm events, a transmission system, a processing device, a medium, and a product. Background Art

[0002] When the management controller monitors that the hardware status of the server is abnormal, it generates an alarm event and reports the alarm event to the remote alarm server in real time through the out-of-band management channel, so that the remote alarm server can perform alarm processing on the alarm event in the first time. If the network connection between the management controller and the remote alarm server is interrupted, the alarm events corresponding to the transmission process and the alarm events that are not transmitted and generated at the local end of the management controller will be lost. When the network connection is restored, the lost data cannot be retrieved, resulting in a reduction in the server's fault diagnosis ability.

[0003] Therefore, how to avoid the loss of alarm events when the network connection is interrupted is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for processing alarm events, a transmission system, a processing device, a medium, and a product, so as to solve the problem that the fault diagnosis ability of the server is reduced due to the loss of alarm events when the network connection is interrupted.

[0005] To solve the above technical problems, the present invention provides a method for processing alarm events, including: Pre-divide a volatile storage area and a non-volatile storage area for the management controller from the server memory area; When detecting a network interruption with the remote server, receive the untransmitted alarm events and store them in the memory areas corresponding to the volatile storage area and the non-volatile storage area; When the network is restored, obtain the target alarm events in the memory area and transmit them to the remote server to receive the returned alarm event fingerprint list; If the target alarm event fingerprint in the alarm event fingerprint list is not in the sending queue of the management controller, resend the target alarm event.

[0006] On the one hand, storing the untransmitted alarm events in the memory areas corresponding to the volatile storage area and the non-volatile storage area includes: Storing the untransmitted alarm events in the volatile storage area; When the storage space of the volatile storage area is less than the preset storage space, migrate the untransmitted alarm events that have been written to the volatile storage area to the non-volatile storage area.

[0007] On the other hand, migrating the untransmitted alarm events that have been written to the volatile storage area to the non-volatile storage area includes: Regarding the untransmitted alarm events that have been written to the volatile storage area as the first alarm events; Performing data compression processing on the first alarm events to obtain the compressed first alarm events; Migrating the compressed first alarm events to the non-volatile storage area.

[0008] On the other hand, the cache structure of the volatile storage area is a circular buffer structure; When writing to the volatile storage area, the untransmitted alarm events are sequentially written to the corresponding data block positions of the circular buffer structure.

[0009] On the other hand, regarding the untransmitted alarm events that have been written to the volatile storage area as the first alarm events includes: Regarding the alarm events within the data block positions written to the circular buffer structure as the second alarm events; Determining the priority levels of the second alarm events according to the urgency of the second alarm events; Selecting the second alarm events corresponding to the first preset priority level as the first alarm events.

[0010] On the other hand, obtaining the target alarm events in the memory area includes: Determining the current second alarm events corresponding to multiple priority levels stored in the current volatile storage area; Determining the current first alarm events corresponding to multiple priority levels stored in the current non-volatile storage area; Selecting the target alarm events based on the second preset priority level among the current first alarm events and the current second alarm events.

[0011] On the other hand, selecting the target alarm events based on the second preset priority level among the current first alarm events and the current second alarm events includes: Sorting the priority levels corresponding to the current first alarm events and the current second alarm events respectively; Generating a transmission queue for the sorted alarm events according to the second preset priority level; Reading in the transmission queue to obtain the target alarm events according to the storage areas of the sorted alarm events.

[0012] On the other hand, selecting the target alarm events based on the second preset priority level among the current first alarm events and the current second alarm events includes: Sort the priority levels corresponding to the current first alarm event and the current second alarm event respectively to obtain the first sorted alarm event; Determine the storage times corresponding to the current first alarm event and the current second alarm event respectively, and sort the current first alarm event and the current second alarm event according to the storage times to obtain the second sorted alarm event; Determine the final third alarm event according to the first sorted alarm event and the second sorted alarm event; Generate a sending queue for the third alarm event according to the second preset priority level; Read the target alarm event from the sending queue according to the storage area of the third alarm event.

[0013] On the other hand, when the target alarm event is obtained from the non-volatile storage area, it is transmitted to the remote server, including: Perform decompression processing on the target alarm event to obtain the decompressed target alarm event; Perform data block partitioning on the decompressed target alarm event to obtain partitioned data; Perform compression processing on the partitioned data to send it to the remote server; wherein, the compression processing method of the partitioned data is different from the data compression processing method of the first alarm event.

[0014] On the other hand, migrating the compressed first alarm event to the non-volatile storage area includes: Obtain the blocks corresponding to multiple sectors of the non-volatile storage area; Write the compressed first alarm event sequentially according to the blocks of multiple sectors to complete the migration.

[0015] On the other hand, the network detection process with the remote server includes: Detect the network status corresponding to the management controller and the remote server through the double heartbeat packet detection mechanism; When the network status is marked as disconnected, it is determined that the network is interrupted; When the network status is marked as restored, it is determined that the network is restored.

[0016] On the other hand, detecting the network status corresponding to the management controller and the remote server through the double heartbeat packet detection mechanism includes: Detect the network connection status between the management controller and the remote server through the first heartbeat packet detection mechanism; Detect the log service running status of the remote log service listening port of the remote server through the second heartbeat packet detection mechanism; If responses are received for both the network connection status and the log service running status, determine that the network status is marked as normal; If a response is not received for either the network connection status or the log service running status, determine that the network status is marked as disconnected.

[0017] On the other hand, the process of receiving responses for the network connection status and the log service running status includes: Obtain a first preset sending period and a first preset number of times; Send the first heartbeat packet detection mechanism and the second heartbeat packet detection mechanism according to the first preset sending period to obtain a first current sending number; If the first current sending number reaches the first preset number of times and no response is received, obtain a second preset sending period and a second preset number of times, where the second preset sending period is less than the first preset sending period; Send the first heartbeat packet detection mechanism and the second heartbeat packet detection mechanism according to the second preset sending period to obtain a second current sending number; If any one of the heartbeat packet detection mechanisms fails to be sent continuously in the second current sending number, record the number of consecutive sending failures; When the number of consecutive sending failures reaches the second preset number of times, determine that the network status is marked as disconnected; If the first current sending number reaches a third preset number of times and the number of consecutive successful sends reaches a fourth preset number of times, determine that the network status is marked as restored.

[0018] On the other hand, the first heartbeat packet detection mechanism is implemented by network layer protocol transmission detection; the second heartbeat packet detection mechanism is implemented by transmission control protocol transmission detection.

[0019] On the other hand, the process of determining the target alarm event fingerprint includes: Process the target alarm event according to the hash algorithm to obtain a corresponding event fingerprint, and encapsulate the event fingerprint and the target alarm event in a data structure.

[0020] On the other hand, resending the target alarm event includes: Increase the priority level of the target alarm event so that it reaches the highest priority level in the next sending queue for resending.

[0021] To solve the above technical problems, the present invention also provides a transmission system for alarm events, including a management controller and a remote server; the management controller and the remote server are connected through a network; The management controller is used to execute the steps of the above-mentioned method for processing alarm events to transmit the alarm events to the remote server.

[0022] To solve the above technical problems, the present invention also provides a device for processing alarm events, including: A memory for storing computer programs; A processor for implementing the steps of the method for processing alarm events as described above when executing the computer program.

[0023] To solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for processing alarm events as described above are implemented.

[0024] To solve the above technical problems, the present invention also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method for processing alarm events are implemented.

[0025] The beneficial effects of the present invention are as follows. On the one hand, a storage area belonging to itself is pre-divided for the management controller, and a secondary storage architecture is established in its own storage area: a volatile storage area and a non-volatile storage area. When the network is normal, alarm events are transmitted to the remote server in real time. When a network interruption is detected, the untransmitted alarm events are stored in the storage area exclusive to the management controller. If they are stored in the memory area of the volatile storage area, the read and write speed of the alarm events is increased, the access latency is reduced, and a short caching process is achieved, ensuring that alarm events are not lost in the case of a network interruption during the server power-on process. If they are stored in the memory area of the non-volatile storage area, the alarm events can be permanently not lost, realizing an efficient caching strategy for the hybrid storage area of the volatile storage area and the non-volatile storage area, and preventing the loss of alarm events during network interruptions. On the other hand, when the network is restored, the target alarm events are obtained from the memory area to be transmitted to the remote server. During this process, it is necessary to receive the returned list of alarm event fingerprints to avoid duplicate transmission by recording the unique fingerprint identification. At the same time, the target alarm event fingerprints in the received list of alarm event fingerprints are compared with the sending queue. If they are not in the sending queue, they need to be resent to prevent the omission of the target alarm events, realizing the resumed transmission of alarm events in the case of network restoration, ensuring that the remote server can completely receive the alarm events, and improving the fault diagnosis ability of the remote server.

[0026] Secondly, two mitigation storage mechanisms are first written to the volatile storage area, and the extremely high writing speed of the volatile storage area is utilized to quickly respond to data writing requests. When the storage space of the volatile storage area is less than the preset storage space, it is migrated to the non-volatile storage area. The alarm events are efficiently read from the volatile storage area to ensure the high utilization rate of the storage space of the volatile storage area while improving the efficiency of reading the alarm events. According to the current storage space usage, dynamic alarm event migration is realized to improve flexibility. When migrating to the non-volatile storage area, a compression method is used to reduce the amount of read data of the alarm event while improving the reading and writing speed. The untransmitted alarm events are sequentially written to the data block positions corresponding to the ring buffer structure, and low latency and continuous storage are achieved through the simple pointer movement of the read and write operations of the ring buffer, thereby improving the writing efficiency of the alarm events. The first alarm event is determined based on the priority level of the alarm event to ensure that the alarm events with high priority levels are persisted and prevented from being lost. The target alarm event is selected from the priority levels corresponding to all alarm events in each memory area based on the second preset priority level, so as to ensure that the alarm events with higher priority levels are transmitted first when the network is restored later, so as to ensure that the alarm events can be transmitted to the remote server for timely diagnosis. The second preset priority level is set based on the priority level of all alarm events, so that the generation process of the sending queue is simple and fast. Only the priority level determined by the urgency of the alarm event is read to achieve the transmission of the alarm event as soon as possible. Based on the urgency and storage time of the alarm event, the priority level of the alarm event can be determined more accurately. In order to improve the efficiency and accuracy of alarm management, it is ensured that key issues can be handled in time. During the transmission process, the data is compressed and transmitted in blocks to reduce the amount of data, improve the transmission efficiency, reduce the network load, and significantly improve the data transmission efficiency by compression, avoiding network congestion and delay.

[0027] In addition, the present invention also provides an alarm event transmission system, an alarm event processing device, a computer-readable storage medium, and a computer program product, which have the same beneficial effects as the above-mentioned alarm event processing method. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0029] Figure 1 A flowchart of a method for processing an alarm event provided by an embodiment of the present invention; Figure 2Flowchart of another method for processing alarm events provided by an embodiment of the present invention; Figure 3 Structural diagram of a transmission system for alarm events provided by an embodiment of the present invention; Figure 4 Structural diagram of a processing device for alarm events provided by an embodiment of the present invention; Figure 5 Structural diagram of a processing apparatus for alarm events provided by an embodiment of the present invention. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0031] The core of the present invention is to provide a method, a transmission system, a processing apparatus, a medium, and a product for processing alarm events to solve the problem that the loss of alarm events caused by network connection interruption reduces the server fault diagnosis ability.

[0032] To enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0033] A management controller, such as a Baseboard Management Controller (BMC), is a core component for server hardware management, responsible for monitoring the server hardware status (such as temperature, voltage, fan speed, etc.), and generating alarm events when abnormalities are detected. The traditional BMC reports alarm events to a remote alarm server (such as a Simple Network Management Protocol (SNMP) service, a Redfish (modern protocol for server management) server, a cloud management platform, etc.) in real time through an Out-of-Band channel. However, when the network connection is interrupted, alarm events may be lost due to being unable to be transmitted, resulting in operation and maintenance personnel being unable to comprehensively grasp the server status and affecting the efficiency of fault diagnosis and recovery. The method for processing alarm events provided by the present invention can solve the above technical problems.

[0034] The management controller can be set on a computer room server. Both the remote server and the computer room server are connected to the Internet, and the management controllers in the remote server and the computer room server are connected through network communication.

[0035] Figure 1 The flowchart of a method for processing alarm events provided by an embodiment of the present invention is as follows Figure 1 As shown, the method includes: S11: Pre-divide the volatile storage area and the non-volatile storage area for the management controller from the server memory area; S12: When detecting a network interruption with the remote server, receive the untransmitted alarm events and store them in the memory areas corresponding to the volatile storage area and the non-volatile storage area; S13: When the network is restored, obtain the target alarm events in the memory area and transmit them to the remote server to receive the returned list of alarm event fingerprints; S14: Determine whether the target alarm event fingerprint in the list of alarm event fingerprints is in the sending queue of the management controller; if not, proceed to step S15, if so, return to step S13 to obtain the next target alarm event; S15: If the target alarm event fingerprint in the list of alarm event fingerprints is not in the sending queue of the management controller, re-transmit the target alarm event.

[0036] Specifically, the server in step S11 is the server to which the management controller belongs, which is different from the remote server. Combining the above interaction scenarios between the management controllers in the remote server and the computer room server, the memory area of the server corresponds to the memory area of the computer room server. Regarding the management controller, in addition to storing the corresponding firmware and configuration data, an independent storage area can be used to store the alarm events of the present invention. Additionally, confirm whether the current management controller supports dividing the storage area from the server memory. If it can be divided, log in to the management interface of the management controller (usually through a Web interface or a command-line tool). In the management interface, find the option related to storage configuration, such as "storage partition" or "memory management", and configure the size and usage of the storage area according to the documentation of the management controller. Configure the storage space for the management controller in the server operating system. Tools provided by the operating system can be used (such as fdisk for Linux or disk management tools for Windows) to create an independent partition or logical volume and mount it to a path accessible by the management controller. Or, use logical volume management to create a logical volume for the management controller to use, that is, allocate an independent logical volume to the management controller.

[0037] Regarding the size of the storage area to be divided, there is no limitation, and it can be set according to the actual situation. It should be noted that in this embodiment, the storage area is divided into a hybrid storage area of a volatile storage area and a non-volatile storage area. Considering that volatile storage (such as Dynamic Random Access Memory (DRAM)) has extremely high read and write speeds and is suitable for frequent access and fast data processing. For example, in database transaction processing, storing hot data (frequently read and written records) in DRAM can significantly reduce data access latency. Non-volatile storage (such as Solid State Drive (SSD)), although slightly slower, can persistently store data. The hybrid storage system can reasonably allocate data between the two, making use of the high-speed characteristics of DRAM and ensuring data persistence.

[0038] In step S12, when a network interruption is detected, the detection mechanism here can check whether the network connection is normal through Ping, or monitor the interface status of the network using a network management tool, or monitor the status of network devices, servers, and applications through an open-source network monitoring tool. It can also be detected through a programming interface, and hardware devices such as the port status of network switches and routers can be used, or network monitoring devices can be used. When an interruption is detected, directly receive the untransmitted alarm events. Here, the untransmitted alarm events can be alarm events detected at the local end of the management controller that have not had time to be transmitted in the network, and also include alarm events being transmitted during network transmission. Store them in the memory areas corresponding to the volatile storage area and the non-volatile storage area.

[0039] Regarding storing in the memory areas corresponding to the two storage areas, it mainly involves different storage strategies. It can be stored in the volatile storage area first, considering that alarm events are likely to accumulate in an instant, and the write speed needs to be increased to achieve fast processing of alarm events through a high read and write speed. When the remaining storage space in the volatile storage area is small, the alarm events in the volatile storage area can be migrated to the non-volatile storage area.

[0040] Or, store in the volatile storage area first, sort the priority levels of each alarm event in this storage area, and then migrate the alarm events with higher priorities to the non-volatile storage area, etc. The setting of the hybrid storage area improves the performance of storing data and realizes an efficient caching strategy. Compared with the conventional technical solution where alarm events are lost during network interruption, in this embodiment, they are stored through a caching method to prevent loss.

[0041] At step S13, when the network is restored, the alarm events directly stored in the storage area are read and transmitted to the remote server. The detection of network restoration can be the same as or different from the detection mechanism in step S12 above, which is not limited here and can be set according to the actual situation. To obtain the target alarm event in the memory area, it should be noted that the acquisition of the alarm event can be read according to the storage sequence or set according to the urgency of the alarm event. Which specific memory area, the volatile storage area or the non-volatile storage area, to obtain from can be determined according to the priority level corresponding to the urgency of the alarm event. If it is stored in a certain memory area, it will be obtained from that memory area.

[0042] Regarding the storage process according to the urgency of the alarm event, considering that when storing and migrating to the non-volatile storage area, more urgent alarm events may be stored in the volatile storage area at any time. Therefore, at this time, all alarm events in the two memory areas need to be compared to find the one with the highest priority level and obtain and transmit it first.

[0043] In addition, not only looking at the priority level, but also combining the factor of the storage time length, so that the alarm events with a longer storage time and a higher priority level are transmitted to the remote server first.

[0044] After transmitting to the remote server, receive the returned alarm event fingerprint list. Considering the corresponding data restoration transmission mechanism in this embodiment, whether there will be missed transmission or duplicate transmission during the alarm time transmission process, the fingerprints corresponding to each alarm event in the alarm event fingerprint list are used for comparison and distinction. Ensure that the alarm events are not transmitted repeatedly and not missed.

[0045] In step S14, compare the target alarm event fingerprint with the fingerprint of the management controller's send queue. If they are the same, it means that it has been sent and transmitted. If they are different, it means that there is a miss at the local end and it needs to be resent. This resending can be placed at any position in the send queue, or it can be placed at the first position for priority sending, etc.

[0046] The beneficial effects of the embodiments of the present invention are as follows. On the one hand, the management controller pre-divides its own storage area and establishes a secondary storage architecture in its own storage area: a volatile storage area and a non-volatile storage area. Under normal network conditions, alarm events are transmitted to the remote server in real time. When a network interruption is detected, the untransmitted alarm events are stored in the storage area exclusive to the management controller. If they are stored in the memory area of the volatile storage area, the read and write speed of the alarm events is increased and the access latency is reduced, thus achieving a short caching process to ensure that alarm events are not lost in the case of a network interruption during the server power-on process. If they are stored in the memory area of the non-volatile storage area, the alarm events can be permanently stored without loss, realizing an efficient caching strategy for the hybrid storage area of the volatile storage area and the non-volatile storage area and preventing the loss of alarm events during network interruptions. On the other hand, when the network is restored, the target alarm events are obtained from the memory area for transmission to the remote server. During this process, it is necessary to receive the returned list of alarm event fingerprints to avoid duplicate transmission by recording the unique fingerprint identifiers. At the same time, the target alarm event fingerprints in the received list of alarm event fingerprints are compared with the sending queue. If they are not in the sending queue, they need to be resent to prevent the omission of the target alarm events, realizing the resumed transmission of alarm events in the case of network restoration and ensuring that the remote server can fully receive the alarm events to improve the fault diagnosis ability of the remote server.

[0047] In some embodiments, the network detection process with the remote server includes: Detecting the network status corresponding to the management controller and the remote server through a dual heartbeat packet detection mechanism; When the network status is marked as disconnected, it is determined that the network is interrupted; When the network status is marked as restored, it is determined that the network is restored.

[0048] Specifically, in this embodiment, a dual heartbeat packet detection mechanism is adopted to detect the network status between the management controller and the remote server. It can be detected through two-way heartbeat detection, or a heartbeat packet can be set for the network connection status between the management controller and the remote server, and a heartbeat packet can be set for the receiving port of the remote server. There is no limitation on the two methods.

[0049] If the former method is used for heartbeat detection, both the sender and the receiver regularly send heartbeat packets, which can effectively avoid misjudgment caused by the loss of unidirectional heartbeat packets. This mechanism can ensure that in a complex network environment, even if the heartbeat packets in one direction are lost, the heartbeat packets in the other direction can still provide accurate connection status information. The network status can be detected at a short time interval to quickly discover network connection interruptions or abnormalities. Once one end does not receive the heartbeat packet from the other end within the specified time, the fault detection mechanism can be triggered to take timely measures.

[0050] If the latter is used for heartbeat detection, the same network connection status can be determined through different dimensions or different detection points, and the accuracy of heartbeat detection can be improved through multi-dimensional detection.

[0051] If the network status is marked as disconnected, it is determined that the network is disconnected and the steps of step S12 are performed; if the network status is marked as restored, it is determined that the network is restored and the steps of step S13 are performed.

[0052] The setting of the dual heartbeat detection mechanism provided in this embodiment can quickly detect the network status while improving the accuracy of detection and judgment, and can also be applied to various network protocols and complex network environments.

[0053] In some embodiments, the network status corresponding to the management controller and the remote server is detected through the dual heartbeat packet detection mechanism, including: Detect the network connection status between the management controller and the remote server through the first heartbeat packet detection mechanism; Detect the running status of the log service of the remote log service listening port of the remote server through the second heartbeat packet detection mechanism; If responses are received for both the network connection status and the log service running status, it is determined that the network status is marked as normal; If no response is received for the network connection status or the log service running status, it is determined that the network status is marked as disconnected.

[0054] Specifically, the first heartbeat packet detects the network connection status between the two, and the second heartbeat packet detects the listening port of the remote server. If responses are received for both heartbeat packets, it indicates that the network status is marked as normal and all transmissions are normal. If a response is not received for one of the heartbeat packets, it indicates that the network status is marked as disconnected.

[0055] The first heartbeat packet detects the network connection status between the two, which can promptly discover problems such as network interruption, packet loss, or excessive delay. For example, if a response is not received for the heartbeat packet within the specified time, it can be determined that the network connection may have been interrupted.

[0056] The second heartbeat packet detects the listening port of the remote server. Even if the network connection is normal, a specific service of the remote server may stop listening due to a fault. Through the second heartbeat packet, the situation where the port does not respond can be promptly discovered, avoiding misjudgment of the network connection status.

[0057] The two heartbeat packets provided in this embodiment respectively correspond to detection mechanisms to comprehensively judge the network status mark, which can significantly improve the accuracy, rapidity, and reliability of network status detection. It is applicable to various complex network environments and application scenarios, can quickly discover faults, optimize resource utilization, enhance the fault tolerance of the system, and improve the user experience.

[0058] In some embodiments, the first heartbeat packet detection mechanism is implemented by network layer protocol transmission detection; the second heartbeat packet detection mechanism is implemented by Transmission Control Protocol (TCP) transmission detection.

[0059] Specifically, the protocols used by the two heartbeat packet detection mechanisms can be the same or different. To more accurately determine the network status flag, the protocols are set to be different.

[0060] The first heartbeat packet detection mechanism is implemented by network layer protocol transmission detection. This protocol can be the Internet Control Message Protocol (ICMP), which mainly detects the connectivity of the network link. It can quickly determine whether the network connection between the local device and the remote device is normal. It is suitable for quickly detecting network connectivity. ICMP detection does not depend on the upper-layer application protocol and can run independently of the application layer service to quickly discover network-level problems.

[0061] The second heartbeat packet detection mechanism is implemented by Transmission Control Protocol (TCP) transmission detection. This protocol can be the Transmission Control Protocol (TCP), which is mainly used to detect whether a specific port of the remote server is in a listening state. Even if the network link is normal, a certain service on the remote server may stop listening due to a fault. Through the TCP heartbeat packet, it can be detected whether the port responds, thereby determining whether the service is available. TCP is a connection-oriented protocol that can provide reliable data transmission and error detection mechanisms to ensure the sending and receiving of heartbeat packets.

[0062] The two heartbeat packet detection mechanisms provided in this embodiment use different protocols, which can quickly locate the fault point. Through the dual guarantees of the network layer and the transport layer, the troubleshooting time is reduced and the availability of the system is improved. At the same time, the network layer detection realizes low overhead, and the transport layer detection realizes on-demand detection, ensuring the optimization of resource utilization while ensuring the detection effect.

[0063] In some embodiments, the process of receiving responses for the network connection status and the log service running status includes: Obtain a first preset sending period and a first preset number of times; Send the first heartbeat packet detection mechanism and the second heartbeat packet detection mechanism according to the first preset sending period to obtain a first current sending number of times; If no response is received when the first current sending number of times reaches the first preset number of times, obtain a second preset sending period and a second preset number of times, where the second preset sending period is less than the first preset sending period; Send the first heartbeat packet detection mechanism and the second heartbeat packet detection mechanism according to the second preset sending period to obtain the second current sending count; If there is any heartbeat packet detection mechanism that fails to be sent continuously in the second current sending count, record the continuous failure count; When the continuous failure count reaches the second preset count, determine that the network status flag is disconnected; When the first current sending count reaches the third preset count and the continuous successful sending count reaches the fourth preset count, determine that the network status flag is restored.

[0064] Specifically, in the response receiving process, first send two heartbeat packet detection mechanisms according to the first preset sending period. The sending frequencies of the two heartbeat packet detection mechanisms can be the same or different. For the convenience of counting heartbeat packet responses, the sending frequencies of the two heartbeat packet detection mechanisms are the same. During the sending process, count the first current sending count. When the first current sending count reaches the first preset count and no response is received from either of the two heartbeat packet detection mechanisms, it indicates that there is a problem with the network during the detection process. To further determine whether it is accurate, continue to send according to the second preset sending period that is less than the first preset sending period to obtain the second current sending count. If there is any heartbeat packet detection mechanism that fails to be sent continuously, the continuous failure count needs to be recorded. When it reaches the second preset count, it is necessary to determine that the network status flag is disconnected. When the first current sending count reaches the third preset count and the continuous successful sending count reaches the fourth preset count, determine that the network status flag is restored.

[0065] Example: The network status monitoring module starts running after the BMC is started. It adopts a dual-heartbeat detection mechanism. One is ICMP detection, and the other is TCP port scanning of the remote log service. The detection logic is as follows: 1. Send an ICMP heartbeat packet to detect the network connection with the remote log server; 2. Send a TCP detection packet to the listening port of the remote log service to detect the running status of the log service; 3. The normal mode of the sending period of the above detection packet is 5s. If no response is received from the heartbeat packet continuously for 2 times, switch to the intensive mode, and the sending period is changed to 1s. If any detection fails, the failure count is incremented by one. If it fails continuously for 5 times, it is marked as network disconnection. After confirming the network disconnection, the detection period is restored from the intensive mode to the normal mode to avoid resource waste. When both detections are successful continuously for 10 times, mark the network status as network restoration.

[0066] The response process of the specific heartbeat detection mechanism provided in this embodiment is used to determine different network status flags, which improves the accuracy of detection. At the same time, through the recording process of the response mechanism, the accuracy of determining the network status flag is improved.

[0067] In some embodiments, storing untransmitted alarm events in memory areas corresponding to the volatile storage area and the non-volatile storage area includes: Storing untransmitted alarm events in the volatile storage area; When the storage space of the volatile storage area is less than the preset storage space, the untransmitted alarm events that have been written to the volatile storage area are migrated to the non-volatile storage area.

[0068] Specifically, the untransmitted alarm events are first stored in the volatile storage area. Considering the limited storage space of the volatile storage area, when the storage space of the volatile storage area is less than the preset storage space, it means that the utilization rate of the volatile storage area is relatively low. At this time, the alarm events that have been written to the volatile storage area need to be migrated to the non-volatile storage area to leave storage space for subsequent generated alarm events and ensure that the storage utilization rate of the volatile storage area always remains at a relatively high level.

[0069] Regarding migrating the alarm events that have been written to the volatile storage area to the non-volatile storage area, it can be a part or all of the currently written alarm events. For the migration of a corresponding part of the alarm events, which part can be determined by the storage time length stored in the volatile storage area or by the priority level. Regarding how to migrate, considering the relatively slow reading speed of the non-volatile storage area, it can be migrated by using a compression method or other methods, which is not limited herein.

[0070] The two mitigation storage mechanisms provided in this embodiment are first written to the volatile storage area. Utilizing the extremely high writing speed of the volatile storage area, it can quickly respond to data writing requests. When the storage space of the volatile storage area is less than the preset storage space, it is then migrated to the non-volatile storage area. Efficiently reading alarm events from the volatile storage area can ensure a high utilization rate of the storage space of the volatile storage area while improving the efficiency of reading alarm events. According to the current storage space usage situation, dynamic alarm event migration is realized, improving flexibility.

[0071] In some embodiments, migrating the untransmitted alarm events that have been written to the volatile storage area to the non-volatile storage area includes: Regarding the untransmitted alarm events that have been written to the volatile storage area as the first alarm events; Performing data compression processing on the first alarm events to obtain compressed first alarm events; Migrating the compressed first alarm events to the non-volatile storage area.

[0072] Specifically, considering that the read and write speed of the non-volatile storage area is relatively slow and can persistently store alarm events, when writing to the non-volatile storage area, a compression method for alarm events is adopted. Regarding the specific compression mechanism corresponding to the compression process here, this embodiment does not make any limitations and can be selected according to the actual situation.

[0073] The first alarm event can be part of the data that has been written to the volatile storage area or all of the data, and can be set according to the actual situation without limitation here.

[0074] When migrating to the non-volatile storage area provided in this embodiment, a compression method is adopted to reduce the amount of read data of alarm events while improving the read and write speed.

[0075] In some embodiments, the cache structure of the volatile storage area is a circular buffer structure; When writing to the volatile storage area, the untransmitted alarm events are sequentially written to the corresponding data block positions of the circular buffer structure.

[0076] Specifically, in the circular buffer structure, when a network interruption occurs, real-time generated alarm events are received, supporting high-concurrency writing and low-latency reading. The circular buffer structure is a first-in-first-out data structure. When the buffer is full, the conventional technical solution will overwrite the earliest old data with new alarm events to achieve cyclic utilization of the storage space. In this embodiment, when the buffer is about to be full, it will be migrated to the non-volatile storage area. When writing the untransmitted alarm events to the storage area of the circular buffer structure, the fixed size of the circular buffer structure is divided into data block positions and sequentially written to the corresponding data block positions.

[0077] In this embodiment, the untransmitted alarm events are sequentially written to the corresponding data block positions of the circular buffer structure, and through simple pointer movement of the read and write operations of the circular buffer, low latency and continuous storage are achieved, improving the writing efficiency of alarm events.

[0078] In some embodiments, taking the untransmitted alarm events that have been written to the volatile storage area as the first alarm events includes: Taking the alarm events within the data block positions written to the circular buffer structure as the second alarm events; Determining the priority level of the second alarm events according to the urgency of the second alarm events; Selecting the second alarm events corresponding to the first preset priority level as the first alarm events.

[0079] Specifically, in combination with the above embodiments, in this embodiment, some alarm events are selected for migration. Therefore, regarding how to select some alarm events, all the untransmitted alarm events written to the volatile storage area are used as the second alarm events, and the corresponding priority levels are determined according to the urgency of the alarm events. The determination of the urgency here can be based on a comprehensive value set according to the weights corresponding to various factors such as the severity, impact range, and potential consequences of the alarm events, and the corresponding priority levels are determined by comparing the sizes of the comprehensive values.

[0080] Select the second alarm events corresponding to the first preset priority level. Here, the first preset priority level can select the first few priority levels, and regarding the first few here, it can be set according to the actual situation.

[0081] The first alarm events provided in this embodiment are determined based on the priority levels of the alarm events to ensure the persistent storage of high-priority alarm events and avoid loss.

[0082] In some embodiments, obtaining target alarm events in the memory area includes: Determine the current second alarm events corresponding to multiple priority levels stored in the current volatile storage area; Determine the current first alarm events corresponding to multiple priority levels stored in the current non-volatile storage area; Select target alarm events based on the second preset priority level among the current first alarm events and the current second alarm events.

[0083] Specifically, obtaining target alarm events in the memory area. As mentioned in the above embodiments, the memory area is multiple (volatile storage area and non-volatile storage area). Regarding how to obtain them, it can be determined according to the priority levels. In the above embodiments, it is mentioned that alarm events with high priority levels have been migrated to the non-volatile storage area, but the migration occurs when the storage space corresponding to the volatile storage area is relatively small. The current storage process may involve newly received untransmitted alarm events stored in the volatile storage area, and the priority levels may be higher than those of the alarm events stored in the non-volatile storage area. Therefore, in this case, a comprehensive judgment is required.

[0084] Sort the priority levels corresponding to all alarm events in each memory area, and select target alarm events based on the second preset priority level.

[0085] The target alarm events are selected from the priority levels corresponding to all alarm events in each memory area based on the second preset priority level in this embodiment to ensure that alarm events with higher priority levels are transmitted first during subsequent network recovery, so as to ensure that the alarm events can be diagnosed in a timely manner when transmitted to the remote server.

[0086] In some embodiments, selecting a target alarm event from a current first alarm event and a current second alarm event based on a second preset priority level includes: Sort the priority levels corresponding to the current first alarm event and the current second alarm event respectively; Generating a sending queue for the sorted alarm events according to a second preset priority level; The target alarm event is obtained by reading from the storage area of the sorted alarm events in the sending queue.

[0087] In combination with the above embodiment, the second preset priority level is the priority level after the priority levels of all alarm events are sorted to generate a sending queue.

[0088] The second preset priority level provided in this embodiment is set based on the priority level of all alarm events, so that the generation process of the sending queue is simple and fast. Only the priority level determined by the urgency of the alarm event is read to achieve the fastest transmission of the alarm event.

[0089] In some other embodiments, selecting a target alarm event from the current first alarm event and the current second alarm event based on a second preset priority level includes: sorting the priority levels corresponding to the current first alarm event and the current second alarm event respectively to serve as the first sorted alarm event; Determine storage times corresponding to the current first alarm event and the current second alarm event, respectively, and sort the current first alarm event and the current second alarm event according to the storage times to serve as second sorted alarm events; Determine a final third alarm event according to the first sorted alarm events and the second sorted alarm events; generating a sending queue for the third alarm event according to the second preset priority level; The target alarm event is obtained by reading from the storage area of the third alarm event in the sending queue.

[0090] Specifically, on the basis of the priority level determined by the urgency of the alarm event, the storage time of the alarm event is also added to characterize the severity of the alarm event according to the urgency. The storage time characterizes the duration of the alarm event or the historical record time. A short-term alarm is an alarm event that has just occurred and lasts for a short time. A long-term alarm is an alarm event that lasts for a long time, which may indicate that the problem is more serious or difficult to solve.

[0091] In this embodiment, the final third alarm event is determined according to the first sorted alarm event and the second sorted alarm event; it can be comprehensively determined through a priority matrix or the corresponding comprehensive value can be determined by a weighting method. The second preset priority level is selected according to the comprehensive value to generate a sending queue.

[0092] Based on the urgency and storage time of the alarm event provided in this embodiment, the priority level of the alarm event can be determined more accurately. To improve the efficiency and accuracy of alarm management and ensure that key problems can be processed in a timely manner.

[0093] In some embodiments, when the target alarm event is obtained from the non-volatile storage area and transmitted to the remote server, it includes: The target alarm event is decompressed to obtain the decompressed target alarm event; The decompressed target alarm event is processed by data block chunking to obtain chunked data; The chunked data is compressed and sent to the remote server; wherein, the compression processing method of the chunked data is different from the data compression processing method of the first alarm event.

[0094] Specifically, in combination with the above embodiments, considering that the alarm event migrated to the non-volatile storage area is stored after compression processing, so in this process, the target alarm event needs to be decompressed, and the decompression processing method here is used in combination with the compression processing method in the above embodiments. The decompression algorithm is not limited and can be used according to the actual situation.

[0095] The decompressed target alarm event is processed by data block chunking to obtain chunked data. Here, data block chunking is to prevent a large number of alarm events from being transmitted at one time, resulting in a slow transmission process and possible omissions. Therefore, the number of alarm events sent in each batch is limited. The chunked data is compressed and sent to the remote server. The compression processing here is different from the data compression processing method of the above first alarm event. Considering the transmission factors on the link during the transmission process, other compression methods can be selected and compressed according to the actual situation.

[0096] During the transmission process, the target alarm event can be encrypted, and the payload part is encrypted using an encryption algorithm to improve the security of data transmission.

[0097] The transmission method provided in this embodiment compresses and then transmits in a chunked manner during the transmission process, reducing the data volume, improving the transmission efficiency, reducing the network load. Compression significantly improves the data transmission efficiency and avoids network congestion and delay.

[0098] In some embodiments, the process of determining the target alarm event fingerprint includes: Process the target alarm event according to the hash algorithm to obtain the corresponding event fingerprint, and encapsulate the event fingerprint and the target alarm event in a data structure.

[0099] Specifically, regarding the generation of the alarm event, based on the monitored hardware failure, add a nanosecond-level timestamp, event level, device unique identifier, event code, and a detailed description of the alarm event to generate structured event data.

[0100] The structured event data is as follows: { "timestamp": 1718000000123456, / / Nanosecond-level timestamp; "severity": "CRITICAL", / / Event level; "device_id": "SVR-Node-01", / / Device unique identifier; "event_code": 0x1A3B, / / IPMI standard event code; "payload": "CPU1 Temp=98℃" / / Detailed description of the alarm event; }

[0101] Generate a unique event fingerprint according to the hash algorithm, and encapsulate the event fingerprint and the target alarm event in a data structure. Its specific structure is as follows: typedef struct{ char fingerprint

[65] ; / / 64-bit SHA-256 string + terminator; uint8_t event_data

[1024] ; / / Original event content; uint8_t transmission_flag; / / 0 = not sent, 1 = sent pending confirmation, 2 = confirmed; }cached_event_t;

[0102] The target alarm event fingerprint provided by this embodiment is determined by the hash algorithm, and is used to verify whether the data is tampered with during transmission or storage, maintain data consistency verification, and prevent data tampering. The unique event fingerprint indicates that each alarm event has a unique fingerprint, which is convenient for tracking and management.

[0103] In some embodiments, migrating the compressed first alarm event to the non-volatile storage area includes: Obtain the blocks corresponding to multiple sectors of the non-volatile storage area; Write the compressed first warning event successively into blocks of multiple sectors to complete the migration.

[0104] Specifically, the non-volatile storage area is stored in units of sectors. One sector is usually 512 bytes or 4096 bytes (4KB). An application or an operating system accesses the storage device through a logical block address (LBA). The storage device internally uses a physical address to manage the actual storage location of data. The storage device (such as an SSD) internally maintains a mapping table that maps the logical address to the physical address. When writing data, the mapping table is updated to reflect the actual storage location of the data. The operating system or the application initiates a write request, specifying the data to be written and the target logical address. The storage device controller looks up the mapping table to determine the physical address corresponding to the target logical address. If the block where the target physical address is located is free, the data is directly written to the target page.

[0105] Considering that if writing is performed in one sector, when the compressed first warning event arrives and is written, it is only written in that sector, resulting in a relatively high wear level. Therefore, in order to balance to other sectors, in this embodiment, the migration is completed in such a rotating manner of writing successively into blocks of multiple sectors.

[0106] The migration process of writing the compressed first warning event successively into blocks of multiple sectors provided by this embodiment disperses its write operations to different physical blocks, achieving wear leveling and extending the service life.

[0107] In some embodiments, resending the target warning event includes: Increasing the priority level of the target warning event so as to resend it with the highest priority level in the next sending queue.

[0108] Specifically, if it is not in the sending queue, it needs to be resent. At this time, it is necessary to mark that it has been sent and waiting for confirmation, re-add the unconfirmed target warning event to the sending queue, increase the priority, and wait for the next retransmission. In addition, through the marking setting, breakpoint resumption is supported, and when the network is disconnected again, it can continue to transmit from the position where the last successful transmission occurred.

[0109] The data structure of the sending queue is as follows: typedef struct{ cached_event_t events; / / Array of event pointers; uint32_t queue_size; / / Current queue length; uint32_t next_index; / / Next position to be sent; }retransmission_queue_t。

[0110] In this embodiment, the priority level of the target alarm event is increased and resent in the next transmission queue with the highest priority level to avoid omission during the next transmission and improve the reliability of transmission.

[0111] Figure 2 As shown in the flowchart of another method for processing alarm events provided by an embodiment of the present invention, Figure 2 as shown, the steps include: S21: Generate an alarm event; S22: Detect whether the network is normal. If so, go to step S23; if not, go to step S24; S23: Upload the alarm event in real time; S24: Encode the alarm event; S25: Write the encoded alarm event into the circular buffer; S26: Determine whether the storage space of the circular buffer is less than the preset storage space. If not, go to step S27; if so, go to step S28; S27: Wait for the next alarm event; S28: Start the compression mechanism to compress the alarm event to obtain the compressed alarm event; S29: Migrate the compressed alarm event to the non-volatile storage area; S30: When the network is restored, sort the alarm events in the circular buffer corresponding to the non-volatile storage area and the volatile storage area by priority level, and select the target alarm event with the highest priority level; S31: Transmit the target alarm event to the remote server in blocks.

[0112] Furthermore, the present invention also provides a transmission system for alarm events, Figure 3 As shown in the structural diagram of a transmission system for alarm events provided by an embodiment of the present invention, Figure 3 as shown, it includes a management controller and a remote server; the management controller and the remote server are connected through a network; The management controller is used to execute the steps of the above-mentioned method for processing alarm events to transmit the alarm events to the remote server.

[0113] For the introduction of a transmission system for alarm events provided by the present invention, please refer to the above method embodiment. The present invention will not be elaborated here, and it has the same beneficial effects as the above-mentioned method for processing alarm events.

[0114] The above has described in detail each embodiment corresponding to the method for processing alarm events. On this basis, the present invention also discloses a device for processing alarm events corresponding to the above method. Figure 4 It is a structural diagram of a device for processing alarm events provided by an embodiment of the present invention. As Figure 4 shown, the device for processing alarm events includes: A partitioning module 11, configured to pre-partition a volatile storage area and a non-volatile storage area for a management controller from a server memory area; A receiving module 12, configured to receive untransmitted alarm events when detecting a network interruption with a remote server, and store them in memory areas corresponding to the volatile storage area and the non-volatile storage area; A transmitting module 13, configured to, when the network is restored, obtain target alarm events from the memory area and transmit them to the remote server to receive a returned list of alarm event fingerprints; A sending module 14, configured to re-send the target alarm event if the target alarm event fingerprint in the list of alarm event fingerprints is not in the sending queue of the management controller.

[0115] Since the embodiments of the device part correspond to the above embodiments, the embodiments of the device part are described with reference to the embodiments of the above method part and will not be elaborated here.

[0116] For the introduction of a device for processing alarm events provided by the present invention, please refer to the above method embodiments. The present invention will not elaborate here, and it has the same beneficial effects as the above method for processing alarm events.

[0117] Figure 5 It is a structural diagram of a device for processing alarm events provided by an embodiment of the present invention. As Figure 5 shown, the device includes: A memory 21, configured to store a computer program; A processor 22, configured to implement the steps of the method for processing alarm events when executing the computer program.

[0118] The device for processing alarm events provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer, etc.

[0119] Among them, the processor 22 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 22 may be implemented in at least one hardware form of a Digital Signal Processor (DSP), a Field-Programmable Gate Array (FPGA), or a Programmable Logic Array. The processor 22 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the Central Processing Unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 22 may be integrated with a Graphics Processing Unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 22 may further include an Artificial Intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.

[0120] The memory 21 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 21 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 21 is at least used to store the following computer program 211. After the computer program is loaded and executed by the processor 22, it can implement the relevant steps of the processing method of the alarm event disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 21 may also include an operating system 212 and data 213, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 212 may include Windows, Unix, Linux, etc. The data 213 may include, but is not limited to, the data involved in the processing method of the alarm event, etc.

[0121] In some embodiments, the alarm event processing device may further include a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27.

[0122] Those skilled in the art can understand that Figure 5 the structure shown in

[0123] does not constitute a limitation on the alarm event processing device, and may include more or fewer components than shown in the figure.

[0124] For the introduction of a processing device for alarm events provided by the present invention, please refer to the above method embodiments. The present invention will not repeat it here, and it has the same beneficial effects as the above method for processing alarm events.

[0125] Furthermore, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor 22, the steps of the method for processing alarm events as described above are implemented.

[0126] It can be understood that if the method in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0127] For the introduction of a computer-readable storage medium provided by the present invention, please refer to the above method embodiments. The present invention will not repeat it here, and it has the same beneficial effects as the above method for processing alarm events.

[0128] Further, the present invention also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method for processing alarm events are implemented.

[0129] For the introduction of a computer program product provided by the present invention, please refer to the above method embodiments. The present invention will not repeat it here, and it has the same beneficial effects as the above method for processing alarm events.

[0130] The above has introduced in detail a method for processing alarm events, a transmission system, a processing device, a medium and a product provided by the present invention. The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

[0131] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including an..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

Claims

1. A method for processing an alarm event, characterized in that Including: Pre-divide a volatile storage area and a non-volatile storage area for the management controller from the server memory area; When a network interruption with the remote server is detected, receive the untransmitted alarm events and store them in the memory areas corresponding to the volatile storage area and the non-volatile storage area; When the network is restored, obtain the target alarm events in the memory area and transmit them to the remote server to receive the returned list of alarm event fingerprints; If the target alarm event fingerprint in the list of alarm event fingerprints is not in the sending queue of the management controller, resend the target alarm event.

2. The method for processing an alarm event according to claim 1, wherein Storing the untransmitted alarm events in the memory areas corresponding to the volatile storage area and the non-volatile storage area includes: Storing the untransmitted alarm events in the volatile storage area; When the storage space of the volatile storage area is less than the preset storage space, migrate the untransmitted alarm events that have been written to the volatile storage area to the non-volatile storage area.

3. The method for processing an alarm event according to claim 2, wherein, Migrating the untransmitted alarm events that have been written to the volatile storage area to the non-volatile storage area includes: Regarding the untransmitted alarm events that have been written to the volatile storage area as the first alarm events; Performing data compression processing on the first alarm events to obtain the compressed first alarm events; Migrating the compressed first alarm events to the non-volatile storage area.

4. The method for processing an alarm event according to claim 2, wherein, The cache structure of the volatile storage area is a circular buffer structure; When writing to the volatile storage area, sequentially write the untransmitted alarm events to the corresponding data block positions of the circular buffer structure.

5. The method for processing an alarm event according to claim 4, wherein, Regarding the untransmitted alarm events that have been written to the volatile storage area as the first alarm events includes: Regarding the alarm events in the data block positions written to the circular buffer structure as the second alarm events; Determining the priority level of the second alarm events according to the urgency of the second alarm events; Selecting the second alarm events corresponding to the first preset priority level as the first alarm events.

6. The method for processing an alarm event according to claim 5, wherein, Obtaining the target alarm events in the memory area includes: Determining the current second alarm events corresponding to multiple priority levels stored in the current volatile storage area; Determining the current first alarm events corresponding to multiple priority levels stored in the current non-volatile storage area; Selecting the target alarm events based on the second preset priority level among the current first alarm events and the current second alarm events.

7. The method for processing an alarm event according to claim 6, wherein Selecting the target alarm events based on the second preset priority level among the current first alarm events and the current second alarm events includes: Sorting the priority levels corresponding to the current first alarm events and the current second alarm events respectively; Generating a sending queue for the sorted alarm events according to the second preset priority level; Reading the target alarm events from the sending queue according to the storage areas of the sorted alarm events.

8. The method for processing an alarm event according to claim 6, wherein Selecting the target alarm events based on the second preset priority level among the current first alarm events and the current second alarm events includes: Sort the priority levels corresponding to the current first warning event and the current second warning event respectively to obtain the first sorted warning event; Determine the storage times corresponding to the current first warning event and the current second warning event respectively, and sort the current first warning event and the current second warning event according to the storage times to obtain the second sorted warning event; Determine the final third warning event according to the first sorted warning event and the second sorted warning event; Generate a sending queue for the third warning event according to the second preset priority level; Read the target warning event from the sending queue according to the storage area of the third warning event.

9. The method for processing an alarm event according to claim 3, wherein, When the target warning event is obtained from the non-volatile storage area and transmitted to the remote server, it includes: Perform decompression processing on the target warning event to obtain the decompressed target warning event; Perform data block chunking on the decompressed target warning event to obtain chunked data; Perform compression processing on the chunked data to send it to the remote server; wherein, the compression processing method of the chunked data is different from the data compression processing method of the first warning event.

10. The method for processing an alarm event according to claim 3, wherein Migrate the compressed first warning event to the non-volatile storage area, including: Obtain the blocks corresponding to multiple sectors of the non-volatile storage area; Write the compressed first warning event sequentially according to the blocks of multiple sectors to complete the migration.

11. The method for processing an alarm event according to claim 1, wherein The network detection process with the remote server includes: Detect the network status corresponding to the management controller and the remote server through a dual heartbeat packet detection mechanism; When the network status is marked as disconnected, it is determined that the network is interrupted; When the network status is marked as restored, it is determined that the network is restored.

12. The method for processing an alarm event according to claim 11, wherein Detect the network status corresponding to the management controller and the remote server through a dual heartbeat packet detection mechanism, including: Detect the network connection status between the management controller and the remote server through the first heartbeat packet detection mechanism; Detect the log service running status of the remote log service listening port of the remote server through the second heartbeat packet detection mechanism; When responses are received for both the network connection status and the log service running status, it is determined that the network status is marked as normal; When no response is received for either the network connection status or the log service running status, it is determined that the network status is marked as disconnected.

13. The method for processing an alarm event according to claim 12, wherein The response receiving process for the network connection status and the log service running status includes: Obtain a first preset sending period and a first preset number of times; Send the first heartbeat packet detection mechanism and the second heartbeat packet detection mechanism according to the first preset sending period to obtain the first current sending number; When the first current sending number reaches the first preset number of times and no response is received, obtain a second preset sending period and a second preset number of times, where the second preset sending period is less than the first preset sending period; Send the first heartbeat packet detection mechanism and the second heartbeat packet detection mechanism according to the second preset sending period to obtain the second current sending number; If there is any heartbeat packet detection mechanism that fails to send continuously in the second current transmission count, record the number of consecutive transmission failures; When the number of consecutive transmission failures reaches the second preset count, determine that the network status flag is disconnected; When the first current transmission count reaches the third preset count and the number of consecutive successful transmissions reaches the fourth preset count, determine that the network status flag is restored.

14. The method for processing an alarm event according to claim 12, wherein The first heartbeat packet detection mechanism is implemented by network layer protocol transmission detection; the second heartbeat packet detection mechanism is implemented by transmission control protocol transmission detection.

15. The method for processing an alarm event according to claim 1, wherein The process of determining the target alarm event fingerprint includes: Process the target alarm event according to the hash algorithm to obtain the corresponding event fingerprint, and encapsulate the event fingerprint and the target alarm event in a data structure.

16. The method for processing an alarm event according to claim 1, wherein Resending the target alarm event includes: Increase the priority level of the target alarm event so that it reaches the highest priority level in the next transmission queue for resending.

17. A transmission system for warning events, characterized in that, It includes a management controller and a remote server; the management controller and the remote server are connected through a network; The management controller is used to execute the steps of the alarm event processing method described in any one of claims 1 to 16 above to transmit the alarm event to the remote server.

18. A processing device for an alarm event, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the alarm event processing method described in any one of claims 1 to 16 when executing the computer program.

19. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the alarm event processing method described in any one of claims 1 to 16 are implemented.

20. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the alarm event processing method described in any one of claims 1 to 16 are implemented.

Citation Information

Patent Citations

  • Alarm recovery message reporting method

    CN101159634A

  • Alarm event management system, method and device and storage medium

    CN110688280A

  • Cloud collaboration method and system for simplifying field network diagnosis of operation and maintenance personnel

    CN113485220A

  • Alarm event processing method and device and computer readable storage medium

    CN114443429A

  • Alarm grading method, device and equipment and readable storage medium

    CN117421188A

Cited By

  • Alarm control method and device of baseboard management controller

    CN121116762A

  • Alarm control method and device of baseboard management controller

    CN121116762B