BMC alarm event processing method and device, equipment and medium

By creating a scheduled task using BMC to detect the status of the SNMP Trap server and sending cached alarm information when it recovers, the problem of alarm information loss caused by the unavailability of the SNMP Trap server is solved, ensuring system stability and reliability.

CN118740583BActive Publication Date: 2026-04-10INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2024-06-28
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The SNMP Trap server may fail to receive Trap fault messages due to network failures, server maintenance, or configuration errors, resulting in alarm information not being delivered in a timely manner and affecting system stability and reliability.

Method used

BMC creates a scheduled task to detect the status of the SNMP Trap server. When the server is detected to be back to normal, it sends any unsent alarm information to the SNMP Trap server. Through caching and queue management, it ensures the effective processing and delivery of important alarm information.

Benefits of technology

Even when the SNMP Trap server is unavailable, important alarm information can be effectively processed and transmitted, maintaining the stable operation of the system and reducing system downtime and data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118740583B_ABST
    Figure CN118740583B_ABST
Patent Text Reader

Abstract

The application provides a BMC alarm event processing method, device, equipment and medium, the method comprises the following steps: when the BMC determines that the SNMP Trap server exists running abnormity, a timing task is constructed; the timing task is used for detecting the running state of the SNMP Trap server; when the timing task is triggered, the running state of the SNMP Trap server is detected; when the running state of the SNMP Trap server is normal, alarm information is sent to the SNMP Trap server, so that even in the case that the SNMP Trap server is unavailable, important alarm information can be effectively processed and delivered, thereby maintaining the stable operation of the system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of server management, and particularly relates to a BMC management alarm event processing method and device, equipment and a medium. BACKGROUND

[0002] In a complex computing environment, various devices and applications in the system may face various failures and problems. In order to ensure the stability and reliability of the system, failure management and alarm notification become a crucial link. When a device or application in the system fails or has an abnormal situation, timely issuing an alarm and taking corresponding measures can effectively reduce system downtime and data loss.

[0003] SNMP (Simple Network Management Protocol) is a widely used protocol in network management, and the Trap (self-trap) message in it is used to report device state changes and failure events to the management system. However, even such a key communication mechanism has certain uncertainties and risks. The SNMP Trap server may not be able to receive the Trap failure message due to various reasons, such as network failure, server maintenance, configuration error, etc., which will cause the SNMP Trap server to be unable to receive the message. SUMMARY

[0004] In view of the problems in the prior art, the present application provides a BMC management alarm event processing method, device, equipment and medium.

[0005] The present application provides a BMC management alarm event processing method, comprising:

[0006] When the BMC determines that the SNMP Trap server has a running exception, a timing task is constructed; the timing task is used to detect the running state of the SNMP Trap server;

[0007] When it is determined that the timing task is triggered, the running state of the SNMP Trap server is detected;

[0008] When it is determined that the SNMP Trap server is running normally, alarm information is sent to the SNMP Trap server.

[0009] According to the BMC management alarm event processing method provided by the present application, before the alarm information is sent to the SNMP Trap server, the method further comprises:

[0010] Obtain the important types and cache times of each alarm information in the cache;

[0011] The alarm information of different importance types is classified;

[0012] The alarm information of each importance type is sorted based on the generation time of the alarm information, and a first sending queue corresponding to each importance type is determined;

[0013] The first sending queues are sorted based on the priority of the importance types, and a second sending queue is determined;

[0014] Correspondingly, when it is determined that the SNMP Trap server is running normally, the alarm information is sent to the SNMP Trap server based on the second sending queue.

[0015] According to the application, a BMC management alarm event processing method is provided, and the BMC determines that the SNMP Trap server is running abnormally, comprising:

[0016] Alarm information at the current time is obtained, and when the number of failures of sending the alarm information at the current time to the SNMP Trap server reaches a first number, it is determined that the SNMP Trap server is running abnormally; correspondingly, based on an interval time, a first timing task after the SNMP Trap server is running abnormally is constructed.

[0017] When the timing task is triggered, it is determined that the SNMP Trap server is running abnormally when the number of failures of sending a detection signal to the SNMP Trap server reaches a second number; correspondingly, based on the current time and the interval time, an nth timing task after the SNMP Trap server is running abnormally is constructed, wherein n is a positive number greater than 2.

[0018] According to the application, a BMC management alarm event processing method is provided, and the method further comprises:

[0019] A plurality of time length ranges are obtained, and each time length range is connected to each other;

[0020] A plurality of time change strategies are obtained, wherein the adjustment change rate of each time change strategy is different; the adjustment change rate represents the increase percentage of the value before adjustment and the value after adjustment;

[0021] Based on the order from small to large of the adjustment change rate of the time change strategy, a corresponding time change strategy is configured for each time length range, and a corresponding relationship between the time length range and the time change strategy is constructed;

[0022] The determination time of the SNMP Trap server running abnormally corresponding to the first timing task and the current time are recorded.

[0023] According to the determined time when the SNMP Trap server is abnormal and the current time, a duration of the abnormal running of the SNMP Trap server is calculated;

[0024] The duration of the abnormal running of the SNMP Trap server is matched with each time range, so that a time range corresponding to the duration of the abnormal running of the SNMP Trap server is obtained.

[0025] According to the correspondence between the time range and the time change strategy, a corresponding time change strategy is determined.

[0026] The interval time is changed based on the determined time change strategy, so that a new interval time is obtained.

[0027] When it is determined that the SNMP Trap server is abnormal, a timing task is constructed based on the new interval time and the current time.

[0028] According to the method for processing the alarm event of the BMC management provided by the application, in the operation process of the power-on of the BMC, the method further comprises:

[0029] The alarm information in the cache is synchronized to the Flash in the BMC.

[0030] The alarm information that is successfully sent is deleted in the cache, and the alarm information that is successfully sent is deleted in the Flash.

[0031] According to the method for processing the alarm event of the BMC management provided by the application, before the first timing task after the abnormal running of the SNMP Trap server is constructed, the method further comprises:

[0032] Each alarm information in a first historical time period is obtained.

[0033] Based on each alarm information in the first historical time period, a shortest interval duration of the same alarm information is obtained.

[0034] If the alarm information first appears in the first historical time period, the shortest interval duration is determined based on the generation time of the alarm information first appearing and the current time when the abnormal running of the SNMP Trap server is determined.

[0035] Based on the shortest interval duration of each alarm information, an interval time required for the first timing task after the abnormal running of the SNMP Trap server is constructed is determined.

[0036] According to the method for regulating and controlling the stratum grouting provided by the application, the method further comprises:

[0037] When the duration of the abnormal running of the SNMP Trap server reaches the alarm duration, a maintenance warning information of the SNMP Trap server is sent out.

[0038] The application further provides a BMC alarm event processing device, comprising:

[0039] A construction module is configured to construct a timing task when the SNMP Trap server is abnormal, and the timing task is configured to detect the running state of the SNMP Trap server.

[0040] A detection module is configured to detect the running state of the SNMP Trap server when the timing task is triggered, and a processing module is configured to send alarm information to the SNMP Trap server when the SNMP Trap server is normal.

[0041] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the above-mentioned any one of the BMC alarm event processing method when executing the program.

[0042] The application further provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the above-mentioned any one of the BMC alarm event processing method.

[0043] The application further provides a computer program product comprising a computer program, and the computer program is executable on a processor to implement the above-mentioned any one of the BMC alarm event processing method.

[0044] The application provides a BMC alarm event processing method, device, equipment and medium, which creates a timing task to detect the state of the SNMP Trap server through the BMC after the server is abnormal, and sends the unsent alarm event to the SNMP Trap server when detecting that the SNMP Trap server is normal, so as to ensure that important alarm information can be effectively processed and delivered even in the case that the SNMP Trap server is unavailable, thereby maintaining the stable operation of the system. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0046] Figure 1 is a flowchart of a BMC alarm event processing method provided by the present application.

[0047] Figure 2 is a specific flowchart of a BMC alarm event processing method provided by the present application.

[0048] Figure 3 is a structural diagram of a BMC alarm event processing device provided by the present application.

[0049] Figure 4 is a structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0051] The BMC alarm event processing method, device, equipment and medium of the present application will be described below. Figures 1-3

[0052] Figure 1 A flowchart of a BMC alarm event processing method provided by the present application is shown, referring to Figure 1 The method comprises the following steps:

[0053] Step 11, when the BMC determines that the SNMP Trap server has a running exception, a timing task is constructed; the timing task is used to detect the running state of the SNMP Trap server;

[0054] Step 12, when the timing task is triggered, the running state of the SNMP Trap server is detected;

[0055] Step 13, when the SNMP Trap server is running normally, the alarm information is sent to the SNMP Trap server.

[0056] It should be noted that for steps 11-13, in a complex computing environment, various devices and application programs in the system may face various faults and problems. In order to ensure the stability and reliability of the system, fault management and alarm notification become a crucial link. When the device or application program in the system has a fault or an abnormal situation, timely alarm and corresponding measures can effectively reduce system downtime and data loss. ​

[0057] SNMP (Simple Network Management Protocol) is a widely used protocol in network management, and the Trap (self-trapping) message is used to report device state changes and fault events to the management system. However, even such a key communication mechanism has certain uncertainties and risks. The SNMP Trap server may not be able to receive the Trap fault message for various reasons, such as network failure, server maintenance, configuration error, etc., which will cause the SNMP Trap server to fail to receive the message.

[0058] In order to provide a BMC management alarm event processing method, when the SNMP Trap server is abnormal, the alarm information cannot be sent to the SNMP Trap server at this time. For this purpose, the BMC (Baseboard Manager Controller, i.e. baseboard management controller) creates a timing task. The timing task is considered to start from the time when the SNMP Trap server is abnormal, and after a certain time, the preset task is started. After the task is executed, the running state of the SNMP Trap server after the exception is detected to determine whether the SNMP Trap server has recovered to normal before the detection is executed. When it is determined that the SNMP Trap server has recovered to normal, the alarm information is sent to the SNMP Trap server at this time.

[0059] The BMC management alarm event processing method provided by the present application creates a timing task to detect the state of the SNMP Trap server after the server is abnormal, and sends the unsent alarm event to the SNMP Trap server when it is detected that the SNMP Trap server has recovered to normal, so as to ensure that important alarm information can be effectively processed and delivered even in the case that the SNMP Trap server is unavailable, thereby maintaining the stable operation of the system.

[0060] In a further method of the above method, there may be more failures of the SNMP Trap server during the time when the SNMP Trap server cannot receive information due to abnormality, and the alarm information of these failures is recorded at this time. Each alarm information is divided into different importance and configured as a corresponding fault type. For example, extremely important is type A, important is type B, and relatively important is type C, etc.

[0061] In the present application, the alarm information is stored in the cache, and after the SNMP Trap server recovers to normal, the alarm information is obtained from the cache and sent to the SNMP Trap server.

[0062] Each alarm information has its own cache time when stored in the cache.

[0063] Different alarm information is divided into the same important type from the perspective of importance.

[0064] In the present application, when the SNMP Trap server recovers normally, the alarm information of different important types is sent according to the importance level. For the alarm information of the same important type, it is sent according to the occurrence time of the alarm information. When the SNMP Trap server is abnormal, the alarm information is stored in the cache once generated. That is, for the alarm information of the same important type, it is sent according to the cache time of the alarm information.

[0065] Based on the above description, the alarm information of each important type in the cache is sorted based on the generation time of the alarm information, and the first sending queue corresponding to each important type is determined. The first sending queue here is the arrangement of the alarm information in each important type.

[0066] Then, based on the priority of the important type, the first sending queue is sorted to determine the second sending queue. The second sending queue at this time is the arrangement of all alarm information in the cache.

[0067] Correspondingly, when the SNMP Trap server is determined to be normal, the alarm information is sent to the SNMP Trap server based on the second sending queue.

[0068] Referring to Figure 2 The complete processing procedure is for the above processing procedure. The specific procedure will not be described.

[0069] The further method of the present application is to distinguish the importance and cache time of the alarm information to be sent after the SNMP Trap server recovers normally, set the queue, and facilitate the sequential sending.

[0070] In the further method of the above method, the process of determining that the SNMP Trap server has an operation exception is mainly explained, and the specific process is as follows:

[0071] There are two cases of SNMP Trap server exception:

[0072] When there is no alarm information or all the alarm information in the cache is sent, the SNMP Trap server is normal at this time. Then, the SNMP Trap server is abnormal due to abnormal reasons. Therefore, it is necessary to determine whether the SNMP Trap server is abnormal in this case.

[0073] Obtain the alarm information at the current time, and send the alarm information to the SNMP Trap server. The SNMP Trap server supports a retry mechanism to ensure that the alarm information can be successfully sent. Therefore, when the number of failures in sending the alarm information at the current time to the SNMP Trap server reaches a first number of times, it can be determined that the SNMP Trap server has a running exception.

[0074] Correspondingly, based on the interval time, a first timing task after the running exception of the SNMP Trap server is constructed. The interval time here is the time length between the time when the exception occurs and the time when the task is triggered for execution.

[0075] When the timing task is triggered to detect whether the SNMP Trap server is restored, the detection result is that the SNMP Trap server is not restored. This is also the case of the SNMP Trap server exception. Therefore, after the timing task is triggered, the SNMP Trap server is detected to determine whether it is restored, and preconfigured detection information is sent to the SNMP Trap server. Here, no alarm information is sent to avoid the possibility of sending failure of the alarm information in the case where it is unknown whether the SNMP Trap server is restored, resulting in loss of the alarm information.

[0076] When the number of failures in sending the detection signal to the SNMP Trap server reaches a second number of times, it can be determined that the SNMP Trap server has a running exception. Correspondingly, based on the current time and the interval time, an nth timing task after the running exception of the SNMP Trap server is constructed, where n is a positive number greater than 2.

[0077] Further methods of the application divide different determination methods for the SNMP Trap server exception detection in different cases, ensure that the timing task is created in time, and improve the secure storage and sending of the alarm information.

[0078] In the further method of the above method, the application can also dynamically adjust the interval time of the timing task according to the detection result of whether the SNMP Trap server is restored. Specifically, if the SNMP Trap server is detected to be offline for a continuous period of time, the BMC will gradually increase the interval time of the timing task to reduce unnecessary occupation of BMC resources, thereby avoiding resource waste.

[0079] Obtain a plurality of time length ranges, and each time length range is connected to each other. For example, [5min, 10min], [11min, 30min], [31min, 60min], [1.1h, 2h], [2.1h, 5h], etc.

[0080] a plurality of time change strategies are acquired, wherein an adjustment change rate of each time change strategy is different; the adjustment change rate represents an increase percentage of a pre-adjustment value and a post-adjustment value. For example, the pre-adjustment value is 10 minutes, the post-adjustment value is 30 minutes, and the adjustment change rate is 200%.

[0081] Based on the adjustment change rate of the time change strategy in the order from small to large, a corresponding time change strategy is configured for each time range, and a corresponding relationship between the time range and the time change strategy is constructed. That is, the adjustment change rate of the first time range is configured with the smallest time change strategy, and the adjustment change rate of the subsequent time range is configured in turn.

[0082] The determination time and the current time corresponding to the abnormal running of the SNMP Trap server of the first timing task are recorded. For example, when the abnormal running of the server is determined, the first timing task is constructed. At this time, the time when the abnormal running of the server is determined is the determination time.

[0083] According to the determination time and the current time of the abnormal running of the SNMP Trap server, the duration of the abnormal running of the SNMP Trap server is calculated.

[0084] The duration of the abnormal running of the SNMP Trap server is matched with each time range to obtain the time range corresponding to the duration of the abnormal running of the SNMP Trap server.

[0085] Based on the corresponding relationship between the time range and the time change strategy, the corresponding time change strategy is determined.

[0086] Based on the determined time change strategy, the interval time is changed to obtain a new interval time.

[0087] At this time, when the abnormal running of the SNMP Trap server is determined, the nth timing task after the abnormal running of the SNMP Trap server is constructed based on the new interval time and the current time.

[0088] In the further method of the above method, the operation process of the BMC power-on synchronizes each alarm information in the cache to the Flash in the BMC.

[0089] When the BMC startup stage judges whether there is a failed cache queue in the flash, and there is a subscription of the SNMP Trap server in the BMC, a timing task is started to repeat the alarm information. Even after the BMC is powered on again, the un-sent alarm information can be preserved, and the risk of losing the alarm information is reduced.

[0090] In the sending process, the sent alarm information is deleted in the cache, and the sent alarm information is synchronized and deleted in the flash.

[0091] In the further method of the above method, each alarm information in a first historical time period is obtained before the first timing task after the SNMP Trap server running abnormality is constructed. Based on each alarm information in the first historical time period, a shortest interval duration of the same alarm information is obtained.

[0092] If the alarm information first appears in the first historical time period, the shortest interval duration is determined based on the generation time of the alarm information first appearing and the current time when the SNMP Trap server running abnormality is determined.

[0093] The interval time required for the first timing task after the SNMP Trap server running abnormality is constructed is determined based on the shortest interval duration of each alarm information.

[0094] It should be noted that when the alarm often appears, it indicates that the failure rate of the server is high, and the recovery time is uncertain, which is long or short. When the alarm does not often appear, it indicates that the failure rate of the server is low, and it is easier to recover. Therefore, based on the historical record to determine the initial interval time, the server usage can be more suitable, so that the sending of the alarm information can be more timely or the real usage of the server can be more timely discovered.

[0095] Based on the above description, when the duration of the SNMP Trap server running abnormality is determined to reach the alarm duration, a maintenance warning information of the SNMP Trap server is sent.

[0096] In the further method of the above method, the construction of the time change strategy is mainly explained, and specifically as follows: for the setting of the initially determined interval time, the alarm information based on the historical record can be analyzed and set. For the adjustment of the subsequent interval time, the alarm type of the alarm information cached in the memory can be judged. If the alarm type indicates that there is more serious alarm information in the alarm information cached in the memory, such alarm information needs to be learned in time after the server is recovered. Therefore, when the second timing task after the server abnormality is constructed, a shorter adjustment value needs to be added to the initially determined interval time to obtain the interval time suitable for the second timing task. Therefore, the interval time of the subsequent timing task is determined in the form of an arithmetic sequence based on the initially determined adjustment value. At this time, the form of an arithmetic sequence can be used as a time change strategy. When the duration is adjusted for many times and is in another time range, another time change strategy can be replaced.

[0097] For another time change strategy, it can be explained that new alarm information can continue to be cached in the cache during the abnormal running of the server. For this purpose, a time change strategy in one of the time ranges is designed based on the duration and the number of newly cached alarm information in the cache. For example, the duration is divided into n parts, and then m minutes are added for each newly added alarm information. At this time, the adjustment value is the duration / n+m. At this time, the above calculation method can be used as the time change strategy. When the duration is in another time range after adjustment for several times, another time change strategy can be replaced.

[0098] For the time change strategy again, it can be explained that when the duration is long and exceeds a certain value. The subsequent time range can be a huge value. At this time, the adjustment value of the time change strategy can be calculated in hours, days, weeks, etc.

[0099] For more time ranges, the difference value of the arithmetic sequence can be increased. The time intervals of the timing tasks in a time range are calculated by using the arithmetic sequence, and only the difference value is different. Similarly, the n value of "the duration is divided into n parts" and the m value of "the newly added alarm information is increased by m minutes" can be adjusted, and the time change strategy can be changed to adapt to multiple time ranges.

[0100] The further method of the application can well construct a suitable timing task by constructing the time change strategy in different ways, which can detect in time after the server is recovered in a short time, and can remind to check and maintain in time when the server cannot be recovered in a long time.

[0101] The BMC alarm event processing device provided by the application is described below. The BMC alarm event processing device described below can be correspondingly referred to the BMC alarm event processing method described above.

[0102] Figure 3 A flowchart of a BMC alarm event processing device provided by the application is shown, referring to Figure 3 The device comprises a construction module 31, a detection module 32 and a processing module 33.

[0103] The construction module 31 is used to construct a timing task when it is determined that the SNMP Trap server exists running abnormity. The timing task is used to detect the running state of the SNMP Trap server.

[0104] The detection module 32 is used to detect the running state of the SNMP Trap server when it is determined that the timing task is triggered.

[0105] The processing module 33 is configured to send the alarm information to the SNMP Trap server when the SNMP Trap server is running normally.

[0106] In the further apparatus, before sending the alarm information to the SNMP Trap server, the processing module is further configured to:

[0107] obtain the importance type and the cache time of each alarm information in the cache;

[0108] divide the alarm information in different importance types;

[0109] sort the alarm information in each importance type based on the generation time of the alarm information, and determine a first sending queue corresponding to each importance type;

[0110] sort the first sending queues based on the priority of the importance type, and determine a second sending queue;

[0111] Correspondingly, the alarm information is sent to the SNMP Trap server based on the second sending queue when the SNMP Trap server is running normally.

[0112] In the further apparatus, the construction module is configured to, in the process of determining that the SNMP Trap server is running abnormally:

[0113] obtain the alarm information at the current time, and determine that the SNMP Trap server is running abnormally when the number of failures of sending the alarm information at the current time to the SNMP Trap server reaches a first number; and correspondingly, based on the interval time, construct a first timing task after the SNMP Trap server is running abnormally.

[0114] determine that the SNMP Trap server is running abnormally when the number of failures of sending the detection signal to the SNMP Trap server reaches a second number when the timing task is triggered; and correspondingly, based on the current time and the interval time, construct an nth timing task after the SNMP Trap server is running abnormally, wherein n is a positive number greater than 2.

[0115] In the further apparatus, after constructing the first timing task, the construction module is further configured to:

[0116] obtain a plurality of time ranges, and each time range is connected to each other;

[0117] obtain a plurality of time change strategies, wherein each time change strategy has a different adjustment change rate; the adjustment change rate represents the increase percentage of the value before adjustment and the value after adjustment;

[0118] The adjustment change rate of the time change strategy is arranged from small to large, a corresponding time change strategy is configured for each time range, and a corresponding relationship between the time range and the time change strategy is constructed;

[0119] A determination time corresponding to the first timing task and a current time are recorded when the SNMP Trap server runs abnormally;

[0120] The duration of the SNMP Trap server running abnormally is calculated according to the determination time and the current time of the SNMP Trap server running abnormally;

[0121] The duration of the SNMP Trap server running abnormally is matched with each time range to obtain a time range corresponding to the duration of the SNMP Trap server running abnormally;

[0122] The corresponding time change strategy is determined based on the corresponding relationship between the time range and the time change strategy;

[0123] The interval time is changed based on the determined time change strategy to obtain a new interval time;

[0124] When it is determined that the SNMP Trap server is abnormal, the timing task is constructed based on the new interval time and the current time.

[0125] In further devices of the above device, during the power-on operation of the BMC, the device further includes a synchronization module for:

[0126] Synchronize each alarm information in the cache to the Flash in the BMC;

[0127] Delete the alarm information that has been successfully sent in the cache, and synchronize to delete the alarm information that has been successfully sent in the Flash.

[0128] In further devices of the above device, before constructing the first timing task after the SNMP Trap server running abnormally, the construction module is further used for:

[0129] Obtain each alarm information in a first historical time period;

[0130] Based on each alarm information in the first historical time period, obtain the shortest interval duration of the same alarm information;

[0131] If the alarm information first appears in the first historical time period, the shortest interval duration is determined based on the generation time of the first appearing alarm information and the current time when the SNMP Trap server is determined to be abnormal;

[0132] The interval time required for constructing the first timing task after the abnormal running of the SNMP Trap server is determined based on the shortest interval length of each alarm information.

[0133] In the further device of the above device, the processing module is further used for:

[0134] When the duration of the abnormal running of the SNMP Trap server reaches the alarm duration, a maintenance warning information of the SNMP Trap server is sent.

[0135] The processing device for BMC management of alarm events provided by the application creates a timing task to detect the state of the SNMP Trap server after the abnormal running of the server, and sends the unsent alarm events to the SNMP Trap server when the SNMP Trap server is detected to be normal, so as to ensure that important alarm information can be effectively processed and delivered even in the case that the SNMP Trap server is unavailable, thereby maintaining the stable running of the system.

[0136] Figure 4 An example of a schematic diagram of the physical structure of an electronic device is shown in Figure 4 The electronic device can include a processor 41, a communications interface 42, a memory 43 and a communications bus 44, wherein the processor 41, the communications interface 42 and the memory 43 complete the communication with each other through the communications bus 44. The processor 41 can call the logical instructions in the memory 43 to execute the processing method for BMC management of alarm events, which includes: when the SNMP Trap server exists abnormal running, constructing a timing task; the timing task is used to detect the running state of the SNMP Trap server; when the timing task is triggered, detecting the running state of the SNMP Trap server; when the SNMP Trap server is normal, sending the alarm information to the SNMP Trap server.

[0137] In addition, the logic instructions in the memory 43 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0138] In another aspect, the present application also provides a computer program product, the computer program product comprising a computer program, the computer program being stored on a non-transitory computer readable storage medium, and the computer program being executable by a processor to cause a computer to perform the method of managing an alarm event of a BMC provided by the above-mentioned methods, the method comprising: determining that the SNMP Trap server has a running exception, constructing a timing task; the timing task is used to detect the running state of the SNMP Trap server; determining that the timing task is triggered, detecting the running state of the SNMP Trap server; determining that the SNMP Trap server is running normally, sending alarm information to the SNMP Trap server.

[0139] In another aspect, the present application also provides a computer program product, the computer program product comprising a computer program, the computer program being stored on a non-transitory computer readable storage medium, and the computer program being executable by a processor to cause a computer to perform the method of managing an alarm event of a BMC provided by the above-mentioned methods, the method comprising: determining that the SNMP Trap server has a running exception, constructing a timing task; the timing task is used to detect the running state of the SNMP Trap server; determining that the timing task is triggered, detecting the running state of the SNMP Trap server; determining that the SNMP Trap server is running normally, sending alarm information to the SNMP Trap server.

[0140] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0142] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A BMC management alarm event processing method, characterized in that, The method comprises the following steps: The BMC determines that the SNMP Trap server is abnormal, and a timing task is constructed; The timing task is used to detect the running state of the SNMP Trap server; When the timing task is triggered, the running state of the SNMP Trap server is detected; When the SNMP Trap server is normal, alarm information is sent to the SNMP Trap server; The BMC determines that the SNMP Trap server is abnormal, and the method comprises the following steps: Alarm information at the current time is obtained, and when the number of failures of sending the alarm information at the current time to the SNMP Trap server reaches a first number, it is determined that the SNMP Trap server is abnormal; Correspondingly, based on the interval time, a first timing task after the SNMP Trap server is abnormal is constructed; When the first timing task is triggered, it is determined that the number of failures of sending a detection signal to the SNMP Trap server reaches a second number, and it is determined that the SNMP Trap server is abnormal; Correspondingly, based on the current time and the interval time, an nth timing task after the SNMP Trap server is abnormal is constructed, wherein n is an integer greater than 2; After the first timing task is constructed, the method further comprises the following steps: A plurality of time ranges are obtained, and each time range is connected to each other; A plurality of time change strategies are obtained, wherein the adjustment change rate of each time change strategy is different; the adjustment change rate represents the increase percentage of the value before adjustment and the value after adjustment; Based on the order of the adjustment change rate of the time change strategy from small to large, each time range is configured with a corresponding time change strategy, and a corresponding relationship between the time range and the time change strategy is constructed; The determination time of the abnormal SNMP Trap server and the current time corresponding to the first timing task are recorded; The duration of the abnormal SNMP Trap server is calculated according to the determination time of the abnormal SNMP Trap server and the current time; The duration of the abnormal SNMP Trap server is matched with each time range to obtain the time range corresponding to the duration of the abnormal SNMP Trap server; Based on the corresponding relationship between the time range and the time change strategy, the corresponding time change strategy is determined; Based on the determined time change strategy, the interval time is changed to obtain a new interval time; When it is determined that the SNMP Trap server is abnormal, a timing task is constructed based on the new interval time and the current time.

2. The BMC management alarm event processing method according to claim 1, wherein, Before the alarm information is sent to the SNMP Trap server, the method further comprises the following steps: The important types and cache times of each alarm information in the cache are obtained; Alarm information under different important types is divided; Based on the generation time of the alarm information, the alarm information under each important type is sorted to determine the first sending queue corresponding to each important type; The first sending queues are sorted based on the priority of the important types, and the second sending queue is determined; Correspondingly, when the SNMP Trap server is determined to be in normal operation, the alarm information is sent to the SNMP Trap server based on the second sending queue.

3. The BMC management alarm event processing method of claim 1, wherein, In the operation process of powering on the BMC, the method further comprises: Synchronizing each alarm information in the cache to the Flash in the BMC; Deleting the alarm information successfully sent in the cache, and synchronously deleting the alarm information successfully sent in the Flash.

4. The BMC management alarm event processing method of claim 1, wherein, Before constructing the first timing task after the SNMP Trap server is determined to be in abnormal operation, the method further comprises: Obtaining each alarm information in a first historical time period; Based on each alarm information in the first historical time period, obtaining the shortest interval duration of the same alarm information; If the alarm information first appears in the first historical time period, determining the shortest interval duration based on the generation time of the alarm information first appearing and the current time when the SNMP Trap server is determined to be in abnormal operation; Based on the shortest interval duration of each alarm information, determining the interval time required for constructing the first timing task after the SNMP Trap server is determined to be in abnormal operation.

5. The BMC management alarm event processing method of claim 1, wherein, The method further comprises: When the duration of the SNMP Trap server being in abnormal operation reaches the alarm duration, issuing a maintenance warning information of the SNMP Trap server.

6. A BMC management alarm event processing apparatus, characterized by comprising: Comprise: A construction module, configured to construct a timing task when the SNMP Trap server is determined to be in abnormal operation; The timing task is used to detect the operation state of the SNMP Trap server; A detection module, configured to detect the operation state of the SNMP Trap server when the timing task is triggered; A processing module, configured to send alarm information to the SNMP Trap server when the SNMP Trap server is determined to be in normal operation; In the process of determining that the SNMP Trap server is in abnormal operation, the construction module is specifically configured to: obtain alarm information at the current time, and determine that the SNMP Trap server is in abnormal operation when the number of failures of sending the alarm information at the current time to the SNMP Trap server reaches a first number; and correspondingly, based on the interval time, construct the first timing task after the SNMP Trap server is determined to be in abnormal operation; When the first timing task is triggered, it is determined that the SNMP Trap server is in abnormal operation when the number of failures of sending the detection signal to the SNMP Trap server reaches a second number; and correspondingly, based on the current time and the interval time, construct the nth timing task after the SNMP Trap server is determined to be in abnormal operation, wherein n is an integer greater than 2. The construction module is further configured to, after constructing the first timing task, acquire a plurality of time length ranges, each of which is contiguous to another; acquire a plurality of time change strategies, wherein each time change strategy has a different adjustment change rate; the adjustment change rate represents an increase percentage of a value before adjustment and a value after adjustment; based on an order of the adjustment change rates of the time change strategies from small to large, configure a corresponding time change strategy for each time length range, construct a corresponding relationship between the time length range and the time change strategy; record a determination time and a current time corresponding to the SNMP Trap server running exception; calculate a duration of the SNMP Trap server running exception according to the determination time and the current time of the SNMP Trap server running exception; match the duration of the SNMP Trap server running exception with each time length range to obtain a time length range corresponding to the duration of the SNMP Trap server running exception; determine a corresponding time change strategy based on the corresponding relationship between the time length range and the time change strategy; change the interval time based on the determined time change strategy to obtain a new interval time; and construct a timing task based on the new interval time and the current time when it is determined that the SNMP Trap server has a running exception.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the BMC alarm event processing method of any one of claims 1-5 when executing the program.

8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the BMC alarm event processing method of any one of claims 1-5 when executed by the processor.

Citation Information

Patent Citations

  • Alarm information transmission method, device and system

    CN109889373A

  • Alarm notification method and device

    CN115865626A