Operation and maintenance alarm information pushing method and system
By monitoring the waiting time in the message queue in real time in the operation and maintenance alarm system and sending fault notifications when timeouts occur, the message queue congestion problem caused by high write latency of solid-state drives and large-scale concurrent alarms is solved, thereby improving the reliability and fault response capability of the operation and maintenance system.
Patent Information
- Application Number
- CN202511211051.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-11
AI Technical Summary
In complex and high-concurrency operation and maintenance scenarios, the high write latency of solid-state drives and the influx of large-scale concurrent alarms can cause message queue congestion, preventing the operation and maintenance alarm system from pushing alarm information normally, thereby affecting the reliability of operation and maintenance.
By obtaining the timestamps of pending maintenance alarm information in the message queue, the waiting time is calculated in real time, and when the waiting time exceeds the threshold, a fault notification is sent to the target terminal through the backup channel to ensure that maintenance personnel can handle and push faults in a timely manner.
This reduces message queue congestion, ensures that maintenance personnel can reliably receive alarm information, and improves the reliability and fault response capability of the maintenance system.
Smart Images

Figure CN120934983A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of operation and maintenance alarm technology, and in particular to a method and system for pushing operation and maintenance alarm information. Background Technology
[0002] In the modern information technology environment, the stable operation of various systems relies on efficient operation and maintenance management. Operation and maintenance alarms, as a crucial feedback on system operational status, are essential for timely and accurate delivery to operation and maintenance personnel for rapid response and fault handling.
[0003] However, in complex and high-concurrency operation and maintenance scenarios, when the solid-state drives (SSDs) upon which the operation and maintenance alarm system relies experience intermittent high write latency due to approaching their design lifespan, a sudden surge of concurrent alarms caused by an upstream infrastructure failure may simultaneously block the disk persistence update of the entry message queue used to manage pending alarm information. This blockage can further suspend the main alarm processing thread, preventing the operation and maintenance alarm system from pushing any new alarm information. In this situation, the operation and maintenance alarm system itself fails to generate a meta-alarm regarding this internal processing capacity loss, thus falling into a state of "silent paralysis." From an external monitoring perspective, the operation and maintenance alarm system may still appear to be "running," with active programs and normal resource usage. However, in reality, all subsequently generated alarm information will be indefinitely piled up at the message queue entry point, unable to be pushed out. At this point, because the operation and maintenance personnel are unaware of the alarm information push failure, they will be unable to intervene and repair the unpushable alarm information in a timely manner, thereby affecting the reliability of operations and maintenance. Summary of the Invention
[0004] This application provides a method and system for pushing operation and maintenance alarm information, which can reduce the situation where message queues are blocked and operation and maintenance alarm information cannot be pushed normally due to high write latency of solid-state drives and large-scale concurrent alarm influx, so that operation and maintenance personnel can receive operation and maintenance alarm information stably, thereby improving operation and maintenance reliability.
[0005] The first aspect of this application provides a method for pushing operation and maintenance alarm information, including:
[0006] Obtain the timestamp of the pending operation and maintenance alarm information in the message queue, wherein the timestamp contains the timestamp when the pending operation and maintenance alarm information enters the message queue;
[0007] The waiting time for the pending maintenance alarm information is calculated in real time based on the timestamp.
[0008] Determine whether the waiting time is greater than a preset processing time threshold;
[0009] If so, a fault notification is sent to the target terminal through the backup channel. The fault notification is used to indicate that the pending maintenance alarm information has failed to be pushed.
[0010] Optionally, after obtaining the timestamps of the operation and maintenance alarm information to be processed in the message queue, the operation and maintenance alarm information push method further includes:
[0011] When it is detected that the pending operation and maintenance alarm information has been removed from the message queue, the pending operation and maintenance alarm information is marked as processed operation and maintenance alarm information;
[0012] Remove the timestamps from the processed maintenance alarm information.
[0013] Optionally, before sending the fault notification to the target terminal via the backup channel, the operation and maintenance alarm information push method further includes:
[0014] Remove the timestamps from the pending maintenance alarm information and store the pending maintenance alarm information in the verification information database;
[0015] The real-time processing status of each pending operation and maintenance alarm information in the pending verification information database is queried from the historical operation and maintenance alarm database. The real-time processing status includes pending processing and processed.
[0016] When the real-time processing status is "processed", the corresponding pending operation and maintenance alarm information will be removed from the pending verification information database.
[0017] When the real-time processing status is pending, the step of sending a fault notification to the target terminal through the backup channel is executed.
[0018] Optionally, after storing the pending maintenance alarm information in the verification information database, the maintenance alarm information push method further includes:
[0019] Calculate the verification duration of each pending maintenance alarm in the pending verification information database. The verification duration is the time difference between the current time and the initial time when each pending maintenance alarm enters the pending verification information database.
[0020] When the verification duration exceeds the preset verification duration threshold, the pending maintenance alarm information corresponding to the verification duration will be removed from the pending verification information database, and a verification timeout log will be generated.
[0021] Optionally, before determining whether the waiting time is greater than a preset processing time threshold, the operation and maintenance alarm information push method further includes:
[0022] The service type for obtaining the pending operation and maintenance alarm information;
[0023] Generate an initial processing time threshold based on the aforementioned business type;
[0024] The initial processing time threshold is updated based on the running load index of the message queue to obtain the preset processing time threshold.
[0025] Optionally, after determining whether the waiting time is greater than a preset processing time threshold, the operation and maintenance alarm information push method further includes:
[0026] Obtain real-time judgment results;
[0027] Based on the real-time judgment result, a real-time updated beacon file is generated in the local path. The beacon file contains the update time and the real-time judgment result.
[0028] Deploy a separate external monitoring entity and use the external monitoring entity to detect the accessibility of the beacon file at a fixed frequency;
[0029] If the beacon file is detected to be inaccessible, it is determined that there is an error in the real-time judgment result, and a monitoring failure notification is sent to the target terminal through an independent external channel.
[0030] Optionally, after the external monitoring entity detects the accessibility of the beacon file at a fixed frequency, the operation and maintenance alarm information push method further includes:
[0031] If the beacon file is detected to be accessible, the beacon file's inactivity period is calculated, where the inactivity period is the time difference between the current time and the beacon file's update time.
[0032] When the duration of the interruption exceeds a preset interruption duration threshold, it is determined that there is an error in the real-time judgment result, and the step of sending a monitoring failure notification to the target terminal through an independent external channel is executed.
[0033] Optionally, after calculating the inactivity duration of the beacon file, the operation and maintenance alarm information push method further includes:
[0034] When the pause duration is less than or equal to the preset pause duration threshold, the real-time judgment result is determined to be correct.
[0035] Based on the real-time judgment result, the running status of the pending operation and maintenance alarm information is determined, and the running status includes normal waiting for processing and waiting timeout.
[0036] When the running status is in the waiting timeout state, the step of sending a fault notification to the target terminal through the backup channel is executed.
[0037] Optionally, after sending the fault notification to the target terminal via the backup channel, the operation and maintenance alarm information push also includes:
[0038] Continuously monitor the duration since the fault notification or monitoring failure notification was sent;
[0039] When the duration of the fault notification or the monitoring failure notification is within the preset cooling period, the corresponding independent external channel or backup channel is shut down.
[0040] The second aspect of this application provides an operation and maintenance alarm information push system, including:
[0041] The acquisition unit is used to acquire the timestamp of the pending operation and maintenance alarm information in the message queue. The timestamp includes the timestamp when the pending operation and maintenance alarm information enters the message queue.
[0042] The calculation unit is used to calculate the waiting time of the pending operation and maintenance alarm information in real time based on the time tag;
[0043] The judgment unit is used to determine whether the waiting time is greater than a preset processing time threshold.
[0044] The sending unit is used to send a fault notification to the target terminal through a backup channel when the waiting time exceeds the preset processing time threshold. The fault notification is used to indicate that the pending maintenance alarm information has failed to be pushed.
[0045] As can be seen from the above technical solutions, this application has the following effects:
[0046] First, the timestamps of pending maintenance alarm messages in the message queue are obtained. These timestamps contain the timestamps of when the alarm messages entered the message queue. The waiting time for these alarm messages is calculated in real-time based on the timestamps. It is then determined whether the waiting time exceeds a preset processing time threshold. If so, a fault notification is sent to the target terminal via a backup channel, indicating a push failure in the pending maintenance alarm message. This allows for continuous monitoring of the waiting time of pending maintenance alarm messages in the message queue to identify potential push failures. A fault notification is then sent to the target terminal held by maintenance personnel via a backup channel, prompting them to address the push failure promptly. This reduces the likelihood of message queue congestion and inability to push maintenance alarm messages due to high write latency from SSDs and a large influx of concurrent alarms, ensuring maintenance personnel can reliably receive alarm messages and thus improving operational reliability. Attached Figure Description
[0047] Figure 1This is a schematic diagram of an embodiment of a method for pushing operation and maintenance alarm information in this application;
[0048] Figure 2-1 and Figure 2-2 This is a schematic diagram of another embodiment of the operation and maintenance alarm information push method in this application;
[0049] Figure 3 This is a schematic diagram of another embodiment of the operation and maintenance alarm information push method in this application;
[0050] Figure 4 This is a schematic diagram of an embodiment of an operation and maintenance alarm information push system in this application. Detailed Implementation
[0051] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0052] It should be understood that, when used in this application specification, the term "comprising" indicates the presence of the described feature, integral, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.
[0053] It should also be understood that the term “and / or” as used in this application specification means any combination of one or more of the associated listed items, as well as all possible combinations, and includes such combinations.
[0054] As used in this application specification, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [the described condition or event] is detected," or "in response to detection of [the described condition or event]."
[0055] Furthermore, in the description of this application, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0056] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0057] In existing technologies, when the solid-state drives (SSDs) upon which the operation and maintenance (O&M) alarm system relies experience intermittent high write latency due to nearing their design lifespan, a simultaneous surge of concurrent alarms caused by an upstream infrastructure failure can block disk persistence updates in the entry message queue used to manage pending alarm information. This blockage can further suspend the main alarm processing thread, preventing the O&M alarm system from distributing any new alarm information. In this situation, the O&M alarm system itself fails to generate meta-alarms regarding this internal processing capacity loss, thus entering a state of "silent paralysis." From an external monitoring perspective, the O&M alarm system may still appear to be "running," with active programs and normal resource usage. However, in reality, all subsequently generated alarm information will be indefinitely accumulated at the message queue entry point, unable to be distributed. At this point, O&M personnel, unaware of the alarm information distribution failure, will be unable to intervene and repair the undistributed alarm information in a timely manner, thereby affecting O&M reliability.
[0058] Based on this, this application discloses a method and system for pushing operation and maintenance alarm information, which can reduce the situation where message queues are blocked and operation and maintenance alarm information cannot be pushed normally due to high write latency of solid-state drives and large-scale concurrent alarm influx, so that operation and maintenance personnel can receive operation and maintenance alarm information stably, thereby improving operation and maintenance reliability.
[0059] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0060] The operation and maintenance alarm information push method described in this application is applied to systems, terminals, servers or other devices with logical analysis and processing capabilities, and is not specifically limited here.
[0061] Please see Figure 1 As shown, one embodiment of the operation and maintenance alarm information push method in this application includes:
[0062] 101. Obtain the timestamp of the pending operation and maintenance alarm information in the message queue. The timestamp contains the timestamp when the pending operation and maintenance alarm information enters the message queue.
[0063] In this embodiment, newly generated operation and maintenance alarm information is identified as pending operation and maintenance alarm information and placed into a message queue. Simultaneously, a specific timestamp is assigned to the pending operation and maintenance alarm information, recording the start time of its entry into the message queue. Specifically, pending operation and maintenance alarm information in the message queue can be monitored in real time. Once a new pending operation and maintenance alarm information is detected, a timestamp recording the current time is generated and attached to the new alarm information, and this timestamp is embedded into the timestamp. This timestamp can be accurate to milliseconds or microseconds to ensure the accuracy of subsequent calculations. It is understood that the timestamp format can be set as: alarm type-server IP-timestamp; the specific format is not limited here.
[0064] 102. Calculate the waiting time for pending maintenance alarm information in real time based on time tags;
[0065] In this embodiment, when the timestamp of a pending maintenance alarm message in the message queue is obtained, the start time node in the timestamp is continuously compared with the real-time time to determine the length of time the pending alarm message has remained in the message queue. This length of time is the waiting time. For example, a scheduled task can be set up to traverse the message queue at fixed time intervals, perform the operation and maintenance alarm message operation and maintenance timekeeping operation and maintenance timekeeping operation and maintenance information by subtracting its timestamp from the current time, and determine the result of the operation and maintenance alarm message operation and maintenance timekeeping operation and maintenance information waiting timekeeping operation and maintenance timekeeping operation and maintenance information waiting timekeeping operation and maintenance alarm message waiting ...
[0066] 103. Determine whether the waiting time is greater than the preset processing time threshold. If so, proceed to step 104.
[0067] In this embodiment, the calculated waiting time for each pending maintenance alarm message is compared with a pre-set time threshold. This preset processing time threshold can be flexibly set according to different types of alarm messages. For example, for alarm messages indicating insufficient database disk space, the preset processing time threshold is set to 3 seconds; while for alarm messages indicating network fluctuations, the preset processing time threshold is set to 1 second. By comparing the waiting time with the preset processing time threshold, an acceptable alarm processing delay upper limit can be set for each pending maintenance alarm message in the message queue. If the waiting time is less than or equal to the preset processing time threshold, it indicates that the pending maintenance alarm message is in a normal push processing flow; however, if the waiting time is less than or equal to the preset processing time threshold, it can be considered that the push processing flow of the pending maintenance alarm message may be abnormal.
[0068] 104. Send a fault notification to the target terminal through the backup channel. The fault notification is used to indicate that the push of pending operation and maintenance alarm information has failed.
[0069] It's important to note that the backup channel is a message push channel independent of the main channel. The main channel is used to push pending maintenance alarms removed from the message queue to the target terminal, while the backup channel is used to push fault notifications generated when the push of these pending maintenance alarms fails. This backup channel can take various forms, such as sending SMS notifications through a separate SMS gateway, sending emails through a separate mail server, or sending messages through a separate instant messaging application interface. The target terminal can be the maintenance personnel's mobile phone, email address, or a specific alarm receiving platform. When the waiting time exceeds a preset processing time threshold, a fault notification is generated based on the corresponding pending maintenance alarm and sent to the target terminal held by the maintenance personnel through the backup channel. This allows the maintenance personnel to promptly resolve the fault caused by the failed push of maintenance alarms. The fault notification may include specific maintenance alarm information, the waiting time, and a timestamp, enabling the maintenance personnel to quickly understand the push failure situation.
[0070] In this embodiment, the timestamps of pending maintenance alarm information in the message queue are first obtained. These timestamps contain the timestamps of when the pending maintenance alarm information entered the message queue. The waiting time for the pending maintenance alarm information is calculated in real time based on the timestamps. It is then determined whether the waiting time exceeds a preset processing time threshold. If so, a fault notification is sent to the target terminal through a backup channel. This fault notification indicates a push failure of the pending maintenance alarm information. In this way, by continuously monitoring the waiting time of pending maintenance alarm information in the message queue, it is possible to identify whether a push failure will occur. A fault notification is then sent to the target terminal held by the maintenance personnel through a backup channel to prompt them to address the push failure in a timely manner. This reduces the likelihood of message queue congestion and inability to push maintenance alarm information normally due to high write latency of solid-state drives and a large influx of concurrent alarms, allowing maintenance personnel to reliably receive maintenance alarm information and thus improving maintenance reliability.
[0071] Please see Figure 2-1 and Figure 2-2 As shown, another embodiment of the operation and maintenance alarm information push method in this application includes:
[0072] 201. Obtain the timestamp of the pending operation and maintenance alarm information in the message queue. The timestamp contains the timestamp when the pending operation and maintenance alarm information enters the message queue.
[0073] Step 201 in this embodiment is the same as described above. Figure 1 Step 101 in the illustrated embodiment is similar and will not be described again here.
[0074] 202. When it is detected that the pending operation and maintenance alarm information has been removed from the message queue, mark the pending operation and maintenance alarm information as the processed operation and maintenance alarm information.
[0075] 203. Remove the timestamps from processed maintenance alarm messages;
[0076] Optionally, in this embodiment, if a pending maintenance alarm message has been removed from the message queue and processed during the monitoring period, it may still be tracked or have its waiting time calculated, leading to unnecessary resource consumption or false alarm notifications. Therefore, when it is detected that a pending maintenance alarm message has been removed from the message queue, the status of the removed pending maintenance alarm message can be updated and its timestamp removed. Specifically, mechanisms such as polling, event listening, or message callbacks can be used to determine whether the pending maintenance alarm message still exists in the message queue. If it is determined that it has been removed from the message queue, it means that the pending maintenance alarm message has entered the normal push process and no longer needs to be monitored for its waiting time. At this time, the status identifier of the pending maintenance alarm message can be updated, such as remarking the pending maintenance alarm message as a processed maintenance alarm message. At the same time, the timestamp associated with the processed maintenance alarm message is deleted. In this way, it is ensured that only alarm messages that are truly in a pending state and are still in the message queue will be continuously monitored and have their waiting time calculated. This effectively avoids invalid tracking of processed alarm information, reduces the waste of network resources, improves the accuracy of fault notification, and prevents false alarms caused by alarm information that has been processed but is still being monitored.
[0077] 204. Real-time calculation of waiting time for pending maintenance alarm information based on time tags;
[0078] Step 204 in this embodiment is the same as described above. Figure 1 Step 102 in the illustrated embodiment is similar and will not be described again here.
[0079] 205. Business types for acquiring pending operation and maintenance alarm information;
[0080] 206. Generate an initial processing time threshold based on the business type;
[0081] 207. Update the initial processing time threshold based on the message queue's runtime load metrics to obtain the preset processing time threshold;
[0082] Optionally, in this embodiment, upon receiving a pending maintenance alarm, the business type to which it belongs can be identified from the metadata, message body, or predefined rules of the alarm. For example, the business type may include "core business alarm," "non-core business alarm," "emergency alarm," and "general alarm," which typically reflect the importance or processing priority of the alarm. Then, for different business types, a basic initial processing time threshold is pre-configured or dynamically calculated. For example, for the "emergency alarm" business type, the initial processing time threshold can be set to a shorter time, such as 5 seconds; while for the "general alarm" business type, the initial processing time threshold can be set to a longer time, such as 30 seconds. Next, the operational load metrics of the message queue are continuously monitored, such as the current length of the message queue, CPU utilization, memory utilization, and message processing rate. When the message queue's operational load metric is high, the initial processing time threshold can be appropriately extended to allow more time for processing pending maintenance alarm information, avoiding misjudgments of push failures due to sudden high load. Conversely, when the operational load metric is low, the initial processing time threshold can be shortened to improve alarm sensitivity. Based on this, the initial processing time threshold can be dynamically updated according to the message queue's operational load metric to obtain a predicted processing time threshold. This ensures that the preset processing time threshold can adaptively adjust when the operational load metric changes, thus avoiding misjudgments or missed detections caused by fixed thresholds. In this way, the preset processing time threshold can be dynamically determined based on the business characteristics of pending maintenance alarm information and the real-time operational status of the message queue. This significantly improves the accuracy and flexibility of fault diagnosis, avoiding false alarms or missed detections caused by improper threshold settings, thereby ensuring timely and accurate identification of maintenance alarm information push failures and effectively guaranteeing the stability and reliability of system operations.
[0083] 208. Determine whether the waiting time is greater than the preset processing time threshold. If so, proceed to step 209.
[0084] 209. Remove the timestamps from pending maintenance alarm information and store the pending maintenance alarm information in the verification information database;
[0085] 210. Query the real-time processing status of each pending operation and maintenance alarm in the pending verification information database from the historical operation and maintenance alarm database. The real-time processing status includes pending and processed.
[0086] 211. When the real-time processing status is "processed", remove the corresponding pending maintenance alarm information from the pending verification information database;
[0087] 212. When the real-time processing status is pending, a fault notification is sent to the target terminal through the backup channel. This fault notification is used to indicate that the pending operation and maintenance alarm information has failed to be pushed.
[0088] Optionally, in this embodiment, when the waiting time for a pending maintenance alarm exceeds a preset processing time threshold, a fault notification is not immediately sent. Instead, the timestamp of the pending maintenance alarm is removed to avoid repeatedly triggering subsequent waiting time checks. Simultaneously, the pending maintenance alarm is temporarily stored in a verification database. The verification database can be understood as a temporary storage area for pending maintenance alarms that may require fault notifications but whose true status still needs further confirmation. Then, the real-time processing status of each pending maintenance alarm in the verification database is queried from the historical maintenance alarm database. This historical maintenance alarm database records the entire lifecycle status of all maintenance alarms from generation to completion, including whether they have been processed and the processing result. The real-time processing status can include "pending" and "processed," indicating whether the pending maintenance alarm is still in an unresolved state. If the query results show that the real-time processing status of the pending maintenance alarm is "processed," it indicates that although the alarm has been waiting in the message queue for too long, it has actually been successfully processed by the main processing flow. In this case, the corresponding pending maintenance alarm will be removed from the verification database, thus avoiding the sending of unnecessary fault notifications. However, if the query results show that the real-time processing status of the pending maintenance alarm is still "pending," it means that the alarm has indeed not been processed for a long time, and the main processing flow has not processed it either. In this case, a fault notification can be sent to the target terminal through a backup channel, ensuring that only faults that truly require attention are notified to maintenance personnel. This intermediate verification process allows for status verification before sending fault notifications, ensuring that only maintenance alarms requiring intervention are pushed, thereby optimizing the fault response process and effectively reducing invalid fault notifications caused by misjudgment or information delays.
[0089] However, in actual operation, if certain pending maintenance alarm information is not removed from the verification information database in a timely manner due to various reasons (such as abnormal historical database queries, data synchronization delays, or the complexity of the alarm's own status), invalid or outdated data may continuously accumulate in the verification information database, thereby occupying system resources and potentially affecting the efficiency of subsequent troubleshooting. Therefore, in another possible implementation, after storing the pending maintenance alarm information in the verification information database, the maintenance alarm information push method in this application may further include: calculating the verification duration of each pending maintenance alarm information in the verification information database, where the verification duration is the time difference between the current time and the initial time when each pending maintenance alarm information enters the verification information database; when the verification duration exceeds a preset verification duration threshold, the pending maintenance alarm information corresponding to the verification duration is removed from the verification information database, and a verification timeout log is generated.
[0090] Specifically, the verification duration refers to the length of time an pending maintenance alarm message remains in the pending verification database after it has been stored. This verification duration is calculated by determining the time difference between the current time and the initial time when the alarm message entered the database. For example, if an alarm message enters the database at 10:00 AM and the current time is 10:30 AM, its verification duration is 30 minutes. The preset verification duration threshold is a pre-defined upper limit that defines the maximum allowed time for an alarm message to remain in the database. When the calculated verification duration exceeds the preset threshold, it indicates that the alarm message has remained in the database for too long, potentially indicating an anomaly or that further verification is unnecessary. In this case, the alarm message can be removed from the database to avoid wasting resources on invalid data. Simultaneously, a verification timeout log can be generated to record and track such anomalies. This log can contain information such as the identifier of pending maintenance alarm information, the time it entered the verification database, the verification duration, and the removal time, which facilitates subsequent troubleshooting and optimization. Based on this, this embodiment can effectively manage the lifecycle of the verification database and avoid resource consumption and data redundancy problems caused by the long-term retention of pending maintenance alarm information.
[0091] In another alternative embodiment, please refer to Figure 3 As shown, after determining whether the waiting time exceeds a preset processing time threshold, the operation and maintenance alarm information push method in this application may further include:
[0092] 301. Obtain real-time judgment results;
[0093] 302. Generate a real-time updated beacon file in the local path based on the real-time judgment result. The beacon file contains the update time and the real-time judgment result.
[0094] 303. Deploy an independent external monitoring entity to detect the accessibility of beacon files at a fixed frequency;
[0095] 304. If the beacon file is found to be inaccessible, it is determined that there is an error in the real-time judgment result, and a monitoring failure notification is sent to the target terminal through an independent external channel;
[0096] Optionally, in this embodiment, the real-time judgment result refers to a Boolean result, such as "yes" or "no," obtained by comparing the waiting time for the pending operation and maintenance alarm information with a preset processing time threshold. This indicates whether a fault notification needs to be sent. The beacon file can be understood as a lightweight status indicator file, generated and stored in a specific path on the local file system. The beacon file contains key metadata, such as its most recent update timestamp and the current real-time judgment result. The independent external monitoring entity can be a process, daemon, or remote monitoring agent independent of the main alarm push service. This entity is configured to attempt to access or read the beacon file at a preset fixed frequency, such as every 5 or 10 seconds. Its purpose is to verify whether the main alarm push service is running normally and continuously updating its status. Furthermore, if the external monitoring entity encounters inaccessible conditions when attempting to access the beacon file, such as the file not existing, incorrect file permissions, insufficient disk space, or the main service process crashing and preventing the file from being updated, it can be determined that the real-time judgment result is incorrect or the main service itself has failed. In this scenario, to ensure timely transmission of fault information, a monitoring failure notification will be sent to the target terminal via an external channel independent of the main alarm notification channel. This channel could be such as an SMS gateway, a separate email service, or a third-party alarm platform. This monitoring failure notification explicitly indicates an anomaly that occurred while monitoring the waiting time of pending maintenance alarm information in the message queue, requiring manual intervention for verification. This provides a reliable means of verifying the accuracy of the judgment and the reliability of the notification mechanism, reducing the likelihood of untimely responses to anomalies in monitoring waiting time or the notification link.
[0097] 305. If the beacon file is detected to be accessible, calculate the beacon file's inactivity period, which is the time difference between the current time and the beacon file's update time.
[0098] 306. When the duration of the interruption exceeds the preset interruption duration threshold, it is determined that there is an error in the real-time judgment result, and a monitoring failure notification is sent to the target terminal through an independent external channel;
[0099] Optionally, when the internal logic or data source that generates the beacon file malfunctions, the beacon file may not be updated for an extended period. Even if the beacon file itself is accessible, external monitoring entities cannot detect such potential errors in a timely manner based solely on accessibility, potentially delaying responses to anomalies in the maintenance alarm push system. Therefore, detection of the beacon file's update pause duration can be introduced to compensate for the shortcomings of relying solely on beacon file accessibility for monitoring. Specifically, the beacon file's update pause duration refers to the time difference between the current system time and the update time recorded within the beacon file. When the calculated pause duration exceeds this preset threshold, it indicates that the beacon file content is severely outdated. In this case, even if the beacon file itself is accessible, it should be considered an error in the real-time judgment result, triggering a monitoring failure notification sent to the target terminal via an independent external channel. This ensures that maintenance personnel can be promptly informed of the abnormal status when monitoring the maintenance alarm push status. This effectively identifies and responds to abnormal beacon file content updates; that is, even when the beacon file is accessible but its content has not been updated for an extended period, potential errors in the real-time judgment result can be detected and reported promptly. This greatly enhances the ability to detect latent faults and avoids misjudgments or missed reports due to outdated data.
[0100] 307. When the pause duration is less than or equal to the preset pause duration threshold, the real-time judgment result is confirmed to be correct.
[0101] 308. Determine the running status of pending operation and maintenance alarm information based on real-time judgment results. The running status includes normal waiting for processing and waiting timeout.
[0102] 309. When the running status is "waiting for timeout", execute the step of sending a fault notification to the target terminal through the backup channel.
[0103] Optionally, when the downtime is less than or equal to a preset downtime threshold, it indicates that the beacon file has been updated in a timely manner, and the real-time judgment results it contains are reliable and valid, thus confirming that the real-time judgment results are correct. Further, after confirming that the real-time judgment results are correct, the operating status of the pending maintenance alarm information will be determined based on these results. This operating status can be divided into two cases: normal waiting for processing and waiting for timeout. Normal waiting for processing indicates that the pending maintenance alarm information is still in the normal processing flow; while waiting for timeout indicates that there may be processing delays or faults. When the operating status is determined to be waiting for timeout, the system will send a fault notification to the target terminal through a backup channel. This ensures that even if there is a problem with the alarm push from the main channel, and the external monitoring system itself is operating normally, fault alarms can still be triggered in a timely manner through effective interpretation of the beacon file content, avoiding the omission or delayed processing of alarm information.
[0104] In another possible implementation, after sending a fault notification or monitoring failure notification to the target terminal through a backup channel, the operation and maintenance alarm information push method in this application may further include: continuously monitoring the duration of the fault notification or monitoring failure notification; and closing the corresponding independent external channel or backup channel when the duration of the fault notification or monitoring failure notification is within a preset cooling period.
[0105] Specifically, after sending a fault notification or monitoring failure notification, a timer is started to record the elapsed time from the moment the notification is successfully sent. This elapsed duration can be continuously updated and tracked to ensure its accuracy. For example, the elapsed duration can be calculated by adding a sending timestamp to the notification sending record and periodically comparing it with the current time. The preset cooldown period can be understood as a pre-defined time window within which the sending and receiving process of the notification is considered to be in a "cooling-off" or "observation" state. This preset cooldown period can be flexibly configured based on actual business needs, the importance of the notification, the urgency of the fault, and the response speed of the target terminal. For example, for urgent fault notifications, the cooldown period can be set shorter to release resources as quickly as possible; for non-urgent notifications, the cooldown period can be appropriately extended. When the elapsed duration of a fault notification or monitoring failure notification is within the preset cooldown period, the corresponding independent external channel or backup channel is closed, i.e., resending of previously sent fault notifications or monitoring failure notifications is prohibited within the preset cooldown period to reduce system resource consumption. This time-managed channel release mechanism reduces unnecessary network traffic and server load, and avoids duplicate alarms or false alarms that may result from the channel being open for a long time.
[0106] Please see Figure 4 As shown, one embodiment of the operation and maintenance alarm information push system in this application includes:
[0107] The acquisition unit 401 is used to acquire the timestamp of the pending operation and maintenance alarm information in the message queue. The timestamp contains the timestamp when the pending operation and maintenance alarm information enters the message queue.
[0108] The calculation unit 402 is used to calculate the waiting time of pending operation and maintenance alarm information in real time based on time tags;
[0109] The judgment unit 403 is used to determine whether the waiting time is greater than the preset processing time threshold.
[0110] The sending unit 404 is used to send a fault notification to the target terminal through a backup channel when the waiting time exceeds a preset processing time threshold. The fault notification is used to indicate that the push failure of the pending operation and maintenance alarm information has occurred.
[0111] In this embodiment, the acquisition unit 401 acquires the timestamps of pending maintenance alarm information in the message queue, which includes the timestamp when the pending maintenance alarm information entered the message queue; the calculation unit 402 calculates the waiting time of the pending maintenance alarm information in real time based on the timestamps; the judgment unit 403 judges whether the waiting time is greater than a preset processing time threshold; and the sending unit 404 sends a fault notification to the target terminal through a backup channel when the waiting time is greater than the preset processing time threshold. This fault notification indicates that the pending maintenance alarm information has experienced a push failure. In this way, by continuously monitoring the waiting time of the pending maintenance alarm information in the message queue, it is possible to identify whether a push failure will occur, and a fault notification is sent to the target terminal held by the maintenance personnel through a backup channel to prompt the maintenance personnel to handle the failure of the maintenance alarm information push in a timely manner. This reduces the situation where message queue congestion and inability to push maintenance alarm information normally are caused by high write latency of solid-state drives and a large influx of concurrent alarms, allowing maintenance personnel to receive maintenance alarm information stably, thereby improving maintenance reliability.
[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0113] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0114] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0115] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0116] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for pushing operation and maintenance alarm information, characterized in that, include: Obtain the timestamp of the pending operation and maintenance alarm information in the message queue, wherein the timestamp contains the timestamp when the pending operation and maintenance alarm information enters the message queue; The waiting time for the pending maintenance alarm information is calculated in real time based on the timestamp. Determine whether the waiting time is greater than a preset processing time threshold; If so, a fault notification is sent to the target terminal through the backup channel. The fault notification is used to indicate that the pending maintenance alarm information has failed to be pushed.
2. The method for pushing operation and maintenance alarm information according to claim 1, characterized in that, After obtaining the timestamps of the operation and maintenance alarm information to be processed in the message queue, the operation and maintenance alarm information push method further includes: When it is detected that the pending operation and maintenance alarm information has been removed from the message queue, the pending operation and maintenance alarm information is marked as processed operation and maintenance alarm information; Remove the timestamps from the processed maintenance alarm information.
3. The method for pushing operation and maintenance alarm information according to claim 1, characterized in that, Before sending the fault notification to the target terminal via the backup channel, the operation and maintenance alarm information push method further includes: Remove the timestamps from the pending maintenance alarm information and store the pending maintenance alarm information in the verification information database; The real-time processing status of each pending operation and maintenance alarm information in the pending verification information database is queried from the historical operation and maintenance alarm database. The real-time processing status includes pending processing and processed. When the real-time processing status is "processed", the corresponding pending operation and maintenance alarm information will be removed from the pending verification information database. When the real-time processing status is pending, the step of sending a fault notification to the target terminal through the backup channel is executed.
4. The method for pushing operation and maintenance alarm information according to claim 3, characterized in that, After storing the pending maintenance alarm information into the verification information database, the maintenance alarm information push method further includes: Calculate the verification duration of each pending maintenance alarm in the pending verification information database. The verification duration is the time difference between the current time and the initial time when each pending maintenance alarm enters the pending verification information database. When the verification duration exceeds the preset verification duration threshold, the pending maintenance alarm information corresponding to the verification duration will be removed from the pending verification information database, and a verification timeout log will be generated.
5. The method for pushing operation and maintenance alarm information according to claim 1, characterized in that, Before determining whether the waiting time is greater than a preset processing time threshold, the operation and maintenance alarm information push method further includes: The service type for obtaining the pending operation and maintenance alarm information; Generate an initial processing time threshold based on the aforementioned business type; The initial processing time threshold is updated based on the running load index of the message queue to obtain the preset processing time threshold.
6. The method for pushing operation and maintenance alarm information according to claim 1, characterized in that, After determining whether the waiting time is greater than a preset processing time threshold, the operation and maintenance alarm information push method further includes: Obtain real-time judgment results; Based on the real-time judgment result, a real-time updated beacon file is generated in the local path. The beacon file contains the update time and the real-time judgment result. Deploy a separate external monitoring entity and use the external monitoring entity to detect the accessibility of the beacon file at a fixed frequency; If the beacon file is detected to be inaccessible, it is determined that there is an error in the real-time judgment result, and a monitoring failure notification is sent to the target terminal through an independent external channel.
7. The method for pushing operation and maintenance alarm information according to claim 6, characterized in that, After utilizing the external monitoring entity to detect the accessibility of the beacon file at a fixed frequency, the operation and maintenance alarm information push method further includes: If the beacon file is detected to be accessible, the beacon file's inactivity period is calculated, where the inactivity period is the time difference between the current time and the beacon file's update time. When the duration of the interruption exceeds a preset interruption duration threshold, it is determined that there is an error in the real-time judgment result, and the step of sending a monitoring failure notification to the target terminal through an independent external channel is executed.
8. The method for pushing operation and maintenance alarm information according to claim 7, characterized in that, After calculating the inactivity duration of the beacon file, the operation and maintenance alarm information push method further includes: When the pause duration is less than or equal to the preset pause duration threshold, the real-time judgment result is determined to be correct. Based on the real-time judgment result, the running status of the pending operation and maintenance alarm information is determined, and the running status includes normal waiting for processing and waiting timeout. When the running status is in the waiting timeout state, the step of sending a fault notification to the target terminal through the backup channel is executed.
9. The method for pushing operation and maintenance alarm information according to claim 6 or 7, characterized in that, After sending the fault notification to the target terminal via the backup channel, the operation and maintenance alarm information push also includes: Continuously monitor the duration since the fault notification or monitoring failure notification was sent; When the duration of the fault notification or the monitoring failure notification is within the preset cooling period, the corresponding independent external channel or backup channel is shut down.
10. A maintenance alarm information push system, characterized in that, include: The acquisition unit is used to acquire the timestamp of the pending operation and maintenance alarm information in the message queue. The timestamp includes the timestamp when the pending operation and maintenance alarm information enters the message queue. The calculation unit is used to calculate the waiting time of the pending operation and maintenance alarm information in real time based on the time tag; The judgment unit is used to determine whether the waiting time is greater than a preset processing time threshold. The sending unit is used to send a fault notification to the target terminal through a backup channel when the waiting time exceeds the preset processing time threshold. The fault notification is used to indicate that the pending maintenance alarm information has failed to be pushed.