Fault self-switching device of high-reliability computer monitoring system
Patent Information
- Application Number
- CN202610678839.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-18
AI Technical Summary
[0005]针对现有技术的不足,本发明提供了高可靠性计算机监控系统的故障自切换装置,解决了缺乏针对时序数据的专属校验逻辑,无法精准识别切换过程中出现的时序数据丢失、读写错误、时序错乱等问题
本发明通过主备机切换的精准化、智能化触发,有效避免误切换与漏切换,保障监控系统连续运行。通过实时监测中心对主机CPU占用、内存占用、网络参数等关键运行参数的全方位实时监测,为切换决策提供了全面、实时的数据支撑;切备处理中心通过设定不同运行项的专属波动值Bi,结合波动特征、波动时长、波动占比及跟随波动项的综合评定,构建了多维度、多层次的切换判定机制,既避免了单一参数瞬时波动导致的误切换,也防止了主机多参数异常叠加引发故障扩大而未及时切换的问题,同时通过总和占比的量化判定的方式,使切换决策更具客观性和可操作性,兼顾了系统运行的稳定性与故障响应的及时性,还可通过波动项展示为人工干预提供参考,进一步降低主机过度负载的风险;
Smart Images

Figure CN122594050A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer monitoring technology, specifically to a fault self-switching device for a high-reliability computer monitoring system. Background Technology
[0002] In high-reliability computer monitoring systems, master-slave switching is a core component to ensure uninterrupted system operation. The accuracy of the switching and the integrity of the data after the switching directly determine the reliability of the monitoring system. It is widely used in critical scenarios such as industry, power, and transportation where the system has extremely high requirements for continuity and stability.
[0003] Currently, existing primary / standby fault self-switching technology still has many technical shortcomings that urgently need to be addressed: The switching trigger determination method is relatively simple, relying mainly on the threshold of a single operating parameter for judgment. It lacks comprehensive consideration of the correlation between fluctuations in multiple operating items such as CPU usage, memory usage, and network parameters. It is easy to cause false switching due to instantaneous fluctuations of a single parameter, or to miss switching due to the failure to identify abnormal superposition of multiple parameters in time, which in turn causes the monitoring system to be interrupted and affects the normal operation of critical scenarios. The data verification mechanism during the switchover process is imperfect. Traditional verification methods often use simple parity checks or single data comparisons, lacking dedicated verification logic for time-series data. This makes it impossible to accurately identify problems such as time-series data loss, read / write errors, and time-series disorder during the switchover process, resulting in damage to the continuity and accuracy of monitoring data, making it difficult to support subsequent monitoring analysis and control decisions. Some switching devices have unreasonable functional module designs, poor coordination between modules, and limited adaptability. Furthermore, the switching and verification process requires a lot of manual intervention, which not only increases the operation and maintenance costs but also further reduces the system's operational reliability, making it unable to meet the actual application needs of various high-reliability monitoring scenarios.
[0004] In view of the shortcomings of the existing technology, there is an urgent need for a fault self-switching device that can achieve precise switching triggering and efficient data verification, so as to solve the pain points of the existing technology and improve the reliability and practicality of computer monitoring system. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a fault self-switching device for a highly reliable computer monitoring system, which solves the problems of lacking dedicated verification logic for time-series data and being unable to accurately identify time-series data loss, read / write errors, and timing disorder during the switching process.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a fault self-switching device for a high-reliability computer monitoring system, comprising: The real-time monitoring center monitors the operating parameters associated with the computer monitoring host in real time. The failover processing center, based on real-time monitored operating parameters, confirms whether different operating items within those parameters are experiencing abnormal fluctuations, and executes the failover process based on the confirmation results to complete the automatic switchover of the primary and standby machines. Specifically: Confirm the category to which the operating parameters belong, and extract the fluctuation value B set for the corresponding category. i Where i represents different running items, based on the real-time generated running parameters, the fluctuation characteristics of the current time compared to the previous time are determined, and the running parameters corresponding to the running items at the current time are set as Yx1, and the running parameters corresponding to the running items at the previous time are set as Yx2. Their fluctuation characteristics = |Yx1-Yx2|, if the fluctuation characteristics ≤ B i If the fluctuation characteristic is greater than B, then no action is taken. i If the duration of the fluctuation signal exceeds 10 seconds, the current running item is marked as a fluctuation item and the fluctuation duration is recorded synchronously. If the duration of the fluctuation signal does not exceed 10 seconds, no marking is performed. Confirm the volatility percentage associated with the volatility term; its volatility percentage = |Volatility characteristic - B i |÷B i Confirm the fluctuation percentage of the corresponding fluctuation item, and then identify whether there are other fluctuation signals within the recorded fluctuation duration. If they exist, mark the corresponding running item as a following fluctuation item; otherwise, do not mark it. The method of confirming the same fluctuation percentage for fluctuation items is adopted. The fluctuation percentage associated with the following fluctuation item is identified, and the confirmed fluctuation percentages are summed to lock the total percentage. If the total percentage is ≥1, the switchover process is executed to switch the running host to the running standby host. If the total percentage is <1, monitoring continues and the marked fluctuation items and their following fluctuation items are displayed synchronously. The data feature confirmation center, based on the switchover time of the primary and backup machines, identifies a set of negotiation time intervals and confirms the time-series data generated by the primary machine within these intervals. It then confirms the time change characteristics associated with different time-series data, identifies regular patterns from the confirmed sets of time change characteristics, and records these patterns as primary machine characteristics. The specific method is as follows: Mark the confirmed switching time as T, and based on the set time threshold S1, confirm a set of negotiation time intervals [T-S1, T+S1]; Confirm the timing data generated by the host within the negotiation time interval, confirm the timing associated with the timing data, generate a timing sequence associated with the timing data according to the chronological order, confirm the timing difference between adjacent timings within the timing sequence, confirm the timing difference between several groups of adjacent timings in sequence, identify the timing difference with the same value from the confirmed groups of timing difference, confirm the same number of different timing differences, lock the timing difference associated with the largest number of the same number, and record this timing difference as the host feature; The timing verification processing center, based on the confirmed negotiation time interval and the characteristics of the primary machine, confirms the actual timing data generated by the standby machine within the negotiation time interval, and locks the actual timing sequence from the actual timing data. Then, based on the characteristics of the primary machine, it generates a standard timing sequence generated by the standby machine. The actual timing sequence is compared and verified with the standard timing sequence to identify whether there are errors or omissions in the timing data during the primary / standby machine switchover process, and corrects them accordingly. The specific method is as follows: The confirmed host characteristics are marked as Tz, and a standard time sequence is generated based on the confirmed switching time T. The standard time sequence is {T, T+Tz, T+2Tz, ..., T+nTz}, and all time sequences of the standard time sequence are within the negotiation time interval, and T+nTz≤T+S1, and n is a positive integer. Then, the timing data generated by the standby machine within the negotiation time interval are confirmed in sequence, and based on the timing associated with different timing data, and according to the chronological order, several sets of actual timing sequences associated with the timing data are generated. Identify the timing differences associated with adjacent timing sequences within the actual timing sequence, and from the confirmed sets of timing differences, identify timing differences with the same value, and confirm the number of times different timing differences are the same, lock the timing difference associated with the largest number of the same value, and record the confirmed timing differences as standby characteristics. Identify whether the primary characteristics and standby characteristics are consistent. If they are consistent, execute the comparison process between the actual timing sequence and the standard timing sequence. If they are not consistent, directly generate a data read / write error signal for display. The system identifies whether the actual time series sequence and the standard time series sequence have the same sorting position. If they are the same, no processing is required. If they are different, the actual time series sequence is recorded as an abnormal sequence and a re-read / write signal is generated. Based on the re-read / write signal, the standby machine reads and writes the subsequent associated time series data in sequence, starting from the switching time.
[0007] Preferably, the operating parameters include CPU usage, memory usage, network parameters, voltage parameters, and temperature parameters.
[0008] This invention provides a fault-switching device for a highly reliable computer monitoring system. Compared with the prior art, it has the following advantages: This invention effectively avoids erroneous and missed switching by precisely and intelligently triggering the primary / standby switchover, ensuring the continuous operation of the monitoring system. Through comprehensive real-time monitoring of key operating parameters such as host CPU usage, memory usage, and network parameters by the real-time monitoring center, it provides comprehensive and real-time data support for switchover decisions. The switchover processing center constructs a multi-dimensional and multi-level switchover judgment mechanism by setting exclusive fluctuation values Bi for different operating items and combining fluctuation characteristics, fluctuation duration, fluctuation proportion, and comprehensive evaluation of following fluctuation items. This avoids erroneous switching caused by instantaneous fluctuations of a single parameter, and also prevents the problem of failure to switch over in time due to the amplification of faults caused by the superposition of multiple abnormal host parameters. Furthermore, the quantitative judgment method based on the total proportion makes the switchover decision more objective and operable, balancing system stability and timely fault response. It also provides a reference for manual intervention through fluctuation item display, further reducing the risk of host overload. Precise verification of the integrity of time-series data after primary / standby switchover effectively prevents data loss, read / write errors, and other problems, ensuring the accuracy of data transmission and storage. The data feature confirmation center extracts the time-series characteristics (host characteristics) of the host machine during normal operation by locking onto the negotiation time interval associated with the switchover, providing a precise reference standard for post-switchover data verification. This avoids the shortcomings of traditional verification methods, such as lack of specificity and low verification accuracy. The time-series verification processing center compares the generated standard time-series sequence with the actual time-series sequence of the standby machine. It first verifies the consistency between the standby machine characteristics and the host characteristics, and then compares the matching degree of the time-series sequences. This achieves dual verification of data errors and missing data after switchover. Simultaneously, it generates corresponding signals for verification anomalies, allowing for timely replenishment of missing data and correction of read / write errors through standby machine re-reading and writing or manual intervention. This ensures seamless data continuity during primary / standby switchover, guarantees the continuity and accuracy of monitoring data, and provides reliable data support for subsequent monitoring analysis and control decisions. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of the principle framework of the present invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] First Embodiment Please see Figure 1This application provides a fault self-switching device for a high-reliability computer monitoring system, including a real-time monitoring center, a backup processing center, a data feature confirmation center, and a timing verification processing center, wherein the real-time monitoring center, the backup processing center, the data feature confirmation center, and the timing verification processing center are electrically connected sequentially from the output node to the input node. The real-time monitoring center monitors the operating parameters associated with the computer monitoring host in real time and transmits the monitored operating parameters to the backup processing center. These operating parameters include CPU usage, memory usage, network parameters, voltage parameters, and temperature parameters. The switching and standby processing center, based on real-time monitoring of operating parameters, confirms whether different operating items within the operating parameters are in an abnormal fluctuation state, and executes the switching and standby processing process based on the confirmation results to complete the self-switching process of the primary and standby machines. Specifically, when the operating parameters associated with the primary machine fluctuate significantly, some parameter characteristics will deviate from the original parameter curves, which means that the corresponding parameters are in a deviation state and need to be confirmed in time. Then, based on the confirmed comprehensive characteristics, it is comprehensively evaluated whether the primary and standby machines need to perform the self-switching process. The specific method used by the preparation and processing center to confirm whether different operating items are in an abnormal fluctuation state is as follows: Confirm the category to which the operating parameters belong, and extract the fluctuation value B set for the corresponding category. i (Different operating items have different fluctuation values), where i represents different operating items. Based on the real-time generated operating parameters, the fluctuation characteristics of the current time compared to the previous time are determined. The operating parameters corresponding to the operating item at the current time are set as Yx1, and the operating parameters corresponding to the operating item at the previous time are set as Yx2. Their fluctuation characteristics = |Yx1-Yx2|. If the fluctuation characteristics ≤ B i If the fluctuation characteristic is greater than B, then no action is taken. i If the duration of the fluctuation signal exceeds 10 seconds, the current running item is marked as a fluctuation item and the fluctuation duration is recorded synchronously. If the duration of the fluctuation signal does not exceed 10 seconds, no marking is performed. Confirm the volatility percentage associated with the volatility term; its volatility percentage = |Volatility characteristic - B i |÷B i Confirm the fluctuation percentage of the corresponding fluctuation item, and then identify whether there are other fluctuation signals within the recorded fluctuation duration. If they exist, mark the corresponding running item as a following fluctuation item; otherwise, do not mark it. The system uses a confirmation method where fluctuation items are confirmed to have the same fluctuation percentage. It identifies the fluctuation percentage associated with the following fluctuation items, sums up the confirmed fluctuation percentages, and locks the total percentage. If the total percentage is ≥1, the system performs a switchover process, switching the running host to the running standby host. If the total percentage is <1, the system continues to monitor and displays the marked fluctuation items and their following fluctuation items. External operators can use the displayed fluctuation items and their following fluctuation items to perform the host / standby switchover process based on their actual operating experience, thus avoiding excessive load on the host. Specifically, during the monitoring process, there are operating parameters for different operating items. When a certain operating parameter fluctuates, the duration of the fluctuation is used to comprehensively assess whether the corresponding operating item is a fluctuating item. If it is a fluctuating item, the fluctuation of other operating items is confirmed. When a certain operating item fluctuates, other operating items are very likely to be affected by the fluctuation. Therefore, there are multiple different fluctuation processes, and it is necessary to comprehensively confirm the fluctuation processing process of multiple fluctuation items to carry out the comprehensive confirmation process of the primary and backup machines.
[0012] Second Embodiment In the specific implementation process, compared with the above embodiments, this embodiment mainly focuses on the data verification process after the primary and backup machine switching process, identifies whether there is a missing state in the timing data after the switch, and takes timely countermeasures. The data feature confirmation center, based on the switching time of the primary and backup machines, locks down a set of negotiation time intervals and confirms the time-series data generated by the primary machine within the negotiation time interval. It also confirms the time change characteristics associated with different time-series data, locks down the regular characteristics from the confirmed sets of time change characteristics, and records them as primary machine characteristics. Specifically, within the locked negotiation time interval, it is the delay at the switching time and the specific associated time at the front end. The specific associated time is the time corresponding to different time-series data before and after. In order to confirm whether there is data loss during the primary and backup machine switching process, it is necessary to confirm the time-series characteristics generated by the corresponding time-series data between the switching times, and then confirm the time-series characteristics generated by the corresponding time-series data after the switching time, so as to comprehensively evaluate whether the time-series characteristics are consistent. The data feature verification center verifies host features using the following specific methods: Mark the confirmed switching time as T, and based on the set time threshold S1, confirm a set of negotiation time intervals [T-S1, T+S1]; Confirm the timing data generated by the host within the negotiation time interval (that is, the timing data generated between T-S1 and T; between T and T+S1, the host has stopped running, and its standby is running, so the host did not generate corresponding timing data). Confirm the timing associated with the timing data, and generate the timing sequence associated with the timing data according to the chronological order. Confirm the timing difference between adjacent timings within the timing sequence. Confirm the timing difference between several groups of adjacent timings in sequence. Identify the timing difference with the same value from the confirmed groups of timing differences. Confirm the same number of times different timing differences occur. Lock the timing difference associated with the largest number of identical occurrences. Record this timing difference as the host feature. Specifically, the confirmed timing differences of the same number of times may not be a set, but when the host is in normal operation, the timing data generated generally will not have large errors. Therefore, the timing data in normal operation is generated according to the set program. Then, the timing difference with the most identical times is the timing feature with the most obvious pattern, which belongs to the host feature generated by the host.
[0013] The timing verification processing center, based on the confirmed negotiation time interval and the characteristics of the host, confirms the actual timing data generated by the standby machine during the negotiation time interval, locks the actual timing sequence from the actual timing data, generates the standard timing sequence generated by the standby machine based on the characteristics of the host, compares and verifies the actual timing sequence with the standard timing sequence, identifies whether the timing data is incorrect or missing during the host-standby machine switchover, and fills it in. The specific method for comparison and verification is as follows: The confirmed host characteristics are marked as Tz, and a standard time sequence is generated based on the confirmed switching time T. The standard time sequence is {T, T+Tz, T+2Tz, ..., T+nTz}, and all time sequences of the standard time sequence are within the negotiation time interval, and T+nTz≤T+S1, and n is a positive integer. Then, the timing data generated by the standby machine within the negotiation time interval are confirmed in sequence, and based on the timing associated with different timing data, and according to the chronological order, several sets of actual timing sequences associated with the timing data are generated. The system identifies the timing differences associated with adjacent timing sequences within the actual timing sequence, identifies timing differences with the same value from several confirmed sets of timing differences, confirms the number of times different timing differences are the same, locks the timing difference associated with the largest number of the same value, and records the confirmed timing differences as standby characteristics. It then checks whether the primary and standby characteristics are consistent. If they are consistent, it performs a comparison process between the actual timing sequence and the standard timing sequence. If they are inconsistent, it directly generates a data read / write error signal for display. The timing data read / written by the standby machine may be inconsistent with that of the primary machine. In this case, manual intervention is required to correct the timing data being read / written by the standby machine. The system identifies whether the actual time series sequence and the standard time series sequence have the same sorting position. If they are the same, no processing is required. If they are different, the actual time series sequence is recorded as an abnormal sequence and a re-read / write signal is generated. Based on the re-read / write signal, the standby machine reads and writes the subsequent associated time series data in sequence from the switching time to ensure accurate data reading and writing and correct the original reading and writing method. Specifically, actual timing data refers to the timing data associated with the standby machine during actual read and write operations. The associated timing characteristics are the associated actual timing sequences. Based on the marked host characteristics, the standard timing sequence is confirmed. By following the specific timing comparison and verification process, it is possible to effectively confirm whether there are errors in the data read and written by the standby machine, and take timely countermeasures to ensure that no data is lost during the switchover process between the primary and standby machines, thereby improving the overall accuracy of the primary and standby machine switchover process.
[0014] Some of the data in the above formulas are numerical calculations with dimensions removed, and the contents not described in detail in this specification are all prior art known to those skilled in the art.
[0015] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A fault-switching device for a high-reliability computer monitoring system, characterized in that, include: The real-time monitoring center monitors the operating parameters associated with the computer monitoring host in real time. The switching and standby processing center, based on real-time monitored operating parameters, confirms whether different operating items within the operating parameters are in an abnormal fluctuation state, and executes the switching and standby processing process based on the confirmation results to complete the automatic switching process of the primary and standby machines. The data feature confirmation center, based on the switching time of the primary and backup machines, locks down a set of negotiation time intervals, confirms the time series data generated by the primary machine within the negotiation time interval, confirms the time change characteristics associated with different time series data, locks down the regular characteristics from the confirmed sets of time change characteristics, and records them as primary machine characteristics. The timing verification processing center, based on the confirmed negotiation time interval and the characteristics of the primary machine, confirms the actual timing data generated by the standby machine during the negotiation time interval, and locks the actual timing sequence from the actual timing data. Then, based on the characteristics of the primary machine, it generates a standard timing sequence generated by the standby machine. The actual timing sequence is compared and verified with the standard timing sequence to identify whether the timing data is incorrect or missing during the primary-standby machine switchover process, and then corrects it.
2. The fault self-switching device for a high-reliability computer monitoring system according to claim 1, characterized in that, The operating parameters include CPU usage, memory usage, network parameters, voltage parameters, and temperature parameters.
3. The fault self-switching device for a high-reliability computer monitoring system according to claim 1, characterized in that, The specific method by which the switching processing center performs the primary / standby automatic switchover process is as follows: Confirm the category to which the operating parameters belong, and extract the fluctuation value B set for the corresponding category. i Where i represents different running items, based on the real-time generated running parameters, the fluctuation characteristics of the current time compared to the previous time are determined, and the running parameters corresponding to the running items at the current time are set as Yx1, and the running parameters corresponding to the running items at the previous time are set as Yx2. Their fluctuation characteristics = |Yx1-Yx2|, if the fluctuation characteristics ≤ B i If the fluctuation characteristic is greater than B, then no action is taken. i If the duration of the fluctuation signal exceeds 10 seconds, the current running item is marked as a fluctuation item and the fluctuation duration is recorded synchronously. If the duration of the fluctuation signal does not exceed 10 seconds, no marking is performed. Confirm the volatility percentage associated with the volatility term; its volatility percentage = |Volatility characteristic - B i |÷B i Confirm the fluctuation percentage of the corresponding fluctuation item, and then identify whether there are other fluctuation signals within the recorded fluctuation duration. If so, directly mark the corresponding running item as a follower fluctuation item. The method of confirming the same fluctuation percentage for fluctuation items is adopted to identify the fluctuation percentage associated with the fluctuation item, and the confirmed fluctuation percentages are summed to lock the total percentage. If the total percentage is ≥1, the switchover process is executed to switch the running host to the running standby host.
4. The fault self-switching device for a high-reliability computer monitoring system according to claim 3, characterized in that, If no other fluctuation signals exist within the recorded fluctuation duration, no marking is made.
5. The fault self-switching device for a high-reliability computer monitoring system according to claim 3, characterized in that, If the total percentage is less than 1, we will continue to monitor and simultaneously display the marked fluctuation items and the following fluctuation items.
6. The fault self-switching device for a high-reliability computer monitoring system according to claim 1, characterized in that, The data feature verification center verifies host features in the following specific way: Mark the confirmed switching time as T, and based on the set time threshold S1, confirm a set of negotiation time intervals [T-S1, T+S1]; Confirm the timing data generated by the host within the negotiation time interval, confirm the timing associated with the timing data, generate a timing sequence associated with the timing data according to the chronological order, confirm the timing difference between adjacent timings within the timing sequence, confirm the timing difference between several groups of adjacent timings in sequence, identify the timing difference with the same value from the confirmed groups of timing difference, confirm the same number of times different timing differences occur, lock the timing difference associated with the largest number of identical occurrences, and record this timing difference as the host feature.
7. The fault self-switching device for a high-reliability computer monitoring system according to claim 1, characterized in that, The timing verification processing center compares and verifies the actual timing sequence with the standard timing sequence in the following specific way: The confirmed host characteristics are marked as Tz, and a standard time sequence is generated based on the confirmed switching time T. The standard time sequence is {T, T+Tz, T+2Tz, ..., T+nTz}, and all time sequences of the standard time sequence are within the negotiation time interval, and T+nTz≤T+S1, and n is a positive integer. Then, the timing data generated by the standby machine within the negotiation time interval are confirmed in sequence, and based on the timing associated with different timing data, and according to the chronological order, several sets of actual timing sequences associated with the timing data are generated. Identify the timing differences associated with adjacent timing values within the actual timing sequence, and from the confirmed sets of timing differences, identify timing differences with the same value, and confirm the number of times different timing differences are the same, lock the timing difference associated with the largest number of times the same value is the same, and record the confirmed timing differences as standby characteristics, identify whether the primary characteristics and standby characteristics are consistent, and if they are consistent, execute the comparison process between the actual timing sequence and the standard timing sequence; The system identifies whether the actual time series sequence and the standard time series sequence have the same sorting position. If they are the same, no processing is required. If they are different, the actual time series sequence is recorded as an abnormal sequence and a re-read / write signal is generated. Based on the re-read / write signal, the standby machine reads and writes the subsequent associated time series data in sequence, starting from the switching time.
8. The fault self-switching device for a high-reliability computer monitoring system according to claim 7, characterized in that, If the characteristics of the primary machine and the standby machine are inconsistent, a data read / write error signal will be generated and displayed directly.