Master-slave system switching method and device, equipment and storage medium

By analyzing the heartbeat packets and operation logs of the master-slave system from multiple dimensions, calculating the fault risk value and prioritizing detection, the problem of misjudgment caused by single-dimensional detection is solved, and accurate master-slave system switching is achieved, ensuring system stability.

CN120856544APending Publication Date: 2025-10-28创优数字科技(广东)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511090173.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing master-slave system switching methods, single-dimensional heartbeat detection is prone to high false alarm rates, leading to unnecessary master-slave switching and affecting business stability.

Method used

By acquiring the synchronization heartbeat packets and operation logs between the master and slave systems, multi-dimensional analysis is performed, including CPU utilization, critical service response time, heartbeat detection information, and error log information. The fault risk value is calculated, and priority detection is performed to determine the optimal slave system.

Benefits of technology

It achieves precise master-slave system switching, reduces the false alarm rate, and ensures system stability and business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856544A_ABST
    Figure CN120856544A_ABST
Patent Text Reader

Abstract

The invention discloses a master-slave system switching method and apparatus, a device and a storage medium. The method comprises the steps of obtaining a synchronous heartbeat packet between a master system and each slave system, a running log of the master system and a running log of each slave system; performing real-time analysis on the operation log of the master system and the operation log of each slave system to determine abnormal information; judging whether the main system has a fault risk based on the heartbeat packet and the abnormal information; if yes, performing priority detection on each slave system to determine an optimal slave system; and performing master-slave switching on the master system and the optimal slave system. Compared with a mode only based on heartbeat detection in the prior art, detection can be carried out from multiple dimensions, the high misjudgment rate caused by single-dimension detection is prevented, whether the main system has risks or not can be accurately judged, priority detection is carried out to determine the optimal slave system, accurate switching of the slave systems is achieved, and the switching efficiency of the slave systems is improved. Unnecessary master-slave switching is not caused, and the system stability is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of master-slave system switching technology, specifically to a master-slave system switching method, apparatus, device, and storage medium. Background Technology

[0002] Master-slave system architecture is widely used in various scenarios such as databases and server clusters. In a master-slave system architecture, the master system is responsible for processing business requests, executing core operations, and synchronizing data to the slave system. The slave system serves as a backup, taking over the work when the master system fails, in order to ensure business continuity.

[0003] Currently, the most common master-slave system switching methods are based on heartbeat detection mechanisms. For example, if the slave system does not receive a response from the master system within a certain period of time, it is determined that the master system is faulty and the master-slave switching process is triggered. However, this single-dimensional detection can lead to a high false alarm rate, causing unnecessary master-slave switching and affecting business stability. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus, device and storage medium for switching master-slave systems to solve the problem that single-dimensional detection can lead to a high false positive rate, causing unnecessary master-slave switching and affecting business stability.

[0005] To achieve the above objectives, the following solution is proposed:

[0006] Firstly, a method for switching between master and slave systems includes:

[0007] Obtain the heartbeat packets synchronized between the master system and each slave system, the operation log of the master system, and the operation log of each slave system;

[0008] The operation logs of the main system and each of the slave systems are analyzed in real time to identify abnormal information.

[0009] Based on the heartbeat packets and abnormal information, determine whether the main system has a risk of failure;

[0010] If so, priority detection is performed on each of the slave systems to determine the optimal slave system;

[0011] Perform a master-slave switch between the master system and the optimal slave system.

[0012] Preferably, determining whether the main system has a risk of failure based on the heartbeat packet and abnormal information includes:

[0013] Extract CPU utilization, response time of each critical service, and heartbeat detection information from the heartbeat packets;

[0014] Extract error log information from the anomaly information;

[0015] Calculate the first fault sub-score corresponding to the CPU utilization, the second fault sub-score corresponding to the response time of each critical service, the third fault sub-score corresponding to the heartbeat detection information, and the fourth fault sub-score corresponding to the error log information.

[0016] Set the weights for CPU utilization, response time of each critical service, heartbeat detection information, and error log information respectively;

[0017] The first fault sub-score, the second fault sub-score, the third fault sub-score, and the fourth fault sub-score are weighted and summed according to their respective weights to obtain the fault risk value of the main system.

[0018] The fault risk value is compared with a preset risk threshold.

[0019] If the fault risk value is greater than the preset risk threshold, then the main system is determined to have a fault risk.

[0020] Preferably, the calculation of the first fault sub-score corresponding to the CPU utilization rate includes:

[0021] Determine whether the CPU utilization rate continuously exceeds a preset utilization threshold within a preset first time period;

[0022] If so, then determine the continuous time;

[0023] Divide the continuous time by the first time period to obtain the continuous timeout ratio;

[0024] The first fault sub-score is determined based on the continuous timeout ratio.

[0025] Preferably, the process of calculating the second fault sub-score corresponding to the response time of each critical service includes:

[0026] Calculate the average response time based on the response time of each critical service described.

[0027] If the average response time exceeds a preset response time threshold, the average response time is divided by the response time threshold to obtain the timeout response ratio.

[0028] The second fault sub-score is determined according to the timeout response ratio.

[0029] Preferably, the process of calculating the third fault sub-score corresponding to the heartbeat detection information includes:

[0030] Extract the time interval between every two heartbeats from the heartbeat detection information;

[0031] Determine whether the time interval between two heartbeats exceeds a preset interval threshold.

[0032] If there are N consecutive time intervals exceeding the interval threshold, then the number of consecutive times is determined, and the proportion of times exceeding the threshold is calculated based on the number of consecutive times.

[0033] The third fault sub-score of the heartbeat detection information is calculated based on the proportion of times the heartbeat exceeds the specified number.

[0034] Preferably, calculating the fourth fault sub-score corresponding to the error log information includes:

[0035] Determine the surge in the number of error logs within a preset second time period from the error log information;

[0036] If the number of surges exceeds a preset threshold, then the number of surges is calculated to exceed a certain proportion based on the threshold.

[0037] The fourth fault sub-score of the error log information is calculated based on the number exceeding the proportion.

[0038] Preferably, the step of prioritizing each of the slave systems to determine the optimal slave system includes:

[0039] For each of the slave systems, the data synchronization delay time of that slave system is calculated based on the anomaly information to generate a first priority score;

[0040] Calculate the load utilization of the system to generate a second priority score;

[0041] Calculate the handover success rate of the system within the first historical time period to generate a third priority score;

[0042] Calculate the network quality score of the system;

[0043] Set coefficients for the data synchronization delay time, load utilization, handover success rate, and network quality score, respectively;

[0044] The first priority score, the second priority score, the third priority score, and the fourth priority score are weighted and summed according to their respective coefficients to obtain the target priority score;

[0045] The slave system with the highest target priority score is selected as the optimal slave system.

[0046] Secondly, a master-slave system switching device includes:

[0047] The acquisition module is used to acquire the heartbeat packets synchronized between the master system and each slave system, the operation log of the master system, and the operation log of each slave system.

[0048] An anomaly information determination module is used to perform real-time analysis on the operation logs of the main system and the operation logs of each slave system to determine anomaly information.

[0049] The judgment module is used to determine whether there is a risk of failure in the main system based on the heartbeat packet and abnormal information;

[0050] The priority detection module is used to perform priority detection on each of the slave systems if the condition is met, so as to determine the optimal slave system.

[0051] The master-slave switching module is used to switch the master system and the optimal slave system between master and slave.

[0052] Thirdly, a master-slave system switching device, including a memory and a processor;

[0053] The memory is used to store programs;

[0054] The processor is configured to execute the program to implement the steps of the master-slave system switching method as described in any of the first aspects.

[0055] Fourthly, a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the master-slave system switching method as described in any of the first aspects.

[0056] As can be seen from the above technical solution, this application obtains the heartbeat packets synchronized between the master system and each slave system, the operation logs of the master system, and the operation logs of each slave system; it performs real-time analysis on the operation logs of the master system and each slave system to identify abnormal information; based on the heartbeat packets and abnormal information, it determines whether the master system has a fault risk; if so, it performs priority detection on each slave system to determine the optimal slave system; and it performs a master-slave switch between the master system and the optimal slave system. This application obtains heartbeat packets, the operation logs of the master and slave systems, and other multi-dimensional information for real-time analysis to identify abnormal information. Compared with the existing technology that only relies on heartbeat detection, it can perform detection from multiple dimensions, preventing the problem of high false positive rates caused by single-dimensional detection. Therefore, it can accurately determine whether the master system has a risk and perform priority detection to determine the optimal slave system, achieving accurate slave system switching without triggering unnecessary master-slave switches and ensuring system stability. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0058] Figure 1 An optional flowchart of a master-slave system switching method provided in an embodiment of this application;

[0059] Figure 2 This application provides a schematic diagram of a master-slave system switching process.

[0060] Figure 3 A schematic diagram of the structure of a master-slave system switching device provided in an embodiment of this application;

[0061] Figure 4 This is a schematic diagram of the structure of a master-slave system switching device provided in an embodiment of this application. Detailed Implementation

[0062] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0063] This invention can be used in a wide variety of general-purpose or special-purpose computing environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor devices, distributed computing environments including any of the above devices, etc.

[0064] This invention provides a master-slave system switching method, which can be applied to various computer terminals or smart terminals. The executing entity can be the processor or server of the computer terminal or smart terminal. The method flowchart is shown below. Figure 1 As shown, specifically including:

[0065] S1: Obtain the heartbeat packets synchronized between the master system and each slave system, the master system's operation log, and the operation logs of each slave system.

[0066] The master-slave system switching method provided in this application can be applied to a monitoring center. Taking a database master-slave architecture as an example, the slave system periodically sends heartbeat packets to the master system. After receiving the packets, the master system returns a response. If the slave system does not receive a response from the master system within a certain period of time (e.g., three consecutive heartbeat timeouts), the master system is deemed to be faulty, triggering the master-slave system switching process. During the switch, the slave system first updates its own data to be as consistent as possible with the master system, such as by copying logs that the master system has not yet synchronized. Then, it promotes itself to the master system to provide services to the outside world. At the same time, the new slave system is started and connected to the new master system, and the master-slave relationship is re-established.

[0067] Unlike traditional heartbeat detection mechanisms, the heartbeat packets obtained in this application carry other information besides the heartbeat identifier, such as key indicator data of the main system, including CPU utilization, response time of each key service, heartbeat detection information, request interface status codes, etc. This information can be sent to the monitoring center for monitoring and analysis.

[0068] The operation logs of the main system and each slave system include log information, operation records, etc.

[0069] This data can then be comprehensively analyzed to enable master-slave failover. This application can acquire heartbeat packets synchronized between the master system and each slave system, the master system's operation log, and the operation logs of each slave system every minute for subsequent calculations and failover decisions, ensuring real-time performance.

[0070] S2: Perform real-time analysis on the operation logs of the main system and each of the slave systems to identify abnormal information.

[0071] Real-time analysis can identify anomalies, such as the number of abnormal logs and abnormal operation records. Similarly, this data can also be sent to the monitoring center.

[0072] S3: Determine whether there is a risk of failure in the main system based on the heartbeat packet and abnormal information.

[0073] It is understandable that some data can be extracted from the abnormal information of the heartbeat and the main system. This data can determine whether there is a fault in the main system. If there is a fault, the master-slave system switching mechanism can be triggered.

[0074] Because heartbeat packets and abnormal information contain multi-dimensional information, they can cover comprehensive, critical, and necessary information that indicates whether the main system is malfunctioning, thereby enabling multi-dimensional comprehensive judgment and preventing misjudgment.

[0075] For example, in the database system of an e-commerce platform, the master database handles user operations such as order placement and payment, and synchronizes data changes to the slave database. The slave database can be used for auxiliary business such as data analysis and report generation. At the same time, it can switch to the master database and continue to provide services when the master database fails. However, the existing switching method that relies solely on the heartbeat detection mechanism has many drawbacks. For example, network fluctuations and short-term high system loads may cause heartbeat packets to be lost or delayed, which may cause the slave database to misjudge the master database as failing, triggering unnecessary master-slave switching and affecting the stability of e-commerce business. For example, during periods of network congestion, even if the master database is running normally, the slave database may switch incorrectly because it does not receive heartbeat packets.

[0076] Therefore, the multi-dimensional detection mechanism provided in this application can comprehensively evaluate the actual operating status of the main system / main database. For example, although the main system can respond to heartbeats normally, some critical services may be abnormal and unable to process business requests normally. In this case, the switching method provided in this application can detect this and will not trigger the switching mechanism.

[0077] S4: If so, perform priority detection on each of the slave systems to determine the optimal slave system.

[0078] If a comprehensive analysis and monitoring reveals that the main system does indeed have a risk of failure, a switching mechanism can be triggered. This judgment method is both accurate and comprehensive, thereby determining an optimal slave system and using this optimal slave system as the new main system.

[0079] S5: Perform a master-slave switch between the master system and the optimal slave system.

[0080] After determining the optimal slave system, considering the business stability and the stability of the master-slave system after the new master-slave relationship is established, an optimal slave system is selected as the new master system. In addition, a new slave system corresponding to the new master system is determined, and a new master-slave system is reorganized to ensure stability and low failure risk.

[0081] To achieve autonomy and flexibility, after determining that the main system has a risk of failure and the optimal slave system is determined, an early warning notification can be sent to the operation and maintenance personnel for confirmation before the switchover. This can be configured to be executed after manual confirmation or automatically.

[0082] As can be seen from the above technical solution, this application obtains the heartbeat packets synchronized between the master system and each slave system, the operation logs of the master system, and the operation logs of each slave system; performs real-time analysis on the operation logs of the master system and each slave system to identify abnormal information; determines whether the master system has a fault risk based on the heartbeat packets and abnormal information; if so, performs priority detection on each slave system to determine the optimal slave system; and performs a master-slave switch between the master system and the optimal slave system. This application obtains heartbeat packets, the operation logs of the master and slave systems, and other multi-dimensional information for real-time analysis to identify abnormal information. Compared to the existing technology that only relies on heartbeat detection, this method can detect from multiple dimensions, preventing the high false positive rate caused by single-dimensional detection. This allows for accurate determination of whether the master system has a risk and performs priority detection to determine the optimal slave system, achieving precise slave system switching without triggering unnecessary master-slave switches. This ensures that the master and slave systems can quickly and accurately complete the switch under various complex conditions, guaranteeing system stability.

[0083] The method provided in this embodiment of the invention, which determines whether the main system has a risk of failure based on the heartbeat packet and abnormal information, is described in detail below:

[0084] Extract CPU utilization, response time of each critical service, and heartbeat detection information from the heartbeat packets;

[0085] Extract error log information from the anomaly information;

[0086] Calculate the first fault sub-score corresponding to the CPU utilization, the second fault sub-score corresponding to the response time of each critical service, the third fault sub-score corresponding to the heartbeat detection information, and the fourth fault sub-score corresponding to the error log information.

[0087] Set the weights for CPU utilization, response time of each critical service, heartbeat detection information, and error log information respectively;

[0088] The first fault sub-score, the second fault sub-score, the third fault sub-score, and the fourth fault sub-score are weighted and summed according to their respective weights to obtain the fault risk value of the main system.

[0089] The fault risk value is compared with a preset risk threshold.

[0090] If the fault risk value is greater than the preset risk threshold, then the main system is determined to have a fault risk.

[0091] Specifically, this application assesses the main system for potential failure risks from multiple dimensions. Information from various dimensions, such as CPU utilization, response time of critical services, and heartbeat detection information, can be extracted from the heartbeat packets. This information can reflect the main system's operation from multiple levels and angles. Furthermore, standard indicators can be preset or acquired for comparison to determine whether these data are abnormal or have problems.

[0092] In addition to extracting data from heartbeat packets, other information different from heartbeat packets can also be extracted from confirmed abnormal information, such as error log information. Error log information contains log information corresponding to the abnormality in the main system or slave system, which is more comprehensive and can determine exactly when and what content caused the abnormality.

[0093] This multi-dimensional monitoring mechanism comprehensively judges the status of the main system from multiple perspectives, avoiding misjudgments caused by single factors such as network fluctuations, making master-slave switching more accurate and reliable, and improving business stability. At the same time, the multi-dimensional monitoring mechanism can promptly detect problems that are difficult to detect by traditional solutions, such as anomalies in some critical services in the main system, triggering switching earlier and ensuring normal business operation.

[0094] When jointly analyzing this data, it can be quantified by scores to calculate a total score, which will serve as the final evaluation criterion. This involves calculating the first fault sub-score corresponding to CPU utilization, the second fault sub-score corresponding to the response time of each critical service, the third fault sub-score corresponding to heartbeat detection information, and the fourth fault sub-score corresponding to error log information. Weights are then assigned, and a weighted sum is calculated to obtain the fault risk value of the main system. This fault risk value is compared with a preset risk threshold. If the fault risk value is greater than the risk threshold, the main system is considered to have a fault risk. If the fault risk value is not greater than the risk threshold, the main system is considered to have no fault risk (more precisely, there is currently no fault risk).

[0095] This application does not limit the form or format of the fault risk value, which can be a percentage system, a normalized system or a proportional system, and the risk threshold can be set to 70 points, 0.7 points, etc.

[0096] Using a score-based method to determine whether the main system has a risk of failure is more intuitive, simple, and greatly improves the efficiency of the assessment.

[0097] The process of calculating the first fault sub-score corresponding to the CPU utilization rate is explained in detail below.

[0098] Determine whether the CPU utilization rate continuously exceeds a preset utilization threshold within a preset first time period;

[0099] If so, then determine the continuous time;

[0100] Divide the continuous time by the first time period to obtain the continuous timeout ratio;

[0101] The first fault sub-score is determined based on the continuous timeout ratio.

[0102] Specifically, since CPU utilization is a critical and important indicator, the utilization threshold can be set to a high value, such as 90%, and the first time period can also be set to a long time, such as 5 minutes, 10 minutes, etc. If the CPU utilization continuously exceeds the utilization threshold within the first time period, it is considered abnormal. However, a score needs to be calculated, so a certain percentage exceeding the threshold needs to be determined as the basis for the score calculation. This means determining the continuous time, which must be greater than the length of the first time period. Therefore, the continuous time can be divided by the first time period to obtain the percentage exceeding the first time period, which is the continuous timeout ratio. The larger the continuous timeout ratio, the more abnormal the CPU utilization. Thus, the first fault score can be determined based on the continuous timeout ratio. Specifically, in one example, a table corresponding to the continuous timeout ratio and the first fault score can be pre-set, as shown in Table 1 below:

[0103] Table 1

[0104]

[0105] The following embodiments provide a detailed explanation of the steps involved in calculating the second fault sub-score corresponding to the response time of each critical service.

[0106] Calculate the average response time based on the response time of each critical service described.

[0107] If the average response time exceeds a preset response time threshold, the average response time is divided by the response time threshold to obtain the timeout response ratio.

[0108] The second fault sub-score is determined according to the timeout response ratio.

[0109] Specifically, the response time of critical services also represents the response speed and efficiency of the main system to a certain extent. Therefore, this application also uses the response time of critical services as an indicator for judging the risk of failure. Since there will be multiple responses of critical services, in order to make accurate judgments, an average value is calculated. That is, the response time of each critical service is added together and divided by the number to obtain the average response time. Similarly, the average response time is divided by a preset response time threshold to obtain the proportion of the average response time exceeding the response time threshold, and the second failure sub-score is calculated accordingly.

[0110] In one example, the critical service is HTTP requests, and the response time threshold is set to 5 seconds.

[0111] Furthermore, the process of calculating the third fault sub-score corresponding to the heartbeat detection information may include the following steps:

[0112] Extract the time interval between every two heartbeats from the heartbeat detection information;

[0113] Determine whether the time interval between two heartbeats exceeds a preset interval threshold.

[0114] If there are N consecutive time intervals exceeding the interval threshold, then the number of consecutive times is determined, and the proportion of times exceeding the threshold is calculated based on the number of consecutive times.

[0115] The third fault sub-score of the heartbeat detection information is calculated based on the proportion of times the heartbeat exceeds the specified number.

[0116] Specifically, the time interval between two heartbeats is called the heartbeat interval. It is the time interval between two consecutive heartbeat signals sent between the master system and the slave system. It is one of the core parameters in the heartbeat detection mechanism and directly affects the fault detection speed and resource consumption of the master system. The standard value (interval threshold) is generally a few seconds, such as 3 seconds. If the time interval is greater than 3 seconds for multiple consecutive times, it is considered to be abnormal. Similarly, the proportion of times the interval threshold is exceeded is calculated to obtain the timeout ratio, which is used to obtain the third fault sub-score.

[0117] Optionally, the process of calculating the fourth fault sub-score corresponding to the error log information may include the following steps:

[0118] Determine the surge in the number of error logs within a preset second time period from the error log information;

[0119] If the number of surges exceeds a preset threshold, then the number of surges is calculated to exceed a certain proportion based on the threshold.

[0120] The fourth fault sub-score of the error log information is calculated based on the number exceeding the proportion.

[0121] Specifically, this application uses the surge frequency of error logs as a benchmark for judgment. That is, a second time period is set, and the number of surges in error logs during the second time period is compared with a preset threshold. If the number of surges exceeds the threshold, it is determined that there is an anomaly. Then, the anomaly ratio is determined. Similarly, the number of surges is divided by the threshold to obtain the ratio of the number of surges, and the fourth fault sub-score is determined according to the ratio of the number of surges.

[0122] The sub-scores for the first, second, third, and fourth faults can all be pre-set and matched according to the corresponding over-limit ratio, over-time ratio, etc., or specific scores can be calculated. For example, if a standard value of 100 is set, the standard value can be divided by the over-limit ratio or over-time ratio to obtain the corresponding sub-score.

[0123] Furthermore, the process of prioritizing each of the slave systems described in this application to determine the optimal slave system is as follows:

[0124] For each of the slave systems, the data synchronization delay time of that slave system is calculated based on the anomaly information to generate a first priority score;

[0125] Calculate the load utilization of the system to generate a second priority score;

[0126] Calculate the handover success rate of the system within the first historical time period to generate a third priority score;

[0127] Calculate the network quality score of the system;

[0128] Set coefficients for the data synchronization delay time, load utilization, handover success rate, and network quality score, respectively;

[0129] The first priority score, the second priority score, the third priority score, and the fourth priority score are weighted and summed according to their respective coefficients to obtain the target priority score;

[0130] The slave system with the highest target priority score is selected as the optimal slave system.

[0131] Specifically, the historical handover success rate refers to the probability of a successful handover in previous master-slave handover processes. For example, a count of 10 can be set, and the number of successful handovers within those 10 can be extracted. Dividing this count by 10 gives the handover success rate. System load utilization refers to the combined usage of CPU and memory. The network quality score is calculated based on Ping latency and packet loss rate. Ping latency is the time required for a data packet to travel from the source device to the target device and return. It is an important indicator for measuring network response speed and connection quality. Packet loss rate is the proportion of data packets that fail to reach the target device during network transmission. It is one of the important indicators for measuring network stability and reliability, usually expressed as a percentage. The weighted sum of the two yields the network quality score. The weights of Ping latency and packet loss rate can be set according to actual conditions or user needs. This embodiment does not impose any restrictions on this.

[0132] It is understandable that the time it takes for the system to synchronize data A minus the time it takes for data A to be updated in the main system is the data synchronization delay time. Therefore, a first priority score can be generated based on the data synchronization delay time. The longer the data synchronization delay time, the smaller the first priority score, and the two are inversely proportional. Similarly, the higher the load utilization rate, the smaller the second priority score, and the two are inversely proportional. Similarly, the lower the switchover success rate of the system in the first historical time period, the lower the third priority score, and the two are directly proportional. Similarly, network quality scores can be used to calculate the target priority score.

[0133] Based on the above direct or inverse proportional scenarios, calculate their respective priority scores. The formula for calculating the target priority score is as follows:

[0134] ;

[0135] in, Indicates the target priority score; Indicates the first priority score. Indicates the second priority score. Indicates the third priority score. This indicates the fourth priority score; , , , These are the coefficients for the first priority score, the second priority score, the third priority score, and the fourth priority score, respectively.

[0136] In one example, the switching process between master-slave architecture systems is as follows: Figure 2 As shown, one master system corresponds to two slave systems. The master system synchronizes data with the two slave systems respectively. The monitoring center is responsible for data monitoring, analysis and decision-making, such as analyzing how to perform master-slave switching, which slave system to select as the optimal slave system, etc. After the decision is made, a switching command needs to be sent.

[0137] and Figure 1 Corresponding to the method described above, embodiments of the present invention also provide a master-slave system switching device for switching between master and slave systems. Figure 1 In the specific implementation of the method, the master-slave system switching device provided in this embodiment of the invention can be integrated into a computer terminal or various mobile devices. Figure 3 The switching mechanism of the master-slave system is introduced, such as... Figure 3 As shown, the device may include:

[0138] The acquisition module 10 is used to acquire the heartbeat packets synchronized between the master system and each slave system, the operation log of the master system, and the operation log of each slave system;

[0139] The anomaly information determination module 20 is used to perform real-time analysis on the operation logs of the main system and the operation logs of each slave system to determine anomaly information.

[0140] The judgment module 30 is used to determine whether there is a risk of failure in the main system based on the heartbeat packet and abnormal information;

[0141] The priority detection module 40 is used to perform priority detection on each of the slave systems if the condition is met, so as to determine the optimal slave system.

[0142] The master-slave switching module 50 is used to switch the master system and the optimal slave system between master and slave.

[0143] As can be seen from the above technical solution, this application obtains the heartbeat packets synchronized between the master system and each slave system, the operation logs of the master system, and the operation logs of each slave system; it performs real-time analysis on the operation logs of the master system and each slave system to identify abnormal information; based on the heartbeat packets and abnormal information, it determines whether the master system has a fault risk; if so, it performs priority detection on each slave system to determine the optimal slave system; and it performs a master-slave switch between the master system and the optimal slave system. This application obtains heartbeat packets, the operation logs of the master and slave systems, and other multi-dimensional information for real-time analysis to identify abnormal information. Compared with the existing technology that only relies on heartbeat detection, it can perform detection from multiple dimensions, preventing the problem of high false positive rates caused by single-dimensional detection. Therefore, it can accurately determine whether the master system has a risk and perform priority detection to determine the optimal slave system, achieving accurate slave system switching without triggering unnecessary master-slave switches and ensuring system stability.

[0144] Furthermore, embodiments of this application provide a switching device for a master-slave system. Optionally, Figure 4 The hardware structure block diagram of the master-slave system switching device is shown. (Refer to...) Figure 4 The hardware structure of the master-slave system switching device may include: at least one processor 01, at least one communication interface 02, at least one memory 03 and at least one communication bus 04.

[0145] In this embodiment, the number of processor 01, communication interface 02, memory 03 and communication bus 04 is at least one, and processor 01, communication interface 02 and memory 03 communicate with each other through communication bus 04.

[0146] Processor 01 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0147] Memory 03 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device.

[0148] The memory stores a program, which the processor can call to execute the master-slave system switching method, including:

[0149] Obtain the heartbeat packets synchronized between the master system and each slave system, the operation log of the master system, and the operation log of each slave system;

[0150] The operation logs of the main system and each of the slave systems are analyzed in real time to identify abnormal information.

[0151] Based on the heartbeat packets and abnormal information, determine whether the main system has a risk of failure;

[0152] If so, priority detection is performed on each of the slave systems to determine the optimal slave system;

[0153] Perform a master-slave switch between the master system and the optimal slave system.

[0154] Optionally, the refined and extended functions of the program can be found in the description of the master-slave system switching method in the method embodiment.

[0155] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor. When the program runs, it controls the device where the storage medium is located to execute the following master-slave system switching method, including:

[0156] Obtain the heartbeat packets synchronized between the master system and each slave system, the operation log of the master system, and the operation log of each slave system;

[0157] The operation logs of the main system and each of the slave systems are analyzed in real time to identify abnormal information.

[0158] Based on the heartbeat packets and abnormal information, determine whether the main system has a risk of failure;

[0159] If so, priority detection is performed on each of the slave systems to determine the optimal slave system;

[0160] Perform a master-slave switch between the master system and the optimal slave system.

[0161] Specifically, the storage medium can be a computer-readable storage medium, which can be an electronic storage device such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM.

[0162] Optionally, the refined and extended functions of the program can be found in the description of the master-slave system switching method in the method embodiment.

[0163] Furthermore, the functional modules in the various embodiments of this disclosure can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a live streaming device, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this disclosure.

[0164] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0165] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0166] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for switching between master and slave systems, characterized in that, include: Obtain the heartbeat packets synchronized between the master system and each slave system, the operation log of the master system, and the operation log of each slave system; The operation logs of the main system and each of the slave systems are analyzed in real time to identify abnormal information. Based on the heartbeat packets and abnormal information, determine whether the main system has a risk of failure; If so, priority detection is performed on each of the slave systems to determine the optimal slave system; Perform a master-slave switch between the master system and the optimal slave system.

2. The method according to claim 1, characterized in that, The determination of whether the main system has a failure risk based on the heartbeat packets and abnormal information includes: Extract CPU utilization, response time of each critical service, and heartbeat detection information from the heartbeat packets; Extract error log information from the anomaly information; Calculate the first fault sub-score corresponding to the CPU utilization, the second fault sub-score corresponding to the response time of each critical service, the third fault sub-score corresponding to the heartbeat detection information, and the fourth fault sub-score corresponding to the error log information. Set the weights for CPU utilization, response time of each critical service, heartbeat detection information, and error log information respectively; The first fault sub-score, the second fault sub-score, the third fault sub-score, and the fourth fault sub-score are weighted and summed according to their respective weights to obtain the fault risk value of the main system. The fault risk value is compared with a preset risk threshold. If the fault risk value is greater than the preset risk threshold, then the main system is determined to have a fault risk.

3. The method according to claim 2, characterized in that, The calculation of the first fault sub-score corresponding to the CPU utilization rate includes: Determine whether the CPU utilization rate continuously exceeds a preset utilization threshold within a preset first time period; If so, then determine the continuous time; Divide the continuous time by the first time period to obtain the continuous timeout ratio; The first fault sub-score is determined based on the continuous timeout ratio.

4. The method according to claim 2, characterized in that, The process of calculating the second fault sub-score corresponding to the response time of each critical service includes: Calculate the average response time based on the response time of each critical service described. If the average response time exceeds a preset response time threshold, the average response time is divided by the response time threshold to obtain the timeout response ratio. The second fault sub-score is determined according to the timeout response ratio.

5. The method according to claim 2, characterized in that, The process of calculating the third fault sub-score corresponding to the heartbeat detection information includes: Extract the time interval between every two heartbeats from the heartbeat detection information; Determine whether the time interval between two heartbeats exceeds a preset interval threshold. If there are N consecutive time intervals exceeding the interval threshold, then the number of consecutive times is determined, and the proportion of times exceeding the threshold is calculated based on the number of consecutive times. The third fault sub-score of the heartbeat detection information is calculated based on the proportion of times the heartbeat exceeds the specified number.

6. The method according to claim 2, characterized in that, Calculating the fourth fault sub-score corresponding to the error log information includes: Determine the surge in the number of error logs within a preset second time period from the error log information; If the number of surges exceeds a preset threshold, then the number of surges is calculated to exceed a certain proportion based on the threshold. The fourth fault sub-score of the error log information is calculated based on the number exceeding the proportion.

7. The method according to any one of claims 1 to 6, characterized in that, The step of prioritizing each slave system to determine the optimal slave system includes: For each of the slave systems, the data synchronization delay time of that slave system is calculated based on the anomaly information to generate a first priority score; Calculate the load utilization of the system to generate a second priority score; Calculate the handover success rate of the system within the first historical time period to generate a third priority score; Calculate the network quality score of the system; Set coefficients for the data synchronization delay time, load utilization, handover success rate, and network quality score, respectively; The first priority score, the second priority score, the third priority score, and the fourth priority score are weighted and summed according to their respective coefficients to obtain the target priority score; The slave system with the highest target priority score is selected as the optimal slave system.

8. A switching device for a master-slave system, characterized in that, include: The acquisition module is used to acquire the heartbeat packets synchronized between the master system and each slave system, the operation log of the master system, and the operation log of each slave system. An anomaly information determination module is used to perform real-time analysis on the operation logs of the main system and the operation logs of each slave system to determine anomaly information. The judgment module is used to determine whether there is a risk of failure in the main system based on the heartbeat packet and abnormal information; The priority detection module is used to perform priority detection on each of the slave systems if the condition is met, so as to determine the optimal slave system. The master-slave switching module is used to switch the master system and the optimal slave system between master and slave.

9. A master-slave system switching device, characterized in that, Including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the master-slave system switching method as described in any one of claims 1-7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the master-slave system switching method as described in any one of claims 1-7.