Site switching method and device, electronic equipment and storage medium
By monitoring multiple metrics of the main site in real time, quickly identifying faults and switching to alternative sites in batches, the problem of long service interruption time after cloud host failure is solved, and high availability and business continuity is achieved.
Patent Information
- Application Number
- CN202510570315.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, when a cloud host switches to an alternative site after a failure, the service interruption time is long and the risk of data loss is high, making it difficult to meet real-time and automation needs.
By monitoring multiple metrics of the main site in real time, quickly identify faults, and switch sessions to alternative sites in batches based on exception metrics and delay priorities to ensure business continuity.
It realizes rapid identification and switching of faulty sites, ensures high availability and business continuity of services, and reduces the risk of data loss.
Smart Images

Figure CN120358135A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of cloud computing, and in particular, to a method, device, electronic device and storage medium for switching sites. Background Art
[0002] With the rapid development of cloud computing technology, cloud hosts, as an important part of enterprise infrastructure, their stability and reliability have attracted much attention. In a large-scale distributed system, a single point of failure may cause service interruption and affect business continuity. Therefore, in a large-scale distributed system, usually in addition to the cloud host (i.e., the primary site), there are other alternative hosts or disaster recovery hosts (i.e., alternative sites), so that after the primary site fails, it can be switched to the alternative site to maintain the continuity of services and operations on the primary site.
[0003] Traditional disaster recovery switching methods often rely on static configuration or manual intervention to switch the cloud host to an alternative host, but this method is difficult to meet the requirements of real-time and automation. In addition, this method will only perform site switching after the primary site fails and requires manual intervention for switching, resulting in a relatively long service or operation interruption time and a high risk of data loss. Summary of the Invention
[0004] The present application provides a method, device, electronic device and storage medium for switching sites, so as to at least solve the problem in the related art that in the scenario of disaster recovery switching to an alternative site, the service or operation interruption time on the primary site is relatively long and the risk of data loss is high.
[0005] The present application provides a method for switching sites, the method includes: obtaining a plurality of metrics of a primary site, each metric being used to indicate the hardware resource occupancy status of the primary site, the network status of the primary site or the service availability status of the primary site; detecting whether there is at least one abnormal metric among the plurality of metrics; if there is at least one abnormal metric, determining at least one failure corresponding to the at least one abnormal metric according to the at least one abnormal metric and the duration of each abnormal metric, each failure being determined based on a combination of one abnormal metric or a plurality of abnormal metrics; judging whether all of the at least one failure are recoverable failures; if so, obtaining a plurality of sessions in the primary site and the delay of the transmission path required for each session, and determining the priority corresponding to each delay; based on the priority corresponding to the delay, batch-switching the plurality of sessions to one or more alternative sites, so that each alternative site re-establishes a communication connection with at least one network device that initiates the plurality of sessions, and the delay is negatively correlated with the priority.
[0006] The present application also provides a switching device for a site, including: an acquisition module, configured to acquire a plurality of metrics of a primary site, each of the metrics being used to indicate the occupancy status of the hardware resources of the primary site, the network status of the primary site, or the service availability status of the primary site;
[0007] a processing module, configured to detect whether there is at least one abnormal metric among the plurality of metrics; if there is the at least one abnormal metric, determine at least one fault corresponding to the at least one abnormal metric according to the at least one abnormal metric and the duration of each abnormal metric, each of the faults being determined based on one of the abnormal metrics or a combination of a plurality of the abnormal metrics; determine whether all of the at least one fault are recoverable faults; if so, acquire a plurality of sessions in the primary site and the delay of the transmission path required for each session, and determine the priority corresponding to each delay; based on the priority corresponding to the delay, batch-switch the plurality of sessions to one or more alternative sites, so that each alternative site re-establishes a communication connection with at least one network device of the plurality of sessions, and the delay is negatively correlated with the priority.
[0008] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above site switching methods when executing the computer program.
[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above site switching methods are implemented.
[0010] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above site switching methods are implemented.
[0011] Through the present application, a plurality of metrics of a primary site can be acquired; when it is detected that there is at least one abnormal metric among the plurality of metrics, at least one fault corresponding to the at least one abnormal metric is determined according to the at least one abnormal metric and the duration of each abnormal metric; when all of the at least one fault are recoverable faults, a plurality of sessions in the primary site and the delay of the transmission path required for each session are acquired, and the priority corresponding to each delay is determined; based on the priority corresponding to the delay, the plurality of sessions are batch-switched to one or more alternative sites.
[0012] Since multiple metrics of the primary site can indicate the hardware resource occupancy status of the primary site, the network status of the primary site, or the service availability status of the primary site, the health status of the primary site can be detected from multiple perspectives. When an abnormal metric is detected, at least one fault corresponding to the primary site is determined in a timely manner based on the abnormal metric, and then the sessions on the primary site are switched in a timely manner to ensure business continuity during the switching process, thereby achieving rapid identification and switching of the faulty site and ensuring high availability of the service. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 Topological structure diagram of a site switching system provided by an embodiment of the present application;
[0015] Figure 2 Flow schematic diagram of a site switching method provided by an embodiment of the present application;
[0016] Figure 3 Flow schematic diagram of another site switching method provided by an embodiment of the present application;
[0017] Figure 4 Structural block diagram of a site switching device provided by an embodiment of the present application;
[0018] Figure 5 Hardware structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0020] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0021] To enable those skilled in the art of this technology to better understand the solution of this application, the following further details this application in conjunction with the accompanying drawings and specific embodiments.
[0022] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the site switching method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0023] The embodiments of this application are applied to a large-scale distributed system in the scenario of disaster recovery switching of the primary site to an alternative site.
[0024] In the related art, to improve the high availability and disaster recovery ability of cloud host services, the primary site is usually switched to an alternative site through manual intervention or static configuration. This method will only perform site switching after the primary site fails and requires manual intervention for switching, resulting in a relatively long service or business interruption time and a high risk of data loss, making it difficult to meet the requirements of real-time and automation.
[0025] To solve the above technical problems, the embodiments of this application provide a site switching method, which realizes the rapid identification and switching of a failed site by real-time monitoring multiple metrics of the primary site, ensuring the high availability of the service.
[0026] The following takes the site switching system as an example to describe the method provided by the embodiments of this application.
[0027] As Figure 1 shown, Figure 1 This is the topology diagram of a site switching system provided by the embodiments of this application. Figure 1 In it, the site switching system 100 includes: a site switching device 101, a primary site 102, an alternative site 103, and a network device 104. Among them, the site switching device 101 is wirelessly connected to the primary site 102, the site switching device 101 is wirelessly connected to the alternative site 103, and the network device 104 is wirelessly connected to the primary site 102.
[0028] Among them, the site switching device 101, also known as the cloud platform, can be any device with computing and communication capabilities. For example, the site switching device 101 can be a server or a cloud server. Optionally, the site switching device 101 also has a display interface or a display screen.
[0029] Among them, the site switching device 101 includes a primary site health monitoring module, a disaster recovery site resource preloading module, and a global dynamic routing and scheduling module. The primary site health monitoring module is used to obtain multiple metrics of the primary site; the disaster recovery site resource preloading module is used to update the routing table and preload the routing table to the alternative site; the global routing and scheduling module is used to switch multiple sessions of the primary site to one or more alternative sites.
[0030] The primary site 102 and the alternative site 103 can be any device with computing and communication capabilities. For example, the primary site 102 or the alternative site 103 can be a server or a cloud server. The same resource pool is deployed at both sites.
[0031] The network device 104 can be any network device with communication functions. For example, the network device 104 can be a server, a router, a switch, etc.
[0032] In addition, it should be understood that Figure 1 only one alternative site and one device are shown in the example. In fact, the system may also include more alternative sites and network devices, and this embodiment does not limit this.
[0033] An embodiment of the present application provides a site switching method, which is applied to a site switching device, such as Figure 2 shown in Figure 2 is a schematic flowchart of a site switching method provided by an embodiment of the present application. The site switching method includes the following steps:
[0034] S201: Obtain multiple metrics of the primary site.
[0035] Among them, each metric is used to indicate the hardware resource occupancy status of the primary site, the network status of the primary site, or the service availability status of the primary site.
[0036] The multiple metrics include multiple hardware resource occupancy status metrics, multiple network status metrics, and multiple application service status metrics of the primary site.
[0037] The multiple hardware resource occupancy status metrics indicate the hardware resource occupancy status of the primary site. The multiple hardware resource occupancy status metrics include the usage rate of the central processing unit (CPU), the memory occupancy rate, the disk throughput, or the disk throughput rate.
[0038] Multiple network status indicators indicate the network status of the primary site. The multiple network status indicators include the end-to-end delay, packet loss rate, and bandwidth utilization of the transmission path required for each session among multiple sessions on the primary site.
[0039] Multiple application service status indicators indicate the service availability status of the primary site. The multiple application service status indicators include the application process survival status of each application and the service port response status of each application.
[0040] Exemplarily, the primary site periodically collects multiple indicators of the primary site through a built-in health monitoring module at regular intervals, and sends the multiple indicators to the switching device of the site through Wi-Fi or a wireless module (Long Range, LoRa); the switching device of the site obtains the multiple indicators, performs deduplication processing on the multiple indicators, and stores the structured data of the multiple indicators using a structured database (MySQL) and stores the time series data using a non-relational database (MongoDB).
[0041] It can be understood that the multiple indicators may also include more indicators without limitation. For example, GPU temperature, response time of the service port, etc.
[0042] S202: Detect whether there is at least one abnormal indicator among the multiple indicators.
[0043] Among them, for each abnormal indicator in the at least one abnormal indicator, the value of each abnormal indicator has a large gap from the corresponding standard value or standard state.
[0044] In some alternative embodiments, if it is detected that any one of the multiple hardware resource occupancy status indicators and the multiple network status indicators is continuously greater than a preset value within a preset time period, then determine any one of the indicators as an abnormal indicator; or, if it is detected that any one of the application service status indicators is in a non-running state, then determine the application process survival status as an abnormal indicator.
[0045] Among them, the preset thresholds corresponding to different hardware resource occupancy status indicators or network status indicators are different. The preset threshold can be set according to actual needs without limitation. For example, the threshold corresponding to CPU utilization is 70%. Another example is that the preset threshold corresponding to the end-to-end delay duration is 100 ms.
[0046] The non-running state may be a sleep state, a pause state, a suspend state, a stop state, a standby state. It can be understood that the non-running state may also include more states without limitation.
[0047] Exemplarily, taking the preset threshold as 70% and the index to be detected as the CPU usage rate as an example, if the switching device of the site detects that the CPU usage rate is greater than or equal to 70%, it determines that the abnormal index is the CPU usage rate; or, taking the preset threshold as 80% and the index to be detected as the memory occupancy rate as an example, if the switching device of the site detects that the memory occupancy rate is greater than or equal to 80%, it determines that the abnormal index is the memory occupancy rate. Or, taking the preset threshold as 100 ms and the index to be detected as the end-to-end delay duration as an example, if the switching device of the site detects that the end-to-end delay duration is greater than or equal to 100 ms, it determines that the abnormal index is the end-to-end delay duration.
[0048] In some alternative embodiments, before detecting whether there is at least one abnormal index among multiple indexes, the switching device of the site obtains a first weight corresponding to the processor utilization rate, a second weight corresponding to the disk throughput rate, and a third weight corresponding to the end-to-end delay duration; calculates the health score of the primary site according to the processor utilization rate, the first weight, the disk throughput rate, the second weight, the end-to-end delay duration, and the third weight; and when the health score is greater than the health threshold, starts the detection step of whether there is at least one abnormal index among the multiple indexes.
[0049] Among them, the health threshold can be set according to actual needs. For example, the health threshold can be 0.8.
[0050] In one example, the switching device of the site calculates the health score of the primary site based on a preset formula, the processor utilization rate, the first weight, the disk throughput rate, the second weight, the end-to-end delay duration, and the third weight. Among them, the preset formula can be H = A * α + B * β + C * γ.
[0051] Among them, H represents the health score; A represents the processor utilization rate; α represents the first weight; B represents the disk throughput rate, β represents the second weight; C represents the end-to-end delay duration, and γ represents the third weight.
[0052] It can be understood that when the health score of the primary site is greater than the health threshold, it indicates that the processor utilization rate of the primary site is relatively high, the disk throughput rate is relatively high, and the end-to-end delay duration is relatively long, that is, the load of the primary site is relatively high or the possibility of a failure occurring is relatively high. At this time, it is necessary to detect multiple indexes in order to determine whether it is necessary to switch multiple sessions to the alternative site based on the multiple indexes in a timely manner.
[0053] S203: If there is at least one abnormal index, determine at least one failure corresponding to the at least one abnormal index according to the at least one abnormal index and the duration of each abnormal index.
[0054] Among them, each failure is determined based on a combination of one abnormal index or multiple abnormal indexes.
[0055] In some alternative embodiments, if at least one abnormal indicator includes a first abnormal indicator, the switching device of the site obtains a first duration during which the first abnormal indicator has an abnormality, detects whether the first duration is greater than or equal to a first threshold value to obtain a first detection result, and determines at least one fault based on the first detection result. Wherein, the first abnormal indicator is any one of a plurality of indicators.
[0056] Exemplarily, taking that at least one abnormal indicator includes the service port response time and the first threshold value is 3 minutes as an example, if at least one abnormal indicator includes the service port response status, the switching device of the site obtains a first duration of 5 minutes during which the service port response status has an abnormality (the service port response status is a non-operating state), detects whether the first duration is greater than or equal to 3 minutes, obtains a first detection result that the duration of the service port response status having an abnormality is greater than 3 minutes, and determines at least one fault as a service port fault based on this first detection result.
[0057] In some alternative embodiments, if at least one abnormal indicator includes a second abnormal indicator and a third abnormal indicator, the switching device of the site obtains a second duration during which the second abnormal indicator has an abnormality and a third duration during which the third abnormal indicator has an abnormality; detects whether the second abnormal indicator is greater than or equal to a second threshold value and whether the third abnormal indicator is greater than or equal to a third threshold value to obtain a second detection result, and determines at least one fault based on the second detection result.
[0058] Wherein, the second abnormal indicator is any one of a plurality of hardware resource occupancy status indicators.
[0059] The third abnormal indicator is any one of a plurality of network status indicators or any one of a plurality of application service status indicators.
[0060] Exemplarily, taking that at least one abnormal indicator includes a second abnormal indicator and a third abnormal indicator, the second abnormal indicator is that the CPU utilization rate is 80%, the third abnormal indicator is that the end-to-end delay duration is 120 ms, the second threshold value is 3 minutes, and the third threshold value is 5 minutes as an example, if at least one abnormal indicator includes the CPU utilization rate and the end-to-end delay duration, the switching device of the site obtains a second duration of 5 minutes during which the CPU utilization rate has an abnormality (that is, the CPU utilization rate is greater than 70%) and a third duration of 7 minutes during which the third abnormal indicator has an abnormality (that is, the end-to-end delay duration is greater than 100 ms); detects that the second detection result is that the second abnormal indicator is greater than the second threshold value of 3 minutes and the third abnormal indicator is greater than the third threshold value of 5 minutes, and determines at least one fault as a CPU overload fault based on this second detection result.
[0061] Optionally, if at least one abnormal indicator includes two fourth abnormal indicators, the handover device of the site may further obtain the duration of the abnormality occurrence in each fourth abnormal indicator; detect whether each fourth abnormal indicator is greater than or equal to the corresponding threshold to obtain a third detection result, and determine at least one fault based on the third detection result.
[0062] Among them, the two fourth abnormal indicators may be two indicators among multiple hardware resource occupancy status indicators, or the two fourth abnormal indicators may be two indicators among multiple network status indicators; or the two fourth abnormal indicators may be two indicators among multiple application service status indicators. For example, the two fourth abnormal indicators may be that the bandwidth utilization rate exceeds 80% and the end-to-end delay duration is greater than 100 ms. Correspondingly, the handover device of the site determines that at least one fault is a network congestion fault based on the bandwidth utilization rate exceeding 80% and the end-to-end delay duration being greater than 100 ms.
[0063] Optionally, the handover device of the site may further obtain a fourth weight corresponding to the bandwidth utilization rate, a fifth weight corresponding to the packet loss rate, and a sixth weight corresponding to the end-to-end delay duration; calculate a first health score of the transmission link required for each session on the primary site according to the bandwidth utilization rate, the fourth weight, the packet loss rate, the fifth weight, the end-to-end delay duration, and the sixth weight; when the first health score is less than the target threshold, determine that the transmission link required for each session fails, that is, determine that at least one fault is a transmission link fault.
[0064] Among them, the first health score can be set according to actual needs. For example, the first health score is 0.6.
[0065] Among them, the fourth weight, the fifth weight, and the sixth weight can be set according to actual needs. For example, the fourth weight is 0.5, the fifth weight is 0.2, and the sixth weight is 0.3.
[0066] It can be understood that when the handover device of the site determines that the transmission link required for each session fails, it can call the shortest path algorithm (Dijkstra algorithm) to recalculate the optimal path, update the routing table, and preload the routing table to the alternative site.
[0067] S204: Determine whether all of the at least one fault are recoverable faults.
[0068] Among them, a recoverable fault is a temporary fault. For example, a recoverable fault may be a CPU overload fault or a network congestion fault caused by a large number of sessions or excessive access requests on the primary site.
[0069] It is understandable that at least one of the faults may include an unrecoverable fault, such as a hard disk damage fault, an operating system fault that is subject to a malicious attack, or a motherboard hardware fault.
[0070] In an example, the switching device of the site may query the fault type of at least one fault in a preset mapping table to determine whether each fault is a recoverable fault or an unrecoverable fault.
[0071] The preset mapping table includes a plurality of corresponding relationships, each corresponding relationship being a corresponding relationship between a fault and a fault type. The fault type includes a recoverable fault or an unrecoverable fault.
[0072] S205: If yes, obtain multiple sessions in the primary site and the delay of the transmission path required for each session, and determine the priority corresponding to each delay.
[0073] Among them, delay is negatively correlated with priority.
[0074] Exemplarily, the switching device of the site obtains that the delay of the transmission path required for the first session among multiple sessions in the main site is 10ms, and the delay of the transmission path required for the second session is 20ms, compares the delay of the transmission path required for the first session (10ms) and the delay of the transmission path required for the second session (20ms), determines that the delay of the transmission path required for the first session is less than the delay of the transmission path required for the second session, and therefore determines that the priority of the first session is greater than the priority of the second session.
[0075] S206: Switching the multiple sessions to one or more candidate sites in batches based on the priorities corresponding to the delays, so that each candidate site re-establishes a communication connection with at least one network device of the multiple sessions.
[0076] In some optional implementations, the site switching device modifies the original address of the main site in the domain name corresponding to the main site to the new address of one or more alternative sites; obtains the session data of each session from the main site; determines the migration order of multiple sessions based on the priority corresponding to the delay; and according to the migration order, uses an encrypted transmission protocol to switch the session identifier and session status of each session in the multiple sessions in batches to the new address of one or more alternative sites.
[0077] The session data includes the session identifier and session status of each session. The session identifier (SessionID) is used to ensure the continuity of the same user request at the primary site and the alternative site.
[0078] Exemplarily, the site switching device may update the domain name resolution record, and modify the original address of the main site in the domain name corresponding to the main site in the domain name resolution record to the new address of one or more candidate sites.
[0079] Understandably, the site switching device adopts a session retention and encrypted transmission protocol or a key verification mechanism, which can ensure service continuity and data transmission security during the switching process. The above site switching logic can also be called a progressive switching mode, and data compensation can also be based on incremental logs.
[0080] In some alternative embodiments, if at least one of the faults is not a recoverable fault, the site switching device directly switches multiple sessions to one or more alternative sites, and migrates all session data of the multiple sessions on the primary site to one or more alternative sites through an incremental log compensation mechanism. This site switching logic can be called a forced switching mode.
[0081] Understandably, during the process of batch-switching multiple sessions to one or more alternative sites, the site switching device can adopt a smooth switching method, that is, a method of first synchronizing data and then switching. After waiting for the last round of incremental data synchronization to complete, multiple sessions are batch-switched to one or more alternative sites to ensure data consistency. Or, in an emergency failure scenario (that is, when at least one of the faults is not a recoverable fault), the unsynchronized data is ignored, and multiple sessions are directly batch-switched to one or more alternative sites, so that the alternative sites take over the sessions (that is, take over the services).
[0082] Furthermore, after the site switching device migrates multiple sessions to one or more alternative sites, it can monitor the processor utilization rate and end-to-end delay duration of the primary site in real time; when it detects that the processor utilization rate is less than the fourth threshold and the end-to-end delay duration is less than the fifth threshold, all sessions on one or more alternative sites are switched back to the primary site.
[0083] Among them, the fourth threshold and the fifth threshold can be set according to actual needs. For example, the fourth threshold is 70%, and the fifth threshold is 90 ms.
[0084] Optionally, before the site switching device batch-switches multiple sessions to one or more alternative sites, it can also generate a fault analysis report and an alarm message based on at least one fault, and send the fault analysis report and the alarm message to the display interface. The alarm message is used to indicate that a fault has occurred on the primary site and site switching is required.
[0085] Based on the above Figure 2The method shown can obtain multiple metrics of the primary site; detect whether there is at least one abnormal metric among the multiple metrics; if there is at least one abnormal metric, determine at least one fault corresponding to the at least one abnormal metric according to the at least one abnormal metric and the duration of each abnormal metric; and when all of the at least one fault are recoverable faults, obtain multiple sessions in the primary site and the delay of the transmission path required for each session, and determine the priority corresponding to each delay; based on the priority corresponding to the delay, batch-switch the multiple sessions to one or more alternative sites.
[0086] Since the multiple metrics of the primary site can indicate the hardware resource occupancy status of the primary site, the network status of the primary site, or the service availability status of the primary site, the health status of the primary site can be detected from multiple perspectives. When an abnormal metric is detected, at least one fault corresponding to the primary site can be determined in a timely manner based on the abnormal metric, and then the sessions on the primary site can be switched in a timely manner to ensure business continuity during the switching process, thereby realizing the rapid identification and switching of the faulty site and ensuring the high availability of the service.
[0087] An embodiment of the present application provides a method for switching sites, which is applied to a site switching device, such as Figure 3 shown Figure 3 is a schematic flowchart of another method for switching sites provided by an embodiment of the present application. The method for switching sites includes the following steps:
[0088] S301: Obtain multiple metrics of the primary site.
[0089] For the specific implementation process, reference can be made to the foregoing step S201, which will not be elaborated here.
[0090] S302: Detect whether there is at least one abnormal metric among the multiple metrics.
[0091] For the specific implementation process, reference can be made to the foregoing step S202, which will not be elaborated here.
[0092] S302: If not, there is no need to switch to an alternative site.
[0093] S304: If so, batch-switch the multiple sessions to one or more alternative sites.
[0094] For the specific implementation process, reference can be made to the foregoing steps S203 - S206, which will not be elaborated here.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0096] An embodiment of the present application further provides a site switching device, such as Figure 4 shown in Figure 4 FIG. 4 is a structural block diagram of a site switching device provided by an embodiment of the present application; the site switching device includes:
[0097] An acquisition module 401, configured to acquire multiple metrics of a primary site, and each metric is used to indicate the hardware resource occupancy status of the primary site, the network status of the primary site, or the service availability status of the primary site.
[0098] A processing module 402, configured to detect whether there is at least one abnormal metric among the multiple metrics. If there is at least one abnormal metric, then determine at least one fault corresponding to the at least one abnormal metric according to the at least one abnormal metric and the duration of each abnormal metric, and each fault is determined based on a combination of one abnormal metric or multiple abnormal metrics. Determine whether all of the at least one fault are recoverable faults. If so, acquire multiple sessions in the primary site and the delay of the transmission path required for each session, and determine the priority corresponding to each delay. Based on the priority corresponding to the delay, batch-switch the multiple sessions to one or more alternative sites, so that each alternative site re-establishes a communication connection with at least one network device of the multiple sessions, and the delay is negatively correlated with the priority.
[0099] In some alternative embodiments, the multiple metrics include: multiple hardware resource occupancy status metrics of the primary site, multiple network status metrics, and multiple application service status metrics; the processing module 402 is specifically configured to determine any one of the metrics as an abnormal metric if it is detected that any one of the multiple hardware resource occupancy status metrics and the multiple network status metrics is continuously greater than a preset value within a preset time period; or, if it is detected that any one of the application service status metrics is in a non-running state, determine that the application process survival status is an abnormal metric.
[0100] In some alternative embodiments, the processing module 402 is further specifically configured to, if the at least one abnormal metric includes a first abnormal metric, acquire a first duration during which the first abnormal metric occurs abnormally, and detect whether the first duration is greater than or equal to a first threshold to obtain a first detection result, and determine at least one fault based on the first detection result, where the first abnormal metric is any one of the multiple metrics;
[0101] The processing module 402 is further specifically configured to, if the at least one abnormal metric includes a second abnormal metric and a third abnormal metric, acquire a second duration during which the second abnormal metric occurs abnormally and a third duration during which the third abnormal metric occurs abnormally, where the second abnormal metric is any one of the multiple hardware resource occupancy status metrics, and the third abnormal metric is any one of the multiple network status metrics or any one of the multiple application service status metrics;
[0102] The processing module 402 is further specifically configured to detect whether the second abnormal index is greater than or equal to the second threshold, and whether the third abnormal index is greater than or equal to the third threshold, obtain a second detection result, and determine at least one fault based on the second detection result.
[0103] In some alternative embodiments, the multiple hardware resource occupancy status indicators include processor utilization and disk throughput, and the multiple network status indicators include end-to-end delay duration; before detecting whether there is at least one abnormal indicator among the multiple indicators, the obtaining module 401 is further configured to obtain a first weight corresponding to the processor utilization, a second weight corresponding to the disk throughput rate, and a third weight corresponding to the end-to-end delay duration; the processing module 402 is further configured to calculate a health score of the primary site according to the processor utilization, the first weight, the disk throughput rate, the second weight, the end-to-end delay duration, and the third weight; the processing module 402 is further configured to start the detection step of whether there is at least one abnormal indicator among the multiple indicators when the health score is greater than the health threshold.
[0104] In some alternative embodiments, the processing module 402 is further specifically configured to modify the original address of the primary site in the domain name corresponding to the primary site to the new addresses of one or more alternative sites; obtain session data of each session from the primary site, where the session data includes a session identifier and a session status of each session; determine a migration order of multiple sessions based on the priority corresponding to the delay; and in accordance with the migration order, use an encrypted transmission protocol to batch-switch the session identifier and the session status of each session among the multiple sessions to the new addresses of one or more alternative sites.
[0105] In some alternative embodiments, if not all of the at least one fault are recoverable faults, the processing module 402 is further configured to directly switch multiple sessions to one or more alternative sites, and migrate all session data of the multiple sessions of the primary site to one or more alternative sites through an incremental log compensation mechanism.
[0106] In some alternative embodiments, after migrating multiple sessions to one or more alternative sites, the processing module 402 is further configured to monitor the processor utilization and the end-to-end delay duration of the primary site in real time; when it is detected that the processor utilization is less than a fourth threshold and the end-to-end delay duration is less than a fifth threshold, switch all sessions on one or more alternative sites back to the primary site.
[0107] For the description of the features in the corresponding embodiments of the site switching device, reference may be made to the relevant description of the corresponding embodiments of the site switching method, which will not be elaborated here one by one.
[0108] An embodiment of the present application further provides an electronic device, as Figure 5 shown Figure 5A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. The electronic device may be Figure 1 The handover device 101 of the site shown in the figure. The electronic device includes a processor 10 and a memory 20. A computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any of the above-described method embodiments for switching sites.
[0109] An embodiment of the present application also provides a computer-readable storage medium in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described method embodiments for switching sites when running.
[0110] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0111] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described method embodiments for switching sites.
[0112] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described method embodiments for switching sites.
[0113] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0114] The above has introduced in detail a method, device, electronic device and storage medium for switching stations provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for switching stations, characterized in that, The method includes: Obtaining a plurality of metrics of the primary site, each metric being used to indicate the hardware resource occupancy status of the primary site, the network status of the primary site, or the service availability status of the primary site; Detecting whether there is at least one abnormal metric among the plurality of metrics; If there is the at least one abnormal metric, determining at least one fault corresponding to the at least one abnormal metric according to the at least one abnormal metric and the duration of each abnormal metric, each fault being determined based on one of the abnormal metrics or a combination of a plurality of the abnormal metrics; Judging whether all of the at least one fault are recoverable faults; If so, obtaining a plurality of sessions in the primary site and the delay of the transmission path required for each session, and determining the priority corresponding to each delay; Based on the priority corresponding to the delay, batch-switching the plurality of sessions to one or more alternative sites, so that each alternative site re-establishes a communication connection with at least one network device that initiates the plurality of sessions, and the delay is negatively correlated with the priority.
2. The method according to claim 1, wherein The plurality of metrics include: a plurality of hardware resource occupancy status metrics of the primary site, a plurality of network status metrics, and a plurality of application service status metrics; The detecting whether there is at least one abnormal metric among the plurality of metrics includes: If it is detected that any one of the plurality of hardware resource occupancy status metrics and the plurality of network status metrics is continuously greater than a preset value within a preset time period, determining the any one of the metrics as one of the abnormal metrics; or, If it is detected that any one of the application service status metrics is in a non-running state, determining the application process survival status as one of the abnormal metrics.
3. The method according to claim 2, wherein The determining at least one fault corresponding to the at least one abnormal metric according to the at least one abnormal metric and the duration of each abnormal metric includes: If the at least one abnormal metric includes a first abnormal metric, obtaining a first duration during which the first abnormal metric has an abnormality, and detecting whether the first duration is greater than or equal to a first threshold to obtain a first detection result, and determining the at least one fault based on the first detection result, the first abnormal metric being any one of the plurality of metrics; If the at least one abnormal metric includes a second abnormal metric and a third abnormal metric, obtaining a second duration during which the second abnormal metric has an abnormality and a third duration during which the third abnormal metric has an abnormality, the second abnormal metric being any one of the plurality of hardware resource occupancy status metrics, and the third abnormal metric being any one of the plurality of network status metrics or any one of the plurality of application service status metrics; Detecting whether the second abnormal metric is greater than or equal to a second threshold and whether the third abnormal metric is greater than or equal to a third threshold to obtain a second detection result, and determining the at least one fault based on the second detection result.
4. The method according to claim 3, characterized in that The multiple hardware resource occupancy status indicators include processor utilization rate and disk throughput rate, and the multiple network status indicators include end-to-end delay duration; before detecting whether there is at least one abnormal indicator among the multiple indicators, the method further includes: Obtaining a first weight corresponding to the processor utilization rate, a second weight corresponding to the disk throughput rate, and a third weight corresponding to the end-to-end delay duration; Calculating a health score of the primary site according to the processor utilization rate, the first weight, the disk throughput rate, the second weight, the end-to-end delay duration, and the third weight; When the health score is greater than a health threshold, starting a detection step of whether there is the at least one abnormal indicator among the multiple indicators.
5. The method according to claim 4, wherein The batch-switching the multiple sessions to one or more alternative sites based on the priority corresponding to the delay includes: Modifying the original address of the primary site in the domain name corresponding to the primary site to the new addresses of one or more of the alternative sites; Obtaining session data of each of the sessions from the primary site, where the session data includes a session identifier and a session status of each of the sessions; Determining a migration order of the multiple sessions based on the priority corresponding to the delay; According to the migration order, using an encrypted transmission protocol, batch-switching the session identifier and the session status of each of the multiple sessions to the new addresses of one or more of the alternative sites.
6. The method according to claim 5, wherein The method further includes: If not all of the at least one fault are recoverable faults, directly switching the multiple sessions to one or more of the alternative sites, and migrating all session data of the multiple sessions of the primary site to the one or more alternative sites through an incremental log compensation mechanism.
7. The method according to claim 6, characterized in that, After migrating the multiple sessions to one or more of the alternative sites, the method further includes: Real-time monitoring the processor utilization rate and the end-to-end delay duration of the primary site; When it is detected that the processor utilization rate is less than a fourth threshold and the end-to-end delay duration is less than a fifth threshold, switching all sessions on one or more of the alternative sites back to the primary site.
8. A switching device for a site, characterized in that, The switching device of the site includes: An obtaining module, configured to obtain multiple indicators of a primary site, where each of the indicators is used to indicate a hardware resource occupancy status of the primary site, a network status of the primary site, or a service availability status of the primary site; A processing module, configured to detect whether there is at least one abnormal metric among the multiple metrics; if there is the at least one abnormal metric, determine at least one fault corresponding to the at least one abnormal metric according to the at least one abnormal metric and the duration of each abnormal metric, where each fault is determined based on one of the abnormal metrics or a combination of multiple abnormal metrics; determine whether all of the at least one fault are recoverable faults; if so, obtain multiple sessions in the primary site and the latency of the transmission path required for each session, and determine the priority corresponding to each latency; based on the priority corresponding to the latency, batch-switch the multiple sessions to one or more alternative sites, so that each alternative site re-establishes a communication connection with at least one network device that initiates the multiple sessions, and the latency is negatively correlated with the priority.
9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the site switching method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the site switching method according to any one of claims 1 to 7 when executed by a processor.