Switching method, device and equipment of dual-computer hot standby system and storage medium
By using machine learning models to determine potential host failures and switch business traffic in dual-machine hot standby systems, the problems of misswitching and resource waste in traditional systems are solved, and more efficient failover and system stability are achieved.
Patent Information
- Application Number
- CN202510209049.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
The traditional dual-machine hot standby system has problems such as error switching and switching delay during fault detection and switching, resulting in waste of resources and business interruption.
By obtaining the original running parameters of the host system, and using the pre-trained target analysis judgment model (based on machine learning) to determine whether the host has a potential failure, switch the service processing traffic to the backup system when a failure is detected.
Improves the accuracy and efficiency of failover, reduces unnecessary switching operations, and improves system stability and resource utilization.
Smart Images

Figure CN120144370A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of system operation and maintenance, and particularly to a switching method, device, equipment and storage medium for a dual-machine hot standby system. Background Art
[0002] In today's information society, with the continuous development of IT infrastructure, high availability and system fault tolerance have become core requirements indispensable in modern computer systems. In order to improve the stability and reliability of the system, a dual-machine hot standby system has emerged. Traditional dual-machine hot standby systems provide a certain degree of fault tolerance through simple switching between the primary machine and the standby machine, but have defects such as low resource utilization rate and inaccurate switching timing.
[0003] Currently, existing dual-machine hot standby systems usually rely on basic methods such as heartbeat detection and fault detection to determine whether the primary machine has a fault. Once the primary machine has an abnormality, the system will be switched. Although these traditional solutions ensure the high availability of the system, in actual use, there are often problems such as false switching and switching delay, resulting in low system efficiency. Especially when the load is high, the standby system is in an idle state, causing serious resource waste; and when a fault occurs, the traditional system switching method cannot achieve a quick response, resulting in service interruption or performance degradation. Summary of the Invention
[0004] The present invention provides a switching method, device, equipment and storage medium for a dual-machine hot standby system to accurately predict whether there are potential faults in the primary machine system, improve the switching timing between the primary and standby systems when the primary machine system has a fault, and at the same time can intelligently schedule the load sharing between the primary and standby systems to improve work efficiency.
[0005] According to one aspect of the present invention, a switching method for a dual-machine hot standby system is provided. The method includes:
[0006] Obtaining the original operation parameters generated during the operation of the primary machine system in the dual-machine hot standby system;
[0007] Determining whether there is a potential fault in the primary machine system according to the original operation parameters and a target analysis and judgment model obtained through pre-training, wherein the target analysis and judgment model is obtained through training of a machine learning model;
[0008] In the case where the primary machine system has a potential fault, guiding the service processing traffic of the primary machine system to the standby system in the dual-machine hot standby system.
[0009] According to another aspect of the present invention, a switching device for a dual-machine hot standby system is provided. The device includes:
[0010] An operating parameter acquisition module, configured to acquire the original operating parameters generated during the operation of the host system in the dual - machine hot standby system;
[0011] A potential fault determination module, configured to determine whether there is a host potential fault in the host system according to the original operating parameters and a target analysis and judgment model obtained by pre - training, wherein the target analysis and judgment model is obtained by training a machine learning model;
[0012] A standby system switching module, configured to direct the service processing traffic of the host system to the standby system in the dual - machine hot standby system when there is a potential fault in the host system.
[0013] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the switching method of the dual - machine hot standby system according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, there is provided a computer - readable storage medium storing computer instructions for causing a processor to implement the switching method of the dual - machine hot standby system according to any embodiment of the present invention when executed.
[0018] The technical solution of the embodiment of the present invention obtains the original operating parameters generated during the operation of the host system in the dual - machine hot standby system. According to the original operating parameters and a target analysis and judgment model obtained by pre - training, it is determined whether there is a host potential fault in the host system, wherein the target analysis and judgment model is obtained by training a machine learning model. When there is a potential fault in the host system, the service processing traffic of the host system is directed to the standby system in the dual - machine hot standby system, avoiding mis - switching based on heartbeat detection in the traditional method, reducing unnecessary switching operations, and improving system stability and switching accuracy.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0021] Figure 1 is a flowchart of a switching method for a dual-machine hot standby system according to Embodiment 1 of the present invention;
[0022] Figure 2 is a flowchart of a switching method for a dual-machine hot standby system according to Embodiment 2 of the present invention;
[0023] Figure 3 is a structural diagram of a switching device for a dual-machine hot standby system according to Embodiment 3 of the present invention;
[0024] Figure 4 is a schematic structural diagram of an electronic device for implementing the switching method of the dual-machine hot standby system of the embodiments of the present invention. Detailed Embodiments
[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0027] Embodiment 1
[0028] Figure 1The following is a flowchart of a switching method for a dual - machine hot - standby system provided in the first embodiment of the present invention. This embodiment is applicable to the situation where the standby server can quickly take over when the primary server fails. This method can be executed by a switching device of the dual - machine hot - standby system. The switching device of the dual - machine hot - standby system can be implemented in the form of hardware and / or software, and the switching device of the dual - machine hot - standby system can be configured in an electronic device. As Figure 1 shown, the method includes:
[0029] S101. Obtain the original operation parameters generated during the operation of the host system in the dual - machine hot - standby system.
[0030] It should be noted that the dual - machine hot - standby system ensures that when the primary server system fails, the standby server system can quickly take over through two server systems (one primary server system and one standby server system), reducing service interruption. The dual - machine hot - standby system ensures service continuity through two servers and is applicable to scenarios with high requirements for high availability. Although there are costs and complexities, its advantages are significant.
[0031] Among them, the original operation parameters are multiple operation parameters generated during the operation of the host system. Exemplarily, the original operation parameters may include CPU usage rate and trend, memory usage rate and trend, network latency and fluctuation, error count and type distribution, service response time and ratio, and system load and load ratio.
[0032] Specifically, the dual - machine hot - standby system monitors and records multiple original operation parameters of the host system in real - time to obtain the original operation parameters, which provide real - time health status information for the dual - machine hot - standby system.
[0033] S102. Determine whether there is a potential host failure in the host system according to the original operation parameters and a pre - trained target analysis and judgment model.
[0034] Among them, the target analysis and judgment model is obtained through training of a machine learning model. It should be noted that in the present invention, the target analysis and judgment model adopts the random forest algorithm because of its excellent performance in high - dimensional data processing and classification tasks. Random forest can effectively process a large number of features and enhance the robustness of the model through the way of ensemble learning, reducing the risk of overfitting.
[0035] Specifically, the original operation parameters can be directly input into the target analysis and judgment model for failure judgment, and according to the output result of the target analysis and judgment model, it is determined whether there is a potential host failure in the host system.
[0036] Exemplarily, determining whether there is a potential host fault in the host system according to the original operating parameters and the target analysis and judgment model obtained through pre-training includes: performing data preprocessing on the original operating parameters to obtain target operating parameters; and determining whether there is a potential host fault in the host system according to the target operating parameters and the target analysis and judgment model obtained through pre-training.
[0037] Among them, the data preprocessing includes but is not limited to data denoising, data standardization, data normalization, and time series processing.
[0038] It should be noted that before the target analysis and judgment model analyzes the monitoring data, it is necessary to preprocess the original operating parameters to ensure their quality and accuracy. Data denoising: The original operating parameters may contain some noise, such as instantaneous abnormal fluctuations or sensor errors, which may affect subsequent analysis. Therefore, it is necessary to remove the noise through means such as filtering and smoothing to ensure that the data truly reflects the actual state of the system. Data standardization and normalization: Different original operating parameters (such as CPU utilization, memory occupancy, response time, etc.) may have different dimensions and ranges. In order to ensure that each indicator is compared under the same importance, it is necessary to perform standardization or normalization processing on the data. Time series processing: The original operating parameters are usually time series data, so it is necessary to perform time window segmentation on the data, analyze the data trend in each time period, and capture the change pattern of the system load.
[0039] Specifically, after performing data preprocessing on the original operating parameters, target operating parameters can be obtained. The target operating parameters are directly input into the target analysis and judgment model for fault judgment, and according to the output result of the target analysis and judgment model, it is determined whether there is a potential host fault in the host system.
[0040] Exemplarily, determining whether there is a potential host fault in the host system according to the target operating parameters and the target analysis and judgment model obtained through pre-training includes: extracting data features from the target operating parameters to obtain characteristic operating parameters; inputting the characteristic operating parameters into the target analysis and judgment model for fault judgment, and determining whether there is a potential host fault in the host system according to the output of the target analysis and judgment model.
[0041] It should be noted that after the data preprocessing is completed, it is necessary to extract the key features that can reflect the health status of the system. The goal of the feature extraction process is to transform the original data into high-dimensional information that can effectively distinguish normal and abnormal states. Among them, the data features at least include parameter trend, peak fluctuation, jitter frequency, and abnormal pattern.
[0042] Parameter trend analysis: Such as the change trends of CPU utilization and memory occupancy. A long-term continuous increase may indicate a performance bottleneck in the system, while occasional small fluctuations may be normal load fluctuations. Peak fluctuations: Capturing sudden changes in system metrics, especially extreme values that occur within a short period, which may indicate potential hardware failures or software anomalies. Jitter frequency: The fluctuations or instabilities of network response times, which may indicate network anomalies or transmission failures. Abnormal patterns: Patterns extracted from error logs, such as hard disk failures, memory leaks, service crashes, etc. These abnormal information usually foreshadows system failures in advance.
[0043] Specifically, after obtaining the characteristic operating parameters, the characteristic operating parameters are directly input into the target analysis and judgment model for fault judgment. According to the output result of the target analysis and judgment model, it is determined whether there is a potential host fault in the host system. The target analysis and judgment model performs multi-dimensional analysis by combining multiple monitoring parameters to comprehensively evaluate the health status of the system. This can avoid making mis-switches due to the abnormality of a single indicator.
[0044] S103. In the case where there is a potential fault in the host system, guide the service processing traffic of the host system to the standby system in the dual-machine hot standby system.
[0045] Among them, the service processing traffic may refer to all data requests, tasks, or service calls being processed in the host system. The service processing traffic may include user request traffic, data processing traffic, and service-to-service call traffic, etc.
[0046] Specifically, when the intelligent analysis module in the dual-machine hot standby system detects a potential fault in the host system, it immediately triggers the standby system takeover operation. The dual-machine hot standby system guides the service processing traffic from the primary system to the standby system through the switching control module, achieving fault switching at the millisecond level to ensure service continuity.
[0047] Exemplarily, assume that during a major promotion period, the primary system of an e-commerce website has an excessive load and a potential fault in the host system is detected. At this time, the switching control module will switch the following traffic to the standby system: requests for users to browse products, place orders, make payments, etc. (user request traffic); background tasks such as order processing and inventory updates (data processing traffic); interactions with payment gateways and logistics systems (service-to-service call traffic). In this way, the system can quickly switch to the standby system when the primary system fails or has a high load, ensuring that users are unaware and service continuity is guaranteed.
[0048] The technical solution of the embodiment of the present invention obtains the original operation parameters generated during the operation of the host system in the dual-active hot standby system. According to the original operation parameters and the target analysis and judgment model obtained by pre-training, it is determined whether there is a potential host failure in the host system, where the target analysis and judgment model is obtained by training a machine learning model. When there is a potential failure in the host system, the service processing traffic of the host system is guided to the standby system in the dual-active hot standby system, avoiding the mis-switching based on heartbeat detection in the traditional method, reducing unnecessary switching operations, and improving the system stability and switching accuracy.
[0049] Based on the above embodiments, the training process of the target analysis and judgment model includes:
[0050] Obtain the sample operation parameters and the operation status results corresponding to the sample operation parameters. Based on the random forest algorithm, input the sample operation parameters into a preset machine learning model for failure judgment, and obtain an output judgment result based on the output of the preset machine learning model. Determine the training error based on the output judgment result and the operation status result, and backpropagate the training error into the preset machine learning model to adjust the network parameters in the preset machine learning model. When the preset convergence condition is met, it is determined that the training of the preset machine learning model is completed, and the target analysis and judgment model is obtained.
[0051] Among them, the sample operation parameters can be the sample parameters after transforming the historical original operation parameters and / or historical target operation parameters and / or historical feature operation parameters. The sample operation parameters include normal operation parameters and abnormal operation parameters; correspondingly, the operation status results include normal operation status and failure operation status.
[0052] Specifically, based on the training function, the training error can be determined according to the output judgment result and the operation status result of the preset machine learning model, and the training error is backpropagated into the preset machine learning model to adjust the network parameters in the preset machine learning model until the preset convergence condition is met. For example, when the number of iterations reaches the preset number or the training error converges, it is determined that the training of the preset machine learning model is completed. At this time, the preset machine learning model with the training completed can be used as the target analysis and judgment model. By using the sample operation parameters and the operation status results for model training, the model can not only identify the current failure state but also predict potential failures through trends. For example, when the CPU load is continuously increasing, the model can give an early warning of the upcoming performance bottleneck or crash.
[0053] Based on the above embodiments, after determining whether there is a potential host failure in the host system according to the original operating parameters and the target analysis and judgment model obtained through pre-training, the method further includes: obtaining historical operating parameters within a preset time period; and optimizing and training the target analysis and judgment model based on the historical operating parameters and the corresponding system operating status.
[0054] That is to say, the dual-active hot standby system retrains the model by using the historical operating parameters within a preset time period at regular intervals or under trigger conditions, enabling the target analysis and judgment model to adapt to the new system environment and load characteristics, thereby continuously optimizing the fault judgment ability. In addition, the target analysis and judgment model will automatically adjust its parameters according to new data. Especially when new types of faults occur, it can better identify and respond. If there are cases of incorrect switching or missed switching, the dual-active hot standby system will feedback the results into the training data to further improve the accuracy of the model.
[0055] Embodiment 2
[0056] Figure 2 The flowchart of a switching method for a dual-active hot standby system provided by Embodiment 2 of the present invention. Based on the above embodiments, the present invention also provides a load sharing technical solution between the primary and standby systems. As Figure 2 shown, the method includes:
[0057] S201. Obtain the original operating parameters generated during the operation of the host system in the dual-active hot standby system.
[0058] S202. Determine whether there is a potential host failure in the host system according to the original operating parameters and the target analysis and judgment model obtained through pre-training.
[0059] S203. In the case where the host system has a potential failure, direct the service processing traffic of the host system to the standby system in the dual-active hot standby system.
[0060] S204. Determine the host load data of the host system according to the original operating parameters.
[0061] Specifically, the dual-active hot standby system reflects the current working states of the primary system and the standby system by collecting and analyzing the original operating parameters in real time, and determines the host load data as the basis for load balancing decisions. Through precise monitoring, the system can continuously track load fluctuations, quickly identify potential bottlenecks and fault risks, and then make corresponding load adjustments.
[0062] The system collects various monitoring metrics at high frequencies and uses an efficient transmission mechanism to upload the data to the central monitoring platform in real time for analysis. For each metric, the system defines a normal fluctuation range and a fault threshold, and the monitoring platform automatically triggers an alarm or adjusts operations based on these thresholds.
[0063] Combined with the time series characteristics of the data, the system can analyze the instantaneous change trend of the load, timely identify potential high-load areas, and make predictions about the future change trend of the load.
[0064] S205. Determine the load sharing decision of the standby system according to the host load data, a preset load threshold, and a fault critical threshold set in advance.
[0065] To ensure the stable operation of the system and avoid resource waste, the setting of the preset load threshold and the fault critical threshold plays a key role in the load balancing mechanism. By monitoring metrics such as CPU, memory, and response time, the system dynamically sets the preset load threshold and the fault critical threshold to determine the timing of task allocation and switching.
[0066] Specifically, compare the host load data with the preset load threshold and the fault critical threshold set in advance, and then determine the load sharing decision of the standby system according to the comparison results.
[0067] Exemplarily, the determining of the load sharing decision of the standby system according to the host load data, a preset load threshold, and a fault critical threshold set in advance includes: when the host load data is lower than the preset load threshold, determine that the standby system does not participate in load sharing; when the host load data is higher than or equal to the preset load threshold and lower than the fault critical threshold, determine that the standby system participates in non-critical task load sharing; when the host load data is higher than or equal to the fault critical threshold, determine that the standby system releases non-critical task loads and prepares to take over the critical task loads of the host system.
[0068] In the present invention, when the load of the host system is low and resources are sufficient (i.e., when the host load data is lower than the preset load threshold), the standby system does not participate in the load sharing of the main system to avoid unnecessary resource waste. The system will maintain the normal operation state of the host system to ensure that resources are utilized maximally. This strategy effectively prevents the idle of resources in the host system and maximally improves the execution efficiency of the main system when resources are idle.
[0069] When the host load data of the host system reaches the set high threshold (i.e., higher than or equal to the preset load threshold and lower than the fault critical threshold, for example, the CPU usage rate exceeds 70%), the standby system starts to share non-critical tasks (such as log analysis, data backup, etc.).
[0070] When the host load data of the host system approaches the failure threshold (i.e., higher than or equal to the critical failure threshold, such as the CPU usage rate reaching more than 90%), the standby system is ready to take over the critical tasks of the main system, ensuring seamless system switching and high availability.
[0071] Only when the load of the main system is close to saturation or there is a potential risk of failure, the standby system will intervene to share resources. By dynamically adjusting task allocation, the standby system can reasonably allocate non-critical tasks, such as log processing, data backup, data cleaning, etc., according to the current load situation, thereby reducing the burden on the main system and improving the overall system performance. The task allocation of the standby system is automatically executed through an intelligent scheduling algorithm to ensure that it does not affect the core business processing of the main system.
[0072] When the load of the host system is high or a failure is about to occur, the standby system will actively release resources and focus on preparing to take over the critical tasks of the main system. This resource optimization strategy ensures that when the main system fails, the standby system can quickly take over all services, minimizing the service interruption time and ensuring the high availability of the system.
[0073] When the system detects a load peak or a failure warning in the main system, the standby system will immediately release the resources of non-critical tasks (such as idle computing resources, storage space, etc.), convert these resources to a standby state, and prepare to take over the tasks of the main system. This process is automatically carried out through an intelligent resource scheduling algorithm to avoid the delay caused by human intervention. The standby system will give priority to taking over the critical business processes in the main system, such as database access, user request processing, etc. The takeover process is optimized to ensure that at the moment of failure, the standby system can seamlessly take over and continue to provide services, minimizing the service interruption time.
[0074] The resource optimization of the standby system also includes the dynamic management of redundant resources. The system will intelligently adjust the resource redundancy level between the main and standby systems according to the real-time load situation. To cope with possible hardware failures, software anomalies, network delays, etc., the standby system makes resource backups in advance. Once a failure occurs, the resources can be quickly transferred to the standby system to ensure the continuity of services.
[0075] The technical solution of the embodiment of the present invention determines the host load data of the host system according to the original operation parameters; determines the load sharing decision of the standby system according to the host load data, a preset load threshold and a critical failure threshold, realizing an adaptive load balancing mechanism, which can dynamically adjust the task allocation between the main and standby systems according to the real-time system load situation and optimize the use of resources.
[0076] Embodiment III
[0077] Figure 3 This is a schematic structural diagram of a switching device for a dual - machine hot - standby system provided in Embodiment 3 of the present invention. As Figure 3 shown, the device includes:
[0078] An operating parameter acquisition module 301, configured to acquire the original operating parameters generated during the operation of the host system in the dual - machine hot - standby system;
[0079] A potential fault determination module 302, configured to determine whether there is a potential host fault in the host system according to the original operating parameters and a target analysis and judgment model obtained through pre - training, wherein the target analysis and judgment model is obtained through training of a machine learning model;
[0080] A standby system switching module 303, configured to, when there is a potential fault in the host system, direct the service - processing traffic of the host system to the standby system in the dual - machine hot - standby system.
[0081] The technical solution of the embodiment of the present invention acquires the original operating parameters generated during the operation of the host system in the dual - machine hot - standby system. According to the original operating parameters and a target analysis and judgment model obtained through pre - training, it is determined whether there is a potential host fault in the host system, wherein the target analysis and judgment model is obtained through training of a machine learning model. When there is a potential fault in the host system, the service - processing traffic of the host system is directed to the standby system in the dual - machine hot - standby system, avoiding mis - switching based on traditional heartbeat detection, reducing unnecessary switching operations, and improving system stability and switching accuracy.
[0082] Optionally, the potential fault determination module 302 includes:
[0083] A data pre - processing unit, configured to perform data pre - processing on the original operating parameters to obtain target operating parameters, where the data pre - processing includes but is not limited to data denoising, data standardization, data normalization, and time - series processing;
[0084] A potential fault determination unit, configured to determine whether there is a potential host fault in the host system according to the target operating parameters and a target analysis and judgment model obtained through pre - training.
[0085] Optionally, the potential fault determination unit includes:
[0086] A data feature extraction sub - unit, configured to perform data feature extraction on the target operating parameters to obtain feature operating parameters, where the data features at least include parameter trend, peak fluctuation, jitter frequency, and abnormal pattern;
[0087] A potential fault determination subunit, configured to input the characteristic operating parameters into the target analysis and judgment model for fault analysis, and determine whether there is a potential host fault in the host system according to the output of the target analysis and judgment model.
[0088] Optionally, the device further includes a model training module.
[0089] The model training module is configured to:
[0090] Obtain sample operating parameters and the corresponding operating status results of the sample operating parameters, where the sample operating parameters include normal operating parameters and abnormal operating parameters.
[0091] Based on the random forest algorithm, input the sample operating parameters into a preset machine learning model for fault judgment, and obtain an output judgment result based on the output of the preset machine learning model.
[0092] Determine a training error based on the output judgment result and the operating status result, and backpropagate the training error to the preset machine learning model to adjust the network parameters in the preset machine learning model.
[0093] When a preset convergence condition is satisfied, determine that the training of the preset machine learning model is completed, and obtain a target analysis and judgment model.
[0094] Optionally, the model training module is further configured to:
[0095] Obtain historical operating parameters within a preset time period.
[0096] Based on the historical operating parameters and the corresponding system operating status, optimize and train the target analysis and judgment model.
[0097] Optionally, the device further includes a load sharing decision module. Among them,
[0098] The load sharing decision module is configured to:
[0099] Determine the host load data of the host system according to the original operating parameters.
[0100] Determine the load sharing decision of the standby system according to the host load data, a preset load threshold, and a fault critical threshold.
[0101] Optionally, the load sharing decision module is specifically configured to:
[0102] In the case where the host load data is lower than the preset load threshold, determine that the standby system does not participate in load sharing.
[0103] When the host load data is higher than or equal to the preset load threshold and lower than the failure critical threshold, it is determined that the standby system participates in non-critical task load sharing;
[0104] When the host load data is higher than or equal to the failure critical threshold, it is determined that the standby system releases non-critical task load and prepares to take over the critical task load of the host system.
[0105] The switching device of the dual-machine hot standby system provided by the embodiments of the present invention can execute the switching method of the dual-machine hot standby system provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0106] Embodiment Four
[0107] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0108] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0109] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0110] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the switching method of the dual-machine hot standby system.
[0111] In some embodiments, the switching method of the dual-machine hot standby system can be implemented as a computer program, which is tangibly included in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the switching method of the dual-machine hot standby system described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the switching method of the dual-machine hot standby system by any other suitable means (e.g., by means of firmware).
[0112] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0113] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0114] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0115] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0116] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0117] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0118] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0119] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A switching method for a dual-machine hot standby system, characterized in that: include: Obtaining original operating parameters generated by the host system in the dual-machine hot standby system during operation; Determine whether the host system has a potential host failure according to the original operating parameters and a pre-trained target analysis and judgment model, wherein the target analysis and judgment model is obtained through machine learning model training; In the event that a potential failure occurs in the host system, the service processing traffic of the host system is directed to the standby system in the dual-machine hot standby system.
2. The method according to claim 1, characterized in that The determining whether the host system has a potential host failure according to the original operating parameters and the pre-trained target analysis and judgment model includes: Performing data preprocessing on the original operating parameters to obtain target operating parameters, wherein the data preprocessing includes but is not limited to data denoising, data standardization, data normalization and time series processing; According to the target operating parameters and the pre-trained target analysis and judgment model, it is determined whether the host system has a potential host failure.
3. The method according to claim 2, characterized in that The determining whether the host system has a potential host failure according to the target operating parameters and the pre-trained target analysis and judgment model includes: Extracting data features of the target operating parameters to obtain characteristic operating parameters, wherein the data features at least include parameter trends, peak fluctuations, jitter frequencies, and abnormal patterns; The characteristic operation parameters are input into the target analysis and judgment model to perform fault analysis, and based on the output of the target analysis and judgment model, it is determined whether there is a potential host fault in the host system.
4. The method according to claim 1, characterized in that: The training process of the target analysis and judgment model includes: Acquire sample operating parameters and operating status results corresponding to the sample operating parameters, wherein the sample operating parameters include normal operating parameters and abnormal operating parameters; Based on the random forest algorithm, the sample operation parameters are input into a preset machine learning model for fault judgment, and an output judgment result is obtained based on the output of the preset machine learning model; Determine a training error based on the output judgment result and the operating status result, and back-propagate the training error to the preset machine learning model to adjust network parameters in the preset machine learning model; When the preset convergence conditions are met, it is determined that the training of the preset machine learning model is completed and the target analysis and judgment model is obtained.
5. The method according to claim 1, characterized in that After determining whether the host system has a potential host failure according to the original operating parameters and the pre-trained target analysis and judgment model, the method further includes: Obtain historical operating parameters within a preset time period; Based on the historical operating parameters and the system operating status corresponding to the historical operating parameters, the target analysis and judgment model is optimized and trained.
6. The method according to claim 1, characterized in that The method further comprises: Determining host load data of the host system according to the original operating parameters; The load sharing decision of the backup system is determined according to the host load data, a preset load threshold and a critical failure threshold.
7. The method according to claim 6, characterized in that The step of determining the load sharing decision of the backup system according to the host load data, a preset load threshold and a critical failure threshold comprises: When the host load data is lower than the preset load threshold, determining that the standby system does not participate in load sharing; In the case where the host load data is higher than or equal to the preset load threshold and lower than the critical failure threshold, determining that the standby system participates in non-critical task load sharing; In a case where the host load data is higher than or equal to the critical failure threshold, it is determined that the backup system releases the non-critical task load and is ready to take over the critical task load of the host system.
8. A switching device for a dual-machine hot standby system, characterized in that: include: An operation parameter acquisition module is used to acquire original operation parameters generated by the host system in the dual-machine hot standby system during operation; A potential fault determination module, used to determine whether the host system has a potential host fault according to the original operating parameters and a pre-trained target analysis and judgment model, wherein the target analysis and judgment model is obtained through machine learning model training; The standby system switching module is used to guide the service processing traffic of the host system to the standby system in the dual-machine hot standby system when there is a potential failure in the host system.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the switching method of the dual-machine hot standby system according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the switching method of the dual-machine hot standby system according to any one of claims 1 to 7 when executed.