A method and device for automatically migrating a channel in a multi-data center environment to recover from disasters
By using monitoring clusters and network hooks to trigger migration modules in a multi-data center environment, the problem of untimely channel migration in the event of data center failure is solved, and the fault response speed and operation and maintenance efficiency are improved.
Patent Information
- Application Number
- CN202410028863.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-01-08
AI Technical Summary
In a multi-data center environment, when a data center is unavailable, all channels relying on the data center are unavailable, resulting in operation and maintenance personnel needing to manually migrate channels, which is costly and untimely migration, affecting the delay and user experience of SMS traffic.
By monitoring the cluster to trigger a webhook when a failure is discovered, the migration module will be triggered to automatically migrate, including obtaining channel metrics, determining faults and starting data synchronization tasks, and migrating the channel to a normal data center.
It realizes automatic migration of channels in a multi-data center environment, improves fault response speed, reduces manual operation of operation and maintenance, reduces information delay problems, and improves operation and maintenance efficiency.
Smart Images

Figure CN118055009B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disaster recovery migration of data centers, and in particular to the field of a method and device for automatically migrating channels for disaster recovery in a multi-data center environment. Background Art
[0002] A data center is a physical or virtual facility used to centrally store, manage, and process large amounts of data. It usually includes servers, storage devices, network devices, backup devices, and other components, and can provide efficient, secure, and reliable data storage and processing capabilities. Data centers are used in a wide range of applications. In the scenario of online text message delivery, multiple data centers are often built to divert traffic to different access customers or operators, so that text messages can flow in and out as close as possible. At the same time, multiple data centers can also improve the high availability of services. That is, when one of the data centers is unavailable (including abnormal situations caused by factors such as data center network outages, disk damage, or single-channel-operator link unavailability), the routing module will direct traffic to an available data center so that text messages can be delivered normally.
[0003] Most channels are single-connected and cannot use multiple processes to connect to operators, otherwise the operator will reject SMS messages, so the channel program can only be deployed in a certain data center.
[0004] Based on the above background, since the channel and the operator are connected through a single data center, when a data center is unavailable, all channels that rely on the data center are unavailable. At this time, it is necessary to migrate all channels of the current fault center to an available data center and activate the routing strategy of the migrated channels so that subsequent SMS traffic is forwarded to the available data center. However, the common practice now is to rely on operation and maintenance personnel to manually migrate the channels of the fault center to an available data center. However, when there are hundreds or thousands of channels on the faulty data center, the labor cost of operation and maintenance is high, and the channels cannot be activated in time, which will cause a lot of delays and unavailability, especially for SMS messages that require low latency (such as verification codes within 60 seconds), which will greatly reduce the user experience. Summary of the invention
[0005] Based on this, the purpose of the present invention is to provide a method and device for automatic migration of disaster recovery channels in a multi-data center environment, which uses a monitoring cluster to trigger a migration module through a configured webhook to perform automatic migration when a fault is discovered, thereby improving fault response and reducing manual operations in operation and maintenance.
[0006] The present invention provides a method for automatic channel migration and disaster recovery in a multi-data center environment, which includes:
[0007] Initialize the system and preset rules that can trigger webhook requests;
[0008] Get the channel indicators of each channel, and determine whether the rules for triggering the webhook request are met based on the channel indicators. If so, trigger the webhook request;
[0009] According to each network hook request, determine whether the corresponding data center is faulty and whether data migration is required. If data migration is required, start the data synchronization task, copy the channels that need to be transferred to the normal data center, and shut down all channels in the faulty data center;
[0010] Establish a connection between the migrated channel and the routing module of the data center where it is located after migration to complete the migration and reconnection of the channel.
[0011] Further, judging whether the corresponding data center is faulty and whether data migration is required according to each network hook request specifically includes:
[0012] Log in to the data center and determine the items to be verified based on the webhook request;
[0013] Perform user data simulation for each item to be tested, and test the data status of the node to be tested for each item to be tested;
[0014] A fault is determined according to the data status of the node, and whether data migration is required is determined according to the fault.
[0015] Furthermore, a plurality of rules that can trigger a webhook request are preset, and each rule that triggers a webhook request is provided with a matching warning rule.
[0016] Furthermore, in the process of initializing the system, an early warning rule is preset, and the early warning rule is used to issue an early warning signal if it is determined that a failure is about to occur in the channel after the channel indicator of each channel is obtained.
[0017] Furthermore, the warning signal is associated with a timing rule. When the warning signal is not processed within the event range preset by the timing rule, a data migration task is initiated to migrate the channels in the data center that issued the warning signal to a normal data center.
[0018] On the other hand, the present invention also provides a channel for automatically migrating disaster recovery equipment in a multi-data center environment, which includes:
[0019] Initialization module: used to initialize the system and preset rules that can trigger webhook requests;
[0020] Monitoring module: used to obtain the channel indicators of each channel, and determine whether the rules for triggering the webhook request are met based on the channel indicators. If so, the webhook request is triggered;
[0021] Migration module: used to determine whether the corresponding data center is faulty and whether data migration is required based on each network hook request. If data migration is required, the data synchronization task is started to copy the channels that need to be transferred to the normal data center, and all channels in the faulty data center are shut down;
[0022] Routing connection module: used to establish a connection between the migrated channel and the routing module of the data center where it is located after migration, so as to complete the migration and reconnection of the channel.
[0023] Furthermore, the monitoring module also includes an early warning unit, which is used to issue an early warning signal if it is determined that a failure is about to occur in the channel after obtaining the channel index of each channel according to a preset early warning rule.
[0024] Furthermore, the monitoring module also includes a timing unit, which is connected to the early warning unit; the timing unit is used to start a data migration task when the early warning signal is not processed within the event range preset by the timing rule, and migrate the channels in the data center that issued the early warning signal to a normal data center.
[0025] Furthermore, the monitoring module also includes a timing unit, which is connected to the early warning unit; the timing unit is used to start a data migration task when the early warning signal is not processed within the event range preset by the timing rule, and migrate the channels in the data center that issued the early warning signal to a normal data center.
[0026] In another aspect, the present invention further provides an electronic device, comprising:
[0027] at least one memory and at least one processor;
[0028] The memory is used to store one or more programs;
[0029] When the one or more programs are executed by the at least one processor, the at least one processor implements the steps of any one of the above-mentioned methods for automatic channel migration and disaster recovery in a multi-data center environment.
[0030] On the other hand, the present invention also provides a computer-readable storage medium, which stores a computer program, characterized in that when the computer program is executed by a processor, the steps of a method for automatic migration and disaster recovery of a channel in a multi-data center environment as described in any one of the above are implemented.
[0031] The present invention uses a monitoring cluster to monitor the data status of each channel, combined with a plurality of preset triggering rules. When the channel index detected by the monitoring cluster meets any triggering rule, the migration module is triggered through the corresponding configured network hook. The migration module further verifies the abnormality of the data center based on the abnormal information carried by the network hook, thereby avoiding the situation of mis-migration and automatically migrating all channels in the abnormal data center to a normal data center when the migration conditions are met, thereby improving the efficiency of operation and maintenance and solving the problem of information delay caused by abnormalities. In addition, the present invention also adds early warning rules and early warning rules on this basis, so that the present invention can not only automatically migrate when a fault that requires migration occurs, but also report some abnormal conditions that occur in the data center, realize pre-inspection and maintenance of the data center, prevent the occurrence of faults, and reduce the scope of fault inspection for operation and maintenance personnel, thereby improving the efficiency of operation and maintenance.
[0032] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is a schematic diagram of the existing client SMS sending process;
[0034] Figure 2 A flow chart of a method for automatic channel migration and disaster recovery in a multi-data center environment provided by the present invention;
[0035] Figure 3 For execution Figure 2 A structural block diagram of a channel automatic migration disaster recovery device of a channel automatic migration disaster recovery method in a multi-data center environment; DETAILED DESCRIPTION
[0036] In the data transmission of daily life, considering the efficiency of data distribution, multiple data centers are often built to exchange data with different access customers or operators. The present invention takes the scenario of SMS distribution as an example for explanation. The channel in the present invention refers to the data channel. The data channel refers to the transmission channel from the data source to the data user, which can be physical or virtual. The data channel can be a network, a disk, a mobile device, or a paper document. The data channel can be unidirectional or bidirectional, and they can support the transmission, storage, processing and analysis of data. The data channel can support a variety of data formats, such as text, images, video, audio, etc., and they can support a variety of data transmission protocols, such as HTTP, FTP, SMTP, etc. The data channel can support a variety of data security technologies, such as encryption, digital signatures, data integrity, etc., to ensure the security and reliability of the data.
[0037] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the existing client text message sending process. It can be seen that after the client sends a text message, the routing module receives the text message for wireless communication. The routing module is usually selected according to the communication distance, and then the data center selects an internal channel to transmit the client text message to the corresponding operator. When data center 1 fails, the text messages sent through the channels in data center 1 cannot reach the operator. The existing operation and maintenance methods are mostly to manually migrate the channels in the failed data center by hand. When there are hundreds or thousands of channels in the failed data center, the labor cost of operation and maintenance is high, and the channels cannot be enabled in time. In addition, this kind of failure is different from the conventional disaster recovery problem. The conventional disaster recovery problem is to back up data in advance, and restore it in the backup of another data center after a data center fails or is damaged due to some force majeure factors. The present invention is only for automatic migration disaster recovery of data channels, which has high requirements for real-time performance. In the case of a large number of data centers, it is impossible for each data center to perform data backup. Therefore, a new automatic migration solution is needed to complete the automatic migration disaster recovery of channels.
[0038] The inventor has found through research that it is possible to use a monitoring cluster (which can be a third-party monitoring middleware or a self-built monitoring middleware) to trigger a migration module to automatically migrate when a fault is found through a configured hook (webhook), thereby improving fault response and reducing manual operations of operation and maintenance. The present invention takes the automatic migration of channels between two data centers (data center 1 and data center 2) as an example, and combines Figure 2 , Figure 3 To explain, Figure 2 A flow chart of a method for automatic migration and disaster recovery of a channel in a multi-data center environment provided by the present invention, Figure 3 For execution Figure 2 A structural block diagram of a channel automatic migration disaster recovery device of a channel automatic migration disaster recovery method in a multi-data center environment, wherein the channel automatic migration disaster recovery method in a multi-data center environment specifically comprises the following steps:
[0039] S10: Initialize the system and preset rules that can trigger the webhook request. Step S10 is performed by the initialization module 10.
[0040] A monitoring cluster is set up to detect the operating status of each data center, and the connection status between the data center and the operator is determined through the preset rules of the monitoring cluster. For example, when data exchange occurs between the channel and the operator, it is necessary to try to connect the channel and the operator first, but due to network links, data transmission speed and other issues, there are occasional connection failures; therefore, in this embodiment, the exemplary setting of Rule 1 for triggering the network hook is: when the number of connection failures between the channel and the operator is greater than 10, it is determined that there may be a link failure between the data center and the operator; or Rule 2: when the disk is inaccessible during the channel status pull, it is determined that the channel may have a fault and cannot be used, and Rule 3: the monitoring cluster cannot pull the reported status of channel 2, and it is determined that channel 2 may have a fault and cannot be used.
[0041] The webhook is a method of adding or changing the performance of a web page through a custom callback function. These callbacks can be saved, modified and managed by third-party users and developers who may be related to the original website or application. In the present invention, the values returned by all webhooks correspond to the data indicators received by the monitoring cluster, that is, the values of the webhooks are associated with the status of the data center.
[0042] S20: Obtain the channel index of each channel, and determine whether the rule for triggering the web hook request is met according to the channel index, and if so, trigger the web hook request. Step S20 is performed by the monitoring module 20.
[0043] In the present invention, since the target migration object is the channel, it is very complicated to monitor the status of the data center, and a lot of invalid information will be obtained; for example, although an abnormality occurs in the data center, its abnormality does not affect the sending of text messages. At this time, the abnormality will also be judged by rules, but after investigation, migration is not required, which wastes system resources. Therefore, the present invention uses a monitoring cluster to monitor each channel, only obtains data indicators related to the sending of text messages, and performs targeted abnormality detection, which saves system resources on the one hand and improves efficiency on the other.
[0044] S30: According to each webhook request, determine whether the corresponding data center is faulty and whether data migration is required. If data migration is required, start the data synchronization task, copy the channels that need to be transferred to the normal data center, and shut down all channels in the faulty data center. Step S30 is performed by the migration module 30.
[0045] When the monitoring cluster detects a triggering warning signal, it only preliminarily determines that there is an abnormality in the connection between the data center and the operator. However, not all abnormalities that trigger warnings require channel migration. For example, if the link between the current data center and the operator is not connected, migration is required. If there is a problem with the channel process itself, it is not the abnormality of the data center that hinders the data transmission between the channel and the operator, and channel migration is not required. Or the operator interface has been closed, and the channel cannot obtain the object to establish a connection with the operator, and there is no need to connect at this time.
[0046] Therefore, some measures need to be taken to further verify it. Specifically,
[0047] Log in to the data center and determine the items to be verified based on the webhook request;
[0048] Perform user data simulation for each item to be tested, and test the data status of the node to be tested for each item to be tested;
[0049] A fault is determined according to the data status of the node, and whether data migration is required is determined according to the fault.
[0050] Take Rule 1 mentioned in step S10 as an example, combined with Figure 1 As an example, log in to all data centers through SSH and use the operator connection tool to connect to the operator. The following situations may occur:
[0051] Data Center 1 is successfully connected (the link to Data Center 1 is connected, which may be a problem with the Channel 2 process and does not require migration).
[0052] Data center 1 fails to connect, and data center 2 also fails to connect (all data center links are blocked, possibly because the carrier interface is closed and no migration is required).
[0053] Data Center 1 connection failed, but Data Center 2 connection succeeded (the link between Data Center 1 and Carrier 2 was disconnected, and Channel 2 was migrated to Data Center 2).
[0054] There are multiple items to be verified in the signal of the same network hook. Only after further verification can the situations where migration is not required be eliminated to avoid the occurrence of mistaken migration.
[0055] S40: Establish a connection between the migrated channel and the routing module of the data center where the migrated channel is located, so as to complete the migration and reconnection of the channel. Step S40 is performed by the routing reconnection module 40.
[0056] Since the channel may be deployed in many ways, you must first confirm the deployment method of the channel before shutting down and migrating it. If the channel is a Java process deployed on a virtual machine, first log in to data center 2 through SSH, copy the jar package (using the Linux command scp or other methods, such as svn co, git clone, etc.) to the normal data center 2, and then use the startup command (java-jar channel 2.jar) to start it.
[0057] If the channel is deployed on a k8s container, first log in to data center 2 through ssh, copy the k8s channel 2.yaml file deployed in the channel to the node of data center 2 (using scp, svn co, git clone, etc.), and then use the k8s command kubectl apply-f channel 2.yaml to start it.
[0058] After the shutdown and migration are completed, the migrated channel will be connected to the routing module of the data center where it is migrated, so that the client can use the channel normally.
[0059] In another embodiment, in order for the operation and maintenance personnel to promptly perform maintenance on the data center anomalies, multiple rules that can trigger network hook requests are preset, and each rule that triggers the network hook request is provided with a matching early warning rule. On the one hand, even if the migration is not triggered, the data center can be checked and maintained in advance to prevent the occurrence of failures. On the other hand, it can reduce the scope of troubleshooting for the operation and maintenance personnel and improve the operation and maintenance efficiency.
[0060] In another embodiment, an early warning rule is preset during the system initialization process, and the early warning rule is used to issue an early warning signal based on a possible fault that is about to occur in the channel.
[0061] For automatic migration disaster recovery, the focus is on the timeliness and reliability of migration. In this embodiment, the inventor sets some early warning rules, which can determine that the connection between the data center or the channel and the data center may reach the critical point of failure based on the data obtained by the monitoring cluster, and remind the operation and maintenance personnel to make a judgment or directly migrate manually through early warning signals, further avoiding the problem that the channel cannot be activated in time, causing a large number of delays and unavailability. Furthermore, a timing rule is preset. After a period of waiting for a response, if the early warning signal is not processed, the automatic migration program is started to migrate the corresponding channel to a normal data center.
[0062] The present invention uses a monitoring cluster to monitor the data status of each channel, combined with a plurality of preset triggering rules. When the channel index detected by the monitoring cluster meets any triggering rule, the migration module is triggered through the corresponding configured network hook. The migration module further verifies the abnormality of the data center based on the abnormal information carried by the network hook, thereby avoiding the situation of mis-migration and automatically migrating all channels in the abnormal data center to a normal data center when the migration conditions are met, thereby improving the efficiency of operation and maintenance and solving the problem of information delay caused by abnormalities. In addition, the present invention also adds early warning rules and early warning rules on this basis, so that the present invention can not only automatically migrate when a fault that requires migration occurs, but also report some abnormal conditions that occur in the data center, realize pre-inspection and maintenance of the data center, prevent the occurrence of faults, and reduce the scope of fault inspection for operation and maintenance personnel, thereby improving the efficiency of operation and maintenance.
[0063] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for automatic channel migration and disaster recovery in a multi-data center environment described in any one of the above embodiments is implemented.
[0064] The present invention may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0065] The above-mentioned embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, and the present invention is also intended to include these modifications and modifications.
Claims
1. A method for automatic migration and disaster recovery of a channel in a multi-data center environment, characterized in that: include: Initialize the system and preset rules that can trigger webhook requests; Get the channel indicators of each channel, and determine whether the rules for triggering the webhook request are met based on the channel indicators. If so, trigger the webhook request; Log in to the data center and determine the items to be checked according to the network hook request; perform user data simulation for each item to be checked and check the data status of the node to be checked for each item to be checked; determine the fault according to the data status of the node, and finally determine whether data migration is required according to the fault; If data migration is required, start the data synchronization task, copy the channels that need to be transferred to the normal data center, and shut down all channels in the faulty data center; Establish a connection between the migrated channel and the routing module of the data center where it is located after migration to complete the migration and reconnection of the channel.
2. According to claim 1, a method for automatic channel migration and disaster recovery in a multi-data center environment is characterized by: Multiple rules that can trigger webhook requests are preset, and each rule that triggers webhook requests has a matching warning rule.
3. The method for automatic channel migration and disaster recovery in a multi-data center environment according to claim 2, characterized in that: In the process of initializing the system, an early warning rule is also preset. The early warning rule is used to issue an early warning signal if it is determined that a failure is about to occur in the channel after the channel indicators of each channel are obtained.
4. The method for automatic channel migration and disaster recovery in a multi-data center environment according to claim 3, characterized in that: The warning signal is associated with a timing rule. When the warning signal is not processed within the time range preset by the timing rule, a data migration task is initiated to migrate the channels in the data center that issued the warning signal to a normal data center.
5. A channel automatically migrates disaster recovery equipment in a multi-data center environment, characterized in that: include: Initialization module: used to initialize the system and preset rules that can trigger webhook requests; Monitoring module: used to obtain the channel indicators of each channel, and determine whether the rules for triggering the webhook request are met based on the channel indicators. If so, the webhook request is triggered; Migration module: used to log in to the data center and determine the items to be verified according to the network hook request; perform user data simulation for each item to be detected and detect the data status of the node to be detected for each item to be detected; determine the fault according to the data status of the node, and finally determine whether data migration is required according to the fault. If data migration is required, start the data synchronization task, copy the channels to be transferred to the normal data center, and shut down all channels in the faulty data center; Routing connection module: used to establish a connection between the migrated channel and the routing module of the data center where it is located after migration, so as to complete the migration and reconnection of the channel.
6. According to claim 5, a channel automatically migrates disaster recovery equipment in a multi-data center environment, characterized in that: The monitoring module further comprises an early warning unit, which is used to send out an early warning signal if it is determined that a fault is about to occur in a channel after acquiring the channel index of each channel according to a preset early warning rule.
7. A channel according to claim 6 automatically migrates disaster recovery equipment in a multi-data center environment, characterized in that: The monitoring module also includes a timing unit, which is connected to the early warning unit; the timing unit is used to start a data migration task when the early warning signal is not processed within a preset time range, and migrate the channels in the data center that issued the early warning signal to a normal data center.
8. An electronic device, characterized in that: include: at least one memory and at least one processor; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the at least one processor implements the steps of a channel automatic migration disaster recovery method in a multi-data center environment as described in any one of claims 1 to 4.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, the steps of a method for automatic channel migration and disaster recovery in a multi-data center environment as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Webhook notification method, device and equipment based on cloud platform and storage medium
CN112311593A
Switching method and device of multi-active data center
CN114285864A