Disaster Tolerance Method, Device and Storage Medium for Business Management Platform

By sending test messages between multiple sub-platforms of the business management platform and generating switching instructions, and automatically switching IP addresses using GSLB, the problems of high costs and geographical restrictions in the existing technology are solved, and efficient disaster recovery switching is achieved.

CN114253774BActive Publication Date: 2025-06-24CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111391905.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2025-06-24
Estimated Expiration
2041-11-23

AI Technical Summary

Technical Problem

The disaster recovery solutions of existing business management platforms have problems such as high software development and labor costs, manual switching causes business congestion and loss, and restricted deployment areas of backup platforms.

Method used

By sending test messages between multiple sub-platforms, determining the detection results, and generating switching instructions to send them to the global load balancing device GSLB, automatically switching IP addresses is achieved to avoid regional restrictions.

Benefits of technology

Automatic switching is achieved, reducing software development and labor costs, avoiding business blockages, and getting rid of geographical restrictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114253774B_ABST
    Figure CN114253774B_ABST
Patent Text Reader

Abstract

The present application provides a disaster recovery method, device and storage medium for a service management platform. The method includes, for each of multiple sub-platforms, sending a test message to the sub-platform and determining a detection result according to the response result returned by the sub-platform, generating a first switching instruction according to the detection results of the multiple sub-platforms, and sending the first switching instruction to the GSLB of the sub-platform that fails among the multiple sub-platforms, so that the GSLB, according to the first switching instruction, switches the IP address of the failed sub-platform from its own production IP address to the production IP address of the non-failed sub-platform. The production IP address is the IP address assigned to the sub-platform for accessing its own sub-platform. The present application saves the development cost of network management software, realizes automatic switching, improves efficiency, and since the GSLB is used for IP address switching, the sub-platforms do not need to be set in the same region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a disaster recovery method, device, and storage medium for a service management platform. Background Art

[0002] With the development of communication technologies, the service volume carried by the service management platform is increasing. Users have higher and higher requirements for the security of the service management platform.

[0003] In the prior art, the service management platform can adopt a 1+1 disaster recovery method to ensure service security. The main implementation methods are as follows: The first is load sharing, where two platforms simultaneously undertake services. When the network management system monitors that one of them fails, it notifies the operation and maintenance personnel to modify the resolution address on the Domain Name Server (DNS), and points the DNS resolution address to the fault-free platform. The second is active-standby disaster recovery. There are two platforms, one is the active platform and the other is the standby platform. The two platforms share the same IP. After the active platform fails, the standby platform undertakes the service through the same IP.

[0004] However, in the process of implementing this application, the inventor found that there are at least the following problems in the prior art: In the first load sharing solution, it is necessary to develop dedicated network management software and manually modify the DNS resolution address, which incurs high software development costs and labor costs. Moreover, with the growth of service volume, a large amount of service congestion and loss will occur during the manual switchover. In the second active-standby disaster recovery solution, both the active platform and the standby platform use the same IP address. Due to the IP address segment allocation problem, it is generally required that the devices be set under the same network device, which limits the deployment area of the standby platform. Summary of the Invention

[0005] This application provides a disaster recovery method, device, and storage medium for a service management platform to reduce software development costs and labor costs, achieve automatic switchover between platforms, and not limit the deployment area of the platforms.

[0006] In a first aspect, this application provides a disaster recovery method for a service management platform. The service management platform includes multiple sub-platforms; the data of the multiple sub-platforms is synchronized; the method includes:

[0007] For each of the multiple sub-platforms, send a test message to the sub-platform, and determine a detection result according to the response result returned by the sub-platform based on the test message;

[0008] Generate a first switching instruction based on the detection results of multiple said sub-platforms, and send the first switching instruction to the Global Server Load Balancing (GSLB) device of the sub-platform that has failed among the multiple said sub-platforms, so that the GSLB switches the IP address of the failed sub-platform from its own production IP address to the production IP address of the sub-platform that has not failed; the production IP address is the IP address assigned to the sub-platform for accessing its own sub-platform.

[0009] In a possible design, multiple said sub-platforms are respectively assigned multiple production IP addresses; the multiple production IP addresses include a first production IP address and a second production IP address; wherein the first production IP address is used for access when its own sub-platform is fault-free, and the second production IP address is used to switch the IP address of other sub-platforms to the second production IP address when other sub-platforms fail.

[0010] In a possible design, the test message is any one of the following: a notification message sent down, an upload of a multimedia file, a download of a multimedia file.

[0011] In a possible design, sending the test message to the sub-platform includes:

[0012] Sending the test message to the sub-platform periodically.

[0013] In a possible design, determining the detection result according to the response result returned by the sub-platform based on the test message includes:

[0014] If a correct response packet is not received within a first preset time, it is determined that the message sending fails, and a detection result of not passing the detection is obtained;

[0015] If a correct response packet is received within a first preset time, wait to receive the delivery result report of the test message;

[0016] If the delivery result report is not received within a second preset time, it is determined that the message reception fails, and a detection result of not passing the detection is obtained.

[0017] In a possible design, sending the test message to the sub-platform includes:

[0018] Sending the test message to the sub-platform through multiple accounts; the multiple accounts respectively belong to networks in different regions;

[0019] Determining the detection result according to the response result returned by the sub-platform based on the test message includes:

[0020] If the number of accounts that fail the detection among the multiple accounts is greater than a preset threshold, it is determined that the detection result of the sub-platform fails the detection.

[0021] In a possible design, the multiple sub-platforms include a first sub-platform and a second sub-platform; the generating a first switching instruction according to the detection result includes:

[0022] If the consecutive number of times the detection result of the first sub-platform fails the detection is greater than a first preset number of times, and the consecutive number of times the detection result of the second sub-platform passes the detection is greater than a second preset number of times, it is determined that the first sub-platform fails and the second sub-platform does not fail, and a first switching instruction is generated to enable GSLB to switch the IP address of the first sub-platform from the production IP address of the first sub-platform to the production IP address of the second sub-platform.

[0023] In a possible design, after generating the first switching instruction according to the detection results of the multiple sub-platforms, it further includes:

[0024] If the consecutive time that the detection result of the first sub-platform passes the detection is greater than a third preset time, it is determined that the failure of the first sub-platform is resolved, and a second switching instruction is generated to enable GSLB to switch the IP address of the first sub-platform from the production IP address of the second sub-platform back to the production IP address of the first sub-platform according to the second switching instruction.

[0025] In a second aspect, the present application provides a disaster recovery device for a service management platform, including:

[0026] A detection module, configured to send a test message to each of the multiple sub-platforms, and determine a detection result according to a response result returned by the sub-platform based on the test message;

[0027] A switching module, configured to generate a first switching instruction according to the detection results of the multiple sub-platforms, and send the first switching instruction to the global load balancing device GSLB of the sub-platform that fails among the multiple sub-platforms, so that GSLB switches the IP address of the failed sub-platform from its own production IP address to the production IP address of the non-failed sub-platform according to the first switching instruction; the production IP address is the IP address assigned to the sub-platform for accessing its own sub-platform.

[0028] In a third aspect, the present application provides a disaster recovery device for a service management platform, including: at least one processor and a memory;

[0029] The memory stores computer-executable instructions;

[0030] The at least one processor executes the computer-executable instructions stored in the memory, such that the at least one processor executes the method as described in the first aspect above and various possible designs of the first aspect.

[0031] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method as described in the first aspect above and various possible designs of the first aspect.

[0032] In a fifth aspect, the present application provides a computer program product including a computer program, which, when executed by a processor, implements the method as described in the first aspect above and various possible designs of the first aspect.

[0033] The disaster recovery method, device, and storage medium of the service management platform provided by the present application. The method includes, for each of the multiple sub-platforms, sending a test message to the sub-platform, determining a detection result according to the response result returned by the sub-platform based on the test message, generating a first switching instruction according to the detection results of the multiple sub-platforms, and sending the first switching instruction to the global load balancing device GSLB of the sub-platform that fails among the multiple sub-platforms, so that the GSLB, according to the first switching instruction, switches the IP address of the failed sub-platform from its own production IP address to the production IP address of the non-failed sub-platform, where the production IP address is the IP address assigned to the sub-platform for accessing its own sub-platform. The disaster recovery switching method of the service management platform provided by the present application obtains the detection results of each sub-platform by sending test messages to each sub-platform respectively, and generates a switching instruction according to the detection results of each sub-platform and sends it to the GSLB, so as to automatically switch the IP address of the failed sub-platform to the IP address of the non-failed sub-platform through the GLSB. It saves the development cost of network management software, realizes automatic switching, saves labor costs, improves efficiency, and avoids service blockage during switching. Since the GSLB is used for IP address switching, each sub-platform does not need to be set in the same region, getting rid of the regional restriction. Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a schematic diagram of an application scenario of a disaster recovery method for a service management platform provided by an embodiment of the present application;

[0036] Figure 2 It is a schematic flow chart of the disaster recovery method for the service management platform provided by the embodiments of the present application;

[0037] Figure 3 It is a schematic flow chart of the working process of GSLB provided by the embodiments of the present application;

[0038] Figure 4 Provided by the embodiments of the present application Figure 2 The specific flow schematic diagram of step 201 in

[0039] Figure 5 It is a schematic structural diagram of the disaster recovery device of the service management platform provided by the embodiments of the present application;

[0040] Figure 6 It is a block diagram of a disaster recovery device for a service management platform provided by the embodiments of the present application. Detailed implementation manners

[0041] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0042] With the development of communication technology, the service volume carried by the service management platform is increasing. Users' requirements for the security of the service management platform are getting higher and higher.

[0043] In the prior art, the business management platform can adopt the 1+1 disaster recovery method to ensure business security. Taking Messaging as a Platform (MaaP) as an example, there are the following two implementation methods for the 1+1 disaster recovery of MaaP: The first is load sharing. Specifically, two sets of MaaP platforms undertake the business simultaneously. When the first set of MaaP platform fails, the second set of MaaP platform can learn about the failure through heartbeats and notify the operation and maintenance personnel, so that the operation and maintenance personnel can modify the resolution address on the Domain Name Server (DNS) and point the DNS resolution address to the healthy MaaP platform, that is, the second set of MaaP platform. After the failure is recovered, the DNS resolution address is pointed back to the first set of MaaP platform. The second is primary and standby disaster recovery. Specifically, for two sets of MaaP, one is used as the primary MaaP and the other is used as the standby MaaP. When working normally, only the primary MaaP platform undertakes the business. When the primary MaaP fails, the standby MaaP starts and takes over the business using the same IP address. However, in the first load sharing scheme, it is necessary to manually modify the DNS resolution address. On the one hand, this increases the workload of the operation and maintenance personnel. On the other hand, with the growth of the 5G message traffic, a large amount of business congestion and loss will occur during the manual switching. In the second primary and standby disaster recovery scheme, the primary MaaP platform and the standby MaaP platform both use the same IP address, which can achieve seamless switching of the business. However, due to the IP address segment allocation problem, it is generally required that the devices be set under the same network device, which limits the deployment area of the standby device.

[0044] To solve the above technical problems, the inventors of the present application have found that by deploying at least two sets of platforms in different regions, the two sets of platforms are used as primary and standby for each other, and the two sets of platforms keep the data (such as local data and Chatbot data of the chat robot) completely consistent. The failure is discovered by sending test messages to each sub-platform to obtain the detection results, and the IP address switching during domain name access is realized through the domain name resolution method of GSLB. Based on this, this embodiment provides a disaster recovery method for a business management platform. By sending test messages to each sub-platform respectively to obtain the detection results of each sub-platform, and generating a switching instruction according to the detection results of each sub-platform and sending it to the GSLB of the failed sub-platform, so as to automatically switch the IP address of the failed sub-platform to the IP address of the non-failed sub-platform through GLSB. This saves the development cost of the network management software, realizes automatic switching, saves labor costs, improves efficiency, and avoids business blockage during the switching. Since the GSLB is used for IP address switching, each sub-platform does not need to be set in the same region, getting rid of the regional restriction.

[0045] Figure 1 It is a schematic diagram of the application scenario of a disaster recovery method for a business management platform provided by an embodiment of the present application. AsFigure 1 As shown, the scenario includes a detection and switching device and a service management platform. Among them, the service management platform, such as the MaaP platform, may include multiple sub-platforms. For example, it may include a first sub-platform set in Computer Room A and a second sub-platform set in Computer Room B. Each sub-platform includes a database server DB, a core server, etc. Server Load Balancer (SLB), Global Server Load Balancer (GSLB), a firewall (not shown), a switch (not shown), and other network devices are provided for both the first sub-platform and the second sub-platform. During normal operation, the sub-platforms in the two computer rooms, as well as network devices such as the firewall, SLB, and GSLB, are all in an active state, each processing the services it accesses to achieve load balancing. The switches (not shown) between the two locations are connected by a dedicated line for service balance distribution and data synchronization. The detection and switching device is used to send test messages to each sub-platform for fault detection and generate a switching instruction according to the detection result to enable the GSLB to perform an IP switch. It should be noted that the detection and switching device may include multiple devices. Some devices are used to send test messages to each sub-platform for fault detection, and some devices are used to generate a switching instruction according to the detection result to enable the GSLB to perform an IP switch. This embodiment does not limit this.

[0046] In the specific implementation process, when external network elements such as the chatbot Chatbot, 5G Message Center (5GMC), and User Equipment (UE) access the service management platform and the MaaP platform through domain names, domain name resolution of the Domain Name Server (DNS) needs to be performed, that is, the IP address corresponding to the domain name of the platform is found, so that after the external network element establishes a connection with the IP address, it can send service requests such as portal access or interface calls. Specifically, the GSLB of the platform can be used as the authorized domain name resolution server, and the access switch is realized through its global load balancing function. The process is as follows: The external network element sends a query request for domain name resolution to the local DNS. The local DNS reads the NS record preset on the upper-level DNS of the GSLB through a series of internal DNS queries, and hands over the domain name resolution work to the GSLB for processing. Among them, the NS record is used to point to the IP address located in the GSLB. After receiving the query request sent by the local DNS, the GSLB calculates and returns the IP address of the selected sub-platform to the local DNS. The IP address can be the production IP address of the platform itself or the production IP address of other sub-platforms switched to after the platform fails. Based on the IP address of the selected sub-platform returned by the local DNS, the external network element performs subsequent operations such as portal access or interface calls. When the above work is carried out normally, the detection and switching device can send test messages to each sub-platform to detect whether each sub-platform fails, and generate a switching instruction according to the detection result, and send the switching instruction to the GSLB. The GSLB switches the production IP address of the failed sub-platform to the production IP address of other non-failed sub-platforms. For example, if the first sub-platform fails, the GSLB can switch the production IP address of the first sub-platform to the production IP address of the non-failed second sub-platform, so that when the external network element accesses the first sub-platform, it actually accesses the non-failed second sub-platform, avoiding the normal service processing of the external network element being affected by the failure of the first sub-platform.

[0047] The disaster recovery method of the service management platform provided in this embodiment automatically switches the IP address of the failed sub-platform to the IP address of the non-failed sub-platform through the GLSB. It saves the development cost of the network management software, realizes automatic switching, saves labor costs, improves efficiency, and avoids service blockage during switching. Since the GSLB is used for IP address switching, each sub-platform does not need to be set in the same region, getting rid of the geographical restriction.

[0048] The technical solution of the present application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0049] Figure 2 This is a schematic flowchart of the disaster recovery method for the service management platform provided by the embodiments of the present application. As Figure 2 shown, the service management platform includes multiple sub-platforms; the data of the multiple sub-platforms is synchronized. The method includes:

[0050] 201. For each of the multiple sub-platforms, send a test message to the sub-platform, and determine a detection result according to the response result returned by the sub-platform based on the test message.

[0051] In this embodiment, the test message may be any one of the following: sending a notification message, uploading a multimedia file, and downloading a multimedia file.

[0052] In this embodiment, the sending the test message to the sub-platform may include: sending the test message to the sub-platform through multiple accounts; the multiple accounts respectively belong to networks in different regions; the determining the detection result according to the response result returned by the sub-platform based on the test message includes: if the number of accounts whose detection results are unqualified among the multiple accounts is greater than a preset threshold, it is determined that the detection result of the sub-platform is unqualified.

[0053] Specifically, the execution subject of this step may be a data processing device installed with probe server software. Taking the Figure 1 MaaP platform shown as an example, the probe server can simulate 2 Chatbots, which are respectively registered on the first sub-platform in computer room A and the second sub-platform in computer room B of 2 sets of MaaP platforms. Usually, one Chatbot is used to inspect the A system of the first sub-platform, and the other inspects the B system of the second sub-platform. To avoid the GSLB returning the IP of the nearest sub-platform according to the principle of proximity, the probe server can be set to access the first sub-platform and the second sub-platform without using a domain name, but to perform fault detection through the real IP addresses of the first sub-platform and the second sub-platform themselves. After obtaining the detection result, send the detection result to the switching server so that the switching server generates a switching instruction according to the detection result. The test message may include sending a notification message, uploading and downloading multimedia files. The purpose is for the probe server to simulate an external network element to complete a complete communication process.

[0054] One or multiple detection servers can be deployed. For example, one detection server can be deployed in each of two different regions. By deploying multiple detection servers, false alarms caused by network problems of the detection servers themselves can be avoided. In the case of deploying multiple detection servers, when switching servers, if only the detection results indicating non - passing the detection from one detection server are received, the first switching instruction is not directly generated. Only when the detection results indicating non - passing the detection are reported by both sets of detection servers, will the first switching instruction be generated.

[0055] In addition, for the Chatbot account used to send test messages and the test mobile phone number used to receive test messages, they can be preset as test numbers in the MaaP platform in advance, only generating message logs and no charging bills.

[0056] 202. Generate a first switching instruction based on the detection results of multiple said sub - platforms, and send the first switching instruction to the Global Server Load Balancing (GSLB) device of the faulty sub - platform among the multiple said sub - platforms, so that the GSLB, according to the first switching instruction, switches the IP address of the faulty sub - platform from its own production IP address to the production IP address of the non - faulty sub - platform; the production IP address is the IP address assigned to the sub - platform for accessing its own sub - platform.

[0057] Optionally, multiple production IP addresses are respectively assigned to the multiple said sub - platforms; the multiple production IP addresses include a first production IP address and a second production IP address; where the first production IP address is used for access when its own sub - platform is fault - free, and the second production IP address is used to switch the IP address of other sub - platforms to the second production IP address when other sub - platforms have faults.

[0058] Optionally, the multiple said sub - platforms include a first sub - platform and a second sub - platform; generating the first switching instruction according to the detection results may include: if the consecutive number of times the detection result of the first sub - platform indicates non - passing the detection is greater than a first preset number of times, and the consecutive number of times the detection result of the second sub - platform indicates passing the detection is greater than a second preset number of times, then it is determined that the first sub - platform has a fault and the second sub - platform is fault - free, and a first switching instruction is generated, so that the GSLB switches the IP address of the first sub - platform from the production IP address of the first sub - platform to the production IP address of the second sub - platform.

[0059] Optionally, after generating the first switching instruction according to the detection results of the multiple sub-platforms, it may also include: if the detection result of the first sub-platform is that the continuous time of passing the detection is greater than a third preset time, it is determined that the fault of the first sub-platform is resolved, and a second switching instruction is generated, so that GSLB switches the IP address of the first sub-platform from the production IP address of the second sub-platform back to the production IP address of the first sub-platform according to the second switching instruction.

[0060] Specifically, generating the first switching instruction based on the detection results of the multiple sub-platforms can be completed by the switching server software. The switching server and the detection server can run on different terminal devices respectively. The first switching instruction or the second switching instruction is generated according to the detection results sent by the detection server, so that the GSLB performs the IP address switching operation according to the switching instruction.

[0061] The switching server can generate corresponding switching instructions according to pre-set policy rules. In order to avoid false positives of detection results caused by network reasons, the switching server will collect detection results from multiple detection servers at the same time and comprehensively determine whether the switching operation should be triggered.

[0062] Exemplarily, the basic parameters that can be used to set the policy may include at least one of the following: a detection script ID, a maximum allowed number of consecutive failures, a detection server ID, an ID of the MaaP platform being tested, and the like.

[0063] Exemplarily, the pre-set policy rules may include the following possible designs:

[0064] In the first possible design, it is assumed that two sets of detection servers are used to detect the first sub-platform and the second sub-platform respectively, and the first sub-platform is detected three times in a row and all the results are failed, that is, the detection fails, and the second sub-platform is detected three times in a row and all the results are passed, that is, the detection is successful, then the switching server generates a first switching instruction, so that GSLB switches the IP address of the first sub-platform from its own production IP address to the production IP address of the second sub-platform according to the first switching instruction.

[0065] In the second possible design, it is assumed that the first sub-platform and the second sub-platform are detected respectively through two sets of detection servers, and one of the first sub-platform and the second sub-platform is detected three times in a row, and the results are all failed, that is, the detection failed, and the second sub-platform is detected three times in a row and all pass the detection, that is, the detection is successful. Since it is unknown which sub-platform failed the detection, an alarm is triggered, and after manual review, it is decided whether to generate the first switching instruction.

[0066] In a third possible design, if an IP address switch has occurred on the first sub-platform or the second sub-platform of the service management platform, for example, the service of the first sub-platform has been switched to the second sub-platform. If the detection server detects that the service of the first sub-platform has automatically resumed, that is, the detection result is passed, and the detection is continuously performed for 30 minutes with all passed detection results, then a second switching instruction can be generated, so that the GSLB automatically switches the IP address of the first sub-platform from the production IP address of the second sub-platform back to the production IP address of the first sub-platform itself according to the second switching instruction.

[0067] The following is an example to illustrate Figure 3 the working process of the GSLB.

[0068] As Figure 3 shown, for an external network element, such as a Chatbot, to perform service processing, the access process to the service management platform may include the following steps:

[0069] 301. The Chatbot sends a domain name query request to the local DNS.

[0070] 302. The local DNS sends a domain name query request to the upper-level DNS of the GSLB.

[0071] 303. The upper-level DNS of the GSLB returns an NS record to the local DNS. This NS record is obtained by registering the IP address of the GLSB in the upper-level DNS of the GSLB.

[0072] 304. The local DNS sends a domain name query request to the GSLB according to the NS record.

[0073] 305. The GSLB returns the IP address of the sub-platform corresponding to the domain name to the local DNS.

[0074] 306. The local DNS returns the IP address returned by the GSLB to the Chatbot.

[0075] In this embodiment, when an external network element accesses the local DNS to query the IP address of the domain name of the service management platform, if the IP address of the domain name of the service management platform does not exist in the DNS cache, the root DNS is queried. If the root DNS is not the upper-level DNS of the GLSB, the query continues until the upper-level DNS of the GSLB is queried, the NS record is obtained, and based on the NS record, the GSLB is queried. The GSLB determines the target IP according to the matching rule and the address of the local DNS, and returns the target IP to the local DNS. The local DNS returns the target IP to the external network element.

[0076] Specifically, there are two common types of DNS: an authoritative resolution server and a recursive resolution server. The recursive resolution server is also the local DNS.

[0077] The authoritative resolution server stores data for some regions in the domain name space. When a DNS is responsible for governing one or more regions, this DNS is called the authoritative server for these regions. The resource records (Name Server, NS) in the root authoritative DNS or secondary authoritative server mark the DNS designated as the regional authoritative server. Through the servers listed in the NS record, other servers consider it to be the authoritative server for that region. This means that any server specified in the NS record is regarded as an authoritative server by other servers and can respond to queries for names contained within the region.

[0078] The recursive server, i.e., the local server, normally has no domain name resolution data initially. All the domain name resolution data in it comes from the query results from its queries to the authoritative resolution server. Once the query is completed, the recursive server will form a cache record locally according to the Time to Live (TTL) and provide DNS resolution query services for users. This is the function of the recursive server.

[0079] GSLB is first of all an authoritative resolution server. The domain name of the corresponding service management platform is saved on it. Suppose it is botplatform.rcs.chinaunicom.com. Then the corresponding NS record must be configured in the upper-level DNS of GSLB, pointing to the IP where GSLB is located (to ensure security, multiple GSLBs can be deployed, that is, there are multiple IP addresses of GSLBs). When the local DNS queries the NS record, it will come to GSLB to query the corresponding IP of botplatform.rcs.chinaunicom.com. And GSLB is an intelligent DNS, which can determine what content to return to the client according to the source IP address. At this time, GSLB can determine to return the IP address of the first sub-platform or the second sub-platform according to the pre-configured IP matching relationship.

[0080] GSLB can also open an interface through which the IP can be dynamically modified to perform IP address switching when receiving a switching instruction sent by the switching server. Specifically, the switching server sends a switching instruction to GSLB. After receiving the switching instruction, GSLB can modify the IP pointing of the corresponding botplatform.rcs.chinaunicom.com without restarting. And the data will be saved so that the data will not be lost after the device where GSLB is located restarts.

[0081] In some embodiments, a primary and standby configuration can also be set inside GSLB for load balancing. To ensure the reliability of the service. Modifying data is executed on the primary device, and the modified records will be automatically synchronized to the standby device.

[0082] In some embodiments, it is assumed that 2 GSLB servers are installed (IP1 and IP2 are two external network IPs providing services). First, the IP addresses (IP1 and IP2) of the GSLB can be registered with the upper-level DNS of the GSLB. Configure the addresses of IP1 and IP2 on the GSLB, and set the access rules according to the client IP address (usually set according to the principle of proximity), and return IPA (the production IP address of the first sub-platform) or IPB (the production IP address of the second sub-platform). Since the GSLB also needs to perform a disaster recovery function, the DNS cache time cannot be too long, and a 1-minute timeout can be configured. In this way, external network elements will not cause too long a disaster recovery switchover time due to the long-term use of cached addresses.

[0083] In some embodiments, the deployment of the detection server, the switching server, and the GSLB can be carried out with reference to the following table.

[0084]

[0085] The disaster recovery method of the service management platform provided in this embodiment obtains the detection results of each sub-platform by sending test messages to each sub-platform respectively, and generates a switching instruction according to the detection results of each sub-platform and sends it to the GSLB, so as to automatically switch the IP address of the faulty sub-platform to the IP address of the non-faulty sub-platform through the GLSB. It saves the development cost of network management software, realizes automatic switching, saves labor costs, improves efficiency, and avoids service blockage during switching. Since the GSLB is used for IP address switching, each sub-platform does not need to be set in the same region, getting rid of the regional restriction.

[0086] Figure 4 For the Figure 2 specific flowchart of step 201 provided in the embodiments of this application. As Figure 4 shown, step 201 may specifically include:

[0087] 401. Periodically send test messages to the sub-platform.

[0088] 402. Determine whether a response message is received within a first preset time. If so, execute step 403; if not, execute step 406.

[0089] 403. Determine whether the received response message is correct. If so, execute step 404; if not, execute step 406.

[0090] 404. Determine whether the delivery result report is received within a second preset time. If so, execute step 405; if not, execute step 407.

[0091] 405. Determine that both the message sending and receiving are successful, obtain the detection result that passes the detection, and send the detection result to the handover server.

[0092] 406. Determine that the message sending fails, obtain the detection result that fails to pass the detection, and send the detection result to the handover server.

[0093] 407. Determine that the message receiving fails, obtain the detection result that fails to pass the detection, and send the detection result to the handover server.

[0094] In the specific detection process, the probe server regularly starts the inspection tour, begins to simulate Chatbot to send messages to the service management platform, and sets the need to deliver the result report; checks whether the MaaP platform responds normally and whether the returned result is correct. If it is correct, wait for the delivery of the result report; if the delivery result report is received within the preset time, check whether the result of the delivery result report is successfully delivered. The inspection tour ends. Send the detection result to the handover server. The probe server is responsible for simulating the sending of the message and checking the returned message result. Whether the detection is successful or failed, the probe server needs to report the detection result to the handover server in real time so that the handover server can generate the corresponding handover instruction according to the preset policy.

[0095] The disaster recovery method of the service management platform provided in this embodiment can comprehensively detect the uplink and downlink functions of the platform by sending test messages and receiving the delivery result report, realizing the comprehensiveness of fault detection. So as to generate the handover instruction in time and avoid affecting the normal communication of users.

[0096] Figure 5 It is a schematic structural diagram of the disaster recovery device of the service management platform provided in the embodiment of the present application. As Figure 5 shown, the disaster recovery device 50 of the service management platform includes: a detection module 501 and a handover module 502.

[0097] The detection module 501 is used to send a test message to each of the multiple sub-platforms and determine the detection result according to the response result returned by the sub-platform based on the test message;

[0098] The handover module 502 is used to generate a first handover instruction according to the detection results of the multiple sub-platforms and send the first handover instruction to the global load balancing device GSLB of the sub-platform that fails in the multiple sub-platforms, so that the GSLB switches the IP address of the failed sub-platform from its own production IP address to the production IP address of the non-failed sub-platform according to the first handover instruction; the production IP address is the IP address assigned to the sub-platform for accessing its own sub-platform.

[0099] The disaster recovery device of the service management platform provided by the embodiment of the present application obtains the detection results of each sub-platform by sending test messages to each sub-platform respectively, and generates a switching instruction according to the detection results of each sub-platform and sends it to the GSLB, so as to automatically switch the IP address of the faulty sub-platform to the IP address of the non-faulty sub-platform through the GLSB. This saves the development cost of network management software, realizes automatic switching, saves labor costs, improves efficiency, and avoids service blockage during switching. Since the GSLB is used for IP address switching, each sub-platform does not need to be set in the same region, getting rid of the geographical restriction.

[0100] The disaster recovery device of the service management platform provided by the embodiment of the present application can be used to execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0101] Figure 6 It is a block diagram of a disaster recovery device of a service management platform provided by an embodiment of the present application. The device can be a data processing device such as a computer, a message transceiver device, a tablet device, etc.

[0102] The device 60 may include one or more of the following components: a processing component 601, a memory 602, a power supply component 603, a multimedia component 604, an audio component 605, an input / output (I / O) interface 606, a sensor component 607, and a communication component 608.

[0103] The processing component 601 generally controls the overall operation of the device 60, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 601 may include one or more processors 609 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 601 may include one or more modules to facilitate the interaction between the processing component 601 and other components. For example, the processing component 601 may include a multimedia module to facilitate the interaction between the multimedia component 604 and the processing component 601.

[0104] The memory 602 is configured to store various types of data to support the operation of the device 60. Examples of these data include instructions for any application or method operating on the device 60, contact data, phone book data, messages, pictures, videos, etc. The memory 602 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0105] The power supply component 603 provides power for various components of the device 60. The power supply component 603 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 60.

[0106] The multimedia component 604 includes a screen that provides an output interface between the device 60 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 604 includes a front camera and / or a rear camera. When the device 60 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0107] The audio component 605 is configured to output and / or input audio signals. For example, the audio component 605 includes a microphone (MIC) that is configured to receive external audio signals when the device 60 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 602 or transmitted via the communication component 608. In some embodiments, the audio component 605 further includes a speaker for outputting audio signals.

[0108] The I / O interface 606 provides an interface between the processing component 601 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0109] The sensor assembly 607 includes one or more sensors for providing a status assessment of various aspects of the device 60. For example, the sensor assembly 607 can detect the on / off state of the device 60, the relative positioning of components, such as the display and keypad of the device 60. The sensor assembly 607 can also detect a change in the position of the device 60 or a component of the device 60, the presence or absence of user contact with the device 60, the orientation or acceleration / deceleration of the device 60, and the temperature change of the device 60. The sensor assembly 607 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 607 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 607 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0110] The communication component 608 is configured to facilitate communication between the device 60 and other devices in a wired or wireless manner. The device 60 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 608 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 608 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0111] In an exemplary embodiment, the device 60 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0112] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 602 including instructions, and the above instructions can be executed by the processor 609 of the device 60 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0113] The above-mentioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0114] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0115] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical disks.

[0116] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor, implements the disaster recovery method of the service management platform executed by the disaster recovery device of the service management platform as described above.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A disaster recovery switching method for a service management platform, characterized in that The business management platform includes multiple sub-platforms; data of the multiple sub-platforms are synchronized; and the method includes: For each of the plurality of sub-platforms, a test message is sent to the sub-platform, and a detection result is determined according to a response result returned by the sub-platform based on the test message; Generate a first switching instruction according to the detection results of the plurality of sub-platforms, and send the first switching instruction to the global load balancing device GSLB of the sub-platform that has failed among the plurality of sub-platforms, so that the GSLB switches the IP address of the sub-platform that has failed from its own production IP address to the production IP address of the sub-platform that has not failed according to the first switching instruction; the production IP address is the IP address assigned to the sub-platform for accessing its own sub-platform; The determining the detection result according to the response result returned by the sub-platform based on the test message includes: If a correct response message is received within the first preset time, waiting to receive a delivery result report of the test message; If the delivery result report is not received within the second preset time, it is determined that the message reception has failed, and a detection result of failing the detection is obtained.

2. The method according to claim 1, wherein The multiple sub-platforms are respectively assigned multiple production IP addresses; the multiple production IP addresses include a first production IP address and a second production IP address; wherein the first production IP address is used for access when the sub-platform itself is fault-free, and the second production IP address is used to switch the IP addresses of other sub-platforms to the second production IP address when other sub-platforms fail.

3. The method according to claim 1, characterized in that, The test message is any one of the following: sending a notification message, uploading a multimedia file, and downloading a multimedia file.

4. The method according to claim 1, characterized in that, The sending of the test message to the sub-platform includes: A test message is periodically sent to the sub-platform.

5. The method according to claim 1, wherein The sending of the test message to the sub-platform includes: Sending a test message to the sub-platform through multiple accounts; the multiple accounts belong to networks in different regions; The determining the detection result according to the response result returned by the sub-platform based on the test message includes: If the number of accounts whose detection results are failed in the multiple accounts is greater than a preset threshold, the detection result of the sub-platform is determined to be failed.

6. The method according to any one of claims 1-5, characterized in that, The multiple sub-platforms include a first sub-platform and a second sub-platform; and generating a first switching instruction according to the detection result includes: If the detection result of the first sub-platform is that the consecutive number of times of failure is greater than a first preset number, and the detection result of the second sub-platform is that the consecutive number of times of passing is greater than a second preset number, it is determined that the first sub-platform has failed, and the second sub-platform has not failed, and a first switching instruction is generated to enable GSLB to switch the IP address of the first sub-platform from the production IP address of the first sub-platform to the production IP address of the second sub-platform.

7. The method according to claim 6, wherein After the first switching instruction is generated according to the detection results of the plurality of sub-platforms, the method further includes: If the continuous time during which the detection result of the first sub-platform passes the detection is greater than the third preset time, it is determined that the failure of the first sub-platform is resolved, and a second switching instruction is generated, so that the GSLB switches the IP address of the first sub-platform from the production IP address of the second sub-platform back to the production IP address of the first sub-platform according to the second switching instruction.

8. A disaster recovery device for a service management platform, characterized in that, The service management platform includes multiple sub-platforms; the data of the multiple sub-platforms is synchronized; it includes: A detection module, configured to send a test message to each of the multiple sub-platforms and determine a detection result according to the response result returned by the sub-platform based on the test message. A switching module, configured to generate a first switching instruction according to the detection results of the multiple sub-platforms and send the first switching instruction to the global load balancing device GSLB of the sub-platform that has failed among the multiple sub-platforms, so that the GSLB switches the IP address of the failed sub-platform from its own production IP address to the production IP address of the sub-platform that has not failed according to the first switching instruction; the production IP address is the IP address assigned to the sub-platform for accessing its own sub-platform. Specifically, the detection module is configured to wait for the delivery result report of the test message if a correct response message is received within the first preset time; if the delivery result report is not received within the second preset time, it is determined that the message reception fails, and a detection result of not passing the detection is obtained.

9. A disaster recovery device for a service management platform, characterized in that, It includes: At least one processor and a memory; The memory stores computer execution instructions; The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the disaster recovery method of the service management platform according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the processor executes the computer execution instructions, the disaster recovery method of the service management platform according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Master-slave switching method and device for database cluster nodes, equipment and medium

    CN111200532A