Operation and Maintenance Method, System and Related Devices

By designing an operation and maintenance system that includes operation and maintenance production and disaster recovery systems, and using the operation and maintenance subsystem with main and multi-active relationships, the problems of high disaster recovery time requirements and insufficient scalability of the disaster recovery architecture in the existing technology are solved, and efficient disaster recovery switching and business continuity are achieved.

CN115883341BActive Publication Date: 2025-05-27CHINA CONSTRUCTION BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211506694.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-05-27
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

When designing business continuity solutions in the prior art, disaster recovery time requirements are high, resulting in increased design difficulty, and the dual data center's disaster recovery architecture scalability and switching capabilities are insufficient.

Method used

Design an operation and maintenance system, including the application layer and the procurement and control layer. The application layer involves the operation and maintenance production system and the operation and maintenance disaster recovery system. Each system includes N operation and maintenance subsystems, which are mutually reserved or multi-active relationships. The procurement and control layer includes multiple procurement and control management services, connecting production and disaster recovery systems. By obtaining transaction requests, the system determines the associated operation and maintenance subsystem, and switches to the disaster recovery system or diverts to other operation and maintenance subsystems when a failure occurs, and manages the procurement and control layer services.

Benefits of technology

It realizes efficient disaster recovery architecture scalability and fast switching response, can add operation and maintenance subsystems as needed, and quickly switch and manage in case of failures, meeting the requirements of high business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115883341B_ABST
    Figure CN115883341B_ABST
Patent Text Reader

Abstract

The present invention discloses an operation and maintenance method, system and related device, which can obtain a transaction request sent by a business system; determine a first operation and maintenance subsystem associated with the transaction request according to the transaction request; for any pair of first operation and maintenance subsystems that are in a primary-backup relationship, in the case of a failure of the first operation and maintenance subsystem in the operation and maintenance production system, switch to the operation and maintenance disaster recovery system to process the target transaction through the operation and maintenance disaster recovery system and manage the acquisition and control management service through the operation and maintenance disaster recovery system; for any pair of first operation and maintenance subsystems that are in a multi-active relationship, in the case of a failure of one of the first operation and maintenance subsystems, divert the data stream processed by the failed first operation and maintenance subsystem to another first operation and maintenance subsystem, and manage the acquisition and control management service through the other first operation and maintenance subsystem. The disaster recovery architecture of the present invention has good scalability, and the switching response between operation and maintenance subsystems is fast and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fintech, and particularly to an operation and maintenance method, system and related device. Background Art

[0002] According to the requirements of the industry for business continuity, the disaster recovery time for basic and important services is 2 - 4 hours. The high requirement for recovery time increases the difficulty of designing the business continuity design plan.

[0003] Currently, many enterprises have planned dual data centers. Generally, the scenarios of multi - active and disaster recovery in different locations for dual data centers are commonly studied. Existing solutions mainly include methods such as multi - active in different locations, disaster recovery in different locations, data backup in different locations, disaster recovery in the same city, and multi - active in the same city. This single - layer application architecture mode does not suit the application characteristics of the operation and maintenance system for local acquisition and control of each data center, and both the scalability and switching ability of the disaster recovery architecture are insufficient. Summary of the Invention

[0004] In view of the above problems, the present invention provides an operation and maintenance method, system and related device that overcome or at least partially solve the above problems.

[0005] In a first aspect, an operation and maintenance method is applied to an operation and maintenance system. The operation and maintenance system includes an application layer and an acquisition and control layer. Among them, the application layer involves an operation and maintenance production system and an operation and maintenance disaster recovery system. Both the operation and maintenance production system and the operation and maintenance disaster recovery system include N operation and maintenance subsystems. The N operation and maintenance subsystems of the operation and maintenance production system and the N operation and maintenance subsystems of the operation and maintenance disaster recovery system correspond one by one. Any pair of operation and maintenance subsystems are in a primary - standby relationship or a multi - active relationship with each other. The acquisition and control layer includes multiple acquisition and control management services, and each acquisition and control management service is connected to at least one operation and maintenance subsystem in the operation and maintenance production system and the corresponding at least one operation and maintenance subsystem in the operation and maintenance disaster recovery system;

[0006] The operation and maintenance method includes:

[0007] Obtain a transaction request sent by a business system;

[0008] Determine a first operation and maintenance subsystem associated with the transaction request according to the transaction request;

[0009] For any pair of first operation and maintenance subsystems in a primary - standby relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the operation and maintenance disaster recovery system to process the target transaction through the operation and maintenance disaster recovery system, and manage the acquisition and control management services of the acquisition and control layer through the operation and maintenance disaster recovery system;

[0010] For any pair of first operation and maintenance subsystems with a multi-active relationship, in the event that one of the first operation and maintenance subsystems fails, the data stream processed by the failed first operation and maintenance subsystem is diverted to the other first operation and maintenance subsystem, and the operation and maintenance management service of the acquisition and control layer is managed by the other first operation and maintenance subsystem.

[0011] Combined with the first aspect, in some optional embodiments, the operation and maintenance method further includes:

[0012] In the event that any of the operation and maintenance management services fails, switch to other operation and maintenance management services to manage the external managed machines managed by the failed operation and maintenance management service through the other operation and maintenance management services.

[0013] Combined with the first aspect, in some optional embodiments, for any pair of first operation and maintenance subsystems with a primary-backup relationship, in the event that the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the operation and maintenance disaster recovery system to process the target transaction through the operation and maintenance disaster recovery system and manage the operation and maintenance management service of the acquisition and control layer through the operation and maintenance disaster recovery system, including:

[0014] For any pair of first operation and maintenance subsystems with a primary-backup relationship, in the event that the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system to perform corresponding processing on the target transaction through the second operation and maintenance subsystem and perform corresponding management on the operation and maintenance management service of the acquisition and control layer through the second operation and maintenance subsystem, where the first operation and maintenance subsystem and the second operation and maintenance subsystem have a primary-backup relationship.

[0015] Combined with the previous embodiment, in some optional embodiments, for any pair of first operation and maintenance subsystems with a primary-backup relationship, in the event that the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system to perform corresponding processing on the target transaction through the second operation and maintenance subsystem and perform corresponding management on the operation and maintenance management service of the acquisition and control layer through the second operation and maintenance subsystem, including:

[0016] For any pair of first operation and maintenance subsystems with a primary-backup relationship, in the event that the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system by means of IP addressing to perform corresponding processing on the target transaction through the second operation and maintenance subsystem and perform corresponding management on the operation and maintenance management service of the acquisition and control layer through the second operation and maintenance subsystem.

[0017] In combination with the first aspect, in some alternative embodiments, determining the first operation and maintenance subsystem associated with the transaction request according to the transaction request includes:

[0018] Determine at least one of the first operation and maintenance subsystems involved in executing the transaction request according to the IP identifier, user identifier, or transaction code identifier carried in the transaction request.

[0019] In a second aspect, an operation and maintenance system includes an application layer and a collection and control layer. Among them, the application layer involves an operation and maintenance production system and an operation and maintenance disaster recovery system. The operation and maintenance production system and the operation and maintenance disaster recovery system both include N operation and maintenance subsystems. The N operation and maintenance subsystems of the operation and maintenance production system and the N operation and maintenance subsystems of the operation and maintenance disaster recovery system correspond one by one. Any pair of operation and maintenance subsystems are in a primary / backup relationship or a multi-active relationship with each other. The collection and control layer includes multiple collection and control management services. Each collection and control management service is connected to at least one operation and maintenance subsystem in the operation and maintenance production system and the corresponding at least one operation and maintenance subsystem in the operation and maintenance disaster recovery system;

[0020] The application layer includes: a transaction request acquisition unit, an associated subsystem determination unit, a primary / backup switching unit, and a multi-active switching unit;

[0021] The transaction request acquisition unit is configured to acquire a transaction request sent by a business system;

[0022] The associated subsystem determination unit is configured to determine a first operation and maintenance subsystem associated with the transaction request according to the transaction request;

[0023] The primary / backup switching unit is configured to, for any pair of first operation and maintenance subsystems in a primary / backup relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the operation and maintenance disaster recovery system to process the target transaction through the operation and maintenance disaster recovery system and manage the collection and control management services of the collection and control layer through the operation and maintenance disaster recovery system;

[0024] The multi-active switching unit is configured to, for any pair of first operation and maintenance subsystems in a multi-active relationship, in the case where one of the first operation and maintenance subsystems fails, divert the data stream processed by the failed first operation and maintenance subsystem to another first operation and maintenance subsystem and manage the collection and control management services of the collection and control layer through the other first operation and maintenance subsystem.

[0025] In combination with the second aspect, in some alternative embodiments, the collection and control layer further includes: a first switching unit;

[0026] The first switching unit is configured to switch to other acquisition and control management services in the event of a failure of any one of the acquisition and control management services, so as to manage the external managed machines managed by the acquisition and control management service with a failure through the other acquisition and control management services.

[0027] In combination with the second aspect, in some optional embodiments, the primary and standby switching unit includes: a second switching unit;

[0028] The second switching unit is configured to, for any pair of first operation and maintenance subsystems that are in a primary and standby relationship, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system in the event of a failure of the first operation and maintenance subsystem in the operation and maintenance production system, so as to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management services of the acquisition and control layer through the second operation and maintenance subsystem, where the first operation and maintenance subsystem and the second operation and maintenance subsystem are in a primary and standby relationship.

[0029] In a third aspect, a computer-readable storage medium stores a program thereon, and when the program is executed by a processor, the operation and maintenance method described in any one of the above is implemented.

[0030] In a fourth aspect, an electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory communicate with each other through the bus; the processor is configured to call program instructions in the memory to execute the operation and maintenance method described in any one of the above.

[0031] With the above technical solution, an operation and maintenance method, system and related device provided by the present invention, the operation and maintenance method is applied to an operation and maintenance system, the operation and maintenance system includes an application layer and a collection and control layer. Among them, the application layer involves an operation and maintenance production system and an operation and maintenance disaster recovery system. The operation and maintenance production system and the operation and maintenance disaster recovery system both include N operation and maintenance subsystems. The N operation and maintenance subsystems of the operation and maintenance production system and the N operation and maintenance subsystems of the operation and maintenance disaster recovery system correspond one by one. Any pair of operation and maintenance subsystems are in a primary-backup relationship or a multi-active relationship with each other. The collection and control layer includes multiple collection and control management services. Each collection and control management service is connected to at least one operation and maintenance subsystem in the operation and maintenance production system and at least one corresponding operation and maintenance subsystem in the operation and maintenance disaster recovery system. The operation and maintenance method includes: obtaining a transaction request sent by a business system; determining a first operation and maintenance subsystem associated with the transaction request according to the transaction request; for any pair of first operation and maintenance subsystems in a primary-backup relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switching to the operation and maintenance disaster recovery system to process the target transaction through the operation and maintenance disaster recovery system and manage the collection and control management services of the collection and control layer through the operation and maintenance disaster recovery system; for any pair of first operation and maintenance subsystems in a multi-active relationship, in the case where one of the first operation and maintenance subsystems fails, diverting the data stream processed by the failed first operation and maintenance subsystem to the other first operation and maintenance subsystem and managing the collection and control management services of the collection and control layer through the other first operation and maintenance subsystem. It can be seen from this that the disaster recovery architecture of the operation and maintenance system of the present invention has good scalability, can add corresponding operation and maintenance subsystems as needed, and the switching response between operation and maintenance subsystems is fast.

[0032] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific implementation manners of the present invention. Brief Description of the Drawings

[0033] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0034] Figure 1 A schematic diagram of a deployment architecture provided by the present invention is shown;

[0035] Figure 2 A schematic diagram of a classification method provided by the present invention is shown;

[0036] Figure 3 shows a flowchart of an operation and maintenance method provided by the present invention;

[0037] Figure 4 shows a schematic structural diagram of an operation and maintenance system provided by the present invention;

[0038] Figure 5 shows a schematic structural diagram of a dual-active AQ mode provided by the present invention;

[0039] Figure 6 shows a schematic structural diagram of a dual-active AA mode provided by the present invention;

[0040] Figure 7 shows a schematic structural diagram of another dual-active AA mode provided by the present invention;

[0041] Figure 8 shows a schematic structural diagram of a dual-active AS mode provided by the present invention;

[0042] Figure 9 shows a schematic diagram of a deployment mode provided by the present invention;

[0043] Figure 10 shows a schematic structural diagram of another operation and maintenance system provided by the present invention;

[0044] Figure 11 shows a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0045] According to the requirements of the industry for business continuity, the disaster recovery time of the basic and important services of the business system is 2-4 hours. The high requirement for recovery time increases the difficulty of designing the business continuity design solution. Application multi-live, application disaster recovery, data synchronization, and switching strategies have become the hotspots of technical research.

[0046] Many enterprises have carried out the planning of dual data centers. Generally, the scenarios of remote multi-live and disaster recovery dual data centers are commonly studied, and most of them do not consider the disaster recovery solutions for multiple locations and multiple data centers in the digital age of explosive data volume.

[0047] Especially in the banking industry, most of the existing solutions address the data synchronization and consistency issues between dual centers, and there is relatively little research on the dedicated architecture deployment and switching strategies of the operation and maintenance system in terms of recovery timeliness. In particular, the operation and maintenance system is essential for the daily operation and maintenance of the data center of the business system, and the automated operation and maintenance capabilities and monitoring capabilities at all levels of the operation and maintenance system are necessary dependencies for the switching of the business system. In terms of the disaster recovery mode of applications, the existing solutions mainly include methods such as multi-site active-active, off-site disaster recovery, off-site data backup, on-site disaster recovery, and on-site multi-site active-active. This single-layer application architecture mode does not suit the application characteristics of the operation and maintenance tools for localized acquisition and control of each data center, and both the scalability and switching capabilities of the disaster recovery architecture are insufficient.

[0048] The present invention is applicable to the field of bank operation and maintenance and supports three-level disaster recovery switching for multiple data centers. The operation and maintenance system usually includes core applications such as automated operation and maintenance, monitoring of different-level objects (including hardware, platforms, and applications), configuration management, alarm handling, and IT service processes. The operation and maintenance system described in this solution supports the operation and maintenance of multiple regions and multiple data centers. In terms of the deployment architecture, the existing operation and maintenance systems generally adopt single-center deployment. User access, inter-system calls, and the access of managed machines are all connected through the same application node. Among them, different from the commonly referred to applications, there are two types of interactions between the operation and maintenance system and the managed machines. One is that an AGENT is installed on the managed machine, and the other is to directly connect and interact through network communication.

[0049] Optionally, as Figure 1 shown, different from the deployment architectures of the above existing operation and maintenance systems: the operation and maintenance system of the present invention is a multi-site and multi-center deployment architecture without high availability. For example, assume there are n applications, namely S1…Sn, where some applications have no managed machines, and some operation and maintenance subsystems need to connect to managed machines (M1, Mp, Mq, and Mr, etc. can all be understood as managed machines). C1 and C2 are two regions, and three data centers are planned (respectively: AZ1, AZ2, and AZ3. The data center can also be called an availability zone). Among them, AZ1 and AZ2 are deployed in region C1, and AZ3 is deployed in region C2. Of course, the present invention does not make specific restrictions on the number of regions and data centers. For example, regions can be increased to C3 and C4, etc., and each region can also have multiple data centers.

[0050] Under the above deployment architecture of the operation and maintenance system, if a disaster recovery operation and maintenance system is to be built, consider the following key indicators: ① Whether the disaster scenarios (regional-level, availability zone-level, and system-level disaster scenarios) can meet the requirements for supporting business recovery. ② Whether the business of the disaster scenario itself can be restored in a timely manner. ③ In terms of system capacity, whether resources can be effectively utilized. Based on the above three requirements, the present invention first classifies the operation and maintenance system, and then uses a hierarchical architecture to achieve three-level disaster recovery capabilities and establish a disaster switching model.

[0051] Based on the above three requirements, since the operation and maintenance system has different classification positions during the business switching process, the operation and maintenance system of the present invention is divided into switching operation type, monitoring type, continuous operation and maintenance type, and management analysis type systems, and the classification method is as follows Figure 2 shown.

[0052] For Figure 2 the operation and maintenance services in, most systems have relatively low requirements for RTO (Recovery Time Objective of information system: refers to the time requirement from interruption to the time when the business or information system must be restored after an emergency) and RPO (Recovery Point Objective of information system: refers to the time point to which the data must be restored after an emergency causes the business or information system to be interrupted, that is, the maximum time length of tolerable data loss). Generally speaking, it is sufficient if the recovery time is within a few hours. However, in the overall context of banking business continuity, due to the increasing scale of the cloud infrastructure in the data center and the scale of the servers under the single-system distributed architecture, it is unrealistic to perform switching and monitoring manually. Therefore, the classification method of the present invention can increase the perspective of the combination of the operation and maintenance system and the banking business, identify more reasonable switching scenarios and switching indicators, and then design a reasonable disaster recovery operation and maintenance system architecture accordingly.

[0053] The disaster recovery mentioned in the present invention includes three levels: system level, availability zone level, and regional level. In the architecture of two regions and three data centers shown above Figure 1 the core operation and maintenance subsystems in the primary and standby mode are mainly distributed in one of the data centers, while the multi-active systems are distributed in multiple data centers. The primary center (operation and maintenance production system, also known as the production side) of the operation and maintenance system of the present invention is in C1, and C1 has two availability zones AZ1 and AZ2, and the disaster recovery center (operation and maintenance disaster recovery system, also known as the disaster recovery side) is built in C2. The operation and maintenance subsystems mainly have three deployment forms: remote multi-active, local multi-active, and remote disaster recovery.

[0054] When the above three-level fault scenarios exist in the business system, the requirements for the operation and maintenance system are analyzed as follows:

[0055] 1. System fault scenario: There is an internal fault in the business system, and the operation and maintenance system runs normally without being affected, and no special disaster recovery design is required.

[0056] 2. Availability zone fault scenario: If the operation and maintenance system is also in the faulty availability zone (data center), the operation and maintenance system itself needs to perform disaster recovery switching and then support the recovery of the business system.

[0057] 3. Regional fault scenario: If the operation and maintenance system is also in the faulty region, the operation and maintenance system itself needs to perform disaster recovery switching and then support the recovery of the business system.

[0058] According to the above Figure 2Based on the analysis of the classification switching requirements of the operation and maintenance system and the above analysis of the business system failure scenarios, the operation and maintenance system of the present invention proposes the following deployment principles and architecture design solutions.

[0059] Deployment principle: To reduce the probability of simultaneous failures of the operation and maintenance system and the business system, in principle, the production side in the primary / backup mode of the operation and maintenance system or the central unit in the multi-active mode can be deployed remotely from the business system. In the example of the present invention, the operation and maintenance production system of the operation and maintenance system should be in Region C2.

[0060] Architecture design solution and deployment principle: Deploy according to a hierarchical multi-active and primary / backup architecture to achieve the goals of three-level disaster recovery and ensure the timeliness of switching.

[0061] Hereinafter, exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0062] As Figure 3 shown, the present invention provides an operation and maintenance method applied to an operation and maintenance system. The operation and maintenance system includes an application layer and a collection and control layer. Among them, the application layer involves an operation and maintenance production system and an operation and maintenance disaster recovery system. Both the operation and maintenance production system and the operation and maintenance disaster recovery system include N operation and maintenance subsystems. The N operation and maintenance subsystems of the operation and maintenance production system and the N operation and maintenance subsystems of the operation and maintenance disaster recovery system correspond one by one and are in a primary / backup relationship with each other. The collection and control layer includes multiple collection and control management services, and each collection and control management service is connected to at least one operation and maintenance subsystem in the operation and maintenance production system and the corresponding at least one operation and maintenance subsystem in the operation and maintenance disaster recovery system;

[0063] The operation and maintenance method includes: S100, S200, S300, and S400;

[0064] Optionally, the operation and maintenance system described in the present invention is as Figure 4As shown in the figure, taking the operation and maintenance subsystem M and the operation and maintenance subsystem N as examples, the operation and maintenance subsystem M involved in the operation and maintenance production system and the operation and maintenance subsystem M involved in the operation and maintenance disaster recovery system are in a primary / standby relationship with each other. The operation and maintenance subsystem N involved in the operation and maintenance production system and the operation and maintenance subsystem N involved in the operation and maintenance disaster recovery system are also in a primary / standby relationship with each other. There is another situation: the operation and maintenance subsystem N involved in the operation and maintenance production system is deployed in the data center in Region C2, and the operation and maintenance subsystem N involved in the operation and maintenance disaster recovery system is deployed in the data center in Region C1, thus forming a multi-active relationship (that is, the operation and maintenance subsystem N in Region C2 and the operation and maintenance subsystem N in Region C1 process the same logic and can respectively undertake the processing work of the corresponding data streams at the same time). It should be noted that for any pair of operation and maintenance subsystems, they are either in a primary / standby relationship or in a multi-active relationship. There is no situation where a pair of operation and maintenance subsystems is both in a primary / standby relationship and in a multi-active relationship. However, in the entire operation and maintenance system, there can be operation and maintenance subsystems in a primary / standby relationship and operation and maintenance systems in a multi-active relationship at the same time. Under the above deployment architecture, the operation and maintenance subsystem M involved in the operation and maintenance production system can synchronize data to the operation and maintenance subsystem M involved in the operation and maintenance disaster recovery system. Similarly, the operation and maintenance subsystem N involved in the operation and maintenance production system can also synchronize data to the operation and maintenance subsystem N involved in the operation and maintenance disaster recovery system, as Figure 4 shown by the dashed line in

[0065] Optionally, the present invention does not specifically limit the running relationship between the operation and maintenance subsystems involved in the application layer, and the serial running relationship and parallel running management can be set according to actual operation and maintenance needs. Taking Figure 4 as an example, the operation and maintenance subsystem M and the operation and maintenance subsystem N are in a serial running relationship.

[0066] Optionally, as mentioned above, some operation and maintenance subsystems in the application layer need to interact with the managed machine, and some operation and maintenance subsystems have no need to interact with the managed machine. For example, Figure 4 the operation and maintenance subsystem N in

[0067] Optionally, the application layer can provide the following services to external applications and external systems: ① Provide various access interfaces to the outside world and correctly route access transactions. That is, according to the diversion logic, through the IP network segment or user ID and other identifiers, a certain algorithm is used to determine which data center node should be used to execute the transaction. ② Complete cross-data center summary calculations to form an overall application view. That is, obtain local calculation data and integrate them into a unified summary view, such as monitoring the localized transaction volume of each data center. If you want to form an overall view of multiple data centers, you must summarize them at the application layer. ③ Complete the switching of the system itself in a disaster scenario, and complete the routing switching of external services and acquisition and control layer access. The switching of the system itself generally includes DNS switching, switching of transaction routing strategies, switching of database replication relationships, switching of disaster recovery status to production status, and master-slave switching of mirrored NAS. According to different system implementations, different switching actions can be selected to complete the switching of the system itself.

[0068] Optionally, the acquisition and control layer provides services to the managed machines to complete the following localized functions: ① Management of the status of the managed machines and the client status. The status of the managed machine usually includes: whether it is active, system status indicators, and performance indicators. Client status: whether it is active, whether the thread is normal, whether the execution permission is normal, the amount of tasks being executed and the task status, etc. ② Issuance and execution of jobs. ③ Collection and upload of data. ④ Local computing, for some statistical analysis work of the local client, calculate local monitoring indicators, such as transaction volume, log volume, and alarm information. Also calculate the local task issuance path, task failure retransmission and other scheduling logic. ⑤ Takeover of managed machines in disaster scenarios and service switching at the application layer.

[0069] S100, obtaining a transaction request sent by a business system;

[0070] Optionally, the business system mentioned in the present invention can be understood as an external system other than the operation and maintenance system. As described above, the operation and maintenance system of the present invention can correctly route the access transaction. Therefore, the present invention can obtain the transaction request sent by the external system to facilitate the subsequent routing of the transaction request and the execution of corresponding decisions. For example, Figure 4 The operation and maintenance subsystem M in the operation and maintenance production system shown can obtain the transaction request sent by the business system, and the present invention does not impose any limitation on this.

[0071] Optionally, the external system may resolve the IP address of the operation and maintenance subsystem M in the operation and maintenance production system based on a domain name resolver, and then send a corresponding transaction request based on the IP address, which is not limited in the present invention.

[0072] S200: Determine, according to the transaction request, a first operation and maintenance subsystem associated with the transaction request;

[0073] Optionally, the present invention can correctly route transaction requests. Therefore, the present invention can determine which operation and maintenance subsystems are required to participate in the execution of the corresponding process for processing the transaction request according to the transaction request. That is, the first operation and maintenance subsystem mentioned in the present invention is one or more operation and maintenance subsystems related to the transaction request, which are collectively referred to as the first operation and maintenance subsystem here, and the present invention does not limit this. It should be noted that, as mentioned above, the operation and maintenance subsystems of the present invention can be deployed in different data centers. Therefore, determining the operation and maintenance subsystems related to the transaction request determines the corresponding data centers, and the present invention does not limit this.

[0074] Optionally, as mentioned above, for operation and maintenance subsystems in a primary-backup relationship, the operation and maintenance subsystem on the operation and maintenance production system side is the primary system. Therefore, first, the first operation and maintenance subsystem associated with the transaction request is obtained from the operation and maintenance production system. Since the operation and maintenance subsystems involved in the operation and maintenance production system correspond one-to-one with the operation and maintenance subsystems involved in the operation and maintenance disaster recovery system and are in a primary-backup relationship. Therefore, the standby first operation and maintenance subsystem associated with the transaction request in the operation and maintenance disaster recovery system is indirectly determined. When a certain first operation and maintenance subsystem involved in the operation and maintenance production system fails, it can be switched to the standby first operation and maintenance subsystem involved in the operation and maintenance disaster recovery system, so as to smoothly execute the corresponding processing process for the transaction request.

[0075] Optionally, for operation and maintenance subsystems in a multi-active relationship, whether it is the operation and maintenance subsystem on the operation and maintenance production system side or the operation and maintenance subsystem on the operation and maintenance disaster recovery system side, the corresponding logic can be executed at the same time and the corresponding data flow processing work can be undertaken. Therefore, the present invention determines the corresponding first operation and maintenance subsystems in a multi-active relationship from both the operation and maintenance production system and the operation and maintenance disaster recovery system.

[0076] Optionally, the present invention does not limit the specific process of determining the first operation and maintenance subsystem associated with the transaction request. For example, in combination with Figure 3 the embodiment shown, in some optional embodiments, the S200 includes:

[0077] Determine at least one of the first operation and maintenance subsystems involved in executing the transaction request according to the IP identifier, user identifier or transaction code identifier carried in the transaction request.

[0078] Optionally, the IP identifier mentioned in the present invention can be understood as an IP network segment and an IP address, etc., and the user identifier can be understood as a user account, a user name and a user ID, etc., and the present invention does not limit this.

[0079] Optionally, at least one of the first operation and maintenance subsystems involved in executing the transaction request mentioned in the present invention is the operation and maintenance subsystem required for processing the transaction request.

[0080] S300. For any pair of first operation and maintenance subsystems that are in a primary / standby relationship with each other, when a failure occurs in the first operation and maintenance subsystem in the operation and maintenance production system, switch to the operation and maintenance disaster recovery system to process the target transaction through the operation and maintenance disaster recovery system, and manage the acquisition and control management service of the acquisition and control layer through the operation and maintenance disaster recovery system;

[0081] Optionally, as described above, processing a transaction request may require multiple first operation and maintenance subsystems in the operation and maintenance production system. When a certain first operation and maintenance subsystem fails, it can be switched to the operation and maintenance disaster recovery system, and the corresponding operation and maintenance subsystem in the operation and maintenance disaster recovery system executes the corresponding processing process.

[0082] Optionally, Figure 4 For example, if the operation and maintenance subsystem N in the operation and maintenance production system fails, switch to the operation and maintenance subsystem N in the operation and maintenance disaster recovery system, and the operation and maintenance subsystem N in the operation and maintenance disaster recovery system executes the corresponding processing process. Also, since the operation and maintenance subsystem N needs to interact with the managed machine, the operation and maintenance subsystem N in the operation and maintenance disaster recovery system can send corresponding instructions to the corresponding acquisition and control management service of the acquisition and control layer.

[0083] For example, in combination with Figure 3 the embodiment shown, in some optional embodiments, the S300 includes:

[0084] For any pair of first operation and maintenance subsystems that are in a primary / standby relationship with each other, when a failure occurs in the first operation and maintenance subsystem in the operation and maintenance production system, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management service of the acquisition and control layer through the second operation and maintenance subsystem, where the first operation and maintenance subsystem and the second operation and maintenance subsystem are in a primary / standby relationship with each other.

[0085] Optionally, the initiator of the action to switch to the operation and maintenance disaster recovery system can be initiated by the disaster recovery control system with the switching plan configured. That is, after the application layer receives the switching instruction initiated by the disaster recovery control system, it can execute the corresponding switching process.

[0086] Optionally, Figure 4 For example, the operation and maintenance subsystem N in the operation and maintenance production system can be understood as the first operation and maintenance subsystem, and the operation and maintenance subsystem N in the operation and maintenance disaster recovery system can be understood as the second operation and maintenance subsystem. The present invention does not limit this.

[0087] Combined with the previous embodiment, in some alternative embodiments, for any pair of first operation and maintenance subsystems that are in a primary-backup relationship, when a failure occurs in the first operation and maintenance subsystem in the operation and maintenance production system, it is switched to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system, so as to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management service of the acquisition and control layer through the second operation and maintenance subsystem, including:

[0088] For any pair of first operation and maintenance subsystems that are in a primary-backup relationship, when a failure occurs in the first operation and maintenance subsystem in the operation and maintenance production system, it is switched to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system by means of IP addressing, so as to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management service of the acquisition and control layer through the second operation and maintenance subsystem.

[0089] Optionally, the present invention does not make specific restrictions on the switching method. In addition to using the IP addressing method, other feasible methods can also be used. It should be noted that: if a failure occurs in a certain operation and maintenance subsystem of the operation and maintenance production system, it can be switched to the corresponding operation and maintenance subsystem in the operation and maintenance disaster recovery system to execute the corresponding process. After the corresponding operation and maintenance subsystem in the operation and maintenance disaster recovery system finishes executing, it can be switched to other subsequent operation and maintenance subsystems in the operation and maintenance production system to execute the corresponding process.

[0090] S400. For any pair of first operation and maintenance subsystems that are in a multi-active relationship, when a failure occurs in one of the first operation and maintenance subsystems, the data stream processed by the failed first operation and maintenance subsystem is split to the other first operation and maintenance subsystem, and the acquisition and control management service of the acquisition and control layer is managed by the other first operation and maintenance subsystem.

[0091] Optionally, the above describes the situation where a failure occurs in the application layer, including the situation where an individual operation and maintenance subsystem of the operation and maintenance production system fails and the situation where the entire operation and maintenance production system fails (if the entire operation and maintenance production system fails, it is directly switched to the entire operation and maintenance disaster recovery system). The situation where a failure occurs in the acquisition and control layer will be described below.

[0092] That is, combined with Figure 3 the shown embodiment, in some alternative embodiments, the operation and maintenance method further includes:

[0093] When any acquisition and control management service fails, it is switched to another acquisition and control management service, so as to manage the external managed machines managed by the failed acquisition and control management service through the other acquisition and control management service.

[0094] Optionally, when any of the above acquisition and control management services fails, the acquisition and control layer can receive a corresponding switching instruction and control other acquisition and control management services to take over the corresponding managed machines according to the switching instruction. For example, Figure 4 Taking Figure 4 as an example, assume that acquisition and control management service 1 is deployed in an availability zone in Region C2, acquisition and control management service 2 is deployed in availability zone 1 in Region C1, and acquisition and control management service 3 is deployed in availability zone 2 in Region C1. If acquisition and control management service 1 fails, acquisition and control management service 2 can take over the managed machines managed by acquisition and control management service 1.

[0095] To further clearly describe the solution of the present invention, the daily process and switching process of the operation and maintenance system are analyzed below:

[0096] 1. Normal state: The access mode of the operation and maintenance system is through DNS addressing. Among them, for the application layer: the operation and maintenance production systems of the primary and standby systems are deployed in C2 (C2 refers to Region C2). The operation and maintenance disaster recovery system is deployed in C1-AZ1 (C1 refers to Region C1, and AZ1 refers to availability zone 1); both data centers of the active-active system provide services externally. In the active-active AQ mode (the mode definition will be analyzed in detail later), one side only has query capabilities, and in the active-active AA mode, it has the ability to distribute core transaction traffic. For the acquisition and control layer: for the operation and maintenance subsystems that interact with the same managed machine, acquisition and control management services (which can be understood as servers) are deployed locally to manage the communication between the operation and maintenance subsystems and the managed machines and complete local calculations.

[0097] 2. Data synchronization: The system performs data synchronization according to the RPO requirements, and the data synchronization can select the platform general solution or the personalized solution.

[0098] 3. Disaster switching state. For the application layer: rely on the DNS server for routing switching, and cooperate with the opening of its own disconnector, state switching, and traffic distribution strategy adjustment to complete the switching. For the acquisition and control layer: when the application layer switches, the nodes of the acquisition and control layer are taken over by the disaster recovery or other multi-active nodes; when a disaster occurs to the acquisition and control layer services themselves, the acquisition and control management services in other data centers take over the managed machines. Database switching: Complete database switching and adjust the data replication relationship, and perform consistency checks on important data (such as configuration management data, personnel permission data, etc.) according to its own business logic, and perform data chasing.

[0099] It should be noted that: for the multi-active relationship, there are three deployment modes in the application layer, namely the active-active AQ mode, the active-active AA mode, and the active-active AS mode. The acquisition and control layer is in a multi-active mode, and each logical availability zone is an equivalent deployment unit. The connection between the acquisition and control layer and the application layer is divided into two modes.

[0100] Optionally, taking Figure 5Take the architecture of the active-active AQ mode shown as an example. LB refers to the load balancer device, WEB refers to the front end, AP refers to the server, and DB refers to the database. 1. Normal state: Users or systems provide services externally through the general domain name, and the general domain name points to the central unit (primary central unit); there is also a sub-domain name that points to the ordinary unit (disaster recovery center), and query transactions can also be executed by accessing through the sub-domain name.

[0101] 2. Data synchronization: Data is maintained in the central unit and synchronized to the ordinary unit at the required frequency. The database status of the ordinary unit is read-only or read-write, and the AP server controls access to the database.

[0102] 3. Disaster switchover: When a disaster occurs in the central unit, no switching action is required. As long as access is made through the sub-domain name, the core capabilities (such as login, query, and verification) can be obtained; other service capabilities are enabled through switches, and the general domain name switchover is completed synchronously to resume all external services.

[0103] It should be noted that: If the system load is not high and the main transactions are query transactions, the architecture shown can be adopted, such as the operation and maintenance portal and user authentication system, etc. Figure 5 as shown.

[0104] Optionally, take Figure 6 the architecture of another active-active AA mode shown as an example. 1. Normal state: Ordinary users or systems provide services externally through the general domain name, and DNS distributes requests to the central unit and the ordinary unit according to the established strategy. Deployment strategies are implemented in LB or AP to split transactions, and the transactions are distributed to the correct deployment unit for processing. 2. Data synchronization: Global data is maintained in the central unit and synchronized to the ordinary unit at the required frequency. Localized data (generally login logs, status information, operation records, etc.) is not synchronized. Cross-regional holistic analysis is completed on the big data platform. 3. Disaster switchover: When a disaster occurs in the central unit or the ordinary unit, through the automatic detection of DNS, the traffic is redistributed, and the transaction splitting strategy is adjusted. It is necessary to stop data synchronization by configuration:

[0105] It should be noted that: (1) The deployment of the splitting strategy can be advanced to correct traffic errors in a timely manner and avoid complex call logics. (2) In the scenario of multiple data centers, the ordinary unit can be expanded to achieve multi-activity; if the system load is extremely large and considering system and availability zone disasters, and cross-regional transaction drainage is not possible, a cross-availability zone multi-activity deployment structure can be added. (3) If the system data can be completely sliced, there is no central unit, and all are peer units.

[0106] Optionally, take Figure 7 the architecture of another active-active AA mode shown as an example. 1. Normal state: Services are provided externally through the sub-domain name daily, and external systems route transaction traffic to the fixed center through configured fixed sub-domain names.

[0107] 2. Data synchronization: Localized data (usually login logs, status information, operation records, etc.) is not synchronized, and global data that needs to be taken over after switching supports two-way synchronization.

[0108] 3. Disaster switching: When a disaster occurs in a data center, the domain name is switched through DNS, and the domain name of the faulty data center is pointed to the IP address of the non-faulty data center to achieve re-distribution of traffic and stop data synchronization.

[0109] It should be noted that: (1) For systems that can slice data according to regions, the Figure 7 architecture can be adopted, and the call logic and data synchronization are relatively simple.

[0110] Optionally, taking the architecture of the active-active AS mode shown in Figure 8 as an example. 1. Daily access: Services are provided externally through the domain name daily, and DNS points to the IP of the production side; the disaster recovery side is not started and is in a standby state. 2. Data synchronization: Data can be synchronized through the platform product solution or the system can export and synchronize data by itself. 3. Disaster switching: When a disaster occurs on the production side, start the AP and WEB on the disaster recovery side, switch DNS, and enable the disaster recovery side system.

[0111] It should be noted that: (1) Each system designs the RPO according to its own business recovery requirements. (2) Each system needs to design disaster tolerance scenarios or propose data dependencies on surrounding systems to avoid unknown system states caused by different data recovery time points of each system.

[0112] Next, taking the typical switching scenario (regional fault switching of switching operation systems) as an example, the solution of the present invention will be further described. The switching operation systems referred to in the present invention include the disaster recovery management and control system (DIMS), the automated operation and maintenance platform (AOP), and the security operation and maintenance management system (SOM). The switching operation system refers to the operation and maintenance system serving the switching of the business system, and the switching timeliness requirement of this type of system is the highest. We will take this scenario as an example to analyze the switching timing sequence under this architecture.

[0113] The disaster recovery management and control system initiates the execution of the one-key switching script, the automated operation and maintenance platform provides the automated ability for switching on the managed system, and the security operation and maintenance management system provides the network device login password and user authentication function.

[0114] Among them, there is a deployment of the acquisition and control layer in the automated operation and maintenance platform, and only the application layer is deployed in the other two systems. There are two deployment modes for the switching operation system: the active-active AQ mode (DIMS and SOM) and the active-active AA mode (AOP). The deployment diagrams are shown in Appendix Figure 9 .

[0115] 1. Normal state: Users log in to the system using the total domain name, and access between systems also uses the total domain name. DIMS and SOM are normally externally served through the A-end node of C2. Some functions of automated operation and maintenance (such as switching functions with higher priorities like network device switching) are externally served through two sub-domain name resolutions to dual-center nodes, and some functions are externally served through the total domain name resolution to the A-end node of C2. The sub-domain names of all systems are called by the disaster recovery control system to complete the switching of this system and quickly support the switching of business systems. The following analyzes the system switching process for failure scenarios:

[0116] 1. Failure in Area C1: (1) Scenario analysis. If there is an overall failure in C1, the application layer and the acquisition and control layer will have an overall failure. The business system switches from C1 to C2, and there is no business system in C1 that needs operation and maintenance. (2) Switching steps. The three switching operation systems provide normal external services, and the switching time is 0. Just stop the data synchronization and replication relationship.

[0117] 2. Failure in Area C2: (1) Scenario analysis. In case of a failure in C2, the business system in C2 needs to be switched to the standby center in C1. The switching operation system needs to be restored first to support business switching. (2) In the case of a failure of the primary center of the non-operation and maintenance system, the operation and maintenance system can provide the disaster switching ability for the business system within 5 minutes.

[0118] 3. Other failure scenarios: (1) Availability zone-level failure. For the application layer, when a failure occurs in the available zone where it is located, the switching process is the same as that of a regional failure; for systems with an acquisition and control layer, the acquisition and control layer and the application layer achieve cross-available zone takeover through domain name switching or based on the control mechanism between the two layers. (2) System-level failure: When a failure occurs in the application layer, the acquisition and control layer and the application layer achieve cross-available zone takeover through domain name switching or based on the control mechanism between the two layers; when a failure occurs in the acquisition and control layer, other acquisition and control layer units need to take over the managed machines. This scenario is time-consuming, but for system-level failures, since there is no need to support business switching and it is only for its own switching, RTO < 4 hours is sufficient.

[0119] As Figure 10 shown, the present invention provides an operation and maintenance system, the operation and maintenance system includes an application layer and an acquisition and control layer, wherein, the application layer involves an operation and maintenance production system and an operation and maintenance disaster recovery system, both the operation and maintenance production system and the operation and maintenance disaster recovery system include N operation and maintenance subsystems, the N operation and maintenance subsystems of the operation and maintenance production system and the N operation and maintenance subsystems of the operation and maintenance disaster recovery system correspond one by one, and any pair of operation and maintenance subsystems are in a primary-standby relationship or a multi-active relationship with each other, the acquisition and control layer includes multiple acquisition and control management services, and each acquisition and control management service is connected to at least one operation and maintenance subsystem in the operation and maintenance production system and the corresponding at least one operation and maintenance subsystem in the operation and maintenance disaster recovery system;

[0120] The application layer includes: a transaction request acquisition unit 100, an associated subsystem determination unit 200, a primary / standby switching unit 300, and a multi-active switching unit 400;

[0121] The transaction request acquisition unit 100 is configured to acquire a transaction request sent by a service system;

[0122] The associated subsystem determination unit 200 is configured to determine a first operation and maintenance subsystem associated with the transaction request according to the transaction request;

[0123] The primary / standby switching unit 300 is configured to, for any pair of first operation and maintenance subsystems that are in a primary / standby relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the operation and maintenance disaster recovery system, so as to process the target transaction through the operation and maintenance disaster recovery system, and manage the acquisition and control management service of the acquisition and control layer through the operation and maintenance disaster recovery system;

[0124] The multi-active switching unit 400 is configured to, for any pair of first operation and maintenance subsystems that are in a multi-active relationship, in the case where one of the first operation and maintenance subsystems fails, split the data stream processed by the failed first operation and maintenance subsystem to another first operation and maintenance subsystem, and manage the acquisition and control management service of the acquisition and control layer through the other first operation and maintenance subsystem.

[0125] Combined with Figure 10 the embodiment shown, in some alternative embodiments, the acquisition and control layer further includes: a first switching unit;

[0126] The first switching unit is configured to, in the case where any acquisition and control management service fails, switch to another acquisition and control management service, so as to manage the external managed machines managed by the failed acquisition and control management service through the other acquisition and control management service.

[0127] Combined with Figure 10 the embodiment shown, in some alternative embodiments, the primary / standby switching unit 300 includes: a second switching unit;

[0128] The second switching unit is configured to, for any pair of first operation and maintenance subsystems that are in a primary / standby relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system, so as to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management service of the acquisition and control layer through the second operation and maintenance subsystem, where the first operation and maintenance subsystem and the second operation and maintenance subsystem are in a primary / standby relationship.

[0129] Combined with the previous embodiment, in some alternative embodiments, the second switching unit includes: a third switching unit;

[0130] The third switching unit is configured to, for any pair of first operation and maintenance subsystems that are in a primary-backup relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system through IP addressing, so as to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management service of the acquisition and control layer through the second operation and maintenance subsystem.

[0131] Combined Figure 10 In the embodiment shown, in some alternative embodiments, the associated subsystem determination unit 200 includes: an associated subsystem determination subunit;

[0132] The associated subsystem determination subunit is configured to determine at least one of the first operation and maintenance subsystems involved in executing the transaction request according to the IP identifier, user identifier, or transaction code identifier carried in the transaction request.

[0133] The present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the operation and maintenance method described in any one of the above is implemented.

[0134] As Figure 11 shown, the present invention provides an electronic device 70, and the electronic device 70 includes at least one processor 701, at least one memory 702 connected to the processor 701, and a bus 703; wherein, the processor 701 and the memory 702 complete communication with each other through the bus 703; the processor 701 is configured to call program instructions in the memory 702 to execute the operation and maintenance method described in any one of the above.

[0135] In the present invention, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0136] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for relevant details.

[0137] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined in the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in the present invention, but will conform to the widest scope consistent with the principles and novel features disclosed in the present invention.

[0138] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.

Claims

1. An operation and maintenance method, characterized in that, it is applied to an operation and maintenance system, the operation and maintenance system includes an application layer and a collection and control layer. Among them, the application layer involves an operation and maintenance production system and an operation and maintenance disaster recovery system. The operation and maintenance production system and the operation and maintenance disaster recovery system both include N operation and maintenance subsystems. The operation and maintenance system deploys corresponding data centers in different regions. The N operation and maintenance subsystems of the operation and maintenance production system and the N operation and maintenance subsystems of the operation and maintenance disaster recovery system correspond one by one. Any pair of operation and maintenance subsystems are in a primary-backup relationship or a multi-active relationship with each other. The collection and control layer includes multiple collection and control management services. Each collection and control management service is connected to at least one operation and maintenance subsystem in the operation and maintenance production system and at least one corresponding operation and maintenance subsystem in the operation and maintenance disaster recovery system; The operation and maintenance method includes: Obtaining a transaction request sent by a business system; Determining at least one first operation and maintenance subsystem involved in executing the transaction request according to the IP identifier, user identifier or transaction code identifier carried in the transaction request; For any pair of first operation and maintenance subsystems in a primary-backup relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switching to the operation and maintenance disaster recovery system to process the target transaction through the operation and maintenance disaster recovery system and manage the collection and control management services of the collection and control layer through the operation and maintenance disaster recovery system; For any pair of first operation and maintenance subsystems in a multi-active relationship, in the case where one of the first operation and maintenance subsystems fails, shunting the data stream processed by the failed first operation and maintenance subsystem to the other first operation and maintenance subsystem and managing the collection and control management services of the collection and control layer through the other first operation and maintenance subsystem; In the case where any one of the collection and control management services fails, switching to other collection and control management services to manage the external managed machines managed by the failed collection and control management service through the other collection and control management services; The step of, for any pair of first operation and maintenance subsystems in a primary-backup relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switching to the operation and maintenance disaster recovery system to process the target transaction through the operation and maintenance disaster recovery system and manage the collection and control management services of the collection and control layer through the operation and maintenance disaster recovery system, includes: For any pair of first operation and maintenance subsystems in a primary-backup relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switching to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system to perform corresponding processing on the target transaction through the second operation and maintenance subsystem and perform corresponding management on the collection and control management services of the collection and control layer through the second operation and maintenance subsystem, where the first operation and maintenance subsystem and the second operation and maintenance subsystem are in a primary-backup relationship.

2. The operation and maintenance method according to claim 1, characterized in that, For any pair of first operation and maintenance subsystems that are in a primary / standby relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system, so as to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management service of the acquisition and control layer through the second operation and maintenance subsystem, including: For any pair of first operation and maintenance subsystems that are in a primary / standby relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system by means of IP addressing, so as to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management service of the acquisition and control layer through the second operation and maintenance subsystem.

3. An operation and maintenance system, the operation and maintenance system includes an application layer and an acquisition and control layer, wherein, the application layer involves an operation and maintenance production system and an operation and maintenance disaster recovery system. Both the operation and maintenance production system and the operation and maintenance disaster recovery system include N operation and maintenance subsystems. The operation and maintenance system deploys corresponding data centers in different regions. The N operation and maintenance subsystems of the operation and maintenance production system and the N operation and maintenance subsystems of the operation and maintenance disaster recovery system correspond one by one. Any pair of operation and maintenance subsystems are in a primary / standby relationship or a multi-active relationship. The acquisition and control layer includes multiple acquisition and control management services. Each acquisition and control management service is connected to at least one operation and maintenance subsystem in the operation and maintenance production system and the corresponding at least one operation and maintenance subsystem in the operation and maintenance disaster recovery system; The application layer includes: a transaction request obtaining unit, an associated subsystem determining unit, a primary / standby switching unit, and a multi-active switching unit; The transaction request obtaining unit is configured to obtain a transaction request sent by a service system; The associated subsystem determining unit is configured to determine at least one first operation and maintenance subsystem involved in executing the transaction request according to the IP identifier, user identifier, or transaction code identifier carried in the transaction request; The primary / standby switching unit is configured to, for any pair of first operation and maintenance subsystems that are in a primary / standby relationship, in the case where the first operation and maintenance subsystem in the operation and maintenance production system fails, switch to the operation and maintenance disaster recovery system, so as to process the target transaction through the operation and maintenance disaster recovery system, and manage the acquisition and control management service of the acquisition and control layer through the operation and maintenance disaster recovery system; The multi-active switching unit is configured to, for any pair of first operation and maintenance subsystems that are in a multi-active relationship, in the case where one of the first operation and maintenance subsystems fails, divert the data stream processed by the failed first operation and maintenance subsystem to another first operation and maintenance subsystem, and manage the acquisition and control management service of the acquisition and control layer through the other first operation and maintenance subsystem; The acquisition and control layer further includes: a first switching unit; The first switching unit is configured to, in the case where any acquisition and control management service fails, switch to other acquisition and control management services, so as to manage the external managed machines managed by the failed acquisition and control management service through the other acquisition and control management services; The primary and standby switching unit includes: a second switching unit; The second switching unit is configured to, for any pair of first operation and maintenance subsystems that are in a primary and standby relationship, switch to the corresponding second operation and maintenance subsystem in the operation and maintenance disaster recovery system when the first operation and maintenance subsystem in the operation and maintenance production system fails, so as to perform corresponding processing on the target transaction through the second operation and maintenance subsystem, and perform corresponding management on the acquisition and control management service of the acquisition and control layer through the second operation and maintenance subsystem, where the first operation and maintenance subsystem and the second operation and maintenance subsystem are in a primary and standby relationship.

4. A computer-readable storage medium, on which a program is stored, characterized in that, when the program is executed by a processor, it implements the operation and maintenance method described in any one of claims 1 to 2.

5. An electronic device, characterized in that, the electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory communicate with each other through the bus; the processor is configured to call program instructions in the memory to execute the operation and maintenance method described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Disaster tolerance switching method and system, storage medium and computer equipment

    CN112463440A