A distributed IT automation operation and maintenance system
By designing a distributed IT automation operation and maintenance system, the problem of decentralized management of enterprise IT automation systems is solved, the high availability and scalability of IT systems are achieved, the agility requirements of IT architecture in the digital age are met, and a secure and reliable self-service capability for enterprise-level IT infrastructure operation is provided.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2026-04-07
AI Technical Summary
Currently, enterprises have multiple IT automation systems and tools built in a decentralized manner, lacking unified operation and management of IT infrastructure. This makes it impossible to meet the agility requirements of IT architecture in the digital age, and to provide secure, reliable, and flexible enterprise-level IT infrastructure operation self-service capabilities for ITSM, DevOps, and AI/MLOps systems.
A distributed IT automated operation and maintenance system was designed, including an operation layer, a control layer, a coordination layer, an application layer, and a capability opening layer. Through components such as a proxy module, an operation service module, a service orchestration module, a timed scheduling module, a scenario module, an API gateway, a service registration and configuration module, and an operation and maintenance management module, high availability and scalability are achieved, ensuring high reliability and unified management of IT resource operations.
It realizes a unified automated operation platform for IT systems in large-scale network environments, meets the IT automation needs of various business scenarios, provides secure, reliable, and flexible self-service capabilities for enterprise-level IT infrastructure operations, and supports high availability and performance scalability.
Smart Images

Figure CN116723077B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large-scale network environments, and more particularly to a distributed IT automated operation and maintenance system. Background Technology
[0002] Digital transformation is bringing profound changes to business philosophies in terms of organizational models, internal processes, and upstream and downstream collaborations, in order to cope with increasingly uncertain, complex, and personalized internal and external environments. Agile enterprise management requires a matching agile IT architecture. Dual-mode IT architecture, distributed microservice application architecture, DevOps management principles, and cloud-native technologies are playing an increasingly important role in building IT systems adapted to agile management in the digital age. IT automation is a catalyst for these digital transformation support technologies in terms of quality and efficiency, and an engine driving the value creation of digital technologies.
[0003] The current fragmented deployment of multiple IT automation systems and tools in enterprises lacks unified operation and management of IT infrastructure. It cannot provide secure, reliable, and flexible enterprise-level IT infrastructure operation self-service capabilities for ITSM, DevOps, and AI / MLOps systems, and cannot meet the agility requirements of IT architecture in the digital age. Summary of the Invention
[0004] In view of the above problems, the present invention is proposed to provide a distributed IT automated operation and maintenance system that overcomes or at least partially solves the above problems.
[0005] According to one aspect of the present invention, a distributed IT automated operation and maintenance system is provided, the system comprising: an operation layer, a control layer, a coordination layer, an application layer, and a capability opening layer, specifically including:
[0006] The agent module, as an operation layer component, is used to implement specific automated operation functions;
[0007] The operation service module is a control layer component that implements agent control and group management;
[0008] The service orchestration module is a coordination layer component that enables the orchestration of automated work processes, accepts instructions from various scenario modules, and executes automated work processes.
[0009] The timed scheduling module is a coordination layer component that enables the triggering and execution of all timed tasks in the entire system.
[0010] The scenario module is an application layer component that implements functions for specific application scenarios;
[0011] API gateway is a component in the capability open layer that provides automated service capabilities to the outside world;
[0012] The service registration and configuration module is a global management component that provides service registration and centralized configuration management for all module instances except for the agent.
[0013] The operation and maintenance management module is a global management component that monitors the health status of all module instances.
[0014] The data storage module is used to store data, and includes a cache module and a database module;
[0015] This invention provides a distributed IT automated operation and maintenance system, the system comprising:
[0016] (1) The proxy module establishes a long connection with the operation service module as a socket client. It configures two or more operation service module addresses to achieve high availability in the primary and backup mode. When communication with the current operation service module fails, the proxy module automatically switches to the backup operation service module.
[0017] (2) For IT resource objects operated via remote protocols, two or more proxy modules can be configured to operate on these IT resource objects to ensure high reliability of proxy operations. After receiving the operation instructions for the target IT resource object from the service orchestration module or scenario module, the operation service module selects an available proxy to execute the operation.
[0018] (3) For operations on the host server where the agent is located, multiple servers that require high availability or load balancing can be grouped into a group. The service orchestration module or scenario module sends the target device group to the operation service module. The operation service module allocates tasks among multiple servers in the same device group according to the policy to achieve high availability of operations.
[0019] (4) The above three points ensure high availability of the communication link from the operation service module to the agent to the IT resources.
[0020] (5) Multiple operation service modules implement domain-based management of agent modules, expanding the scale of automated operations. Each operation service module is responsible for maintaining IT resource objects and agent modules within its managed domain, and storing the communication relationships between the three in a cache module. When an agent starts, it registers itself with the operation service module and periodically reports the online status of the IT resources it is responsible for operating. When communication between the agent and the currently connected operation service module fails, it automatically switches to a backup operation service module, which automatically updates the connection relationships between the agent and the operation service module in the cache module.
[0021] (6) The operation and maintenance management module periodically checks the online status of the operation service module through a heartbeat mechanism. When the operation service module is offline, it will delete the operation service module and all its agents and IT resource communication relationships from the cache module.
[0022] (7) Multiple service orchestration modules enable parallel computation of automated processes. The service orchestration module periodically updates task load information to the service registration and configuration module. Before calling the orchestration service module to execute an automated process, the scenario module, API gateway module, and timed scheduling module first request the service orchestration module with the lowest load from the service registration and configuration module. When executing each automated task in the automated process, the service orchestration module finds the operation service module based on the target IT resource and issues execution instructions to it. The service orchestration module writes automated process instance information, execution status, and result information to the database and also caches them in the cache module. Under normal circumstances, the operation service module returns execution information to the service orchestration module that sent the automated task. If the service orchestration module that sent the automated task fails, the operation service module obtains a backup service orchestration module through the service registration and configuration module and returns the automated task execution information. The newly taken-over service orchestration module obtains the automated process instance information from the cache module and drives the process instance execution.
[0023] (8) High availability of the scheduled task module. A global task scheduling mechanism is adopted, in which the scheduled task module starts each scheduled task according to the set scheduling strategy. The specific task execution is completed by the operation service module, service orchestration module, and scenario module. Structurally, a master-slave or master-multiple-slave architecture is adopted to ensure high availability. Combined with the service registration and configuration module, this invention proposes a simplified master election algorithm to realize real-time master election among multiple scheduled task modules.
[0024] (9) The service registration and configuration module provides service registration and centralized configuration management services for all modules of the distributed IT automated operation and maintenance system of the present invention, except for the agent module, and adopts a multi-module cluster structure.
[0025] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of the system structure provided in an embodiment of the present invention;
[0028] Figure 2This is a schematic diagram of distributed operation provided in an embodiment of the present invention;
[0029] Figure 3 This is a schematic diagram of the addressing information carried by data during data transmission between layers, provided in an embodiment of the present invention;
[0030] Figure 4 This is a schematic diagram of the distributed orchestration service provided in an embodiment of the present invention;
[0031] Figure 5 This is a flowchart of the leader selection algorithm provided in an embodiment of the present invention. Detailed Implementation
[0032] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0033] The terms "comprising" and "having," and any variations thereof, in the specification, embodiments, claims, and drawings of this invention are intended to cover non-exclusive inclusion, such as including a series of steps or units.
[0034] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0035] 1. For example Figure 1 The diagram shown is a schematic diagram of the system structure according to an embodiment of the present invention.
[0036] Agent module 1 is deployed on a physical server or virtual machine to implement specific automated operation functions.
[0037] Proxy module 1 receives target IT resource information (address, account, password, etc.), parameters, commands, or scripts, executes shell commands or scripts, or calls APIs to operate the host server's hardware environment, OS, database, middleware, and applications. It can also operate remote servers, network devices, storage devices, cloud environments, databases, middleware, and application systems via access protocols (such as SSH, HTTP, JDBC, etc.). It returns operation process data and result data in a pre-defined, unified format. When communication with the operation service module fails, it caches the operation process data and result data in local files and re-reports them after communication is restored. It periodically checks and reports the online status of the remote devices it is responsible for.
[0038] Operations Service Module 2 is deployed on a physical server or virtual machine to implement agent control and group management. One operations service module manages N agents, and horizontal performance scaling and high availability are achieved by deploying multiple operations service modules. Deploying operations service modules on a server spanning multiple network domains enables communication with IT resources across multiple network domains.
[0039] The operation service module mainly implements the following functions:
[0040] (1) Remote management of the agent module. Remotely deploy the agent module, start, stop, and restart the agent module, centrally set agent module parameters, and upgrade the agent module online.
[0041] (2) Data communication. Provides a unified API interface for proxy operations to the outside world, and acts as a communication bridge between the proxy module, the service orchestration module, and various scenario modules, while shielding the service orchestration module and various scenario modules from the operational details.
[0042] (3) Flow control. Flow control is implemented based on job priority and the busyness of target IT resources.
[0043] (4) IT resource management. This enables unified management of managed devices, including synchronizing managed devices and their operation accounts from the configuration management database (cmdb) and setting up the operation relationship between the agent and the managed devices.
[0044] (5) Centralized script management. Enables centralized management, online editing, security checks and testing of scripts, and supports encapsulating scripts into atomic tasks with input / output parameters and specific functions for user convenience.
[0045] Service Orchestration Module 3 is deployed on physical servers or virtual machines to automate job orchestration and execution. The Service Orchestration Module uses a cluster architecture, allowing for the deployment of N instances to achieve performance scaling and high availability for job execution.
[0046] The service orchestration module mainly implements the following functions:
[0047] (1) Graphical workflow orchestration. Provides a graphical tool for automating job workflow orchestration. Supports serial, parallel, branching, concurrent, aggregation, and loop workflow semantics, as well as nested sub-workflows. Supports manual jobs, time-based jobs, and automated jobs. Supports the transfer of workflow variable values between jobs and between parent and child workflows throughout the entire workflow. Each job can define automatic processing actions linked to its status, such as notifications and rework in case of job exceptions, and set flow control indicators such as priority and concurrency.
[0048] (2) Process Execution. It receives instructions from various scenario modules, executes automated work processes, and drives the execution of manual, time-based, and automated tasks based on process semantics. Manual tasks allow operators to input information through an interface. Time-based tasks implement delays or timed intervals between tasks. Automated tasks call the operation service module to send parameters, commands, or scripts to the target IT resources and receive information on the automated task execution process and results.
[0049] (3) Process monitoring. Monitor the work process and the execution information and results of each automated job, and automatically execute some notification or control actions based on the result status. When the job execution is abnormal, the operator is allowed to perform human intervention such as pausing, stopping, redoing, or skipping.
[0050] The scheduled task 4 is deployed on a physical server or virtual machine, using a master-slave mode to trigger and execute all scheduled tasks in the entire system.
[0051] Module 5 implements functions for specific application scenarios, such as job scheduling, inspection, compliance checks, application deployment, disaster recovery switching, configuration management, environment preparation, and fault self-healing. Module 5 calls upon automated process services provided by the service orchestration module and automated operation services provided by the operation service module to complete various automated tasks.
[0052] Scene modules are deployed on physical servers or virtual machines, and each scene module can be deployed with multiple instances to achieve a cluster.
[0053] API Gateway Module 6 provides automated service capabilities to the outside world through RESTful interfaces, including:
[0054] (1) Automated process information query interface;
[0055] (2) Automated process execution interface;
[0056] (3) Interface for querying the execution results of automated processes;
[0057] (4) Job information query interface;
[0058] (5) Job information execution interface;
[0059] (6) Job information execution result query interface;
[0060] (7) Agent information query interface.
[0061] The API gateway module is deployed on a physical server or virtual machine, and multiple instances can be deployed to achieve a cluster.
[0062] Service registration and configuration module 7 provides service registration and centralized configuration management for all module instances except the agent module. It can be deployed on physical servers or virtual machines, and multiple instances can be deployed to achieve a cluster.
[0063] The operations and maintenance management module 8 monitors the health status of all module instances. It is deployed on physical servers or virtual machines and uses a master-slave configuration to achieve high availability.
[0064] Cache module 9 implements global memory storage of status data related to automated processes and operations, connection relationships between agent and operation service modules, timed scheduling and scheduling status, and running status information of each module instance, providing global data sharing access for multiple instances of the same module.
[0065] The caching module provides a publish-subscribe mechanism, which enables asynchronous message communication between its clients.
[0066] The caching module adopts a cluster architecture with multiple instances, deployed on physical servers or virtual machines.
[0067] Database module 10 uses an SQL database to persistently store various configuration parameters, process and result data related to automated workflows and operations, and information such as the execution status of scheduled tasks. It adopts a cluster architecture and is deployed on physical servers or virtual machines.
[0068] 2. For example Figure 2 The diagram shown is a schematic representation of a distributed operation provided in an embodiment of the present invention.
[0069] 2.1 Agent Module Registration
[0070] (1) Figure 2 The connection marked 1 in the middle represents the proxy module. Figure 2 When (211....2nn) starts, it acts as a socket client and interacts with the operation service module ( Figure 2 Establish a TCP connection (31....3n), each proxy module has a unique ID identifier, and configure two operation service module addresses (one primary and one backup).
[0071] (2) The operation service module provides a graphical management interface for configuring the target device for operation. Figure 2 The proxy module of device 111.....1nn, as well as the operation protocol, account, and password information, or the target device and account information can be automatically synchronized from a third-party configuration management database (cmdb).
[0072] like Figure 2Agent module 211 can operate from device 111 to device 11n, and agent module 2n1 can operate from device 1n1 to device 1nn. To achieve high availability, two or more agent modules can be configured for the same device. As shown in the figure, device 11n can be operated by agent module 211 and agent module 221 at the same time. The system sends the operation commands to one of the online agent modules according to the configuration order.
[0073] (3) The operation service module saves its connection relationship with the agent module and the device relationship that the agent module is responsible for operating into the cache module, forming a global tree structure.
[0074] (4) The information in the cache module is accessed by all service orchestration modules and operation service modules.
[0075] (5) The agent module periodically monitors the online status of the target device and actively reports it to the operation service module, which then deletes the offline device from the cache module.
[0076] (6) The operation service module periodically sends heartbeat packets to the agent module to check whether it is online. The agent module also uses the heartbeat packets to determine whether the connection with the operation service module is normal.
[0077] (7) After the operation service module determines that the agent module is offline, it deletes the agent module and all device information under it from the cache module.
[0078] (8) After the proxy module determines that the communication with the current operation service module is abnormal, it will actively switch to another operation service module. The latter will update the relationship between the operation service module and the proxy in the cache module, delete the proxy module and its device information from the original operation service module and then add it to the new operation service module.
[0079] 2.2 Instruction Execution
[0080] All automation scenarios can be transformed into automated operation processes. For ease of explanation, the following will use... Figure 2 The "Example Automated Operation Flow" in the document is used as an example to illustrate the instruction execution flow of this invention.
[0081] (1) Operation Service Module Figure 2 (31 to 3n) and orchestration service module ( Figure 2 When starting up, (41 to 4n) register service information (service name, service instance IP, service instance port, service instance status) with the service registration configuration module.
[0082] (2) The operation and maintenance management module obtains information on all operation service modules and orchestration service modules from the service registration and configuration module, periodically calls the heartbeat interface of these modules to check their online status, and updates it to the service registration and configuration module. The service registration and configuration module stores the online status information of all operation service modules and orchestration service modules.
[0083] (3) Example automated operation process has 3 automated task nodes connected in series: task J1 operates device 111 (through agent module 211), task J2 operates device 11n (both agent modules 211 and 221 can operate device 11n, and the two are mutually redundant to achieve high availability), and task J3 operates device 12n (through agent module 221).
[0084] (4) Scene Module Figure 2 Execute automated operations by calling the service registration configuration module to obtain the service orchestration module with the lightest load, assuming it is service orchestration module 41, and send a request to service orchestration module 41 to execute the example automated operation process.
[0085] (5) Service Orchestration Module 41 Instantiates Example Automated Operation Process (Creates an execution instance of an automated process in the database, loads the entire process model and execution parameter information, etc.). The requesting scenario module records its service ID and service instance ID, as well as automated operation process instance information (such as process instance ID, instance status, etc.) in the cache module and the database.
[0086] (6) Service orchestration module 41 executes J1. It queries the cache module to find the operation path of target device 111, which is operation service module 31 -> proxy module 211 -> device 111. Service orchestration module 41 sends the process instance ID, J1 task instance ID, J1 execution command and parameter information, target device 111 address information, target proxy module 221 address information, and the service instance ID registered by service orchestration module 41 in service registration configuration module 6 to operation service 31. Service orchestration module 41 saves the task instance ID and execution status of task J1 to cache module 6 and the database.
[0087] (7) The operation service module 31 performs flow control according to the system configuration strategy. After the flow control conditions are met, it sends the J1 execution command to the agent module 221.
[0088] (8) The agent module 221 performs automated operations on the device 111 and returns the execution results.
[0089] (9) Under normal circumstances, the proxy module 221 will return the execution result to the operation service module 31, which will then return it to the service orchestration module 41.
[0090] (10) Service orchestration module 41 saves the execution status and result information of task J1 to cache module 6 and database system.
[0091] (11) Service orchestration module 41 executes J2 tasks. From the cache module ( Figure 2 There are two operation paths for querying the target device 11n: Operation Service Module 31 -> Agent Module 211 -> Device 11n, and Operation Service Module 32 -> Agent Module 221 -> Device 11n. Service Orchestration Module 41 can choose any one of them to execute. Other operation procedures are the same as those in steps (6) to (10) above.
[0092] (12) Service orchestration module 41 executes J3 task, the steps are the same as (6) to (10).
[0093] (13) Figure 2 The service registration configuration module and the operation management module are both deployed in a cluster to avoid single points of failure.
[0094] 2.3 Operation service module failure during instruction execution
[0095] In steps (6) and (9) of 2.2 above, if the operation service module malfunctions, the implementation example of the present invention will handle it as follows:
[0096] (1) In step (6) of 2.2, the service orchestration module 41 sends a task execution command to the operation service module 31 and finds that the operation service module 31 is faulty. Since it takes a period of time (e.g., 3 heartbeat intervals) for the proxy module connected to the operation service module 31 to determine the fault of the operation service module 31 and switch to the backup operation service module, the operation path obtained by the service orchestration module 41 from the cache module 8 is still the operation service module 31 -> proxy module 211 -> device 111. The service orchestration module 41 will detect the communication abnormality when sending the command to the operation service module 31, and then return the abnormality information to the scenario module, and then to the user. The user can manually intervene on the graphical management interface, such as waiting for a period of time to redo the J1 task. At this time, the proxy module originally connected to the operation service module 31 will reconnect to the backup operation service module and rebuild the new operation path in the cache module 8.
[0097] (2) In step (8) of 2.2, the agent module 221 returns the result of the automated task execution to the operation service module 31. At this time, the operation service module 31 is found to be faulty. The agent module 221 first caches the task execution result information in a local file, then reconnects to the backup operation service module, such as the operation service module 32, and then returns the task execution result to the operation service module 32.
[0098] (3) The operation service module 32 returns the result to the service orchestration module 41 based on the service instance ID number of the service orchestration module in the result response message returned by the agent module.
[0099] 2.4 Service orchestration module failure during instruction execution
[0100] In step (9) of 2.2 above, the handling logic for service orchestration module failures when the operation service module returns task result information to the service orchestration module is described in detail in point 5 below.
[0101] 3. For example Figure 3 The figure shows the addressing information carried by the data during data transmission between layers in an embodiment of the present invention.
[0102] Each scenario module has a scenario service ID and a scenario service instance ID, which respectively identify the type of service provided and the specific service instance.
[0103] The scenario module executes an operation and maintenance scenario function, which calls the service orchestration module. The latter instantiates the corresponding automated process instance, identifies each process instance with a process instance ID and the service orchestration instance ID of the currently executing process instance. The scenario module establishes a correspondence between this information and the scenario service instance ID and the scenario service ID and saves it in the cache module.
[0104] When the service orchestration module executes a task in an automated process, it generates a task instance ID and packages the task instance ID, process instance ID, service orchestration instance ID, scenario service instance ID, and scenario service ID together and sends them to the operation service module. The service orchestration module saves this information in the cache module so that other service orchestration modules can access it.
[0105] The operation service module will forward the task instance ID, process instance ID, service orchestration instance ID, scenario service instance ID, scenario service ID, task-related parameters, and target device information to the agent module.
[0106] The agent returns task execution information, including the task instance ID, process instance ID, service orchestration instance ID, scenario service instance ID, and scenario service ID information sent when the task request was made.
[0107] The operation service module returns task execution information to the corresponding service orchestration module based on the service orchestration instance ID in the message returned by the agent. If the original service orchestration instance fails at this time, the operation service module requests a backup instance of the service orchestration instance from the service registration and configuration module and returns the task execution information.
[0108] When a process instance completes execution, the service orchestration module returns the scenario module instance ID based on the returned message. If the original scenario module instance fails at this time, it requests a backup instance of the scenario module instance from the service registration and configuration module and returns the execution information.
[0109] 4. For example Figure 4 The diagram shown is a schematic of the distributed orchestration service provided in an embodiment of the present invention.
[0110] (1) Scenario module 51 executes a scenario function and obtains service orchestration instance 401 from the service registration configuration module, assuming it is service orchestration module 41. Scenario module 51 sends an automation process execution request 402 to service orchestration module 41.
[0111] (2) Service orchestration module 41 creates an automated process instance and saves the instance information to cache module 403 and database module 404.
[0112] (3) The service orchestration module 41 instantiates the task information in the automation process, saves the task information to the cache module 403 and the database module 404, and sends it to the operation agent module 31 on the target device path.
[0113] (4) The operation service module 31 receives the task execution result information from the agent module and normally returns it to the requesting service orchestration module (identified by 406, i.e.) Figure 4 The service orchestration module 41 in the document (see point 3 for the specific addressing process)
[0114] (5) If the operation service module 31 finds that the original requested service orchestration module 41 is abnormal, it obtains the backup service orchestration module of service orchestration module 41 from the service registration configuration module (see point 5 below for the specific acquisition algorithm), which is assumed to be service orchestration module 4n.
[0115] (6) The operation service module 31 returns task execution result information 408 to the service orchestration module 4n. After receiving the information, the service orchestration module 4n updates the result information of the task instance and the status information of the automated process instance in the cache module 409 and the database module 410, and drives the execution of the next process node according to the process semantics.
[0116] (7) If the current automation task is the last node of the process, it means that the automation process instance has been completed and the service orchestration instance 4n returns the automation process instance result information to the scenario module 51.
[0117] 5. Backup relationships between multiple instances of the service module
[0118] In points (6) and (7) of the above item 4, both involve the problem of selecting a backup instance when the original service instance is abnormal in the case of N service instances. When a service has N service instances and one of the service instances fails, the tasks originally assigned to this service instance need to be taken over by another service instance. To make the takeover orderly, the following backup relationship is agreed between service instances: It is agreed that the numerical value of the service instance's IP address plus the port number is used as the unique identifier of the service instance. All available service instances of the same service are arranged in a ring in ascending order according to this unique identifier. The latter is defaulted to be the backup of the former. For example, if id1 < id2 < id3, id2 is the backup of id1, id3 is the backup of id2, and id1 is the backup of id3.
[0119] 6. Refer to Figure 5 , which is the flowchart of the master selection algorithm provided by the embodiment of the present invention
[0120] The timing scheduling module 4 can only execute the global scheduling of timing tasks by the master module instance at the same moment, and other instances are in the hot standby state. It is necessary to implement master selection among multiple instances.
[0121] Similarly, the operation and maintenance management module 8 can only monitor the health status of each module by the master module instance at the same moment, and other instances are in the hot standby state. It is necessary to implement master selection among multiple instances.
[0122] Common master selection algorithms include Paradox, Raft, etc. These algorithms not only implement master selection among multiple instances, but also implement state synchronization among multiple instances to ensure consistency. The state synchronization among multiple instances of the timing scheduling module and the operation and maintenance management module in the embodiment of the present invention can be realized through the cache module and the database system. Therefore, a Figure 5 simplified master selection algorithm is adopted.
[0123] The master selection process is triggered and executed when there are changes (service instances go online or offline) in the service instances of the service module (specifically, the timing scheduling or operation and maintenance management module in the embodiment of the present invention) in the system initialization and service registration and configuration module.
[0124] Step 1, read all instance data of the current service from the service registration and configuration module.
[0125] If at least one of the online service instances of the current service has a role of leader, then:
[0126] Step 2, register itself with the service registration and configuration module, and set role = "follower" and registerTimeStamp = "current time".
[0127] The process ends.
[0128] If the current online service instance is empty, or if it is not empty but the service instance with the role "leader" is empty, then:
[0129] Step 3: Register yourself with the service registration configuration module, setting role="leader" and registerTimeStamp="current time".
[0130] Step 4: Read all instance data of the current service from the service registration configuration module.
[0131] If only one of the online service instances in the current service has the role of "leader", then the process ends.
[0132] If multiple service instances in the current online service instance have the role of "leader", then:
[0133] If your registerTimeStamp is the smallest among all service instances with the role of "leader", and no other service instance has a registerTimeStamp equal to it, then the process ends.
[0134] If its own registerTimeStamp is not the smallest among all service instances with the role "leader", it updates itself with the service registration configuration module, setting role="follower" and registerTimeStamp="current time", and the process ends.
[0135] If your registerTimeStamp is the smallest among all service instances with the role of "leader", but another service instance has a registerTimeStamp equal to it, then:
[0136] Wait for a random period of time (e.g., 100 milliseconds).
[0137] Update itself with the service registration configuration module by setting registerTimeStamp="current time".
[0138] Jump to step 4.
[0139] Beneficial effects: It features high reliability and horizontal scalability, enabling a unified automated operation platform for IT systems in large-scale network environments, and meeting the IT automation requirements of various business scenarios.
[0140] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A distributed IT automated operation and maintenance system, characterized in that, The operation and maintenance system includes: an operation layer, a control layer, a coordination layer, an application layer, and a capability opening layer, specifically including: The agent module, as an operation layer component, is used to implement specific automated operation functions; The operation service module is a control layer component that implements agent control and group management; The service orchestration module is a coordination layer component that enables the orchestration of automated work processes, accepts instructions from various scenario modules, and executes automated work processes. The timed scheduling module is a coordination layer component that enables the triggering and execution of all timed tasks in the entire system. The scenario module is an application layer component that implements functions for specific application scenarios; API gateway is a component in the capability open layer that provides automated service capabilities to the outside world; The service registration configuration module is a global management component that provides services to all module instances except for the agent. The operation and maintenance management module is a global management component that monitors the health status of all module instances. The caching module is used to store state data related to automated processes and operations in memory. The database module is used to store various configuration parameters, automated processes, and operation-related process and result data; The system consists of an automated task operation subsystem composed of an operation service module and an agent module, and is distributed. The distributed IT automated operation and maintenance system consists of several automated task operation subsystems, enabling horizontal scaling of the operation scale; Each automated task operation subsystem contains at least two operation service modules and several agent modules; Each IT resource is operated by one or more agents simultaneously; The operation path is composed of the operation service module, the agent module, and IT resources. The operation service module is responsible for maintaining it and caching it in the cache module to build the global operation path of the entire system. The service orchestration module sends the task execution command to the target operation service module based on the global operation path; High availability of proxy operations can be achieved by setting up multiple operation paths for an IT resource; The proxy establishes a TCP connection with the operation service module as a socket client, and the two perform heartbeat checks on each other. When the operation service module fails, the proxy module will actively switch to the backup operation service module, which will automatically update the global operation tree in the cache module. During the failover process, the agent module will first cache the task execution results to the local file system and then re-report them after the switchover is complete.
2. The distributed IT automated operation and maintenance system according to claim 1, characterized in that, The system's service orchestration modules supporting various scenario functions are distributed, with multiple automated process instances executed in parallel across multiple service orchestration module instances. The execution information of the automated process instances is globally shared through a caching module. The scenario module can send an automated process execution request to any service orchestration module; During the execution of an automated process instance, if the service orchestration module of the executing process instance crashes, the backup service orchestration module will automatically take over the execution of the automated process instance. The failover mechanism between multiple service orchestration modules is orderly and can be completed automatically without prior configuration.
3. A distributed IT automated operation and maintenance system according to claim 1, characterized in that, The system’s timed scheduling module and operation and maintenance management module adopt a master-slave mode, and only the master module instance can be executed at any given time. It is necessary to implement the selection of master among multiple instances. State synchronization between multiple instances is achieved through the caching module and the database module; with the help of the service registration configuration module, the distributed IT automated operation and maintenance system adopts a simplified leader election algorithm. When a new instance joins or leaves, the roles of all current instances are obtained from the service registration configuration module. Then, based on the leader and follower roles of each instance, the appropriate role for the current instance is determined, and a proposal is submitted to the service registration configuration module. If there is a conflict, the proposal is submitted again after a random waiting period of time until the role can be determined.
Citation Information
Patent Citations
Monitoring system and method for employing multi-process applications in container cluster
CN106776212A
Operation and maintenance system, operation and maintenance method, electronic equipment and storage medium
CN111526049A