Operation and maintenance system and method, electronic equipment, storage medium and program product

By building an operation and maintenance system, using API gateways and unified device models to standardize and abstract heterogeneous network devices, and combining knowledge graphs for fault analysis and automated workflow management, the complexity and inefficiency of traditional network management methods are solved, achieving efficient network operation and maintenance.

CN121125442APending Publication Date: 2025-12-12INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511032350.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

In existing technologies, traditional network management methods are complex, inefficient, and prone to errors when operating heterogeneous network devices, making it difficult to effectively maintain and operate networks.

Method used

An operation and maintenance system is built using an API gateway, a heterogeneous device abstraction module, an intelligent operation and maintenance module, and an orchestration module. The system standardizes and abstracts heterogeneous network devices through a unified device model, uses knowledge graphs for root cause analysis of faults, and automatically generates and executes workflows.

Benefits of technology

It enables unified and standardized management of heterogeneous networks, improves fault handling efficiency, builds an automated operation and maintenance closed loop, and enhances the agility and reliability of network services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125442A_ABST
    Figure CN121125442A_ABST
Patent Text Reader

Abstract

The invention provides an operation and maintenance system and method, electronic equipment, a storage medium and a program product, and relates to the technical field of information processing, and the system comprises an API gateway, a heterogeneous equipment abstraction module, an intelligent operation and maintenance module and an arrangement module; the API gateway is used for providing and managing network capability APIs, and the network capability APIs correspond to resource management, service configuration and equipment monitoring capability of the underlying network management system; a heterogeneous device abstraction module configured to call a network capability API through an API gateway; the heterogeneous device abstraction module is used for providing a real-time telemetry data stream standardized by the unified device model; the intelligent operation and maintenance module is used for receiving the standardized real-time telemetering data flow, analyzing the standardized real-time telemetering data flow based on knowledge graph correlation analysis, and generating a fault root cause analysis report; and the arrangement module is used for analyzing the fault root cause analysis report into a workflow and calling the heterogeneous equipment abstraction module to execute the workflow.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information processing, and in particular to an operation and maintenance system and method, an electronic device, a storage medium and a program product. BACKGROUND

[0002] In the process of enterprise digital transformation, the scale and complexity of industry customers' private networks are growing. These private networks often deploy network devices from different manufacturers, forming a complex heterogeneous network environment. Traditional network management methods usually rely on network management systems for specific manufacturer devices or manual configuration and monitoring through command line interfaces. This approach has the problems of complex operation, low efficiency and error-prone.

[0003] Therefore, how to effectively perform network operation and maintenance has become a problem to be solved in the industry. SUMMARY

[0004] The present application provides an operation and maintenance system and method, an electronic device, a storage medium and a program product to solve the problem of how to effectively perform network operation and maintenance in the prior art.

[0005] The present application provides an operation and maintenance system, comprising: an API gateway, a heterogeneous device abstraction module, an intelligent operation and maintenance module, and an orchestration module. The API gateway is configured to provide and manage network capability APIs corresponding to resource management, service configuration and device monitoring capabilities of underlying network management systems. The heterogeneous device abstraction module is configured to call the network capability APIs through the API gateway, and the heterogeneous device abstraction module is configured to provide standardized real-time telemetry data streams via a unified device model. The intelligent operation and maintenance module is configured to receive the standardized real-time telemetry data streams, analyze the standardized real-time telemetry data streams based on knowledge graph correlation analysis, and generate a fault root cause analysis report. The orchestration module is configured to parse the fault root cause analysis report into a workflow, and call the heterogeneous device abstraction module to execute the workflow.

[0006] According to the operation and maintenance system provided by the present application, the system further comprises a multi-tenant service module configured to allow tenants to build and submit declarative service definitions to the orchestration module. The orchestration module is further configured to: receive externally submitted declarative service definitions, parse the declarative service definitions into the workflow, and call the heterogeneous device abstraction module to execute the workflow.

[0007] According to the operation and maintenance system provided by the application, the system further comprises a multi-tenant management module. The multi-tenant management module is configured to enforce an access control policy and data isolation on the service definition received by the orchestration module and the data accessed by the intelligent operation and maintenance module based on a tenant identity.

[0008] According to the operation and maintenance system provided by the application, the orchestration module is specifically configured to: In the case that the orchestration module receives the fault root cause analysis report, the orchestration module is configured to automatically generate the workflow for fault isolation.

[0009] According to the operation and maintenance system provided by the application, the heterogeneous device abstraction module is specifically configured to: The heterogeneous device abstraction module is configured to load a pluggable vendor adaptation driver corresponding to the heterogeneous network device, and realize conversion of instructions and data between the unified device model and a private interface of a device vendor.

[0010] The application further provides an operation and maintenance method based on the operation and maintenance system, comprising the following steps: The API gateway provides and manages a network capability API, and the network capability API corresponds to resource management, service configuration and device monitoring capability of an underlying network management system. The network capability API is called through the heterogeneous device abstraction module, telemetry data of the heterogeneous network device is collected, and the telemetry data is standardized to generate a standardized real-time telemetry data stream. The intelligent operation and maintenance module is configured to perform real-time analysis on the standardized real-time telemetry data stream, and generate a fault root cause analysis report. The orchestration module is configured to parse the fault root cause analysis report into a workflow, and call the heterogeneous device abstraction module to execute the workflow.

[0011] According to the operation and maintenance method provided by the application, the method further comprises the following steps: The method further comprises the following steps:

[0012] The application further provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the operation and maintenance method according to any one of the above when executing the computer program.

[0013] The application further provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the operation and maintenance method according to any one of the above.

[0014] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements the operation and maintenance method according to any one of the above.

[0015] The operation and maintenance system, method, electronic device, storage medium and program product provided by the application, the heterogeneous device abstraction module, and the unified device model are used to standardize and abstract the devices of different manufacturers at the bottom layer. The design shields the differences between the bottom layer hardware, provides a unified and regular data view and control interface for the upper layer application, and thus realizes unified and standardized management of the heterogeneous network. The intelligent operation and maintenance module can directly process the standardized telemetry data provided by the heterogeneous device abstraction module, and actively performs intelligent analysis of abnormal detection and fault root cause. This improves the traditional passive and lagging operation and maintenance mode to an active and predictive intelligent operation and maintenance mode, and significantly improves the fault processing efficiency. Finally, the arrangement module can automatically translate the business requirements at the high layer into the bottom layer network configuration, realizes rapid and automatic deployment of the business, breaks through the whole link from the business intention to the network execution and from the intelligent analysis of the fault to the automatic repair, builds a closed loop of the automatic operation and maintenance, and greatly improves the agility and reliability of the network service. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 is a schematic diagram of the operation and maintenance system structure provided by the application; Figure 2 is a schematic diagram of the operation and maintenance method flow provided by the application; Figure 3 is a schematic diagram of the structure of the electronic device provided by the application. DETAILED DESCRIPTION

[0018] In order to make the objects, technical solutions and advantages of the application clearer, the technical solutions in the application will be described clearly and completely below with reference to the drawings in the application. Obviously, the described embodiments are some embodiments of the application, but not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0019] Figure 1 is a schematic diagram of the operation and maintenance system structure provided by the application, as Figure 1As shown, including: API gateway 11, heterogeneous device abstraction module 12, intelligent operation and maintenance module 13, orchestration module 14; The API gateway 11 is used for providing and managing network capability API, the network capability API corresponds to the resource management, service configuration and device monitoring capability of the underlying network management system; In the application, the API gateway is a middleware component, which acts as an intermediary between client applications and backend services, and plays a unified entry role in the system, all client requests first pass through the API gateway, and then are forwarded to the corresponding backend service by the API gateway.

[0020] In the application, all externally provided APIs are centrally managed, and the client only needs to interact with the API gateway without directly dealing with multiple backend services. According to the URL, method and other information of the request, the request is routed to the correct backend service. Moreover, the API gateway can combine the requests of multiple backend services and assemble them into a response returned to the client.

[0021] The network capability API is a set of defined interfaces through which external applications can access and operate the functions of the underlying network management system.

[0022] The resource management capability enables the API to effectively allocate, schedule and recycle various resources in the network, such as devices, ports, bandwidth, etc., to ensure rational use and optimal configuration of resources; The service configuration capability allows the API to quickly deploy, adjust and optimize network services to meet changing business needs; The device monitoring capability ensures that the running state of network devices can be obtained in real time, and potential fault risks can be discovered and handled in a timely manner.

[0023] Another important function of the API gateway is to finely manage these network capability APIs. The management process covers the entire life cycle of the API, from the initial creation, release to subsequent maintenance and update, and each link needs to be strictly controlled to ensure the stability and reliability of the API. In the API release stage, the gateway needs to accurately set the access path, request method and other basic information of the API, and reasonably plan the version iteration of the API to avoid unnecessary impact on existing business due to version update.

[0024] On the one hand, strict identity authentication is required for each access API request to verify whether the requestor has legal access qualification; On the other hand, according to the permission level of the requestor, the API range and specific function permission that can be accessed are accurately authorized, so as to effectively prevent unauthorized access and malicious operation.

[0025] In addition, during system operation, the API gateway also shoulders the heavy responsibility of flow limiting and monitoring. The flow limiting measure can prevent a certain API or a certain client from occupying too many system resources due to excessive calls, thereby affecting the performance and stability of the entire system. At the same time, the gateway monitors the access of the API in real time, including recording the frequency of requests, the time of responses, the error rate and other key indicators.

[0026] The heterogeneous device abstraction module 12 is configured to call the network capability API through the API gateway; and the heterogeneous device abstraction module is configured to provide a real-time telemetry data stream standardized through a unified device model; In the present application, the heterogeneous device abstraction module does not directly interact with the underlying network management system, but uses the network capability API provided by the API gateway to realize the call of the underlying network function through the intermediate link of the API gateway, so that the heterogeneous device abstraction module can interact with different network management systems in a more standardized and unified manner.

[0027] And the heterogeneous device abstraction module is configured to provide a real-time telemetry data stream standardized through a unified device model. This indicates that the heterogeneous device abstraction module undertakes the important task of processing and standardizing telemetry data from different devices.

[0028] Telemetry data refers to various running state data generated by network devices in real time, and these data have extremely important value for network monitoring and operation and maintenance. However, due to the existence of various types of devices in the network, the telemetry data generated by these devices may have great differences in format, semantics and other aspects, which brings difficulties to the unified processing and analysis of data.

[0029] In order to solve this problem, the heterogeneous device abstraction module introduces the concept of unified device model. The unified device model is a model that abstracts and generalizes the common features of different manufacturers and different types of devices.

[0030] By mapping the telemetry data of various devices to this unified device model, the heterogeneous device abstraction module can convert the originally scattered and heterogeneous data into a data stream with unified format and semantics.

[0031] This standardized real-time telemetry data stream not only facilitates subsequent analysis and processing, but also improves the interoperability and scalability of the entire system.

[0032] The heterogeneous device abstraction module calls network capability APIs through the API gateway, realizes the call of underlying network functions, and utilizes the unified device model to standardize the telemetry data from different devices, thereby providing high-quality, easily processed real-time data basis for the entire operation and maintenance system.

[0033] The intelligent operation and maintenance module 13 is configured to receive the standardized real-time telemetry data stream, and analyze the standardized real-time telemetry data stream based on knowledge graph correlation analysis, to generate a fault root cause analysis report. In the present application, the intelligent operation and maintenance module first undertakes the task of receiving the real-time telemetry data stream after standardization processing.

[0034] The real-time telemetry data stream is generated after conversion by the heterogeneous device abstraction module through the unified device model, and has a unified format and semantics.

[0035] The real-time telemetry data stream contains various types of key information generated by network devices during operation, such as device performance indicators, port traffic, system logs, etc. These key information is an important basis for monitoring network operation status and timely discovering potential problems.

[0036] The knowledge graph is a powerful semantic modeling tool that can represent various entities in the network (such as devices, services, fault phenomena, configuration parameters, etc.) and their complex relationships in the form of a graph.

[0037] The heterogeneous device abstraction module can link seemingly isolated data points together and mine potential associations and patterns by constructing a knowledge graph.

[0038] In the analysis process, the intelligent operation and maintenance module utilizes the semantic relationships and reasoning rules in the knowledge graph to perform multi-dimensional correlation analysis on real-time telemetry data. For example, when the CPU usage of a certain device suddenly increases, the module can quickly locate the service carried by the device, the associated other devices, and the part of the network topology that may be affected with the help of the knowledge graph. At the same time, combined with historical data and known fault patterns, the module can comprehensively evaluate the current situation, judge whether a fault is likely to occur, and the possible type and scope of the fault.

[0039] After completing the analysis, the intelligent operation and maintenance module can generate a fault root cause analysis report. The fault root cause analysis report records the analysis results of the real-time telemetry data stream, clearly pointing out the root cause of the fault, such as device hardware failure, software defects, configuration errors, network congestion, etc. The report may also contain the impact range of the fault, severity assessment, and recommended repair measures, etc. Through such a fault root cause analysis report, the operation and maintenance personnel can quickly understand the overall situation of the fault, so as to take precise and effective measures to handle the fault, reduce the impact of the fault on the business, and ensure the stable operation of the network.

[0040] The orchestration module 14 is used to parse the fault root cause analysis report into a workflow, and call the heterogeneous device abstraction module to execute the workflow.

[0041] In the present application, the orchestration module plays a key role in automated processing and coordinated execution in the operation and maintenance system. Its main function is to convert the fault root cause analysis report generated by the intelligent operation and maintenance module into a series of operational task sets, and coordinate system resources to complete the execution of tasks.

[0042] Firstly, the orchestration module receives the fault root cause analysis report from the intelligent operation and maintenance module. The report describes in detail the root cause of the fault, the impact range, related devices, and recommended repair measures, etc. The orchestration module needs to structure the report content and extract executable operation steps. This step involves understanding and converting the report content into a workflow definition that the system can recognize and execute.

[0043] The workflow is essentially a sequence of tasks, containing multiple interrelated operation steps, each of which may correspond to different devices, configuration changes or data operations. The orchestration module generates a specific workflow according to the fault root cause analysis report, combined with pre-defined workflow templates and rules in the system. The generation process of the workflow needs to consider the type of the fault, the range of affected devices, business priority, etc., to ensure that the generated workflow can efficiently and accurately solve the problem.

[0044] After the workflow is generated, the orchestration module is responsible for calling the heterogeneous device abstraction module to execute the workflow. The orchestration module interacts with the heterogeneous device abstraction module through a standardized interface, sending execution instructions and parameters. The heterogeneous device abstraction module uses its abstract models and API calling capabilities for different devices to execute specific operations according to the received instructions. This may include configuring device parameters, executing fault recovery processes, restarting services, etc.

[0045] During execution, the orchestration module continuously monitors the progress and execution status of the workflow. It needs to handle abnormal situations that may occur, such as the failure of a certain task, the unresponsiveness of a device, etc. According to the preset exception handling strategy, the orchestration module decides whether to retry the task, skip the task or terminate the entire workflow. At the same time, the orchestration module records log information during the execution process, so as to facilitate subsequent audit and analysis.

[0046] In addition, the orchestration module is also responsible for coordinating the execution of multiple workflows, avoiding resource conflicts or task blocking. It arranges the execution order of tasks through scheduling algorithms to ensure the stability and efficiency of the system. After the execution of the workflow is completed, the orchestration module generates an execution report, recording the execution result, completion time, resource usage, etc. of the workflow, so that the operation and maintenance personnel can understand the effect and efficiency of the entire fault handling process.

[0047] In the present application, the heterogeneous device abstraction module abstracts the underlying devices of different manufacturers using a unified device model. This design shields the differences between the underlying hardware and provides a unified and standardized data view and control interface for the upper layer application, thereby realizing the unified and standardized management of heterogeneous networks. The intelligent operation and maintenance module can directly process the standardized telemetry data provided by the heterogeneous device abstraction module, actively detect abnormalities and intelligently analyze fault root causes. This improves the traditional passive and lagging operation and maintenance mode to an active and predictive intelligent operation and maintenance mode, significantly improving the fault handling efficiency. Finally, the orchestration module can automatically translate high-level business requirements into underlying network configurations, realizing the rapid and automated deployment of services; it breaks through the whole link from business intent to network execution and from intelligent fault analysis to automatic repair, builds a closed loop of automated operation and maintenance, and greatly improves the agility and reliability of network services.

[0048] Optionally, the system further comprises a multi-tenant service module, which is configured to allow tenants to build and submit declarative service definitions to the orchestration module. The orchestration module is further configured to: receive externally submitted declarative service definitions and parse the declarative service definitions into the workflow, and call the heterogeneous device abstraction module to execute the workflow.

[0049] In the present application, the multi-tenant service module enables different tenants to build and submit declarative service definitions to the orchestration module. It meets the individual needs of each tenant for network services in a multi-tenant environment, while ensuring resource isolation and security between tenants.

[0050] The multi-tenant service module provides a user-friendly interface or API, allowing tenants to define the required services according to their business needs without needing to understand the complex configuration details of the underlying network devices.

[0051] The multi-tenant service module takes into account the diverse needs of tenants and the scalability of the system. It allows each tenant to define services within their own logical space, which can cover network topology, device configuration requirements, quality of service parameters, access control policies, and other aspects. Through a declarative approach, tenants only need to specify the desired service state without worrying about the specific implementation process. For example, a tenant can define a service that requires a high-priority virtual private network connection between specified network nodes without needing to know the specific configuration commands of the underlying devices.

[0052] After receiving these declarative service definitions, the orchestration module parses and converts them into specific workflows. The orchestration module has a built-in service definition parser that can understand the semantics of declarative service definitions and map them to a series of executable operation steps. This process involves syntax analysis, semantic verification, and matching with system resources of service definitions. For example, the orchestration module will identify the network device types, configuration parameter requirements, and other information involved in the service definition, and generate corresponding workflow tasks based on this information, including device configuration changes, resource allocation or release, and other operations.

[0053] Next, the orchestration module calls the heterogeneous device abstraction module to execute these workflows. The heterogeneous device abstraction module uses its abstract models and API calling capabilities for different vendor devices to convert the tasks in the workflow into specific device operation instructions. These instructions are sent to the underlying network management system through the API gateway, which completes the actual device configuration and resource management operations. During execution, the orchestration module also monitors the execution status of the workflow to ensure smooth completion of tasks and handles exceptions when necessary.

[0054] In addition, the collaborative work of the multi-tenant service module and the orchestration module also involves resource management and permission control. The multi-tenant service module ensures that each tenant can only access and operate resources within their own permission scope, preventing resource conflicts and security leaks. The orchestration module also limits operations according to tenant permissions when parsing and executing workflows, ensuring the security and stability of the system.

[0055] In this invention, the combination of the multi-tenant service module and the orchestration module provides the system with strong multi-tenant support and flexible service definition and execution capabilities. This design not only improves the user-friendliness and scalability of the system, but also improves the operation and maintenance efficiency through automated workflow processing, reducing the need for manual intervention, so that the entire network operation and maintenance system can better adapt to complex and changing business environments.

[0056] Optionally, the system further comprises a multi-tenant management module; The multi-tenant management module is configured to enforce access control policies and data isolation based on tenant identity for service definitions received by the orchestration module and data accessed by the intelligent operation and maintenance module.

[0057] In the present application, the multi-tenant management module enforces access control policies and data isolation based on tenant identity for service definitions received by the orchestration module and data accessed by the intelligent operation and maintenance module. The core role of the multi-tenant management module is to ensure that each tenant can only operate and access resources within its authorized scope, thereby maintaining the security of the system and the confidentiality of the data.

[0058] The multi-tenant management module first needs to accurately identify and verify the tenant identity. This is usually achieved by integrating with identity authentication services (such as LDAP, OAuth, etc.), ensuring that the identity of each tenant in the system is unique and traceable. Once the tenant identity is confirmed, the multi-tenant management module begins to monitor and manage all operation requests of the tenant.

[0059] For service definitions received by the orchestration module, the multi-tenant management module verifies whether the tenant has the right to submit and execute specific service definitions according to the pre-defined access control policy. The access control policy may be based on role (RBAC), attribute (ABAC), or other custom rules. For example, a certain tenant may be authorized to define and submit network configuration services related to its own business, but it has no right to operate the service definitions of other tenants. The multi-tenant management module ensures that only service definitions that meet the policy can be accepted and processed by the orchestration module, thereby preventing unauthorized service definitions from potentially affecting the system and other tenants.

[0060] Similarly, for the data accessed by the intelligent operation and maintenance module, the multi-tenant management module enforces data isolation policies. This means that the data of each tenant can only be accessed by the users and system components authorized by the tenant, and cannot be directly obtained or modified by other tenants. At the data storage level, the system may use a multi-tenant database architecture to ensure the independence of tenant data through physical or logical isolation. For example, independent database instances or tenant identification fields in the database are used to distinguish the tenant to which the data belongs. In addition, during data transmission and processing, the multi-tenant management module also ensures the confidentiality and integrity of the data to prevent data leakage or tampering.

[0061] In the present application, the multi-tenant management module closely cooperates with the orchestration module and the intelligent operation and maintenance module to ensure the security and stability of the entire system in a multi-tenant environment.

[0062] Optionally, the orchestration module is specifically configured to: In the case that the orchestration module receives the fault root cause analysis report, the orchestration module is configured to automatically generate the workflow of fault isolation.

[0063] In the case that the orchestration module receives the fault root cause analysis report, the orchestration module is configured to automatically generate the workflow of fault isolation.

[0064] The orchestration module has the ability to automatically generate the workflow of fault isolation after receiving the fault root cause analysis report, which is a key link for the intelligent operation and maintenance system to realize automatic fault processing. When the orchestration module obtains the fault root cause analysis report generated by the intelligent operation and maintenance module, it first comprehensively analyzes the report content. This includes identifying key information such as the specific type of the fault, the location of the fault, the scope of the impact, and the network devices involved.

[0065] Based on the analysis result, the orchestration module will automatically trigger the fault isolation mechanism. It has multiple fault isolation strategies pre-installed inside, which are designed according to common fault patterns and network topology structures. For example, if the network problem is caused by device hardware failure, the orchestration module will generate a workflow to guide the heterogeneous device abstraction module to call the functions of the underlying network management system through the API gateway, and perform isolation operations on the device, such as closing the fault port or rerouting traffic to the standby link.

[0066] In addition, the orchestration module is also responsible for coordinating the multi-tenant management module to ensure that the resources and services of different tenants are not affected during the fault isolation process, and to maintain the isolation and security between tenants. Through cooperation with the multi-tenant management module, it can adjust the isolation strategy according to the tenant's permissions and resource allocation, to avoid unnecessary interference to the business of other tenants.

[0067] During the automatic execution of the fault isolation workflow, the orchestration module continuously monitors the execution status. If a certain isolation step fails or an exception occurs, it will adjust according to the pre-set exception handling strategy, such as retry, rollback or manual intervention. At the same time, detailed execution logs are recorded for subsequent audit and analysis.

[0068] Optionally, the heterogeneous device abstraction module is specifically used for: Loading the plug-in vendor adaptation driver corresponding to the heterogeneous network device to realize the conversion of instructions and data between the unified device model and the private interface of the device vendor.

[0069] In the present application, the unified device model defines a set of general attributes, methods and data structures for representing and operating network devices.

[0070] The unified device model masks the specific differences of underlying devices, allowing upper-layer applications (such as orchestration modules and intelligent operation and maintenance modules) to interact with various devices in a unified manner. For example, the unified device model may define a general method to query the interface state of a device, and different vendor adaptation drivers are responsible for converting this general method into specific device commands.

[0071] When the heterogeneous device abstraction module needs to interact with a certain device, it selects and loads the corresponding vendor adaptation driver according to the type and vendor information of the device. After the adaptation driver is loaded, the module converts the instructions in the unified device model into private instructions that the device can understand by calling the interfaces provided by the adaptation driver. For example, the unified device model may indicate that the interface traffic statistics information of the device needs to be obtained, and the adaptation driver will convert this indication into a query command for a specific vendor device, such as querying the interface traffic of a Cisco device through SNMP OID or querying the interface traffic of a Huawei device through NetConf protocol.

[0072] Similarly, when the device returns data, the adaptation driver is responsible for parsing and converting these raw data into a data structure that conforms to the unified device model. This may involve operations such as data format conversion, field mapping, and unit conversion of numerical values. The converted data can be directly used by the upper-layer application without needing to worry about the specific format of the data source device.

[0073] Figure 2 The operation and maintenance method provided by the present application provides a flowchart of the operation and maintenance method, as shown in Figure 2 The operation and maintenance method provided by the present application provides a flowchart of the operation and maintenance method, as shown in Step 210, the API gateway provides and manages network capability APIs, which correspond to the resource management, service configuration and device monitoring capabilities of the underlying network management system; In the present application, first, the API gateway provides network capability APIs, covering the resource management, service configuration and device monitoring functions of the underlying system. This includes full life cycle management of resources, rapid configuration adjustment of services and real-time monitoring of devices.

[0074] Secondly, the API gateway manages these APIs to ensure their safe, stable and efficient use. Management measures involve API creation, release, authentication and authorization, as well as flow control and monitoring. Through the authentication and authorization mechanism, the API gateway verifies the identity and permissions of the requester, ensuring that only legitimate requests can access the corresponding functions. Flow control measures prevent APIs from being called excessively, ensuring system stability. At the same time, the API gateway monitors API access and collects relevant indicators to promptly discover and solve problems. The API gateway realizes the opening and management of the capabilities of the underlying network management system, providing a convenient access method for upper-layer applications and supporting the unified management and automatic operation and maintenance of special network devices.

[0075] Step 220, through the heterogeneous device abstraction module, calling the network capability API, collecting telemetry data of heterogeneous network devices, and standardizing the telemetry data to generate standardized real-time telemetry data stream; The heterogeneous device abstraction module calls the network capability API to collect telemetry data of heterogeneous network devices, which is provided and managed by the API gateway.

[0076] The heterogeneous device abstraction module communicates with devices of different manufacturers by loading pluggable manufacturer adaptation drivers, collects raw telemetry data including device performance indicators, interface traffic, system logs, etc. The collected data is then standardized to unify data format and semantics, generating standardized real-time telemetry data stream.

[0077] Step 230, through the intelligent operation and maintenance module, real-time analysis of the standardized real-time telemetry data stream, generating fault root cause analysis report; The intelligent operation and maintenance module receives and analyzes these standardized telemetry data streams in real time.

[0078] Using knowledge graph correlation analysis technology, combining preset rules and machine learning algorithms, the data is deeply mined to identify abnormal patterns and potential faults. Based on the analysis results, the intelligent operation and maintenance module generates a fault root cause analysis report detailing the fault cause, impact range, and recommended repair measures.

[0079] Step 240, through the orchestration module, the fault root cause analysis report is parsed into a workflow, and the heterogeneous device abstraction module is called to execute the workflow.

[0080] The orchestration module receives the fault root cause analysis report and parses it into a specific workflow. According to the report content and the pre-defined workflow template, the orchestration module generates a series of automated tasks such as configuration changes, device restarts, or traffic rerouting, etc.

[0081] The orchestration module then calls the heterogeneous device abstraction module to execute these tasks to achieve automatic isolation and repair of faults. During execution, the orchestration module monitors task progress to ensure smooth completion of the workflow.

[0082] In the present application, the heterogeneous device abstraction module standardizes and abstracts different manufacturers' devices at the bottom layer by using a unified device model. This design shields the differences in the bottom layer hardware and provides a unified and regular data view and control interface for the upper layer application, thereby realizing the unified and standardized management of the heterogeneous network. The intelligent operation and maintenance module can directly process the standardized telemetry data provided by the heterogeneous device abstraction module, and actively perform intelligent analysis of abnormal detection and fault root cause. This improves the traditional passive and lagging operation and maintenance mode to an active and predictive intelligent operation and maintenance mode, and significantly improves the fault processing efficiency. Finally, the orchestration module can automatically translate the high-level business requirements into the bottom layer network configuration, realizes the rapid and automatic deployment of the business, and breaks through the whole link from the business intention to the network execution and from the intelligent analysis of the fault to the automatic repair, thereby constructing the closed loop of the automatic operation and maintenance and greatly improving the agility and reliability of the network service.

[0083] Optionally, according to the operation and maintenance method of the present application, the method further comprises: receiving an externally submitted declarative service definition, and parsing the declarative service definition into the workflow, and calling the heterogeneous device abstraction module to execute the workflow.

[0084] In the present application, the system receives declarative service definitions from different tenants through the multi-tenant service module. These service definitions describe the network service state expected by the tenants in a declarative manner, including network topology, device configuration, service quality requirements, etc., without detailing the specific implementation steps.

[0085] After the orchestration module obtains these declarative service definitions, it parses them and converts them into a series of specific workflow tasks. This process involves semantic understanding of the service definition and matching with system resources to determine the specific operations required to achieve the service state expected by the tenants.

[0086] The orchestration module then calls the heterogeneous device abstraction module to execute the corresponding tasks according to the parsed workflow. The heterogeneous device abstraction module uses the manufacturer adaptation driver loaded by it to convert the tasks in the workflow into instructions for specific manufacturer devices, and calls the network capability API through the API gateway to realize the configuration and management of the devices.

[0087] Figure 3 is a structural schematic diagram of an electronic device provided by the present application, as Figure 3As shown, the electronic device can include a processor 310, a communications interface 320, a memory 330 and a communications bus 340, wherein the processor 310, the communications interface 320 and the memory 330 complete the communication with each other through the communications bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the operation and maintenance method, which includes that the API gateway provides and manages the network capability API, and the network capability API corresponds to the resource management, service configuration and device monitoring capability of the underlying network management system; Through the heterogeneous device abstraction module, the network capability API is called, the telemetry data of the heterogeneous network device is collected, and the telemetry data is standardized to generate a standardized real-time telemetry data stream; Through the intelligent operation and maintenance module, the standardized real-time telemetry data stream is analyzed in real time to generate a fault root cause analysis report; Through the arrangement module, the fault root cause analysis report is parsed into a workflow, and the heterogeneous device abstraction module is called to execute the workflow.

[0088] In addition, the logical instructions in the memory 330 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk and various program code storage media.

[0089] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, and the computer can execute the operation and maintenance method provided by the above-mentioned method, which includes that the API gateway provides and manages the network capability API, and the network capability API corresponds to the resource management, service configuration and device monitoring capability of the underlying network management system; The heterogeneous device abstraction module calls the network capability API, collects telemetry data of the heterogeneous network device, and performs standardization processing on the telemetry data to generate a standardized real-time telemetry data stream. The intelligent operation and maintenance module performs real-time analysis on the standardized real-time telemetry data stream to generate a fault root cause analysis report. The orchestration module parses the fault root cause analysis report into a workflow, and calls the heterogeneous device abstraction module to execute the workflow.

[0090] In another aspect, the application also provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the operation and maintenance method provided by the above method, the method comprising: an API gateway providing and managing network capability APIs corresponding to resource management, service configuration and device monitoring capabilities of an underlying network management system; The heterogeneous device abstraction module calls the network capability API, collects telemetry data of the heterogeneous network device, and performs standardization processing on the telemetry data to generate a standardized real-time telemetry data stream. The intelligent operation and maintenance module performs real-time analysis on the standardized real-time telemetry data stream to generate a fault root cause analysis report. The orchestration module parses the fault root cause analysis report into a workflow, and calls the heterogeneous device abstraction module to execute the workflow.

[0091] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments. Those skilled in the art can understand and implement without creative labor.

[0092] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0093] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An operation and maintenance system, characterized in that, include: API gateway, heterogeneous device abstraction module, intelligent operation and maintenance module, orchestration module; The API gateway is used to provide and manage network capability APIs, which correspond to the resource management, service configuration and device monitoring capabilities of the underlying network management system. The heterogeneous device abstraction module is configured to call the network capability API through the API gateway; Furthermore, the heterogeneous device abstraction module is used to provide real-time telemetry data streams standardized through a unified device model; The intelligent operation and maintenance module is used to receive the standardized real-time telemetry data stream, and analyze the standardized real-time telemetry data stream based on knowledge graph association analysis to generate a fault root cause analysis report. The orchestration module is used to parse the fault root cause analysis report into a workflow and call the heterogeneous device abstraction module to execute the workflow.

2. The operation and maintenance system according to claim 1, characterized in that, The system further includes a multi-tenant service module, which is used for tenants to construct and submit declarative service definitions to the orchestration module; The orchestration module is also used for: Receive externally submitted declarative service definitions, parse the declarative service definitions into the workflow, and call the heterogeneous device abstraction module to execute the workflow.

3. The operation and maintenance system according to claim 1, characterized in that, The system also includes: a multi-tenant management module; The multi-tenant management module is used to enforce access control policies and data isolation on the service definitions received by the orchestration module and the data accessed by the intelligent operation and maintenance module based on the tenant identity.

4. The operation and maintenance system according to claim 1, characterized in that, The orchestration module is specifically used for: Upon receiving the fault root cause analysis report, the orchestration module is configured to automatically generate the fault isolation workflow.

5. The operation and maintenance system according to claim 1, characterized in that, The heterogeneous device abstraction module is specifically used for: Load pluggable vendor-compatible drivers corresponding to heterogeneous network devices to realize instruction and data conversion between the unified device model and the device vendor's private interface.

6. A method for maintaining and operating a system based on any one of claims 1-5, characterized in that, include: The API gateway provides and manages network capability APIs, which correspond to the resource management, service configuration, and device monitoring capabilities of the underlying network management system. Through the heterogeneous device abstraction module, the network capability API is called to collect telemetry data from heterogeneous network devices, and the telemetry data is standardized to generate a standardized real-time telemetry data stream. The standardized real-time telemetry data stream is analyzed in real time through the intelligent operation and maintenance module to generate a fault root cause analysis report. The orchestration module parses the fault root cause analysis report into a workflow and calls the heterogeneous device abstraction module to execute the workflow.

7. The operation and maintenance method according to claim 6, characterized in that, The method further includes: Receive externally submitted declarative service definitions, parse the declarative service definitions into the workflow, and call the heterogeneous device abstraction module to execute the workflow.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the operation and maintenance method as described in any one of claims 6 to 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the operation and maintenance method as described in any one of claims 6 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the operation and maintenance method as described in any one of claims 6 to 7.