Operation and maintenance method, device and equipment based on integrated operation and maintenance portal platform and medium

By integrating data collection, alarm output, fault location and unified management functions on the operation and maintenance portal platform, the problems of dispersion and lack of automation of operation and maintenance tools in the existing technology are solved, and more efficient operation and maintenance management and more stable infrastructure are achieved.

CN119996171APending Publication Date: 2025-05-13CHINA SOUTHERN POWER GRID DIGITAL GRID GROUP (GUANGDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510094828.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology operation and maintenance tools are scattered and lack of automated and intelligent means, which leads to operation and maintenance personnel need to directly manage each infrastructure object, which is very labor-intensive and inefficient.

Method used

An operation and maintenance method based on an integrated operation and maintenance portal platform is proposed. Data collection of infrastructure equipment is collected through the data acquisition module, alarm information is output, fault cause and location are located, and the configuration information, firmware version and storage space of the equipment are uniformly managed.

Benefits of technology

It improves the operation and maintenance efficiency of infrastructure equipment, reduces operation and maintenance costs, and improves the stability and security of infrastructure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996171A_ABST
    Figure CN119996171A_ABST
Patent Text Reader

Abstract

The invention discloses an operation and maintenance method, device and equipment based on an integrated operation and maintenance portal platform and a medium, and relates to the technical field of computers, and the method comprises the steps: collecting different types of infrastructure equipment through a data collection module through the integrated operation and maintenance portal platform to obtain collection data; outputting alarm information according to the collected data; locating the reason and position of the fault according to the collected data; and uniformly managing at least one of the configuration information, the firmware version and the storage space of each infrastructure equipment. Based on the integrated operation and maintenance portal platform, the method has remarkable advantages in the aspects of heterogeneous equipment compatibility, multi-mode acquisition of various equipment data, integrated operation and maintenance management, unified display management, comprehensive creativity and the like, so that the problem of low efficiency in the aspect of infrastructure equipment management in the prior art can be better solved; and the operation and maintenance efficiency is improved, the operation and maintenance cost is reduced, and the stability and safety of infrastructures are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an operation and maintenance method, device, equipment and medium based on an integrated operation and maintenance portal platform. Background Art

[0002] As enterprise informatization continues to deepen, the scale of infrastructure equipment is growing, and operation and maintenance work is becoming more and more complicated. The existing operation and maintenance tools are scattered and lack automation and intelligent means. Operation and maintenance personnel are required to directly manage each infrastructure object, which is a large workload and inefficient. In addition, with the rapid development of technologies such as cloud computing, big data, and artificial intelligence, the demand for automation and intelligence in operation and maintenance work is also increasing. Summary of the invention

[0003] The main purpose of the embodiments of the present application is to propose an operation and maintenance method, device, equipment and medium based on an integrated operation and maintenance portal platform to improve the operation and maintenance efficiency of infrastructure equipment.

[0004] To achieve the above-mentioned purpose, one aspect of an embodiment of the present application proposes an operation and maintenance method based on an integrated operation and maintenance portal platform, the method is applied to the integrated operation and maintenance portal platform, and the method comprises the following steps:

[0005] The data acquisition module collects data from different types of infrastructure equipment;

[0006] Outputting alarm information according to the collected data;

[0007] Locate the cause and location of the fault according to the collected data;

[0008] Centrally manage at least one of the configuration information, firmware version and storage space of each of the infrastructure devices.

[0009] In some embodiments, the collecting data from different types of infrastructure equipment by the data collection module includes the following steps:

[0010] The data collection module collects device hardware configuration, network configuration, health status, energy consumption data, performance and capacity data of server devices, network devices and storage devices according to server out-of-band management, network device SNMP protocol, storage device API, SSH, SMIS and SNMP as the collected data.

[0011] In some embodiments, the step of outputting warning information according to the collected data comprises the following steps:

[0012] Compare thresholds corresponding to device hardware configuration, network configuration, health status, energy consumption data, performance and capacity data of server devices, network devices and storage devices to obtain comparison results;

[0013] The alarm information is output according to the comparison result; wherein the alarm information includes an alarm description, alarm location information and suggested steps to resolve the alarm.

[0014] In some embodiments, locating the cause and location of the fault according to the collected data includes the following steps:

[0015] The causes of the faults located according to the collected data include hardware failure, software failure and network failure;

[0016] Based on the topological relationship between the network device port and other devices and the topological relationship between the fiber optic switch and the storage device and the host, and according to the cause of the fault, the fault location and the impact range are located.

[0017] In some embodiments, the step of uniformly managing the configuration information of each of the infrastructure devices comprises the following steps:

[0018] Centrally manage the hardware configuration, network settings, access control and centralized reading and modification of settings of servers in each of the infrastructure devices;

[0019] Unified management of network devices and security devices in each of the infrastructure devices so that the network devices and the security devices regularly back up corresponding device data;

[0020] The network devices in each of the infrastructure devices are managed uniformly so that the network devices can recover data from the most adjacent multiple infrastructure devices.

[0021] In some embodiments, the step of uniformly managing the firmware versions of each of the infrastructure devices comprises the following steps:

[0022] Regularly scan the firmware version of each of the infrastructure devices and compare it with the latest firmware library; when it is detected that the firmware version is lower than the set version number, output upgrade information to the operation and maintenance personnel;

[0023] Batch firmware version upgrade for multiple infrastructure devices of the same model;

[0024] Scan the firmware versions of each of the infrastructure devices to determine the infrastructure devices with firmware risks, and then inform the operation and maintenance personnel in a centralized page information manner or email manner.

[0025] In some embodiments, the step of uniformly managing the storage space of each of the infrastructure devices comprises the following steps:

[0026] Collecting information about storage space of SAN storage, NAS storage, distributed storage, fiber switches, and backup devices from each of the infrastructure devices, and then managing each of the storage spaces in a unified manner;

[0027] Compare the IPOS, throughput and response time of each storage space, and then select a number of storage spaces with the highest utilization rate;

[0028] The capacity usage trend of each of the storage spaces is analyzed, and the capacity exhaustion date is predicted.

[0029] To achieve the above-mentioned purpose, another aspect of the embodiment of the present application proposes an operation and maintenance device based on an integrated operation and maintenance portal platform, the device comprising:

[0030] A data collection unit, used to collect data from different types of infrastructure equipment through a data collection module;

[0031] An alarm output unit, used to output alarm information according to the collected data;

[0032] A fault locating unit, used to locate the cause and position of the fault according to the collected data;

[0033] The unified management unit is used to uniformly manage at least one of the configuration information, firmware version and storage space of each of the infrastructure devices.

[0034] To achieve the above objective, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above method when executing the computer program.

[0035] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0036] The embodiments of the present application include at least the following beneficial effects:

[0037] This application can collect data from different types of infrastructure equipment through the data collection module through the integrated operation and maintenance portal platform; output alarm information based on the collected data; locate the cause and location of the fault based on the collected data; and uniformly manage at least one of the configuration information, firmware version, and storage space of each infrastructure device. This application is based on an integrated operation and maintenance portal platform, and has significant advantages in terms of compatibility with heterogeneous devices, multi-mode collection of data from various types of equipment, integrated operation and maintenance management, unified display management, and comprehensive information innovation. The above advantages enable this application to better solve the low efficiency problem of existing technologies in infrastructure equipment management, and improve operation and maintenance efficiency, reduce operation and maintenance costs, and enhance the stability and security of infrastructure. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0039] Figure 1 A flowchart of an operation and maintenance method based on an integrated operation and maintenance portal platform provided in an embodiment of the present application;

[0040] Figure 2 A schematic diagram of the system structure of the infrastructure automated operation and maintenance of the integrated operation and maintenance portal platform provided in the embodiment of the present application;

[0041] Figure 3 A flowchart of an embodiment of centralized monitoring and alarming provided by an embodiment of the present application;

[0042] Figure 4 An example diagram of device scanning provided in an embodiment of the present application;

[0043] Figure 5 An example diagram of a monitoring list provided in an embodiment of the present application;

[0044] Figure 6 An example diagram of a single device monitoring overview provided in an embodiment of the present application;

[0045] Figure 7 An example diagram of the details of a single device monitoring indicator provided in an embodiment of the present application;

[0046] Figure 8 An example diagram of an overview of alarm events provided in an embodiment of the present application;

[0047] Fig. 9 An example diagram of a current alarm event list provided in an embodiment of the present application;

[0048] Fig.10 An example diagram of the alarm notification rules provided in the embodiment of the present application;

[0049] Fig.11 An example diagram of an ARP information list provided in an embodiment of the present application;

[0050] Fig.12 A flowchart of an embodiment of automatic operation and maintenance of a network topology provided in an embodiment of the present application;

[0051] Fig.13 An example diagram of creating a network topology provided in an embodiment of the present application;

[0052] Fig.14 An example diagram of the real-time status of the network topology provided in the embodiment of the present application;

[0053] Fig.15 A flowchart of an embodiment of network configuration backup and distribution provided in an embodiment of the present application;

[0054] Fig.16 An example diagram of a configuration template provided in an embodiment of the present application;

[0055] Fig.17 An example diagram of a backup task provided in an embodiment of the present application;

[0056] Fig.18 A schematic diagram of the structure of an operation and maintenance device based on an integrated operation and maintenance portal platform provided in an embodiment of the present application;

[0057] Fig.19 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.

[0059] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".

[0060] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0062] Reference Figure 1 The embodiment of the present application provides an operation and maintenance method based on an integrated operation and maintenance portal platform. The method is applied to the integrated operation and maintenance portal platform. The method may include but is not limited to S100 to S130, which are as follows:

[0063] S100: Collecting data from different types of infrastructure equipment through a data collection module.

[0064] Further, S100 may include the following steps:

[0065] The data collection module collects device hardware configuration, network configuration, health status, energy consumption data, performance and capacity data of server devices, network devices and storage devices according to server out-of-band management, network device SNMP protocol, storage device API, SSH, SMIS and SNMP as the collected data.

[0066] S110: Outputting alarm information according to the collected data.

[0067] Further, S110 may include the following steps:

[0068] Compare thresholds corresponding to device hardware configuration, network configuration, health status, energy consumption data, performance and capacity data of server devices, network devices and storage devices to obtain comparison results;

[0069] The alarm information is output according to the comparison result; wherein the alarm information includes an alarm description, alarm location information and suggested steps to resolve the alarm.

[0070] S120: Locate the cause and position of the fault according to the collected data.

[0071] Further, S120 may include the following steps:

[0072] The causes of the faults located according to the collected data include hardware failure, software failure and network failure;

[0073] Based on the topological relationship between the network device port and other devices and the topological relationship between the fiber optic switch and the storage device and the host, and according to the cause of the fault, the fault location and the impact range are located.

[0074] S130: Centrally manage at least one of the configuration information, firmware version, and storage space of each of the infrastructure devices.

[0075] As an optional implementation, the step of uniformly managing the configuration information of each of the infrastructure devices includes the following steps:

[0076] Centrally manage the hardware configuration, network settings, access control and centralized reading and modification of settings of servers in each of the infrastructure devices;

[0077] Unified management of network devices and security devices in each of the infrastructure devices so that the network devices and the security devices regularly back up corresponding device data;

[0078] The network devices in each of the infrastructure devices are managed uniformly so that the network devices can recover data from the most adjacent multiple infrastructure devices.

[0079] As an optional implementation, the step of uniformly managing the firmware versions of the infrastructure devices includes the following steps:

[0080] Regularly scan the firmware version of each of the infrastructure devices and compare it with the latest firmware library; when it is detected that the firmware version is lower than the set version number, output upgrade information to the operation and maintenance personnel;

[0081] Batch firmware version upgrade for multiple infrastructure devices of the same model;

[0082] Scan the firmware versions of each of the infrastructure devices to determine the infrastructure devices with firmware risks, and then inform the operation and maintenance personnel in a centralized page information manner or email manner.

[0083] As an optional implementation, the step of uniformly managing the storage space of each of the infrastructure devices includes the following steps:

[0084] Collecting information about storage space of SAN storage, NAS storage, distributed storage, fiber switches, and backup devices from each of the infrastructure devices, and then managing each of the storage spaces in a unified manner;

[0085] Compare the IPOS, throughput and response time of each storage space, and then select a number of storage spaces with the highest utilization rate;

[0086] The capacity usage trend of each of the storage spaces is analyzed, and the capacity exhaustion date is predicted.

[0087] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.

[0088] This embodiment aims to provide an automated operation and maintenance solution for infrastructure equipment based on an integrated operation and maintenance portal platform. This embodiment provides an operation and maintenance method and system for centralized integrated operation and maintenance of various infrastructure equipment in the data center and various types of equipment by integrating multiple technical means such as server out-of-band management, network equipment SNMP protocol, storage device API, SSH, SMIS and SNMP. It also improves operation and maintenance efficiency, reduces operation and maintenance costs, and provides strong support for enterprise informatization construction through centralized management, automated monitoring, intelligent identification and other means.

[0089] This embodiment mainly provides a unified operation and maintenance portal to simplify the communication between operation and maintenance applications and infrastructure. The operation and maintenance applications only need to interact with the integrated operation and maintenance portal platform and do not need to communicate directly with each infrastructure object. This can reduce the workload of operation and maintenance application development and improve the convenience and maintainability of operation and maintenance applications.

[0090] By centrally managing operation and maintenance applications through an integrated operation and maintenance portal, you can optimize various types of operation and maintenance processes, quickly count and display data, and realize automated operation and maintenance in multiple scenarios, thereby reducing production costs and improving production efficiency.

[0091] While optimizing operation and maintenance, this operation and maintenance portal can provide effective data support for users' operation and maintenance decisions. It has strong applicability and replicability and can be widely used.

[0092] As operation and maintenance data and applications continue to increase, users will become more and more dependent on the operation and maintenance portal. The project will have continuity, and will continuously optimize and enrich the operation and maintenance portal to generate continuous benefits.

[0093] The content of this embodiment scheme includes:

[0094] Integrated operation and maintenance portal platform: It can be used to receive, process and analyze data transmitted by the collector, provide user interaction interfaces for major functional modules, and an integrated portal platform for data content display.

[0095] Data acquisition module: mainly responsible for receiving data transmitted by the collector, processing and analyzing device status, triggering alarms, generating reports, and providing user operation interfaces. Compatible with multiple acquisition protocols for different types of equipment.

[0096] Automated monitoring and alerting: By setting thresholds for multiple monitoring indicators and performing comparisons, intelligent anomaly detection and alert management can be achieved.

[0097] Centralized risk warning: Intelligently identify the operating status of equipment components, collect and summarize the key components: disk, memory life and media error number, and present them to users in a visual way. Risk notification is also issued for the risk of component removal.

[0098] Unified configuration management: The automated operation and maintenance system can centrally manage the configuration information of various types of equipment.

[0099] Unified firmware version management: The system regularly scans device firmware versions, network device software versions, and storage device firmware (microcode) versions, and compares them with the firmware baseline version. It also provides batch firmware version upgrades.

[0100] Troubleshooting and automatic tracing: Through intelligent tracing analysis algorithms, the system can quickly locate the cause and location of the fault, determine the source of the fault, and the scope of impact.

[0101] Unified storage management: fully compatible with various brands and models of SAN storage, NAS storage, distributed storage, fiber switches, and backup devices to achieve unified storage management.

[0102] Next, the specific content of this embodiment is described as follows:

[0103] 1. Integrated operation and maintenance portal platform.

[0104] The integrated operation and maintenance portal platform is the integrated management and monitoring platform of this embodiment. It is mainly responsible for unifying the interfaces and protocols of various heterogeneous devices, automatically identifying and adapting devices of different brands, and achieving unified configuration, unified monitoring, unified alarm and statistical analysis. Its core functions and features include:

[0105] Unified monitoring interface: The operation and maintenance platform provides a graphical user interface to display the real-time status of the device, historical data trends, alarm information, etc. Administrators can view the health status of the device through the portal system and respond in a timely manner.

[0106] Intelligent threshold matching and alarm management: The operation and maintenance system compares the collected data with the set threshold rules. When the data exceeds the set threshold, the system automatically triggers an alarm and notifies relevant personnel via email, SMS, pop-up windows, etc.

[0107] Data analysis and fault prediction: The operation and maintenance platform can analyze historically collected data, identify equipment performance bottlenecks, abnormal trends and other information, and predict possible equipment failures in advance, thereby providing a basis for equipment repair and maintenance.

[0108] 2. Data acquisition module.

[0109] This embodiment mainly supports different devices and protocols, including servers, network devices, security devices, storage devices and other heterogeneous devices. By integrating server out-of-band management, network device SNMP protocol, storage device API, SSH, SMIS and SNMP and other technical means, the following information can be collected: device hardware configuration, network configuration, health status, energy consumption data, performance and capacity data and other information. Specific collection technologies include:

[0110] Server equipment: The main collection method is through out-of-band.

[0111] Network equipment, security equipment: mainly based on SNMP protocol collection method;

[0112] Storage devices: including collection protocols and interfaces such as API, SSH, SMIS and SNMP.

[0113] 3. Automated monitoring and alarm.

[0114] This embodiment mainly supports the setting of alarm thresholds and levels for collected indicators, supports 7*24 hours of uninterrupted automatic monitoring, and compares the real-time status of collected indicators to achieve automated anomaly detection and alarm management. Specific functions include:

[0115] Autonomous collection of equipment data: The system uses the above-mentioned technical means through the built-in collection engine to collect comprehensive health data, configuration management information, capacity and performance indicators of servers, network devices and storage devices;

[0116] Real-time monitoring: Through automated operation and maintenance tools, the operating status of server equipment, network equipment, storage and other equipment is monitored in real time 24 hours a day, 7 days a week, including hardware health status, configuration information, performance indicators, capacity usage, etc.

[0117] Smart alarm: When monitoring of storage devices shows abnormalities or potential risks, the automated operation and maintenance system will immediately trigger an alarm and notify the operation and maintenance personnel to handle it in time. The alarm information should include a detailed description of the abnormality, location information, and recommended solution steps.

[0118] 4. Troubleshooting and automatic tracing.

[0119] This embodiment mainly supports the use of intelligent source tracing analysis algorithms, so that the system can quickly locate the cause and location of the fault, determine the source of the fault, and the scope of impact. Specific functions include:

[0120] Fault detection: The automated operation and maintenance system can automatically detect faults of various types of equipment, including hardware faults, software faults, and network faults. Through intelligent source tracing analysis algorithms, the system can quickly locate the cause and location of the fault and determine the source of the fault.

[0121] Locating the fault point: Through the intelligent tracing analysis algorithm, the system can quickly locate the cause and location of the fault, determine the source of the fault, and the scope of impact.

[0122] Intelligent tracing algorithm: The implementation principle of the algorithm is mainly based on the automatic acquisition of the topological relationship between the network device port and other devices, as well as the automatic identification of the topological relationship between the optical switch and storage devices and the host.

[0123] 5. Unified configuration management.

[0124] This embodiment mainly supports unified management of configuration information of various types of equipment, including server configuration management, network equipment, configuration backup and recovery of security equipment, etc. Specific functions include:

[0125] Server equipment: centralized reading and modification of server hardware configuration, network settings, and access control settings.

[0126] Configuration backup: Network devices, security devices, regularly back up the data of network devices to prevent data loss or damage.

[0127] Configuration recovery: Once network devices need to restore data from the most recent backup, it can be done in batches to minimize data loss and recovery time.

[0128] 6. Unified firmware version management.

[0129] This embodiment mainly supports periodic scanning of device firmware versions and comparison with the firmware baseline version. It also provides batch firmware version upgrades. Specific functions include:

[0130] Firmware version management: The automated operation and maintenance system will regularly scan the server device firmware version and compare it with the latest firmware library. If the firmware version is found to be outdated, the system will prompt the operation and maintenance personnel to upgrade it.

[0131] Batch upgrade: Support batch firmware upgrade for multiple devices of the same model to ensure that all devices are running on the latest firmware version. During the upgrade process, the system will conduct sufficient testing and verification to ensure the stability and compatibility of the new firmware.

[0132] Risk warning: Supports scanning device firmware versions, screening out devices with firmware risks, and notifying users through centralized page information, emails, etc.

[0133] 7. Unified storage management.

[0134] This embodiment mainly supports different brands and models of SAN storage, NAS storage, distributed storage, fiber switches, and backup devices to achieve unified storage management. Specific functions include:

[0135] Storage resource management: Automatically collect the resource status of storage devices, hosts, optical switches and other devices, and centrally manage various resources such as devices, storage pools, volumes, and Luns.

[0136] Storage performance management: real-time statistics, comparison of IPOS, throughput, response time of various storage resources, screening out hot resources;

[0137] Storage device capacity management: Trend analysis of capacity usage and prediction of capacity exhaustion dates.

[0138] It should be noted that:

[0139] a. Heterogeneous brand compatibility: Ensure that the collector supports different devices and protocols, including: server out-of-band, SNMP protocol of network devices, and API, SSH, SMIS and SNMP of storage devices on the same platform.

[0140] b. Wide applicability: The integrated automatic operation and maintenance platform can manage devices across regions and of different brands, breaking down brand barriers and achieving unified management of heterogeneous devices. This greatly reduces the management complexity caused by different device brands.

[0141] c. Integrated operation and maintenance: By connecting to different interfaces and protocols of various types of equipment, the platform automatically identifies and adapts to equipment of different brands, achieving unified configuration, unified monitoring and unified alarm, thereby simplifying the management process. The integrated operation and maintenance portal platform provides a unified entrance for infrastructure operation and maintenance management, enabling various equipment managers to perform centralized management on one platform, improving management efficiency.

[0142] d. Autonomous and controllable: The core technologies and deployment environment adopted by the integrated operation and maintenance platform can all be realized in the form of trust-based innovation. The core technologies and source codes of trust-based innovation products are in the hands of domestic enterprises and are not affected by external technology blockades and sanctions.

[0143] e. System security: Security is fully considered in the design of trusted computing products, and a variety of security technologies and means are adopted, such as encryption, access control, vulnerability scanning, etc., to ensure that the system is protected from attacks and damage.

[0144] f. The trusted computing products have good compatibility with mainstream domestic software and hardware systems and can be seamlessly integrated into existing systems, reducing the cost and risk of system upgrades and migrations.

[0145] Beneficial effects:

[0146] This embodiment provides an operation and maintenance solution for an integrated automatic operation and maintenance portal platform for data center infrastructure, which has significant advantages in terms of compatibility with heterogeneous equipment, multi-mode collection of data from various types of equipment, integrated operation and maintenance management, unified display management, and comprehensive information innovation. The above advantages enable the integrated operation and maintenance portal platform of this embodiment to better solve the pain points of existing users in infrastructure equipment management, improve operation and maintenance efficiency, reduce operation and maintenance costs, and enhance the stability and security of infrastructure.

[0147] Advantages:

[0148] Verification results: The system of this embodiment has significant advantages in compatibility with heterogeneous devices, multi-mode data collection, integrated operation and maintenance management, unified display management, and information technology deployment solutions.

[0149] Evaluation: The advantages are fully described, reflecting the unique value of the system in this embodiment.

[0150] Next, the solution of the embodiment of the present application will be described in more specific implementation manner.

[0151] Embodiment 1: System structure of automated operation and maintenance of infrastructure of integrated operation and maintenance portal.

[0152] See also Figure 2 , Figure 2 : is a schematic diagram of the system structure of the infrastructure automated operation and maintenance of the integrated operation and maintenance portal platform provided in this embodiment, such as Figure 2 As shown, the system structure of this embodiment adopts a layered architecture design, which mainly includes a data object layer, a collection layer, a processing layer, and a presentation layer. Each layer communicates through a standardized interface to ensure the scalability and flexibility of the system. The details are as follows:

[0153] Object layer: The monitored objects include: servers, hyper-convergence, minicomputers, storage, fiber switches, network equipment, security equipment, etc., taking into account all kinds of mainstream manufacturers' models on the market;

[0154] Collection layer: collects monitoring information of IT hardware equipment to provide data foundation for the upper layer of the platform. The collected information includes: alarm data, configuration data, and equipment health status data.

[0155] Business layer: processes the monitoring data obtained by the collection layer to achieve comprehensive and real-time monitoring of IT hardware equipment. 1) The monitoring engine includes an alarm monitoring module, a performance capacity monitoring module, and a configuration management module: the monitoring module uses various monitoring technologies to monitor hardware alarm information, component warning information, and performance capacity information; the configuration information management module collects and manages hardware configuration information; 2) The system function module implements alarm classification and grading, threshold management, alarm push, data reporting, user management and other functions to provide calls and support for platform operation. 3) Control management: realize the reading of firmware versions of servers and network devices, and batch update of firmware versions; server BMC settings, BiOS settings, and user password setting management;

[0156] Display layer: Provides a unified centralized display portal, and realizes the unified portal presentation of the platform through various methods such as WEB interface and large screen.

[0157] Example 2: Centralized monitoring and alarming.

[0158] See also Figure 3 , Figure 3 FIG. 1 is a flow chart of an embodiment of centralized monitoring and alarming provided by this embodiment. Figure 3 As shown, the device includes multiple collection protocols based on various types of equipment. The system uses a built-in collection engine to perform 7*24 hours of uninterrupted real-time monitoring of the operating status of servers, network devices, storage devices and other equipment; when the storage device is monitored to have an abnormality or potential risk, the automated operation and maintenance system will immediately trigger an alarm.

[0159] The process description is shown in Table 1:

[0160]

[0161]

[0162] Table 1

[0163] The following is a step-by-step description of an implementation example of centralized monitoring and alarming.

[0164] Step 1: Automatic scanning and discovery.

[0165] Log in to the integrated operation and maintenance portal platform, enter the IP address segment and device SNMP or SSH parameters, and automatically scan and discover related devices. Figure 4 The input elements are: IP segment, device brand, SNMP parameters (a few in SSH mode)

[0166] Output: device SN, device IP, brand model, device type, etc.

[0167] Step 2: Centralized monitoring.

[0168] Centrally manage the devices added to the monitoring list and display the health data collected by the devices in real time. The devices display the monitored device resources and various indicator information of the devices in the form of a classified list, including: Figure 5 For monitoring list, Figure 6 For a single device monitoring overview, Figure 7 Details of monitoring indicators for a single device.

[0169] Step 3: Centralized alarm.

[0170] Alarm viewing: The integrated operation and maintenance portal provides a unified alarm platform, integrating events and alarms generated by different nodes in different regions, and integrating events and alarms from different message sources. Different types of events and severity are displayed with different colors / icons, and detailed information such as the source, time, and cause of the event are displayed in the same window. Figure 8 For an overview of alarm events, Fig. 9 It is the current alarm event list.

[0171] Alarm triggering: The system provides preset alarm thresholds. Once the threshold conditions are reached, an alarm is automatically generated.

[0172] Alarm processing: The management personnel can adjust the alarm level of the threshold according to the actual operation and maintenance needs, close or confirm the alarm, and confirm that it is a hardware failure, and then further report it for repair;

[0173] Alarm notification: Administrators can set up a variety of alarm notification methods, including: traditional email, SMS, third-party platforms and other alarm methods, to comprehensively and promptly notify relevant users. Fig.10 Alarm notification rules.

[0174] Embodiment 3: Intelligent tracing algorithm.

[0175] This embodiment supports multiple modes of device collection, including SSH, telnet, SNMPv1, SNMPv2 and SNMPv3. SNMPv3 supports a combination of encryption and authentication.

[0176] It supports automatic identification of the device name, manufacturer, type, version and other information of network and security devices and automatic classification; it supports detailed device information collection through SSH and telnet, including device operation status information, port configuration, port status, physical neighbor relationship, mac address table, ARP information, etc.

[0177] The platform supports self-learning to form terminal resource information, conducts refined combing of online terminals in the entire network, automatically identifies the manufacturer and model of the networked terminals through the MAC address prefix library, and quickly generates terminal access topology and locates terminal access locations through terminal resource information.

[0178] The platform supports self-learning to form ARP table resource information, automatically sort out the ARP resource table on each gateway, and analyze the static and dynamic ARP binding status of the entire network. Fig.11 ARP information list.

[0179] Through the resource list of the ARP table, the corresponding terminal device is addressed from the MAC address library, and the connection relationship between the terminal device and the network device port is automatically established. This is the core element of realizing the intelligent tracing algorithm.

[0180] Example 4: Automatic operation and maintenance of network topology.

[0181] See also Fig.12 , Fig.12 It is a flowchart of an embodiment of automatic operation and maintenance of network topology provided by this embodiment.

[0182] The process description is shown in Table 2:

[0183]

[0184]

[0185]

[0186] Table 2

[0187] The following are the steps for network topology operation and maintenance management:

[0188] Step 1: Automatically generate a topology view.

[0189] It automatically generates multi-level topologies, including WAN topology and internal LAN topology, by combining LLDP, SNMP, MAC, ARP and other protocols. It also monitors important ports and lines for device interconnection and provides intuitive visual alarms. It discovers topologies in a variety of scenarios, including but not limited to device IP discovery, discovery through IP address segment scanning, and deep detection of unknown IP addresses.

[0190] Create a topology view settings page, see Fig.13 Create a network topology.

[0191] Step 2: Port traffic monitoring.

[0192] The platform can monitor the traffic associated with each node and port on the topology map. Click on the device to display the current operating status. When a device or port fails, it can issue a real-time alarm and visualize it on the topology map. Draw a monitoring topology for specific applications to achieve full-link monitoring capabilities for specific application access, including monitoring of parameters such as traffic, ping packets, delay, and packet loss.

[0193] Step 3: Automatic identification + manual creation.

[0194] The automatically generated topology map supports manual creation and maintenance, manual configuration of the physical connection relationship between two devices, definition of different areas, display of device names and port information in the topology, and association of device asset locations based on custom device naming or other key information.

[0195] Step 4: Full-link monitoring capability.

[0196] The system updates the node status, connection traffic information, and topology relationship changes in the network topology in real time according to the collection frequency (generally every 5 minutes, which can be customized). It provides intuitive and visual automatic operation and maintenance data for device nodes, port stages, and traffic indicators. Once an indicator alarm is generated, the fault level and fault point are clearly displayed on the topology map, and the alarm information can be directly viewed and processed.

[0197] Network topology real-time monitoring diagram, see Fig.14 The real-time status of the network topology.

[0198] Embodiment 5: Network configuration management.

[0199] See also Fig.15 , Fig.15 This is a flow chart of an embodiment of network configuration backup and distribution provided by this embodiment. The device includes backing up the configuration of network equipment according to the scheduled backup task and establishing a configuration baseline template for a certain model of network equipment.

[0200] The process description is shown in Table 3:

[0201]

[0202] Table 3

[0203] The following are the steps for configuring and managing network devices:

[0204] Step 1: Configure backup.

[0205] Backup operation: The integrated operation and maintenance portal provides automatic or manual configuration backup operations. After the operation is completed, the inspection results can be viewed and exported. For the non-compliant options, rectification can be quickly and automatically completed to complete the rectification closed loop.

[0206] Step 2: Configure template maintenance.

[0207] It supports setting configuration templates for network devices of different manufacturers and models, and can customize baseline standards, define multiple inspection items, and combine multiple inspection items to create configuration templates. The inspection templates are compatible with mainstream network devices. Fig.16 Configuration template for .

[0208] Step 3: Backup task management.

[0209] Backup task management: freely configure backup tasks by organization and device type; achieve the purpose of batch scheduled automatic backup through scheduled tasks and real-time configuration check tasks. Fig.17 backup tasks.

[0210] Step 4: Baseline check.

[0211] It supports checking the current configuration information of network devices, and matching them one by one with the configuration baseline files of devices of the same model, performing baseline configuration checks on designated devices, filtering out configuration files that do not meet the baseline, and marking non-compliant items.

[0212] Reference Fig.18 The embodiment of the present application further provides an operation and maintenance device based on the integrated operation and maintenance portal platform, which can implement the above-mentioned operation and maintenance method based on the integrated operation and maintenance portal platform, and the device includes:

[0213] A data collection unit, used to collect data from different types of infrastructure equipment through a data collection module;

[0214] An alarm output unit, used to output alarm information according to the collected data;

[0215] A fault locating unit, used to locate the cause and position of the fault according to the collected data;

[0216] The unified management unit is used to uniformly manage at least one of the configuration information, firmware version and storage space of each of the infrastructure devices.

[0217] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0218] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned operation and maintenance method based on the integrated operation and maintenance portal platform when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0219] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0220] See also Fig.19 , Fig.19 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0221] The processor 1901 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0222] The memory 1902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1902 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1902, and the processor 1901 calls and executes the operation and maintenance method based on the integrated operation and maintenance portal platform of the embodiment of this application;

[0223] Input / output interface 1903, used to implement information input and output;

[0224] Communication interface 1904, used to realize communication interaction between the device and other devices, which can be realized through wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0225] A bus 1905 that transmits information between the various components of the device (e.g., the processor 1901, the memory 1902, the input / output interface 1903, and the communication interface 1904);

[0226] The processor 1901 , the memory 1902 , the input / output interface 1903 and the communication interface 1904 are connected to each other in communication within the device via the bus 1905 .

[0227] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned operation and maintenance method based on the integrated operation and maintenance portal platform.

[0228] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0229] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0230] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0231] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0232] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0233] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0234] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0235] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0236] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0237] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0238] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0239] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0240] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. The operation and maintenance method based on the integrated operation and maintenance portal platform is characterized by: The method is applied to the integrated operation and maintenance portal platform, and the method comprises the following steps: The data acquisition module collects data from different types of infrastructure equipment; Outputting alarm information according to the collected data; Locate the cause and location of the fault according to the collected data; Centrally manage at least one of the configuration information, firmware version and storage space of each of the infrastructure devices.

2. The operation and maintenance method based on the integrated operation and maintenance portal platform according to claim 1 is characterized in that: The method of collecting data from different types of infrastructure equipment by using a data collection module includes the following steps: The data collection module collects device hardware configuration, network configuration, health status, energy consumption data, performance and capacity data of server devices, network devices and storage devices according to server out-of-band management, network device SNMP protocol, storage device API, SSH, SMIS and SNMP as the collected data.

3. The operation and maintenance method based on the integrated operation and maintenance portal platform according to claim 1, characterized in that: The step of outputting the alarm information according to the collected data comprises the following steps: Compare thresholds corresponding to device hardware configuration, network configuration, health status, energy consumption data, performance and capacity data of server devices, network devices and storage devices to obtain comparison results; The alarm information is output according to the comparison result; wherein the alarm information includes an alarm description, alarm location information and suggested steps to resolve the alarm.

4. The operation and maintenance method based on the integrated operation and maintenance portal platform according to claim 1, characterized in that: The method of locating the cause and position of the fault according to the collected data comprises the following steps: The causes of the faults located according to the collected data include hardware failure, software failure and network failure; Based on the topological relationship between the network device port and other devices and the topological relationship between the fiber optic switch and the storage device and the host, and according to the cause of the fault, the fault location and the impact range are located.

5. The operation and maintenance method based on the integrated operation and maintenance portal platform according to claim 1, characterized in that: The step of uniformly managing the configuration information of each of the infrastructure devices comprises the following steps: Centrally manage the hardware configuration, network settings, access control and centralized reading and modification of settings of servers in each of the infrastructure devices; Unified management of network devices and security devices in each of the infrastructure devices so that the network devices and the security devices regularly back up corresponding device data; The network devices in each of the infrastructure devices are managed uniformly so that the network devices can recover data from the most adjacent multiple infrastructure devices.

6. The operation and maintenance method based on the integrated operation and maintenance portal platform according to claim 1, characterized in that: The steps of uniformly managing the firmware versions of the infrastructure devices include the following steps: Regularly scan the firmware version of each of the infrastructure devices and compare it with the latest firmware library; when it is detected that the firmware version is lower than the set version number, output upgrade information to the operation and maintenance personnel; Batch firmware version upgrade for multiple infrastructure devices of the same model; Scan the firmware versions of each of the infrastructure devices to determine the infrastructure devices with firmware risks, and then inform the operation and maintenance personnel in a centralized page information manner or email manner.

7. The operation and maintenance method based on the integrated operation and maintenance portal platform according to claim 1, characterized in that: The step of uniformly managing the storage space of each of the infrastructure devices comprises the following steps: Collecting information about storage space of SAN storage, NAS storage, distributed storage, fiber switches, and backup devices from each of the infrastructure devices, and then managing each of the storage spaces in a unified manner; Compare the IPOS, throughput and response time of each of the storage spaces, and then select a number of the storage spaces with the highest utilization rate; The capacity usage trend of each of the storage spaces is analyzed, and the capacity exhaustion date is predicted.

8. The operation and maintenance device based on the integrated operation and maintenance portal platform is characterized by: The device comprises: A data collection unit, used to collect data from different types of infrastructure equipment through a data collection module; An alarm output unit, used to output alarm information according to the collected data; A fault locating unit, used to locate the cause and position of the fault according to the collected data; The unified management unit is used to uniformly manage at least one of the configuration information, firmware version and storage space of each of the infrastructure devices.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.