Implementation method, system and device of integrated monitoring operation and maintenance platform and storage medium
Patent Information
- Application Number
- CN202610960210.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本发明要解决的技术问题是:在现有技术中,不同信息技术设施和相关业务系统通常遵循不同通信协议并使用相互独立的监控系统,导致数据分散且格式不一,形成信息孤岛,缺乏全局可视化面板;各系统配置独立告警机制需人工跨系统排查故障,导致故障发现和修复时间过长;日常运维缺乏跨系统的自动协同与追踪机制,管理效率与风险控制水平较低
[0026]本公开采用数据采集整合、分析告警、可视化与运维管理的分层解耦架构,打破了异构设施间的信息壁垒,实现了多协议运行数据的统一清洗与标准化存储;借助监控大屏使全栈设施的运行态势透明化,支持直观的全局掌控;利用告警压缩与抑制机制消减了底层设备级联故障引发的虚假告警;结合问题待办的自动触发与风险追踪联动机制,缩短了平均故障发现时间和平均修复时间,降低了团队跨系统维护的运营成本,提升了信息技术基础设施的管理效率与抗风险能力。
Smart Images

Figure CN122838142A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semiconductor technology, and in particular to a method, system, device, and storage medium for implementing an integrated monitoring and maintenance platform. Background Technology
[0002] Currently, IT equipment, data center environmental monitoring equipment, and other related systems follow different protocols and are managed using their respective monitoring systems. Data is scattered across these systems, stored in different formats, and uses separate visualization dashboards. Each system is configured with alarms (SMS, email, app), requiring users to log into each system to troubleshoot when a fault occurs. Daily maintenance and operational data statistics require querying each system.
[0003] In conclusion, there is an urgent need for an integrated platform with full-stack monitoring capabilities. Summary of the Invention
[0004] The technical problem this invention aims to solve is that, in the prior art, different information technology facilities and related business systems typically follow different communication protocols and use independent monitoring systems, resulting in scattered and inconsistent data formats, forming information silos and lacking a global visualization panel; each system has an independent alarm mechanism, requiring manual cross-system troubleshooting, which leads to excessively long fault discovery and repair times; daily operation and maintenance lacks cross-system automatic collaboration and tracking mechanisms, resulting in low management efficiency and risk control levels.
[0005] To address the aforementioned technical problems, this invention provides a method for implementing an integrated monitoring and maintenance platform, comprising:
[0006] Step 1: Collect status and performance data of heterogeneous facilities;
[0007] Step 2: Clean and standardize the status data and performance data to obtain standard operating data, and store the standard operating data in the database;
[0008] Step 3: Analyze the standard operating data, identify anomalies, and trigger a unified alarm;
[0009] Step 4: Generate maintenance tasks based on the anomalies, and visualize the standard operating data and the anomalies.
[0010] Preferably, in step one, the heterogeneous facilities include at least one of hardware devices, data center environment devices, operating systems, network devices, applications, databases, and middleware.
[0011] Preferably, in step one, the collection of status data and performance data of heterogeneous facilities includes: collection using a multi-protocol adaptation engine; the multi-protocol adaptation engine supports at least one of the following collection methods: Simple Network Management Protocol, Serial Communication Protocol, Intelligent Platform Management Interface, Windows Management Specification, Agent Program, and Application Programming Interface.
[0012] Preferably, in step two, the cleaning and standardization of the status data and the performance data includes: completing data cleaning and integration through extraction, transformation, and loading processes.
[0013] Preferably, in step three, the discovery of an anomaly and triggering of a unified alarm includes performing an alarm management operation, which includes alarm compression and alarm suppression; wherein, the alarm compression is configured to merge multiple identical alarms; and the alarm suppression is configured to suppress child device alarms when the parent device fails.
[0014] Preferably, in step three, the alarm management operation further includes a threshold setting operation, wherein the threshold setting operation is configured to set trigger conditions based on continuous monitoring data of a preset period.
[0015] Preferably, in step three, triggering a unified alarm further includes performing an alarm routing operation, wherein the alarm routing operation is configured to push the unified alarm to the responsible person via at least one of email, SMS and voice call media based on business importance, time period and alarm level policy.
[0016] Preferably, in step four, generating maintenance tasks based on the anomaly includes: automatically triggering the generation of a problem-handling process and pushing the problem-handling process to the primary responsible person; in response to the transfer operation of the primary responsible person, notifying the problem handling personnel of the problem-handling process; and in response to the triggering operation, generating risk items and tracking and associating the risk items.
[0017] Preferably, in step four, the visualization of the standard operating data and the anomalies includes: displaying at least one of the following through a monitoring dashboard and a monitoring screen: a topology map, a dynamic panel, a trend curve, and a health status card.
[0018] This invention also provides an integrated monitoring and maintenance system, comprising:
[0019] The data acquisition module is configured to collect status and performance data from heterogeneous facilities;
[0020] The data integration module is configured to clean and standardize the status data and the performance data to obtain standard operating data, and store the standard operating data in the database.
[0021] The analysis and alarm module is configured to analyze the standard operating data, detect anomalies, and trigger a unified alarm.
[0022] The operation and maintenance management module is configured to generate operation and maintenance tasks based on the anomalies, and to visualize the standard operation data and the anomalies.
[0023] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method as described in any of the preceding claims.
[0024] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method as described in any of the preceding claims.
[0025] As described above, the implementation method, system, equipment, and storage medium of the integrated monitoring and maintenance platform of the present invention have the following beneficial effects:
[0026] This disclosure adopts a layered and decoupled architecture for data collection and integration, analysis and alarm, visualization, and operation and maintenance management. It breaks down information barriers between heterogeneous facilities and achieves unified cleaning and standardized storage of multi-protocol operation data. It makes the operation status of the entire stack of facilities transparent through a monitoring dashboard, supporting intuitive global control. It reduces false alarms caused by cascading failures of underlying devices by using alarm compression and suppression mechanisms. Combined with the automatic triggering of problem pending and risk tracking linkage mechanisms, it shortens the average fault discovery time and average repair time, reduces the operating costs of cross-system maintenance by the team, and improves the management efficiency and risk resistance of information technology infrastructure. Attached Figure Description
[0027] Figure 1 The diagram shows a flowchart illustrating the implementation method of the integrated monitoring and maintenance platform of the present invention.
[0028] Figure 2 The diagram shown is a schematic representation of the integrated monitoring and maintenance system structure of the present invention.
[0029] Figure 3 The diagram shown is a schematic representation of the hardware structure of the electronic device of the present invention. Detailed Implementation
[0030] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0031] This disclosure provides a specific implementation of an integrated monitoring and maintenance platform. This platform is typically executed by electronic devices or distributed server clusters equipped with computing and processing capabilities. In practical applications, the platform adopts a layered and decoupled architecture for data acquisition and integration, data analysis and alerting, data visualization, and supporting maintenance management systems. Each core module is physically or logically isolated to facilitate independent maintenance and expansion at different levels.
[0032] refer to Figure 1 An implementation method for an integrated monitoring and maintenance platform includes:
[0033] Step 1: Collect status and performance data from heterogeneous facilities. This step corresponds to the unified data acquisition layer, which serves as the platform's data source and is responsible for extracting real-time status and performance data from various heterogeneous devices. Acquiring multi-source heterogeneous data breaks down the communication protocol limitations between independent systems, enabling centralized aggregation of underlying information technology resources, improving the completeness and real-time performance of data acquisition, and reducing the time loss caused by scattered cross-system queries.
[0034] In some embodiments, in step one, the heterogeneous facilities include at least one of hardware devices, data center environment equipment, operating systems, network devices, applications, databases, and middleware. Hardware devices encompass physical server hosts, storage arrays, or network security gateways. Data center environment equipment includes power and environmental facilities such as uninterruptible power supplies, temperature and humidity sensors, or air conditioning cooling units. The operating system is compatible with various open-source or commercial system versions. Network devices include core communication nodes such as switches, routers, or firewalls. Applications, databases, and middleware constitute the software foundation for upper-layer business operations. By being compatible with different levels of hardware and software objects, the platform can achieve full-stack, three-dimensional monitoring.
[0035] In some embodiments, step one, collecting status and performance data from heterogeneous facilities, includes: using a multi-protocol adaptation engine for collection; the multi-protocol adaptation engine supports at least one of the following collection methods: Simple Network Management Protocol, Serial Communication Protocol, Intelligent Platform Management Interface, Windows Management Specification, Agent Program, and Application Programming Interface. Developing or integrating a multi-protocol-adaptive data collection engine ensures rapid adaptation to various new and old devices and systems through its open design, providing high flexibility and scalability. When new models of equipment or special systems are introduced within an enterprise, data access can be completed without reconstructing the overall architecture, reducing secondary development and maintenance costs. Employing diverse collection methods such as agent programs or application programming interfaces allows for flexible selection of resource-efficient collection methods based on the performance sensitivity of the monitored object, ensuring the stable operation of the monitored business system.
[0036] Step two involves cleaning and standardizing the status and performance data to obtain standard operating data, which is then stored in the database. This step corresponds to the high-performance data integration and storage layer, where the collected raw data is cleaned, organized, standardized, and stored to provide stable and efficient data services for upper-layer applications. Cleaning includes removing garbled characters, filtering duplicate messages, and correcting invalid data fragments. Standardization involves converting unstructured or semi-structured data from various heterogeneous devices into a structured format that is uniformly recognized within the platform. This standardization process eliminates format differences in cross-system data interaction, preventing interference with subsequent anomaly detection.
[0037] In some embodiments, step two, cleaning and standardizing status and performance data, includes: data cleaning and integration through extraction, transformation, and loading processes. A high-performance database is used as the core storage to handle the high-concurrency writing and fast querying requirements of massive amounts of monitoring data. Complex data association logic can be implemented through code development. Through the extraction, transformation, and loading processes, the system improves the throughput and processing efficiency of the data storage layer, effectively coping with the pressure of massive data influx in a factory-level environment.
[0038] Step 3: Analyze standard operating data to identify anomalies and trigger unified alarms. This step corresponds to the intelligent analysis and alarm engine, which performs real-time analysis of the processed data and detects anomalies based on preset rules or intelligent algorithms, triggering alarms. Detecting anomalies at the data source transforms traditional passive manual inspections into proactive system-level early warnings, significantly reducing the average fault detection time and average repair time.
[0039] In some embodiments, step three, detecting an anomaly and triggering a unified alarm, includes performing alarm management operations, which include alarm compression and alarm suppression. Alarm compression is configured to merge multiple identical alarms; alarm suppression is configured to suppress child device alarms when the parent device fails. Alarm management supports flexible dependency configuration, and alarm compression effectively reduces the attention drain on maintenance personnel from repetitive and redundant information. In the event of a core network outage or a cabinet power failure cascading fault, alarm suppression can mask false alarms from Haiquanzi devices caused by the parent node's offline state, preventing a systemic alarm storm and helping on-duty personnel quickly locate the source fault node.
[0040] In some embodiments, step three of the alarm management operation further includes a threshold setting operation, which is configured to set trigger conditions based on continuous monitoring data over a preset period. The preset period can be one week, one month, or other time periods suitable for reflecting the operating patterns of the equipment. By combining long-term continuous monitoring trends with dynamic threshold setting, compared to simply comparing fixed absolute values, it can more accurately reflect the true health baseline of the equipment operation, filter out occasional indicator fluctuations, and thus reduce the false alarm rate of the monitoring system.
[0041] In some embodiments, step three, triggering a unified alarm, further includes performing an alarm routing operation. This alarm routing operation is configured to push the unified alarm to the responsible party via at least one of email, SMS, and voice call media, based on business importance, time period, and alarm level policy. This flexible alarm routing mechanism with multiple channels ensures that alarms are accurately pushed to the relevant responsible parties, avoiding missed fault reports or response delays caused by a single communication channel being blocked.
[0042] Step 4: Generate maintenance tasks based on anomalies and visualize standard operational data and anomalies. This step corresponds to a unified visualization and interaction layer and a supporting maintenance management system, providing users with intuitive data display and operation interface. The maintenance team no longer needs to maintain multiple monitoring systems, significantly reducing learning and management costs.
[0043] In some embodiments, step four, generating maintenance tasks based on anomalies, includes: automatically triggering the generation of a problem-handling process and pushing it to the primary responsible person; responding to the primary responsible person's transfer operation, notifying the problem-handling personnel of the problem-handling process; and responding to the trigger operation to generate risk items and track and associate them. Utilizing modern web front-end technology combined with a platform database, the corresponding maintenance management system can be built and flexibly expanded according to management requirements. The automatic generation of problem tasks ensures that the problem handling process is recorded and traceable. The design of triggering the generation of risk items makes subsequent risk tracking more convenient and improves the overall resilience of the enterprise's IT environment.
[0044] In some embodiments, step four, visualizing standard operating data and anomalies, includes displaying at least one of the following: topology diagrams, dynamic panels, trend graphs, and health status cards through monitoring dashboards and large monitoring screens. Customizable monitoring dashboards are built using modern web front-end technologies to display overall situation and detailed indicators. Large monitoring screen mode is supported to meet the needs of leadership decision-making and on-duty monitoring scenarios.
[0045] refer to Figure 2 An integrated monitoring and maintenance system includes:
[0046] The data acquisition module is configured to collect status and performance data from heterogeneous facilities;
[0047] The data integration module is configured to clean and standardize status data and performance data to obtain standard operating data, and store the standard operating data in the database.
[0048] The analysis and alarm module is configured to analyze standard operating data, detect anomalies, and trigger unified alarms.
[0049] The operation and maintenance management module is configured to generate operation and maintenance tasks based on anomalies and to visualize standard operation data and anomalies.
[0050] The various modules in the system are manifested as purpose-specific integrated circuit hardware, or as software program code snippets or container microservices executed by a general-purpose computing core. This decoupled system architecture reduces code coupling between functional domains, ensuring that subsequent independent upgrades and performance scaling of individual data acquisition nodes or analysis modules will not negatively interfere with the normal operation of other business modules, thus improving the system's iteration efficiency.
[0051] refer to Figure 3 This disclosure provides an electronic device, which may include: a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504, wherein the processor 1501, the communication interface 1502, and the memory 1503 communicate with each other through the communication bus 1504.
[0052] Memory 1503 is used to store computer programs;
[0053] When the processor 1501 executes the computer program stored in the memory 1503, it implements the steps of the above-mentioned integrated monitoring and maintenance platform implementation method.
[0054] The communication bus mentioned in electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.
[0055] For ease of illustration, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between electronic devices and other devices. Memory may include random access memory (RAM) or non-volatile memory, such as at least one disk drive.
[0056] Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor. The aforementioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0057] This disclosure also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the implementation method of any of the above-described integrated monitoring and maintenance platforms.
[0058] This disclosure also provides a computer program product containing instructions that, when run on a computer, causes the computer to execute any of the above-described integrated monitoring and maintenance platform implementation methods.
[0059] The above implementations can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, they can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions corresponding to this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive).
[0060] It should be noted that the illustrations provided in this embodiment are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0061] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A method for implementing an integrated monitoring and maintenance platform, characterized in that, include: Step 1: Collect status and performance data of heterogeneous facilities; Step 2: Clean and standardize the status data and performance data to obtain standard operating data, and store the standard operating data in the database; Step 3: Analyze the standard operating data, identify anomalies, and trigger a unified alarm; Step 4: Generate maintenance tasks based on the anomalies, and visualize the standard operating data and the anomalies.
2. The implementation method of the integrated monitoring and maintenance platform according to claim 1, characterized in that: In step one, the heterogeneous facilities include at least one of hardware devices, data center environment devices, operating systems, network devices, applications, databases, and middleware.
3. The implementation method of the integrated monitoring and maintenance platform according to claim 1, characterized in that: In step one, the collection of status and performance data of heterogeneous facilities includes: collection using a multi-protocol adaptation engine; the multi-protocol adaptation engine supports at least one of the following collection methods: Simple Network Management Protocol, Serial Communication Protocol, Intelligent Platform Management Interface, Windows Management Specification, Agent Program, and Application Programming Interface.
4. The implementation method of the integrated monitoring and maintenance platform according to claim 1, characterized in that: In step two, the cleaning and standardization of the status data and the performance data includes: completing data cleaning and integration through extraction, transformation, and loading processes.
5. The implementation method of the integrated monitoring and maintenance platform according to claim 1, characterized in that: In step three, the discovery of an anomaly and triggering of a unified alarm includes performing alarm management operations, which include alarm compression and alarm suppression; wherein, alarm compression is configured to merge multiple identical alarms; and alarm suppression is configured to suppress child device alarms when the parent device fails.
6. The implementation method of the integrated monitoring and maintenance platform according to claim 5, characterized in that: In step three, the alarm management operation also includes a threshold setting operation, which is configured to set trigger conditions based on continuous monitoring data of a preset period.
7. The implementation method of the integrated monitoring and maintenance platform according to claim 1, characterized in that: In step three, triggering a unified alarm also includes performing an alarm routing operation, wherein the alarm routing operation is configured to push the unified alarm to the responsible person via at least one of email, SMS and voice call media based on business importance, time period and alarm level policy.
8. The method for implementing the integrated monitoring and maintenance platform according to claim 1, characterized in that: In step four, generating maintenance tasks based on the anomaly includes: automatically triggering the generation of a problem task process and pushing the problem task process to the primary responsible person; in response to the transfer operation of the primary responsible person, notifying the problem handling personnel of the problem task process; and in response to the trigger operation, generating risk items and tracking and associating the risk items.
9. The implementation method of the integrated monitoring and maintenance platform according to claim 1, characterized in that: In step four, the visualization of the standard operating data and the anomalies includes: displaying at least one of the following through a monitoring dashboard and a monitoring screen: a topology map, a dynamic panel, a trend curve, and a health status card.
10. An integrated monitoring and maintenance system, characterized in that, include: The data acquisition module is configured to collect status and performance data from heterogeneous facilities; The data integration module is configured to clean and standardize the status data and the performance data to obtain standard operating data, and store the standard operating data in the database. The alarm analysis module is configured to analyze the standard operating data, detect anomalies, and trigger unified alarms. The operation and maintenance management module is configured to generate operation and maintenance tasks based on the anomalies, and to visualize the standard operation data and the anomalies.
11. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 9.