A distributed operation and maintenance management and predictive maintenance system and method for multi-intelligent gateway clusters
Patent Information
- Application Number
- CN202611040306.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
AI Technical Summary
各网关独立运行,缺乏统一的可观测性平台,无法实时掌握集群整体健康状态
[0037]全局可视,主动运维:提供了一种对多网关集群的统一实时监控,运维人员可在一个面板查看所有网关健康度、设备分布、预测性告警,从被动响应转向主动运维。
Smart Images

Figure CN122845453A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of edge computing and distributed system operation and maintenance technology for industrial Internet of Things (IoT), specifically a distributed operation and maintenance management and predictive maintenance system and method for multi-intelligent gateway clusters. Background Technology
[0002] As the scale of the Industrial Internet of Things (IIoT) expands, a single smart gateway often cannot cover all field devices, typically requiring the deployment of multiple gateway nodes (such as edge gateways and protocol conversion gateways). These gateways are distributed across different geographical locations and carry different services, presenting existing operation and maintenance solutions with the following challenges:
[0003] First, there are monitoring blind spots. Each gateway operates independently, lacking a unified observability platform, making it impossible to monitor the overall health of the cluster in real time. Maintenance personnel need to log into each gateway individually to check its status, which is inefficient.
[0004] Second, fault response is delayed. Gateway faults (such as process crashes, network interruptions, and memory leaks) are usually only discovered after users report them, lacking proactive early warning and self-healing capabilities. The time between the occurrence of a fault and its discovery often lasts for hours or even days.
[0005] Third, there is uneven resource utilization. Some gateways are overloaded (CPU utilization remains high for extended periods), while others are idle and unable to dynamically schedule device connections or data collection tasks, resulting in low overall cluster resource utilization.
[0006] Fourth, predictive maintenance is lacking. Existing solutions only provide threshold alarms (such as alarms when disk usage exceeds 90%), and cannot predict future failures based on historical trends (such as predicting when the disk will be full based on the disk usage growth curve), leading to frequent unplanned downtime.
[0007] Fifth, configuration management is chaotic. The configurations (device drivers, collection policies, reporting targets) of multiple gateways are scattered across various nodes, making unified distribution, version alignment, and rollback difficult. Manually logging into each gateway to modify the configuration is prone to errors and makes it impossible to trace the change history.
[0008] In existing technologies, container orchestration platforms such as Kubernetes are mainly designed for stateless applications and are not suitable for stateful gateways that need to be directly connected to physical devices via serial ports / CAN. IoT platforms such as ThingsBoard and AWS IoT provide gateway management functions, but lack predictive analysis based on historical trends and self-healing capabilities, and cannot achieve dynamic device migration across gateways. Summary of the Invention
[0009] The purpose of this invention is to overcome or at least partially solve the above problems by proposing a distributed operation and maintenance management and predictive maintenance system and method for multi-intelligent gateway clusters.
[0010] To achieve the above objectives, the present invention provides a distributed operation and maintenance management and predictive maintenance system for multiple intelligent gateway clusters, including a central service layer and a gateway node layer. The central service layer is deployed on a central server, and the gateway node layer contains multiple distributed industrial IoT gateways.
[0011] The central service layer includes a cluster management agent and a predictive maintenance agent;
[0012] Each gateway in the gateway node layer is an independently running program with built-in multi-protocol drivers and exposes standardized REST APIs, including health check interfaces, status query interfaces, device management interfaces, and log query interfaces.
[0013] The cluster management agent communicates with each gateway node via HTTP REST API, periodically polls to obtain the gateway health status and operating indicators, maintains the gateway registry, device-gateway mapping table and configuration version repository, and performs configuration management, task scheduling and operation and maintenance actions;
[0014] The predictive maintenance agent is coupled with the time-series database. Based on the gateway's historical operating indicators, it performs trend analysis and fault prediction, identifies potential faults such as disk full, memory leak, connection overload, and deterioration of collection quality, generates self-healing strategies, and pushes them to the cluster management agent for execution.
[0015] The cluster management agent, based on the self-healing strategy, calls the gateway REST API to perform tasks such as restarting the data collection process, cleaning up logs, migrating across gateway devices, or rolling back configurations.
[0016] Preferably, the cluster management agent also performs the following operations:
[0017] Receive gateway registration requests, record gateway identifier, network address and protocol capability tags, and maintain the global gateway registry;
[0018] The system polls the health check interface and status query interface of each gateway at a preset cycle, collects system indicators, business indicators and process status, and writes the data into the time series database.
[0019] Maintain configuration version records, support batch distribution of device configurations, configuration activation reload, historical version query and configuration rollback.
[0020] Preferably, the predictive maintenance agent uses a trend analysis model based on historical data to predict indicators such as disk usage, memory usage, number of connected devices, and data collection success rate. When the predicted value exceeds a preset threshold, the corresponding self-healing strategy is triggered.
[0021] Preferably, when the cluster management agent performs cross-gateway device migration, it adopts an atomic API call sequence: first, it creates the device configuration on the target gateway and reloads it to take effect; after verifying that the data collection is normal, it deletes the corresponding configuration on the source gateway to ensure that the data collection is not interrupted during the migration process.
[0022] Preferably, the predictive maintenance agent has root cause analysis capabilities, and when multiple gateways malfunction simultaneously, it can locate the common upstream fault through correlation analysis.
[0023] Preferably, when the cluster management agent detects that the gateway has no response after multiple consecutive polling attempts, it marks it as suspicious or offline and triggers an alarm notification.
[0024] Preferably, the system is compatible with a single gateway multi-agent system. Each gateway runs an automatic device access agent to perform device detection, protocol identification, and driver configuration. The cluster management agent achieves unified management of devices and configurations within the gateway nodes by calling the interfaces provided by the agent.
[0025] Preferably, the cluster management agent accesses a large language model, supporting queries of cluster status in natural language via instant messaging tools or a web interface.
[0026] The present invention also provides a predictive maintenance method based on the above system, comprising the following steps:
[0027] The cluster management agent periodically polls the system and business metrics of each gateway and writes them into the time-series database;
[0028] The future health status of each gateway is assessed using a trend analysis model based on historical data;
[0029] When the predicted value exceeds the threshold, the corresponding self-healing strategy is matched according to the preset mapping relationship between fault type and self-healing strategy.
[0030] The cluster management agent calls the gateway REST API to execute the self-healing strategy and records the operation and maintenance logs.
[0031] The present invention also provides a gateway cluster load balancing method based on the above system, comprising the following steps:
[0032] The cluster management agent polls the REST API of each gateway to collect real-time load metrics.
[0033] Identify high-load and low-load gateways within the cluster;
[0034] Select the set of devices to be migrated from high-load gateways;
[0035] The device is migrated from a high-load gateway to a low-load gateway through an atomic configuration operation, and the global device-gateway mapping table is updated.
[0036] Compared with existing technologies, this invention provides a distributed operation and maintenance management and predictive maintenance system and method for multi-intelligent gateway clusters, which has the following beneficial effects:
[0037] Global visibility and proactive operation and maintenance: It provides a unified real-time monitoring of multi-gateway clusters. Operation and maintenance personnel can view the health status of all gateways, device distribution, and predictive alarms on a single panel, shifting from passive response to proactive operation and maintenance.
[0038] Predictive maintenance reduces downtime: By predicting time-series trends, potential problems such as full disks, memory leaks, and connection overloads can be detected in advance, and self-healing or migration can be performed automatically, significantly reducing unplanned downtime.
[0039] Dynamic load balancing: Automatically migrates devices based on the real-time load of the gateway to avoid single-point overload and improve the overall resource utilization of the cluster.
[0040] Configuration consistency and traceability: Centralized configuration management avoids configuration drift caused by manual login to various gateways. Version records and rollback capabilities make change risks controllable, and all configuration changes are traceable.
[0041] Natural language interaction lowers the barrier to entry: Operations and maintenance personnel do not need to master professional commands; they can complete cluster status queries and operations and maintenance through natural language, reducing training costs. Attached Figure Description
[0042] Figure 1 This is the overall architecture diagram of the multi-gateway cluster operation and maintenance management system of the present invention;
[0043] Figure 2 This is a flowchart of the predictive maintenance agent workflow of the present invention;
[0044] Figure 3 This is a timing diagram for cross-gateway device migration in this invention;
[0045] Figure 4 A flowchart for configuring batch distribution and version rollback for this invention is provided. Detailed Implementation
[0046] The present invention will be further described in detail below with reference to the accompanying drawings.
[0047] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this description, those skilled in the art can make creative modifications to this embodiment as needed, but as long as they are within the scope of the claims of the present invention, they are protected by patent law.
[0048] This invention provides a distributed operation and maintenance management and predictive maintenance system and method for multi-intelligent gateway clusters, which solves the technical problems in the prior art. The overall concept is as follows:
[0049] Example 1
[0050] Please see Figures 1-4 A distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters, comprising:
[0051] Orchestrator layer: Deployed on the central server, it runs the cluster management agent and predictive maintenance agent, communicates with each gateway node through the HTTP REST API, and is responsible for metric collection, status management, task scheduling and predictive inference.
[0052] Gateway Node Layer: Each gateway node is an independent industrial IoT gateway program (e.g., implemented in Go), with built-in multi-protocol drivers (Modbus RTU / TCP, CAN, DL / T 645, IEC 60870-5-104, OPC UA, etc.). It exposes REST APIs, including health check interfaces (GET / health), status query interfaces (GET status), device management interfaces ( / devices series), log query interfaces (GET / logs), data acquisition process restart interfaces (POST / process / restart), and log cleanup interfaces (POST / logs / cleanup). The gateway itself does not have a built-in agent; all intelligent logic is handled by the agent in the Orchestrator layer.
[0053] Physical device layer: Industrial field devices connected to each gateway, including electricity meters, sensors, PLCs, frequency converters, etc.
[0054] The cluster management agent obtains health status and operational metrics by periodically polling the REST API of each gateway, and performs maintenance actions such as configuration distribution and device migration by calling the gateway's device management API.
[0055] The core functions of the cluster management agent:
[0056] Gateway registration and discovery support manual input of gateway addresses by operations and maintenance personnel, or automatic registration confirmation by the cluster management agent during the first polling after the gateway specifies the Orchestrator address in its configuration file. Upon startup, the gateway sends a registration request to the cluster management agent, carrying the gateway ID, IP address, and capability tags (supported protocol types, maximum number of devices, etc.). The cluster management agent maintains the gateway registry, recording the address, status, and capability information of each gateway.
[0057] Health status polling: Periodically (every 30 seconds by default), polls the / health and status interfaces of each gateway to obtain the following metrics:
[0058] System metrics include CPU utilization, memory utilization, and remaining disk space;
[0059] Business metrics include the number of connected devices, data acquisition success rate, error count, and average response latency.
[0060] Process status includes the gateway process's liveness status and MQTT connection status.
[0061] Data is written to a time-series database (such as InfluxDB) to retain historical data for trend analysis. If there is no response after N consecutive polling attempts (default 3), the gateway is marked as "suspicious." If the threshold is exceeded, it is marked as "offline" and an alarm notification is triggered (supports Lark, DingTalk, and email).
[0062] Unified configuration management maintains the configuration version records of each gateway, supporting the following operations:
[0063] Batch distribution of device configurations to a specified gateway (calling the gateway's REST API for device creation / update).
[0064] Call the gateway's / devices / reload interface to apply the configuration;
[0065] Record the version number and content of each configuration change, and support querying historical versions;
[0066] Configure rollback to restore the target gateway's device configuration to a specified historical version and reload it.
[0067] Natural language interaction, integrated with a large language model, allows users to query cluster status in natural language via Lark robot or Web Chat (such as "show all gateways with CPU usage exceeding 70%" or "which gateway had the most failures last week"). The agent then queries the internal status database and replies in natural language.
[0068] The core functionality of predictive maintenance agents:
[0069] For data storage, the cluster management agent writes the metrics obtained from each polling into a time-series database (such as InfluxDB) to retain historical data for trend analysis.
[0070] Fault prediction types are identified by using trend analysis models (such as linear regression, exponential smoothing, etc.) based on historical indicators in the time series database.
[0071] Disk full prediction: Based on the historical growth trend of disk usage, estimate the time point when it reaches 90% and trigger the log cleanup task in advance (call the gateway log cleanup interface or execute logrotate remotely).
[0072] Memory leak prediction: If the detected memory usage shows a continuous increase without a downward trend, it is judged as a suspected leak. The time when the dangerous threshold is predicted is then set, and it is recommended to restart the data collection process (by calling the gateway process restart interface).
[0073] Connection overload prediction: Based on the growth trend of the number of devices connected to the gateway, predict the time when the gateway will reach its limit and schedule some devices to migrate to an idle gateway in advance.
[0074] Predicting data acquisition quality degradation and analyzing the downward trend of acquisition success rate, proactively triggering alarms and suggesting checking network or device connectivity when the success rate is predicted to fall below a threshold in the future.
[0075] The self-healing action execution function allows the predictive maintenance agent to invoke the following actions through the cluster management agent.
[0076] The software self-healing mechanism performs operations such as restarting the data collection process, cleaning log files, and reloading configurations via the gateway's REST API.
[0077] Device migration involves migrating the devices under the responsibility of the problematic gateway to other healthy gateways within the same cluster.
[0078] Alarm notifications are sent via Lark, DingTalk, and email, providing predictive maintenance reports that describe the predicted risks and the actions already taken.
[0079] The root cause analysis function automatically executes the following correlation analysis process when the cluster management agent detects that multiple gateway nodes are abnormal at the same time (such as simultaneously reporting high error rates, lost heartbeats, and failed data collection).
[0080] Abnormal event aggregation collects abnormal events from all gateways in time windows (e.g., the most recent 5 minutes) and groups them according to the type of abnormality (CPU overload, collection failure, network timeout, etc.).
[0081] Common dependency detection analyzes the common dependencies of these abnormal gateways, including shared upstream services (such as MQTT broker, time-series database, configuration center), shared network infrastructure (such as the same switch, router, VPN gateway), and shared physical environment (such as the same rack power supply, the same data center air conditioning).
[0082] Association rule matching uses predefined characteristic rules for common upstream failures. For example, if more than half of the gateways simultaneously experience an "MQTT connection lost" event, it is inferred that the MQTT broker may be down or the network may be experiencing a partition; if all gateways under the same switch simultaneously lose heartbeats, it is inferred that the switch is faulty; if multiple gateways simultaneously report a decrease in data collection success rate and all error codes are "connection timeout", it is inferred that the network device (router, firewall) is abnormal.
[0083] The root cause analysis output, along with the results of the correlation analysis as the root cause conclusion, a list of affected gateways, and recommended troubleshooting steps, is sent to the cluster management agent. The cluster management agent can then send the root cause analysis report to operations and maintenance personnel via alert notifications.
[0084] Optional automatic repair: For upstream failures that can be handled programmatically (such as detecting an MQTT broker anomaly and attempting to restart the service), the root cause analysis results can directly trigger the cluster management agent to execute the corresponding repair actions.
[0085] This root cause analysis capability can quickly distinguish between individual gateway faults and shared upstream faults, avoiding the need for maintenance personnel to check each gateway one by one, and greatly improving the fault diagnosis efficiency of large-scale gateway clusters.
[0086] Example 2
[0087] This embodiment is an application scenario of unified health monitoring of a multi-gateway cluster in the system described in Embodiment 1.
[0088] The deployment scenario involves a factory deploying 5 IoT smart gateways, each responsible for collecting data from Modbus TCP devices in different workshops, all connected to the central server via a local area network.
[0089] The registration process involves each gateway calling the registration interface of the cluster management agent after startup, and reporting the gateway ID, IP address, and list of supported protocols.
[0090] In the polling step, the cluster management agent polls the / health interface of each gateway every 30 seconds to obtain metrics such as CPU, memory, disk, number of online devices, and collection success rate, and writes them to InfluxDB.
[0091] The alarm procedure is as follows: If a gateway fails to respond for three consecutive times, the cluster management agent marks it as offline and notifies the operations and maintenance personnel via Lark robot: "Gateway GW-03 is offline. Last online time was 14:32. Please check the network connection."
[0092] The operation and maintenance personnel inquired via Lark, "Which gateways currently have CPU usage exceeding 70%?" The cluster management agent checked the latest metrics and replied, "GW-02 currently has 78% CPU usage, with 32 connected devices; the rest of the gateways are normal."
[0093] Example 3
[0094] This embodiment is an application scenario of disk full prediction and automatic cleanup of the system described in Embodiment 1.
[0095] The predictive maintenance agent analyzed the disk usage data of GW-04 over the past 7 days and found that it increased from 60% to 82%, and linear extrapolation predicted that it would reach 95% in 3 days.
[0096] The predictive maintenance agent sends alerts to the cluster management agent, recommending log cleanup.
[0097] The cluster management agent notified the operations and maintenance personnel via Lark that "the GW-04 disk is predicted to be full within 3 days. It is recommended to clean up the historical logs. Do you confirm the execution?"
[0098] After user confirmation, the cluster management agent calls the log cleanup interface of GW-04, and the disk utilization rate drops to 45%.
[0099] The predictive maintenance agent records this event and updates the predictive baseline.
[0100] Example 4
[0101] This embodiment describes a load balancing application scenario for cross-gateway device migration in the system described in Embodiment 1.
[0102] The cluster management agent detected that GW-02 CPU usage remained above 85% (connected to 32 devices), while GW-05 CPU usage was only 15% (connected to 8 devices).
[0103] The cluster management agent automatically generates a migration plan to migrate 10 Modbus TCP devices on GW-02 to GW-05.
[0104] The user is notified that "GW-02 is under high load and we plan to migrate 10 devices to GW-05. Data collection will not be interrupted during the migration. Do you confirm?"
[0105] After user confirmation, the migration will be executed according to the following atomic process.
[0106] Query the complete configuration of the device to be migrated from the source gateway (call GET / devices / :uid and GET / devices / :uid / registers).
[0107] Create the same device configuration on the target gateway (by calling POST / devices and PUT / devices / :uid / registers), and then reload to make it effective.
[0108] Verify that the device on the target gateway is collecting data normally (check the status to confirm that the device is online).
[0109] Remove the device configuration from the source gateway (by calling DELETE / devices / :uid) and reload.
[0110] Update the device-gateway mapping table inside the cluster management agent and record migration logs.
[0111] After the migration was completed, the CPU usage of GW-02 dropped to 42%, while that of GW-05 rose to 48%, resulting in a more balanced cluster load. The entire migration process was completed through an atomic sequence of API calls, with the target gateway taking over first during the migration to ensure continuous data collection.
[0112] Example 5
[0113] This embodiment describes the configuration version management and rollback application scenario of the system described in Embodiment 1.
[0114] The operations and maintenance personnel changed the Modbus collection cycle of all gateways from 5 seconds to 1 second through the cluster management agent (batch distribution of configuration).
[0115] After the change, it was found that the error rate of GW-01 data collection increased (the 1-second cycle of this gateway device responded too quickly).
[0116] The operations and maintenance personnel instructed the cluster management agent to roll back GW-01 to the previous version configuration.
[0117] The cluster management agent calls the GW-01 device update interface to restore the original configuration and reloads, and the GW-01 data collection returns to normal.
[0118] The above description of the embodiments is provided to facilitate understanding and use of the present invention by those skilled in the art. It is obvious to those skilled in the art that various modifications can be made to the embodiments, and the general principles described herein can be applied to other embodiments without creative effort. Therefore, the present invention is not limited to the above embodiments. Improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the present invention should be within the protection scope of the present invention.
Claims
1. A distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters, comprising a central service layer and a gateway node layer, wherein the central service layer is deployed on a central server, and the gateway node layer includes multiple distributed industrial IoT gateways, characterized in that: The central service layer includes a cluster management agent and a predictive maintenance agent; Each gateway in the gateway node layer is an independently running program with built-in multi-protocol drivers and exposes standardized REST APIs, including health check interfaces, status query interfaces, device management interfaces, and log query interfaces. The cluster management agent communicates with each gateway node via HTTP REST API, periodically polls to obtain the gateway health status and operating indicators, maintains the gateway registry, device-gateway mapping table and configuration version repository, and performs configuration management, task scheduling and operation and maintenance actions; The predictive maintenance agent is coupled with the time-series database. Based on the gateway's historical operating indicators, it performs trend analysis and fault prediction, identifies potential faults such as disk full, memory leak, connection overload, and deterioration of collection quality, generates self-healing strategies, and pushes them to the cluster management agent for execution. The cluster management agent, based on the self-healing strategy, calls the gateway REST API to perform tasks such as restarting the data collection process, cleaning up logs, migrating across gateway devices, or rolling back configurations.
2. The distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters according to claim 1, characterized in that: The cluster management agent also performs the following operations: Receive gateway registration requests, record gateway identifier, network address and protocol capability tags, and maintain the global gateway registry; The system polls the health check interface and status query interface of each gateway at a preset cycle, collects system indicators, business indicators and process status, and writes the data into the time series database. Maintain configuration version records, support batch distribution of device configurations, configuration activation reload, historical version query and configuration rollback.
3. The distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters according to claim 1, characterized in that: The predictive maintenance agent uses a trend analysis model based on historical data to predict indicators such as disk usage, memory usage, number of connected devices, and data collection success rate. When the predicted value exceeds a preset threshold, the corresponding self-healing strategy is triggered.
4. The distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters according to claim 1, characterized in that: When the cluster management agent performs cross-gateway device migration, it adopts an atomic API call sequence: first, it creates the device configuration on the target gateway and reloads it to take effect; after verifying that the data collection is normal, it deletes the corresponding configuration on the source gateway to ensure that the data collection is not interrupted during the migration process.
5. A distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters according to claim 1, characterized in that: The predictive maintenance agent has root cause analysis capabilities. When multiple gateways malfunction simultaneously, it can locate the common upstream fault through correlation analysis.
6. The distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters according to claim 1, characterized in that: When the cluster management agent detects that the gateway has no response after multiple consecutive polling attempts, it marks it as suspicious or offline and triggers an alarm notification.
7. A distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters according to claim 1, characterized in that: The system is compatible with single-gateway multi-agent systems. Each gateway runs an automatic device access agent to perform device detection, protocol identification, and driver configuration. The cluster management agent achieves unified management of devices and configurations within the gateway nodes by calling the interfaces provided by the agent.
8. The distributed operation and maintenance management and predictive maintenance system for multi-intelligent gateway clusters according to claim 1, characterized in that: The cluster management agent is connected to the large language model, which supports querying the cluster status in natural language through instant messaging tools or web interfaces.
9. A predictive maintenance method based on the system according to any one of claims 1-8, characterized in that, Includes the following steps: The cluster management agent periodically polls the system and business metrics of each gateway and writes them into the time-series database; The future health status of each gateway is assessed using a trend analysis model based on historical data; When the predicted value exceeds the threshold, the corresponding self-healing strategy is matched according to the preset mapping relationship between fault type and self-healing strategy. The cluster management agent calls the gateway REST API to execute the self-healing strategy and records the operation and maintenance logs.
10. A gateway cluster load balancing method based on the system according to any one of claims 1-8, characterized in that, Includes the following steps: The cluster management agent polls the REST API of each gateway to collect real-time load metrics. Identify high-load and low-load gateways within the cluster; Select the set of devices to be migrated from high-load gateways; The device is migrated from a high-load gateway to a low-load gateway through an atomic configuration operation, and the global device-gateway mapping table is updated.