Server cluster monitoring system based on multi-node collaboration and implementation method thereof

By employing a multi-node collaborative architecture and intelligent decision-making modules, the high availability and dynamic load balancing issues of the server cluster monitoring system are resolved, enabling efficient anomaly detection and business continuity while ensuring data privacy and resource utilization.

CN121077908APending Publication Date: 2025-12-05四川华鲲振宇智能科技有限责任公司
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202511184915.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing server cluster monitoring systems have shortcomings in high availability, dynamic load balancing, and anomaly detection, leading to system paralysis, uneven resource utilization, high false alarm rate, high false negative rate, poor business continuity, and risks of data privacy leakage.

Method used

A multi-node collaborative architecture is adopted, which realizes collaborative reasoning and dynamic task sharding among nodes through dynamic topology network module, cross-level indicator collection module and intelligent collaborative decision-making module. Combined with federated learning and knowledge graph construction, an intelligent multi-level response mechanism is built to ensure business continuity.

Benefits of technology

It achieves high availability and elastic scaling, improves monitoring efficiency, reduces service downtime, increases resource utilization, and ensures data privacy and business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121077908A_ABST
    Figure CN121077908A_ABST
Patent Text Reader

Abstract

The invention relates to a server cluster monitoring system based on multi-node collaboration and an implementation method thereof, a dynamic topology network module is configured to reconstruct a connection topology among monitoring nodes in real time according to node performance and link quality, support mixed configuration of a star type, a ring type and a net structure, and realize multi-node collaboration. Multi-dimensional data capture from a physical layer to an application layer is realized through a cross-level index acquisition module based on an integrated hardware sensor interface and a virtualization layer probe, and each node is enabled to perform collaborative reasoning through parameter encryption sharing through a decision model based on federated learning. A monitoring task fragmentation strategy is dynamically adjusted through an adaptive elastic fragmentation unit according to network delay and load fluctuation, and an abnormal event association rule base is updated in real time through an incremental knowledge graph construction unit. High availability and elastic expansion are realized through a multi-node collaborative architecture, the monitoring efficiency is improved in combination with dynamic load balancing and hybrid detection, and an intelligent multi-level response mechanism is constructed to guarantee the service continuity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of server cluster monitoring, and particularly relates to a server cluster monitoring system based on multi-node cooperation and an implementation method thereof. BACKGROUND

[0002] With the rapid development of cloud computing and distributed computing, the scale of modern server clusters is increasingly large, and the number of nodes and business complexity are significantly improved. Traditional monitoring systems mostly adopt centralized architecture, rely on a single control node for data collection and analysis, and have the following technical bottlenecks: if the master node in the centralized monitoring architecture fails, the entire system will be paralyzed, and it is difficult to meet the high availability requirement. The static task allocation strategy cannot adapt to the dynamically changing node load, and manual intervention is required to adjust the monitoring range when a new node is added. Existing anomaly detection mostly relies on threshold alarm or a single algorithm model (such as statistical method), and has insufficient recognition ability for complex and variable cluster anomaly patterns (such as cascading failure), and has high false alarm rate and omission rate. The alarm strategy is mostly fixed grade division, and cannot dynamically adjust the response priority according to the business impact degree, and lacks automatic fault isolation and recovery mechanism. The monitoring data transmission and storage generally adopt general encryption method, and are not optimized for the characteristics of massive time series data, and have leakage risk.

[0003] In the improvement attempt of the prior art, for example, a patent proposes a monitoring architecture based on distributed agents, which alleviates the single point failure problem to some extent, but its load balancing strategy is still based on static weight allocation, resulting in uneven utilization of node resources. The literature "Intelligent Monitoring System Design in Cloud Computing Environment" optimizes anomaly detection by using machine learning algorithm, but the model training relies on centralized data aggregation, which has privacy leakage risk and poor real-time performance. In addition, the existing scheme usually needs to interrupt the monitoring service during disaster recovery, and it is difficult to realize seamless switching.

[0004] In view of the above problems, a new type of server cluster monitoring system and method are needed, which realizes high availability and elastic expansion through a multi-node cooperation architecture, improves monitoring efficiency by combining dynamic load balancing and a hybrid detection model, and builds an intelligent multi-level response mechanism to ensure business continuity. SUMMARY

[0005] The purpose of the present application is to provide a server cluster monitoring system based on multi-node cooperation and an implementation method thereof, which realizes high availability and elastic expansion through a multi-node cooperation architecture, improves monitoring efficiency by combining dynamic load balancing and a hybrid detection model, and builds an intelligent multi-level response mechanism to ensure business continuity.

[0006] To solve the above technical problems, the technical solutions adopted by the present application are as follows:

[0007] In a first aspect, a server cluster monitoring system based on multi-node cooperation is provided, comprising a dynamic topology network module, a cross-level index collection module, and an intelligent cooperative decision module, wherein the intelligent cooperative decision module comprises a decision model based on federated learning, an adaptive elastic fragmentation unit, and an incremental knowledge graph construction unit.

[0008] The dynamic topology network module is configured to reconstruct the connection topology between monitoring nodes in real time according to node performance and link quality, and support a hybrid configuration of star type, ring type, and mesh structure.

[0009] The cross-level index collection module is used to capture multi-dimensional data from the physical layer to the application layer based on integrated hardware sensor interfaces and virtualization layer probes.

[0010] The decision model based on federated learning is used to enable each node to perform collaborative reasoning through parameter encryption sharing.

[0011] The adaptive elastic fragmentation unit is used to dynamically adjust the monitoring task fragmentation strategy according to network delay and load fluctuation.

[0012] The incremental knowledge graph construction unit is used to update the abnormal event correlation rule base in real time.

[0013] Preferably, the dynamic topology network module comprises a topology awareness submodule, a link reorganization submodule, and an energy consumption optimization submodule.

[0014] The topology awareness submodule generates a topology weight matrix by tracking node resource utilization and network round-trip delay in real time.

[0015] The link reorganization submodule automatically enables an alternative routing path when the communication delay between nodes exceeds a set threshold.

[0016] The energy consumption optimization submodule predicts the energy consumption distribution of different topology structures based on Monte Carlo simulation and preferentially selects a networking method with an energy efficiency ratio better than a preset value.

[0017] Preferably, the topology awareness submodule collects key indicators of node CPU / GPU utilization, memory occupancy, and network interface throughput in real time through a distributed probe cluster, constructs a three-dimensional weight matrix model in combination with round-trip delay monitoring implemented by a low-orbit satellite communication network, and updates the weight distribution once every specified time using a sliding window algorithm, wherein the matrix dimensions of the three-dimensional weight matrix model include node computing power weight, link quality weight, and topology stability weight.

[0018] The link reorganization submodule includes a path prediction model constructed based on deep learning. When it is detected that the end-to-end delay exceeds a preset threshold, a three-level response mechanism is automatically triggered. The three-level response mechanism is to preferentially switch to a Mesh network composed of regional edge nodes, secondly select SDN rerouting of cross-regional backbone nodes, and in an emergency, activate a satellite communication backup link. The measured data shows that under the impact of burst traffic in the intelligent manufacturing scene, the service interruption time can be compressed to within 12 ms, which is 73% faster than the recovery speed of the traditional OSPF protocol.

[0019] The energy consumption optimization submodule adopts an improved Monte Carlo simulation algorithm and an energy consumption prediction model based on core parameters such as dynamic transmit power adjustment coefficient, topology node distribution density, protocol stack overhead factor, and heat dissipation system bearing threshold. Through specified iteration simulation, a Pareto optimal solution set is generated.

[0020] Preferably, the topology awareness submodule, link reorganization submodule, and energy consumption optimization submodule interact with each other through an event bus, and a hierarchical energy efficiency management architecture is used to realize closed-loop control. The hierarchical energy efficiency management architecture includes a perception layer, a decision layer, and an execution layer. The perception layer deploys a lightweight data collection agent. The hierarchical energy efficiency management architecture, the execution layer, is driven by OpenFlow / proprietary dual protocols.

[0021] Preferably, the cross-layer index collection module performs hardware-level index fusion collection and synchronously acquires CPU microarchitecture counters, memory RAS characteristic data, and GPU shader unit load. It performs virtualization layer index correlation analysis, establishes a mapping relationship topology graph between virtual machines and host resources, and performs application layer transaction tracking. Through an injected probe, it captures distributed transaction call chain correlations to underlying resource consumption.

[0022] Preferably, the intelligent collaborative decision-making module constructs a distributed gradient aggregation architecture, aggregates and updates the abnormal detection model parameters trained locally by each node through a homomorphic encryption method, implements a double-layer load balancing strategy, divides the monitoring domain using an improved consistent hashing algorithm at the global layer, and uses an ant colony optimization algorithm for task scheduling at the local layer. It also deploys a self-healing knowledge graph engine that automatically starts a graph neural network for knowledge distillation when detecting rule conflicts.

[0023] Preferably, in the distributed gradient aggregation architecture, each edge node builds a lightweight abnormal detection model based on the PyTorch framework. After processing the model parameters generated by local training using a homomorphic encryption algorithm, it performs cross-domain transmission through a multi-cloud computing power scheduling platform. After the control node aggregates and encrypts the parameters, it executes a federated average update strategy to ensure that the original data does not leave the local node during model iteration;

[0024] The double-layer load balancing strategy comprises a global layer and a local layer, the global layer adopts an improved consistent hashing algorithm, divides a specified number of monitoring nodes into a plurality of virtual monitoring domains, configures a primary and backup dual controller for each virtual monitoring domain, realizes a second-level fault switching through a heartbeat detection mechanism, and introduces a load factor to dynamically adjust the weight distribution of nodes on a hash ring, and the local layer is in a single virtual monitoring domain and generates a task scheduling scheme based on an ant colony optimization algorithm.

[0025] Preferably, the system further comprises a causal reasoning alarm module, the causal reasoning alarm module comprises a Bayesian network, a multi-modal alarm fusion unit and a dynamic priority adjustment unit, the Bayesian network is a root cause positioning unit, a causal dependency graph of multi-dimensional monitoring indicators is constructed, the multi-modal alarm fusion unit integrates log features, performance indicators and transaction tracking data for joint analysis, and the dynamic priority adjustment unit automatically upgrades the alarm response level according to the business SLA level.

[0026] In a second aspect, an implementation method of a server cluster monitoring system based on multi-node cooperation is provided, and the implementation of the server cluster monitoring system based on multi-node cooperation comprises the following steps:

[0027] S1: deploying a containerized monitoring agent and constructing a dynamic topology network, and reconstructing the connection topology between monitoring nodes in real time according to node performance and link quality, supporting a mixed configuration of star type, ring type and mesh structure;

[0028] S2: reconstructing the connection topology between monitoring nodes in real time according to node performance and link quality, supporting a mixed configuration of star type, ring type and mesh structure;

[0029] S3: making each node perform collaborative reasoning through parameter encryption sharing through a decision model of federated learning;

[0030] S4: dynamically adjusting the monitoring task fragmentation strategy according to network delay and load fluctuation;

[0031] S5: updating an abnormal event association rule library in real time;

[0032] S6: automatically performing service degradation, traffic rerouting and hot patch loading operations when a key abnormality is detected.

[0033] The beneficial effects of the present application include:

[0034] The application provides a server cluster monitoring system based on multi-node cooperation and an implementation method thereof.

[0035] Firstly, the distributed gradient aggregation architecture homomorphic encryption technology is used to protect the data privacy of model parameters in the aggregation process, avoid the risk of original data leakage, and ensure the consistency of strategies in the rule conflict resolution process through the self-healing knowledge graph engine, and realize the protection of data integrity.

[0036] Secondly, the double-layer load balancing strategy is combined with the improved hash algorithm and the ant colony optimization to realize cross-level resource scheduling, improve the resource utilization rate in the node cluster, and reduce the task scheduling delay.

[0037] Finally, the graph neural network of the knowledge graph engine realizes real-time knowledge distillation, automatically solves the rule conflict, and realizes the minute-level horizontal expansion through the stateless characteristics of the node designed by the distributed architecture, and shortens the node service interruption time. BRIEF DESCRIPTION OF DRAWINGS

[0038] Fig. 1 It is the architecture schematic diagram of the server cluster monitoring system based on multi-node cooperation of the application.

[0039] Fig. 2 It is the flowchart of the implementation method of the server cluster monitoring system based on multi-node cooperation of the application. DETAILED DESCRIPTION

[0040] The application will be further described below with reference to the accompanying drawings. Figs. 1-2 Further detailed description of the application:

[0041] The server cluster monitoring system based on multi-node cooperation includes a dynamic topology network module, a cross-level index collection module, and an intelligent cooperative decision module. The intelligent cooperative decision module includes a decision model based on federated learning, an adaptive elastic fragmentation unit, and an incremental knowledge graph construction unit. The dynamic topology network module is configured to reconstruct the connection topology between monitoring nodes in real time according to node performance and link quality, and supports a hybrid configuration of star, ring, and mesh structure. The cross-level index collection module is used to capture multi-dimensional data from the physical layer to the application layer based on integrated hardware sensor interfaces and virtualization layer probes. The decision model based on federated learning is used to enable nodes to perform collaborative reasoning through parameter encryption sharing. The adaptive elastic fragmentation unit is used to dynamically adjust the monitoring task fragmentation strategy according to network delay and load fluctuations. The incremental knowledge graph construction unit is used to update the abnormal event association rule base in real time.

[0042] In the present embodiment, the dynamic topology network module includes a topology awareness submodule, a link reorganization submodule, and an energy consumption optimization submodule. The topology awareness submodule generates a topology weight matrix by tracking node resource utilization and network round-trip delay in real time. The link reorganization submodule automatically enables an alternative routing path when the communication delay between nodes exceeds a set threshold. The energy consumption optimization submodule predicts the energy consumption distribution of different topology structures based on Monte Carlo simulation and preferentially selects a networking method with an energy efficiency ratio better than a preset value.

[0043] The dynamic topology network module dynamically generates a topology weight matrix based on a node performance scoring model, and the scoring formula is as follows:

[0044] S i =w1·CPU free / CPU total +w2·BW avail / BW max +w3·Latency base / Latency current ;

[0045] wherein w1 is the CPU index weight, w2 is the BW index weight, and w3 is the Latency index weight. CPU free is the CPU resource idle rate, CPU total is the total CPU resources, BW avail is the available bandwidth, BW max is the maximum bandwidth, Latency base is the baseline delay, and Latency current is the current delay.

[0046] Topology switching is realized by a distributed consistency protocol. In a star topology, the core node selects the three nodes with the highest scores as regional centers. In a mesh topology, multi-path redundancy is automatically activated when the link packet loss rate is greater than 5%. In a ring topology, the timing data synchronization between low-power edge devices is realized.

[0047] Embodiment 2

[0048] On the basis of embodiment 1, the topology awareness submodule collects key indicators of node CPU / GPU utilization, memory occupancy, and network interface throughput in real time through a distributed probe cluster, combines the round-trip delay monitoring realized by the low-orbit satellite communication network, constructs a three-dimensional weight matrix model, the matrix dimension of the three-dimensional weight matrix model includes node computing power weight, link quality weight, and topology stability weight, and adopts a sliding window algorithm to update the weight distribution once every specified time. The link recombination submodule includes a path prediction model constructed based on deep learning. When it is detected that the end-to-end delay exceeds a preset threshold, a three-level response mechanism is automatically triggered. The three-level response mechanism is to preferentially switch to a Mesh network composed of regional edge nodes, secondly select an SDN rerouting of cross-regional backbone nodes, and in an emergency, activate a satellite communication backup link. The measured data shows that under the impact of sudden traffic in an intelligent manufacturing scenario, this module can compress the service interruption time to within 12 ms, which improves the recovery speed by 73% compared with the traditional OSPF protocol. The energy consumption optimization submodule adopts an improved Monte Carlo simulation algorithm and an energy consumption prediction model based on core parameters such as a dynamic transmission power adjustment coefficient, a topology node distribution density, a protocol stack overhead factor, and a heat dissipation system bearing threshold. Through specified iterations, a Pareto optimal solution set is generated.

[0049] In this embodiment, the topology awareness submodule deploys a distributed probe cluster containing 128 lightweight Agent nodes, which collects node resources, network state parameters in real time, and calculates a three-dimensional weight matrix. The node resources include CPU / GPU utilization and memory occupancy, and the network state includes interface throughput and satellite communication round-trip delay, which is synchronized and calibrated through a GNSS clock. The three-dimensional weight matrix includes node computing power weight, link quality weight, and topology stability weight. The node computing power weight is calculated by (0.6 x CPU utilization + 0.4 x GPU utilization) / 100, and its update period is 500 ms. The link quality weight is calculated by 1-(delay / 100 ms + packet loss rate x 10), and its update period is 200 ms. The topology stability weight is calculated by the reciprocal of the standard deviation of the node online rate in the sliding window, and its update period is 1 s.

[0050] The topology-aware sub-module, the link reorganization sub-module and the energy consumption optimization sub-module interact with each other through an event bus, and a hierarchical energy efficiency management architecture is adopted to realize closed-loop control, wherein the hierarchical energy efficiency management architecture comprises a perception layer, a decision layer and an execution layer, the perception layer is provided with a lightweight data collection agent, the hierarchical energy efficiency management architecture, the execution layer is driven by OpenFlow / proprietary dual protocols.

[0051] The perception layer comprises a lightweight agent, the data collection frequency is 50Hz, the decision layer uses a DRL-PathNet inference engine, the model inference delay is 2.3ms / request, the execution layer uses OpenFlow / proprietary dual stack driving, and the flow table issuing rate is 10000 per second.

[0052] The cross-layer index collection module collects hardware-level index fusion, synchronously acquires CPU micro-architecture counters, memory RAS characteristic data and GPU shader unit load, performs virtualization layer index correlation analysis, establishes a mapping relationship topology graph of virtual machines and host resources, and performs application layer transaction tracking, and captures the correlation between distributed transaction call chains and underlying resource consumption through an injected probe. The CPU micro-architecture counter reads the RDTSC instruction through the PMC (performance monitoring counter), and samples every clock cycle (3.4GHz). The memory RAS characteristic data captures the ECC correction / uncorrected error count through the EDAC controller, and records through event triggering. The GPU shader unit load samples the SM utilization rate every 100μs.

[0053] The intelligent collaborative decision module constructs a distributed gradient aggregation architecture, aggregates and updates the abnormal detection model parameters trained locally by each node through a homomorphic encryption method, implements a double-layer load balancing strategy, divides the monitoring domain using an improved consistent hashing algorithm at the global layer, and uses an ant colony optimization algorithm for task scheduling at the local layer; and deploys a self-healing knowledge graph engine, which automatically starts a graph neural network for knowledge distillation when detecting rule conflicts.

[0054] In the distributed gradient aggregation architecture, each edge node constructs a lightweight abnormal detection model based on the PyTorch framework, processes the model parameters generated by local training through a homomorphic encryption algorithm, transmits them across domains through a multi-cloud computing power scheduling platform, aggregates and encrypts the parameters at the control node, executes a federated average update strategy, and ensures that the original data does not leave the local node during model iteration. In the distributed gradient aggregation architecture, a ResNet-18 lightweight model is used for local training of edge nodes, a Paillier algorithm is used to encrypt the gradient ΔW, and the key length is 2048bit. The Kubernetes multi-cloud scheduler selects the lowest delay path, and the time delay is less than 50ms.

[0055] The double-layer load balancing strategy includes a global layer and a local layer, the global layer adopts an improved consistent hashing algorithm, divides a specified number of monitoring nodes into a plurality of virtual monitoring domains, configures primary and backup dual controllers for each virtual monitoring domain, realizes second-level fault switching through a heartbeat detection mechanism, introduces a load factor to dynamically adjust the weight distribution of nodes on a hash ring, and the local layer generates a task scheduling scheme based on an ant colony optimization algorithm within a single virtual monitoring domain.

[0056] When the same entity attribute appears contradictory, such as "node state = normal" but "CPU temperature > 90℃", graph neural network (GNN) distillation is performed, a conflict subgraph is constructed, including the contradictory nodes and the associated entities within 3 hops, a node correction probability is calculated through a graph attention network, a correction candidate set is generated, and after manual review, the main knowledge graph is updated.

[0057] It also includes a causal reasoning alarm module, which includes a Bayesian network, a multi-modal alarm fusion unit and a dynamic priority adjustment unit, the Bayesian network root cause positioning unit, a causal dependency graph of multi-dimensional monitoring indicators is constructed, the multi-modal alarm fusion unit integrates log features, performance indicators and transaction tracking data for joint analysis, and the dynamic priority adjustment unit automatically upgrades the alarm response level according to the business SLA level.

[0058] The implementation method of the server cluster monitoring system based on multi-node cooperation is based on the server cluster monitoring system based on multi-node cooperation, and includes the following steps:

[0059] S1: Deploy a containerized monitoring agent and build a dynamic topology network, and reconstruct the connection topology between monitoring nodes in real time according to node performance and link quality, supporting a mixed configuration of star, ring and mesh structure;

[0060] S2: Reconstruct the connection topology between monitoring nodes in real time according to node performance and link quality, supporting a mixed configuration of star, ring and mesh structure;

[0061] S3: Through the decision model of federated learning, each node performs collaborative reasoning through parameter encryption sharing;

[0062] S4: Dynamically adjust the monitoring task fragmentation strategy according to network delay and load fluctuation;

[0063] S5: Real-time update of abnormal event association rule base;

[0064] S6: When a critical abnormality is detected, automatically perform service degradation, traffic rerouting and hot patch loading operations.

[0065] In summary, the server cluster monitoring system based on multi-node cooperation and the implementation method thereof provided by the application reconstruct the connection topology between monitoring nodes in real time according to node performance and link quality through a dynamic topology network module, support the mixed configuration of star type, ring type and mesh structure, realize multidimensional data capture from the physical layer to the application layer based on integrated hardware sensor interfaces and virtualization layer probes through a cross-level index collection module, enable each node to perform collaborative reasoning through parameter encryption sharing based on a decision model of federated learning, dynamically adjust the monitoring task fragmentation strategy according to network delay and load fluctuation through an adaptive elastic fragmentation unit, and update the abnormal event association rule library in real time through an incremental knowledge graph construction unit. Through the multi-node cooperation architecture, high availability and elastic expansion are realized, the monitoring efficiency is improved in combination with dynamic load balancing and mixed detection, and an intelligent multi-level response mechanism is constructed to guarantee business continuity.

[0066] Through the distributed gradient aggregation architecture homomorphic encryption technology, the data privacy of model parameters in the aggregation process is ensured, the risk of original data leakage is avoided, the strategy consistency is maintained in the rule conflict resolution process through the self-healing knowledge graph engine, and the data integrity protection is realized. Through the double-layer load balancing strategy combined with the improved hash algorithm and the ant colony optimization, cross-level resource scheduling is realized, the resource utilization in the node cluster is improved, and the task scheduling delay is reduced. Through the graph neural network of the knowledge graph engine, real-time knowledge distillation is realized, rule conflicts are automatically solved, and minute-level horizontal expansion can be realized through the stateless characteristics of the node designed by the distributed architecture, and the node service interruption time is shortened.

Claims

1. A server cluster monitoring system based on multi-node collaboration, characterized in that, The application relates to a dynamic topology network module, a cross-layer index collection module and an intelligent collaborative decision module, wherein the intelligent collaborative decision module comprises a decision model based on federal learning, an adaptive elastic fragmentation unit and an incremental knowledge graph construction unit. The dynamic topology network module is configured to reconstruct the connection topology between monitoring nodes in real time according to node performance and link quality, and supports a mixed configuration of star type, ring type and mesh structure. The cross-layer index collection module is used for realizing multidimensional data capture from the physical layer to the application layer based on an integrated hardware sensor interface and a virtualization layer probe. The decision model based on federal learning is used for enabling nodes to perform collaborative reasoning through parameter encryption sharing. The adaptive elastic fragmentation unit is used for dynamically adjusting the monitoring task fragmentation strategy according to network delay and load fluctuation. The incremental knowledge graph construction unit is used for updating an abnormal event correlation rule library in real time. 2.The multi-node coordination based server cluster monitoring system according to claim 1, wherein, The dynamic topology network module comprises a topology awareness submodule, a link recombination submodule and an energy consumption optimization submodule. The topology awareness submodule generates a topology weight matrix by tracking node resource utilization and network round-trip delay in real time. The link recombination submodule automatically enables an alternative routing path when the communication delay between nodes exceeds a set threshold. The energy consumption optimization submodule predicts the energy consumption distribution of different topology structures based on Monte Carlo simulation and preferentially selects a networking mode with an energy efficiency ratio better than a preset value. 3.The multi-node coordination based server cluster monitoring system according to claim 2, wherein, The topology awareness submodule collects key indicators of node CPU / GPU utilization, memory occupancy and network interface throughput in real time through a distributed probe cluster, combines low-orbit satellite communication network implemented round-trip delay monitoring, constructs a three-dimensional weight matrix model, and updates the weight distribution once every specified time by using a sliding window algorithm, wherein the matrix dimension of the three-dimensional weight matrix model comprises node computing power weight, link quality weight and topology stability weight. The link recombination submodule comprises a path prediction model constructed based on deep learning, and automatically triggers a three-level response mechanism when detecting that the end-to-end delay exceeds a preset threshold, the three-level response mechanism is preferentially switched to a Mesh network composed of regional edge nodes, secondly selects an SDN rerouting of cross-regional backbone nodes, and activates a satellite communication backup link in an emergency, and measured data shows that, under the impact of burst traffic in an intelligent manufacturing scene, the module can compress the service interruption time to less than 12ms, and the recovery speed is improved by 73% compared with a traditional OSPF protocol. The energy consumption optimization submodule adopts an improved Monte Carlo simulation algorithm, and is based on an energy consumption prediction model of core parameters such as a dynamic transmission power adjustment coefficient, a topology node distribution density, a protocol stack overhead factor and a heat dissipation system bearing threshold, and generates a Pareto optimal solution set through specified iteration simulation.

4. The multi-node coordination based server cluster monitoring system according to claim 3, wherein, The topology-aware sub-module, the link reorganization sub-module and the energy consumption optimization sub-module interact with each other through an event bus, and a hierarchical energy efficiency management architecture is adopted to realize closed-loop control, the hierarchical energy efficiency management architecture includes a perception layer, a decision layer and an execution layer, the perception layer deploys a lightweight data collection agent, the hierarchical energy efficiency management architecture, the execution layer, and openFlow / proprietary dual-protocol driving.

5. The multi-node coordination based server cluster monitoring system according to claim 1, wherein, The cross-layer index collection module performs hardware-level index fusion collection and synchronously acquires CPU micro-architecture counters, memory RAS characteristic data and GPU shader unit loads; performs virtualization layer index correlation analysis, establishes a mapping relationship topology diagram of virtual machines and host resources, and performs application layer transaction tracking, and through an injected probe, captures distributed transaction call chain correlation to underlying resource consumption.

6. The multi-node coordination based server cluster monitoring system according to claim 1, wherein, The intelligent collaborative decision module constructs a distributed gradient aggregation architecture, so that the abnormal detection model parameters locally trained by each node are aggregated and updated through a homomorphic encryption method, and a double-layer load balancing strategy is implemented, an improved consistent hashing algorithm is used in the global layer to divide the monitoring domain, and an ant colony optimization algorithm is used in the local layer for task scheduling; And deploy a self-healing knowledge graph engine, which automatically starts a graph neural network for knowledge distillation when a rule conflict is detected.

7. The multi-node coordination based server cluster monitoring system according to claim 6, wherein, In the distributed gradient aggregation architecture, each edge node builds a lightweight abnormal detection model based on the PyTorch framework, processes the model parameters generated by local training through a homomorphic encryption algorithm, performs cross-domain transmission through a multi-cloud computing power scheduling platform, aggregates and encrypts the parameters at the control node, and executes a federated average update strategy to ensure that the original data does not leave the local node during model iteration; The double-layer load balancing strategy includes a global layer and a local layer, the global layer uses an improved consistent hashing algorithm to divide a specified number of monitoring nodes into a plurality of virtual monitoring domains, each virtual monitoring domain is configured with a primary and backup dual controller, and a heartbeat detection mechanism is used to realize second-level fault switching, a load factor is introduced to dynamically adjust the weight distribution of the nodes on the hash ring, and the local layer generates a task scheduling scheme based on an ant colony optimization algorithm within a single virtual monitoring domain.

8. The multi-node coordination based server cluster monitoring system according to claim 1, wherein, It also includes a causal reasoning alarm module, which includes a Bayesian network, a multi-modal alarm fusion unit and a dynamic priority adjustment unit, the Bayesian network root cause positioning unit, a causal dependency graph of multi-dimensional monitoring indicators is constructed, the multi-modal alarm fusion unit, log features, performance indicators and transaction tracking data are integrated for joint analysis, and the dynamic priority adjustment unit automatically upgrades the alarm response level according to the business SLA level.

9. The implementation method of the server cluster monitoring system based on multi-node cooperation, according to any one of claims 1-8, characterized in that, The steps include: S1: deploying a containerized monitoring agent and building a dynamic topology network, and reconstructing the connection topology between monitoring nodes in real time according to node performance and link quality, supporting a mixed configuration of star type, ring type and mesh structure; S2: according to node performance and link quality, real-time reconstruction of the connection topology between monitoring nodes supports a mixed configuration of star type, ring type and mesh structure; S3: through the decision model of federated learning, each node performs collaborative reasoning through parameter encryption sharing; S4: dynamically adjust monitoring task fragmentation strategy according to network delay and load fluctuation; S5: real-time update abnormal event association rule base; S6: automatically perform service degradation, traffic rerouting and hot patch loading operation when detecting key abnormality.

Citation Information

Cited By

  • Multi-port cooperative control method and system combined with hardware framework structure

    CN121332887A

  • Communication networking method, device and equipment of computing power server and storage medium

    CN121603385A

  • Water consumption terminal abnormal state monitoring and dynamic service authorization method and system

    CN121682650A

  • Metadata processing method based on event driving, electronic equipment and computer program product

    CN121935016A

  • Cluster CPU cooperative frequency regulation method and system, computer equipment and medium

    CN122019189A