Network card real-time monitoring and fault diagnosis system under distributed architecture
By deploying a real-time monitoring and fault diagnosis system for network cards in distributed networks, the problem of difficulty in real-time monitoring and rapid diagnosis of network cards in the existing technology is solved, and accurate monitoring of network cards status and rapid identification of faults is achieved, and network reliability and operation and maintenance efficiency are improved.
Patent Information
- Application Number
- CN202510377583.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to monitor network card status in real time and accurately and diagnose faults quickly, resulting in difficulty in capturing and positioning network faults in a timely manner, affecting business continuity and user experience.
A network card real-time monitoring and fault diagnosis system under a distributed architecture is designed, including data acquisition module, data aggregation node, real-time monitoring module, fault diagnosis module, alarm module, visualization module, remote control module and configuration management module. The system deploys acquisition probes at network nodes, collects network card data in real time, and uses hash algorithms to perform data shunt and load balancing, combining sliding window algorithms and machine learning models for real-time monitoring and fault diagnosis.
Real-time and accurate monitoring of network card status and rapid diagnosis of faults are achieved, with early warnings being more than 30% ahead of schedule and fault identification accuracy exceeding 90%, greatly shortening the inspection time, reducing business interruption time, and improving administrator response efficiency through multi-channel alarm and visual modules.
Smart Images

Figure CN120200893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network card monitoring and diagnosis, and particularly to a real-time monitoring and fault diagnosis system for network cards under a distributed architecture. Background Art
[0002] In today's digital age, the network has been deeply integrated into all walks of life and has become an indispensable infrastructure for enterprise operations and social operation. Due to its advantages such as high efficiency, flexibility, and strong scalability, the distributed architecture is widely used in large enterprises, data centers, and cloud computing environments. However, with the continuous expansion of the network scale and the continuous increase in complexity, as a key component of network connection, the stable operation of network cards faces many challenges.
[0003] Most traditional network card monitoring methods are relatively simple and often rely on basic tools provided by the system or manual regular inspections. These methods have obvious defects: on the one hand, the functions of the system-provided tools are limited, and they can only provide basic information such as traffic statistics, and it is difficult to deeply analyze potential network card faults and cannot meet the requirements of fine management of complex distributed networks; on the other hand, manual inspections are inefficient, not only consuming a large amount of human resources, but also having an inspection interval, making it difficult to quickly capture sudden faults. Once a network card fails, such as hardware damage, driver program anomalies, network configuration errors, etc., if it cannot be quickly detected and located, it will cause network jams and data transmission interruptions, bringing serious losses to network-dependent services, such as financial transaction delays and online service disconnections, affecting user experience and even causing economic losses.
[0004] At the same time, fault diagnosis is even more difficult, lacking intelligent means. When a problem occurs, technicians need to rely on experience to check possible causes one by one, which is a cumbersome and time-consuming process. In the face of a complex distributed network environment, the fault points may be hidden among numerous nodes, and the positioning difficulty is extremely high. In addition, network cards of different brands and models have different characteristics, and traditional monitoring and diagnosis methods lack universality and are difficult to adapt to diverse network card configurations. To sum up, there is an urgent need for an intelligent system that can monitor the status of network cards in real time and accurately and quickly diagnose faults to ensure the reliable operation of distributed networks. Summary of the Invention
[0005] A real-time monitoring and fault diagnosis system for network cards under a distributed architecture proposed by the present invention is to solve the problems mentioned in the above-mentioned prior art.
[0006] To achieve the above object, the present invention adopts the following technical solution: A real-time monitoring and fault diagnosis system for network cards under a distributed architecture, comprising:
[0007] Data acquisition module: Deploy acquisition probes at distributed network nodes to collect network card traffic, packet capacity, and transmission rate data. Let the current network load be L, which represents the ratio of the data volume transmitted in the current network to the network bandwidth, and the sampling frequency adjustment coefficient be k, which is preset according to the network environment and used to adjust the change of the sampling frequency. The sampling frequency is dynamically adjusted according to the formula, where is the initial sampling frequency, that is, the sampling frequency when the network is at the standard load, is the standard load value, representing the typical load level when the network is running normally. The data is transmitted to the data aggregation node in real time. By optimizing the circuit layout, the energy consumption is reduced. An overvoltage protection circuit is built-in. When encountering an abnormal high-voltage impact in the network and the voltage V exceeds the threshold, this circuit quickly cuts off the power supply to protect the internal components, and at the same time sends a fault warning signal to the data aggregation node;
[0008] Data aggregation node: Receive data from the acquisition probes and perform data shunting through the hash algorithm. Let the hash function for data shunting be, calculate the hash value of the input data to determine the shunting direction. The shunting direction is adjusted according to the load conditions of the backend storage nodes. Deploy a load monitoring module at the backend storage nodes to obtain the load indicators such as CPU usage rate and memory occupancy rate of the nodes in real time, and calculate the average load. For each batch of newly received data, determine the shunting direction according to the hash function, and then according to the load balancing factor, when the load of the storage node is higher than 120% of the average load, adjust the shunting probability according to the formula. By adjusting the data flow direction, ensure the pressure balance of the storage nodes;
[0009] Real-time monitoring module: Analyze the aggregated data and use the sliding window algorithm to monitor the network card traffic trend in real time. Let the sliding window size be W, which represents the time span for statistical traffic data, and the time step be, which represents the time interval for each sliding window movement. Calculate the average traffic within the window. The upper limit of the traffic threshold is, which represents the reasonable upper limit of the traffic when the network card is working normally, and the lower limit is, which represents the reasonable lower limit of the traffic when the network card is working normally. When or, trigger an alarm, and combine the packet error rate E and retransmission rate R indicators to judge the running state of the network card. If and (, are preset thresholds, representing the upper limits of the packet error rate and retransmission rate when the network card is working normally), determine whether the network card has a fault;
[0010] Fault diagnosis module: diagnose faults based on rule base and machine learning model. The rule base includes hardware faults and driver conflict rules. The model input is the real-time data vector of the network card, which represents the real-time performance index data of the network card. These data are obtained from the acquisition probe and reflect the current working status of the network card; the output is the fault type vector, which represents the type of fault that occurs, including hardware short circuit fault, driver crash fault, network connection timeout fault and probability vector; the integrated learning method is adopted to integrate the advantages of decision tree and neural network sub-models. In the training stage, the collected fault sample data is divided into training set, validation set and test set in proportion. For the input real-time data vector of the network card, each sub-model is processed in parallel and its own fault probability is output. Finally, the final fault probability is calculated by weighted average, where is the weight of the sub-model, which is determined according to the accuracy and recall rate indicators of the sub-model on the validation set;
[0011] Alarm module: If an abnormality is detected or a fault is confirmed, an alarm is issued through SMS, email, or system pop-up window. The alarm information includes the location of the faulty network card, the fault type, and the estimated impact range, and is pushed to the network administrator. The SMS alarm delay time is, the email alarm delay time is, and the system pop-up window delay time is, which represent the time interval from the occurrence of the fault to the sending of the SMS, email, and pop-up window, and the total response time respectively; the email alarm is accompanied by a troubleshooting guide, which is stored in the knowledge base in the form of pictures and texts, and is automatically extracted and attached from the knowledge base when the email is sent;
[0012] Visualization module: Network card performance data and fault information are displayed in the form of charts, which is convenient for administrators to understand the health status of the network. The chart update interval is set to represent the time difference between two adjacent chart data refreshes. When the system time passes, the data is automatically refreshed;
[0013] Remote management and control module: The administrator accesses the system through the Web or mobile terminal to perform network card restart and parameter adjustment operations; suppose the mobile APP interface operation response time is and the Web operation response time is, which respectively represent the time interval from issuing an operation instruction on the mobile terminal and the Web terminal to receiving execution feedback, and the total operation feedback delay; the mobile APP adopts a simple interface design, and the operation instructions on the APP are transmitted to the server through the HTTPS protocol. After verification by the server, the instructions are forwarded to the node where the target network card is located for execution.
[0014] Furthermore, it also includes a configuration management module: it supports the configuration of collection probes, monitoring rules, and alarm strategies. The administrator can modify the traffic threshold and alarm recipients as needed. The configuration changes take effect in real time without restarting the system. The configuration modification effective time is set to , which means that the configuration takes effect after the modification.
[0015] Furthermore, it also includes a self - repair function module: for software failures, it can automatically detect and repair according to a preset plan. Let the software failure detection period be \(T_d\), which represents the time interval between two adjacent software failure detections. When a failure is detected, a repair program is started, and the repair time is \(T_r\), which represents the time required from the discovery of the failure to the completion of the repair. The repair success rate is \(P\). Optimize \(T_d\), \(T_r\), and the repair plan through simulation tests.
[0016] Furthermore, in the alarm module, SMS alarms support international language templates. The system has a built - in language text library, and the administrator can set the corresponding language template according to the region where the network node is located or user requirements.
[0017] Furthermore, in the visualization module, the charts support interactive operations. When the administrator clicks on the chart area, the front - end interface sends a data request to the background through Ajax technology. The background queries the database to obtain the details of the data packet corresponding to the clicked position at that moment.
[0018] Furthermore, in the remote control module, the mobile APP has a simple interface design and is adapted to the mobile phone system. When operating, the operation instructions of the administrator on the APP are transmitted to the server through the HTTPS protocol. After the server verifies, the instructions are forwarded to the node where the target network card is located for execution.
[0019] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0020] First of all, the real - time monitoring is comprehensive and accurate. The multi - dimensional data collection is combined with intelligent algorithms to quickly capture subtle changes such as network card traffic and data packets, accurately judge the running state, and the fault warning is advanced by more than 30% on average, killing potential risks in the bud.
[0021] The fault diagnosis is efficient and intelligent. It integrates a rule base and a machine - learning model, combines the advantages of multiple sub - models, and the accuracy rate of identifying difficult faults exceeds 90%, greatly shortening the troubleshooting time and reducing the business interruption duration.
[0022] The alarms are timely and diverse. Multiple channels of notifications are accompanied by detailed information to ensure that the administrator is aware of the failure within 30 seconds and can quickly respond according to the guidance. The visualization is intuitive and convenient. The chart interaction allows the administrator to easily understand the network health and master the latest dynamics within 5 minutes.
[0023] The remote control is flexible. The mobile and Web ends can be operated at any time, and the instruction feedback is within 5 seconds, improving the operation and maintenance efficiency. The configuration management is convenient and secure, taking effect in real - time and recording logs, ensuring that the system is optimized as needed and comprehensively improving the management efficiency of distributed network network cards. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic block diagram of a network card real - time monitoring and fault diagnosis system under a distributed architecture proposed by the present invention. Detailed implementation manners
[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0027] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined. In addition, the terms "mounted", "connected" and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations. The present invention will be further described in detail below with reference to the accompanying drawings.
[0028] Refer to Figure 1 : A real-time monitoring and fault diagnosis system for network cards under a distributed architecture, including:
[0029] Data Acquisition Module: Acquisition probes are deployed at each node of the distributed network to collect data such as network card traffic, packet size, and transmission rate at a frequency not lower than 100 Hz. Let the current network load be L, which represents the ratio of the data being transmitted in the current network to the upper limit of the network bandwidth and is used to measure the network busyness; the sampling frequency adjustment coefficient is k, which is preset according to the network environment and is used to adjust the amplitude of the sampling frequency change; the sampling frequency f is dynamically adjusted according to the formula f = f0 + k×(L - L0), where f0 is the initial sampling frequency, that is, the sampling frequency when the network is at the standard load, and L0 is the standard load value, which represents the typical load level when the network is running normally. The data is transmitted to the data aggregation node in real time; the acquisition probe adopts a low-power design, selects a low-power chipset, and reduces energy consumption by optimizing the circuit layout to ensure that the power consumption does not exceed 5W; and it has a self-protection function with an overvoltage protection circuit built in. When encountering an abnormal high-voltage impact in the network and the voltage V exceeds the threshold V th , the circuit quickly cuts off the power supply to protect the internal components, and at the same time sends a fault warning signal to the data aggregation node. The warning signal contains the probe position information for quick positioning and troubleshooting.
[0030] Data Aggregation Node: Receives data from each acquisition probe and uses the hash algorithm for data shunting. Let the hash function for data shunting be H(x), which calculates the hash value of the input data x to determine the shunting direction; the shunting direction is dynamically adjusted according to the load conditions of the backend storage nodes. A load monitoring module is deployed at the backend storage nodes to obtain load indicators such as CPU usage rate and memory occupancy rate of each node in real time and calculate the average load For each batch of new data received, the shunting direction is initially determined according to the hash function H(x), and then according to the load balancing factor β, when the load L of storage node i i is higher than 120% of the average load , the shunting probability is adjusted according to the formula . By dynamically adjusting the data flow direction, the pressure on each storage node is balanced to avoid data congestion; at the same time, data caching technology is adopted, and the caching duration is not less than 10 minutes to ensure data continuity for subsequent analysis.
[0031] Real-time Monitoring Module: Performs multi-dimensional analysis on the aggregated data and uses the sliding window algorithm to monitor the network card traffic trend in real time. Let the sliding window size be W, which represents the time span for statistical traffic data, and the time step is, which represents the time interval for each sliding window movement. The average traffic is calculated within the window The upper limit of the traffic threshold is F max , which represents the reasonable upper limit of the traffic when the network card is working normally, and the lower limit is F min , which represents the reasonable lower limit of the traffic when the network card is working normally. When or Trigger a warning when the time comes, and judge the operating status of the network card by combining the data packet error rate E and the retransmission rate R indicators. If E > E th and R > R th (E th and R th are preset thresholds, representing the upper limits of the data packet error rate and the retransmission rate when the network card is working properly), it is determined that the network card may have a fault.
[0032] Fault diagnosis module: Diagnose faults based on a rule base and a machine learning model. The rule base covers rules for common hardware faults, driver conflicts, etc. The machine learning model is trained with a large number of fault samples. Let the input of the model be the real-time data vector X of the network card = [x1, x2,..., x n , where x1, x2,..., x n respectively represent different real-time performance index data of the network card, such as the network card temperature value, the length of the send queue, the occupancy rate of the receive buffer, etc. These data are obtained from the acquisition probe and are used to comprehensively reflect the current working state of the network card; the output is the fault type vector Y = [y1, y2,..., y m , where y1, y2,..., y m respectively represent different possible fault types, such as hardware short circuit fault, driver program crash fault, network connection timeout fault, etc., and the probability vector P = [p1, p2,..., y m . An integrated learning method is adopted to integrate the advantages of multiple sub-models such as decision trees and neural networks. In the training stage, a large number of collected fault sample data are divided into a training set, a validation set and a test set according to a certain proportion. First, each sub-model is trained with the training set, and then the sub-model parameters are adjusted with the validation set, such as the depth of the decision tree, the number of layers and nodes of the neural network, etc. For the input real-time data vector X of the network card = [x1, x2,..., x n , each sub-model processes it in parallel and outputs its own fault probability p ij , and finally the final fault probability is calculated by weighted average where w j is the weight of the sub-model, which is comprehensively determined according to indicators such as the accuracy rate and recall rate of the sub-model j on the validation set, so that the recognition accuracy of the system for difficult faults is higher, reaching more than 90%.
[0033] Alarm module: Once an anomaly is detected or a fault is diagnosed, it alarms through multiple channels such as text messages, emails, and system pop-ups. The alarm information includes the location of the faulty network card, the fault type, and the estimated impact range, and is pushed to the network administrator, and the response time does not exceed 30 seconds. Let the text message alarm delay time be T sms , the email alarm delay time be T email , and the system pop-up alarm delay time be T popup, respectively representing the time intervals from the occurrence of a fault to the sending of text messages, emails, and pop - ups. The total response time T = max(T sms , T email , T popup ); The text message alarm supports multi - language templates. The system has a built - in multi - language text library. The administrator can set the corresponding language template in the configuration management module according to the region where the network node is located or the user's needs. When a fault occurs and a text message alarm needs to be sent, the system automatically retrieves the corresponding language text, combines information such as the location of the faulty network card, the type of fault, and the estimated impact range to generate an alarm text message, and pushes it to the administrator's mobile phone in a timely manner through the text message gateway; The email alarm comes with a detailed fault troubleshooting guide. This guide is stored in the knowledge base in a graphic and text form and is automatically extracted and attached when the email is sent, facilitating the administrator to quickly locate the cause of the fault and take repair measures.
[0034] Visualization module: Displays network card performance data and fault information in the form of charts. For example, a line chart shows the traffic changes, and a bar chart compares the network card failure rates of different nodes, facilitating the administrator to intuitively understand the network health status. The data update interval does not exceed 5 minutes. Let the chart update time interval be Δt, representing the time difference between two consecutive chart data refreshes. When the system time passes through Δt each time, the data is automatically refreshed; The chart supports interactive operations. The visualization interface is built based on the HTML5 Canvas technology. When the administrator clicks on the chart area, the front - end interface sends a data request to the background through the Ajax technology. The background quickly queries the database to obtain the detailed data packet details corresponding to the clicked position at that moment, such as source IP, destination IP, protocol type, data packet content summary, etc., and then returns it to the front - end in JSON format. The front - end displays it in a pop - up window to improve the information acquisition efficiency.
[0035] Remote control module: The administrator can remotely access the system through the Web end or the mobile end to perform operations such as restarting the network card and adjusting parameters. The operation instructions are encrypted during transmission to ensure network security, and the feedback delay of instruction execution does not exceed 5 seconds. Let the operation response time of the mobile APP interface be T app , and the operation response time of the Web end be T web , respectively representing the time intervals from sending the operation instruction to receiving the execution feedback on the mobile end and the Web end. The total operation feedback delay T op = max(T app , T web) The mobile APP adopts a simple interface design, follows the Material Design or similar mobile design specifications, uses the ReactNative or Flutter cross-platform development framework, and is compatible with mainstream mobile systems. During operation, the operation instructions of the administrator on the APP are encrypted and transmitted to the server via the HTTPS protocol. After the server verifies the identity, the instructions are forwarded to the node where the target network card is located for execution, and the execution result is also encrypted and returned to the APP. An operation feedback area is set in the APP to display the instruction execution status in real time, facilitating the network administrator to control the network status anytime and anywhere.
[0036] In the present invention, there is also a configuration management module: it supports flexible configuration of collection probes, monitoring rules, alarm strategies, etc. The administrator can modify the traffic threshold, alarm recipients, etc. as needed, and the configuration modification takes effect immediately without restarting the system. Let the configuration modification take effect time be T config , T config = 0, indicating that the configuration modification takes effect immediately after modification; the configuration operation has a log recording function, and a database table is specially used to store configuration logs. Each log record contains fields such as the operator, operation time, modified content, etc. When the administrator modifies the configuration such as the traffic threshold and alarm recipients, the system automatically generates a log record in the background and inserts it into the configuration log table. At the same time, it notifies the audit module through the message queue. The audit module regularly reviews the configuration logs for easy traceability and auditing to ensure the security and compliance of the system configuration. The configuration modification takes effect immediately through hot deployment technology, that is, on the basis of not interrupting the system operation, the modified configuration file is dynamically loaded.
[0037] In the present invention, there is also a self-repair function module: for some software-level faults, such as configuration file errors, it can automatically detect and repair according to the preset plan, reducing the need for manual intervention. Let the software fault detection period be T detect , which represents the time interval between two adjacent software fault detections. When a fault is detected, the repair program is started, and the repair time is T repair , which represents the time required from the discovery of the fault to the completion of the repair. The repair success rate is η. T detect , T repair and the repair plan are optimized through multiple simulation tests to improve the value.
[0038] In the present invention, in the alarm module, SMS alarms support multi-language templates. The system has a built-in multi-language text library. The administrator can set the corresponding language template in the configuration management module according to the region where the network node is located or the needs of the user. When a fault occurs and a SMS alarm needs to be sent, the system automatically retrieves the corresponding language text, combines the location of the faulty network card, the type of fault, the estimated impact range and other information to generate an alarm SMS, which is pushed to the administrator's mobile phone in a timely manner through the SMS gateway; the email alarm is accompanied by a detailed troubleshooting guide, which is stored in the knowledge base in the form of pictures and texts, and is automatically extracted and attached from the knowledge base when the email is sent, so that the administrator can quickly locate the cause of the fault and take repair measures.
[0039] In the present invention, in the visualization module, the chart supports interactive operations, and a visualization interface is built based on the Canvas technology of HTML5. When the administrator clicks on the chart area, the front-end interface sends a data request to the background through Ajax technology. The background quickly queries the database to obtain the detailed data packet details at the time corresponding to the click position, such as source IP, destination IP, protocol type, data packet content summary, etc., and then returns it to the front-end in JSON format, which is displayed in a pop-up window to improve the efficiency of information acquisition.
[0040] In the present invention, in the remote control module, the mobile APP adopts a simple interface design, follows Material Design or similar mobile design specifications, uses ReactNative or Flutter cross-platform development framework, and adapts to mainstream mobile phone systems. During operation, the administrator's operation instructions on the APP are encrypted and transmitted to the server via the HTTPS protocol. After the server verifies the identity, it forwards the instructions to the node where the target network card is located for execution, and the execution results are also encrypted and returned to the APP. An operation feedback area is set in the APP to display the command execution status in real time, which is convenient for network administrators to control the network status anytime and anywhere.
[0041] In the present invention, in the configuration management module, the configuration operation has a log recording function, and a database table is used to store the configuration log. Each log record contains fields such as the operator, operation time, and modification content. When the administrator modifies the configuration of the flow threshold, the alarm recipient, etc., the system automatically generates a log record in the background and inserts it into the configuration log table. At the same time, the audit module is notified through the message queue. The audit module regularly reviews the configuration log to facilitate traceability and auditing, ensuring the security and compliance of the system configuration. The configuration modification takes effect in real time and is realized through hot deployment technology, that is, the modified configuration file is dynamically loaded without interrupting the operation of the system.
[0042] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A network card real-time monitoring and fault diagnosis system under a distributed architecture, characterized in that: include: Data collection module: deploys collection probes on distributed network nodes to collect network card traffic, data packet capacity, and transmission rate data; Assume that the current network load is L, which represents the ratio of the amount of data transmitted in the current network to the network bandwidth. The sampling frequency adjustment coefficient is k, which is pre-set according to the network environment and is used to adjust the sampling frequency change. The sampling frequency F is dynamically adjusted according to the formula f=f0+k×(L-L0), where f0 is the initial sampling frequency, that is, the sampling frequency when the network is under standard load, and L0 is the standard load value, which represents the typical load level when the network is operating normally. The data is transmitted to the data aggregation node in real time. By optimizing the circuit layout, the energy consumption is reduced. The overvoltage protection circuit is built in. When the network encounters abnormal high voltage shock, the voltage V exceeds the threshold V th When a fault occurs, the circuit quickly cuts off the power supply to protect the internal components and sends a fault warning signal to the data aggregation node; Data aggregation node: receives data from the collection probes, performs data diversion through the hash algorithm, and assumes that the hash function of data diversion is H(x). The hash value of the input data x is calculated to determine the diversion direction; The diversion direction is adjusted according to the load of the backend storage node. The load monitoring module is deployed on the backend storage node to obtain the CPU usage and memory usage load indicators of the node in real time and calculate the average load. Each time a batch of new data is received, the diversion direction is determined according to the hash function H(x), and then according to the load balancing factor β, when the load L of storage node i i Above average load When 120% of Adjust the diversion probability and ensure balanced pressure on storage nodes by adjusting the data flow direction; Real-time monitoring module: Analyze the aggregated data and use the sliding window algorithm to monitor the network card traffic trend in real time; set the sliding window size to W, which represents the time span of statistical traffic data, and the time step to represent the time interval of each sliding window movement, and calculate the average traffic in the window The upper limit of the flow threshold is F max , which indicates the reasonable upper limit of the traffic when the network card is working normally, and the lower limit is F min , which indicates the reasonable lower limit of traffic when the network card is working normally. or The warning is triggered when the packet error rate E and retransmission rate R are combined to judge the network card operation status. If E>E th And R>R th , where E th , R th It is a preset threshold, which indicates the upper limit of the packet error rate and retransmission rate when the network card is working normally, and determines whether there is a fault in the network card; Fault diagnosis module: diagnose faults based on rule base and machine learning model. The rule base includes hardware fault and driver conflict rules. The model input is the real-time data vector X = [x1, x2, ..., x n ], where x1, x2, …, x n They represent the real-time performance index data of the network card, which are obtained from the acquisition probe and reflect the current working status of the network card; the output is the fault type vector Y = [y1, y2, ..., y m ], where y1, y2, …, y m They represent the types of faults that occur, including hardware short circuit faults, driver crash faults, network connection timeout faults, and probability vectors P = [p1, p2, ..., y m ]; using ensemble learning method, integrating the advantages of decision tree and neural network sub-models, in the training stage, the collected fault sample data is divided into training set, validation set and test set in proportion, for the input network card real-time data vector X = [x1, x2, ..., x n ], each sub-model is processed in parallel and outputs its own failure probability p ij , and finally calculate the final failure probability by weighted average where w j is the weight of the sub-model, which is determined based on the accuracy and recall rate of sub-model j on the validation set; Alarm module: If an abnormality is detected or a fault is confirmed, an alarm is sent through SMS, email, or system pop-up window. The alarm information includes the location of the faulty network card, the fault type, and the estimated impact range. It is pushed to the network administrator. The SMS alarm delay time is set to T. sms , the delay time of email alert is T email , the system pop-up window delay time is T popup , respectively represent the time interval from the occurrence of the fault to the sending of the SMS, email, and pop-up window, and the total response time T = max(T sms , T email , T popup ); The email alert is accompanied by a troubleshooting guide, which is stored in the knowledge base in the form of pictures and texts. It is automatically extracted from the knowledge base and attached when the email is sent; Visualization module: Network card performance data and fault information are displayed in charts, which helps administrators understand the health status of the network. The chart update interval is set to Δt, which indicates the time difference between two adjacent chart data refreshes. When the system time passes Δt, the data is automatically refreshed. Remote control module: The administrator accesses the system through the Web or mobile terminal to perform network card restart and parameter adjustment operations; the mobile terminal APP interface operation response time is set to T app , the Web operation response time is T web , respectively represent the time interval from issuing an operation instruction to receiving execution feedback on the mobile terminal and the Web terminal, and the total operation feedback delay T op =max(T app , T web ); The mobile APP adopts a simple interface design. The operation instructions on the APP are transmitted to the server through the HTTPS protocol. After the server verifies, the instructions are forwarded to the node where the target network card is located for execution.
2. The network card real-time monitoring and fault diagnosis system under a distributed architecture according to claim 1, characterized in that: It also includes a configuration management module: it supports the configuration of collection probes, monitoring rules, and alarm strategies. Administrators can modify traffic thresholds and alarm recipients as needed. Configuration changes take effect in real time without restarting the system. The configuration change takes effect at T config , T config =0, indicating that the configuration takes effect after modification.
3. The network card real-time monitoring and fault diagnosis system under a distributed architecture according to claim 1, characterized in that: It also includes a self-repair function module: for software failures, it can automatically detect and repair them according to the preset plan. Suppose the software failure detection cycle is T detect , represents the time interval between two adjacent software fault detections. When a fault is detected, the repair program is started and the repair time is T repair , which represents the time required from fault discovery to repair completion, and the repair success rate is η. T is optimized through simulation test detect , T repair and repair plans.
4. The network card real-time monitoring and fault diagnosis system under a distributed architecture according to claim 1, characterized in that: In the alarm module, SMS alarms support international language templates. The system has a built-in language text library. Administrators can set corresponding language templates based on the region where the network node is located or user needs.
5. The network card real-time monitoring and fault diagnosis system under a distributed architecture according to claim 1, characterized in that: In the visualization module, the chart supports interactive operations. When the administrator clicks on the chart area, the front-end interface sends a data request to the back-end through Ajax technology. The back-end queries the database to obtain the data packet details at the time corresponding to the click position.
6. The network card real-time monitoring and fault diagnosis system under a distributed architecture according to claim 1, characterized in that: In the remote control module, the mobile APP adopts a simple interface design and is adapted to the mobile phone system. During operation, the administrator's operating instructions on the APP are transmitted to the server via the HTTPS protocol. After verification by the server, the instructions are forwarded to the node where the target network card is located for execution.