Network communication monitoring method and device, electronic equipment and medium
By collecting network card element information in real time and using the Transformer model for online learning, dynamically adjusting the patch size and error threshold, the problem of untimely detection of network communication abnormalities is solved, and efficient and real-time fault node identification and model training efficiency are achieved.
Patent Information
- Application Number
- CN202510591355.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The prior art has problems of untimely and inefficient abnormal detection in network communication monitoring, especially when monitoring with large language models and NVIDIA collective communication library, it affects the model training efficiency and cannot quickly identify faulty nodes.
By collecting meta information when the network card performs a collection of set communication operations in real time, generating a set of evolution rate sequence vectors, and using the Transformer model for online learning, dynamically adjusting the patch size and error threshold, and identifying network communication exceptions.
It realizes efficient and real-time monitoring of network communication status, quickly identifying faulty nodes, ensuring that model training efficiency is not affected, and adapting to complex network environments.
Smart Images

Figure CN120416100A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network communication monitoring, and in particular, to a network communication monitoring method, device, electronic device and medium. Background Art
[0002] The existing C4 Diagnosis technology uses collective communication library logs to real-time check the network communication status. However, the monitoring information it collects is complex, and the abnormal detection process is cumbersome, which will delay the detection of faults, and its real-time performance is inferior to some active detection methods.
[0003] Alternatively, monitor the network communication status by means of (Compute Unified Device Architecture, CUDA) Event marking in the training framework of the Large Language Model (LLM), visualize the three-dimensional parallel training time process along the time axis, and real-time locate the fault to a certain step in the training process. And this active monitoring method itself will affect the efficiency of model training and cannot identify specific fault nodes.
[0004] Alternatively, in the LLaMA3 (Large Language Model Meta AI) model, design the NVIDIA Collective Communications Library (NCCL) Extended to track network activities by sending small information packets and provide a fault snapshot. And these redundant data packets will slow down the transmission rate of the data stream and even cause redundant congestion situations. Therefore, how to efficiently and real-time monitor the network communication status and quickly identify fault nodes has become an urgent problem to be solved. Summary of the Invention
[0005] The present invention provides a network communication monitoring method, device, electronic device and medium, which are used to solve the defects of untimely and low-efficiency abnormal detection of network communication in the prior art, and realize efficient and real-time monitoring of the network communication status and quick identification of fault nodes.
[0006] The present invention provides a network communication monitoring method, including: Real-time collect the meta-information of each data block at the network card outlet when the network card executes a collective communication operation, where the meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; Perform data processing on the meta-information to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation; Input the set of evolution rate sequence vectors into a fault detection model to obtain the predicted rate of each queue pair output by the fault detection model; Compare the rate error between the predicted rate of each queue pair and the corresponding actual traffic rate; Calculate the error mean and standard deviation of all queue pairs based on the rate errors of each queue pair; Perform network communication anomaly judgment based on the error mean and standard deviation, and determine the anomaly point location based on the queue pair number.
[0007] In a possible implementation, the method further includes: Clean the meta information to obtain valid metadata; Normalize the valid metadata, and aggregate the normalized valid metadata to generate a set of evolution rate sequence vectors of all queue pairs of all communication endpoints under the same set communication operation.
[0008] In a possible implementation, the method further includes: When the error mean is greater than the error threshold, it is determined that there is an anomaly in network communication, where the error threshold is dynamically adjusted by the error mean and standard deviation in combination with a preset adjustment parameter; When it is determined that there is an anomaly in network communication, query the queue pair number corresponding to the traffic with the anomaly, and determine the anomaly point location based on the queue pair number; When the error mean is less than or equal to the error threshold, it is determined that the network communication is normal.
[0009] In a possible implementation, the method further includes: The fault detection model is trained through the following steps: Obtain a set of sample evolution rate sequence vectors of all queue pairs created when the network card performs set communication operations; Input the set of sample evolution rate sequence vectors into a Transformer model for online learning, and terminate the model training when the preset early stopping condition is met to obtain the fault detection model, where the preset early stopping condition includes that the validation loss no longer improves, the validation accuracy no longer increases, and the gap between the training loss and the validation loss increases.
[0010] In a possible implementation, the method further includes: Input the set of sample evolution rate sequence vectors into a Transformer model. When the Transformer model learns based on the set of sample evolution rate sequence vectors, adopt an adaptive Patch processing method. By dynamically adjusting the Patch size, optimize the prediction performance of the Transformer model according to the degree of data fluctuation in the time series.
[0011] In a possible implementation, the method further includes: Determine the data change rate of the time series according to the degree of data fluctuation in the time series; When the data change rate is greater than the first fluctuation threshold, use the minimum Patch to capture the changes in the set of sample evolution rate sequence vectors; When the data change rate is less than the second fluctuation threshold, use the maximum Patch to capture the changes in the set of sample evolution rate sequence vectors; When the data change rate is greater than the second fluctuation threshold and less than the first fluctuation threshold, adaptively adjust the size of the Patch using a preset formula.
[0012] The present invention also provides a network communication monitoring device, including the following modules: A data acquisition module, configured to collect in real time the meta-information of each data block at the network card egress when the network card performs a collective communication operation. The meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; perform data processing on the meta-information to generate a set of evolution rate sequence vectors of all queue pairs of all communication endpoints under the same collective communication operation; A prediction module, configured to input the set of evolution rate sequence vectors into a fault detection model to obtain the predicted rate of each queue pair output by the fault detection model; A comparison module, configured to compare the rate error between the predicted rate of each queue pair and the corresponding actual traffic rate; calculate the error mean and standard deviation of all queue pairs based on the rate error of each queue pair; A judgment module, configured to perform network communication anomaly judgment based on the error mean and standard deviation, and determine the anomaly point location based on the queue pair number.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the network communication monitoring method as described in any one of the above.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the network communication monitoring method described in any one of the above.
[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the network communication monitoring method described in any one of the above.
[0016] The network communication monitoring method, device, electronic device and medium provided by the present invention collect, in real time, the meta-information of each data block at the network card egress when the network card performs a collective communication operation, where the meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; perform data processing on the meta-information to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation; input the evolution rate sequence vector set into a fault detection model to obtain the predicted rate of each queue pair output by the fault detection model; compare the rate error between the predicted rate of each queue pair and the corresponding actual traffic rate; calculate the error mean and standard deviation of all queue pairs based on the rate error of each queue pair; perform network communication anomaly judgment based on the error mean and standard deviation, and determine the anomaly point location based on the queue pair number. Compared with the defects of untimely and low-efficiency network communication anomaly detection in the prior art. By this solution, by collecting the metadata information of the collective communication operation at the network card layer and using a deep learning model to predict and detect anomalies in network traffic, it can be applied to high-performance computing clusters in data centers. This method realizes efficient real-time network communication monitoring without affecting model training and can quickly locate faulty nodes or links. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic architecture diagram of the network communication monitoring method provided by the present invention.
[0019] Figure 2 is a schematic flow diagram of the network communication monitoring method provided by the present invention.
[0020] Figure 3 is a schematic structural diagram of the network communication monitoring device provided by the present invention.
[0021] Figure 4 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0023] To facilitate the understanding of the embodiments of the present invention, the following will further explain with specific embodiments with reference to the accompanying drawings. The embodiments do not constitute a limitation to the embodiments of the present invention.
[0024] Figure 1 It is a schematic architecture diagram of the network communication monitoring method provided by the present invention. As Figure 1 shown, the architecture of the network communication monitoring method is used for the network environment of the large model training network. The figure includes a training network, a storage network, a Graphics Processing Unit (GPU) server, and a network switch (switch).
[0025] The training network is connected to multiple GPU servers through a switch (switch). It is responsible for transmitting training data and model parameters between GPU servers to support distributed training tasks.
[0026] The storage network is also connected to the GPU server through a switch (switch). It is responsible for providing data storage and access services. The GPU server can read training data or write model parameters through the storage network.
[0027] The figure shows multiple GPU servers, which are the main computing units for executing model training tasks. Each GPU server is connected to the training network and the storage network for data transmission and storage access.
[0028] The figure shows multiple network switches, which are the core components of the network and are responsible for forwarding data packets between different network devices. The switch connects the training network, the storage network, and the GPU server to ensure efficient data transmission in the network.
[0029] Training data can be transferred from the storage network to the GPU server and then shared across different GPU servers through the training network to support distributed training. Model parameters updated during training can be synchronized across GPU servers through the training network to maintain model consistency. GPU servers can access storage devices through the storage network to read training data or write training results.
[0030] Figure 2 It is a flow chart of the network communication monitoring method provided by the present invention, such as Figure 2 As shown, the method includes the following: S21. Collecting in real time the meta information of each data block exported by the network card when the network card performs the collective communication operation.
[0031] S22. Perform data processing on the meta-information to generate a set of evolution rate sequence vectors of all queue pairs of all communication endpoints under the same collective communication operation.
[0032] The embodiment of the present invention provides a fault detection tool based on a large model training network, which can ensure that the training is seamless and can monitor the network status in real time and identify faulty nodes.
[0033] Specifically, based on passive detection and collective communication operations, we proposed the MegaMon monitoring tool: a large-scale model network monitoring tool that uses the Transformer model to detect time series anomalies by ingesting real-time traffic data from server network cards. MegaMon collects the evolution rate series of all basic communication endpoints (Queue Pairs, QPs) created by the sending network card when performing collective communication operations. This data is then fed into a Transformer model deployed independently of the training network machine for online prediction. If the predicted rate differs from the actual traffic flow by more than a pre-set error threshold, a network communication anomaly is identified and an alarm is issued.
[0034] First, the data collection and processing module in MegaMon collects metadata for each data block at the network interface card (NIC) outlet in real time. This metadata includes the size of each data block, the transmission time window, and the corresponding queue pair number. This metadata is processed to generate a set of evolution rate sequence vectors for all queue pairs across all communication endpoints under the same collective communication operation.
[0035] Specifically, when the network card performs data transmission, the two parties first exchange necessary communication information, including the Global Identifier (GID), Queue Pair Number (QPN), and First In First Out (FIFO) queue. The receiving end prepares a receive buffer and provides the sending end with the data receive address, data block size, and index information. After receiving this information, the sending end starts data transmission by issuing a Work Request (WR). Whenever a WR is completed, the receiving end reclaims the request and obtains the meta information of the transmission based on the feedback of the Completion Queue (CQ). The meta information includes the size of each data block, the transmission time window, and its corresponding QPN. The data acquisition and processing module monitors and manages data transmission by recording this information multiple times. Optionally, the data acquisition and processing module can also perform data cleaning on the meta information to obtain valid metadata; normalize the valid metadata, and aggregate the normalized valid metadata to generate a set of evolution rate sequence vectors for all queue pairs of all communication endpoints under the same set of communication operations.
[0036] Among them, the Global Identifier (GID) is an identifier used to uniquely identify each endpoint node in the network. The GID ensures that in a wide area network or a data center network, each device can be uniquely identified, allowing data packets to be accurately routed to the destination.
[0037] The Queue Pair Number (QPN) is a unique identifier assigned to each queue pair in the channel adapter to identify different communication channels. Each Queue Pair (QP) contains a send queue and a receive queue for processing incoming and outgoing data.
[0038] The First In First Out (FIFO) queue is a data structure where elements always enter from one end and leave from the other end, and the first element to enter will be the first to leave. This structure ensures the sequentiality of data processing.
[0039] S23. Input the set of evolution rate sequence vectors into the fault detection model to obtain the predicted rate of each queue pair output by the fault detection model.
[0040] Input the set of evolution rate sequence vectors into a pre-trained fault detection model. The model predicts the traffic of each queue pair based on the set of evolution rate sequence vectors and outputs the predicted rate of each queue pair.
[0041] Specifically, MegaMon predicts the traffic of all queue pairs at a future moment and generates the The expected rate of a queue pair at time is .
[0042] S24. Compare the prediction rates of all the queue pairs with the errors of the corresponding actual traffic to obtain the mean error and the standard deviation.
[0043] S25. Calculate the mean error and the standard deviation of all the queue pairs based on the rate errors of each queue pair.
[0044] S26. Judge the network communication anomalies based on the mean error and the standard deviation, and determine the anomaly point location based on the queue pair number.
[0045] Then, compare the predicted rate of each queue pair with the actual rate to calculate the prediction error of each queue pair. The formula is as follows: Furthermore, aggregate the errors of all the queue pairs at time to calculate the mean error of all the queue pairs. The formula is as follows: where n is the number of queue pairs. According to the preset error threshold , if the aggregated error at a certain time step exceeds this error threshold, it is marked as an anomaly. To adapt to the dynamically changing data environment, the error threshold can be calculated by calculating the mean error and the standard deviation of all the time steps. The dynamic threshold calculation formula is as follows: where is the adjustment parameter. By predicting the rate of each queue pair separately and aggregating these rate errors to detect anomalies in the collective communication operation, it can effectively identify the anomalies in the communication process and reduce the false alarms caused by network fluctuations.
[0046] Moreover, when it is determined that there are network communication anomalies, the anomaly point location can be determined based on the queue pair number.
[0047] Identify the specific communication channel through the queue pair number (QPN). Combine the global identifier (GID) to determine the device or node where the fault point is located.
[0048] For example, at a certain time step t , the system detects the following situation: Queue Pair 1: The predicted rate is 100 Mbps, the actual rate is 80 Mbps, and the error is 20 Mbps.
[0049] Queue Pair 2: The predicted rate is 120 Mbps, the actual rate is 118 Mbps, and the error is 2 Mbps.
[0050] Queue Pair 3: The predicted rate is 90 Mbps, the actual rate is 60 Mbps, and the error is 30 Mbps.
[0051] Assume the dynamic threshold is 10 Mbps: The errors of Queue Pair 1 and Queue Pair 3 exceed the threshold and are marked as anomalies.
[0052] Through the Queue Pair Number (QPN) and the Global Identifier (GID), the system locates the devices and links where Queue Pair 1 and Queue Pair 3 are located. Further analyzing the transmission path, it is determined that the fault may occur in Link1−2 and Link3−4.
[0053] Advantages of this method: By collecting and analyzing network traffic data in real time, anomalies can be detected quickly. By dynamically adjusting the Patch size and the dynamic threshold, the fault points can be accurately identified. Based on passive detection, it does not affect the efficiency of model training. It can adapt to complex network environments and dynamic data traffic.
[0054] When testing three typical network faults: slow network card, abnormal congestion, and link jitter, MegaMon demonstrates excellent anomaly detection capabilities and shows different response speeds for different types of faults. The slow network card fault causes a significant drop in the transmission rate, and MegaMon can quickly detect the anomaly through the obvious change in the aggregated rate; abnormal congestion is manifested as a violent fluctuation in a short period, and MegaMon flexibly responds to data fluctuations by adjusting the adaptive Patch size to achieve timely feedback; link jitter is difficult to detect due to its small and irregular change amplitude, but MegaMon can still effectively capture this fault through the dynamic monitoring and adaptive adjustment mechanism. Overall, MegaMon demonstrates its efficient and accurate anomaly detection capabilities when dealing with multiple faults in complex network environments.
[0055] The network communication monitoring method provided by the present invention collects, in real time, the meta-information of each data block at the network card egress when the network card performs a collective communication operation. The meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; processes the meta-information to generate a set of evolution rate sequence vectors of all queue pairs of all communication endpoints under the same collective communication operation; inputs the set of evolution rate sequence vectors into a fault detection model to obtain the predicted rate of each queue pair output by the fault detection model; compares the predicted rate of each queue pair with the rate error of the corresponding actual traffic; calculates the error mean and standard deviation of all queue pairs based on the rate error of each queue pair; performs network communication anomaly judgment based on the error mean and standard deviation, and determines the anomaly point location based on the queue pair number. Compared with the defects of untimely and inefficient network communication anomaly detection in the prior art. By this method, through the method of collecting the meta-packet information at the collective communication operation level of the network card layer, and using the deep learning model framework to predict the network traffic and then perform anomaly detection, it can be applied to the high-performance computing network cluster in the data center, which can not only ensure that the model training is imperceptible, but also achieve efficient and real-time monitoring of the network communication status and quickly identify the faulty nodes.
[0056] In the embodiment of the present invention, a fault detection model needs to be pre-trained, and the fault detection model is obtained through the following steps: Obtain a set of sample evolution rate sequence vectors of all queue pairs created when the network card performs a collective communication operation; input the set of sample evolution rate sequence vectors into a Transformer model for online learning, and terminate the model training when the preset early stopping condition is met to obtain a fault detection model. The preset early stopping condition includes that the validation loss no longer improves, the validation accuracy no longer increases, and the gap between the training loss and the validation loss increases.
[0057] During the model training process, an adaptive Patch processing method is adopted, and by dynamically adjusting the Patch size, the prediction performance of the Transformer model is optimized according to the data fluctuation degree in the time series.
[0058] Specifically, to ensure the stability of network transmission and training efficiency, the model not only needs to make predictions quickly and accurately but also needs to have the ability to adapt to dynamic network environments. The traditional Patch processing method slices time series data with a fixed length of Patch and inputs it into the Transformer model for processing, but it has limitations in dealing with network traffic with uneven data changes. Therefore, the embodiment of the present invention proposes an adaptive Patch processing method (AdaptivePatch-Transformer), which optimizes the prediction performance by dynamically adjusting the Patch size according to the severity of data fluctuations in the time series.
[0059] Specifically, the Patch size is determined by the change rate of the time series and the formula is as follows: where is the minimum Patch size, is the maximum Patch size, is the first fluctuation threshold (high fluctuation threshold), is the second fluctuation threshold (low fluctuation threshold), is the adjustment coefficient, is a tiny constant to prevent the denominator from being zero.
[0060] When exceeds , the minimum Patch ( ) is used to finely capture the changes; when it is less than , the maximum Patch ( ) is adopted to focus on the global trend and reduce the computational burden. Between the two, the Patch size is adaptively adjusted according to the formula. This mechanism effectively controls the Patch size through fixed thresholds, ensuring that the model balances accuracy and efficiency when processing complex time series data. In addition, AdaptivePatch-Transformer does not use fixed batch data for training in each iteration but dynamically adjusts the weights through online learning, enabling the model to capture data changes in real time, improving the accuracy and real-time performance of predictions, and is especially suitable for complex network traffic time series prediction tasks.
[0061] During model training, online learning is performed according to the data in the time series, and the training is terminated when the early stopping condition is met.
[0062] Early stopping is a judgment technique used to prevent model overfitting. It monitors the performance of the model on the validation set during the model training process and stops the training early when the model performance starts to decline. The judgment conditions can include: the validation loss no longer improves, the validation accuracy no longer increases, the gap between the training loss and the validation loss increases, etc.
[0063] By collecting the meta-packet information at the collective communication operation level of the network card layer, and using the deep learning model framework to predict the network traffic and then perform anomaly detection, it can be applied in the high-performance computing network cluster of the data center, which can ensure that the model training is imperceptible.
[0064] The network communication monitoring device provided by the present invention will be described below. The network communication monitoring device described below can be mutually referred to the network communication monitoring method described above.
[0065] Figure 3 It is a schematic structural diagram of the network communication monitoring device provided by the present invention, specifically including: The data acquisition module 301 is used to collect the meta-information of each data block at the network card outlet when the network card performs collective communication operations in real time. The meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; perform data processing on the meta-information to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation. For detailed description, refer to the relevant description corresponding to the above method embodiment, which will not be elaborated here.
[0066] The prediction module 302 is used to input the evolution rate sequence vector set into the fault detection model to obtain the predicted rate of each queue pair output by the fault detection model. For detailed description, refer to the relevant description corresponding to the above method embodiment, which will not be elaborated here.
[0067] The comparison module 303 is used to compare the predicted rate of each queue pair with the rate error of the corresponding actual traffic; calculate the error mean and standard deviation of all queue pairs based on the rate error of each queue pair. For detailed description, refer to the relevant description corresponding to the above method embodiment, which will not be elaborated here.
[0068] The judgment module 304 is used to perform network communication anomaly judgment based on the error mean and standard deviation, and determine the anomaly point position based on the queue pair number. For detailed description, refer to the relevant description corresponding to the above method embodiment, which will not be elaborated here.
[0069] Figure 4 It exemplifies a schematic structural diagram of an electronic device, such as Figure 4As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete communication with each other through the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute a network communication monitoring method, which includes: real-time collecting meta-information of each data block at the network card egress when the network card performs a collective communication operation, where the meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; performing data processing on the meta-information to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation; inputting the evolution rate sequence vector set into a fault detection model to obtain the predicted rate of each queue pair output by the fault detection model; comparing the predicted rate of each queue pair with the rate error of the corresponding actual traffic; calculating the error mean and standard deviation of all queue pairs based on the rate error of each queue pair; performing network communication anomaly judgment based on the error mean and standard deviation, and determining the anomaly point location based on the queue pair number.
[0070] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0071] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the network communication monitoring method provided by the above-mentioned various methods. The method includes: collecting in real time the meta-information of each data block at the network card outlet when the network card performs a collective communication operation, where the meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; performing data processing on the meta-information to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation; inputting the evolution rate sequence vector set into a fault detection model to obtain the predicted rate of each queue pair output by the fault detection model; comparing the predicted rate of each queue pair with the rate error of the corresponding actual traffic; calculating the error mean and standard deviation of all queue pairs based on the rate error of each queue pair; performing network communication anomaly judgment based on the error mean and standard deviation, and determining the anomaly point location based on the queue pair number.
[0072] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the network communication monitoring method provided by the above-mentioned various methods. The method includes: collecting in real time the meta-information of each data block at the network card outlet when the network card performs a collective communication operation, where the meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; performing data processing on the meta-information to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation; inputting the evolution rate sequence vector set into a fault detection model to obtain the predicted rate of each queue pair output by the fault detection model; comparing the predicted rate of each queue pair with the rate error of the corresponding actual traffic; calculating the error mean and standard deviation of all queue pairs based on the rate error of each queue pair; performing network communication anomaly judgment based on the error mean and standard deviation, and determining the anomaly point location based on the queue pair number.
[0073] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0074] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A network communication monitoring method, characterized in that, Including: Collecting in real time the meta-information of each data block at the network card egress when the network card performs a collective communication operation, where the meta-information includes the queue pair number corresponding to each data block, and the queue pair includes a send queue and a receive queue; Performing data processing on the meta-information to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation; Inputting the evolution rate sequence vector set into a fault detection model to obtain the predicted rate of each queue pair output by the fault detection model; Comparing the rate error between the predicted rate of each queue pair and the corresponding actual traffic rate; Calculating the error mean and standard deviation of all queue pairs based on the rate error of each queue pair; Judging network communication anomalies based on the error mean and standard deviation, and determining the anomaly point location based on the queue pair number.
2. The method according to claim 1, wherein The performing data processing on the meta-information to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation includes: Performing data cleaning on the meta-information to obtain valid metadata; Normalizing the valid metadata, and aggregating the normalized valid metadata to generate an evolution rate sequence vector set of all queue pairs of all communication endpoints under the same collective communication operation.
3. The method according to claim 1, characterized in that, The judging network communication anomalies based on the error mean and standard deviation, and determining the anomaly point location based on the queue pair number includes: When the error mean is greater than the error threshold, it is determined that there is an anomaly in network communication, where the error threshold is dynamically adjusted by the error mean and standard deviation in combination with a preset adjustment parameter; When it is determined that there is an anomaly in network communication, query the queue pair number corresponding to the traffic with the anomaly, and determine the anomaly point location based on the queue pair number; When the error mean is less than or equal to the error threshold, it is determined that the network communication is normal.
4. The method according to any one of claims 1 to 3, characterized in that The fault detection model is trained through the following steps: Obtaining a sample evolution rate sequence vector set of all queue pairs created when the network card performs a collective communication operation; Inputting the sample evolution rate sequence vector set into a Transformer model for online learning, and terminating the model training when a preset early stopping condition is met to obtain the fault detection model, where the preset early stopping condition includes that the validation loss no longer improves, the validation accuracy no longer increases, and the gap between the training loss and the validation loss increases.
5. The method according to claim 4, characterized in that The inputting the sample evolution rate sequence vector set into a Transformer model for online learning includes: Inputting the sample evolution rate sequence vector set into a Transformer model. When the Transformer model learns based on the sample evolution rate sequence vector set, an adaptive Patch processing method is adopted, and the Patch size is dynamically adjusted to optimize the prediction performance of the Transformer model according to the data fluctuation degree in the time series.
6. The method according to claim 5, characterized in that The adaptive patch processing method optimizes the prediction performance of the Transformer model according to the degree of data fluctuation in the time series by dynamically adjusting the patch size, including: Determining a data change rate of the time series according to a degree of data fluctuation in the time series; When the data change rate is greater than a first fluctuation threshold, a minimum patch is used to capture changes in the sample evolution rate sequence vector set; When the data change rate is less than a second fluctuation threshold, a maximum patch is used to capture changes in the sample evolution rate sequence vector set; When the data change rate is greater than the second fluctuation threshold and less than the first fluctuation threshold, the size of the patch is adaptively adjusted using a preset formula.
7. A network communication monitoring device, characterized in that, include: A data collection module is configured to collect, in real time, metadata of each data block exported by the network card when the network card performs a collective communication operation, the metadata including the queue pair number corresponding to each data block, the queue pair including a sending queue and a receiving queue; process the metadata to generate a set of evolution rate sequence vectors for all queue pairs of all communication endpoints under the same collective communication operation; A prediction module, configured to input the evolution rate sequence vector set into a fault detection model to obtain a predicted rate of each queue pair output by the fault detection model; a comparison module, configured to compare a rate error between the predicted rate of each queue pair and the corresponding actual traffic; and calculate an error mean and a standard deviation for all queue pairs based on the rate error of each queue pair; A judgment module is used to judge network communication anomalies based on the error mean and standard deviation, and to determine the location of anomalies based on the queue pair number.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the network communication monitoring method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the network communication monitoring method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the network communication monitoring method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Traffic data anomaly detection method and device, electronic equipment and storage medium
CN112770112A
Multi-dimensional time sequence anomaly detection method and device, electronic equipment, storage medium and program product
CN119917967A
Wind power prediction method and system based on PatchTST
CN119921317A
Fault positioning method and device, storage medium and program product
CN119922073A