Switch monitoring method and electronic device

CN122802404APending Publication Date: 2026-09-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611259621.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-19
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]本申请提供了交换机监控方法及电子设备,以至少解决相关技术中如何在降低对交换机业务性能的影响的基础上,对交换机进行持续、高精度的性能监控的问题

Benefits of technology

[0011] The switch monitoring method and electronic equipment disclosed in this application avoid competing for resources between monitoring tasks and core services by performing high-frequency data acquisition and model training during the first time period (the period of low business activity). During the second time period (the period of high business activity), the sampling frequency is proactively reduced, and a pre-trained temporal super-resolution model is used to perform high-fidelity reconstruction of the low-frequency sampled data, generating monitoring data sequences with higher temporal resolution and detail accuracy. Thus, while significantly reducing the impact on switch forwarding performance, continuous and high-precision network status monitoring is achieved, reaching an optimal balance between monitoring overhead and monitoring quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802404A_ABST
    Figure CN122802404A_ABST
Patent Text Reader

Abstract

The application discloses a switch monitoring method and electronic equipment, and relates to the technical field of data processing. High-frequency data collection and model training are performed during a business low period, i.e., a first time period, so as to avoid competition for resources between a monitoring task and core business. During a business peak period, i.e., a second time period, the sampling frequency is actively reduced, and a pre-trained time sequence super-resolution model is used to perform high-fidelity reconstruction on low-frequency sampling data, so as to generate a monitoring data sequence with higher time resolution and detail accuracy. Thus, while significantly reducing the influence on the switch forwarding performance, continuous and high-precision network state monitoring is realized, and an optimized balance between monitoring cost and monitoring quality is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a switch monitoring method and electronic device. Background Technology

[0002] In modern data center network environments, switches are critical network devices, and their performance monitoring is of paramount importance. To address the common problems of low sampling frequency and insufficient data accuracy in switch monitoring, high-frequency telemetry is primarily used to collect underlying performance data. However, continuously performing full-speed, high-frequency telemetry on a switch generates a large amount of data flow, which may increase packet forwarding latency, exacerbate jitter, and even reduce the effective throughput of data within the switch, ultimately degrading its service performance.

[0003] Therefore, how to perform continuous and high-precision performance monitoring of switches while minimizing the impact on their service performance is a problem that urgently needs to be solved. Summary of the Invention

[0004] This application provides a switch monitoring method and electronic device to at least address the problem in related technologies of how to perform continuous and high-precision performance monitoring of a switch while minimizing the impact on the switch's service performance.

[0005] This application provides a switch monitoring method, including:

[0006] Within the first time period, the first performance data of the switch to be monitored is collected at the first frequency, and the first performance data is used to train the time series super-resolution model. The first time period is the period when the service load of the switch to be monitored is lower than the preset load threshold. During the second time period, the second performance data of the switch to be monitored is collected at a second frequency lower than the first frequency. The second time period is the period during which the service load of the switch to be monitored is not lower than a preset load threshold. The second performance data is input into the trained temporal super-resolution model, and temporal feature extraction and context information fusion processing are performed to obtain the third performance data of the third frequency, where the third frequency is higher than the second frequency. Real-time monitoring of the switch under test is performed based on the first performance data and / or the third performance data.

[0007] This application also provides a switch monitoring device, including: The acquisition unit is used to acquire the first performance data of the switch to be monitored at a first frequency within a first time period. The training unit is used to train the time-series super-resolution model using the first performance data. The first time period is the period when the service load of the switch to be monitored is lower than the preset load threshold. The acquisition unit is also used to acquire second performance data of the switch to be monitored at a second frequency lower than the first frequency during the second time period, the second time period being the period during which the service load of the switch to be monitored is not lower than a preset load threshold. The processing unit is used to input the second performance data into the trained temporal super-resolution model, perform temporal feature extraction and context information fusion processing, and obtain the third performance data of the third frequency, wherein the third frequency is higher than the second frequency; The monitoring unit is used to perform real-time monitoring of the switch under monitoring based on the first performance data and / or the third performance data.

[0008] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described switch monitoring methods.

[0009] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described switch monitoring methods.

[0010] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described switch monitoring methods.

[0011] The switch monitoring method and electronic equipment disclosed in this application avoid competing for resources between monitoring tasks and core services by performing high-frequency data acquisition and model training during the first time period (the period of low business activity). During the second time period (the period of high business activity), the sampling frequency is proactively reduced, and a pre-trained temporal super-resolution model is used to perform high-fidelity reconstruction of the low-frequency sampled data, generating monitoring data sequences with higher temporal resolution and detail accuracy. Thus, while significantly reducing the impact on switch forwarding performance, continuous and high-precision network status monitoring is achieved, reaching an optimal balance between monitoring overhead and monitoring quality. Attached Figure Description

[0012] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating a switch monitoring method provided in an embodiment of this application; Figure 2 An overall block diagram of switch monitoring provided in this application embodiment; Figure 3 A logical diagram illustrating switch monitoring provided in an embodiment of this application; Figure 4 This is a schematic diagram of data processing for a temporal super-resolution model provided in an embodiment of this application; Figure 5 A schematic diagram illustrating the calculation process of a context vector provided in an embodiment of this application; Figure 6 A schematic diagram of dimensional changes provided in an embodiment of this application; Figure 7 A schematic diagram of model inference for a switch during peak load periods, provided as an embodiment of this application; Figure 8 A schematic diagram illustrating the training of a temporal super-resolution model provided in an embodiment of this application; Figure 9 A schematic diagram of model training for a switch during off-peak hours, provided as an embodiment of this application; Figure 10 This is a schematic diagram of the structure of a switch monitoring device provided in an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0015] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0016] In modern data center networks, service traffic exhibits high dynamism and burstiness. To ensure stable network operation and efficient maintenance, continuous and accurate performance monitoring of switches is necessary. Traditional monitoring methods typically collect performance data at a fixed frequency. However, during periods of high switch load, high-frequency data collection can compete with the switch's core data forwarding services for computing, memory, and bus resources, potentially leading to a decline in forwarding performance and creating a conflict between monitoring overhead and service assurance. Therefore, there is an urgent need for a method that can dynamically adjust monitoring strategies based on the actual load status of the switch, maintaining high-precision monitoring capabilities while ensuring service performance.

[0017] This application provides a switch monitoring method, the core of which lies in distinguishing different service load stages of the switch and adopting differentiated data acquisition and enhancement strategies. Specifically, it makes full use of the abundant system resources when the service load is low to perform high-fidelity data acquisition and model training; while when the service load increases, it switches to lower frequency acquisition and uses a pre-trained intelligent model to perform high-precision reconstruction of the sparsely acquired data, thereby outputting a performance data sequence that meets the monitoring accuracy requirements without increasing the burden on the equipment.

[0018] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] Figure 1 This document provides a flowchart illustrating a switch monitoring method as an embodiment of the present application. The method is described in detail below in conjunction with its execution flow.

[0020] Step 101: During the first time period, collect the first performance data of the switch to be monitored at the first frequency, and use the first performance data to train the time series super-resolution model. The first time period is the time period when the service load of the switch to be monitored is lower than the preset load threshold.

[0021] In the embodiments of this application, the first time period is the period during which the service load of the switch to be monitored is lower than a preset load threshold, which can be understood as a low-load period when services are idle. Service load refers to the workload of the switch to be monitored in handling core tasks such as data forwarding and protocol calculation. It can usually be reflected by a comprehensive evaluation of indicators such as the utilization rate of the central processing unit (CPU), the buffer occupancy rate of the switching chip (Application-Specific Integrated Circuit, ASIC), or the port throughput.

[0022] The preset load threshold is a pre-defined threshold value used to distinguish between off-peak and peak business periods. It can be configured based on historical experience or Service Level Agreement (SLA) requirements. The first frequency is a relatively high data sampling frequency used during this off-peak period, such as sub-millisecond or millisecond level. It can capture subtle changes in network status, thereby providing high-quality, high-resolution raw data samples for model training.

[0023] The primary performance data encompasses multi-dimensional metrics reflecting the operational status of the monitored switch, such as port traffic rate, queue depth, packet error count, and buffer status. The data collection process can be achieved, but is not limited to, through the following methods: based on the IP Flow Information Export (IPFIX) protocol format, efficient data transmission is achieved via a Direct Memory Access (DMA) engine and the Netlink module.

[0024] Temporal super-resolution models are neural network models based on deep learning architectures used to process time-series data, such as Transformer-based Image Super-Resolution (TTSR) models. Their core function is to learn the complex mapping relationship from low-frequency time series to high-frequency time series. During periods of low business activity, this temporal super-resolution model is trained using high-resolution, first-performance data collected at the highest frequency as a supervisory signal. Its goal is to learn the deep temporal patterns and inherent laws governing the changes in network performance metrics over time. This allows for the reconstruction of sequences that are highly similar in trend and detail to high-frequency samples, even when only low-frequency sampled data is input in subsequent applications.

[0025] Step 102: During the second time period, collect the second performance data of the switch to be monitored at a second frequency lower than the first frequency. The second time period is the period during which the service load of the switch to be monitored is not lower than a preset load threshold.

[0026] In the embodiments of this application, the second time period corresponds to the period during which the service load of the monitored switch is not lower than (i.e., equal to or higher than) a preset load threshold, which can be understood as the peak service period when business is busy. During this stage, in order to minimize the resource occupation of critical data forwarding services by the monitoring behavior, the data collection frequency is actively reduced to a second frequency. The second frequency is significantly lower than the first frequency, for example, it may only be a fraction of the first frequency. The data collected at this lower frequency is the second performance data.

[0027] Although the density of the second performance data at different points in time is low and may not directly reveal certain fast transient events, the reduced sampling frequency greatly alleviates the pressure on the switch's processing resources, memory bandwidth, and internal data channels, effectively avoiding problems such as increased packet forwarding latency or decreased throughput caused by excessive monitoring overhead.

[0028] Step 103: Input the second performance data into the trained temporal super-resolution model, perform temporal feature extraction and context information fusion processing to obtain the third performance data of the third frequency, wherein the third frequency is higher than the second frequency.

[0029] In the embodiments of this application, the trained temporal super-resolution model is a temporal super-resolution model that has been trained and reached convergence within a first time period. When the low-frequency second performance data stream is input into the trained temporal super-resolution model, the model first performs temporal feature extraction. This process refers to the model automatically learning and extracting deep, abstract feature representations that characterize the evolution of performance indicators over time from the input discrete time point data through its internal neural network layers (e.g., Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), or temporal convolutional layers). These temporal features go beyond the surface values ​​of the original data and include information such as the trend, periodicity, and correlation of flow changes.

[0030] Next, the model performs contextual information fusion processing. This means that when generating high-frequency estimates for the target time point, the model does not rely solely on the current or most recent input points, but dynamically evaluates and fuses the importance of information from different time points throughout the entire input sequence history using algorithms such as attention mechanisms. The model can determine the state of historical moments that are more relevant to predicting the high-frequency value of the current moment, and accordingly assign different weights to different historical features. This integrates extensive contextual information—the overall context of historical states related to the current moment—into a comprehensive context vector. Based on this context vector rich in historical and current information, the model's decoding part performs super-resolution reconstruction, ultimately outputting third-party performance data.

[0031] The third frequency is higher than the second frequency used in the input data, and can typically reach a level comparable to or even higher than the first frequency used in the training phase. Therefore, although the third performance data is obtained at a low frequency during the physical acquisition phase, after intelligent processing by the model, its data point density and sequence integrity on the time axis are significantly improved, forming a smooth, continuous, and detailed high-precision performance curve.

[0032] Step 104: Perform real-time monitoring of the switch to be monitored based on the first performance data and / or the third performance data.

[0033] In the embodiments of this application, the network monitoring system can flexibly select data sources based on the current time period. During off-peak periods (the first time period), high-frequency collected first performance data is used directly for monitoring and analysis; during peak periods (the second time period), third performance data with high-frequency characteristics, reconstructed from a time-series super-resolution model, is used for monitoring and analysis. In this way, regardless of the switch's load state, maintenance personnel can obtain monitoring data streams that meet the requirements for high-precision analysis, enabling real-time observation of switch status, performance bottleneck analysis, and triggering timely and accurate fault diagnosis and early warning.

[0034] The switch monitoring method presented in this application reduces the direct sampling overhead during periods of high service load. By lowering the physical sampling frequency, it effectively alleviates the competition between monitoring tasks and forwarding services for the limited system resources of the switch, ensuring that the performance of core network services is not affected by monitoring activities. Secondly, while reducing the sampling burden, it intelligently improves the accuracy of monitoring data through a temporal super-resolution model. This allows the output performance data sequence to reach or even exceed the level of high-frequency direct sampling in terms of temporal resolution. Furthermore, the intermediate data points generated through contextual information fusion better conform to the inherent temporal patterns of network traffic, making the monitoring curve smoother and more reasonable, greatly improving the usability and analytical value of the monitoring data. Finally, leveraging the natural fluctuation of service load over time, the resource-intensive model training process is scheduled during resource-sufficient off-peak periods, while only lightweight low-frequency sampling and efficient forward inference are performed during resource-constrained peak periods. This achieves a dynamic optimal balance between system resource utilization and monitoring requirement fulfillment, providing a practical and feasible technical path for implementing lossless high-precision monitoring in large-scale data center networks.

[0035] To facilitate understanding of the overall process of the embodiments of this application, an overall block diagram of switch monitoring is provided in the embodiments of this application, such as... Figure 2 As shown, the management plane sends a data collection task to the business switch, and then the data plane switch transmits the data to the pre-trained model in the control plane for data optimization. The optimized data is then saved to the database, and the display plane displays high-frequency data based on the content of the database.

[0036] In one possible specific implementation of this application, an embodiment of this application also provides a logical diagram of switch monitoring, such as... Figure 3As shown, the switch monitoring process of this application can be implemented in the following way: Performance indicator data of the switch is collected at frequency F1 using an HFT-style High-Speed ​​Telemetry (HFT) system. Performance indicators include, but are not limited to, port traffic, buffer status, queue depth, and number of error packets. The collection process is based on the IPFIX protocol format, and efficient data transmission is achieved through a DMA engine and a Netlink module.

[0037] The collected performance metrics data are input into a trained and converged deep learning oversampling TTSR model. This TTSR model employs an attention-based temporal neural network architecture, specifically optimized for the temporal characteristics of network performance data. Finally, high-precision monitoring data is output at an F2 frequency. The F2 frequency is 2-4 times that of F1. The oversampling process not only increases the number of data points but also fills in reasonable intermediate values, making the monitoring curve smoother. The high-precision, high-smoothness switch performance monitoring data output by the TTSR model can be used for applications such as real-time monitoring, performance analysis, and fault diagnosis.

[0038] In one possible implementation of this application embodiment, when performing temporal feature extraction and context information fusion processing to obtain third performance data at a third frequency, the following methods can be used, but are not limited to: encoding the second performance data into a data matrix, wherein the number of rows in the data matrix is ​​the product of the time interval for collecting the second performance data and the second frequency, and the number of columns in the data matrix is ​​the number of performance index types, and the second performance data contains multiple types of performance indices; for each data point in the data column corresponding to each performance index, temporal features are extracted sequentially according to the time step corresponding to each data point to obtain a first feature sequence reflecting the temporal change pattern of the second performance data; information fusion is performed on the first feature sequence based on an attention mechanism to generate a context vector reflecting the importance information corresponding to each data point; multiple target output time steps are determined based on the time interval corresponding to the third frequency, and oversampled data points corresponding to each of the multiple target output time steps are generated according to the context vector to obtain the third performance data.

[0039] In the embodiments of this application, the second performance data includes various types of performance metrics, such as inbound traffic, outbound traffic, packet loss, packet error, queue depth, etc., with each performance metric representing an independent monitoring dimension. The encoding process specifically involves constructing a two-dimensional matrix, or data matrix, from the second performance data obtained within a collection period and arranged chronologically. The number of rows in this two-dimensional matrix is ​​determined by the product of the total time span (i.e., time interval) during which the second performance data is collected and the second frequency, intuitively reflecting the total number of data points collected within this time period. The number of columns in this two-dimensional matrix is ​​equal to the number of performance metric types, with each column corresponding to a sequence of performance metrics of a specific type changing over time. Through matrix representation, multi-dimensional, time-evolving performance data is integrated into a regular tensor form suitable for deep learning model processing.

[0040] For each column of the data matrix, i.e., for each independent performance index sequence, a time-series feature extraction process is initiated. This extraction is performed sequentially according to the time step corresponding to each data point. A time step is the basic discrete unit when processing time-series data, and each time step typically corresponds to a sampling time. The model utilizes built-in temporal neural network layers (e.g., Recurrent Neural Networks (RNNs) or their variants, Long Short-Term Memory Networks (LSTMs), Gated Recurrent Units (GRUs)) to progressively read and process subsequent data points, starting from the first data point in each column.

[0041] For the input at the current time step, a new hidden state is generated through a nonlinear transformation, combining the hidden state calculated at the previous time step. This new hidden state not only contains information from the current input but also encodes all historical information up to the current moment in a compressed form. Through iterative calculations step-by-step, a corresponding first feature sequence is generated for each input performance index sequence. The first feature sequence is a series of high-dimensional feature vectors, each located at a time step. It is no longer the raw, readable performance value but an abstract representation reflecting the inherent temporal variation pattern of the performance index. For example, it may encode deeper information such as trends of increasing or decreasing traffic, short-term fluctuation patterns, and potential periodicity.

[0042] To generate high-frequency output points, the entire historical information needs to be utilized. Therefore, an attention mechanism is introduced to deeply fuse information from the first feature sequence. The core idea of ​​the attention mechanism is to dynamically evaluate the importance or relevance of information from different historical time points in the sequence for generating the current target output. Specifically, each feature vector in the first feature sequence is converted into a key vector and a value vector using a learnable parameter matrix. Simultaneously, a query vector is generated based on the target time (or query intent) for generating the current data point. By calculating the similarity between the query vector and all historical key vectors, a series of importance information, i.e., attention weights, is obtained. These attention weights, after normalization, represent the importance distribution of features at each historical time point to the current generation task. Finally, the corresponding value vectors are weighted and summed using the attention weights to generate a comprehensive context vector. The context vector is a fixed-dimensional vector; it is not simply an average of historical features, but rather selectively and dynamically fuses the most relevant parts of the first feature sequence to the current task, providing rich and focused contextual information for subsequent generation steps.

[0043] Finally, based on the third frequency of the final desired output, it can be determined how many high-frequency data points need to be inserted between the two low-frequency sampling points of the second performance data, i.e., multiple target output time steps are determined. For each target output time step, the context vector associated with that target output time step is used as a key input. The decoder part of the model (possibly composed of fully connected neural network layers) receives this context vector rich in historical information and learns to map it to the performance index value at the corresponding target time, i.e., the oversampled data points. Oversampling refers to inserting new, reasonable data points between known data points while maintaining the main features of the original signal, thereby improving the resolution of the data sequence on the time axis.

[0044] By generating corresponding oversampled data points for all target output time steps and arranging them in chronological order, high-frequency, continuous third-performance data was successfully reconstructed from low-frequency second-performance data. This process not only increases the number of data points, but more importantly, because each inserted point is generated based on an understanding of the global temporal context, the reconstructed sequence more accurately reflects the true, smooth trajectory of network performance changes, rather than simple linear interpolation.

[0045] This application provides a foundation for multi-dimensional parallel processing by constructing a data matrix, accurately captures dynamic patterns of indicators such as network traffic through time-series feature extraction, and realizes intelligent filtering and weighted fusion of historical information through an attention mechanism. This ensures that the final generated oversampled data points not only meet the requirements of high-frequency monitoring in terms of quantity, but also have a high degree of rationality and fidelity in terms of quality, providing a data foundation for accurate monitoring and analysis.

[0046] In one possible implementation of this application embodiment, when performing temporal feature extraction for each data point in the data column corresponding to each performance index according to the time step corresponding to each data point, it can be implemented in the following way, but is not limited to: combining the target data point with the historical hidden state passed by the trained temporal super-resolution model to obtain combined data, and performing feature extraction on the combined data to obtain the hidden state corresponding to the target data point, wherein the hidden state is used as the historical hidden state for feature extraction of the next data point; until the temporal feature extraction of all data points is completed, the hidden state sequence composed of the hidden states corresponding to all data points is determined as the first feature sequence.

[0047] In the embodiments of this application, the temporal feature extraction process iteratively processes each target data point in the data column sequentially. A target data point refers to the specific data value input to the temporal neural network layer at the current processing moment, representing the original observation value of a certain performance index at a specific sampling time. When processing begins, a historical hidden state is initialized. Initially, the historical hidden state is typically set to a zero vector or some initial value, symbolizing the starting point where the model knows nothing about the past. As processing progresses, the historical hidden state represents the model's generalized memory and understanding of all historical data seen up to that point, retained after processing the previous data point. For each target data point, the model does not view it in isolation but combines it with the current historical hidden state carrying historical information. In practice, this combination typically involves concatenating the numerical representation of the current data point (which may have undergone simple transformations) with the hidden state vector from the previous time step, forming a comprehensive combined data vector.

[0048] Next, the model's computational unit (e.g., an RNN cell or an LSTM unit) performs nonlinear transformations and feature extraction on the combined data. This computational unit contains trainable weight parameters and activation functions, designed to learn how to update its internal memory and extract more refined, higher-level representations from the combination of current input and historical states. After computation, the unit outputs a new hidden state. This newly generated hidden state has a dual identity: first, it serves as the feature representation of the current target data point after being understood by the model, encoding all relevant information and its evolution from the beginning of the sequence to the current time step; second, it is immediately assigned a new role, transforming into the historical hidden state upon which the next data point is processed. In this way, information flow achieves seamless transfer and evolution between time steps.

[0049] This process is repeated cyclically. For each data point processed, the model absorbs new information and updates its internal state summary accordingly. When the last data point in the data column is processed, the model has completed a scan of the entire data column. At this point, the model has calculated and saved a corresponding hidden state for each data point in the data column. Arranging these hidden states in chronological order of their corresponding data points forms a hidden state sequence. Each vector in this hidden state sequence is a summary of its corresponding moment and all its previous history. This final hidden state sequence, carrying complete sequence information, is explicitly defined as the first feature sequence. Therefore, the essence of the first feature sequence is a feature vector with context awareness; it is no longer a simple record of the original data, but a deep feature encoding of the temporal dynamics of the input performance index sequence by the model.

[0050] Furthermore, the process of temporal feature extraction can be represented by, but is not limited to, the following formula:

[0051] in, Indicates the hidden state of the target data point. The time step for the target data point. This indicates that the input data is the target data point. To be hidden in history and These are the model parameters of the trained temporal super-resolution model, and σ is the preset activation function.

[0052] Finally, the high-dimensional features are extracted through the RNN unit of the model to obtain the first feature sequence. During the generation of the first feature sequence, the dimensionality of the sequence changes as follows: the data column corresponding to each performance indicator changes from dimension (…). ), which became the dimension of the first feature sequence ( ),in, The time interval for collecting the second performance data. For the second frequency, It is determined by the model parameters in the trained temporal super-resolution model.

[0053] This application endows the model with memory capabilities, enabling it to understand each new data point based on a complete historical context. This allows it to effectively capture long-term dependencies that may exist in network performance data, such as gradual traffic trends lasting several seconds or periodic small fluctuation patterns. Simultaneously, the recursive processing mechanism possesses inherent causality and sequentiality, highly consistent with the natural process of time-series data generation. This makes the feature extraction process efficient and coherent, with each hidden state incrementally updated based on the previous state, avoiding redundancy in information processing. The resulting first feature sequence (i.e., the hidden state sequence) is a coherent, context-rich feature stream, providing directly operable information units for subsequent attention mechanisms. This allows the model to perform more accurate and effective weight allocation when fusing historical information to generate new data points, ultimately ensuring the authenticity and smoothness of the reconstructed high-frequency performance data in terms of temporal logic.

[0054] In one possible implementation of this application embodiment, when generating a context vector reflecting the importance information corresponding to each data point, the following methods can be used, but are not limited to: projecting the first feature sequence through different preset linear transformation parameters to generate a key vector sequence and a value vector sequence corresponding to the first feature sequence, wherein the trained temporal super-resolution model includes multiple preset linear transformation parameters associated with the attention mechanism; projecting the hidden state into a query vector through a linear transformation; calculating the similarity between the query vector and each key vector in the key vector sequence, and normalizing the similarity to obtain the target attention weight, which is used to characterize the importance features of each data point relative to the generated oversampled data points; and performing a weighted summation of each value vector in the value vector sequence based on the target attention weight to obtain the context vector.

[0055] In the embodiments of this application, the attention mechanism is a computational model that simulates cognitive attention. It allows the model to flexibly assign different importance weights to different parts of the input sequence when processing information, thereby achieving focus on key information.

[0056] First, the first feature sequence needs to be transformed into a form suitable for attention computation. Specifically, the trained temporal super-resolution model internally stores multiple sets of trainable pre-defined linear transformation parameters associated with the attention mechanism. These pre-defined linear transformation parameters are learned during the model training phase, and they define the projection relationship from the feature space to a specific attention subspace. Using one set of parameters (e.g., key projection parameters), each hidden state vector in the first feature sequence is linearly transformed to generate a corresponding key vector sequence. Simultaneously, using another set of parameters (e.g., value projection parameters), the same first feature sequence is transformed to generate a value vector sequence. The key can be understood as an index or identifier of information at each historical moment, used for matching with the query; while the value contains the specific information content to be extracted and used at each historical moment.

[0057] To perform attention computation for a specific generation objective (e.g., generating a new high-frequency data point between the original two sampling points in the second performance data), the model needs to form a focus. This focus originates from the model's current state, specifically represented by the hidden state at the current time step (which summarizes all historical information up to the current moment). Projecting this hidden state through another independent set of linear transformation parameters (e.g., query projection parameters) yields the query vector. The query vector represents the model's focus at the current moment; it is the active query in the attention computation.

[0058] Next, the model calculates the similarity between the query vector and each key vector in the key vector sequence. This similarity calculation is typically implemented using dot product operations or additive attention, essentially measuring the strong association between the current query intent and the information identifiers at each historical moment. The calculated raw similarity scores are then normalized (usually using the Softmax function) to become a probability distribution sequence of weights, i.e., the target attention weights. Each target attention weight corresponds to a data point (and its corresponding hidden state), representing the importance of the features carried by that data point relative to the currently generated oversampled data points. The higher the weight, the more crucial and relevant the information at that historical moment is to the current generation target.

[0059] Finally, using the target attention weights that incorporate importance judgments, a weighted summation is performed on each corresponding value vector in the value vector sequence. This process does not simply take the average of historical information, but rather performs a differentiated fusion based on its relevance to the current task. Value vectors with higher weights occupy a larger proportion of the summation result, while those with lower weights contribute less. Through this weighted aggregation, information from all historical moments is selectively integrated into a fixed-dimensional vector according to importance proportions; this vector is the final context vector. This context vector is a targeted information summary; it does not contain all historical details, but focuses on the historical context most relevant to generating the current goal, thus providing crucial information for the model decoder to generate accurate oversampled data points.

[0060] Furthermore, the process of generating context vectors can be represented by, but is not limited to, the following formula:

[0061]

[0062]

[0063]

[0064] in, The query vector has dimensions of 1. , It is a hidden state. For linear transformation projection parameters in the model (query projection parameters); It is a sequence of key vectors. Preset linear variation parameters (key projection parameters). It is a sequence of value vectors. Preset linear variation parameters (value projection parameters). This is the first characteristic sequence. for The dimensions of the context vector at time step, the key vector sequence, the value vector sequence, the first feature sequence, and the context vector are all [missing information]. h is determined by the model parameters, and Softmax is the activation function.

[0065] This application achieves intelligent and dynamic information fusion. The model can adaptively adjust its focus on historical information based on each specific generation task (i.e., different target output time steps), rather than using a fixed fusion mode, making the data reconstruction process more flexible and accurate. Projecting feature sequences into key-value vectors through different parameter sets enables the model to process information identifiers and content separately in different semantic subspaces, enhancing the expressive and discriminative capabilities of attention computation. By generating context vectors through weighted summation of normalized attention weights, the oversampled data points output by the model can be inferred based on the most relevant historical evidence, effectively avoiding interference from irrelevant historical noise. This ensures that the final reconstructed high-frequency performance data sequence not only has high temporal resolution but is also numerically reasonable and reliable, truly reflecting the evolution logic of the network state at the micro-timescale.

[0066] In one possible specific implementation of this application, in order to facilitate understanding of the process of temporal feature extraction and context information fusion, the following implementation is provided as an example: Regarding the trained temporal super-resolution model, this application provides a data processing schematic diagram of the temporal super-resolution model, such as... Figure 4 As shown, the temporal feature extraction module serves as the backbone network, progressively extracting high-dimensional time-series features from the input data. The attention mechanism module is embedded within the backbone network to dynamically adjust the importance of different historical moments, thereby guiding the model to prioritize which historical information to focus on when generating the next data point.

[0067] Specifically, the time-series feature extraction module processes the input data sequentially to construct a feature matrix containing high-dimensional time-series features. Current time step The original data ( (m) refers to the second performance data. F1 represents the low-frequency sampling period, i.e., the second frequency, and m is the number of performance indicator types. If decomposed into individual performance indicator types, then ( , 1), Since the m dimensions are independent of each other, the multi-dimensional effect can be achieved by processing the m dimensions in parallel at the same time.

[0068] The temporal feature extraction module here can be composed of various time series processing networks, including but not limited to recurrent neural networks, gated recurrent neural networks, and long short-term memory neural networks, based on the current input (target data point) and the hidden state of the previous time step. Together they determined the hidden state after the update. It reflects all historical information up to the current time. When the model needs to generate oversampled data points, the attention mechanism evaluates all hidden states ( , , ..., ), calculate the context vector Regarding the calculation process of context vectors, embodiments of this application provide a schematic diagram of the context vector calculation process, such as... Figure 5 As shown: During the calculation of this context vector, each historical hidden state The key vector sequence is obtained through a series of linear transformations related to the attention mechanism. Sum value vector sequence The query vector associated with the current target output time point It is usually derived from the current hidden state. Through comparison... With each The similarity between the two (e.g., dot product) is used to obtain the attention score, which is then normalized to the target attention weights using the Softmax function. This ultimately forms a weighted context vector. This context vector integrates comprehensive historical information and the latest sequence state, providing rich contextual information for the decoder to use in generating the final oversampled data points.

[0069] Furthermore, regarding the dimensional changes of the sequences in this application, i.e., the data columns corresponding to each performance metric, from the perspective of dimension ( ), which became the dimension of the first feature sequence ( In addition to the process described in this application, embodiments also provide a schematic diagram of dimensional changes, such as... Figure 6 As shown, high-dimensional features are extracted through an RNN unit, with the dimensionality changing from ( )become( ).

[0070] In one possible implementation of this application embodiment, when performing time-series feature extraction, a parallel and independent processing method is adopted, and time-series feature extraction and context information fusion are performed separately for the data column corresponding to each performance index.

[0071] In the embodiments of this application, performance metrics such as port bandwidth utilization, queue depth, and error packet count, although all belonging to monitoring data reflecting the status of a switch, often differ significantly in their physical meaning, numerical units, variation patterns, and dynamic range. For example, traffic data may exhibit continuous fluctuations over a large numerical range, while error packet counts may remain at zero for extended periods, occasionally showing sparse spikes. Treating them as a whole for mixed modeling may make it difficult for the model to learn the unique inherent patterns of each metric, and the numerical differences and pattern conflicts between different metrics may also interfere with the model's training and inference processes.

[0072] Therefore, this application employs a parallel and independent processing architecture. Parallelism primarily refers to logical concurrent processing, allowing the system to configure independent or shared but independently invoked processing pipelines for each type of performance metric data. Parallel and independent processing means that the entire forward propagation process of each performance metric's data flow—from feature extraction and state update to attention calculation and finally the generation of oversampled data points—is isolated and does not interfere with each other in mathematical calculations and data flow. Specifically, for each data column in the data matrix (representing a performance metric), the model performs a separate process of temporal feature extraction and contextual information fusion: starting from that column of data, iterative temporal feature extraction based on hidden states is performed sequentially to obtain a first feature sequence (hidden state sequence) that reflects only the changing patterns of that metric; subsequently, key-value projection under the attention mechanism is performed independently based on this sequence, and an independent query is formed for the generation target to calculate the attention weight specific to the historical sequence of that metric. Finally, the context vector for that metric is fused to reconstruct the high-frequency sequence of that metric.

[0073] This application leverages the independence of different performance metrics through parallel and independent processing, avoiding feature confusion. This allows model paths (or sub-modules within the model corresponding to each metric) optimized for each metric to more accurately learn and capture the unique temporal dynamics of that metric, whether it's smooth changes, step changes, or sudden spikes, thereby improving the reconstruction accuracy and fidelity for each individual metric. Secondly, from a computational efficiency and implementation perspective, the independent processing logic of the data columns is very clear, facilitating efficient batch computation on parallel computing hardware. Multiple processing threads can be started simultaneously, or the broadcast mechanism of tensor operations can be used to complete the same operation steps for all metric columns at once, significantly improving the overall data processing throughput and meeting the low latency requirements of real-time monitoring systems. When it is necessary to add or remove monitoring metric types, the corresponding processing pipeline can be modularly adjusted without redesigning the entire model structure. Simultaneously, anomalies or noise in the data stream of one performance metric are difficult to propagate to the processing results of other performance metrics through shared model parameters, ensuring the stability of the overall monitoring output. In summary, by adopting a parallel and independent processing strategy targeting multiple performance metrics, a good balance is achieved in processing efficiency, model specialization, and system flexibility while ensuring high-precision reconstruction.

[0074] In one possible implementation of this application embodiment, when generating oversampled data points corresponding to multiple target output time steps based on context vectors, the following methods can be used, but are not limited to: inputting the first context vector corresponding to the first output time step into the output layer of the trained temporal super-resolution model for linear transformation and mapping calculation to obtain the oversampled data points corresponding to the first output time step, wherein the first output time step is any one of the multiple target output time steps; combining the oversampled data points corresponding to the multiple target output time steps according to the time order of the multiple target output time steps to form a third performance data sequence of the third frequency; wherein, for each performance index, oversampled data point generation is performed in parallel through an independent output layer to obtain a third performance data sequence containing multi-dimensional performance indexes.

[0075] In the embodiments of this application, for any one of the multiple target output time steps, such as the first output time step, the generation process is as follows: The first context vector calculated corresponding to the first output time step is input into the output layer of the trained temporal super-resolution model. The output layer is the component in the model responsible for generating the final prediction result, and is usually composed of one or more fully connected neural networks, containing trainable weights and bias parameters. This layer performs linear transformation and mapping calculations on the input first context vector. The linear transformation refers to multiplying the input vector by a weight matrix and adding a bias vector. This operation can map high-dimensional context information to the target output dimension. The mapping calculation usually introduces non-linear activation functions (such as Rectified Linear Unit (ReLU), Sigmoid, etc., or not used in some scenarios) on the basis of the linear transformation to enhance the expressive power of the model and map the result to a numerical range that conforms to the actual physical meaning of the performance index (for example, constraining the output between 0 and 1 using the Sigmoid function to represent utilization). After a series of calculations at the output layer, the model generates a performance index value that precisely corresponds to the first output time step, i.e., the oversampled data point of that first output time step. This oversampled data point does not come directly from physical acquisition, but is a high-frequency estimate generated by the model based on a deep understanding and inference of the low-frequency input sequence.

[0076] After generating corresponding oversampled data points for each target output time step, they need to be organized into an ordered monitoring data stream. This is done by combining the data points according to the temporal order of the multiple target output time steps. This means arranging all generated data points from earliest to latest according to their corresponding timestamps, forming a continuous, equally time-interval sequence. The density of data points in this newly constructed sequence (i.e., the third frequency) is much higher than the sampling frequency of the original second performance data, thus forming a complete, high-temporal-resolution third performance data sequence. This third performance data sequence is the final output of the model; it fills the information gaps between low-frequency sampling points and provides a smooth, continuous, and detailed performance change curve.

[0077] Crucially, to maintain clarity and efficiency in the processing logic, the third performance data generation process is performed independently for each performance metric. For each performance metric, the model configures or dynamically invokes an independent output layer (or an independent parameter partition within the output layer). This means that the context vector corresponding to the port traffic metric and the context vector corresponding to the queue depth metric will be processed by different output layer parameters, thereby specifically learning how to map the context information of their respective domains back to their respective domain values. These independent generation processes are executed in parallel, allowing for the simultaneous processing of metric data across all dimensions. Ultimately, the high-frequency sequences generated by all metrics are integrated together to form a third performance data sequence containing multi-dimensional performance metrics, thereby comprehensively and synchronously reconstructing a high-precision status profile of the switch across multiple monitoring dimensions.

[0078] This application employs a dedicated output layer to map each performance metric specifically, ensuring accurate reconstruction of performance metrics with different dimensions and statistical characteristics, thus avoiding prediction bias or pattern confusion that might occur with a shared output layer. Strictly adhering to the chronological order in combining oversampled data points guarantees the temporal consistency and continuity of the output sequence, allowing the reconstructed performance curves to be directly used for timestamp-based analysis, alerting, and visualization, seamlessly compatible with the format and usage of real high-frequency sampled data. Parallel execution improves the overall generation speed from context vectors to the final multi-dimensional performance data sequence, meeting the real-time data processing requirements of high-frequency monitoring scenarios and enabling stable and efficient operation in actual network operation and maintenance systems.

[0079] In one possible implementation of this application, regarding the data processing process of the trained temporal super-resolution model during peak load periods, this application embodiment also provides a schematic diagram of model inference of a switch during peak load periods, such as... Figure 7As shown, during peak business periods, the sampling frequency is proactively reduced, and a pre-trained deep learning oversampling model is used to perform high-precision interpolation on sparse sampling points to reconstruct a complete state sequence equivalent to high-frequency sampling.

[0080] In one possible implementation of this application embodiment, when training a temporal super-resolution model using first performance data, the following methods can be used, but are not limited to: using the first performance data to form a training dataset; inputting the training dataset into the temporal super-resolution model for forward propagation calculation to obtain predicted output data; calculating the loss function value between the predicted output data and the first performance data, and updating the model parameters of the temporal super-resolution model through a backpropagation algorithm until a preset convergence condition is reached to obtain a trained temporal super-resolution model.

[0081] In the embodiments of this application, the training dataset is directly composed of first performance data, which is high-fidelity, high-temporal-resolution data directly collected at a higher first frequency during the low-load period when the switch's service load is below a preset threshold. The first performance data realistically records the subtle changes of the switch under various states, providing a standard answer for the model's learning. When constructing the training dataset, the continuous first performance data stream is typically truncated and divided to form multiple training samples. Each sample may contain a continuous low-frequency subsequence (as simulated input) and a complete high-frequency sequence within the corresponding time period (as the training target), thereby establishing a pairing relationship between input and output. This ensures that the training data and the actual problem the model will solve in the future (i.e., reconstructing high-frequency output from low-frequency input) are completely consistent in data distribution and task objective.

[0082] In each iteration, samples from the prepared training dataset are input in batches into the temporal super-resolution model. The model then performs forward propagation computation. Forward propagation refers to the process where the input data starts from the model's input end and passes through all the pre-defined network layers (such as: temporal feature extraction layer, attention mechanism layer, output layer, etc.), being passed and transformed layer by layer according to the computational rules defined by the network architecture, until the final predicted output data is generated. At this stage, the model processes the input according to its current internal model parameters (including weights and biases) and outputs its prediction results for high-frequency sequences, i.e., the predicted output data.

[0083] Next, the loss function value is calculated between the predicted output data and the first performance data (i.e., the actual high-frequency data). The loss function is a pre-defined function used to quantify the difference or error between the model's predicted values ​​and the true target values. Loss functions include, but are not limited to, Mean Squared Error (MSE) and Mean Absolute Error (MAE), which measure the closeness of the predicted sequence to the true sequence in terms of value and shape from different perspectives. The calculated loss function value is a scalar, and its magnitude directly reflects the inaccuracy of the current model's prediction; the larger the value, the greater the prediction error and the worse the model performance.

[0084] After obtaining the loss value, the model parameters of the temporal super-resolution model are updated using the backpropagation algorithm. Backpropagation is an optimization technique in deep learning. It calculates the gradient of the loss function relative to the parameters of each layer of the model. The gradient indicates the direction and magnitude in which each parameter should be adjusted to reduce the loss value. Starting from the output layer, the algorithm propagates the error gradient layer by layer in the opposite direction to forward propagation, and calculates the gradient of each layer's parameters using the chain rule. Subsequently, gradient descent or its variants are used to make minor adjustments to the model parameters based on the calculated gradient direction. Model parameters are the collective term for all trainable weights and biases in the temporal super-resolution model; they are the concrete carriers of model knowledge. Updating these model parameters allows the model to learn and improve from the current prediction errors.

[0085] The above steps—forward propagation to calculate predictions, loss calculations, and backpropagation to update parameters—constitute a complete training iteration. This process is repeated, allowing the model to continuously learn and correct itself on a large number of training samples. Iteration continues until a preset convergence condition is met. The preset convergence condition is a pre-defined criterion used to determine whether the model has been sufficiently trained. It typically includes: the loss value decreasing to a stable low level and no longer changing significantly (convergence); achieving the expected accuracy on an independent validation dataset; or reaching a preset maximum number of training epochs. When the preset convergence condition is met, the training process terminates, and the resulting model state is the trained temporal super-resolution model. At this point, the model's internal parameters have been sufficiently optimized, enabling it to accurately capture the temporal dynamics of network performance data and reliably perform super-resolution reconstruction tasks from low-frequency inputs to high-frequency outputs.

[0086] This application uses high-frequency performance data collected in real-world scenarios as training data, ensuring that the model learns the true evolution patterns of the network state rather than assumed patterns, thus making the reconstruction results more reliable. Forward propagation and loss calculation provide the model with clear learning objectives and performance metrics. Utilizing the backpropagation algorithm for parameter updates guides the model parameters to evolve in a direction that continuously reduces prediction errors, gradually improving reconstruction accuracy. Finally, pre-defined convergence conditions scientifically control the training process, avoiding undertraining or overfitting, ensuring that the generated model possesses strong generalization ability without excessively relying on noise in the training data. This provides a model foundation for stably generating high-precision third-party performance data during peak business periods.

[0087] In one possible implementation of this application, regarding the training process of the temporal super-resolution model, this application embodiment also provides a schematic diagram of the training of the temporal super-resolution model, such as... Figure 8 As shown, this includes data acquisition, preprocessing, model building, model training, and parameter optimization.

[0088] Regarding the process of model training during off-peak load periods, this application embodiment also provides a schematic diagram of model training for a switch during off-peak load periods, such as... Figure 9 As shown, during periods of low business activity or within a preset training window, data is collected at full rate and high frequency via telemetry. Based on this data, a deep learning oversampling model closely related to port traffic characteristics is trained.

[0089] In one possible implementation of this application, to enable the monitoring system to adapt to business fluctuations and avoid insufficient resource utilization during low-traffic periods or resource contention during high-traffic periods due to fixed collection frequencies, dynamic adaptation of collection frequency switching logic to business load is required. This involves acquiring the business operation status of the monitored switch in real time and comparing it with preset standards to automatically adjust the collection frequency of performance data, ensuring an optimal balance between monitoring needs and resource consumption under different business load scenarios.

[0090] Specifically, but not limited to the following methods may also be used: obtain the service load of the switch to be monitored; in response to the service load being lower than a preset load threshold, collect first performance data at a first frequency; in response to the service load being not lower than the preset load threshold, collect second performance data at a second frequency.

[0091] In the embodiments of this application, service load is a comprehensive quantitative indicator used to characterize the busyness of the data processing plane and control plane of the switch under monitoring. It can be obtained by reading performance counters provided by the switch's operating system, such as: CPU overall utilization, ASIC internal forwarding engine buffer occupancy, and the percentage of instantaneous throughput of high-speed interfaces to port bandwidth. Alternatively, it can be obtained by weighting multiple key performance indicators to obtain a normalized load score.

[0092] Based on the acquired real-time service load information, it is compared with a pre-configured preset load threshold in real time, and corresponding conditional responses are executed. The preset load threshold is a pre-set threshold value that serves as the boundary between two system states: relatively abundant resources (low load period) and relatively scarce resources (high load period).

[0093] When the real-time monitored load value remains consistently below the set threshold, it is determined that the current period is within a model training window that is conducive to in-depth monitoring without affecting the main business. At this time, the high-precision acquisition mode is automatically activated to collect the first performance data at the first frequency. The first frequency, as a higher sampling frequency (e.g., full-rate high-frequency telemetry), can capture the fine ripples and transient events of the network state, providing a rich data source for model training or direct high-precision monitoring.

[0094] Correspondingly, when the real-time load reaches or exceeds the set threshold, it indicates that the data processing pressure on the switch is increasing and system resources are becoming strained. To protect the performance of core forwarding services from interference by monitoring activities, the system immediately switches to energy-saving monitoring mode, collecting secondary performance data at a secondary frequency. The secondary frequency is a sampling frequency significantly lower than the primary frequency (e.g., one-tenth or one-hundredth of the primary frequency). Proactively reducing the sampling frequency directly reduces the CPU computation cycle, memory bandwidth, and internal bus resource consumption of data acquisition, encapsulation, and reporting processes, thereby minimizing the potential impact of monitoring overhead on service forwarding latency and throughput.

[0095] This application tightly couples the behavior of the monitoring system with the actual working state of the switch through a dynamic load awareness and frequency switching mechanism, realizing the transformation from static configuration to dynamic response. This transforms monitoring activities from fixed background overhead into an intelligent agent that senses the environment and adjusts itself accordingly. It creates a resource-sensitive working mode: during off-peak periods, it utilizes idle resources to accumulate high-quality data and knowledge (training the model); during peak periods, it maintains monitoring continuity with minimal necessary overhead and compensates for insufficient data density through intelligent models. This achieves an optimal dynamic balance between ensuring business performance and meeting monitoring accuracy requirements at the system level.

[0096] In one possible implementation of this application, the acquisition process for the second performance data can also be implemented in, but is not limited to, the following manner: Port traffic data of the switch is acquired using a high-speed telemetry HFT system at the F1 frequency (second frequency). This acquisition process strictly follows the IPFIX protocol format and utilizes a DMA engine to directly obtain performance counter data from the switch's ASIC chip. This data is then transmitted to user space via a Netlink module. The user first issues an acquisition task to the switch system, which starts the task according to its configuration and directly accesses the DMA system to write telemetry data. Subsequently, the Linux kernel module assumes the responsibility of transmitting the data to the user-layer counting synchronization process, after which the user process synchronizes the data to the database. The acquired data includes, but is not limited to, key indicators such as port packets per second, port bandwidth utilization, and number of error packets, which have m dimensions. Before being input into the TTSR model, the raw data needs to undergo preprocessing steps, including but not limited to normalization operations.

[0097] In one possible implementation of this application embodiment, when performing real-time monitoring of the switch to be monitored based on the first performance data and / or the third performance data, the following methods can be used, but are not limited to: storing the first performance data and / or the third performance data in a time-series database; reading the first performance data and / or the third performance data from the time-series database for displaying data curves on a preset real-time monitoring interface, or for analyzing abnormal events using a preset fault diagnosis system.

[0098] In the embodiments of this application, both the high-resolution first performance data collected directly during business downtime and the high-resolution third performance data reconstructed through a time-series super-resolution model during business peak times need to be persistently stored for later use. Therefore, a time-series database is used as the storage component. A time-series database is a database optimized for processing time-series data, and it is highly optimized for data arriving in chronological order and indexed by timestamps.

[0099] The system stores the incoming first and / or third performance data, along with their precise timestamps, in a time-series database in real time. This enables long-term data retention, providing the possibility for historical backtracking and trend analysis.

[0100] After data storage, during monitoring, the first performance data and / or the third performance data will be read from the time-series database. This reading process can be performed in real time or near real time as needed.

[0101] The high-precision, high-time-resolution data stream read out should serve at least the following two scenarios: The first scenario is visual monitoring, which involves displaying data curves through a pre-defined real-time monitoring interface. This pre-defined real-time monitoring interface refers to a graphical console. It continuously reads the latest or specified historical performance data by calling a database interface and uses a graphics library to plot it into dynamically updated data curves, such as line charts and area charts. Because the input data frequency is very high (whether it's the first or third performance data point), the plotted curves are exceptionally smooth and detailed, clearly displaying millisecond-level or even microsecond-level fluctuations in metrics such as traffic and queue depth. This allows operations personnel to intuitively perceive subtle changes in network status and promptly detect transient phenomena such as micro-bursts that are difficult to capture with traditional monitoring.

[0102] The second scenario is intelligent analysis, specifically for analyzing abnormal events using a pre-defined fault diagnosis system. This system is an automated analysis engine or algorithm module. It also acquires high-precision performance data streams from a time-series database and applies pre-defined rules, thresholds, or machine learning models to scan the data in real time. Due to the high temporal resolution of the data, it can perform abnormal event analysis, detecting anomaly patterns with extremely short durations and specific amplitudes, such as instantaneous queue congestion, periodic spikes in minor errors, or early signs of link overload. This high-fidelity data provides the foundation for analysis, enabling the fault diagnosis system to issue warnings earlier and more accurately, and even pinpoint the root cause, thereby achieving a shift from a reactive to a proactive prevention-oriented operational model.

[0103] This application utilizes a time-series database for storage, ensuring the smooth and efficient flow of data throughout the system. By supporting high-fidelity data curve displays on a real-time monitoring interface, it enhances the visualization of network status and situational awareness. By providing high-quality data input to the fault diagnosis system, it offers accurate and timely automated anomaly detection and analysis capabilities, effectively shortening the mean time to repair faults and improving the overall reliability and operational efficiency of the data center network.

[0104] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0105] Embodiments of this application also provide a switch monitoring device. Figure 10 This application provides a schematic diagram of the structure of a switch monitoring device, as shown below. Figure 10 As shown, it includes: The acquisition unit 1001 is used to acquire the first performance data of the switch to be monitored at a first frequency during a first time period. Training unit 1002 is used to train the time series super-resolution model using the first performance data. The first time period is the period when the service load of the switch to be monitored is lower than the preset load threshold. The acquisition unit 1001 is also used to acquire second performance data of the switch to be monitored at a second frequency lower than the first frequency during the second time period, wherein the second time period is the period during which the service load of the switch to be monitored is not lower than a preset load threshold. The processing unit 1003 is used to input the second performance data into the trained temporal super-resolution model, perform temporal feature extraction and context information fusion processing, and obtain the third performance data of the third frequency, wherein the third frequency is higher than the second frequency. The monitoring unit 1004 is used to perform real-time monitoring of the switch under monitoring based on the first performance data and / or the third performance data.

[0106] In one embodiment of this application, the processing unit 1003 is specifically used for: The second performance data is encoded into a data matrix. The number of rows in the data matrix is ​​the product of the time interval for collecting the second performance data and the second frequency. The number of columns in the data matrix is ​​the number of types of performance indicators. The second performance data contains multiple types of performance indicators. For each data point in the data column corresponding to each performance index, time-series features are extracted sequentially according to the time step corresponding to each data point to obtain the first feature sequence that reflects the time-series change pattern of the second performance data. Information fusion is performed on the first feature sequence based on the attention mechanism to generate a context vector that reflects the importance information of each data point. Multiple target output time steps are determined based on the time interval corresponding to the third frequency. Oversampled data points corresponding to each of the multiple target output time steps are generated according to the context vector to obtain the third performance data.

[0107] In one embodiment of this application, the processing unit 1003 is specifically used for: The target data point is combined with the historical hidden state passed by the trained temporal super-resolution model to obtain the combined data. The combined data is then used to extract features to obtain the hidden state corresponding to the target data point. The hidden state is used as the historical hidden state for feature extraction of the next data point. Until the temporal features of all data points are extracted, the hidden state sequence consisting of the hidden states corresponding to each data point is determined as the first feature sequence.

[0108] In one embodiment of this application, the processing unit 1003 is specifically used for: The first feature sequence is projected through different preset linear transformation parameters to generate the key vector sequence and value vector sequence corresponding to the first feature sequence. The trained temporal super-resolution model includes a variety of preset linear transformation parameters associated with the attention mechanism. The hidden state is projected into a query vector through a linear transformation. The similarity between the query vector and each key vector in the key vector sequence is calculated separately, and the similarity is normalized to obtain the target attention weight. The target attention weight is used to characterize the importance of the feature corresponding to each data point relative to the generated oversampled data points. The context vector is obtained by weighted summation of each value vector in the value vector sequence based on the target attention weights.

[0109] In one embodiment of this application, the processing unit 1003 is specifically used for: We employ a parallel and independent processing approach, performing time-series feature extraction and contextual information fusion separately for the data columns corresponding to each performance metric.

[0110] In one embodiment of this application, the processing unit 1003 is specifically used for: The first context vector corresponding to the first output time step is input into the output layer of the trained temporal super-resolution model for linear transformation and mapping calculation to obtain the oversampled data points corresponding to the first output time step. The first output time step is any time step among multiple target output time steps. The oversampled data points corresponding to each of the multiple target output time steps are combined according to the time order of the multiple target output time steps to form the third performance data sequence of the third frequency; For each performance metric, oversampling data points are generated in parallel through an independent output layer to obtain a third performance data sequence containing multi-dimensional performance metrics.

[0111] In one embodiment of this application, the training unit 1002 is specifically used for: The training dataset is constructed using the first performance data; The training dataset is input into the temporal super-resolution model for forward propagation calculation to obtain the predicted output data. The loss function value between the predicted output data and the first performance data is calculated, and the model parameters of the temporal super-resolution model are updated through the backpropagation algorithm until the preset convergence condition is met, thus obtaining the trained temporal super-resolution model.

[0112] In one embodiment of this application, the acquisition unit 1001 is further configured to: Obtain the service load of the switch to be monitored; In response to a business load falling below a preset load threshold, first performance data is collected at a first frequency; In response to a business load not falling below a preset load threshold, second performance data is collected at a second frequency.

[0113] In one embodiment of this application, the monitoring unit 1004 is specifically used for: Store the first performance data and / or the third performance data in a time-series database; Read first and / or third performance data from the time-series database for display of data curves on the preset real-time monitoring interface, or for analysis of abnormal events by the preset fault diagnosis system.

[0114] For a description of the features in the embodiment corresponding to the switch monitoring device, please refer to the relevant description of the embodiment corresponding to the switch monitoring method, which will not be repeated here.

[0115] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the switch monitoring method.

[0116] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the switch monitoring method when running.

[0117] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0118] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the switch monitoring method.

[0119] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the switch monitoring method.

[0120] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0121] The present application provides a detailed description of a switch monitoring method and electronic device. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for monitoring a switch, characterized in that, include: Within a first time period, first performance data of the switch to be monitored is collected at a first frequency, and the first performance data is used to train the time series super-resolution model. The first time period is the period when the service load of the switch to be monitored is lower than a preset load threshold. During the second time period, the second performance data of the switch to be monitored is collected at a second frequency lower than the first frequency. The second time period is the period during which the service load of the switch to be monitored is not lower than the preset load threshold. The second performance data is input into the trained temporal super-resolution model, and temporal feature extraction and context information fusion processing are performed to obtain the third performance data of the third frequency, wherein the third frequency is higher than the second frequency; The switch to be monitored is monitored in real time based on the first performance data and / or the third performance data.

2. The switch monitoring method according to claim 1, characterized in that, The step of inputting the second performance data into the trained temporal super-resolution model, performing temporal feature extraction and context information fusion processing to obtain the third performance data of the third frequency includes: The second performance data is encoded into a data matrix, wherein the number of rows in the data matrix is ​​the product of the time interval for collecting the second performance data and the second frequency, and the number of columns in the data matrix is ​​the number of types of performance indicators, and the second performance data contains multiple types of the performance indicators; For each data point in the data column corresponding to each performance index, time-series features are extracted sequentially according to the time step corresponding to each data point to obtain a first feature sequence that reflects the time-series change pattern of the second performance data. Information fusion is performed on the first feature sequence based on an attention mechanism to generate a context vector that reflects the importance information of each data point. Based on the time interval corresponding to the third frequency, multiple target output time steps are determined, and oversampled data points corresponding to each of the multiple target output time steps are generated according to the context vector to obtain the third performance data.

3. The switch monitoring method according to claim 2, characterized in that, For each data point in the data column corresponding to each performance index, the temporal feature extraction is performed sequentially according to the time step corresponding to each data point to obtain a first feature sequence reflecting the temporal change pattern of the second performance data, including: The target data point is combined with the historical hidden state passed by the trained temporal super-resolution model to obtain combined data, and features are extracted from the combined data to obtain the hidden state corresponding to the target data point. The hidden state is used as the historical hidden state for feature extraction of the next data point. Until the temporal features of all data points are extracted, the hidden state sequence formed by the hidden states corresponding to each of the data points is determined as the first feature sequence.

4. The switch monitoring method according to claim 3, characterized in that, The step of fusing information from the first feature sequence based on an attention mechanism to generate a context vector reflecting the importance information of each data point includes: The first feature sequence is projected through different preset linear transformation parameters to generate a key vector sequence and a value vector sequence corresponding to the first feature sequence. The trained temporal super-resolution model includes a variety of preset linear transformation parameters associated with the attention mechanism. The hidden state is projected into a query vector through a linear transformation; The similarity between the query vector and each key vector in the key vector sequence is calculated, and the similarity is normalized to obtain the target attention weight. The target attention weight is used to characterize the importance of the feature corresponding to each data point relative to the generation of the oversampled data point. The context vector is obtained by weighting and summing each value vector in the value vector sequence based on the target attention weight.

5. The switch monitoring method according to claim 3, characterized in that, Using a parallel and independent processing approach, time-series feature extraction and contextual information fusion are performed on the data columns corresponding to each performance index.

6. The switch monitoring method according to claim 2, characterized in that, The step of generating oversampling data points corresponding to each of the multiple target output time steps based on the context vector to obtain the third performance data includes: The first context vector corresponding to the first output time step is input into the output layer of the trained temporal super-resolution model for linear transformation and mapping calculation to obtain the oversampled data point corresponding to the first output time step, wherein the first output time step is any one of the plurality of target output time steps; The oversampling data points corresponding to each of the multiple target output time steps are combined according to the time order of the multiple target output time steps to form the third performance data sequence of the third frequency; For each performance metric, oversampling data points are generated in parallel through an independent output layer to obtain the third performance data sequence containing multi-dimensional performance metrics.

7. The switch monitoring method according to claim 1, characterized in that, The step of training the temporal super-resolution model using the first performance data includes: The training dataset is constructed using the first performance data; The training dataset is input into the temporal super-resolution model for forward propagation calculation to obtain the predicted output data. Calculate the loss function value between the predicted output data and the first performance data, and update the model parameters of the temporal super-resolution model through the backpropagation algorithm until the preset convergence condition is reached, thereby obtaining the trained temporal super-resolution model.

8. The switch monitoring method according to claim 1, characterized in that, The method further includes: Obtain the service load of the switch to be monitored; In response to the service load being lower than the preset load threshold, the first performance data is collected at the first frequency; In response to the service load not being lower than the preset load threshold, the second performance data is collected at the second frequency.

9. The switch monitoring method according to claim 1, characterized in that, The real-time monitoring of the switch to be monitored based on the first performance data and / or the third performance data includes: Store the first performance data and / or the third performance data in a time-series database; The first performance data and / or the third performance data are read from the time-series database for display as data curves on a preset real-time monitoring interface, or for analysis of abnormal events by a preset fault diagnosis system.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the switch monitoring method as described in any one of claims 1 to 9 when executing the computer program.