Upsampling of network traffic trajectories
By pre-processing, denoising and post-processing of network service trajectories, the problem of difficulty in detecting transient high bandwidth requirements in the prior art is solved, and high-precision service-related counter upsampling is realized, and the visibility and accuracy of service quality optimization are improved.
Patent Information
- Application Number
- CN202411699971.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-11-26
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to detect and optimize the transient high bandwidth requirements in network service trajectories, resulting in difficulty in optimizing service quality, and the 5-minute granularity service trajectory cannot effectively capture instantaneous problems.
Through one method, the input value of the service-related counter is obtained using the input rate, preprocessed to generate the scaled value, and applied the trained iterative denoising process to generate the denoising upsampling value, and post-processed to generate the output value, achieving high-precision upsampling of the service-related counter.
It realizes high-precision upsampling of network service trajectories, can effectively capture transient events, improve the visibility and accuracy of service quality optimization, avoid false alarms, and provide a privacy-protected data processing solution.
Smart Images

Figure CN120050202A_ABST
Abstract
Description
Technical Field
[0001] Various example embodiments generally relate to methods and apparatuses for upsampling network traffic traces. Background Art
[0002] Network interfaces (WAN interfaces, gateways, ONTs, OLTs, etc.) can provide traffic-related counters over a period of time. Each traffic-related counter can represent a count of traffic-related parameters measured at one or more points in a data network, such as: the number of bytes (or byte volume), the number of connections, the number of packets, the number of requests, the number of operations, etc.
[0003] Each counter value represents an aggregated amount measured over a period of time (an aggregated amount of bytes, connections, packets, requests, operations, etc.), resulting in average traffic information. Depending on the network interface, this information may be available at a given granularity, i.e., at a given time rate, for example once every 5 minutes.
[0004] Since these data can be the only practical source of usage-related information provided by network devices, due to their privacy protection characteristics, because no additional software or high computational power resources are required, etc., network management / troubleshooting / optimization software typically relies on these data.
[0005] In fact, instantaneous high bandwidth demands may lead to instantaneous problems, so being able to detect such high demands would be a plus for optimizing the quality of service provided among users of a shared medium. However, for traffic traces with a 5-minute granularity, such transient events are typically aggregated out, preventing their detection, and in contrast, even generating false alarms.
[0006] Obtaining traffic-related counters at a finer granularity can be useful. Summary of the Invention
[0007] Independent claims define the scope of protection. Embodiments, examples, and features described in this specification that do not fall within the scope of protection (if any) will be construed as examples to assist in understanding the various embodiments or examples within the scope of protection.
[0008] According to a first aspect, a method includes: obtaining a first time series of input values of a traffic-related counter at an input rate, where the input values are within a first range; preprocessing the first time series of input values to generate a first sequence of scaled values at an output rate equal to the input rate multiplied by an upsampling factor K, where the scaled values are within a second range derived from the first range by a non-linear function; applying a trained iterative denoising process to the first sequence of scaled values stacked with a first noise signal at the output rate to generate a first sequence of denoised upsampled values within the second range at the output rate, where the first noise signal has values within the second range; postprocessing the first sequence of denoised upsampled values to generate a first time series of output values within the first range at the output rate, where each input value in the first time series of input values is equal to the sum of K corresponding consecutive output values.
[0009] Preprocessing the first time series of input values may include: upsampling the first time series of input values by generating K upsampled values from each input value in the input values to generate a first sequence of upsampled values such that the sum of the K upsampled values is equal to the input value under consideration; applying a scaling function to each upsampled value in the upsampled values to generate a first sequence of scaled values within the second range.
[0010] Generating the K upsampled values may include: replacing each input value in the input values with K equal upsampled values, each upsampled value being calculated by dividing the input value under consideration by K.
[0011] The trained iterative denoising process may be based on a denoising diffusion probability model.
[0012] The first sequence of scaled values may be used as a conditioning signal for the trained iterative denoising process.
[0013] The trained iterative denoising process may use an iteratively executed U-net model.
[0014] The method may include: training an iterative denoising process by: obtaining a time series of input training values of a traffic-related counter at a second rate, where the input training values are within a first range; obtaining a sequence of aggregated training values at a first rate within the first range, the aggregated training values corresponding to an aggregated version of the input training values; scaling the input training values to generate a first sequence of scaled training values at a second rate within a second range; preprocessing the sequence of aggregated training values to generate a second sequence of scaled training values at a second rate within the second range; adding a noise signal within the second range to the first sequence of scaled training values to generate a sequence of noisy training values; applying one iteration of the iterative denoising process to the sequence of noisy training values using the first sequence of scaled training values as a target signal and the second sequence of scaled training values as a conditioning signal to generate a sequence of output denoised values; adapting one or more parameters of the iterative denoising process based on a loss function that evaluates the remaining noise in the sequence of output denoised values.
[0015] Post-processing of the first sequence of denoised upsampled values may include: scaling the first sequence of denoised upsampled values to generate a first sequence of scaled upsampled values within the first range; generating a first time series of output values by conditioning the values in the first sequence of scaled upsampled values such that each input value in the first time series of input values is equal to the sum of K corresponding consecutive output values.
[0016] Conditioning the values in the first sequence of scaled upsampled values may include: applying a linear scaling factor to the values in the first sequence of scaled upsampled values, where the linear scaling factor is calculated based on the input value and the sum of K corresponding consecutive output values.
[0017] The method may include: discarding the first P values and the last Q values in the first time series of output values, or signaling that the first P values and the last Q values in the first time series of output values are not trustworthy.
[0018] The method may include: obtaining a second input time series of input values of a service-related counter at an input rate within a second range, wherein each of the first and second series relates to traffic via a respective communication channel on the same physical or logical transmission link; preprocessing the second time series of input values to generate a second series of scaled values at an output rate, wherein the scaled values in the second series of scaled values are within the second range; applying a trained iterative denoising process to the second series of scaled values and a second noise signal at the output rate to generate a second series of denoised upsampled values at the output rate within the second range, wherein the second noise signal has values within the second range, and wherein the trained iterative denoising process is jointly applied to the first series of scaled values, the first noise signal, the second series of scaled values, and the second noise signal; postprocessing the second series of denoised upsampled values to generate a second time series of output values at the output rate within a first range, wherein each input value in the second time series of input values is equal to the sum of K corresponding consecutive output values in the second time series of output values.
[0019] The method may include: performing operations on one or more network devices or network functions based on one or more output values in a first time series of output values.
[0020] According to another aspect, an apparatus includes components for: obtaining a first time series of input values of a service-related counter at an input rate, wherein the input values are within a first range; preprocessing the first time series of input values to generate a first series of scaled values at an output rate equal to the input rate multiplied by an upsampling factor K, wherein the scaled values are within a second range derived from the first range by a non-linear function; applying a trained iterative denoising process to the first series of scaled values stacked with a first noise signal at the output rate to generate a first series of denoised upsampled values at the output rate within the second range, wherein the first noise signal has values within the second range; postprocessing the first series of denoised upsampled values to generate a first time series of output values at the output rate within the first range, wherein each input value in the first time series of input values is equal to the sum of K corresponding consecutive output values.
[0021] The apparatus may include components for performing one or more or all steps of the method according to the first aspect. The components may include at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform one or more or all steps of the method according to the first aspect. The components may include circuitry (e.g., processing circuitry) for performing one or more or all steps of the method according to the first aspect.
[0022] According to another aspect, an apparatus includes at least one processor and at least one memory storing instructions which, when executed by the at least one processor, cause the apparatus to perform: obtaining a first time series of input values of a traffic-related counter at an input rate, wherein the input values are within a first range; preprocessing the first time series of input values to generate a first sequence of scaled values at an output rate equal to the input rate multiplied by an upsampling factor K, wherein the scaled values are within a second range derived from the first range by a non-linear function; applying a trained iterative denoising process to the first sequence of scaled values and a first noise signal at the output rate to generate a first sequence of denoised upsampled values within the second range at the output rate, wherein the first noise signal has values within the second range; and postprocessing the first sequence of denoised upsampled values to generate a first time series of output values within the first range at the output rate, wherein each input value in the first time series of input values is equal to the sum of K corresponding consecutive output values.
[0023] When executed by the at least one processor, the instructions may cause the apparatus to perform one or more or all of the steps of the method according to the first aspect.
[0024] According to another aspect, a computer program includes instructions which, when executed by an apparatus, cause the apparatus to perform: obtaining a first time series of input values of a traffic-related counter at an input rate, wherein the input values are within a first range; preprocessing the first time series of input values to generate a first sequence of scaled values at an output rate equal to the input rate multiplied by an upsampling factor K, wherein the scaled values are within a second range derived from the first range by a non-linear function; applying a trained iterative denoising process to the first sequence of scaled values and a first noise signal at the output rate to generate a first sequence of denoised upsampled values within the second range at the output rate, wherein the first noise signal has values within the second range; and postprocessing the first sequence of denoised upsampled values to generate a first time series of output values within the first range at the output rate, wherein each input value in the first time series of input values is equal to the sum of K corresponding consecutive output values.
[0025] The instructions may cause the apparatus to perform one or more or all of the steps of the method according to the first aspect.
[0026] According to another aspect, a non-transitory computer-readable medium includes program instructions stored thereon that cause a device to at least perform: obtaining a first time series of input values of a traffic-related counter at an input rate, where the input values are within a first range; preprocessing the first time series of input values to generate a first series of scaled values at an output rate equal to the input rate multiplied by an upsampling factor K, where the scaled values are within a second range derived from the first range by a non-linear function; applying a trained iterative denoising process to the first series of scaled values and a first noise signal at the output rate to generate a first series of denoised upsampled values within the second range at the output rate, where the first noise signal has values within the second range; and postprocessing the first series of denoised upsampled values to generate a first time series of output values within the first range at the output rate, where each input value in the first time series of input values is equal to the sum of K corresponding consecutive output values.
[0027] The program instructions may cause the device to perform one or more or all of the steps of the method according to the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Example embodiments will be more fully understood from the following detailed description and the accompanying drawings, which are given by way of illustration only and thus do not limit the present disclosure.
[0029] Figure 1 Shows an input network traffic trace according to an example and an upsampled version of the input network traffic trace obtained by the upsampling method disclosed herein;
[0030] Figure 2 Is a schematic diagram of an upsampling system in the inference phase according to an example;
[0031] Figure 3 Shows the behavior of a compensation function applied by the upsampling method according to an example;
[0032] Figure 4 Shows aspects of the upsampling method according to an example;
[0033] Figure 5 Is a schematic diagram of an upsampling system in the inference phase for upsampling two or more traffic traces according to an example;
[0034] Figures 6A - 6B Shows a curve representing an input sequence according to an example;
[0035] Figures 7A - 7B Shows a curve representing a conditioning signal according to an example;
[0036] Figures 8A - 8B Shows a curve representing an output upsampled signal according to an example;
[0037] Figure 9 illustrates the behavior of a Markov chain according to an example;
[0038] Figure 10 illustrates a noise schedule according to an example
[0039] Figure 11 illustrates aspects of a resampling process according to an example;
[0040] Figure 12 illustrates a high-level view of a U-Net model architecture according to an example;
[0041] Figure 13 is a schematic diagram of an upsampling system during a training phase according to an example;
[0042] Figure 14 illustrates an upsampling traffic counter according to an example;
[0043] Figure 15 illustrates an upsampling traffic counter according to an example;
[0044] Figure 16 is a flowchart showing a method for upsampling network traffic traces according to an example; and
[0045] Figure 17 is a block diagram showing an exemplary hardware structure of a device according to an example.
[0046] It should be noted that these drawings are intended to illustrate various aspects of the devices, methods, and structures used in the example embodiments described herein. The use of like or identical reference numerals in the various drawings is intended to indicate the presence of like or identical elements or features. Detailed Description
[0047] Detailed example embodiments are disclosed herein. However, the specific structural and / or functional details disclosed herein are for the purpose of describing example embodiments only and providing a clear understanding of the basic principles. However, these example embodiments may be practiced without these specific details. These example embodiments may be embodied in many alternative forms and various modifications may be made, and should not be construed as limited to the embodiments described herein. In addition, the drawings and description may have been simplified to illustrate elements and / or aspects relevant to a clear understanding of the present invention, while eliminating many other elements well known in the art or irrelevant to an understanding of the present invention for the sake of clarity.
[0048] A method that only requires aggregating traffic-related data as input artificially enhances and enriches the input low-rate (aggregated) traffic-related data using machine learning-based methods to create traffic-related data that could have been obtained at a higher rate.
[0049] Machine learning (ML) is an application that enables computer systems to perform tasks by reasoning based on patterns found in data analysis without being explicitly programmed. Machine learning explores the research and construction of algorithms (also referred to as tools in this article) that can learn from existing data and make predictions about new data. Such machine learning algorithms operate by building an ML model from example training data, generating data-driven predictions as output.
[0050] An intelligent upsampling system is disclosed that utilizes ML based on a denoising diffusion probability model and a process with network domain-specific training.
[0051] The denoising diffusion probability model is trained such that the upsampling system is adapted to upsample aggregated traffic counters from one rate to another, higher rate. The denoising diffusion probability model embeds specific network domain knowledge through a specific training strategy and is thus able to upsample low-rate traffic-related data in an accurate, reliable, and realistic manner to generate higher-rate traffic-related data.
[0052] The main breakthrough lies in the ability to provide such capabilities with privacy-protected data (i.e., without the need for per-service data, per-device data, per-application data, etc.), without the need for packet inspection or retrieving more data, thereby providing better transient event visibility and enabling better subsequent processing and decision-making. This provides a solution for data production and / or data collection constraints.
[0053] Other advantages include the ability to work in real-time mode (causal) or non-causal mode (processing "future" data), and to work regardless of the granularity of the input data. The method is compatible with any type of network (optical network, wireless network, etc.) and is not linked to a specific network type or transmission link (passive optical network, radio network, copper link, etc.), any communication technology, or any communication protocol.
[0054] The method provides output higher-granularity traffic-related values suitable for any subsequent decision-making process, such as for targeted troubleshooting or network optimization, network configuration, etc. For example, accurate service type identification and quantification can be performed without the privacy and scalability issues of DPI.
[0055] The method constrains the output higher-granularity traffic-related values such that after being aggregated back, they exactly match the input lower-granularity counter traffic-related values.
[0056] The method preprocesses the range and dynamic characteristics of the input values to leverage the performance of the ML model.
[0057] In one or more embodiments, traffic-related input data from different channels (e.g., upstream and downstream channels, optical channels) can be jointly processed by a denoising diffusion probability model (e.g., using strong tensors), with the effect of providing a unified insight to the model and limiting the differences between different outputs.
[0058] Due to domain-driven selection, the hyperparameter space can be simplified for exploration to find the best model configuration within the minimal amount of resources involved.
[0059] The upsampling method is flexible because there are no restrictions on the upsampling factor and / or the granularity of the input data.
[0060] Even when there is no technical limitation on the length of the input data sequence, the upsampling method can predict an accurate upsampled data sequence from a short input data sequence, such as an input data sequence as low as 24 counter values.
[0061] Using the upsampling method disclosed herein, a sequence of 40 counters of traffic data within an aggregated 240-second period can be upsampled 8 times to generate a sequence of 320 counters of traffic data within an aggregated 30-second period.
[0062] According to another example, a sequence of 24 counters of traffic data within an aggregated 300-second period is upsampled 10 times to generate a sequence of 240 counters of traffic data within an aggregated 30-second period.
[0063] These examples are not exhaustive, and the upsampling method disclosed herein can handle more combinations of input granularity, sequence length, and upsampling factor.
[0064] Figure 1 High-granularity traffic data (C2) is shown compared to corresponding lower-granularity / aggregated traffic data (curve C1).
[0065] Classical simple interpolation / upsampling techniques or general ML models would not be able to generate curve C2 from curve C1. The upsampling method disclosed herein allows the generation of values corresponding to curve C2 with the help of curve C1 (e.g., curve C1 is used as a conditioning signal). This enables actual subsequent operations, decision-making processes, troubleshooting analysis, configuration tasks, and optimization tasks based on curve C1.
[0066] The upsampling system disclosed herein relies on a supervised deep learning model that performs denoising. Training of such a model requires training data that includes a large amount of first input data (“fine-grained” data) at a high rate. The training set may include second input data (“coarse-grained” data) at a lower target rate. The low-rate second input data can be obtained by aggregating the first input data at a high rate with an aggregation factor corresponding to the ratio between the high rate and the low rate. The aggregation factor is equal to the upsampling factor applied by the upsampling system.
[0067] Since the upsampling method is flexible and can handle input data with various upsampling factors, a training data set including first input data with very fine granularity can be used to train denoising models for these various upsampling factors.
[0068] In fact, by aggregating the fine-grained data with various aggregation factors, a training data set including coarse-grained data can be generated at various rates. Using the training data including fine-grained data and coarse-grained data, a model can be trained to produce fine-grained data that can provide coarse-grained data.
[0069] The upsampling system is applicable to any upsampling factor. The upsampling factor can be equal to 2 or greater. Several upsampling factors, such as upsampling factors between 10 and 100, can be used in the training phase to tune the denoising model for any upsampling factor.
[0070] In the inference phase, the upsampling system is configured to provide high-rate enhanced data (“fine-grained” data) from low-rate input data (“coarse-grained” data).
[0071] Figure 2 is a schematic diagram of the upsampling system 200 in the inference phase according to an example.
[0072] This figure represents an upsampling method in the inference phase that generates an output sequence Y (i.e., a fine-grained traffic trace) at a second rate given an input sequence X at a first rate (i.e., a coarse-grained traffic trace). Each traffic trace corresponds to a sequence of values of traffic-related parameters, such as a sequence of values of a traffic-related counter.
[0073] The input sequence is submitted to a preprocessing step before being used as an adjustment in a loop of denoising steps performed by an iterative denoising process. Then, the denoised sequence at the output of the denoising process is post-processed to provide an output sequence with well-scaled and upsampled traffic-related counter values with fine granularity.
[0074] The upsampling system includes: a pre-processing block 210, a stacking block 215, a denoising block 220, and a post-processing block 230. The pre-processing block 210 includes two sub-blocks: an up-sampling sub-block 211 and a scaling sub-block 212, and the order of the two sub-blocks 211 and 212 can be reversed. The post-processing block 230 includes two sub-blocks: a compensation sub-block 232 and a scaling sub-block 231, and the order of the two sub-blocks 231 and 232 can be reversed.
[0075] By applying preprocessing functions, including an upsampling function applied by upsampling block 211 and a scaling function applied by scaling block 212 , the input sequence X may be converted into an upsampled sequence S, which is an upsampled and scaled version of the input sequence X.
[0076] The starting point of the denoising process is Gaussian noise (e.g., perfectly isotropic Gaussian noise) N t , which is stacked with the upsampled sequence S by the stacking block 215 to generate a stacked sequence Z t .
[0077] Scaling function 212 is a nonlinear scaling function that performs range adaptation. The role of scaling function 212 is to provide scaling values in the same range as Gaussian noise, and to enable the dynamic characteristics of the signal to take into account small changes in, for example, the input sequence X. Since the scaling function applies a nonlinear transformation, the difference between large values tends to decrease, the difference between small values tends to increase, and the gap between large and small values also decreases. This prevents any characteristic of the signal (small value, large value) from taking priority or being ignored during processing as much as possible. Therefore, dynamic range adaptation is performed in this way, where dynamic range refers to the ratio of the maximum measurable signal to the minimum measurable signal (e.g., in dB).
[0078] The denoising block 220 is configured to apply a denoising process and includes a denoising model that is iteratively applied multiple times to generate a denoised sequence Z 0 The denoising block 220 is configured to generate a Gaussian noise N t Gaussian noise is detected and removed with the help of a conditioning signal in the form of a stacked upsampled sequence S. The upsampled sequence S is used as a conditioning signal for a denoising model to produce realistic fine-grained traffic trajectories. The denoising model can be a diffusion probabilistic denoising model (e.g., a U-Net model). The Gaussian noise has the same dimension as the output sequence.
[0079] By applying a post-processing function, including a scaling function applied by the scaling block 231 and a compensation function applied by the compensation sub-block 232, the denoised sequence Z 0 Convert to output sequence Y.
[0080] The upsampling block 211 can apply various upsampling functions to generate an upsampled sequence U (before scaling). The upsampling function applied by block 211 can be a function that adds zero values between two input values of the input sequence X. The upsampling function applied by block 211 can be a function that repeats each value of the input sequence X K times, where K is the upsampling factor. The upsampling function applied by block 211 can be a convolutional filter applied to the input sequence X to generate a sequence U that includes inter-sample filtered values between two input values.
[0081] The upsampling function is configured to generate an upsampled sequence U such that after the values of the upsampled sequence U are aggregated back by an aggregation factor K, the traffic represented by the values of the upsampled sequence U corresponds to the traffic represented by the values of the input sequence X.
[0082] For example, if the input sequence X includes 48 values and the upsampling factor is K = 10, the upsampling function 211 generates an upsampled sequence U of 480 values by copying each value of the input sequence X 10 times. To avoid a range shift of the traffic trajectory represented by the input values by repeating the same input value K times, each value of the upsampled sequence U is divided by the upsampling factor K, such that after the values of the upsampled sequence U are aggregated back by an aggregation factor K, the traffic represented by the values of the upsampled sequence U corresponds to the traffic represented by the values of the input sequence X.
[0083] The scaling function 212 is configured to generate a scaled value of the upsampled sequence S within the same range as the Gaussian noise for each input value of the upsampled sequence U, and to adapt the dynamic characteristics of the signal in order to be able to take into account, for example, small variations in the input sequence X.
[0084] For example, if the denoising model 220 uses Gaussian noise with a mean of 0 and a variance of 1 as input, the scaled values of the upsampled sequence S are scaled to the same range [-1, 1]. In addition, the dynamic range is also adapted by applying a non-linear scaling function S1 that takes into account the maximum input value.
[0085] For example, if the input values in the input sequence X represent a traffic expression in kilobits and the maximum bandwidth is 10 Gb / s, where the traffic is reported every 5 minutes, the maximum input value is:
[0086] X max =(5*60)[s]10 7 [kbps]=3*10 9 [kb]
[0087] To linearly scale the input kilobit value x within the initial range [0, X max to the noise range (e.g., [-1, 1]), the following logarithmic scaling function S1(x) can be used:
[0088] S1(x) = (log10(x + 1) / g) - 1 (1)
[0089] where the coefficient g can be g = 5 and is determined according to X max to be determined.
[0090] Although Figure 2 it is shown that the scaling function S1 is applied after applying the upsampling function in block 211, the order of these two functions can be reversed, i.e., the scaling function is applied first and then the upsampling function.
[0091] By stacking block 215, a vector including the values of the scaled upsampling sequence S is stacked with a vector including the values of the noise sequence N t to generate a stacked vector. The stacked vector Z t is provided as the input to the denoising model 220 to generate the iterative denoising sequences Z t-1 、Z t-2 and so on. After a given number of iterations, the final denoising sequence D = Z 0 is obtained for the upsampling sequence S.
[0092] After generating the denoising sequence D = Z 0 from the upsampling sequence S, post - processing is applied to the denoising sequence D having values within a noise range (e.g., [-1, 1]). The post - processing includes a scaling function S2 applied by block 231 and a compensation function applied by block 232. The scaling function S2 is the inverse function of the scaling function S1 applied by block 212 and is applied to each value of the denoising sequence D to generate an upsampled sequence P of denoised values that are within the same range [0, X max as the values of the input sequence X and have the same dynamic characteristics.
[0093] For example, if the values of the upsampled and scaled sequence S are within the range [-1, 1], the output values of the denoising sequence Z 0 produced by the model are also within the same range, and these values are scaled back to the input range by applying the inverse scaling function S2 231, for example, to retrieve the values in the k - bit space. This is achieved by inverting the function S1 defined by equation (1):
[0094] S2(z) = 10 5*(z+1) - 1
[0095] The compensation function applied by block 232 is applied for business envelope conservation, with the aim that the output sequence Y matches the coarse - grained input sequence X after being aggregated back by the upsampling factor K. This means that the output sequence Y exactly fits the signal envelope given by the input sequence X.
[0096] Figure 3 Illustrates the compensation function applied by sub - block 232 according to the example.
[0097] After the output of the model has been scaled back, the denoised up - sampled values of sequence P are aggregated back and compared with the corresponding input values in input sequence X: In each group of K consecutive values, the up - sampled values of sequence P are summed to calculate a compensation factor to be applied to each value in the group. The compensation factor is calculated such that after being applied to each value in the group, the sum of these compensated values equals the corresponding input value available in input sequence X.
[0098] If we note the up - sampled values of group p i (where i = 1 to K), and if we note the corresponding input value A of the group, then the compensated value y i can be calculated by the following formula:
[0099]
[0100] such that ∑y i = A.
[0101] Figure 4 Shows the MSE between the target signal and the predicted up - sampled signal generated by the denoising model during the training phase on a logarithmic scale (before scaling and compensation).
[0102] As Figure 4 shown, the start part 410 and end part 420 of the signal include significantly higher error values. In this example, the denoising model is less accurate for the start part 410 and end part 420 than when the denoising model is applied to the middle of the sequence. This can be easily explained by the fact that the denoising model uses the surrounding context of the samples to better generate the possible true denoised sequence. But for the start and end of the sequence, the denoising model may not have information about what happened before and what will happen next. This missing information cannot be utilized, and thus the accuracy of the denoising model drops slightly in these start and end regions, resulting in a higher loss that may be detected during the training phase.
[0103] To compensate for this edge effect, the denoising model can be trained with the entire sequence length, but the post - processing block 230 can discard the start and end parts of each denoised sequence by only retaining the central part corresponding to the most reliable part of the denoised sequence. During the inference phase, the start and end parts of each denoised sequence D (or output sequence Y) can be discarded, and / or it can be signaled that the start and end parts of each denoised sequence are not trustworthy (e.g., signaling which entities of the output will be used).
[0104] For example, the first M values and the last Q values in the denoised sequence D or the output sequence Y can be discarded and / or signaled as untrustworthy, where the number of values of M and Q is determined based on an error or loss value determined during the training phase. For example, M (and the corresponding Q) can correspond to the length of the starting portion (and the corresponding ending portion) within which, during the training phase, the MSE or loss value (e.g., average) is higher than a threshold.
[0105] Figure 5 FIG. 500 is a schematic diagram of an upsampling system 500 for upsampling two or more network traffic traces during the inference phase according to an example.
[0106] Each network traffic trace can correspond to a sequence of values of traffic-related parameters, such as a sequence of values of a traffic-related counter. Two input sequences XA, XB can represent network traffic traces during the same time period. This example can be generalized to any number of input sequences.
[0107] The example upsampling system 500 allows for the joint upsampling of two input sequences XA, XB by applying the same processing chain (i.e., preprocessing, denoising, and postprocessing) to each input sequence, where the two input sequences XA, XB represent traffic traces over corresponding channels on the same physical or logical transmission link, and where the denoising process is performed jointly on the two input sequences, as will be described in detail below. During inference, the upsampling system generates two output upsampled sequences.
[0108] For example, the input sequence XA represents a traffic trace on the downstream channel of a transmission link, and XB represents a traffic trace on the upstream channel of the same transmission link. For example, the input sequence XA represents a traffic trace on a first optical channel (with a first central wavelength), and XB represents a traffic trace on a second optical channel (with a second central wavelength).
[0109] The upsampling system can include (i) two preprocessing blocks (e.g., one preprocessing block for each input sequence) to generate two corresponding preprocessed sequences from the two input sequences XA, XB respectively, and (ii) a stacking block configured to stack the two preprocessed sequences, where the preprocessing applied to the input sequences XA, XB can be the same for each input sequence XA, XB.
[0110] Alternatively, as Figure 5 shown, the upsampling system can include a stacking block 505 configured to stack two input sequences XA, XB to generate a stacked input sequence X, followed by a preprocessing block 510 for generating a preprocessed sequence S from the stacked input sequence X.
[0111] As Figure 5As shown, the upsampling system may include a second stacking block 515, a denoising block 520, a separation block 225, and two post-processing blocks 530A and 530B.
[0112] The preprocessing block 510 is configured to generate a scaled upsampling sequence S from the stacked input sequence X.
[0113] Reference Figure 2 The description of the preprocessing block 210 applies to Figure 5 the preprocessing block 510. As Figure 2 shown, the preprocessing block 510 includes two sub-blocks: an upsampling sub-block and a scaling sub-block, which are not shown here for simplicity. Figure 5 in
[0114] The upsampling function performed by the upsampling sub-block of the preprocessing block 510 may be one of the upsampling functions described with reference to Figure 2 the preprocessing block 210.
[0115] The scaling function performed by the scaling sub-block of the preprocessing block 510 may be one of the scaling functions described with reference to Figure 2 the preprocessing block 210.
[0116] The stacking block 215 is configured to generate a stacked vector by stacking the following:
[0117] - A first vector that includes the values of the preprocessing sequence S,
[0118] - A second vector that includes the values of the noise sequence N t .
[0119] Reference Figure 2 The description of the denoising block 220 applies to Figure 5 the denoising block 520. The stacked vector Z t is provided as the input to the denoising model 520 to generate iterative denoising sequences Z t-1 , Z t-2 and so on. After a given number of iterations, a final denoising sequence D is obtained for the upsampling sequence S. This allows denoising to be performed jointly on the two sequences.
[0120] The separation block 525 is configured to: separate the final denoising sequence into two denoising sequences: a first denoising sequence DA corresponding to the transformation of the first vector SA, and a second denoising sequence DB corresponding to the transformation of the second vector SB.
[0121] The post-processing block 530A (corresponding 530B) is configured to: generate an output sequence YA (corresponding YB) from the denoising sequence DA (corresponding DB).
[0122] Reference Figure 2The description of the post - processing block 230 applies to each of the pre - processing blocks 530A, 530B Figure 5 as shown in Figure 2 . Each of the post - processing blocks 530A, 530B includes two sub - blocks: a compensation sub - block and a scaling sub - block, which for simplicity reasons are not shown in Figure 5 .
[0123] The scaling function performed by the scaling sub - blocks of the post - processing blocks 530A, 530B can be one of the scaling functions described with reference to Figure 2 the post - processing block 230. The scaling function performed by the scaling sub - block of the post - processing block 530A (corresponding 530B) is the inverse function of the scaling function performed by the scaling sub - block of the pre - processing block 510A (corresponding 510B). For example, the denoising sequences DA, DB can be on a logarithmic scale and they are scaled in a linear space.
[0124] The compensation function performed by the compensation sub - blocks of the post - processing blocks 530A, 530B can be the same for each of the post - processing blocks 530A, 530B and can be the compensation function described with reference to Figure 2 the post - processing block 230. The compensation function generates the output sequence YA or YB.
[0125] These two channels can correspond to an upstream channel and a downstream channel. By jointly processing the upstream and downstream traffic trajectories into the same tensor, the following advantages can be obtained.
[0126] The model can construct a stronger global representation of the traffic trajectories with a single latent representation, such as both the upstream and downstream traffic trajectories. Having such a single representation allows the model to sample more realistic pairs of upstream and downstream trajectories because they belong to a single point in the model hyperspace.
[0127] The model can utilize the correlation between the upstream adjustment signal and the downstream adjustment signal to enhance the information contained in a single stream, which greatly helps to identify the noise added to x 0 .
[0128] Figure 6A (corresponding 6B) shows the curves representing the input sequences XA (corresponding XB), i.e., the change of the traffic - related counters over time according to the example.
[0129] Figure 6A The change of the aggregated 5 - minute byte counter value of the upstream traffic corresponding to the communication link over 4 hours, and Figure 6B the change of the aggregated 5 - minute byte counter value of the downstream traffic corresponding to the same communication link over the same time period.
[0130] Figure 7A(Corresponding 7B) shows the adjustment signal SA (corresponding SB) obtained after preprocessing the input sequence XA (corresponding XB).
[0131] Figure 8A (Corresponding 8B) shows the output upsampled signal YA (corresponding YB) obtained after postprocessing the denoised signal DA (corresponding DB). These figures illustrate the ability of the system to generate an upsampled output signal with low-width high peaks from a more stepped adjustment signal.
[0132] In one or more embodiments, a generative diffusion model is used for denoising, such as using a denoising diffusion probability model. A diffusion model or probability diffusion model is a parameterized Markov chain trained using variational inference, which is used to generate samples that match the input data after a finite time.
[0133] A diffusion model is a deep generative model that works by adding noise (e.g., Gaussian noise) to the available training data (also known as the forward diffusion process) and then reversing the process (called denoising or reverse diffusion process) to recover the data. The model gradually learns to remove the noise. This learned denoising process generates new high-quality signals from random noise signals.
[0134] Figure 9 Illustrates the behavior of a Markov chain trained to iteratively denoise the signal x T to generate the denoised signal x 0 of.
[0135] As Figure 9 shown, by applying an iterative (reverse diffusion) denoising process, using an adjustment signal composed of aggregated traffic data, network traffic data can be synthesized from random noise.
[0136] Forward diffusion process
[0137] To train the model to remove noise from a noisy signal, a noisy traffic signal is generated through the forward diffusion process. The principle is to gradually add noise to the target low-granularity traffic signal (x 0 ) based on a noise schedule. After T noise addition steps, it is assumed that the initial traffic signal will be completely destroyed, and the resulting noisy traffic signal x T can be regarded as isotropic Gaussian noise.
[0138] The noise addition Markov chain can be represented as:
[0139]
[0140] where Denotes sampling of \(x\) from a Gaussian distribution with mean \(\mu\) and variance \(\sigma\). The noise schedule \(\beta\) can be chosen such that the initial information contained in \(x\) 0 is not destroyed too quickly. An interesting property of a Markov chain is that the next state of the chain depends only on the previous state in the chain, not on the states before that. Due to this property, there is a direct way to compute \(x_{t}\) 0 directly from \(x_{0}\) t without going through all the intermediate steps (from \(0\) to \(t\)): \(x_{t}\) t can be sampled directly from \(x_{0}\) 0 as follows:
[0141]
[0142] where
[0143] \(\alpha_{t}\) t \(:= 1 - \beta_{t}\) t
[0144] and
[0145]
[0146] and \(I\) is the identity matrix.
[0147] With this notation, the noise schedule \(\beta_{t}\) can be represented relative to the forward noise time step \(t\) as shown in Figure 10 The factor represents the amount of Gaussian noise added to the initial fine-grained business signal \(x_{0}\) along the noise step \(t\). 0
[0148] With the above reparameterization, the sample \(x_{t}\) t can be represented as a function of \(x_{0}\) using equation (3a): 0
[0149]
[0150] where where \(\epsilon\) is a realization of \(\mathcal{N}(0, I)\), i.e., a Gaussian distribution with mean \(0\) and variance \(I\).
[0151] Reverse diffusion process
[0152] At any given step \(t\), a noisy business trajectory can be constructed with a predefined noise schedule. The deep learning model is trained to be able to extract the added noise at a specific step \(t\). In other words, we need to find the \(\theta\) parameters of the distribution \(p_{\theta}\) such that \(x_{t}\) can be sampled from \(x_{0}\) t t-1 . To simplify parameterization, the model can utilize partial knowledge of x by relying on an adjustment signal, which represents a preprocessed aggregation of the input signal x with a factor equal to the upsampling factor K. 0 When the Gaussian noise added between steps is small enough, the distribution function p 0 can also be regarded as a Gaussian function and can be expressed as:
[0153] where μ θ is the mean of the distribution to be predicted, and ∑
[0154]
[0155] is its variance. θ θ θ Using the distribution function p
[0156] Sampling process
[0157] , an inverse Markov chain can be created starting from isotropic Gaussian noise x θ , and x T is recursively sampled, followed by x T-1 , then x T-2 , x T-3 , …, until the denoised signal x 0 . This is called the sampling process.
[0158] An interesting property of the diffusion model is that the steps of the sampling process can be fewer than the range of steps used in the training process.
[0159] Figure 11 illustrates aspects of the resampling process according to an example.
[0160] For example, the optimal number of training steps T can be 1000. This means that during training, x t is sampled at a randomly selected t between 0 and 1000. Then, the goal of the model is to recover the noisy step t by predicting the noise added for that specific step.
[0161] During the sampling process, the diffusion model is flexible enough to have fewer sampling steps than training steps. For example, there can be T2 sampling steps to sample from x T to x 0 . To reduce the sampling steps from T to T2, we use T2 equally spaced real numbers between 1 and T (including the endpoints), and then round each result to the nearest integer.
[0162] For example, if we want to sample x 0 in 10 steps while the model is trained with 1000 steps, we can start from isotropic Gaussian noise x1000 and sample x 999 Start, and then take sample x 999 as x 889 and pass it to the model, and then pass it to sample x 888 and so on until x is generated 0 .
[0163] Figure 11 This resampling step concept is visually supported by an example of 10 steps out of 1000 steps.
[0164] Training objective
[0165] The main goal of the denoising model is to generate x in the fine-grained business signal space 0 . Since x 0 is generated by the distribution p θ , the training can be carried out by optimizing the common variational bound on the negative log-likelihood:
[0166]
[0167] Since x 0 depends on the chain x T , …, x 0 , the log-likelihood is intractable and further simplifications can be used. By means of the variational lower bound and Bayes' equality, an upper bound of (4) can be written, which will be minimized by the model with parameters θ, as follows:
[0168] L vlb :=L 0 +D KL (q(x t-1 |x t , x 0 )||p θ (x t-1 |x t ))+L T (5)
[0169] where D KL is the Kullback-Leibler (KL) divergence, where q is defined by equation (3a). The terms L 0 and L T can be neglected and can be ignored to minimize the error. For more details, see the example reference
[01] .
[0170] Equation (5) uses the q posterior and no longer uses the q prior. In further development of the mathematical derivation, we can obtain:
[0171]
[0172] where ∈ θ is a function approximator that is used to predict ∈ from x t and ∈, ∈ θ is then the prediction of ∈.
[0173] One simplification for obtaining equation (6) is to set the p θ variance (∑ θ ) to the same schedule as q (i.e., β t ).
[0174] This simple objective L simple gives reasonably good results, but the accuracy may degrade rapidly when the number of resampling steps is reduced.
[0175] The model can also be trained by using another objective L hybrid :
[0176] L hybrid = L simple + λL vlb
[0177] where L vlb is mainly the Kullback-Leibler (KL) divergence (equation (5)) between q(x t-1 |x t , x 0 ) and p θ (x t-1 |x t ).
[0178] The KL divergence measures the difference between two probability distributions over the same variable x, and if these two distributions are Gaussian, its KL divergence can be easily computed using their respective means and variances.
[0179] It can be shown that the q posterior mean can be analytically derived from x 0 , x t and , all of which are known during the forward process, while the q posterior variance is simply a combination of .
[0180] For the mean and variance of p θ , it can also be shown that the mean can be derived from ∈ θ (x t , t), while the variance should be predicted by the model because L simple does not provide a learning signal for ∑ θ (x t , t).
[0181] As disclosed in
[01] , the mean μθ (x t , t) can be derived from ∈ θ (x t , t) based on Equation (7):
[0182]
[0183] According to Reference
[02] , it can be proven that ∑ θ (x t , t) has an upper bound β and a lower bound where β t is the variance of the prior q, is the variance of the posterior q.
[0184] It can be shown that the reasonable range of Σ θ (x t , t) is very small, so it is difficult for a neural network to directly predict ∑ θ (x t , t), even in the logarithmic domain. Instead, it is better to parameterize the variance as an interpolation between β t and in the logarithmic domain.
[0185] Specifically, the model outputs a vector v, which contains a component for each dimension, and we convert this output to variance as follows:
[0186]
[0187] Therefore, we have all the means and variances to calculate the L vlb terms, and finally calculate the loss L hybrid of our model.
[0188] Experiments were conducted to evaluate the optimal number of sampling steps for upsampling applied to business trajectories.
[0189] In the case where the model objective is only L simple , the number of sampling steps needs to be almost as high as the number of training steps to achieve the best model accuracy.
[0190] In the case where the model objective is L hybrid , the number of sampling steps can be greatly reduced without affecting the sampling accuracy. It was found that the optimal number of training steps and sampling steps are 1000 and 50, respectively. This balance between the number of training steps and the number of sampling steps can achieve significant results within a limited inference processing time.
[0191] Figure 12 Shows a high-level view of the U-Net model architecture according to the example.
[0192] The U-Net model can be used as a denoising model in each iteration of the iterative denoising process. The U-Net model can include some additional attention layers. In Figure 12 the example of
[0193] Input x t and Cond are the target traffic sequences changed by the noise and conditioning signal (i.e., the coarse-grained traffic counter) added by the forward process at step t. Before entering the U-Net model, these inputs are stacked by channel (the upstream (us) and downstream (ds) channels are stacked in x t and the conditioning signal Cond respectively). The U-Net model can mainly include ResNet blocks (with or without attention layers), followed by a downsampling layer (convolutional layer) or an upsampling layer (nearest interpolation layer).
[0194] The output of the model includes 4 channels, representing the noise (∈ θ,us ; ∈ θ,ds ) added to the forward passes of the upstream and downstream respectively, and the interpolation factors (v us ; v ds ) of the upstream and downstream respectively.
[0195] These outputs are used to calculate Σ θ (x t , t) and μ θ (x t , t) according to Equation (8) and Equation (7) respectively, for both the upstream (where ∈ θ = ∈ θ,us and v θ = v us ) and the downstream (where ∈ θ = ∈ θ,ds and v θ = v ds ).
[0196] Then Σ θ (x t , t) and μ θ (x t , t) are used to calculate the output denoising signal D, which corresponds to x t-1
[0197] x t-1 = μ θ (x t , t) + ∑ θ (x t , t)z
[0198] Among them
[0199] If t > 0, then Otherwise z = 0
[0200] Training strategy
[0201] Figure 13 is a schematic diagram of the upsampling system in the training phase according to the example, where L vlb is taken as ∈ θ and v θ and the distribution of the added Gaussian noise is calculated as a function of. See Example
[02] .
[0202] As shown by the theory behind the diffusion model described herein, the model is trained to learn a representation of real business trajectories, so it can start from isotropic Gaussian noise and, with the help of some conditions, generate business trajectories belonging to the same space as the real trajectories. To construct such a representation, the model largely relies on the business sequence provided as x 0 The training dataset provides examples in the real business trajectory space, which cover all regions of this space as much as possible in a balanced manner.
[0203] To train the diffusion model, hundreds of thousands of long business sequences (one week) can be generated with low-granularity aggregation (30 seconds) upstream and downstream. The long sequences of several days allow for various non-business patterns, covering all human business usage behaviors (weekdays, evenings, nights, weekends).
[0204] Then, the training dataset is constructed by extracting smaller sequences from these long sequences based on randomly smaller time windows (e.g., 6h). Since these smaller sequences are low-granularity aggregations, they can be used in preprocessing to create x 0 sequences for the forward diffusion process. Then, for each x 0 , x t for a random t ∈ [1, T] is generated. Based on these targets, ∈ and v can be derived.
[0205] The conditioning signal can also be constructed from these low-granularity sequences by simply aggregating their values with the upsampling factor. After this is done, the preprocessing will scale the conditioning in the range [-1, 1] by applying the same scaling function S1 as in the inference phase (see Figure 2 ).
[0206] As Figure 13 shown by the training system of, a sequence Y of input training values of business-related counters is obtained at a second rate, where the second rate is equal to the first rate multiplied by the upsampling factor K. The input training values are within the first range.
[0207] Obtain a corresponding sequence X of aggregated training values at a first rate within a first range, where the aggregated training values correspond to an aggregated version of input training values. The sequence X of aggregated training values can be obtained by aggregating (block 1305) a sequence Y of input training values.
[0208] Block 1310 scales the input training values in the sequence Y of input training values at a second rate within a second range to generate a first sequence YS of scaled training values. The scaling function of block 1310 is the same as the scaling function S1 used in the inference phase (e.g., see Figure 2 block 212).
[0209] Block 1320 preprocesses the sequence X of aggregated training values at a second rate within a second range to generate a second sequence XS of scaled training values. The preprocessing function of block 1320 includes an upsampling function (block 1321) and a scaling function (block 1322). The preprocessing function of block 1320 is the same as the preprocessing function used in the inference phase (e.g., see Figure 2 block 210), and uses the same scaling function as the scaling function applied by block 1310 to the sequence Y of input training values (applied by block 1322). Thus, the scaling function of block 1322 is also the same as the scaling function S1 used in the inference phase (e.g., see Figure 2 block 212).
[0210] Add a noise signal N within the second range t to the first sequence YS of scaled training values to generate a first sequence YN of noisy training values.
[0211] Block 1340 stacks a vector including the values of the first sequence YN with a vector including the values of the second sequence XS of scaled training values to generate a second sequence Z of noisy training values in a stacked vector Z.
[0212] Block 1350 applies one iteration of an iterative denoising process (e.g., a U-Net model) to the stacked vector Z to generate a sequence D of output denoised values. Here, the first sequence YN of noisy training values is stacked with the second sequence XS of scaled training values (XS is used as a conditioning signal) and provided as the input to the U-Net to find ∈ that minimizes a loss function θ and v θ .
[0213] Block 1390 calculates a loss based on the added noise N t , and the output ∈ θ and v θOne or more parameters of the U-Net are adapted to minimize the loss function. The loss evaluates the distance between two probability distributions, and the training process aims to minimize this distance. The loss can evaluate the divergence between the q posterior and p, and is the sum of the terms defined by equations (5) and (6). The loss can be calculated per training epoch.
[0214] During the training epoch, the added noise signal N is applied to the signal YS at random step t using equation (3b) driven by the schedule β t and then the base model is trained using this noisy signal YN to denoise the signal YN and predict the signal at step t - 1. The output of the U-net is ∈ θ and v θ , which is used to calculate the output denoised signal D, which corresponds to x t-1 .
[0215] During the training epoch, the added noise signal N varies by selecting random steps t in the range [0, T] to train the base model with more or less noise, and then the base model is used in each iteration of the iterative denoising process during the inference phase from step t = T to t = 0. t to add different levels of noise.
[0216] Results and performance
[0217] The results of upsampling the 5-minute traffic counter to the 30-second traffic counter are as Figure 14 and Figure 15 shown. These examples are from the validation dataset that the model has never processed during training.
[0218] These figures illustrate that the system disclosed herein can generate fine-grained patterns that cannot be generated by any traditional arithmetic means with a coarse-grained pattern as input.
[0219] In Figure 14 and Figure 15 , the model upsamples both upstream (us) and downstream (ds) simultaneously. For each stream, the figure shows the target signal ( Figure 13 Y in), the prediction obtained at the output of the denoising process ( Figure 13 Z-1 in), and the naive prediction corresponding to the conditioning signal generated using sample-and-hold interpolation ( Figure 13 U in).
[0220] Figure 14 shows examples of upsampled traffic counters from 5-minute aggregated traffic to 30-second aggregated traffic, and magnifies the subset corresponding to samples 500 to 800. This figure illustrates that the system moves from a more stepped conditioning signal (Figure 14 the ability to generate an upsampled signal including a single wide peak near sample 550 in
[0221] Figure 15 Another example of the upsampling service counter from 5 - minute aggregated traffic to 30 - second aggregated traffic is shown, and a subset corresponding to samples 500 to 800 is magnified. This example illustrates the system's ability to generate an upsampled signal including a double peak from a conditioning signal with a single platform: see the portion between samples 560 and 570.
[0222] Furthermore, even though it may not be visible on the chart, due to the intelligent post - processing of the model, the predicted 10 - fold aggregation exactly matches the target aggregation.
[0223] Figure 14 Another interesting example of the upsampling system's capabilities is shown, where a diffusion - based upsampling system generates some fine - grained patterns that are not obvious in the aggregated traffic traces.
[0224] The main use case directly relates to the ability to provide a better view and thus detect better transient phenomena occurring in network traffic data. If this traffic data does not have a sufficiently fine granularity, events affecting QoS / QoE tend to be averaged out, or worst of all, tend to generate false alarms. This prevents any subsequent processing from reliably detecting these events and thus acting efficiently or troubleshooting these links. Using the current domain - embedding model, a lot of network - specific knowledge utilizes the internal context information present in low - granularity traces to propose an enhanced upsampled version of the traffic trace that is very close to the underlying ground truth and presents network - specific patterns. This provides a level of detail visibility that is more realistic than classical mathematical upsampling techniques, enabling actual troubleshooting or optimization tasks this time.
[0225] Another direct use case exists in the area of efficient data storage / compression. In fact, there are challenges in retrieving and storing highly fine - grained data in a scalable manner across the network and over long periods of time. Additionally, the amount of TBs to be stored would be huge. In this sense, through the present invention, we provide a solution for continuously aggregating long - term / historical data that can reliably resample it when high - granularity data is needed. Compared to compression, this still allows direct processing of the aggregated data (for some tasks that do not require highly fine - grained data) because this data is still meaningful (rather than compressed in Zip format as the compressed version is no longer understandable).
[0226] Figure 16The figure shows a flowchart of a method for upsampling network traffic traces according to one or more example embodiments. The steps of the method may be implemented by a network device according to any example described herein.
[0227] Although these steps are described in a sequential manner, those skilled in the art will understand that some steps may be omitted, combined, performed in a different order, and / or performed in parallel.
[0228] In step 1610, a first time series of input values of traffic-related counters is obtained at an input rate. The input values are within a first range;
[0229] In step 1620, the first time series of input values is preprocessed to generate a first sequence of scaled values at an output rate equal to the input rate multiplied by an upsampling factor K. The scaled values are within a second range;
[0230] In step 1630, a trained iterative denoising process is applied to the first sequence of scaled values and a first noise signal at the output rate to generate a first sequence of denoised upsampled values within the second range at the output rate. The first noise signal has values within the second range.
[0231] In step 1640, the first sequence of denoised upsampled values is postprocessed to generate a first time series of output values within the first range at the output rate. Each input value in the first time series of input values is equal to the sum of K corresponding consecutive output values.
[0232] In step 1650, one or more operations may be performed on one or more network devices and / or network functions based on one or more output values of the output sequence.
[0233] (Multiple) operations may depend on the context and / or scenario and / or network environment and / or the type of traffic-related counters to be monitored. One or more operations may include at least one of configuration operations, resource management operations, monitoring operations, channel estimation, optimization operations, repair operations, maintenance operations, restart, reboot, software updates, signaling operations, etc.
[0234] Those skilled in the art should understand that any functional, engine, block diagram, flowchart, state transition diagram, process chart, and / or data structure representation described herein represents an illustrative circuit traditional conceptual diagram embodying the principles of the present invention. Similarly, it should be understood that any flowchart, flowchart, state transition diagram, pseudocode, etc. represents various processes.
[0235] Although a flowchart may depict a set of steps as a sequential process, many of the steps may be performed in parallel, concurrently, or simultaneously. Additionally, some steps may be omitted, combined, or performed in a different order. When the steps of a process are complete, the process may terminate, but there may also be additional steps not disclosed in the drawings or the specification.
[0236] Each process, function, engine, block, step described herein may be implemented in hardware, software, firmware, middleware, microcode, or any suitable combination thereof.
[0237] When implemented in software, firmware, middleware, or microcode, the instructions for performing the necessary tasks may be stored in a computer-readable medium, which may or may not be included in a host device or host system configured to execute the instructions. The instructions may be transmitted via the computer-readable medium and loaded onto the host device or host system. The instructions are configured to cause the host device or host system to perform one or more functions disclosed herein. For example, as described above, according to one or more examples, at least one memory may include or store instructions, and the at least one memory and the instructions may be configured to, together with at least one processor, cause the host device or host system to perform one or more functions.
[0238] Figure 17 An example embodiment of an apparatus 9000 is shown as an example host device or host system. The apparatus 9000 may be used to perform one or more or all of the steps of the methods disclosed herein. The apparatus 9000 may be configured to perform one or more functions of the upsampling system disclosed herein.
[0239] The apparatus may be a general-purpose computer, a special-purpose computer, a programmable processing device, a machine, etc. The apparatus may be or include any of the following or be a part thereof: a user device, a client device, a mobile phone, a laptop computer, a computer, a network element, a data server, a network resource controller, a network device, a router, a gateway, a network node, a computer, a cloud-based server, a web server, an application server, a proxy server, etc.
[0240] As schematically shown, the apparatus 9000 may include at least one processor 9010 and at least one memory 9020. The apparatus 9000 may include one or more communication interfaces 9040 (e.g., network interfaces for accessing wired / wireless networks, including Ethernet interfaces, WIFI interfaces, etc.), which are connected to the processor and configured to communicate via (a plurality of) wired / non-wired communication links. The apparatus 9000 may include a user interface 9030 (e.g., keyboard, mouse, display screen, etc.) connected to the processor. The apparatus 9000 may also include one or more media drives 9050 for reading computer-readable storage media (e.g., digital storage disks 9060 (CD-ROM, DVD, Blu-ray, etc.), USB keys 9080, etc.). The processor 9010 is connected to each of the other components 9020, 9030, 9040, 9050 to control their operations.
[0241] The memory 9020 may be or include random access memory (RAM), cache memory, non-volatile memory, backup memory (e.g., programmable or flash memory), read-only memory (ROM), hard disk drive (HDD), solid state drive (SSD), or any combination thereof. The ROM of the memory 9020 may be configured to store one or more computer program codes of the operating system of the apparatus 9000 and / or one or more software applications. The RAM of the memory 9020 may be used by the processor 9010 for temporarily storing data.
[0242] The processor 9010 may be configured to store, read, load, execute, and / or otherwise process instructions 9070 stored in the computer-readable storage media 9060, 9080, and / or the memory 9020 such that when the instructions are executed by the processor, the apparatus 9000 performs one or more or all of the steps of the methods described herein for the relevant apparatus 9000.
[0243] The instructions may correspond to program instructions or computer program codes. The instructions may include one or more code segments. The code segments may represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. By passing and / or receiving information, data, arguments, parameters, or memory contents, the code segments may be coupled to another code segment or hardware circuit. The information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable technique, including memory sharing, message passing, token passing, network transmission, etc.
[0244] When provided by a processor, a function may be provided by a single dedicated processor, a single shared processor, or multiple individual processors (some of which may be shared). The term "processor" should not be construed as referring only to hardware capable of executing software and may implicitly include one or more processing circuits, whether programmable or not. A processor or similar processing circuit may correspond to a digital signal processor (DSP), a network processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a system on a chip (SoC), a central processing unit (CPU), an arithmetic logic unit (ALU), a programmable logic unit (PLU), a processing core, programmable logic, a microprocessor, a controller, a microcontroller, a microcomputer, a quantum processor, or any device capable of responding to and / or executing instructions in a defined manner and / or according to defined logic.
[0245] A computer-readable medium or computer-readable storage medium may be any tangible storage medium suitable for storing computer or processor-readable instructions. A computer-readable medium may more generally be any storage medium capable of storing and / or containing and / or carrying instructions and / or data. A computer-readable medium may be a non-transitory computer-readable medium. The term "non-transitory" as used herein is a limitation on the medium itself (i.e., tangible, rather than a signal), rather than a limitation on the persistence of data storage (e.g., RAM vs. ROM).
[0246] A computer-readable medium may be a portable or fixed storage medium. A computer-readable medium may include one or more storage devices such as a permanent mass storage device, a magnetic storage medium, an optical storage medium, a digital storage disk (CD-ROM, DVD, Blu-ray, etc.), a USB key or dongle or peripheral, a memory suitable for storing computer or processor-readable instructions.
[0247] A memory suitable for storing computer or processor-readable instructions may be, for example: read-only memory (ROM), a permanent mass storage device such as a disk drive, a hard disk drive (HDD), a solid state drive (SSD), a memory card, core memory, flash memory, or any combination thereof.
[0248] The phrase "a component configured to perform one or more functions" or "a component for performing one or more functions" may correspond to one or more functional blocks including circuitry adapted to perform or configured to perform the relevant function(s). The block may perform the function itself or may cooperate and / or communicate with one or more other blocks to perform the function. A "component" may correspond to or be implemented as "one or more modules", "one or more devices", "one or more units", etc.
[0249] The component may include at least one processor and at least one memory, the memory including at least one memory storing instructions which, when executed by the at least one processor, cause the device to perform the contemplated (s) function. As an alternative or in combination, the component may include circuitry (e.g., processing circuitry) configured to perform the contemplated (s) function.
[0250] As used in this application, the term "circuitry" may refer to one or more or all of the following:
[0251] (a) only hardware circuit implementations (e.g., implementations in only analog and / or digital circuitry) and
[0252] (b) combinations of hardware circuits and software, such as, where applicable: (i) combinations of (s) analog and / or digital hardware circuits with software / firmware, and (ii) any portions of (s) hardware processors (including (s) digital signal processors), software, and (s) memories that work together to cause a device such as a mobile phone or a server to perform various functions); and
[0253] (c) (s) hardware circuits and / or (s) processors, such as (s) microprocessors or portions of (s) microprocessors, which require software (e.g., firmware) to operate, but where the software may be absent when not needed to operate.
[0254] This definition of circuitry applies to all uses of the term in this application, including in any claims. As another example, as used in this application, the term circuitry also encompasses implementations of only hardware circuits or processors (or processors) or portions of hardware circuits or processors and their attendant software and / or firmware. For example, if applicable to a particular claim element, the term circuitry also includes integrated circuits for network elements or network nodes or any other computing or network device.
[0255] The term circuitry may encompass digital signal processor (DSP) hardware, network processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. For example, circuitry may be or include hardware, programmable logic, programmable processors that execute software or firmware, and / or any combination thereof (e.g., processors, control units / entities, controllers) for executing instructions or software and controlling the transmission and reception of signals, and memories for storing data and / or instructions.
[0256] The circuitry can also make decisions or determinations, generate frames, packets, or messages for transmission, decode received frames or messages for further processing, and perform other tasks or functions described herein. The circuitry can control the transmission of signals or messages over a radio network and can control the reception of signals or messages, etc., via one or more communication networks.
[0257] Although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present disclosure, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0258] The terms used herein are only for describing specific embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" used herein also include the plural forms. It should also be understood that the terms "comprises", "comprising", "includes", and / or "including" when used herein specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0259] While aspects of the present disclosure have been specifically shown and described with reference to the above embodiments, those skilled in the art will understand that various additional embodiments can be contemplated by modifying the disclosed machines, systems, and methods without departing from the scope of the disclosed subject matter. These embodiments should be understood to fall within the scope of the present disclosure as determined based on the claims and any equivalents thereof.
[0260] List of Main Abbreviations
[0261] AI: Artificial Intelligence
[0262] DDPM: Denoising Diffusion Probability Model
[0263] DPI: Deep Packet Inspection
[0264] KL: Kullback-Leibler
[0265] MAE: Mean Absolute Error
[0266] Mbbp: Megabits per Second
[0267] MAE: Machine Learning
[0268] MSE: Mean Squared Error
[0269] OLT: Optical Line Terminal
[0270] ONT: Optical Network Terminal
[0271] PON: Passive Optical Network
[0272] QoE: Quality of Experience
[0273] QoS: Quality of Service
[0274] VLB: Variational Lower Bound
[0275] WAN: Wide Area Network
[0276] XGS PON: 10 Gigabit Symmetric Passive Optical Network
[0277] Cite References
[0278]
[01] J.Ho, A.Jain and P.Abbeel. ”Denoising diffusion probabilistic models”, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada; arXiv:2006.11239.
[0279]
[02] A.Nichol and P.Dhariwal, ”Improved denoising diffusion probabilistic models”, February 18, 2021; arXiv:2102.09672
Claims
1. A method for communication, comprising: Upsampling a network traffic trace representing a change in a traffic-related counter over time based on: acquiring a first time series of input values of the traffic-related counter at an input rate, wherein the input values are within a first range; preprocessing the first time series of input values to generate a first sequence of scaled values at an output rate equal to the input rate multiplied by an upsampling factor K, wherein the scaled values are within a second range derived from the first range by a nonlinear function; applying a trained iterative denoising process to the first sequence of scaled values stacked with a first noise signal at the output rate to generate a first sequence of denoised upsampled values at the output rate within the second range, wherein the first noise signal has values within the second range; The first sequence of denoised upsampled values is post-processed to generate a first time series of output values at the output rate within the first range, wherein each input value in the first time series of input values is equal to a sum of K corresponding consecutive output values.
2. The method of claim 1 , wherein preprocessing the first time series of input values comprises: upsampling the first time series of input values by generating K upsampled values from each of the input values to generate a first sequence of upsampled values such that a sum of the K upsampled values is equal to the input value under consideration; A scaling function is applied to each of the upsampled values to generate a first sequence of the scaled values within the second range.
3. The method of claim 2, wherein generating the K upsampled values comprises: Each of the input values is replaced by K equal upsampled values, each upsampled value being calculated by dividing the input value under consideration by K. The method of claim 1 , wherein the trained iterative denoising process is based on a denoising diffusion probability model. 5 . The method of claim 1 , wherein the first sequence of scaled values is used as a conditioning signal for the trained iterative denoising process.
6. The method of claim 1, wherein the trained iterative denoising process uses an iteratively executed U-net model.
7. The method according to claim 1, comprising: The iterative denoising process is trained by: acquiring a time series of input training values of the service-related counter at the second rate, wherein the input training values are within the first range; acquiring a sequence of aggregated training values within the first range at the first rate, the aggregated training values corresponding to aggregated versions of the input training values; scaling the input training values to generate a first sequence of scaled training values within the second range at the second rate; preprocessing the sequence of aggregated training values to generate a second sequence of scaled training values at the second rate within the second range; adding a noise signal within the second range to the first sequence of scaled training values to generate a noisy training value sequence; applying one iteration of the iterative denoising process to the sequence of noisy training values using the first sequence of scaled training values as a target signal and using the second sequence of scaled training values as a conditioning signal to generate a sequence of output denoised values; One or more parameters of the iterative denoising process are adapted based on a loss function evaluating the residual noise in the sequence of output denoised values.
8. The method of claim 1 , wherein post-processing of the first sequence of denoised upsampled values comprises: scaling the first sequence of denoised upsampled values to generate a first sequence of scaled upsampled values within the first range; The first time series of output values is generated by adjusting values in the first sequence of scaled upsampled values such that each input value in the first time series of input values is equal to a sum of K corresponding consecutive output values.
9. The method of claim 8, wherein adjusting the values in the first sequence of scaled upsampled values comprises: A linear scaling factor is applied to values in the first sequence of scaled upsampled values, wherein the linear scaling factor is calculated based on the input value, and the sum of the K corresponding consecutive output values.
10. The method according to claim 1, comprising: The first P values and the last Q values in the first time series of the output values are discarded, or it is signaled that the first P values and the last Q values in the first time series of the output values are unreliable.
11. The method according to claim 1, comprising: acquiring a second time series of input values of the traffic-related counter at the input rate within the second range, wherein each of the first series and the second series relates to traffic via a corresponding communication channel on the same physical or logical transmission link; preprocessing the second time series of input values to generate a second sequence of scaled values at the output rate, wherein the scaled values in the second sequence of scaled values are within the second range; applying a trained iterative denoising process to the second sequence of scaled values and a second noise signal at the output rate to generate a second sequence of denoised upsampled values at the output rate within the second range, wherein the second noise signal has values within the second range, wherein the trained iterative denoising process is jointly applied to the first sequence of scaled values, the first noise signal, the second sequence of scaled values, and the second noise signal; The second sequence of denoised upsampled values is post-processed to generate a second time series of output values at the output rate within the first range, wherein each input value in the second time series of input values is equal to a sum of K corresponding consecutive output values in the second time series of output values.
12. The method according to claim 1, comprising: An operation is performed on one or more network devices or network functions based on one or more output values in the first time series of output values.
13. An apparatus for use in a communication network, comprising means for: Upsampling a network traffic trace representing a change in a traffic-related counter over time based on: acquiring a first time series of input values of the traffic-related counter at an input rate, wherein the input values are within a first range; preprocessing the first time series of input values to generate a first sequence of scaled values at an output rate equal to the input rate multiplied by an upsampling factor K, wherein the scaled values are within a second range derived from the first range by a nonlinear function; applying a trained iterative denoising process to the first sequence of scaled values stacked with a first noise signal at the output rate to generate a first sequence of denoised upsampled values at the output rate within the second range, wherein the first noise signal has values within the second range; The first sequence of denoised upsampled values is post-processed to generate a first time series of output values at the output rate within the first range, wherein each input value in the first time series of input values is equal to a sum of K corresponding consecutive output values.