Dynamic adaptive in-band network telemetering arrangement method and device

The dynamic self-adaptive in-band network telemetry method optimizes scheduling through diffusion models and reinforcement learning to balance measurement accuracy and network stability, addressing inefficiencies in existing INT technologies.

CN120321111AInactive Publication Date: 2025-07-15BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510813366.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing in-band network telemetry technology has problems such as excessive resource overhead, difficult to balance measurement accuracy and network stability in dynamic network environments, and lacks general applicability in complex dynamic network scenarios.

Method used

The dynamic adaptive in-band network telemetry orchestration method is adopted, and the target random optimization model is established by obtaining network structure data, combining diffusion model and reinforcement learning to generate INT orchestration scheme, balance measurement accuracy and network stability, and build a dynamic network model to adaptively adjust the orchestration strategy.

Benefits of technology

In a dynamic network environment, the balance between measurement accuracy and network stability is achieved, the efficiency and applicability of the telemetry system are improved, and the changes in complex network environments are adapted to.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321111A_ABST
    Figure CN120321111A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic adaptive in-band network telemetering arrangement method and device, and relates to the technical field of network monitoring, the method comprises the following steps: obtaining network structure data of a network to be telemetered, the network structure data comprising a network node set and a network link set; according to the network structure data, establishing a target stochastic optimization model used for describing the dynamic network environment under a plurality of time slots; the target stochastic optimization model is obtained based on an INT arrangement model, a measurement accuracy model and a network stability model of the network to be telemetered running under a time slot structure; according to the target stochastic optimization model, generating a target INT arrangement scheme through a diffusion model and reinforcement learning; wherein the reverse denoising process in the diffusion model is a strategy function in reinforcement learning. According to the invention, the balance between the measurement accuracy and the network stability in a dynamic network environment is realized, and the efficiency of a telemetering system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network monitoring, and in particular, to a dynamic adaptive in-band network telemetry orchestration method and apparatus. Background Art

[0002] In-band network telemetry (INT) is an emerging network measurement technology that can provide real-time, fine-grained network status information for various network management and control tasks (such as traffic engineering, fault diagnosis, security analysis, etc.), meeting the visualization needs of network administrators for end-to-end transmission information of network flows. The implementation of INT technology relies on programmable data plane devices. Specifically, as Figure 1 shown, the switches in the figure are all programmable data plane switches. Each switch includes multiple telemetry items, which are indicated by different shapes (such as squares, regular hexagons, and circles) to represent different telemetry items, and different colors indicate different switches; when a data packet is forwarded to a programmable switch, the programmable switch matches the pre-coded INT instruction in the data packet header, encodes the corresponding switch metadata into the INT header and inserts it into the data packet. When the data packet is forwarded to the last hop (i.e., switch 3 in the figure), the programmable switch assembles the INT header and the INT instruction in the data packet into an INT report and sends it to the data analyzer, while the original data packet content is forwarded to the destination host. Among them, the data packet format is ETH (Ethernet), IP (Internet Protocol), TCP / UDP (Transmission Control Protocol / User Datagram Protocol), INT instruction, and the inserted INT header (such as INT1, INT2... INTn).

[0003] Although network measurement using INT technology can obtain real-time, fine-grained, and rich network status information, without restrictions, INT technology will consume too much network bandwidth and the overall resources of the telemetry system. Moreover, as the network scale expands, the overhead brought by INT will become even greater.

[0004] At present, the research on reducing the extra overhead of INT and improving the overall efficiency of the telemetry system can be mainly divided into two categories: telemetry mechanism optimization and multi-telemetry task scheduling optimization. However, the INT technical solutions designed for telemetry mechanism optimization do not jointly consider the telemetry information requirements of different management applications in the network, which may lead to problems such as redundant data collection and single-point performance bottlenecks in the scenario of parallel multi-network management applications, and it is difficult to further improve the overall efficiency of the telemetry system. At present, the INT solutions for multi-telemetry task scheduling optimization mainly build models around aspects such as telemetry coverage, measurement accuracy, and bandwidth overhead, but most of them ignore the dynamics of the network. On the other hand, in a dynamic network, many network parameters (such as available link bandwidth, application flow rate, etc.) are time-varying. To ensure measurement accuracy and efficiency, the scheduling scheme needs to be adjusted according to network changes. However, frequently adjusting the telemetry scheduling scheme will lead to a decline in network stability. Moreover, the current telemetry scheduling algorithms are mostly designed for specific optimization objectives and specific system models, lacking generality in complex dynamic network scenarios, and it is difficult to adjust the telemetry scheduling strategy in a timely manner according to the network environment, which in turn leads to a decline in measurement accuracy and a reduction in the overall efficiency of the system. Summary of the Invention

[0005] The purpose of the present invention is to provide a dynamic adaptive in-band network telemetry scheduling method and device to achieve the balance between measurement accuracy and network stability in the network telemetry of a dynamic network environment and improve the efficiency of the telemetry system.

[0006] In a first aspect, the present invention provides a dynamic adaptive in-band network telemetry scheduling method, including: Obtain the network structure data of the network to be telemetered, where the network structure data includes a set of network nodes and a set of network links; According to the network structure data, establish a target stochastic optimization model for multiple time slots describing the dynamic network environment; the target stochastic optimization model is obtained based on the INT scheduling model, measurement accuracy model, and network stability model of the network to be telemetered operating under a time slot structure; According to the target stochastic optimization model, generate a target INT scheduling scheme through a diffusion model and reinforcement learning; where the reverse denoising process in the diffusion model is the policy function in reinforcement learning.

[0007] In an optional embodiment, according to the network structure data, establishing a target stochastic optimization model for multiple time slots describing the dynamic network environment includes: According to the network structure data, establish an INT scheduling model, a measurement accuracy model, and a network stability model of the network to be telemetered operating under a time slot structure; Based on the INT scheduling model, measurement accuracy model, and network stability model, construct an initial stochastic optimization model; The initial stochastic optimization model is decoupled from the stochastic optimization problem by using Lyapunov optimization technology to obtain the target stochastic optimization model.

[0008] In an alternative embodiment, the scheduling scheme for time slot t in the INT scheduling model is expressed as: ; In the measurement accuracy model, within time slot t, the scheduling scheme the total information gain obtained is expressed as: ; In the network stability model, the change in the scheduling scheme between different time slots is expressed as: ; The initial stochastic optimization model is expressed as: ; where, indicates whether flow f collects telemetry item i at switch v in time slot t, represents the set of time slots, represents the set of network nodes, , represents the set of telemetry items that can be collected within switch v, , represents the set of traffic flows in the network to be telemetered, , represents the information gain obtained when telemetry item i at switch v is collected by traffic flow in time slot t, represents the number of bytes occupied when telemetry item i is collected by traffic flow, represents the telemetry item capacity that traffic flow f can carry within time slot t; represents the desired long-term network stability limit.

[0009] In an alternative embodiment, the target stochastic optimization model is expressed as: ; where, represents the queue backlog of the virtual queue used to characterize the change in the scheduling scheme at time slot t, , represents the change in the scheduling scheme between different time slots, represents the desired long-term network stability limit, represents a non-negative control parameter, represents within time slot t, the scheduling scheme the total information gain obtained, Indicates whether the flow f collects the telemetry item i at the switch v in the time slot t. Represents the set of traffic flows in the network to be telemetered. , Represents the set of time slots. Represents the set of network nodes. , Represents the set of telemetry items that can be collected within the switch v. , Represents the number of bytes occupied when the telemetry item i is collected by the traffic flow. Represents the telemetry item capacity that the traffic flow f can carry within the time slot t.

[0010] In an optional embodiment, according to the target stochastic optimization model, a target INT orchestration scheme is generated through a diffusion model and reinforcement learning, including: Construct an agent network, which includes a policy network and an evaluation network. The policy network is used to output the current action according to the current environmental state and the policy function, and the evaluation network is used to calculate the value function according to the current environmental state and the current action. The current environmental state includes the information gain of the telemetry item collection observed in the current time slot, the telemetry item capacity that the traffic flow can carry within the time slot, and the queue backlog of the virtual queue. The current action is the orchestration scheme for the current time slot. Obtain the experience sequence obtained by the agent through the interaction between the initialized agent network and the network environment of the network to be telemetered, and store the experience sequence in the buffer; wherein, the network environment is used to feedback the current reward according to the current action and the reward function corresponding to the target stochastic optimization model, and update the current environmental state to obtain the environmental state of the next time slot; the experience sequence includes the current environmental state, the current action, the current reward, and the environmental state of the next time slot under multiple time slots. Sample the experience sequence in the buffer to obtain training samples. Train the agent network according to the training samples and the objective function corresponding to the target stochastic optimization model to obtain the trained agent network. Determine the current actions in each time slot output by the policy network in the trained agent network as the target INT orchestration scheme.

[0011] In an optional embodiment, the current action is calculated by the agent through the policy network according to the current environmental state, the policy function, and the action exploration parameter. The action exploration parameter is used to add noise to the action, and the action exploration parameter is updated based on the diffusion policy entropy estimated by using the Gaussian mixture model; train the agent network according to the training samples and the objective function corresponding to the target stochastic optimization model to obtain the trained agent network, including: Calculate the value function through the evaluation network according to the current environmental state and the current action in the training samples. Update the network parameters and action exploration parameters of the agent network through the backpropagation algorithm according to the value function, the current environmental state and current reward in the training samples, and the objective function.

[0012] In an alternative embodiment, the evaluation network includes two Q-value networks and their target networks; the objective function of the policy network is: ; The objective function of the evaluation network is: ; The reward function r is: ; Wherein, represents the network parameters of the policy network, represents the environmental state, represents the buffer, represents the final action generated through the reverse denoising process, represents the policy function of the policy network under the network parameters and the environmental state , represents the smaller of the value functions output by the two Q-value networks, represents the network parameters of the Q-value network, represents the next environmental state of, represents the environmental state , action under the reward, represents the discount factor, , represent the value functions output by the two Q-value networks, , represent the value functions output by the target networks of the two Q-value networks, represents the action obtained by inputting into the policy network, represents a non-negative control parameter, represents the total information gain obtained by the scheduling scheme in time slot t, represents the queue backlog representation of the virtual queue at time slot t, represents the change in the scheduling scheme between different time slots, represents a non-negative penalty factor, represents whether flow f collects telemetry item i at switch v in time slot t, , , , Indicates the number of bytes occupied when telemetry item i is collected by the service flow. Indicates the telemetry item capacity that service flow f can carry within time slot t.

[0013] In a second aspect, the present invention provides a dynamic adaptive in-band network telemetry orchestration device, including: An acquisition module, configured to acquire network structure data of the network to be telemetered, where the network structure data includes a network node set and a network link set; A building module, configured to build a target stochastic optimization model for multiple time slots describing a dynamic network environment according to the network structure data; the target stochastic optimization model is obtained based on an INT orchestration model, a measurement accuracy model, and a network stability model when the network to be telemetered operates in a time slot structure; A generation module, configured to generate a target INT orchestration scheme according to the target stochastic optimization model through a diffusion model and reinforcement learning; wherein, the reverse denoising process in the diffusion model is a policy function in reinforcement learning.

[0014] In a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, it implements the dynamic adaptive in-band network telemetry orchestration method in any one of the foregoing embodiments.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the dynamic adaptive in-band network telemetry orchestration method in any one of the foregoing embodiments.

[0016] The dynamic adaptive in-band network telemetry orchestration method and device provided by the present invention can acquire network structure data of the network to be telemetered, where the network structure data includes a network node set and a network link set; according to the network structure data, build a target stochastic optimization model for multiple time slots describing a dynamic network environment; the target stochastic optimization model is obtained based on an INT orchestration model, a measurement accuracy model, and a network stability model when the network to be telemetered operates in a time slot structure; according to the target stochastic optimization model, generate a target INT orchestration scheme through a diffusion model and reinforcement learning; wherein, the reverse denoising process in the diffusion model is a policy function in reinforcement learning. In this way, in the face of a dynamic network environment, when constructing the target stochastic optimization model, the relationship between balancing measurement accuracy and network stability is fully considered, realizing the balance between measurement accuracy and network stability in the network telemetry of a dynamic network environment; and a diffusion model and reinforcement learning are combined to generate a target INT orchestration scheme, which has good applicability and generality in a complex network environment and improves the efficiency of the telemetry system. Description of the Drawings

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is the working principle diagram of INT; Figure 2 It is a schematic flowchart of a dynamic adaptive in-band network telemetry orchestration method provided by an embodiment of the present invention; Figure 3 It is a network structure model provided by an embodiment of the present invention; Figure 4 It is an example of an INT orchestration scheme provided by an embodiment of the present invention; Figure 5 It is a flowchart of an adaptive INT orchestration algorithm provided by an embodiment of the present invention; Figure 6 It is a schematic structural diagram of a dynamic adaptive in-band network telemetry orchestration device provided by an embodiment of the present invention; Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiments

[0019] The following will clearly and completely describe the technical solutions of the present invention in combination with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0020] Currently, the research on reducing the additional overhead of INT and improving the overall efficiency of the telemetry system can be mainly divided into two categories: telemetry mechanism optimization and multi-telemetry task orchestration optimization.

[0021] Optimization of Telemetry Mechanism: The INT solution optimized for the telemetry mechanism mainly starts from the working principle of INT to reduce the additional overhead brought by INT, which can be further divided into filtering of network events, sampling mechanism, and encoding aggregation method. Filtering of network events refers to extracting and filtering key network information during the upload of INT report raw data, thereby reducing CPU usage and the storage cost of status information; the sampling mechanism refers to adjusting the proportion of monitored data packets according to the frequency of significant changes in network information during telemetry to reduce the total amount of INT network status data collected; the encoding aggregation method refers to encoding the requested network status data onto multiple data packets for collection and performing packet-by-packet aggregation, static flow-by-flow aggregation, and dynamic flow-by-flow aggregation according to different application requirements to reduce the overhead of INT for each data packet.

[0022] Optimization of Multi-Telemetry Task Orchestration: The INT solution optimized for multi-telemetry task orchestration mainly considers the parallel scenario of multiple network management applications. The telemetry information requirements of different network management applications are not exactly the same. For example, congestion control focuses on switch port utilization, queue occupancy, etc.; load balancing focuses on device numbers, link utilization, etc.; network fault tracing mainly focuses on the processing time within the switch. Therefore, the main goal of multi-telemetry task orchestration is to collect telemetry data as efficiently as possible under constraints such as bandwidth and computing resources to meet the information requirements of different network management applications. Specifically, the common optimization goals of multi-telemetry task orchestration include minimizing probe flows, maximizing the number of completed telemetry tasks, and the telemetry task completion time, etc.; the constraints include network bandwidth, in-network device resource overhead, and network maximum transmission unit, etc.

[0023] The above two types of optimization methods have the following disadvantages: 1. The INT technical solution optimized for the telemetry mechanism does not jointly consider the telemetry information requirements of different management applications in the network, which may lead to problems such as redundant data collection and single-point performance bottlenecks in the parallel scenario of multiple network management applications, and it is difficult to further improve the overall efficiency of the telemetry system.

[0024] 2. The current INT solutions for optimizing multi-telemetry task orchestration mainly build models around aspects such as telemetry coverage, measurement accuracy, and bandwidth overhead, but mostly ignore the dynamic nature of the network. On the other hand, in a dynamic network, many network parameters are time-varying. To ensure measurement accuracy and efficiency, the orchestration scheme needs to be adjusted according to network changes. However, frequently adjusting the telemetry orchestration scheme will lead to a decline in network stability, and the current orchestration schemes rarely consider the balance between measurement accuracy and network stability.

[0025] 3. Currently, most telemetry orchestration algorithms are designed for specific optimization goals and specific system models, lacking generality in complex dynamic network scenarios. It is difficult to adjust telemetry orchestration strategies in a timely manner according to the network environment, resulting in a decrease in measurement accuracy and the overall efficiency of the system.

[0026] Based on this, a dynamic adaptive in-band network telemetry orchestration method and device provided by an embodiment of the present invention mainly face the INT orchestration problem in a dynamic network environment, construct a dynamic network model and a dynamic INT orchestration model, and fully consider the relationship between balancing measurement accuracy and network stability, so as to realize an INT orchestration strategy that can be adaptively adjusted according to dynamic network changes.

[0027] Specifically, the embodiment of the present invention considers the INT orchestration of long-term planning. It is assumed that the telemetry system (i.e., the network to be telemetered) operates in a time slot structure, and the time axis is discretized into time frames. The network parameters change with time slots, and the network parameters within a certain time frame are regarded as unchanged. This is essentially a stochastic optimization problem. The ultimate goals to be achieved are: (1) Coordinate the INT orchestration schemes in different time slots to maximize the long-term average information gain brought by telemetry. Here, it is assumed that collecting different types of telemetry data in different switches will bring different information gains, and this information gain also changes with time slots as network parameters. (2) Measure the relationship between network measurement accuracy and network stability, and output an adaptive INT orchestration scheme to keep the network in a stable operating environment. And an algorithm for generating an INT orchestration strategy is designed based on the diffusion model, which greatly enhances the adaptability and generality of the orchestration strategy in a complex dynamic network environment.

[0028] For the convenience of understanding this embodiment, a dynamic adaptive in-band network telemetry orchestration method disclosed by the embodiment of the present invention will be introduced in detail first.

[0029] The embodiment of the present invention provides a dynamic adaptive in-band network telemetry orchestration method, which can be executed by an electronic device with data processing capabilities. Refer to Figure 2 the flowchart of a dynamic adaptive in-band network telemetry orchestration method shown in the figure. This method mainly includes the following steps S210 to step S230: Step S210, obtain the network structure data of the network to be telemetered, where the network structure data includes a network node set and a network link set.

[0030] The physical network structure of the network to be telemetered can be modeled as an undirected graph , where is the network node set, is the network link set. For any programmable switch node , it can insert the metadata in various switches into the data packet as telemetry items (such as device ID, queue depth, processing delay, etc.). The switch v can collect a set of telemetry items as . The set of traffic flows in the network is defined as . For any traffic flow , it persists in the network for a certain period according to a certain route. To better capture the dynamics of the network, assume that the telemetry system runs in a time-slot structure and discretizes the time axis into time frames . The traffic flow f can carry a telemetry item capacity of t bytes in a time slot . As the network link condition changes, also changes between different time slots.

[0031] In Figure 3 a network structure model shown, each switch includes 3 optional telemetry items, where different shapes (such as squares, regular hexagons, and circles) indicate different telemetry items, and different colors indicate different switches. When traffic flow 1 passes through switches 1, 3, and 5, three telemetry items of switch 1, two telemetry items of switch 3, and three telemetry items of switch 5 are respectively inserted into the data packet; when traffic flow 2 passes through switches 4, 3, and 2, two telemetry items of switch 4, one telemetry item of switch 3, and two telemetry items of switch 2 are respectively inserted into the data packet.

[0032] Step S220, according to the network structure data, establish a target stochastic optimization model for multiple time slots describing the dynamic network environment; the target stochastic optimization model is obtained based on the INT orchestration model, measurement accuracy model, and network stability model of the network to be telemetered operating in a time-slot structure.

[0033] In the embodiments of the present invention, a stochastic optimization model is established for the dynamic network environment. The stochastic optimization model stipulates that the INT orchestration scheme needs to meet basic network constraints in each time slot: (1) There should be no oversampling of telemetry data, as oversampling will cause waste of bandwidth resources and reduce the efficiency of the orchestration scheme; (2) The bandwidth resources occupied by telemetry data cannot exceed the limit, as excessive bandwidth occupation will cause congestion in the network and affect service transmission in the network. In addition, by describing the changes in the telemetry orchestration scheme in different time slots, a long-term network stability constraint is constructed to reduce the impact of data collection scheme adjustment on network stability. Then, to facilitate the solution of the problem model (i.e., the stochastic optimization model), the Lyapunov optimization technique is used to transform the stochastic optimization problem, and the long-term INT orchestration problem is decoupled and transformed into a real-time online optimization problem.

[0034] In some possible embodiments, the above step S220 includes the following sub-steps S221 to S223: Sub-step S221: According to the network structure data, establish an INT scheduling model, a measurement accuracy model, and a network stability model for the telemetry network operating under a time slot structure.

[0035] In the INT scheduling model, the routing path of the traffic flow remains unchanged. In this embodiment, telemetry items in the switch are collected based on the traffic flow, while optimizing the selection and allocation of telemetry items, and outputting an INT scheduling scheme. Therefore, the INT scheduling scheme can be expressed as follows: (1); Wherein, represents whether the flow f collects telemetry items at the switch t in the time slot v , i , is the set of variables t in the time slot .

[0036] Whenever a scheduling scheme is generated in the time slot t , the traffic flow will continuously collect corresponding telemetry items according to the scheduling scheme until a new scheduling scheme is output in the next time slot. Figure 4 shows an example of an INT scheduling scheme generated according to the Figure 3 shown network topology. Figure 4 describes the INT scheduling schemes generated in time slot 1 and time slot 2. It can be seen from this that in different time slots, different traffic flows pass through different switches and collect corresponding telemetry items. The scheduling scheme is the selection and allocation of telemetry items on different traffic flows.

[0037] To intuitively evaluate the degree to which the scheduling scheme meets the telemetry requirements, this embodiment designs a measurement accuracy model. For a given telemetry item , when it is collected by the traffic flow in the time slot t , it needs to occupy bytes and obtain an information gain . The information gain is used to evaluate the importance of the telemetry item. In the time slot t , the total information gain obtained by the scheduling scheme can be expressed as: (2).

[0038] In the network stability model, consider jointly optimizing the INT scheduling schemes between different time slots. Since , When parameters such as etc. change with time, the scheduling scheme needs to be adjusted to adapt to the network dynamics. However, frequent adjustment of the scheduling scheme may affect the performance of the switch, thus affecting the network stability. To intuitively evaluate the impact of scheduling scheme adjustment on network stability, in this embodiment, the change of the scheduling scheme between different time slots is expressed as: (3); Formula (3) calculates the sum of weighted changes in the entire network with as the weight. Since the data packet lengths occupied by collecting different types of telemetry items are different, the impact of the programmable switch inserting different types of telemetry items into the data packet on its performance is also different. Generally speaking, the higher the change value, the greater the degree of adjustment of the scheduling scheme, which means the network is less stable. On the contrary, the smaller the change value, the more stable the network is.

[0039] Sub-step S222, based on the INT scheduling model, measurement accuracy model and network stability model, constructs an initial random optimization model.

[0040] Based on the above INT scheduling model, measurement accuracy model and network stability model, the dynamic INT scheduling can be characterized as the following stochastic optimization problem (i.e., the initial random optimization model): (4); where the goal of the problem is to maximize the long-term average information gain obtained by the scheduling scheme between different time slots, and the constraint C 1 ensures that in any time slot t for any switch v the telemetry items in i can only be assigned to a single traffic flow, thus preventing redundant data collection caused by oversampling. Constraint C 2 ensures that in any time slot t for the traffic flow f the total overhead of all collected telemetry items does not exceed the available capacity limit. Constraint C 3 establishes a network stability constraint, where the parameter represents the desired long-term network stability limit. is set to a smaller value, corresponding to a higher level of expected network stability, thus restricting the adjustment range of the scheduling scheme. In this problem, in order to adapt to the network dynamic changes and maximize the long-term average information gain, it may lead to very drastic short-term adjustments (changes of the scheduling scheme across adjacent time slots) of the INT scheduling scheme. However, these adjustments are also restricted by the long-term network stability at the same time. This coupling relationship makes the problem challenging.

[0041] Sub-step S223: Use the Lyapunov optimization technique to decouple the initial stochastic optimization model for the stochastic optimization problem, obtaining the target stochastic optimization model.

[0042] Handling long-term network stability constraints C 3 is a major challenge in the optimization problem of the initial stochastic optimization model. To solve this problem, this embodiment uses the Lyapunov optimization theory to convert the long-term network stability constraint into a virtual queue. The virtual queue is defined as follows: (5); Where is the time slot t is the queue backlog at time t = 0, . Based on Equation (5), we can obtain: (6); Summing both sides of Equation (6) for t from 0 to T - 1, we can get: (7); Dividing both sides of Equation (7) by T , and taking the limit value for T , we can obtain: (8); Observing Equation (8), we can find that as long as the virtual queue satisfies rate stability, i.e., , then: (9); That is, it satisfies the network stability constraint C 3.

[0043] To further transform the problem , this embodiment uses the drift-penalty algorithm in the Lyapunov optimization theory to control the queue , and determine the INT scheduling scheme within each time slot. First, define the Lyapunov function as: (10); Where is the set of virtual queues. Based on the Lyapunov function, further define the Lyapunov drift as: (11); Thus, the expression of the drift-penalty algorithm is obtained: (12); ​Among them, is a non - negative control parameter used to balance the importance degree of network stability and the information gain of the INT scheduling scheme, i.e., the measurement accuracy. When minimizing Equation (12), the conditions corresponding to Constraint C3 are satisfied .

[0044] Next, further processing of Equation (12) can obtain the upper bound of the drift penalty expression: (13); Among them, is a constant within time slot t , , is the upper bound of. For , when all variables in the network change, obtains the upper bound . Therefore, the goal of the problem is transformed into minimizing the upper bound of the drift penalty expression within each time slot t . The decoupled problem (i.e., the objective stochastic optimization model) can be expressed as: (14).

[0045] Step S230, according to the objective stochastic optimization model, generate the target INT scheduling scheme through the diffusion model and reinforcement learning; among them, the reverse denoising process in the diffusion model is the policy function in reinforcement learning.

[0046] This embodiment proposes to use an adaptive INT scheduling algorithm to solve the problem model (i.e., the objective stochastic optimization model), and adopt a generative algorithm based on the diffusion model to solve the optimization objective problem for each time slot.

[0047] To efficiently solve the problem t within each time slot , this embodiment designs an INT scheduling policy generation algorithm based on the diffusion model. The diffusion model is an efficient generative tool that transforms data from the original distribution to the Gaussian noise distribution by gradually adding noise, and then gradually removes the noise through the reverse process to reconstruct the original data distribution. This process is usually described as a continuous Markov chain: the forward process gradually increases the noise level, and the reverse process contains a conditional generative model that is trained to predict the optimal reverse transformation for each denoising step. Therefore, the diffusion model can generate data samples starting from pure noise by reversing the diffusion process.

[0048] To effectively utilize the diffusion model to generate INT orchestration policies, this embodiment combines the actor-critic architecture in Reinforcement Learning (RL) to guide the training of the diffusion model, using the reverse denoising process in the diffusion model as the policy function in RL. The adaptive INT orchestration algorithm part will be introduced in detail later.

[0049] Based on the above, the above step S230 may include: constructing an agent network, which includes a policy network and an evaluation network. The policy network is used to output the current action according to the current environmental state and the policy function, and the evaluation network is used to calculate the value function according to the current environmental state and the current action. The current environmental state includes the information gain collected by the telemetry items observed in the current time slot, the telemetry item capacity that the service flow can carry in the time slot, and the queue backlog of the virtual queue. The current action is the orchestration plan for the current time slot; obtaining the experience sequence obtained by the agent through the interaction between the initialized agent network and the network environment of the network to be telemetered, and storing the experience sequence in the buffer; wherein, the network environment is used to feedback the current reward according to the current action and the reward function corresponding to the target stochastic optimization model, and update the current environmental state to obtain the environmental state of the next time slot; the experience sequence includes the current environmental state, the current action, the current reward, and the environmental state of the next time slot under multiple time slots; sampling the experience sequence in the buffer to obtain training samples; training the agent network according to the training samples and the target function corresponding to the target stochastic optimization model to obtain the trained agent network; determining the current actions in each time slot output by the policy network in the trained agent network as the target INT orchestration plan.

[0050] Uniform sampling can be used during sampling in the buffer. To avoid the actions output by the diffusion policy network falling into local optimal solutions, the method of entropy estimation can be used to increase the exploration of the policy. Since the entropy of the diffusion policy cannot be directly estimated, a Gaussian mixture model is used to fit the policy distribution. The action exploration parameter can be updated based on the estimated diffusion policy entropy. The action exploration parameter can be used to add noise to the action. Based on this, the current action is calculated by the agent through the policy network according to the current environmental state, the policy function, and the action exploration parameter. The action exploration parameter is used to add noise to the action and is updated based on the diffusion policy entropy estimated using the Gaussian mixture model. When training the agent network, the value function can be calculated through the evaluation network according to the current environmental state and the current action in the training samples; the network parameters and action exploration parameters of the agent network are updated through the backpropagation algorithm according to the value function, the current environmental state and the current reward in the training samples, and the target function.

[0051] Optionally, to reduce the instability and performance degradation problems caused by overestimation in policy learning, the evaluation network includes two Q-value networks and their target networks; the objective function of the policy network can be: ; The objective function of the evaluation network can be: ; Reward function r can be: ; where represents the network parameters of the policy network, represents the environmental state, represents the buffer, represents the final action generated through the reverse denoising process, represents the policy function of the policy network under the network parameters and the environmental state , represents the smaller value function output by the two Q-value networks, represents the network parameters of the Q-value network, represents 's next environmental state, represents the environmental state and the action 's reward, represents the discount factor, , represent the value functions output by the two Q-value networks, , represent the value functions output by the target networks of the two Q-value networks, represents the action obtained by inputting into the policy network, represents a non-negative control parameter, represents the total information gain obtained by the scheduling scheme t within the time slot , represents the queue backlog representation of the virtual queue at time slot t , represents the change in the scheduling scheme between different time slots, represents a non-negative penalty factor, represents whether the flow f collects the telemetry item t at the switch v in the time slot i , , , , represents a telemetry item i The number of bytes occupied when collected by the service flow, represents the service flow f in a time slot t The capacity of the telemetry item that can be carried within it.

[0052] The network parameters of the policy network can be updated according to the value function and the objective function of the policy network; the network parameters of the evaluation network can be updated according to the value function, the current reward, and the objective function of the evaluation network; the diffusion policy entropy can be calculated according to the current environmental state in the training sample, and the action exploration parameters can be updated based on the diffusion policy entropy.

[0053] The dynamic adaptive in-band network telemetry orchestration method provided by the embodiments of the present invention can obtain the network structure data of the network to be telemetered, and the network structure data includes a network node set and a network link set; according to the network structure data, a target stochastic optimization model for describing the dynamic network environment in multiple time slots is established; the target stochastic optimization model is obtained based on the INT orchestration model, the measurement accuracy model, and the network stability model when the network to be telemetered operates in a time slot structure; according to the target stochastic optimization model, a target INT orchestration scheme is generated through a diffusion model and reinforcement learning; among them, the reverse denoising process in the diffusion model is the policy function in reinforcement learning. In this way, in the face of a dynamic network environment, when constructing the target stochastic optimization model, the relationship between the balanced measurement accuracy and network stability is fully considered, realizing the balance between measurement accuracy and network stability in the network telemetry of the dynamic network environment; and the diffusion model and reinforcement learning are combined to generate the target INT orchestration scheme, which has good applicability and versatility in a complex network environment and improves the efficiency of the telemetry system.

[0054] For ease of understanding, the above adaptive INT orchestration algorithm part will be introduced in detail below.

[0055] In the diffusion model, each step of the data, i.e., as a latent variable, has the same dimension as the original data , where is the probability distribution of the original data. In the forward diffusion process, the noise is gradually introduced and added to the original data N after steps, and the corresponding series of noise variances are (the noise variance is a preset hyperparameter), so the forward process can be expressed as: (15); (16); where denotes the joint probability of the entire forward diffusion sequence under the condition that the original data is , denotes the probability of transitioning to under the condition of denotes follows a Gaussian distribution with a mean of and a variance of .

[0056] When the number of steps in the noise addition process , becomes an isotropic Gaussian distribution. The reverse diffusion process, i.e., the denoising process, can be expressed as: (17); (18); where denotes the joint probability of the entire reverse diffusion process sequence, given by a model with parameters (i.e., the noise prediction network in the following text); denotes the conditional probability from n to at the step in the reverse diffusion process; denotes the mean of the distribution, determined by a model with parameters , an input of and n ; denotes the variance of the distribution, determined by a model with parameters , an input of and n ; is the data distribution of N after noise addition at the step. When , , is the identity matrix.

[0057] Taking the reverse denoising process in the diffusion model as the policy function in RL, it can be expressed as follows: (19); where , is the action distribution obtained after noise addition at the N step, and is also the starting action distribution in the reverse denoising process. And according to the above, it is a Gaussian distribution. is the final action generated through the reverse denoising process and is input into RL for evaluation training. In this embodiment, the optimal action obtained after the policy finally converges is the time slot t INT scheduling policy within , is the environmental state in RL, which is composed of the telemetry item acquisition information gain t observed within each time slot and the virtual queue , that is .

[0058] can be modeled as a Gaussian distribution , and the policy is parameterized, and the variance can be set to , and the mean can be constructed according to the noise prediction model as: (20); where , , is the noise prediction network. To obtain the final action, it is necessary to sample sequentially from N different Gaussian distributions. According to the reparameterization trick, the sampling process can be expressed as: (21); where , n is the reverse number of steps from N to 0, .

[0059] In the policy improvement stage during the training process, the goal of the policy function is to maximize the expected Q value ( Q value, that is, the value function) of the action generated by the diffusion model (i.e., the INT scheduling policy within the time slot) given the environmental state: (22); where is the buffer for storing the experience sequence obtained from the interaction between the RL agent and the environment, is the evaluation network for estimating the Q value. And in the policy evaluation stage during the training process, two Q value networks, namely , and the target network , are used to update the evaluation network, then takes , the smaller of which makes the training more stable; meanwhile, the objective function of the policy evaluation phase (i.e., Q the value network) can be expressed as: (23); wherein, is the state of the next environmental state immediately following, is the state input into the diffusion policy network to obtain the action; represents the discount factor, indicating the degree of discount for future rewards. r is the reward function. To solve the optimization problem , r can be set as: (24); which contains two parts. One part is the negative value of the optimization objective of the problem i.e., , and the other part represents the satisfaction of the constraint C 2, i.e., , is a non - negative penalty factor used to impose a penalty for violating the constraint C 2.

[0060] To avoid the action output by the diffusion policy network falling into a local optimal solution, an entropy estimation method can be used to increase the exploration of the policy. Since the entropy of the diffusion policy cannot be directly estimated, a Gaussian mixture model is used to fit the policy distribution: (25); where J is the number of Gaussian distributions, is the mixing weight of the j th Gaussian distribution, which satisfies , , , and j are the mean and variance of the th Gaussian distribution respectively. In this embodiment, the EM (Expectation Maximization) algorithm is used to calculate the weight parameter . For the state K , action samples are obtained by sampling using the diffusion policy, and then the posterior probability that the action sample j belongs to each Gaussian distribution component (26); Then, use the result calculated by formula (26) to update the Gaussian distribution and the corresponding mixture weights. : (27); (28); (29).

[0061] Through continuous iteration, the convergence of each parameter is finally achieved. After obtaining the converged Gaussian mixture model, based on formula (25) and the empirical sequence buffer the diffusion policy entropy can be calculated: (30); where d is the dimension of the action . Additionally, the action exploration parameter : (31); where g is the update learning rate of the parameter , is the target entropy (the target entropy is a preset hyperparameter). The action exploration parameter can be used to add noise to the action: (32); where, is a preset hyperparameter. Therefore, during the empirical sequence sampling process, the policy entropy can be adjusted by adding noise to the action output by the diffusion policy, increasing the exploration of the policy.

[0062] In summary, the overall process of the adaptive INT orchestration algorithm is as Figure 5 shown. The episode is the round, and the input hyperparameters include: the number of training rounds H , the time slot T , the initialization parameters of the neural network, the action exploration parameter, and the INT orchestration model parameters. The algorithm goes through a total of H rounds, H is a preset hyperparameter. First, the RL agent interacts with the environment for a total of T time slots to collect the empirical sequence. The agent observes the environmental state , and according to the policy in formula (19) and formula (32), adds noise to the action to obtain the action . The environment feeds back the reward based on the action , and the environment status is updated to , the state update includes the random variable Value updates and virtual queues of updates, including It can be updated according to formula (5): A random update that satisfies Poisson distribution or normal distribution can be used. After completing an interaction, the empirical sequence obtained is will be stored in the buffer. T After the time slot interaction is completed, The experience sequence is uniformly sampled, and the corresponding neural network parameters are updated according to the objective function (22) (23) (the parameter update formula is related to the gradient of the objective function) and the action exploration parameters are updated according to formula (31) , and then enter the next round until the maximum number of rounds is reached. After the algorithm reaches convergence after multiple rounds of training, the strategy obtained in each time slot This is the optimal INT scheduling strategy, which completes the time slot t Internal Problems The solution.

[0063] The embodiment of the present invention considers the telemetry task orchestration optimization problem in the parallel scenario of multiple network management applications, describes the dynamic network environment by constructing a random optimization model, and designs an INT orchestration strategy generation algorithm based on the diffusion model, thereby achieving a balance between measurement accuracy and network stability in network telemetry.

[0064] For the INT orchestration strategy generation algorithm, the optimization algorithm based on model and constraint construction is often subject to specific environments, lacks adaptability to network dynamics, and has poor generalization. However, the INT orchestration strategy generation algorithm designed based on the diffusion model in the embodiment of the present invention has high flexibility and strong data distribution fitting ability, and can adapt to the dynamic network environment very well.

[0065] In summary, the key technical points of the embodiments of the present invention include: 1. The INT orchestration strategy is generated by combining the diffusion model with reinforcement learning, which has good applicability and versatility in complex network environments; 2. The Gaussian mixture model is used to estimate the entropy of the diffusion strategy to further enhance the exploratory nature of the diffusion strategy to avoid falling into the local optimal solution; 3. A dynamic telemetry task orchestration model is constructed in a multi-network management application parallel scenario to form a telemetry orchestration stochastic optimization problem in a dynamic network environment; 4. Decouple the stochastic optimization problem and explore the balance between measurement accuracy and network stability through problem optimization objectives and constraints.

[0066] Corresponding to the above-mentioned dynamic adaptive in-band network telemetry orchestration method, an embodiment of the present invention further provides a dynamic adaptive in-band network telemetry orchestration device. Refer to Figure 6 the structural schematic diagram of a dynamic adaptive in-band network telemetry orchestration device shown in The acquisition module 601 is configured to acquire network structure data of the network to be telemetered, and the network structure data includes a network node set and a network link set; The establishment module 602 is configured to establish a target stochastic optimization model for multiple time slots describing the dynamic network environment according to the network structure data; the target stochastic optimization model is obtained based on the INT orchestration model, the measurement accuracy model, and the network stability model when the network to be telemetered operates under the time slot structure; The generation module 603 is configured to generate a target INT orchestration scheme according to the target stochastic optimization model through a diffusion model and reinforcement learning; wherein, the reverse denoising process in the diffusion model is the policy function in reinforcement learning.

[0067] The dynamic adaptive in-band network telemetry orchestration device provided by the present invention can acquire network structure data of the network to be telemetered, and the network structure data includes a network node set and a network link set; establish a target stochastic optimization model for multiple time slots describing the dynamic network environment according to the network structure data; the target stochastic optimization model is obtained based on the INT orchestration model, the measurement accuracy model, and the network stability model when the network to be telemetered operates under the time slot structure; generate a target INT orchestration scheme according to the target stochastic optimization model through a diffusion model and reinforcement learning; wherein, the reverse denoising process in the diffusion model is the policy function in reinforcement learning. In this way, in the face of a dynamic network environment, when constructing the target stochastic optimization model, the relationship between balancing measurement accuracy and network stability is fully considered, realizing the balance between measurement accuracy and network stability in the network telemetry of the dynamic network environment; and a diffusion model and reinforcement learning combination method is used to generate the target INT orchestration scheme, which has good applicability and generality in a complex network environment and improves the efficiency of the telemetry system.

[0068] Further, the above-mentioned establishment module 602 is specifically configured to: establish an INT orchestration model, a measurement accuracy model, and a network stability model when the network to be telemetered operates under the time slot structure according to the network structure data; construct an initial stochastic optimization model based on the INT orchestration model, the measurement accuracy model, and the network stability model; and use the Lyapunov optimization technique to decouple the initial stochastic optimization model for the stochastic optimization problem to obtain the target stochastic optimization model.

[0069] Further, the time slot t orchestration scheme in the above-mentioned INT orchestration model is expressed as: ; In the measurement accuracy model, within the time slot t , the overall information gain obtained by the scheduling scheme is expressed as: ; ; The change in the scheduling scheme between different time slots in the network stability model is expressed as: ; The initial random optimization model is expressed as: ; where indicates whether the flow f is a telemetry item at the collection switch t in the time slot v , i , represents the set of time slots, represents the set of network nodes, , represents the set of telemetry items that can be collected within the switch v , , represents the set of traffic flows in the network to be telemetered, , represents the telemetry item v at the switch i in the time slot t when it is collected by the traffic flow, the information gain obtained, represents the number of bytes occupied when the telemetry item i is collected by the traffic flow, represents the capacity of the telemetry items that the traffic flow f can carry within the time slot t ; represents the expected long-term network stability limit.

[0070] Furthermore, the above target random optimization model is expressed as: ; where represents the queue backlog of the virtual queue used to characterize the change in the scheduling scheme at the time slot t , , represents the change in the scheduling scheme between different time slots, represents the expected long-term network stability limit, represents a non-negative control parameter, represents within the time slot t , the scheduling scheme The total information gain obtained Indicates the flow f Whether in the time slot t Collecting switch v Telemetry items at i , Indicates the set of traffic flows in the network to be telemetered, , Indicates the set of time slots, Indicates the set of network nodes, , Indicates the switch v The set of telemetry items that can be collected inside, , Indicates the telemetry item i The number of bytes occupied when collected by the traffic flow, Indicates the traffic flow f In the time slot t The telemetry item capacity that can be carried inside.

[0071] Furthermore, the above-mentioned generation module 603 is specifically configured to: construct an agent network, the agent network includes a policy network and an evaluation network, the policy network is used to output the current action according to the current environmental state and the policy function, the evaluation network is used to calculate the value function according to the current environmental state and the current action, the current environmental state includes the information gain of the telemetry item collection observed in the current time slot, the telemetry item capacity that the traffic flow can carry in the time slot, and the queue backlog of the virtual queue, and the current action is the scheduling scheme of the current time slot; obtain the experience sequence obtained by the agent through the interaction between the initialized agent network and the network environment of the network to be telemetered, and store the experience sequence in the buffer; wherein, the network environment is used to feedback the current reward according to the current action and the reward function corresponding to the target stochastic optimization model, and update the current environmental state to obtain the environmental state of the next time slot; the experience sequence includes the current environmental state, the current action, the current reward, and the environmental state of the next time slot in multiple time slots; sample the experience sequence in the buffer to obtain training samples; train the agent network according to the training samples and the target function corresponding to the target stochastic optimization model to obtain the trained agent network; determine the current actions in each time slot output by the policy network in the trained agent network as the target INT scheduling scheme.

[0072] Furthermore, the above-mentioned current action is calculated by the agent through the policy network based on the current environmental state, policy function, and action exploration parameter. The action exploration parameter is used to add noise to the action and is updated based on the diffusion policy entropy estimated using the Gaussian mixture model. The above-mentioned generation module 603 is further configured to: calculate the value function through the evaluation network according to the current environmental state and the current action in the training sample; update the network parameters and action exploration parameters of the agent network through the backpropagation algorithm according to the value function, the current environmental state and the current reward in the training sample, and the objective function.

[0073] Furthermore, the above-mentioned evaluation network includes two Q-value networks and their target networks; the objective function of the policy network is: ; The objective function of the evaluation network is: ; Reward function r is: ; Among them, represents the network parameters of the policy network, represents the environmental state, represents the buffer, represents the final action generated after the reverse denoising process, represents the policy network at the network parameters , environmental state under the policy function, represents the smaller of the value functions output by the two Q-value networks, represents the network parameters of the Q-value network, represents the next environmental state of, represents the environmental state , action under the reward, represents the discount factor, , represents the value functions output by the two Q-value networks, , represents the value functions output by the target networks of the two Q-value networks, represents the action obtained by inputting into the policy network, represents the non-negative control parameter, represents at the time slot t within, the scheduling scheme obtained total information gain, represents the time slott represents the queue backlog of the virtual queue at a time represents the change in the scheduling scheme between different time slots represents a non - negative penalty factor represents a flow f whether at time slot t collection switch v the telemetry item at i , , , , represents the telemetry item i the number of bytes occupied when being collected by the traffic flow represents the traffic flow f at time slot t the capacity of the telemetry items that can be carried within

[0074] For the device provided in this embodiment, its implementation principle and the technical effects produced are the same as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment.

[0075] As Figure 7 shown, an electronic device 700 provided by an embodiment of the present invention includes: a processor 701, a memory 702, and a bus. The memory 702 stores a computer program that can run on the processor 701. When the electronic device 700 runs, communication between the processor 701 and the memory 702 is through the bus, and the processor 701 executes the computer program to implement the above - mentioned dynamic adaptive in - band network telemetry scheduling method.

[0076] Specifically, the above - mentioned memory 702 and processor 701 can be general - purpose memory and processor, and no specific limitation is made here.

[0077] An embodiment of the present invention also provides a computer - readable storage medium. A computer program is stored on this computer - readable storage medium, and when the computer program is run by a processor, it executes the dynamic adaptive in - band network telemetry scheduling method in the foregoing method embodiment. The computer - readable storage medium includes: various media such as USB flash drives, mobile hard disks, read - only memory (abbreviated as ROM), RAM, magnetic disks, or optical discs that can store program codes.

[0078] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0079] In all the examples shown and described herein, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0081] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some communication interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0082] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0083] In addition, in each embodiment of the present invention, each functional module may be integrated into one processing module, may exist separately physically for each module, or two or more modules may be integrated into one module.

[0084] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dynamic adaptive in-band network telemetry orchestration method, characterized in that, Including: Obtain the network structure data of the network to be telemetered, where the network structure data includes a network node set and a network link set; According to the network structure data, establish a target stochastic optimization model for multiple time slots describing a dynamic network environment; the target stochastic optimization model is obtained based on the INT orchestration model, measurement accuracy model, and network stability model when the network to be telemetered operates under a time slot structure; According to the target stochastic optimization model, generate a target INT orchestration scheme through a diffusion model and reinforcement learning; wherein, the reverse denoising process in the diffusion model is the policy function in reinforcement learning.

2. The method according to claim 1, wherein The step of establishing a target stochastic optimization model for multiple time slots describing a dynamic network environment according to the network structure data includes: According to the network structure data, establish an INT orchestration model, a measurement accuracy model, and a network stability model when the network to be telemetered operates under a time slot structure; Based on the INT orchestration model, the measurement accuracy model, and the network stability model, construct an initial stochastic optimization model; Use Lyapunov optimization technology to decouple the stochastic optimization problem of the initial stochastic optimization model to obtain the target stochastic optimization model.

3. The method according to claim 2, wherein The time slot in the INT scheduling model t scheduling scheme is expressed as: ; In the measurement accuracy model, within the time slot t the total information gain obtained by the scheduling scheme is expressed as: ​ ; The scheduling scheme changes between different time slots in the network stability model It is expressed as: ; The initial stochastic optimization model is expressed as: ; Among them, represents a flow f whether in a time slot t collection switch v the telemetry item at i , represents a set of time slots, represents a set of network nodes, , represents a switch v the set of telemetry items that can be collected within , represents the set of traffic flows in the network to be telemetered, , represents a switch v the telemetry item at i in the time slot t the information gain obtained when the telemetry item is collected by the traffic flow, represents the telemetry item i the number of bytes occupied when the telemetry item is collected by the traffic flow, represents the traffic flow f in the time slot t the capacity of the telemetry items that can be carried within; represents the expected long-term network stability limit.

4. The method according to claim 2, wherein The target stochastic optimization model is expressed as: ; in, Indicates time slot t The queue backlog of the virtual queue used to represent the change of the orchestration scheme, , Indicates the change of scheduling scheme between different time slots, represents the expected long-term network stability limit, represents a non-negative control parameter, Indicates time slot t Internal, arrangement plan The total information gain obtained is, Representation Flow f Is it in the time slot? t Collection switch v Telemetry items at i , Represents the set of service flows in the network to be telemetered. , represents a collection of time slots, Represents a collection of network nodes, , Indicates a switch v The collection of telemetry items that can be collected within, , Represents a telemetry item i The number of bytes occupied by the business flow when collected, Indicates business flow f In time slot t The telemetry capacity that can be carried within.

5. The method according to claim 1, wherein The step of generating a target INT orchestration scheme through a diffusion model and reinforcement learning according to the target stochastic optimization model includes: Construct an agent network, which includes a policy network and an evaluation network. The policy network is used to output the current action according to the current environmental state and the policy function, and the evaluation network is used to calculate the value function according to the current environmental state and the current action. The current environmental state includes the information gain collected by the telemetry items observed in the current time slot, the capacity of the telemetry items that the traffic flow can carry in the time slot, and the queue backlog of the virtual queue. The current action is the orchestration scheme for the current time slot; Obtain the experience sequence obtained by the agent through the interaction between the initialized agent network and the network environment of the network to be telemetered, and store the experience sequence in a buffer; wherein, the network environment is used to feedback the current reward according to the current action and the reward function corresponding to the target stochastic optimization model, and update the current environmental state to obtain the environmental state of the next time slot; the experience sequence includes the current environmental state, the current action, the current reward, and the environmental state of the next time slot in multiple time slots; Sample the experience sequence in the buffer to obtain training samples; Train the agent network according to the training samples and the objective function corresponding to the target stochastic optimization model to obtain the trained agent network; Determine the current actions in each time slot output by the policy network in the trained agent network as the target INT orchestration scheme.

6. The method according to claim 5, wherein The current action is calculated by the agent through the policy network according to the current environmental state, the policy function, and the action exploration parameter, where the action exploration parameter is used to add noise to the action and is updated based on the diffusion policy entropy estimated using a Gaussian mixture model; Training the agent network according to the training samples and the objective function corresponding to the target stochastic optimization model to obtain a trained agent network, including: Calculating a value function through the evaluation network according to the current environmental state and the current action in the training samples; Updating the network parameters of the agent network and the action exploration parameter through a backpropagation algorithm according to the value function, the current environmental state and the current reward in the training samples, and the objective function.

7. The method according to claim 5, wherein The evaluation network includes two Q-value networks and their target networks; the objective function of the policy network is: ; The objective function of the evaluation network is: ; The reward function r is as follows: ; Among them, represents the network parameters of the said policy network, represents the environmental state, represents the said buffer, represents the final action generated through the reverse denoising process, represents the policy function of the said policy network with network parameters and environmental state ; represents the smaller value function output by the two said Q-value networks, represents the network parameters of the said Q-value network, represents the next environmental state of represents the environmental state and action at which the reward is obtained, represents the discount factor, and represents the value functions output by the two said Q-value networks, and represents the value functions output by the target networks of the two said Q-value networks, represents the action obtained by inputting into the said policy network, represents a non-negative control parameter, t represents the total information gain obtained by the orchestration scheme within the time slot represents the time slot t when the queue backlog of the virtual queue is represented, represents the change in the orchestration scheme between different time slots, represents a non-negative penalty factor, represents whether the flow f collects the telemetry item t at the switch v ; i represents the number of bytes occupied when the telemetry item i is collected by the service flow, represents the telemetry item capacity that the service flow f can carry within the time slot t .

8. A dynamic adaptive in-band network telemetry orchestration device, characterized in that, Including: An acquisition module for acquiring network structure data of a network to be telemetered, where the network structure data includes a network node set and a network link set; A building module for building a target stochastic optimization model for describing a dynamic network environment under multiple time slots according to the network structure data; the target stochastic optimization model is obtained based on an INT orchestration model, a measurement accuracy model, and a network stability model when the network to be telemetered operates under a time slot structure; A generation module for generating a target INT orchestration scheme through a diffusion model and reinforcement learning according to the target stochastic optimization model; where the reverse denoising process in the diffusion model is the policy function in reinforcement learning.

9. An electronic device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, and is characterized in that When the processor executes the computer program, it implements the dynamic adaptive in-band network telemetry orchestration method according to any one of claims 1-7.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is run by a processor, it executes the dynamic adaptive in-band network telemetry orchestration method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Base station energy-saving regulation and control method based on diffusion model and reinforcement learning

    CN119277491A