Method for determining at least one alternative path for mitigation in the event of saturation, or risk of saturation, on at least one link of an autonomous system

A reinforcement learning-based method for determining alternative paths in autonomous systems addresses the inefficiencies of manual traffic routing, automating bypass path creation and removal to prevent saturation, enhancing service quality and reducing operational errors.

WO2026008655A1PCT designated stage Publication Date: 2026-01-08ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/068731
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-05
Filing Date
2025-07-01
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing methods for managing traffic routing in autonomous systems are manual, time-consuming, and prone to errors, leading to potential traffic losses and service degradation due to intra-domain saturation, which is exacerbated by complex and dynamic content delivery strategies.

Method used

A reinforcement learning-based method for determining alternative paths within an autonomous system, using agents trained on data exchanges to create or remove bypass paths dynamically to prevent or reduce saturation, leveraging matrices and vectors for load and saturation states.

Benefits of technology

Enables real-time or near-real-time adjustment of data flow paths to prevent or reduce saturation, ensuring quality of service by automating the creation and removal of bypass paths based on load and saturation data, reducing operational errors and computational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025068731_08012026_PF_FP_ABST
    Figure EP2025068731_08012026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for determining at least one alternative path for routing at least one data stream within an autonomous system implementing an IP routing protocol, in the event of saturation, or risk of saturation, of at least one link of a plurality of target links present on at least one path between an entry point and an exit point of the autonomous system. Such a method comprises: - a reinforcement-learning training phase (11), in which at least one agent (AG) is trained to determine at least one path for routing the at least one data stream within the autonomous system, on the basis of at least one exchange of data between the at least one agent (AG) and a reinforcement-learning environment (ENV), as a function of at least one datum representative of the saturation or risk of saturation; - an inference phase (12), in which the at least one trained agent is used to deliver the at least one alternative path.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] TITLE: Method for determining at least one alternative path for mitigation in case of saturation or risk of saturation on at least one link of an autonomous system.

[0003] technical field

[0004] The invention relates to the field of data flow routing within a communication network. More specifically, the invention relates to techniques aimed at reducing or preventing the occurrence of saturation phenomena on the links used to route traffic within an autonomous system of a communication network.

[0005] Previous art

[0006] A wide area network, such as the Internet, typically comprises numerous autonomous systems. An autonomous system can be defined as a set of subnets and routers managed by a single administrative authority (e.g., a state, an organization, a company such as a service provider, etc.). To maintain overall accessibility and connectivity within the wide area network, including the end-to-end routing of data streams from a source device to a destination device that may not belong to the same autonomous system, the various autonomous systems within the wide area network are interconnected.

[0007] Each autonomous system also decides locally, based on its own internal routing policy, the best paths for data flows within it. Therefore, a path taken by a data flow in a "forward" direction, from the first device to the second device in the network, is generally not the same as the path taken by a data flow in a "return" direction, from the second device to the first. The paths through which data flows are routed in a wide area network such as the internet are thus asymmetric, and an autonomous system cannot predict through which neighbors or entry points a data flow will be delivered.The generalization of content caching within increasingly numerous content delivery networks, for example so that users can access it more quickly, further exacerbates this phenomenon, as it leads to the development of increasingly complex and dynamic content delivery strategies.

[0008] One consequence of the difficulty in estimating and anticipating the amount of traffic to be routed within an autonomous system is the risk of intra-domain saturation. This can occur, in particular, when the amount of traffic routed over one or more links within the autonomous system becomes too close to, equal to, or greater than the maximum capacity of the link(s) in question. If these saturation phenomena persist over time, they can lead to traffic losses (e.g., IP packet loss, where packets never reach the final destination) and a degradation of service quality, with a particularly disappointing effect on users.

[0009] A monitoring team that observes intra-domain congestion can attempt to offload the congested link(s) by rerouting some data flows to other paths within the autonomous system, for example, via temporary bypass paths specifically created for this purpose. However, detecting congested links, identifying the data flows routed through these links, creating potential bypass paths, analyzing the effects of rerouting traffic through these bypass paths, and removing the bypass paths once the congestion episode is over are currently performed manually. These tasks are time-consuming, complex, and tedious, and their implementation is therefore potentially prone to numerous errors.

[0010] Therefore, there is a need for a solution that allows for more effective protection against the risks of saturation of traffic routing links within an autonomous system.

[0011] Summary of the invention

[0012] The present technique offers a solution to address certain drawbacks of the prior art. In one respect, the present technique relates to a method for determining at least one alternative path for routing at least one data stream within an autonomous system.More specifically, such a process includes: a reinforcement learning training phase, in which at least one agent is trained to determine at least one routing path for said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent and a reinforcement learning environment, based on at least one data point representing a saturation or risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system; an inference phase, in which said at least one trained agent is used to deliver said at least one alternative path.

[0013] In this way, the present technique offers an ingenious solution based on reinforcement learning techniques to determine alternative paths for routing data flows within an autonomous system. This allows for real-time or near-real-time adjustment of the paths used to route all or part of the data flows in transit within the autonomous system, preventing or at least reducing the occurrence of saturation phenomena and thus offering greater guarantees in terms of quality of service. As described later, such an adjustment includes, for example, the creation of new bypass paths (or tunnels, typically IP tunnels) and / or the removal of previously created bypass paths, depending on the evolution of saturation phenomena or the identified risks of saturation.

[0014] In a particular embodiment, said at least one data exchange comprises: the execution, by said at least one agent, of at least one action on said learning environment, determined based on current state data comprising data representative of a current load state of said target links and data representative of a current saturation state of said target links, said action being associated with at least one current routing path of at least one target data stream within said autonomous system; the obtaining, by said at least one agent, in response to said execution, of updated state data and reward data from said learning environment.In this way, the proposed technique allows, through a trial-and-error approach, the exploration and testing, without operational impact, of routing target data flows via different paths, or even the creation of specific bypass paths for such routing, based on observational data from a training environment. Thus, an agent is able to learn, before its deployment in production, to identify bypass paths within an autonomous system that minimize saturation phenomena.

[0015] In a particular embodiment, the determination of said action executed on the learning environment takes into account, in addition to said current state data, previous state data, including data representing a previous load state of said target links and data representing a previous saturation state of said target links.

[0016] In this way, the agent has additional information on which to base an action to be performed on the learning environment. In particular, by having access, in addition to an observation at a current time, to an observation at a previous time, the agent has the ability to assess changes or trends such as significant increases in load on certain links, for example, which allows it to anticipate risks of saturation before such saturations become actual.

[0017] According to a particular characteristic, said data representing a previous state of load of said target links and said data representing a current state of load of said target links each take the form of a matrix, one dimension among rows and columns of said matrix being associated with said target links, the other dimension among rows and columns of said matrix being associated with said target flows, each coefficient of said matrix being associated with both one of said target links and one of said target data flows, and having as its value a load rate of said target link due to said target data flow.

[0018] According to a particular characteristic, said data representing a previous state of saturation of said target links and said data representing a current state of saturation of said target links each take the form of a vector, each component of said vector being associated with one of said target links and having as its value a first value representing an absence of saturation of said target link or a second value representing a saturation of said target link.

[0019] In this way, we have a simple yet comprehensive mathematical representation of the observations made on the training environment, in the form of matrices and / or vectors that are easily manipulated within a learning model. Furthermore, having a separate variable to indicate whether a link is saturated or not (i.e., the saturation state), distinct from the one used to define the link load (i.e., the load state), offers great flexibility in implementing the process according to this technique. For example, it allows an administrator of the autonomous system to define the conditions and criteria under which a link is considered saturated or not, with these criteria potentially going beyond simply considering a high link load.

[0020] In a particular embodiment, said action associated with at least one current routing path of at least one target data stream is selected from: maintaining said current routing path of said target data stream; or modifying said current routing path of said target data stream.

[0021] In this way, a current routing path of a target data stream is not necessarily modified, if the agent determines, based on observations of the learning environment obtained, that a modification of the current routing path of this data stream has little or no impact on the occurrence of saturation phenomena within the autonomous system.

[0022] According to a particular feature, said modification of the current routing path of said target data stream includes, where said current routing path corresponds to a path established according to a routing protocol in force within said autonomous system, the creation, alongside said routing protocol, of a bypass path between an entry point and an exit point of said current routing path, and the rerouting of said target data stream onto said bypass path.

[0023] In this way, the proposed technique makes it possible in certain situations to bypass the routing protocol in force within the autonomous system, by allowing the creation of bypass paths (or tunnels, typically IP tunnels) outside the protocol, for example when it is determined that a current routing path established according to the classic routing protocol is not optimal in order to avoid the occurrence of saturation phenomena within the autonomous system.

[0024] According to a particular characteristic, said modification of the current routing path of said target data stream includes, where said current routing path corresponds to a bypass path defined in the margin of a routing protocol in force within said autonomous system, the removal of said bypass path, and the rerouting of said target data stream along a path defined according to said routing protocol.

[0025] In this way, a bypass path (or tunnel) previously created outside the routing protocol in effect within the autonomous system can be automatically removed when it no longer constitutes an optimal path for avoiding or limiting saturation within the autonomous system. This frees up the computing resources used to maintain this bypass path. Consequently, the system reverts to standard methods for routing data flow, compliant with the routing protocol in effect within the autonomous system. This reduces the number of operations managed outside a typical autonomous system configuration, the implementation of which is potentially error-prone and costly in terms of computing power.

[0026] In one particular embodiment, said reward data is determined based on a current number of saturated links within said autonomous system.

[0027] In this way, we have a simple metric to assess the extent to which an autonomous system is impacted by saturation phenomena that could lead to degradations in the quality of services.

[0028] According to a particular characteristic, said reward data is further determined based on a current number of bypass paths used to route said at least data stream within said autonomous system.

[0029] In this way, the cost of maintaining bypass paths or tunnels for routing data flows within the autonomous system is also taken into account for determining alternative paths according to the present technique.

[0030] In one particular embodiment, said alternative path is a bypass path created between said entry point and said exit point of said autonomous system. In another particular embodiment, said alternative path is a path established according to a routing protocol in force within said autonomous system, following the removal of a bypass path previously created between said entry point and said exit point of said autonomous system.

[0031] In this way, the said inference phase allows the creation and / or deletion, by the said trained agent, of at least one bypass path within the said autonomous system.

[0032] In another aspect, the present technique relates to an electronic device for determining at least one alternative path for the routing of at least one data stream within an autonomous system.Such a device includes means (for example at least one processor configured for this purpose) for the implementation of: a reinforcement learning training phase, in which at least one agent is trained to determine at least one routing path for said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent and a reinforcement learning environment, based on at least one data point representing a saturation or risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system; an inference phase, in which said at least one trained agent is used to deliver said at least one alternative path.

[0033] Such an electronic device can, of course, exhibit the various characteristics of the determination method according to the invention, which can be combined or considered individually. Thus, the characteristics and advantages of this device are the same as those of the method for determining at least one alternative path for routing at least one data stream within an autonomous system, and are not described in further detail.

[0034] In another aspect, the proposed technique also relates to a computer program product downloadable from a communication network and / or stored on a computer-readable medium and / or executable by a microprocessor, comprising program code instructions for executing a method for determining at least one alternative path as described above in any of its embodiments, when executed on a computer. The proposed technique also relates to a computer-readable storage medium on which is stored a computer program comprising program code instructions for executing the steps of the method as described above, in any of its embodiments.

[0035] Such a recording medium can be any entity or device capable of storing the program. For example, the medium may include a storage means, such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a USB flash drive or a hard drive.

[0036] On the other hand, such a recording medium can be a transmissible medium such as an electrical or optical signal, which can be transmitted via an electrical or optical cable, by radio, or by other means, so that the computer program it contains can be executed remotely. The program according to the invention can, in particular, be uploaded to a network, for example, the Internet.

[0037] The different embodiments mentioned above can be combined with each other for the implementation of the invention.

[0038] Figures

[0039] Other features and advantages of the invention will become clearer upon reading the following description of a preferred embodiment, given by way of simple illustrative and non-limiting example, and the accompanying drawings, among which:

[0040] [Fig 1] presents an example of an autonomous system in which the present technique can be implemented, in a particular embodiment of the proposed technique;

[0041] [Fig 2] presents an example of an autonomous system exhibiting intra-domain saturation; [Fig 3] illustrates the general principle of a method for determining at least one alternative path for routing at least one data stream within an autonomous system, in a particular embodiment of the proposed technique;

[0042] [Fig 4] presents the data exchanges carried out between an agent and a reinforcement learning environment for training said agent, in a particular embodiment of the proposed technique;

[0043] [Fig 5] presents an example of representing a target link load state within an autonomous system, in a particular embodiment of the proposed technique; [Fig 6] presents an example of representing a target link saturation state within an autonomous system, in a particular embodiment of the proposed technique;

[0044] [Fig 7] presents an example of an autonomous system after implementation of a process for determining at least one alternative path, in a particular embodiment of the proposed technique;

[0045] [Fig 8] describes a simplified architecture of an electronic device for the implementation of the proposed technique, in a particular embodiment.

[0046] Detailed description of the invention

[0047] This application addresses at least some of the aforementioned drawbacks.

[0048] In all figures in this document, elements and steps of the same nature are designated by the same reference.

[0049] In its first aspect, this technique relates to a method for determining at least one alternative path for the routing (or rerouting) of at least one data stream within an autonomous system implemented in a communication network. As described later, this technique can thus lead to modifications of the current routing paths of one or more target data streams. Such modifications can be implemented, for example, by creating bypass paths (or tunnels, typically IP tunnels) and / or removing previously created bypass paths within the autonomous system, depending on the evolution of saturation phenomena or identified saturation risks. Thus, the proposed technique enables automatic mitigation in the event of saturation or the risk of saturation on one or more links of the autonomous system.As described later, such mitigation is implemented by controlled traffic rerouting using a reinforcement learning agent. An example of an autonomous system (AS) in which this technique can be implemented is described in relation to Figure 1 for illustrative purposes.

[0050] Such an autonomous system (AS) comprises a plurality of devices, typically routers (nine routers RI to R9 in the example in Figure 1), responsible for routing traffic within the system using an IP routing protocol (Internet Protocol). These routers include edge routers (routers RI, R2, R3, R5, and R9 in the example in Figure 1) through which traffic to or from other autonomous systems (or other entities in the communication network) enters or leaves the autonomous system.

[0051] Data flows are routed within the autonomous system SA via paths defined by the IP routing protocol, which is implemented in a distributed manner within each router of the system. For example, a data flow fl enters the autonomous system SA at an entry point formed by router RI, and is then successively transmitted hop by hop from router RI to router R2, from router R2 to router R5, and from router R5 to router R9. Router R9 acts as the exit point, directing the data flow fl to another entity in the communication network (e.g., another autonomous system). The routing path R1-R2-R5-R9 of the data flow fl, represented by arrows in Figure 1, is thus formed by a set of links (link R1-R2, link R2-R5, link R5-R9) taken successively from the entry point RI of the autonomous system SA to the exit point R9 of the autonomous system.

[0052] At any given moment, the autonomous system (AS) is potentially traversed by a multitude of data streams, each routed from an entry point to an exit point of the system, via a routing path defined according to the IP routing protocol implemented at the level of the AS's routers. Each path consists of a succession of links between routers or other equipment within the AS. In this context, the paths are calculated in a distributed manner by the routers, without any bandwidth reservation, and saturation phenomena can therefore occur.Thus, as illustrated in Figure 2, the routing of multiple data streams within the autonomous system (AS) results in some links of the autonomous system being more loaded than others at certain times, for example, because they are used simultaneously by several data streams (for clarity, only one data stream fl is referenced in Figure 2, but it is understood that other data streams not shown in this figure are potentially routed at the same time as the fl stream in the autonomous system (AS)). For example, with a representation like the one in Figure 2, in which the thickness of a line associated with a link represents the load factor of the link (in other words, the thicker the line, the more loaded the link), we observe, for example, that the R8-R7 link is more loaded than the R7-R9 link, which itself is more loaded than the R5-R6 link.In a critical situation, when the load on a link becomes too high—that is, too close to, equal to, or even exceeding the link's maximum capacity—saturation phenomena, defined later in this document, can occur. In the example shown in Figure 2, such a saturation phenomenon is observed on the R2-R5 link, where it is illustrated as a very thick link with hatching. Saturation of this link results in disruptions to the routing of data streams using it, such as the fl data stream (packet loss, service degradation, etc.).

[0053] When an occurrence or risk of occurrence of such saturation phenomena is detected, one objective of the present technique is therefore to determine, for at least some of the target data flows routed in the autonomous system, alternative paths, in that they allow, via the rerouting of the target data flows on their respective alternative paths, to minimize these saturation phenomena on at least some of the target links of the autonomous system (for example by eliminating these phenomena, reducing them, or limiting their duration over time).In other words, the present technique relates to the implementation of an automatic mitigation solution in case of saturation or risk of saturation on one or more links between two routers: as detailed later, it aims more particularly to offer an automatic intra-domain desaturation controlled by a reinforcement learning agent, based on a rerouting on an alternative path of at least one target data flow.In this sense, it differs from load balancing or optimal resource allocation solutions (such as those that can be implemented in "software-defined network" or SDN networks), in which a controller centrally allocates paths with a certain amount of bandwidth based on reservation requests made to it, performing a full recalculation of allocation for all requests with each new request or repeatedly to determine if there is an optimal allocation.

[0054] The general principle of the proposed method for determining at least one alternative path for routing at least one data stream within an autonomous system is now illustrated with reference to Figure 3, in a particular embodiment. This method is implemented, for example, within an electronic device described later in this document. In step 31, a reinforcement learning training phase is implemented using a previously instantiated and configured reinforcement learning environment.More specifically, the learning environment is configured according to data collected on the autonomous system that is the object of the process according to the present technique, this collected data including in particular numerous data relating to the topology, configuration, and use of the autonomous system, such as for example the number of routers present in the autonomous system, the number of links defined, the capacity of these links, the routing protocol in force within the autonomous system, the profile of the data flows passing through the autonomous system, etc.Such data can, for example, be entered, in whole or in part, by a human operator using a human-machine interface provided by the electronic device responsible for implementing the alternative path determination process described herein, or it can be automatically obtained from various data sources via a communication network to which the electronic device is connected. Thus, the training environment forms a simulation environment as close as possible to the autonomous system as deployed in production. For example, the training environment can be implemented as a digital twin of the target autonomous system.

[0055] During training phase 31, at least one agent (instantiated, for example, by the electronic device responsible for implementing the process) is trained to determine paths for at least one data stream to traverse the autonomous system from an entry point to an exit point associated with the data stream in question, while respecting predefined constraints. More specifically, the data stream routing paths within the autonomous system are determined taking into account at least one piece of data representing a saturation or a risk of saturation of at least one link among a plurality of target links existing in the autonomous system. As described later, these target links may include all the links of the autonomous system, or only a subset of the links of the autonomous system that an administrator of the autonomous system wishes to protect more specifically (i.e.(to protect as much as possible from any risk of saturation).

[0056] During this reinforcement learning training phase, using a trial-and-error approach, multiple iterations of data exchange are carried out between the agent and the learning environment, with a view to broadly exploring the results obtained in response to different actions or sequences of actions detailed later.

[0057] Each of these exchanges comprises, according to the usual vocabulary of reinforcement learning and as illustrated in relation to Figure 4: the execution, by an agent AG, of at least one action ACT on the learning environment ENV, based on current state data D_ETA C; the reception, by the AG agent, in response to the execution of said ACT actions, of updated state data (or subsequent state data) D_ETAs and reward data D_REC from the ENV learning environment.

[0058] The D_ETA state data can be considered as observations of the ENV training environment. Thus, the current D_ETA state data C are representative of a current state of the ENV reinforcement learning environment, and the updated state data D_ETAs are representative of a new state of the ENV reinforcement learning environment following the execution of ACT actions on that environment. As illustrated in Figure 4, the updated state data D_ETAs become the current state data D_ETA C for the next iteration of a data exchange as previously described.

[0059] The D_REC reward data represents one or more metrics of the effectiveness of the ACT actions performed on the ENV environment, with respect to the intended objective. More specifically, the successive actions executed on the learning environment are determined by a reinforcement algorithm, in order to maximize over time the D_REC reward data received in response to these actions.

[0060] Within the framework of this technique, the objective is to prevent (or at least limit) saturation phenomena on one or more target links, and / or in relation to one or more target data streams, within the autonomous system. The target links may include all existing links within the autonomous system, or conversely, only a subset of these links, previously identified as requiring monitoring and on which an administrator of the autonomous system, for example, has a particular interest in protecting against any saturation phenomena.Following a similar approach, target data flows can include all data flows routed within the autonomous system, or conversely only a part of these data flows, previously identified as needing monitoring and for which an administrator of the autonomous system has, for example, a particular interest in protecting against any saturation phenomenon.

[0061] The D_ETA state data allows the AG agent to obtain at least a partial view (i.e., an "observation") of the state of the ENV learning environment, before and after execution of an ACT action on that environment.

[0062] In a particular embodiment of the present technique, these D_ETA state data include data representing a load state of the target links and data representing a saturation state of the target links.

[0063] The load state of a link corresponds to a load factor (or occupancy rate) of the link. Typically, such a variable can be normalized to take values ​​between 0 and 1 (or between 0% and 100%), with a load factor of 0% indicating that no data flows are passing through the link, and a load factor of 100% indicating that the entire bandwidth of the link (i.e., the entire capacity of the link) is being used to carry data flows. Depending on a particular characteristic, the load state of a link can be broken down by data flow, thus providing a view, for the given link, of the contribution of each data flow using that link to the total load factor of the link.It can thus be determined, for example, for a link Li used by three data streams Fi, F2, and F3 at a current time, that the data stream Fi is responsible for 60% of the current load of link Li, that the data stream F2 is responsible for 12% of the current load of link Li, and that the data stream F3 is responsible for 17% of the current load of link Li, resulting in a total current load rate of link Li of 89%. As illustrated in Figure 5, the load states of the target links of the autonomous system at a current time can be represented in matrix form. More specifically, each row of the matrix is ​​associated with a particular link among n target links (Li, L2, L3, ..., L). n ), each column of the matrix is ​​associated with a particular data stream among m target data streams (Fi, F2, F3, F4, ..., F mThe value in the matrix at the intersection of a row and a column corresponds to the proportion of the link's load due to the data flow in question (expressed as a percentage in the example in Figure 5). A link's saturation state is a parameter defining whether the link is considered saturated or not, based on predefined criteria. For example, a link's saturation state is represented by a Boolean value, where a first Boolean value (e.g., "0" or "false") indicates that the link is not saturated, and a second Boolean value (e.g., "1" or "true") indicates that the link is saturated.Although it generally derives directly from the load state of a link, it is particularly advantageous, as proposed in this technique, to use a variable other than the load state—in this case, the saturation state—to define whether a link is saturated or not. This allows saturation to be assessed in different ways (for example, according to specific criteria chosen by the autonomous system administrator), thus offering flexibility in implementing the method according to this technique. Depending on various alternative specific characteristics, a link can, for example, be considered saturated when the total load rate of the link exceeds a certain threshold (for example, above 70%, 80%, 90%, or 95% load), or when propagation delays on the link become excessive (e.g.,on average) at a certain threshold, or based on an analysis of any other relevant metric with regard to a traffic engineering policy desired by the autonomous system administrator. As illustrated in Figure 6, the saturation states of the target links of the autonomous system at a current time can be represented in a vector form, each component of the vector being associated with a particular link of the target links, a value of 0 indicating that the link in question is not saturated (case of the L2 and L links). n of figure 6) and a value of 1 indicating on the contrary that the link in question is indeed saturated (case of the Li and L3 links in figure 6).

[0064] Starting from a current observation, that is, current state data D_ETA Ccomprising data representing a current load state of the target links and data representing a current saturation state of the target links (taking, for example, respectively, the form of a matrix according to Figure 5 and a vector according to Figure 6), the AG agent determines at least one ACT action to be executed on the ENV learning environment, an action being associated with a current routing path of a data stream randomly selected from a set of target data streams within the autonomous system.

[0065] More specifically, the determined action is selected from: maintaining the current routing path of the target data flow; or modifying the current routing path of the target data flow.

[0066] When the determined action corresponds to maintaining the current routing path of the target data flow, this routing path must not be modified. The AG agent then transmits instructions to this effect to the ENV learning environment, or at the very least does not transmit any instructions to modify the current routing path (such an absence of instructions being in itself characteristic of an action to maintain the current routing path).

[0067] When the determined action corresponds to a modification of the current routing path of the target data flow, two main cases may arise.

[0068] In the first scenario, the current routing path corresponds to a path established according to the routing protocol in effect within the autonomous system, as deployed and implemented within the various routers of the autonomous system. Modifying the current routing path of the target data flow then involves creating, alongside this routing protocol, a bypass path (typically an IP tunnel) between an entry point and an exit point of said current routing path, and rerouting said target data flow onto said bypass path.

[0069] In a second scenario, the current routing path corresponds to a bypass path (typically an IP tunnel) defined outside the scope of a routing protocol in effect within the autonomous system. Modifying the current routing path of the target data flow then involves removing this previously established bypass path. More specifically, such removal results in the rerouting of the target data flow along a path defined by the aforementioned routing protocol, as deployed and implemented within the various routers of the autonomous system.

[0070] The execution of the specified actions by agent AG leads to a new state of the learning environment ENV, associated with updated state data D_ETA Sincluding data representing an updated load state of the target links and data representing an updated saturation state of the target links (taking, for example, respectively the form of a matrix according to Figure 5 and a vector according to Figure 6).

[0071] Optionally, in a particular embodiment, as illustrated in relation to Figure 4, the agent's determination of the actions to be executed on the learning environment takes into account, in addition to the current state data D_ETA C , of so-called previous state data D_ETA P Such previous state data D_ETA PThese correspond to the data used as current state data during at least one previous iteration (for example, during the immediately preceding iteration) of the data exchanges between the agent and the learning environment. This data may have been stored, for example, in the memory of the device responsible for implementing the process according to this technique. The previous state data D_ETA P These include, for example, data representing a previous load state of the target links and data representing a previous saturation state of the target links (taking, for example, the form of a matrix as shown in Figure 5 and a vector as shown in Figure 6, respectively). Such an embodiment is advantageous because it allows the AG agent to have at least two observations (at least one previous observation represented by the previous state data D_ETA). Pand a current observation represented by the current state data D_ETA C ) corresponding to at least two successive states of the autonomous system, that is, having access to data allowing it to assess an evolution of the autonomous system. Thus, the agent is able, in particular, to learn more effectively to detect a risk of future saturation of a target link, even before such saturation occurs, based, for example, on a significantly increasing load rate observed on the link in question, between the previous observation and the current observation of the autonomous system.

[0072] In addition, and based on the updated D_ETA status data SReward data (D_REC) is calculated in response to the execution of at least one ACT (Action Token) on the ENV (Environmental Virtualization). This reward data typically takes the form of a simple numerical value r, usually a real number, whose value is directly correlated to the number of links still saturated in the autonomous system after the ACT actions have been implemented. For example, the value r can be defined using the formula r = -wlrl, where rl is the number of saturated links and wl is a positive weight. Thus, the greater the number of saturated links, the lower the value of the reward r. In other words, the reinforcement learning implemented using this technique allows an agent to learn by receiving rewards if it manages to desaturate one link without saturating another.Depending on a particular characteristic, the value r can alternatively be the result of a sum r = -wlrl - w2r2 in which, in addition to the number rl of saturated links, the number r2 of workarounds still existing after the implementation of the ACT action is also taken into account. Indeed, the creation and maintenance of a workaround are operations that have a computational cost, so it seems legitimate to try to minimize the number of workarounds present in the autonomous system. The positive weights wl and w2 then allow us to weight the relative importance given respectively to the number of saturated links rl or the number of workarounds still existing r2 in the reward calculation. In this way, the agent will seek during learning to maximize its future rewards, thus learning a policy (i.e.a function that indicates what action to take given an observed state of the environment) aimed at jointly minimizing intra-domain saturation and the number of bypass paths in the autonomous system.

[0073] Thus, at the end of the learning phase 31, at least one agent has been trained on a large number of trials - in other words, the weights of a model associated with the agent have been progressively adjusted during the trials - so that the agent is able to determine alternative paths to reroute at least some data flows within the autonomous system in case of risk or occurrence of saturation phenomena on target links of the autonomous system.

[0074] Following the training phase 31, an inference phase 32 is implemented, in which at least one agent trained during the training phase is deployed in production and used to determine and deliver alternative paths for data flows within the autonomous system (AS), in order to minimize saturation on all or part of the existing links in this system (and, potentially, to minimize the number of existing bypass paths in the autonomous system as much as possible). In one particular embodiment, an agent using this technique can, for example, be deployed at the level of an entity overseeing at least one router of the autonomous system. Alternatively, several agents using this technique can also be deployed in a distributed manner across multiple routers of the autonomous system.During this inference phase 32, the agent, thanks to its prior training, selects for at least one target data stream the actions determined to be most appropriate for this objective. These actions include maintaining the current routing path of the target data stream, creating a bypass path for rerouting the target data stream, or removing a bypass path currently used by the target data stream so that it is rerouted along a path defined by the routing protocol in effect within the autonomous system. Figure 7 illustrates an example of the result of implementing the proposed technique on an autonomous system (AS) in an initial situation as shown in Figure 2 previously.More specifically, as seen in Figure 7, the data stream fl initially routed on the path R1-R2-R5-R9 is rerouted on an alternative bypass path R1-R2-R6-R9 specifically created for this purpose on instructions from the agent, in order to avoid the saturation phenomenon initially observed on the R2-R5 link.

[0075] In specific embodiments, various approaches described below may be adopted, optionally and potentially in a complementary manner, to, for example, improve the performance and / or robustness of the process for determining at least one alternative path for routing at least one data stream using this technique. In particular, various techniques for configuring the learning model and / or the way it is trained may be used. By way of illustration and without limitation, such a learning model may, in particular, be a Deep Q-Network (DQN) model.

[0076] According to one approach, it is possible to integrate into the model a discount factor y, between 0 and 1, allowing control over the extent to which the model focuses on distant future rewards. In this particular case, the discount factor y is set to a low value, for example y = 0.1, so that the model concentrates on the immediate future, i.e., on a relatively immediate gain. Such a configuration is indeed well-suited to the present technique, since the actions performed by the agent are associated with operations to maintain or modify current data flow routing paths, which therefore have an immediate impact on the load rate and ultimately on the saturation of the target links of the autonomous system.

[0077] According to a second approach, an Epsilon-Greedy strategy can be implemented to ensure a proper balance between exploration and exploitation. This involves, for example, preventing the agent from simply selecting actions associated with a high future cumulative reward without ever attempting to discover new sequences of actions that could potentially be better for the ultimate goal (e.g., minimizing the occurrence of saturation on certain links). It is therefore proposed to integrate into the model the consideration of an exploration factor E, between 0 and 1, allowing control over the proportion of exploration and exploitation to be implemented during the learning phase.According to a particular characteristic, in this case, the exploration factor e is initialized to e = 1 at the very beginning of learning and configured to decrease, for example linearly, to e = 0.1 at the end of learning. In this way, the agent selects an action to execute randomly or quasi-randomly at the beginning of the learning phase, and then, as learning progresses, gradually selects actions associated with a high future cumulative reward. In other words, learning gradually evolves from an initial strategy of pure exploration to a final strategy of pure exploitation.

[0078] A third approach employs replay memory techniques. This involves storing the agent's past experiences—an experience being defined as a set of associated data including current state data (i.e., an observation at time t), an action selected based on that observation, and subsequent reward and state data (i.e., an observation at time t+1) obtained in response to that action—and using these experiences randomly during the learning phase, with the aim of breaking the correlation between consecutive learning experiences. In a particular feature of this technique, a memory capable of storing 1000 past experiences is used.

[0079] According to a fourth approach, techniques based on the use of a target network in addition to the learning network are implemented. More specifically, the target network is a time-shifted copy of the learning network, whose weights are thus periodically fixed and lagging behind those of the learning network, with the aim of stabilizing the learning process. A particular feature of this technique is that the target network is updated every 100 steps. In another aspect, the proposed technique also relates to a device for determining at least one alternative path for routing at least one data stream within an autonomous system. Such an electronic device is capable of performing the process described above in any of its embodiments.More particularly, such a device according to the present technique comprises, in a particular embodiment: means for implementing a reinforcement learning type training phase, in which at least one agent is trained to determine at least one routing path for said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent and a reinforcement learning environment, based on at least one data point representing a saturation or a risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system; means for implementing an inference phase, in which said at least one trained agent is used to deliver said at least one alternative path.

[0080] Figure 8 schematically and in a simplified manner represents the structure of such an electronic device in a particular embodiment. The device, according to the proposed technique, comprises for example a memory 81 consisting of a buffer memory M, a processing unit 82, equipped for example with a microprocessor pP, and controlled by the computer program Pg 83, implementing steps of the process of determining at least one alternative path for the routing of at least one data stream within an autonomous system, according to at least one embodiment of the invention.To this end, the electronic device also includes, in a particular embodiment, at least one communication interface (for example, an Ethernet communication interface) and / or at least one human-machine interface, enabling it to receive and transmit data to and from other equipment in a communication network to which it is connected and / or from human operators.

[0081] At initialization, the code instructions of computer program 83 are loaded into the buffer before being executed by the processor of processing unit 82. Processing unit 82 receives, for example, configuration data as input E, which it uses to instantiate a reinforcement learning environment. The microprocessor of processing unit 82 then performs the following steps in the process of determining at least one alternative path, according to the instructions of computer program 83. More specifically, processing unit 82 instantiates at least one agent and trains it—based on data exchanges with the reinforcement learning environment—to generate alternative paths for routing at least one data stream within an autonomous system, which minimize the occurrence of saturation phenomena.At the end of this training, the trained agent is then used to generate at least one alternative path for the routing or rerouting of at least one data stream within an autonomous system, which is delivered by the processing unit 82 at output S.

Claims

1. DEMANDS 1. Method for determining at least one alternative path for the routing of at least one data stream within an autonomous system implementing an IP routing protocol, in the event of saturation or risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of said autonomous system, said method being characterized in that it comprises: a reinforcement learning training phase (11) in which at least one agent (AG) is trained to determine at least one routing path of said at least one data stream within said autonomous system, on the basis of at least one data exchange between said at least one agent (AG) and a reinforcement learning environment (ENV), as a function of at least one data representative of said saturation or risk of saturation;an inference phase (12), in which said at least one trained agent is implemented within said autonomous system and is used to deliver said at least one alternative path.

2. A method according to claim 1, characterized in that said at least one data exchange comprises: the execution, by said at least one agent, of at least one action on said learning environment, determined based on current state data comprising data representative of a current load state of said target links and data representative of a current saturation state of said target links, said action being associated with at least one current routing path of at least one target data stream within said autonomous system; the obtaining, by said at least one agent, in response to said execution, of updated state data and reward data from said learning environment.

3. Method according to claim 2, characterized in that the determination of said action performed on the learning environment takes into account, in addition to said current state data, previous state data, including data representing a previous load state of said target links and data representing a previous saturation state of said target links.

4. A method according to claim 3, characterized in that said data representing a previous state of charge of said target links and said data representing a current state of charge of said target links each take the form of a matrix, one dimension among rows and columns of said matrix being associated with said target links, the other dimension among rows and columns of said matrix being associated with said target flows, each coefficient of said matrix being associated with both one of said target links and one of said target data flows, and having as its value a load rate of said target link due to said target data flow.

5. Method according to claim 3, characterized in that said data representing a previous state of saturation of said target links and said data representing a current state of saturation of said target links each take the form of a vector, each component of said vector being associated with one of said target links and having as its value a first value representing an absence of saturation of said target link or a second value representing a saturation of said target link.

6. Method according to claim 2, characterized in that said action associated with at least one current routing path of at least one target data stream is selected from: maintaining said current routing path of said target data stream; or modifying said current routing path of said target data stream.

7. Method according to claim 6, characterized in that said modification of the current routing path of said target data stream comprises, where said current routing path corresponds to a path established according to a routing protocol in force within said autonomous system, the creation, alongside said routing protocol, of a bypass path between an entry point and an exit point of said current routing path, and the rerouting of said target data stream on said bypass path.

8. Method according to claim 6, characterized in that said modification of the current routing path of said target data stream comprises, where said current routing path corresponds to a bypass path defined in the margin of a routing protocol in force within said autonomous system, the removal of said bypass path, and the rerouting of said target data stream according to a path defined according to said routing protocol.

9. Method according to claim 2, characterized in that said reward data are determined as a function of a current number of saturated links within said autonomous system.

10. Method according to claim 2, characterized in that said reward data are further determined as a function of a current number of bypass paths used to route said at least data stream within said autonomous system.

11. A method according to any one of the preceding claims, characterized in that said alternative path is a bypass path created between said entry point and said exit point of said autonomous system.

12. A method according to any one of the preceding claims, characterized in that said alternative path is a path established according to a routing protocol in force within said autonomous system, following the removal of a bypass path previously created between said entry point and said exit point of said autonomous system.

13. An electronic device for determining at least one alternative path for routing at least one data stream within an autonomous system, said device being characterized in that it comprises at least one processor configured to implement: a reinforcement learning training phase, in which at least one agent is trained to determine at least one routing path for said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent and a reinforcement learning environment, as a function of at least one data point representing a saturation or a risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system; an inference phase, in which said at least one trained agent is used to deliver said at least one alternative path.

14. Product computer program downloadable from a communication network and / or stored on a computer-readable medium and / or executable by a microprocessor, characterized in that it includes program code instructions for the execution of a process according to any one of claims 1 to 12, when executed by a computer.

Citation Information

Patent Citations

  • Resource optimization methods and systems based on deep reinforcement learning under SDN architecture

    CN113518039B

  • Autonomous traffic (self-driving) network with traffic classes and passive and active learning

    US20230145097A1