Method for determining at least one optimized path for routing at least one data stream within an autonomous system.

A reinforcement learning-based method for optimizing data stream paths in autonomous systems addresses inefficiencies in manual rerouting, automating the detection and prevention of saturation, enhancing service quality and reducing operational errors.

FR3164336A1Pending Publication Date: 2026-01-09ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
FR2024007391
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-05
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing methods for managing traffic routing in autonomous systems are inefficient and prone to errors, leading to saturation phenomena that result in traffic loss and service degradation, due to the complexity and time-consuming nature of manually detecting and rerouting congested links.

Method used

A reinforcement learning-based method for determining optimized data stream paths within an autonomous system, using an agent trained on a simulation environment to explore and create bypass paths, minimizing saturation by dynamically adjusting routing based on load and saturation states of links.

Benefits of technology

Enables real-time or near-real-time adjustment of data flow paths to prevent or reduce saturation, ensuring higher quality of service by automating the detection and rerouting process, thus reducing operational errors and resource costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for determining at least one optimized path for routing at least one data stream within an autonomous system. The invention relates to a method for determining at least one optimized path for routing at least one data stream within an autonomous system.Such a process comprises: - a reinforcement learning training phase (11), in which an agent (AG) is trained to determine at least one routing path for said at least one data stream within said autonomous system, based on at least one data exchange between said agent (AG) and a reinforcement learning environment (ENV), according to at least one data point representing a saturation or risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system; - an inference phase (12), in which said trained agent is used to deliver said at least one optimized path. Abstract figure: Figure 3.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method for determining at least one optimized path for routing at least one data stream within an autonomous system. Technical field

[0001] The invention relates to the field of data flow routing within a communication network. More particularly, the invention relates to techniques aimed at reducing or preventing the occurrence of saturation phenomena on the links used to route traffic within an autonomous system of a communication network. Previous art

[0002] A wide area communication network, such as the Internet, generally comprises numerous autonomous systems. An autonomous system can be defined as a set of subnets and routers managed by the same administrative authority (e.g., a state, an organization, a company such as a service provider, etc.). In order to maintain overall accessibility and connectivity within the wide area network, including the end-to-end routing of data flows from a source device to a destination device that does not necessarily belong to the same autonomous system, the various autonomous systems of the wide area communication network are interconnected.

[0003] Each autonomous system also decides locally, based on its own internal routing policy, the best paths for the data flows passing through it. Therefore, a path taken by a data flow in a "forward" direction, from the first device to the second device in the network, is generally not the same as the path taken by a data flow in a "return" direction, from the second device to the first device. The paths through which data flows are routed in a wide area communication network such as Intermet are thus asymmetric, and an autonomous system cannot predict through which neighbors or entry points a data flow will be delivered.The widespread adoption of content caching within increasingly numerous content delivery networks, for example to allow users to access content more quickly, further exacerbates this phenomenon, leading to the development of increasingly complex and dynamic content delivery strategies.

[0004] One consequence of the difficulty in estimating and anticipating the amount of traffic to be routed within an autonomous system is the risk of the occurrence of phenomena Intra-domain saturation can occur, particularly when the amount of traffic carried on one or more links within the autonomous system becomes too close to, equal to, or greater than the maximum capacity of the link(s) in question. If these saturation phenomena persist over time, they can lead to traffic loss (e.g., lost IP packets that never reach the final destination equipment) and a degradation of service quality, with a particularly disappointing effect for users.

[0005] A monitoring team that observes intra-domain congestion may attempt to offload the congested link(s) by rerouting certain data flows over other paths within the autonomous system, for example, via temporary bypass paths specifically created for this purpose. However, detecting congested links, identifying the data flows routed through these links, creating potential bypass paths, analyzing the effects of rerouting traffic through these bypass paths, and removing the bypass paths once the congestion episode has ended are currently performed manually, which is time-consuming, complex, and tedious, and whose implementation is therefore potentially prone to numerous errors.

[0006] There is therefore a need for a solution that makes it possible to protect more effectively against the risks of saturation of traffic routing links within an autonomous system. Summary of the invention

[0007] The present technique provides a solution to overcome certain drawbacks of the prior art. In one aspect, the present technique relates to a method for determining at least one optimized path for routing at least one data stream within an autonomous system. More specifically, such a method comprises:

[0008] - a reinforcement learning-type training phase, in which at least one agent is trained to determine at least one routing path of said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent and a reinforcement learning environment, based on at least one data representative of a saturation or risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system;

[0009] - an inference phase, in which said at least one trained agent is used to deliver said at least one optimized path.

[0010] In this way, the present technique offers an ingenious solution based on reinforcement learning techniques to determine optimized paths for routing data flows within an autonomous system. Thus, a real-time or near-real-time adjustment of the paths used to route all or part of the data flows in transit within the autonomous system can be performed, which makes it possible to avoid, or at least reduce, the occurrence of saturation phenomena, and therefore to offer greater guarantees in terms of quality of service.

[0011] In a particular embodiment, said at least one data exchange comprises:

[0012] - the execution, by said at least one agent, of at least one action on said learning environment, determined based on current state data including data representative of a current load state of said target links and data representative of a current saturation state of said target links, said action being associated with at least one current routing path of at least one target data stream within said autonomous system;

[0013] - obtaining, by said at least one agent, in response to said execution, data updated status and reward data from said learning environment.

[0014] In this way, the proposed technique allows, using a trial-and-error approach, the exploration and testing, without operational impact, of routing target data flows via different paths, or even the creation of specific bypass paths for such routing, based on observational data from a training environment. Thus, an agent is able to learn, before its deployment in production, to identify optimal data flow routing paths within an autonomous system, which minimize saturation phenomena.

[0015] In a particular embodiment, the determination of said action executed on the learning environment takes into account, in addition to said current state data, previous state data, including data representing a previous load state of said target links and data representing a previous saturation state of said target links.

[0016] In this way, the agent has additional information on the basis of which to determine an action to be performed on the learning environment. In particular, since the agent thus has access, in addition to an observation at a current time, to an observation at a time preceding the current time, it has the ability to assess changes or trends such as increases in workload important links for example, which allows it in particular to anticipate risks of saturation before such saturations become actual.

[0017] According to a particular feature, said data representing a previous state of charge of said target links and said data representing a current state of charge of said target links each take the form of a matrix, one dimension among rows and columns of said matrix being associated with said target links, the other dimension among rows and columns of said matrix being associated with said target flows, each coefficient of said matrix being associated with both one of said target links and one of said target data flows, and having as its value a load rate of said target link due to said target data flow.

[0018] According to a particular feature, said data representing a previous state of saturation of said target links and said data representing a current state of saturation of said target links each take the form of a vector, each component of said vector being associated with one of said target links and having as its value a first value representing an absence of saturation of said target link or a second value representing a saturation of said target link.

[0019] In this way, a simple yet comprehensive mathematical representation of the observations made on the training environment is obtained, in the form of matrices and / or vectors that are easily manipulated within a learning model. Furthermore, having a separate variable to indicate whether a link is saturated or not (i.e., the saturation state), distinct from the one used to define the link load (i.e., the load state), offers considerable flexibility in implementing the method according to the present technique. For example, it allows an administrator of the autonomous system to define the conditions and criteria under which a link is considered saturated or not, with these criteria potentially going beyond simply considering a high link load.

[0020] In a particular embodiment, the action associated with at least one current routing path of at least one target data stream is selected from:

[0021] - a maintenance of said current routing path of said target data stream; or

[0022] - a modification of said current routing path of said data stream target.

[0023] In this way, a current routing path of a target data stream is not necessarily modified, if the agent determines on the basis of observations of the learning environment obtained that a modification of the current routing path of this data stream has little or no impact on the occurrence of saturation phenomena within the autonomous system.

[0024] According to a particular feature, said modification of the current routing path of said target data stream includes, where said current routing path corresponds to a path established according to a routing protocol in force within said autonomous system, the creation, alongside said routing protocol, of a bypass path between an entry point and an exit point of said current routing path, and the rerouting of said target data stream on said bypass path.

[0025] In this way, the proposed technique makes it possible in certain situations to bypass the routing protocol in force within the autonomous system, by allowing the creation of bypass paths (or tunnels) outside the protocol, for example when it is determined that a current routing path established according to the classic routing protocol is not optimal in order to avoid the occurrence of saturation phenomena within the autonomous system.

[0026] According to a particular feature, said modification of the current routing path of said target data stream includes, where said current routing path corresponds to a bypass path defined in the margin of a routing protocol in force within said autonomous system, the removal of said bypass path, and the rerouting of said target data stream along a path defined according to said routing protocol.

[0027] In this way, a bypass path (or tunnel) previously created outside the routing protocol in effect within the autonomous system can be automatically deleted when it no longer constitutes an optimal path for avoiding or limiting the occurrence of saturation phenomena within the autonomous system. This frees up the computing resources used to maintain this bypass path. Consequently, the system reverts to the usual methods for routing the data flow, in accordance with the routing protocol in effect within the autonomous system. This reduces the number of operations managed outside a typical autonomous system configuration, the implementation of which is potentially error-prone and costly in terms of computing power.

[0028] In a particular embodiment, said reward data is determined based on a current number of saturated links within said autonomous system.

[0029] In this way, we have a simple metric to assess the extent to which an autonomous system is impacted by saturation phenomena likely to lead to degradations in quality of services.

[0030] According to a particular feature, said reward data are further determined as a function of a current number of bypass paths used to route said at least data stream within said autonomous system.

[0031] In this way, the cost of maintaining bypass paths or tunnels for routing data flows within the autonomous system is also taken into account for determining the optimized paths according to the present technique.

[0032] In another aspect, the present technique relates to an electronic device for determining at least one optimized path for routing at least one data stream within an autonomous system. Such a device includes means (for example, at least one processor configured for this purpose) for implementation:

[0033] - of a reinforcement learning-type training phase, in which at least one agent is trained to determine at least one routing path of said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent and a reinforcement learning environment, based on at least one data representative of a saturation or risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system;

[0034] - of an inference phase, in which said at least one trained agent is used to deliver said at least one optimized path.

[0035] Such an electronic device may of course exhibit the various characteristics relating to the determination method according to the invention, which can be combined or considered separately. Thus, the characteristics and advantages of this device are the same as those of the method for determining at least one optimized path for routing at least one data stream within an autonomous system, and are not detailed further.

[0036] According to another aspect, the proposed technique also relates to a computer program product downloadable from a communication network and / or stored on a computer-readable medium and / or executable by a microprocessor, comprising program code instructions for the execution of a method for determining at least one optimized path as described above in any of its embodiments, when executed on a computer.

[0037] The proposed technique also relates to a computer-readable recording medium on which is recorded a computer program comprising program code instructions for executing the steps of the process as described above, in any of its embodiments.

[0038] Such a recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a USB flash drive or a hard drive.

[0039] On the other hand, such a recording medium can be a transmissible medium such as an electrical or optical signal, which can be transmitted via an electrical or optical cable, by radio, or by other means, so that the computer program it contains is executable remotely. The program according to the invention can, in particular, be uploaded to a network, for example, the Internet.

[0040] The different embodiments mentioned above can be combined with each other for the implementation of the invention. Figures

[0041] Other features and advantages of the invention will become more apparent upon reading the following description of a preferred embodiment, given by way of simple illustrative and non-limiting example, and the accompanying drawings, among which:

[0042] [Fig-1] presents an example of an autonomous system in which the present technique can be implemented, in a particular embodiment of the proposed technique;

[0043] [Fig.2] presents an example of an autonomous system exhibiting intra-domain saturation;

[0044] [Fig.3] illustrates the general principle of a method for determining at least one optimized path for routing at least one data stream within an autonomous system, in a particular embodiment of the proposed technique;

[0045] [Fig.4] presents the data exchanges carried out between an agent and a reinforcement learning environment for training said agent, in a particular embodiment of the proposed technique;

[0046] [Fig.5] presents an example of representation of a load state of target links within an autonomous system, in a particular embodiment of the proposed technique;

[0047] [Fig.6] presents an example of representation of a saturation state of target links within an autonomous system, in a particular embodiment of the proposed technique;

[0048] [Fig.7] presents an example of an autonomous system after implementation of a method for determining at least one optimized path, in a particular embodiment of the proposed technique;

[0049] [Fig. 8] describes a simplified architecture of an electronic device for implementing the proposed technique, in a particular embodiment. Detailed description of the invention

[0050] The present application makes it possible to remedy at least some of the aforementioned drawbacks.

[0051] In all figures in this document, elements and steps of the same nature are designated by the same reference.

[0052] According to a first aspect, the present technique relates to a method for determining at least one optimized path for routing at least one data stream within an autonomous system implemented in a communication network. An example of an autonomous system (AS) in which the present technique can be implemented is described by way of illustration with reference to [Fig. 1].

[0053] Such an autonomous system (AS) comprises a plurality of devices, typically routers (nine routers RI to R9 in the example in [Fig. 1]), responsible for routing traffic within the system. These routers include edge routers (routers RI, R2, R3, R5, and R9 in the example in [Fig. 1]) through which traffic to or from other autonomous systems (or other entities in the communication network) enters or leaves the autonomous system (AS).

[0054] Data streams are routed within the autonomous system SA via paths defined by a routing protocol implemented within each router of the system. For example, a data stream fl enters the autonomous system SA at an entry point formed by router RI, then is successively transmitted hop by hop from router RI to router R2, from router R2 to router R5, from router R5 to router R9, with router R9 acting as the exit point directing the data stream fl to another entity in the communication network (e.g., another autonomous system). The routing path R1-R2-R5-R9 of the data stream fl, represented by arrows in [Fig. 1], is thus formed by a set of links (the R1-R2 link, the R2-R5 link, the R5-R9 link) taken successively, from the entry point RI of the autonomous system SA to the exit point R9 of the autonomous system.

[0055] At any given moment, the autonomous system SA is thus potentially traversed by a plurality of data streams, each routed from an entry point to an exit point of the system, via a routing path defined according to the routing protocol implemented at the level of the autonomous system's routers. Each path is formed by a succession of links between routers or other equipment of the autonomous system. As illustrated in [Fig. 2], the routing of multiple data streams within the autonomous system SA means that some links of the autonomous system are more heavily loaded than others during certain time periods, for example, because they are used simultaneously by several data streams (for clarity, only one data stream fl is referenced in [Fig. 2], but it is understood that other data streams not shown in this figure are potentially routed at the same time as the fl stream in the autonomous SA system). For example, with a representation according to [Fig. 2], in which the thickness of a line associated with a link represents a load factor of the link (in other words, the thicker the line, the more loaded the link), we observe, for example, that the R8-R7 link is more loaded than the R7-R9 link, which itself is more loaded than the R5-R6 link. In a critical case, when the load factor of a link becomes too high, that is, too close to, equal to, or even greater than the maximum capacity of the link, saturation phenomena, defined later in this document, can occur. In the example of [Fig.[2], such a saturation phenomenon is observed on the R2-R5 link, where it is illustrated as a very thick link with hatching. The saturation of this link results in disruptions in the routing of data streams using it, such as the fl data stream for example (packet loss, service degradation, etc.).

[0056] When an occurrence or risk of occurrence of such saturation phenomena is detected, an objective of the present technique is therefore to determine, for at least some of the target data flows routed in the autonomous system, so-called "optimized" paths, in that they allow, via the rerouting of the target data flows on their respective optimized paths, to minimize these saturation phenomena on at least some of the target links of the autonomous system (for example by eliminating these phenomena, reducing them, or limiting their duration over time).

[0057] The general principle of the proposed method for determining at least one optimized path for routing at least one data stream within an autonomous system is now illustrated with reference to [Fig. 3], in a particular embodiment. This method is implemented, for example, within an electronic device described later in this document.

[0058] In step 31, a reinforcement learning training phase is implemented using a previously instantiated and configured reinforcement learning environment. More specifically, the learning environment is configured based on data collected on the autonomous system that is the subject of the process according to this technique. This collected data includes, in particular, numerous data relating to the topology, configuration, and use of the autonomous system, such as, for example, the number of routers present in the autonomous system, the number of defined links, the capacity of these links, and the routing protocol in force within the autonomous system. The profile of data flows within the autonomous system, etc. Such data can, for example, be entered, in whole or in part, by a human operator using a human-machine interface provided by the electronic device responsible for implementing the optimized path determination process described herein, or it can be automatically obtained from various data sources via a communication network to which the electronic device is connected. Thus, the training environment forms a simulation environment as close as possible to the autonomous system as deployed in production. For example, the training environment can be implemented as a digital twin of the target autonomous system.

[0059] During the training phase 31, at least one agent (instantiated, for example, by the electronic device responsible for implementing the process) is trained to determine paths to be taken by at least one data stream to traverse the autonomous system from an entry point to an exit point associated with the data stream in question, while respecting pre-defined constraints. More specifically, the data stream routing paths within the autonomous system are determined taking into account at least one piece of data representing a saturation or a risk of saturation of at least one link among a plurality of target links existing in the autonomous system. As described later, these target links may include all the links of the autonomous system, or only a portion of the links of the autonomous system that an administrator of the autonomous system wishes to protect more specifically (i.e.to protect as much as possible from any risk of saturation).

[0060] During this reinforcement learning training phase, using a trial-and-error approach, multiple iterations of data exchange are carried out between the agent and the learning environment, with a view to broadly exploring the results obtained in response to different actions or sequences of actions detailed later.

[0061] Each of these exchanges comprises, according to the usual vocabulary in reinforcement learning and as illustrated in relation to [Fig. 4]:

[0062] - the execution, by an AG agent, of at least one ACT action on the environment ENV learning, based on current state data D_ETAC;

[0063] - the receipt, by agent AG, in response to the execution of said ACT actions, of updated state data (or subsequent state data) D_ETAS and reward data D_REC from the learning environment ENV.

[0064] The state data D_ETA can be considered as observations of the learning environment ENV. Thus, the current state data D_ETAC are representative of a current state of the learning environment by ENV reinforcement, and the updated state data D_ETAS, represent a new state of the ENV reinforcement learning environment following the execution of ACT actions on that environment. As illustrated in [Fig. 4], the updated state data D_ETAS becomes the current state data D_ETAC for the next iteration of a data exchange as previously described.

[0065] The reward data D_REC represents one or more metrics of the effectiveness of the actions ACT performed on the ENV environment, with respect to the objective pursued. More specifically, the successive actions executed on the learning environment are determined by a reinforcement algorithm, so as to tend to maximize over time the reward data D_REC received in response to these actions.

[0066] Within the framework of this technique, the objective is to prevent saturation phenomena from occurring (or at least to limit such phenomena) on one or more target links, and / or in relation to one or more target data streams, within the autonomous system. The target links may include all existing links within the autonomous system, or conversely, only a subset of these links, previously identified as requiring monitoring and on which an administrator of the autonomous system, for example, has a particular interest in protecting against any saturation phenomena.Following a similar approach, target data flows can include all data flows routed within the autonomous system, or conversely only a part of these data flows, previously identified as needing monitoring and for which an administrator of the autonomous system has, for example, a particular interest in protecting against any saturation phenomenon.

[0067] The D_ETA state data allows the AG agent to obtain at least a partial view (i.e. an "observation") of the state of the ENV learning environment, before and after execution of an ACT action on this environment.

[0068] In a particular embodiment of the present technique, these D_ETA state data include data representing a load state of the target links and data representing a saturation state of the target links.

[0069] The load state of a link corresponds to a load factor (or occupancy rate) of the link. Typically, such a variable can be normalized to take values ​​between 0 and 1 (or between 0% and 100%), with a load factor of 0% indicating that no data flows are passing through the link, and a load factor of 100% indicating that the entire bandwidth of the link (i.e., the entire capacity of the link) is being used to carry data flows. Depending on a particular characteristic, the load state of a link can be broken down by data flow, thus providing a view, for the link in question, of the contribution of each data flow. borrowing this link at the total link load rate. It can thus be determined, for example, for a link Li borrowed by three data streams Fb, F2, and F3 at a current time, that the data stream Fi is responsible for 60% of the current load of link Lb, that the data stream F2 is responsible for 12% of the current load of link Li, and that the data stream F3 is responsible for 17% of the current load of link Li, resulting in a total current load rate of link Li of 89%. As illustrated in [Fig. 5], the load states of the target links of the autonomous system at a current time can be represented in matrix form. More specifically, each row of the matrix is ​​associated with a particular link among n target links (Lb, L2, L3, ..., Ln), and each column of the matrix is ​​associated with a particular data stream among m target data streams (Fb, F2, F3, F4, ..., Fm), the value of the matrix at the intersection of a row and a column corresponding to the share of the load rate of the link considered due to the data flow considered (expressed as a percentage in the example of [Fig.5]). .

[0070] The saturation state of a link is a data point defining whether the link is considered saturated or not, with regard to predefined criteria. The saturation state of a link is, for example, represented by a boolean, a first value of the boolean (e.g. the value "0" or the value "false") being representative of an absence of saturation of the link in question, and a second value of the boolean (e.g. the value "1" or the value "true") being representative of a saturation of the link in question.Although it generally derives directly from the load state of a link, it is particularly advantageous, as proposed in this technique, to use a variable other than the load state—in this case, the saturation state—to define whether a link is saturated or not. This allows saturation to be assessed in different ways (for example, according to specific criteria chosen by the autonomous system administrator), thus offering flexibility in implementing the method according to this technique. Depending on various alternative specific characteristics, a link can, for example, be considered saturated when the total load rate of the link exceeds a certain threshold (for example, above 70%, 80%, 90%, or 95% load), or when propagation delays on the link become excessive (e.g.,on average) at a certain threshold, or based on an analysis of any other relevant metric with regard to a traffic engineering policy desired by the autonomous system administrator. As illustrated in [Fig. 6], the saturation states of the target links of the autonomous system at a current time can be represented in vector form, each component of the vector being associated with a particular link of the target links, a value of 0 indicating that the link in question is not saturated (case of the L2 and Ln links in [Fig. 6]) and a value of 1 indicating . on the contrary that the link in question is actually saturated (case of the Li and L3 links in [Fig.6]).

[0071] From a current observation, i.e. current state data D_ETAC comprising data representing a current load state of the target links and data representing a current saturation state of the target links (taking for example respectively the form of a matrix according to [Fig.5] and a vector according to [Fig.6]), the AG agent determines at least one action ACT to be executed on the ENV learning environment, an action being associated with a current routing path of a target data flow within the autonomous system.

[0072] More specifically, the determined action is selected from:

[0073] - maintaining the current routing path of the target data stream; or

[0074] - a modification of the current routing path of the target data stream.

[0075] When the determined action corresponds to maintaining the current routing path of the target data stream, this routing path must not be modified. The AG agent then transmits instructions to this effect to the ENV learning environment, or at the very least does not transmit any instruction to modify the current routing path (such an absence of instructions being in itself characteristic of an action to maintain the current routing path).

[0076] When the determined action corresponds to a modification of the current routing path of the target data flow, two main cases may arise.

[0077] According to a first case, the current routing path corresponds to a path established according to the routing protocol in force within the autonomous system, as deployed and implemented within the various routers of the autonomous system. Modifying the current routing path of the target data flow then involves creating, alongside this routing protocol, a bypass path (or tunnel) between an entry point and an exit point of said current routing path, and rerouting said target data flow onto said bypass path.

[0078] According to a second case, the current routing path corresponds to a bypass path defined outside the scope of a routing protocol in effect within the autonomous system. Modifying the current routing path of the target data stream then involves removing this previously established bypass path. More specifically, such removal results in the rerouting of the target data stream along a path defined according to said routing protocol, as deployed and implemented within the various routers of the autonomous system.

[0079] The execution, by the AG agent, of the determined actions leads to a new state of the ENV learning environment, associated with updated state data D_ETAS comprising data representing an updated load state of the target links and data representing an updated saturation state of the target links (taking for example respectively the form of a matrix according to [Fig.5] and a vector according to [Fig.6]).

[0080] Optionally, in a particular embodiment, as illustrated in relation to [Fig. 4], the agent's determination of the actions to be performed on the learning environment takes into account, in addition to the current state data D_ETAC, so-called previous state data D_ETAP. Such previous state data D_ETAP corresponds to the data used as current state data during a previous iteration (for example, during the immediately preceding iteration) of the data exchanges carried out between the agent and the learning environment, this data having, for example, been stored in a memory of the device responsible for implementing the process according to the present technique.The previous state data D_ETAP includes, for example, data representing a previous load state of the target links and data representing a previous saturation state of the target links (taking, for example, the form of a matrix according to [Fig. 5] and a vector according to [Fig. 6], respectively). Such an embodiment is interesting because it allows the AG agent to have two observations (a previous observation represented by the previous state data D_ETAP and a current observation represented by the current state data D_ETAC) corresponding to two successive states of the autonomous system; that is, to have access to data allowing it to evaluate an evolution of the autonomous system.Thus, the agent is able, in particular, to learn more effectively to detect a risk of future saturation of a target link, even before such saturation occurs, based, for example, on a significantly increasing load rate observed on the link in question, between the previous observation and the current observation of the autonomous system.

[0081] In addition, and based on the updated state data D_ETAS obtained in response to the execution of at least one ACT action on the ENV environment, reward data D_REC is calculated. This reward data typically takes the form of a simple numerical value r, typically a real number, whose value is directly correlated to the number of links still saturated in the autonomous system after the ACT actions have been implemented. For example, the value r can be defined according to the formula r = -wl rl, where rl is the number of saturated links, and wl a positive weight. Thus, the greater the number of saturated links, the lower the value of the reward r.

[0082] According to a particular feature, the value r can alternatively be the result of a sum r = -wlrl - w2r2 in which, in addition to the number rl of saturated links, the number r2 of bypass paths still existing after implementation of the ACT action is also taken into account. Indeed, the creation and maintenance of a bypass path are operations that have a computational cost, so it seems legitimate to seek to minimize the number of bypass paths present in the autonomous system. The positive weights wl and w2 then allow the relative importance given respectively to the number of saturated links rl or the number of bypass paths still existing r2 in the calculation of the reward to be weighted. In this way, the agent will seek during learning to maximize its future rewards, thus learning a policy (i.e.a function that indicates what action to take given an observed state of the environment) aimed at jointly minimizing intra-domain saturation and the number of workaround paths in the autonomous system.

[0083] Thus, at the end of the learning phase 31, at least one agent has been trained on a large number of trials - in other words, the weights of a model associated with the agent have been progressively adjusted during the trials - so that the agent is able to determine optimal paths according to an optimization criterion to route data flows within the autonomous system while respecting certain constraints.

[0084] Following the training phase 31, an inference phase 32 is implemented, in which at least one agent trained in the training phase is deployed in production, and used to determine and deliver optimized paths to be taken by the data flows transiting within the autonomous system SA, in particular to minimize saturation phenomena on all or part of the links existing in this system (and, possibly, also to minimize as much as possible the number of bypass paths existing in the autonomous system).In other words, during this inference phase 32, the agent selects, for each target data stream, based on its prior training, the actions determined to be the most appropriate for this objective. These actions include maintaining the current routing path of the target data stream, creating a bypass path for rerouting the target data stream, or removing a bypass path currently used by the target data stream so that it is rerouted along a path defined by the routing protocol in effect within the autonomous system. Figure 7 illustrates an example of the result of implementing the proposed technique on an autonomous system (AS) in a [system / platform]. The initial situation is as shown in [Fig.2] previously presented. More specifically, as seen in [Fig.7], the data stream fl initially routed on the path R1-R2-R5-R9 is rerouted on an optimized bypass path R1-R2-R6-R9 specifically created for this purpose on the agent's instructions, in order to avoid the saturation phenomenon initially observed on the R2-R5 link.

[0085] In particular embodiments, various approaches presented below may be adopted, optionally and possibly in a complementary manner, in order, for example, to make the process of determining at least one optimized path for routing at least one data stream according to this technique more efficient and / or more robust. More specifically, various techniques for configuring the learning model and / or the way in which it is trained may be used. By way of illustration and without limitation, such a learning model may, in particular, be a DQN (Deep Q-Network) type model.

[0086] According to a first approach, it is possible to integrate into the model the consideration of a discount factor y, between 0 and 1, allowing control over the extent to which the model focuses on distant rewards in the future. According to a particular feature, in the present case, the discount factor y is set to a low value, for example y = 0.1, so that the model focuses on the immediate future, that is, on a relatively immediate gain. Such a configuration is indeed well suited to the present technique, insofar as the actions performed by the agent are associated with operations to maintain or modify current data flow routing paths, which therefore have an immediate impact on the load factor and ultimately on the saturation of the target links of the autonomous system.

[0087] According to a second approach, an Epsilon-Greedy type strategy can be implemented to ensure a proper balance between exploration and exploitation. This involves, for example, preventing the agent from simply selecting actions associated with a high future cumulative reward without ever attempting to discover new sequences of actions that could potentially prove better with respect to the final objective (e.g., minimizing the occurrence of saturation on certain links). It is therefore proposed to integrate into the model the consideration of an exploration factor e, between 0 and 1, allowing control over the proportion of exploration and the proportion of exploitation to be implemented during the learning phase.According to a particular characteristic, in this case, the exploration factor e is initialized to e = 1 at the very beginning of learning, and configured to decrease, for example linearly, to e = 0.1 at the end of learning. In this way, the agent selects an action to execute randomly or quasi-randomly at the beginning of the learning phase, and then, as learning progresses. Progress gradually selects actions associated with a high future cumulative reward. In other words, learning evolves progressively from an initial strategy of pure exploration to a final strategy of pure exploitation.

[0088] According to a third approach, replay memory techniques are implemented. More specifically, this involves storing the agent's past experiences—an experience corresponding to a set of associated data comprising current state data (i.e., an observation at time t), an action selected based on this observation, and subsequent reward and state data (i.e., an observation at time t+1) obtained in response to this action—and using them randomly during the learning phase, with the aim of breaking the correlation between consecutive learning experiences. In a particular feature of the present technique, a memory capable of storing 1000 past experiences is used.

[0089] According to a fourth approach, techniques based on the use of a target network in addition to the learning network are implemented. More specifically, the target network is a time-shifted copy of the learning network, whose weights are thus periodically fixed and lagging behind those of the learning network, with the aim of stabilizing the learning process. According to a particular feature of the present technique, the target network is updated every 100 steps.

[0090] According to another aspect, the proposed technique also relates to a device for determining at least one optimized path for routing at least one data stream within an autonomous system. Such an electronic device is capable of carrying out the process described above in any one of its embodiments. More particularly, such a device according to the present technique comprises, in a particular embodiment:

[0091] - means of implementing a learning-type training phase by reinforcement, in which at least one agent is trained to determine at least one routing path of said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent and a reinforcement learning environment, as a function of at least one data point representing a saturation or a risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system;

[0092] - means for implementing an inference phase, in which said at least a trained agent is used to deliver said at least one optimized path

[0093] Figure 8 schematically and in a simplified manner represents the structure of such an electronic device in a particular embodiment. The device, according to the proposed technique, comprises, for example, a memory 81 consisting of a buffer memory M, a processing unit 82, equipped, for example, with a microprocessor qP, and controlled by the computer program Pg 83, implementing steps of the process for determining at least one optimized path for routing at least one data stream within an autonomous system, according to at least one embodiment of the invention.To this end, the electronic device also includes, in a particular embodiment, at least one communication interface (for example, an Ethernet communication interface) and / or at least one human-machine interface, enabling it to receive and transmit data to and from other equipment in a communication network to which it is connected and / or from human operators.

[0094] At initialization, the code instructions of the computer program 83 are loaded into the buffer memory before being executed by the processor of the processing unit 82. The processing unit 82 receives as input E, for example, configuration data by means of which it can instantiate a reinforcement learning environment.

[0095] The microprocessor of the processing unit 82 then performs the following steps of the process for determining at least one optimized path, according to the instructions of the computer program 83. More specifically, the processing unit 82 instantiates at least one agent and trains it—based on data exchanges with the reinforcement learning environment—to generate optimized paths for routing at least one data stream within an autonomous system, which minimize the occurrence of saturation phenomena. Following this training, the trained agent is then used to generate at least one optimized path for routing at least one data stream within an autonomous system, which is delivered by the processing unit 82 as output S.

Claims

Demands

1. A method for determining at least one optimized path for routing at least one data stream within an autonomous system, said method being characterized in that it comprises: - a reinforcement learning training phase (11), in which at least one agent (AG) is trained to determine at least one routing path for said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent (AG) and a reinforcement learning environment (ENV), as a function of at least one data point representing a saturation or a risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system; - an inference phase (12), in which said at least one trained agent is used to deliver said at least one optimized path.

2. A method according to claim 1, characterized in that said at least one data exchange comprises: - the execution, by said at least one agent, of at least one action on said learning environment, determined based on current state data comprising data representative of a current load state of said target links and data representative of a current saturation state of said target links, said action being associated with at least one current routing path of at least one target data stream within said autonomous system; - the obtaining, by said at least one agent, in response to said execution, of updated state data and reward data from said learning environment.

3. Method according to claim 2, characterized in that the determination of said action performed on the learning environment takes into account, in addition to said current state data, previous state data, including data representing a previous load state of said target links and data representing a previous saturation state of said target links.

4. A method according to claim 3, characterized in that said data representing a previous state of charge of said target links and said data representing a current state of charge of said target links each take the form of a matrix, one dimension among rows and columns of said matrix being associated with said target links, the other dimension among rows and columns of said matrix being associated with said target flows, each coefficient of said matrix being associated with both one of said target links and one of said target data flows, and having as its value a load rate of said target link due to said target data flow.

5. A method according to claim 3, characterized in that said data representing a previous state of saturation of said target links and said data representing a current state of saturation of said target links each take the form of a vector, each component of said vector being associated with one of said target links and having as its value a first value representing an absence of saturation of said target link or a second value representing a saturation of said target link.

6. A method according to claim 2, characterized in that said action associated with at least one current routing path of at least one target data stream is selected from: - maintaining said current routing path of said target data stream; or - modifying said current routing path of said target data stream.

7. The method according to claim 6, characterized in that said modification of the current routing path of said target data stream comprises, where said current routing path corresponds to a path established according to a routing protocol in force within said autonomous system, the creation, alongside said routing protocol, of a bypass path between an entry point and an exit point of said current routing path, and the rerouting of said target data stream on said bypass path.

8. A method according to claim 6, characterized in that said modification of the current routing path of said target data stream comprises, where said current routing path corresponds to a bypass path defined in margin of a routing protocol in force within said autonomous system, the removal of said bypass path, and the rerouting of said target data flow along a path defined according to said routing protocol.

9. Method according to claim 2, characterized in that said reward data are determined as a function of a running number of saturated links within said autonomous system.

10. A method according to claim 2, characterized in that said reward data are further determined as a function of a current number of bypass paths used to route said at least data stream within said autonomous system.

11. Electronic device for determining at least one optimized path for routing at least one data stream within an autonomous system, said device being characterized in that it comprises at least one processor configured to implement: - a reinforcement learning type training phase, in which at least one agent is trained to determine at least one routing path for said at least one data stream within said autonomous system, based on at least one data exchange between said at least one agent and a reinforcement learning environment, as a function of at least one data representative of a saturation or a risk of saturation of at least one link of a plurality of target links included on at least one path between an entry point and an exit point of the autonomous system;- an inference phase, in which said at least one trained agent is used to deliver said at least one optimized path.

12. Product computer program downloadable from a communication network and / or stored on a computer-readable medium and / or executable by a microprocessor, characterized in that it includes program code instructions for the execution of a process according to any one of claims 1 to 10, when executed by a computer.

Citation Information

Patent Citations

  • Resource optimization methods and systems based on deep reinforcement learning under SDN architecture

    CN113518039B

  • Autonomous traffic (self-driving) network with traffic classes and passive and active learning

    US20230145097A1