Device and method for generating disturbances in a network.

The method addresses the challenge of testing shared network infrastructures by generating disturbances based on configuration data and service level objectives, ensuring robustness and compliance with service level agreements while avoiding performance degradation.

FR3148504B1Active Publication Date: 2025-11-07ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023004418
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-05-03
Publication Date
2025-11-07
Estimated Expiration
2043-05-03

AI Technical Summary

Technical Problem

Existing disruption deployment solutions are primarily applicable to single administrative environments and do not account for constraints of network infrastructures shared between service providers and infrastructure managers, especially in cloud and edge network scenarios where knowledge of these constraints is not shared.

Method used

A method for generating disturbances in a network infrastructure to test its robustness, involving obtaining configuration data, determining maximum unavailability values, and deploying disturbances based on these data, with the option to delegate disruption execution to service providers and using private keys for secure transmission.

Benefits of technology

Enables robustness testing of shared network infrastructures by ensuring compliance with service level objectives and avoiding performance degradation or crashes, leveraging AI-based models for intelligent disruption scenario generation and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000022_0000
    Figure 00000022_0000
  • Figure 00000022_0001
    Figure 00000022_0001
  • Figure 00000023_0000
    Figure 00000023_0000
Patent Text Reader

Abstract

The invention relates to a method for generating disturbances in a network to verify its robustness, said network comprising a plurality of infrastructures and services. The method comprises: - obtaining (E1) a database containing configuration data for the plurality of infrastructures and services; - determining (E2) a maximum unavailability value for said infrastructure and services, and if said maximum unavailability value is positive (E3): - deploying (E4) a new service; - transmitting (E5) disturbance data for the network infrastructure and deployed services into the network, the disturbances being determined from at least the configuration data in the database and the calculated maximum unavailability value; - updating (E6) said maximum unavailability value following the disturbances. Figure for the abstract: Fig. 1.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Device and method for generating disturbances in a network. technical field

[0001] The invention relates to the field of transmitting disturbances in network infrastructures in order to test their robustness. Previous technique

[0002] Disruption deployment solutions exist today, but these scenarios are primarily applicable to cloud service providers (CSPs) or infrastructures deployed within a single environment of the same administrative entity. However, with the deployment of remote applications located in the cloud and the deployment of instances on edge networks, it is not possible to launch disruptions that take into account the constraints of both network infrastructures and service providers, particularly because knowledge of these constraints is not shared between service providers and infrastructure managers.

[0003] A solution enabling the testing of a network infrastructure shared between an infrastructure operator and service providers by sending disturbances is therefore desired. Description of the invention

[0004] The present invention relates to a method for generating disturbances in a network to verify its robustness, said network comprising a plurality of infrastructures and services deployed and hosted in said network, said services being provided by at least one service provider, the method comprising: - obtaining at least one database comprising configuration data for the plurality of infrastructures and services, - determining a maximum unavailability value for said infrastructure and services based on a minimum tolerated service level calculated from service indicators of said infrastructure and services, and if said maximum unavailability value is positive: - the transmission in the data network of disturbances to the network infrastructure and deployed services, the disturbances being determined from at least the database configuration data and the calculated maximum unavailability value, - the update of said maximum unavailability value following disruptions.

[0005] According to some embodiments, the method includes the deployment of a new service, said new service being determined from at least the database configuration data and the calculated maximum unavailability value.

[0006] According to some embodiments, the method includes the delegation of the right to execute disruption elements by at least one of the service providers of one of its subdomains.

[0007] According to some embodiments, the delegation of the right to execute disruption elements includes the transmission of private keys relating to the service on which the disruptions are deployed prior to the transmission of the disruptions.

[0008] According to certain embodiments in which said service indicators are obtained from infrastructure and service monitoring software.

[0009] According to certain embodiments, said disturbances comprise one or more of the following disturbances: - a desequencing of data packets transmitted to services, - the introduction of additional delays in the service of requests to service providers; - the modification of the content of incoming requests or responses, - the insertion of new requests unknown to the services.

[0010] According to certain embodiments, said disturbances comprise one or more of the following disturbances: - a desequencing of data packets transmitted to services, - the introduction of additional delays in the service of requests to service providers; - the modification of the content of incoming requests or responses, - the insertion of new requests unknown to the services.

[0011] According to some embodiments, the method includes at least one query to obtain metrics from the infrastructure and services in order to constitute the configuration data of said database, said metrics including performance, the load of one or more processors of the infrastructure, a rate of memory used by the services or the infrastructure.

[0012] According to some embodiments, the database configuration data includes a set of disturbances with associated impacts on infrastructure and services, service level indicators (SLI) and service level objectives for infrastructure and services.

[0013] The features, presented individually in this application in connection with certain embodiments of the process of this application, can be combined with each other according to other embodiments of this process.

[0014] The invention also relates to a computer program comprising instructions for carrying out the steps of the process according to the invention, according to any of its embodiments, when said program is executed by a computer.

[0015] The invention also relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the invention, according to any one of its embodiments.

[0016] The invention also relates to a server for generating disturbances in a network to verify its robustness, comprising one or more processors configured together or separately for instructions for executing the steps of the method according to the invention, according to any one of its embodiments. Thus, the invention relates to a server for generating disturbances in a network to verify its robustness, said network comprising a plurality of infrastructures and services deployed and hosted in said network, said services being provided by at least one service provider, the server comprising one or more processors configured together or separately for - obtain at least one database containing configuration data for the plurality of infrastructures and services - determine a maximum unavailability value for said infrastructure and services based on a minimum tolerated service level calculated from service indicators for said infrastructure and services, and if said maximum unavailability value is positive: - transmitting data on network disruptions to the network infrastructure and deployed services, with the disruptions being determined from at least the database configuration data and the calculated maximum unavailability value, - update said maximum unavailability value following disruptions.

[0017] Other features and advantages of the present invention will become apparent from the description given below, with reference to the accompanying drawings which illustrate an example of an embodiment without any limiting character. Brief description of the drawings

[0018] [Fig-1] The [Fig. 1] represents an example of a system implementing certain embodiments of the present invention.

[0019] [Fig.2] Figure [Fig.2] represents a method according to one embodiment of the invention,

[0020] [Fig. 3] Figure 3 illustrates variations in the error budget according to a mode of realization,

[0021] [Fig.4] Fig.4 represents an embodiment of obtaining the base of data including configuration data for multiple infrastructures and services,

[0022] [Fig.5] The [Fig.5] represents an embodiment of perturbation determination and perturbation launch. Description of the implementation methods

[0023] In the following, a new service means any addition of applications or parts of applications of CSP (aka an application of a CSP or a new CSP) or of the infrastructure (aka a new service in containers) in the infrastructure servers.

[0024] In the following description, the terms margin of error or error budget, service level indicator (SLI) and service level target (SLO) describe:

[0025] SLI is a metric. For example, this metric can correspond to the processing time of an HTTPS request by a data center, defined as the time it takes for the last bit of the response to be sent minus the time it takes for the first bit of the request to be sent.

[0026] SLO is the value to be met over a period of time. For example, 99.99% of requests transmitted in less than 100ms over 30 days.

[0027] SLI (English acronym for "Service Level Indicator") and SLO (English acronym for "Service Level Objective") are two concepts related to the management of service levels in computer systems.

[0028] An SLI is a quantitative measure that defines the performance of a system or service based on a specific criterion. For example, an SLI could be the response time of a server or the number of connection errors.

[0029] An SLO, on the other hand, is a service level objective that one wishes to achieve for a given SLI. It is therefore a performance indicator that one wants to maintain or exceed. For example, an SLO could be that the response time of a server must not exceed 100 milliseconds.

[0030] SLOs are generally based on SLIs. That is to say, to define an SLO, one will rely on one or more quantitative measures to determine the objective to be achieved. SLIs thus allow us to measure the performance of a system or service, while SLOs allow us to define the objectives to be achieved.

[0031] In other words, SLIs measure the performance of a system or service, while SLOs define the objectives to be achieved in terms of service levels. SLOs are based on SLIs and allow performance to be measured against pre-established criteria.

[0032] The margin of error or error budget is the duration that represents 100% minus the SLO (expressed as a percentage). For example, for an SLO of 97%, the margin of error is 3%, meaning that 3% of requests may take more than 100ms to process over 30 days.

[0033] Advantageously, the method allows for the combination of quality or performance indicators (SLOs) of the network infrastructure and application service SLOs within an "Operational Quality Engineering" (IQO) service, also known as "Site Reliability Engineering" or "SRE." Infrastructure IQO techniques measure the infrastructure's resilience to infrastructure-level disruption test scenarios. Application IQO techniques measure the application service's resilience to application-level disruption test scenarios. The application service provider can delegate the management of its disruption scenarios to the infrastructure operator or a third party to manage the security of its service.Such an operational quality assurance (OQ) service for launching disruption scenarios can be hosted by an infrastructure operator or by a third party, such as a software vendor, providing an OQ service common to both the application and the infrastructure. In this way, the infrastructure operator, or this third party, can pool telemetry data from multiple application service providers to train artificial intelligence-based models for generating and configuring disruption scenarios. The process can generate new disruption scenarios (attacks, disruptions, errors injected into the service provider's service) that the service provider cannot design itself due to a lack of data to understand infrastructure constraints or those related to third-party services. It can also combine disruption scenarios from different service providers.The process can also avoid triggering service disruptions that could degrade infrastructure performance or even cause it to crash.

[0034] An embodiment of a system implementing the present invention is illustrated in [Fig. 1]. This system may comprise several modules. The embodiment described in [Fig. 1] corresponds to a Kubernetes environment, also denoted K8S. It includes, in particular, two load balancers (LBs) that distribute applications across the different servers or groups. It also includes an ingress controller. It also includes two groups (clusters). Each group comprises workers (w), here 2, and each worker comprises two services (S) in this example. Each group includes an 8K master. Each cluster may, for example, be located at a different infrastructure provider.

[0035] "Master nodes" and "Worker nodes" are concepts originating from Kubernetes. A cluster is composed of a set of servers. They are generally co-located, but architectures exist where these two functions can be on different PoPs (Points of Presence), also called access points, for example, depending on space and other constraints that would prevent having more servers.

[0036] The "Worker node" supports applications and services.

[0037] The "Master nodes" are in charge of management.

[0038] Ingress here means Ingress Controller which is a function allowing a flow to enter the cluster and be routed to the correct node / service.

[0039] A Kubernetes cluster therefore consists of at least one master node and one or more "worker node" servers supporting the requested services / applications. These applications are instantiated in what are called Pods, which are similar to containers. Containers are a segmentation of the operating system that allows the operator to clearly separate rights and access, as well as resources, between containers. This segmentation is done for security reasons, but also for ease of management. If a container fails, it does not cause the entire system to fail, but only the container. It is thus easy to restart it without impacting the other containers. The term Pod is specific to Kubernetes because it also incorporates Kubernetes management functions, unlike the term container, which is very generic in itself.

[0040] There are fewer master nodes than worker nodes because they do not have to support applications, but only cluster management functions related to Kubernetes. For example, they manage a database of applications running on the cluster as well as control functions, etc.

[0041] Private keys can be managed at different levels: - at the load-balancer level - at the level of Pingre ss controller, - at the container (pod) level - At the VIP level, meaning virtual IP address (Internet Protocol), a virtual IP address allows the identification of a cluster of multiple servers. This avoids the need for SNI snooping. SNI snooping, which means service name indicator scanning, is a function that verifies the SNI field in a TLS (Transport Layer Security) packet, notably in the case of HTTPS (HTTP over TLS). The SNI field contains the domain name targeted by the request. Without this field, since the packet is encrypted, a device in a state of interruption, such as a load balancer or proxy, would not be able to know which site the user is requesting. Here, the SNI field therefore allows the Load-Balancer to direct the request to the correct server with the requested service and to present the correct certificate when establishing the TLS connection.

[0042] An embodiment of a method according to the invention is shown in [Fig. 2].

[0043] The method includes a first step 1E1 of obtaining a database comprising configuration data for the plurality of infrastructures and services. It may, for example, include IQO scenarios. This database includes a data model comprising the infrastructure architecture as well as the configuration of the associated IQO scenarios. The configuration of the IQO scenarios includes the configuration of the IQO scenarios applied by software, for each hosted application service provider, for example in a common Kubernetes cluster, by geographic location. According to some embodiments, obtaining or creating the database of IQO scenarios can be done by querying the different elements (configuration, state, etc.) composing the infrastructure and by querying the different service providers.This database can include detailed configurations, individual scenario templates, common (aka synchronous) scenario templates, and known SLOs. Figure 4 represents one implementation of retrieving configuration data from the database.

[0044] Configuration data may include detailed configurations (e.g., Worker node [identifier, location, processor, memory], certificates and location of private keys and algorithm, key size, etc.), scenario frames for each actor (Each scenario is described in Turbulence, for example: Turbulencel: delay, method, interface, protocol, etc.), common (aka synchronous) scenario frames, known SLOs (e.g., minimum tolerated SLO for a given scenario, e.g., RTT < 150 ms for HTTP, Delay < 60 ms, with a 5% margin of error, etc.). This configuration data is obtained directly from the architecture's equipment if the goal is to inventory configurations, or from access providers or partner service operators if the goal is to create scenarios.

[0045] : During step E2, a maximum combined unavailability value is determined. infrastructure and services based on a minimum tolerated service level (SLO) calculated from infrastructure and service indicators.

[0046] To achieve this, service level objectives (SLOs) can be defined. These service level objectives are derived from the service level indicators (SLIs) of the elements included in the infrastructure, i.e., devices ensuring data routing in the network and storage devices in data centers, as well as the service level indicators (SLIs) of the services or applications. An SLI measures specific aspects of the service levels provided. Service level indicators include request response time, availability, error rate, and throughput.

[0047] The SLO service level objective can be defined as:

[0048] SLI (Service Level Indicator) and SLO (Service Level Objective) are two concepts related to the management of service levels in computer systems.

[0049] An SLI is a quantitative measure that allows the performance of a system or service to be defined according to a specific criterion. For example, an SLI could be the response time of a server or the number of connection errors.

[0050] An SLO, on the other hand, is a service level objective that one wishes to achieve for a given SLI. It is therefore a performance indicator that one wants to maintain or exceed. For example, an SLO could be that the response time of a server must not exceed 100 milliseconds.

[0051] The link between the two is that SLOs are generally based on SLIs. That is to say, to define an SLO, one will rely on one or more quantitative measures to determine the objective to be achieved. SLIs thus allow us to measure the performance of a system or service, while SLOs allow us to define the objectives to be achieved.

[0052] In summary, SLIs measure the performance of a system or service, while SLOs define the objectives to be achieved in terms of service levels. SLOs are based on SLIs and allow performance to be measured against pre-established criteria.

[0053] A common way of expressing the relationship between SLO and SLI is to say that the SLO objective must be greater than or equal to the performance level measured by SLI.

[0054] This could therefore be represented by a formula such as the following, for example:

[0055] SLI < SLO

[0056] This means that the SLI must be kept below or equal to the SLO to meet the defined service level objectives. For example, if an SLI measures the response time of a system and the SLO is set at 100 milliseconds, the formula would indicate that the response time must not exceed 100 milliseconds to meet the intended service level.

[0057] The metrics used can for example be extracted using the software grafana and prometheus: Prometheus and Grafana are two popular tools used in the field of monitoring and observability of computer systems.

[0058] Prometheus is an open-source performance monitoring system that collects performance data from various sources and stores it in an in-memory database. It allows monitoring the status and performance of different real-time systems, by collecting metric data such as CPU load, memory used, and error rates.

[0059] Grafana is an open-source data visualization tool that can be used with many different data sources, including Prometheus. It allows users to create graphs and dashboards to visualize and analyze performance data from different systems in an easy and intuitive way.

[0060] Together, Prometheus and Grafana are often used to monitor and analyze the performance of computer systems, particularly in production and cloud environments. They are used in industry to help development and systems management teams understand how systems operate and to detect and resolve performance issues.

[0061] A maximum downtime value, or error budget (EB), can then be determined for all services and infrastructure. This error budget is calculated from the SLO. For example, the error budget corresponds to the formula 1 - SLO. An illustration of an error budget is shown in [Fig. 3]. The error budget represents a period of downtime determined to be acceptable for a given period. It may be the maximum permissible threshold of errors and interruptions.

[0062] The error budget is available to developers, who can use it when launching a new feature. Based on the service level objective and the error budget, it is possible to determine whether launching a product or service is feasible with the allocated error budget.

[0063] Figure 3 illustrates the evolution of the error budget according to one embodiment. The y-axis represents the error budget (as a percentage) and the x-axis represents time, with each time step corresponding to a month of the year. This error budget can correspond to the error budget in an infrastructure, at a service provider, or both combined. It illustrates the phases during which the error budget is considered insufficient to trigger disruptions and the phases during which the error budget allows for the triggering of disruptions. In Figure 3, it can be noted that the error budget is less than 99.99% between approximately mid-June and early September. This means that it is possible, for example, to deploy new functions or modify the network and service configuration and challenge this modification with network and service disruptions that may be acceptable, because the SLO is respected.

[0064] The error budget is a concept closely related to that of one or more SLOs in service level management. It refers to the amount of errors or downtime that are tolerated relative to a given service level objective (SLO). In other words, the error budget is the percentage of time during which The system may fail or produce results outside the acceptable range for the SLO, while still meeting the overall service level objective. The relationship between the error budget and the SLO is therefore very close; the error budget is calculated by subtracting the percentage of tolerated downtime or errors (defined by the error budget) from the percentage of time allocated to meet the SLO.

[0065] For example, if the SLO is set to an average response time of 100 milliseconds, the error budget could be 0.5%. This would mean that response times greater than 100 milliseconds are permitted for 0.5% of the system's uptime without exceeding the accepted service level.

[0066] In summary, the error budget and the SLO are closely linked because the error budget defines the quantity of errors tolerated in relation to a given SLO and thus makes it possible to maintain the expected level of service in case of problems.

[0067] When response times are for example >100 ms correspond to 0.499% system availability, it is advisable not to modify the configuration by deploying new functions and to work on the reliability of the network and services and not to deploy new functions.

[0068] The error budget EB can be defined as:

[0069] EB= 100%-SLOs

[0070] An SLO is determined per system, and therefore an error budget per system

[0071] Example 1: SLO Application

[0072] Example 2: Cloud Infrastructure SLO

[0073] Example 3: Network SLO

[0074] In some embodiments, an SLO is not expressed as a percentage. Although percentages are commonly used in service level management to measure performance objectives, SLOs can be defined in different ways depending on the needs and objectives of the system or service.

[0075] For example, an SLO can be defined as a time measure, such as a maximum response time in milliseconds, or as an availability measure, such as a percentage of the total time the system must be operational.

[0076] The error budget is generally calculated from the SLO using a formula that subtracts the SLO from 100%, giving the acceptable amount of error or downtime. For example, if the SLO is defined as a maximum response time of 100 milliseconds, the error budget could be calculated using the formula 100% - SLO, which would give an error budget of 0.0% if the target response time must always be met.

[0077] If the error budget is exceeded (threshold reached, in our example, 0.5% error), step E3, then we move to a step E4 and if it is not exceeded (threshold not reached) we move to step E5.

[0078] During step E4, a new service is deployed, the new service being determined from at least the database configuration data and the error budget value reached.

[0079] During a step E5, data on disturbances of the network infrastructure and deployed services are injected into the network, the disturbances being determined from at least the configuration data of the database and the calculated maximum unavailability value.

[0080] During step E6, the maximum unavailability value is updated following the injected disturbances, and the process returns to step E2. It is also possible, according to certain embodiments, to determine countermeasures to improve the SLO.

[0081] If the error budget is exceeded (threshold reached) at step E3, then we proceed to step E7. During this step, we launch an intelligent analysis based on artificial intelligence, commonly known as AISV, an English acronym for "Artificial Intelligence Software Vendor".

[0082] In an example of using the process applied to security, one might, for instance, want to apply a countermeasure proposed by the AISV and deduced from the EB value calculated following the injection of turbulence. Thus, for example, the private key storage strategy is modified to improve its security. Therefore, during step E8, a countermeasure can be applied to services or infrastructure whose error budget has been exceeded.

[0083] Then during a step E10 we can apply the strategy determined for the set considered, the services or infrastructure whose error budget is exceeded And we return to step E2.

[0084] Let us now take an example of the application of the process as described with reference to [Fig.2]. According to this example, the calculation of error budgets carried out during step E2 indicates that the error budget of the network infrastructure is respected as well as that of a location service.

[0085] In order to be able to deploy the disruptions on the location service, the location service provider delegates one of its subdomains (a geographical area, an IP subnet, a cluster or group of clusters, a mobile network access gateway for example of type UPF (English acronym for "User Plane Function") to a network infrastructure operator or to a third party in charge of implementing the disruption scenario.

[0086] An analytics service provided or developed by the Cloud provider and / or the network infrastructure provider is capable of adapting to the turbulence scenario to the contexts of Cloud infrastructure and network infrastructure. Conversely, a service provider could ask a network infrastructure provider to apply the turbulence scenarios specific to that service provider's applications.

[0087] The service provider dynamically delegates one of its subdomains to the operator for the execution of a scenario from a Master managing the pods hosting the service provider's applications. The scenario contains the relevant pods:

[0088] The method may therefore include the delegation by the service provider's application of the right to execute isolated disruption elements in one of its subdomains by at least one of the cloud and / or network infrastructure providers. Since a network or cloud operator is responsible for testing an infrastructure and services implemented on that infrastructure, the service provider may advantageously delegate the testing of a service by implementing disruption scenarios.

[0089] In some embodiments, delegation consists of automating the implementation of the process. Establishing the delegation involves the delegate, here the service provider CSP1, sending a signed request to the delegate, here the cloud and / or network provider OP1. This request contains the parameters defining the precise scope of the disruption scenarios to which the delegation applies. These parameters include, for example: the network / cloud infrastructure elements concerned (e.g., pods, clusters), the subdomains concerned (e.g., NSI, NS2, etc.), the disruption scenarios concerned, credentials, and the validity period of the delegation.

[0090] The delegation methods that can be implemented are for example: - ACME / STAR which uses specific x509 provided by a CA (external), or- Delegated credentials: a faster delegation signed by the delegate.

[0091] The scenario is then executed by one of the infrastructure devices and can be controlled by the infrastructure operator, the service provider or a third party.

[0092] If the service provider controls the deployment of disruptions, then the service provider requests the infrastructure error budgets, provides everything to the third party and obtains an updated scenario.

[0093] If the infrastructure operator controls the deployment of the disruptions, then the infrastructure operator requests the error budgets from the service provider, provides everything to the third party and obtains an updated scenario.

[0094] Examples of perturbations that can be implemented within the framework of certain embodiments of the present invention are given below by way of example.

[0095] According to some embodiments, the disturbances comprise one or more of the following disturbances: - a desequencing of data packets entering the infrastructure or a service, - a desequencing of data packets leaving the infrastructure or a service, - the introduction of additional delays in the flow of incoming requests to service providers, - Activating an imbalance in the routing of a load balancer in order to briefly saturate a group of servers, - the modification of the content of incoming requests or outgoing responses, - the insertion of new requests unknown to the services and infrastructure, - the activation of protection procedures included in the services or infrastructure, procedures which do not trigger in normal operation.

[0096] For example, a desequence of incoming or outgoing data packets in the infrastructure or in a service, may for example consist of the injection of an erroneous message to a kubemetes service such as "etcd" or "grafanas".

[0097] The duration of the imbalance generated during the injection of disturbances in order to briefly affect a group of servers may depend on the error budget.

[0098] Among the protection procedures included in the services or in the infrastructure, we can mention the logging of all packets to a subdomain of a service; the thorough checking of certificates of http requests to a subdomain of a service or coming from an IP subnet.

[0099] First example of turbulence in the context of an Apache HTTPS type data server.

[0100] The infrastructure operator knows that the service provider includes Apache containers from service deployment logs (ArgoCD), or by querying the Kubernetes infrastructure (via a command that allows control of its Kubernetes nodes, i.e., listing the services), or by observing incoming HTTP requests in an ingress controller. ArgoCD is an open-source continuous deployment and continuous delivery tool that facilitates the management of software deployments in production environments. It allows for the automatic deployment of applications to various environments such as Kubernetes clusters, virtual machines, and bare-metal servers.

[0101] This latter mode assumes that the infrastructure has the private keys and therefore acts as a reverse proxy (the element that governs incoming connections in a data center) for the service. It can perform application counting if it possesses the keys encrypting the relevant streams, or only if it has received a delegation allowing it to read at least part of the signaling received by the hosted application.

[0102] Second example of turbulence in the context of a MySQL environment

[0103] The infrastructure operator knows that the service provider has MySQL containers from service deployment logs (ArgoCD), or by querying the K8S infrastructure (via pe, kubectl list services), or by observing incoming HTTP requests in the Ingress controller. This last method assumes that the infrastructure has the private keys and therefore acts as a reverse proxy for the service: it terminates the TLS connection, performs counting, and may even perform SRE actions, and routes the request to the hosted application.

[0104] For example, an SRE scenario for testing the "ability to route requests from a user device to the correct platform and service" could include the following turbulence: - Turbulencel: introduce desequencing in the order of arrival of requests from a user device (and not in the packets); e.g., a 5-second desequencing in the HTTPS protocol for incoming reverse proxy traffic - Turbulence!: Introduce very long delays in the incoming HTTP POST requests from the reverse proxy to the database servers. Example: Delay a MySQL form by 3 minutes and observe the impact.

[0105] Third example of turbulence in the context of a pre-coded scenario in the service provider's code.

[0106] The disruption scenario includes internal application disruption functions (programmed abnormal behavior requiring authentication to activate) enabled using start and stop keys dynamically generated by the infrastructure operator from the scenario's credits. A pre-coded disruption function is included in the application program code of the service provider: for example, the application routing container includes a disruption function which, once activated, generates routing errors to challenge the behavior of the application components. The components are designed to withstand routing errors occurring, in particular, during the container lifecycle (re-trying client requests upon shutdown, starting routing functions, etc.).

[0107] Figure 4 represents an embodiment of obtaining the database comprising configuration data for the plurality of infrastructures and services,

[0108] During a step T1, service providers transmit a reliability test request for one or more services to the infrastructure manager, and more specifically to a module that controls the establishment and management scenarios, the infrastructure monitor. Step T1 is part of starting to establish the inventory of service providers.

[0109] During step T2, the infrastructure monitor sends a request to initiate the disruption procedure to some or all of the service providers present. During step T3, these service providers send a response to the monitor. This response can be an agreement, a rejection, a request to postpone, or the initiation of the disruptions. During step T4, the infrastructure monitor also sends notification messages regarding ongoing disruptions to the infrastructure master (K8S). It receives a response, step T5. This response includes the states of the various elements tested. Steps T2 to T5 are part of the querying of the various components, both service providers and infrastructure modules.

[0110] During a T6 step, the load manager, LB, transmits metrics to the infrastructure monitor. Among the metrics transmitted are network metrics (number of packets lost), application metrics (request processing time by the hosted application), cloud metric (load manager (LB) traverse time).

[0111] During a T7 step, the sidecar proxy (a proxy, a command element, located on the cloud infrastructure (generally co-located with the application, the client, or the service provider)) also transmits its metrics (e.g., ongoing sessions, network traffic, etc.). Thanks to this function, the application is relieved of a number of functions, such as sending its metrics to the infrastructure monitor.

[0112] During a T8 step, service providers also transmit their metrics, in the form of SLIs, to the infrastructure monitor.

[0113] The method specifies 2 new functions: the SRE Monitor capable of creating the SRE database by querying the different elements of the infrastructure and the service and the SRE injector enabling the triggering of disturbances.

[0114] Figure 5 represents an embodiment of determining disturbances and launching disturbances according to certain embodiments of the process according to the invention.

[0115] During a Zl step, the sidecar proxy transmits its metrics to the infrastructure monitor. A sidecar proxy is an application that is added to an existing application within a container.

[0116] During a Z2 step, service providers transmit their metrics to the infrastructure monitor.

[0117] During a Z3 step, the load manager, LB, transmits its metrics to the infrastructure monitor.

[0118] During a Z4 step, the sidecar proxy transmits its metrics to the service provider monitor.

[0119] During a Z5 step, service providers transmit their metrics to the service provider monitor.

[0120] During a Z6 step, the service provider's monitor transmits the calculated SLO to a service provider's SRE injector.

[0121] During a Z7 step, the infrastructure monitor transmits the calculated SLO to an infrastructure SRE injector.

[0122] During a Z8 step, a synchronization is performed between the infrastructure's SRE injector and the service provider's SRE monitor. During this step, the error budget is calculated, as described previously with reference to [Fig.2].

[0123] During a Z9 step, a verification and key generation is performed between the service provider's SRE monitor and the infrastructure's SRE injector.

[0124] During a Z10 step, an activation signal is transmitted from the infrastructure SRE injector to the service provider's SRE injector to execute the disruption scenario(s).

[0125] During a ZI 1 step, the infrastructure's SRE injector transmits disturbances to the load manager or activates turbulence at the load manager, as previously mentioned with regard to step E5 of [Fig.2].

[0126] During a Z12 step, the infrastructure's SRE injector transmits disturbances to the sidecar proxy or activates turbulence at the sidecar proxy as previously mentioned with regard to step E5 of [Fig.2].

[0127] During a Z13 step, the infrastructure SRE injector transmits disturbances to service providers or activates disturbances at the service provider level as previously mentioned with regard to step E5 of [Fig.2].

[0128] We will now describe an example of implementation of the invention.

[0129] Metrics are used to define the availability of a service over a short period t of duration T and comply with a framework to validate the period. Several metrics can be combined to invalidate a period t.

[0130] The number of valid periods in P determines the availability percentage. For example, if 90% of the periods T are valid in period P, the service is available 90% of the time. In this case, the error budget in P is 10%. This margin of error is available without compromising availability. It is used by each actor to test the robustness of their service.

[0131] The following metrics are used:

[0132] The application metric is the processing time of the incoming request by the hosted application.

[0133] The cloud infrastructure metric is the infrastructure load manager traversal time.

[0134] The network metric is the number of packets lost.

[0135] Consider the following three turbulences:

[0136] InjSql23(x,y): Injection of X attacks into Y SQL queries

[0137] Forml5(x,y): Injection of X random web forms into HTML POST requests in Y sessions

[0138] DeSeqlO(x): Unsequencing X web forms from HTML POST requests

[0139] The impacts of these three turbulences on the error budget are as follows (the (Figures shown in the table correspond to the error budget): Actor Application Infrastructure Cloud Network Scenario Turbulences Estimated Impact E Ba (%) Estimated Impact EBi (%) Estimated Impact E Br (%) InjSql23(3,5) 8 7 3 InjSql23(8,9) 6 4 2 Forml5 (6,12) 4 1 1 Forml5(3,6) 2 1 1 DeSeql0(5) 1 1 1 DeSeqlO(lO) 3 3 3

[0140] It is assumed that at time tl, the respective error budgets (EBa, EBi, EBr) are (15,6, 5).

[0141] According to embodiments of the present invention, the method determines the turbulence sequences which respect the error budget of each (application, cloud infrastructure, network) with the constraint that the turbulence deteriorates the error budget of each in an acceptable way, that is to say that the application of the turbulence does not make the error budget negative.

[0142] If we try to launch InjSql23(3,5), we obtain (EBa, EBi, EBr) = (15, 6, 5) - (8, 7, 3)=(7, -2, 3).

[0143] The cloud infrastructure's error budget is negative, so we cannot launch this disruption.

[0144] If we try to launch InjSql23(8,9), we obtain (EBa, EBi, EBr) = (15, 6, 5) - (6, 4, 2) = (9, 2, 3).

[0145] The perturbation Forml5(6, 12) can be executed simultaneously or in series giving (EBa, EBi, EBr) = (9, 2, 3) - (4, 1, 1)= (5, 1, 2).

[0146] DeSeql0(5) can also be executed at the same time or in series with InjSql23(8,9) and Forml5(6, 12)

[0147] We obtain (EBa, EBi, EBr) = (5, 1, 2) - (1, 1, 1) = (4, 0, 1)

[0148] Forml5(3, 6) = can also be executed simultaneously or sequentially as InjSql23(8,9):

[0149] We obtain (EBa, EBi, EBr) = (9,2, 3) - (2, 1, 1) = (7, 1, 2)

[0150] DeSeqlO(lO) cannot be executed simultaneously or sequentially as InjSql23(8,9) because we would obtain (EBa, EBi, EBr) = (9, 2, 3) - (3, 3, 3) = (6, -1, 0):

[0151] A first disruption scenario that can be launched is therefore: - InjSql23(8,9) - Forml5(6, 12) - DeSeql0(5)

[0152] A second scenario that can be launched is therefore: - InjSql23(8,9) - Forml5(3,6)

[0153] The impacts of each parameterized turbulence are estimated in advance, for example, by training on application log errors and activity parameters observed on the infrastructure and network. These impacts may be part of the configuration data present in the database.

[0154] An example is given in the following table: Application Infrastructure Cloud Network Errors and attacks observed in logs Number of requests Number of sessions Estimated impact EBA (%) Ran (GB) iCpu (GFLOPS) Disk (GB) Estimated impact EBI (%) Delay (ms) Loss (integer) Jig (ms) Estimated impact EBr (%) InjSql23 3 5 8 20 17 15 7 20 20 10 3 InjSql23 8 9 6 28 19 10 4 25 28 20 2 Forml5 11 32 4 19 25 20 1 35 25 19 1 Forml5 15 15 2 25 35 28 1 15 35 25 1 DeSeql 0 10 10 1 10 15 10 1 10 15 20 1 DeSeql 0 20 20 3 20 10 20 3 20 10 28 3

Claims

Demands

1. A method for generating disturbances in a network to verify its robustness, said network comprising a plurality of infrastructures and services deployed and hosted in said network, said services being provided by at least one service provider, the method comprising: - obtaining (E1) at least one database comprising configuration data for the plurality of infrastructures and services, - determining (E2) a maximum unavailability value of said infrastructure and services based on a minimum tolerated service level (SLO) calculated from service indicators of said infrastructure and services, and if said maximum unavailability value is positive (E3): - transmitting (E5) into the network disturbance data of the network infrastructure and deployed services,The disruptions being determined from at least the database configuration data and the calculated maximum unavailability value, - the update (E6) of said maximum unavailability value following the disruptions.

2. A method according to claim 1 comprising, prior to said transmission of disturbances, the deployment (E4) of a new service, said new service being determined from at least the database configuration data and the calculated maximum unavailability value.

3. A method according to any one of claims 1 or 2 comprising the delegation of the right to execute the disruption data by at least one of the service providers of one of its subdomains.

4. A method according to any one of the preceding claims wherein said service indicators are obtained from infrastructure and service monitoring software.

5. A method according to any one of the preceding claims, wherein said disturbances comprise one or more of the following: - a desequencing of data packets transmitted to the services, - the introduction of additional delays in the service of requests to service providers; - the modification of the content of incoming requests or responses, - the insertion of new requests unknown to the services.

6. A method according to any one of the preceding claims comprising at least one query for obtaining metrics of the infrastructure and services in order to constitute the configuration data of said database, said metrics comprising performance, the load of one or more processors of the infrastructure, a memory usage rate by the services or the infrastructure.

7. A method according to any one of the preceding claims, wherein the database configuration data includes a set of disturbances associated with their impact on infrastructure and services, service level indicators (SLI), and service level objectives (SLO) for infrastructure and services.

8. Computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 7 when said program is executed by a computer.

9. A computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 7

10. 1 d. / . A server for generating disturbances in a network to verify its robustness, said network comprising a plurality of infrastructures and services deployed and hosted in said network, said services being provided by at least one service provider, the server comprises one or more processors configured together or separately to: - obtain at least one database comprising configuration data for the plurality of infrastructures and services, - determine a maximum unavailability value for said infrastructure and services based on a minimum tolerated service level (SLO) calculated from service indicators of said infrastructure and services, and if said maximum unavailability value is positive: - transmit disturbance data for the network infrastructure and deployed services into the network being determined from at least the database configuration data and the calculated maximum downtime value, - update said maximum unavailability value following disruptions.