Methods and entities for determining the load of a network platform and adjusting platform resources based on this estimate
By using a predictive model to estimate load and latency, the resource adjustment process in network platforms becomes more efficient, addressing the inefficiencies of current techniques and enhancing service quality.
Patent Information
- Application Number
- FR2023012048
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-09
AI Technical Summary
Current techniques for adjusting resources in network platforms based on load are inefficient due to the time required to duplicate or release execution environments and start/stop software components in response to load changes.
A process involving a computer-based model training to predict the load and latency of network platform modules, allowing for proactive resource adjustment based on estimated future needs.
This approach enables more efficient resource management by anticipating and preparing for upcoming load demands, thereby reducing the time needed to adjust resources and improving service quality.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Methods and entities for determining the load of a network platform and adjusting resources of the platform based on this estimation Prior art
[0001] The invention is situated in the context of telecommunications networks, and more precisely in that of optimizing the quality of service provided by network platforms.
[0002] The invention finds a preferred but non-limiting application in the context of so-called new generation networks (NGN, for Next Generation Networks) in which telecommunications network operators offer third parties access to services by making application program interfaces, hereinafter API (Application Programming Interface), available to them.
[0003] The sizing and availability of the resources of these platforms (in terms of computing power, memory, storage capacity for example) are decisive for the quality of the service provided by these platforms.
[0004] In the current state of the art, the resources allocated by these platforms are adjusted to meet their load (in English "autoscalling"). In practice, resources are allocated (respectively released) when the load becomes higher (respectively lower) than predetermined thresholds.
[0005] These techniques, although satisfactory, have certain drawbacks due to the time required for: - duplicate execution environments (e.g. containers) and start software components to respond to a surge in load, and to: - stop software components and release execution environments in the event of a load drop.
[0006] The invention aims at a solution for adjusting the resources of a network platform which does not have these drawbacks. Subject matter and summary of the invention
[0007] Thus, and according to a first aspect, the invention relates to a method for training a model for determining the load of a network platform, this method being implemented by a computer and comprising the following steps: - obtaining a history of requests issued by at least one third party and received by the network platform over at least one past period of time; - obtaining at least one load of at least one module of said network platform during said at least one period of time;
[0008] - training the model with training data comprising at least temporal data representative of said at least one period of time and said at least one load obtained, a use (or inference) of the trained model making it possible to determine the load of said at least one module of said network platform during a period of time provided as input to the model.
[0009] In one embodiment, the load of the platform or of a module of the platform is the number of requests processed per unit of time by this platform or by this module.
[0010] In one embodiment, the training method further comprises: - obtaining at least one latency of said at least one module of said network platform during said at least one period of time; - the training data of the model comprising said at least one latency, a use of the trained model further enabling the latency of said at least one module of said network platform to be determined during the time period provided as input to the network.
[0011] In one embodiment, the latency of the platform or of a module of the platform is the overall response time of a request by this platform or by this module.
[0012] In one embodiment, the training data further comprises an identifier of the third party sending said at least one request.
[0013] This embodiment makes it possible to train the model to predict the latency or load induced by requests from a particular third party.
[0014] In one embodiment, the training data further comprises an identifier of an application within the framework of which said at least one request was issued.
[0015] This embodiment makes it possible to train the model to predict the latency or load induced by the requests of a particular application.
[0016] The model used by the invention may, for example, be a recurrent neural network. This model may also use a forest comprising at least one decision tree.
[0017] The invention also relates to a model for determining the load of a platform over a period of time, this model being obtained by a method as mentioned above.
[0018] According to a second aspect, the invention aims at a method for determining the load of at least one module of a network platform, this method comprising a step of using a load prediction model as mentioned above to determine the load and possibly the latency of at least one module of the platform. network over a period of time provided as input to the network.
[0019] Correlatively, the invention relates to an entity for determining the load of at least one module of a network platform, this entity exposing at least one application program interface making it possible to trigger the execution of at least one use of at least one load determination model trained by a training method as mentioned above to determine the load of at least one module of a network platform during a period of time provided as input to the network.
[0020] According to a third aspect, the invention relates to a method for providing a network service implemented by a network platform and comprising the following steps: - obtaining an estimate of the load of at least one module of said network platform over a period of time by implementing a determination method as mentioned above; - adjusting resources of said at least one module for said period of time based on a result of said determination; - processing at least one request issued by a third party to access said service using said adjusted resources.
[0021] The invention also relates to a network platform comprising: - a module configured to query a load determination entity as mentioned above to obtain an estimate of the load of at least one module of said network platform over a period of time;
[0022] - a resource adjustment module of said at least one module for said period of time depending on a result of said determination;
[0023] - a module for processing at least one request issued by a third party to access a network service provided by said platform using said adjusted resources.
[0024] Thus, and very advantageously, the invention proposes to train a model that can be used to determine, in other words predict, the load and possibly the latency of a network platform or modules of a network platform during a future period of time so as to be able to adjust the resources of the platform or of some of its modules in order to be able to process the requests that will be received during this future period of time with the resources thus adjusted.
[0025] Thus, the platform's resources are sized according to the platform's future needs, and reserved in advance.
[0026] The disadvantages of the prior art related to the time required to duplicate or release execution environments, start or stop software components to respond to an increase or decrease in load are therefore resolved by the invention.
[0027] Advantageously, if the latency of the platform or one of its modules is completed zero or relatively low for a future period of time, the invention proposes, in one embodiment, not to allocate additional resources for this period of time.
[0028] In one embodiment, the training method further comprises: - a step of comparing an actual load or an actual latency of at least one module of said network platform with the determined load or latency of said at least one module during a period of time and; - a step of retraining the model based on the result of this comparison.
[0029] This embodiment advantageously makes it possible to re-train the model when it no longer determines the load or latency satisfactorily.
[0030] In a particular embodiment, the model is only re-trained when such a drift is observed.
[0031] The invention also relates to a computer program comprising instructions for executing the steps of the training method and / or the steps of the load determination method and / or the steps of providing a network service when said program is executed by a computer.
[0032] This program may use any programming language, and be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0033] The invention also relates to a computer-readable information medium, and comprising instructions of a computer program as mentioned above. The information medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM, a non-volatile memory of the flash type or even a magnetic recording means, for example a hard disk. Furthermore, the information medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means. The program according to the invention may in particular be downloaded from a network such as the Internet. Alternatively, the information medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the method in question. Brief description of the drawings
[0034] Other characteristics and advantages of the present invention will emerge from the description given below, with reference to the appended drawings which illustrate exemplary embodiments thereof which are not in any limiting nature. In the figures:
[0035] [Fig.l] [Fig.l] represents a model that can be implemented in a particular embodiment of the invention;
[0036] [Fig.2] [Fig.2] represents the main steps of a training method according to an embodiment of the invention;
[0037] [Fig.3] [Fig.3] illustrates examples of training data that may be used in particular embodiments of the invention;
[0038] [Fig.4] [Fig.4] represents another model which can be implemented in a particular embodiment of the invention;
[0039] [Fig.5] [Fig.5] represents a load determining entity according to a particular embodiment of the invention;
[0040] [Fig.6] [Fig.6] represents a network platform according to a particular embodiment of the invention;
[0041] [Fig.7] [Fig.7] represents in the form of a flowchart the main steps of a load prediction method and the main steps of a service provision method in accordance with particular embodiments of the invention;
[0042] [Fig.8] [Fig.8] represents the hardware architecture of a load determination entity and a network platform according to the invention in particular embodiments of the invention. Description of the embodiments
[0043] [Fig.l] represents a model M0D; implemented by computer and configured to determine or predict the load LOADij and the latency LATlj of a network platform to respond to requests that will be received from a third party T; during a future time period PTj
[0044] In another embodiment of the invention, the model M0D; predicts only the load LOADij.
[0045] This model M0D; is driven by a PTR drive method in accordance with a particular embodiment of the invention and the main steps of which are described with reference to [Fig.2].
[0046] This training method comprises a step E10 of obtaining the history of requests RQir received from this third party T; over past time periods PTr (in other words a history of requests received over past time periods), and a step E20 of obtaining at least one load LOADir (and in this example at least one latency LATir) of the platform during these periods.
[0047] In one embodiment, new query history, load and latency data are received every day and aggregated with previous data (step E30).
[0048] Different models can be used depending on the time step of determination. desired and the amount of historical data to be processed by the model.
[0049] For example, if we want to predict the load (and latency) of the platform for each hour of a day, a recurrent neural network, for example of the LSTM (Long Short Term Memory) type, can be used.
[0050] In a known manner, a method for training a recurrent neural network M0D; may consist of dividing the data into a training data set and a test data set, the training data being used to train the model, the test data being used to evaluate its performance.
[0051] This training and test data includes the temporal data representative of the time periods during which requests were received by the platform, the load and possibly the latency of the platform or of at least some of its modules during these time periods.
[0052] They may also include additional data, for example an identifier of the third party issuing the request or an identifier of an application within the framework of which the third party issued this request.
[0053] The neural network is trained to perform a regression task of predicting the platform load (single-output neural network) and possibly its latency (two-output neural network) over a period of time, based on the time periods (input data).
[0054] The neural network can be validated and fine-tuned on the entire test data set.
[0055] In another embodiment, if one wishes to predict the load (and latency) of the platform for much shorter periods of time, for example of the order of a minute, one can use a linear regression model, for example based on a decision tree or on a forest of decision trees.
[0056] Also in known manner, a method for training a forest of decision trees may consist of training each tree on a sub-part of the training set with random samples and random features (random subset of features), each of the trees being trained on its respective part of the training set to predict the platform load as a function of a time period. The individual predictions may be aggregated (e.g. by taking the average of the predictions for a regression) to determine the final prediction of the forest.
[0057] The test data can be used to evaluate the performance of the decision tree forest, for example using the mean square error.
[0058] A Transformer type network can be used.
[0059] In a particular embodiment, the training method comprises a step E40 of preparing the training data of the model. This preparation fa- optional depends on the type of model.
[0060] [Fig.3](a) represents the format of the data used to train a recurrent neural network, for example of the LSTM type.
[0061] In this embodiment, the temporal data representative of the time periods may be HDr timestamp data extracted directly from the requests.
[0062] [Fig.3](b) represents the format of the data used to train a decision tree or forest of trees. In this embodiment, the temporal data representative of the time periods PTr can be calendar data DCr (time slot, hour, day of the week, month, ...) determined from the timestamp data HDr.
[0063] As mentioned previously, the training data may include an identifier of the third party T; issuing the request or an identifier of an application within the framework of which the request was issued.
[0064] In a particular embodiment, and as shown in [Fig.3](c), the training data may also include complementary contextual data DCC obtained from the time periods PTr. For example, this complementary data includes: - binary information indicating whether the time period PTr corresponds to a weekend; or - information indicating that the time period PTr corresponds to a particular event, for example a sporting event, an advertising campaign, etc.
[0065] This additional data makes it possible to improve the prediction model when the distribution of requests issued by the third party T; is correlated with this additional data, for example if the number of requests increases during the week, decreases on weekends and increases considerably during an advertising campaign of the third party.
[0066] In one embodiment, the model M0D; is trained a first time, then retrained when a drift is observed between the load LOAD^ (or the latency LATij) predicted by this model over a time period PTj and the actual load LOADij (or the actual latency LOADij) of the platform measured during this time period PTj.
[0067] For this purpose, the training method comprises, in this embodiment, a step E50 of obtaining the actual load LOADij and latency LAT^j during the time period PTj, a step E60 of obtaining the actual load LOADij and latency LATij determined by using the model M0D; during this time period PTj, and a step E70 of comparing these actual loads and latencies with the loads LOADij and LATij determined by the model over this time period PTj.
[0068] In the embodiment described here, when the difference between the actual and predicted loads or between the actual and predicted latencies exceeds a determined threshold (for example 10%), the model is re-trained (step E80) on the basis of the training data. aggregated at step E30.
[0069] In a particular embodiment, the training of the recurrent neural network implements a transfer learning mechanism, so as to only retrain the network over short periods, for example three months, and thus reduce the training time.
[0070] In the embodiments described above, the load and latency of the network platform as a whole are considered.
[0071] In another embodiment shown in [Fig.4], the network platform is composed of a plurality of functional modules MFk and a model MOD ki is trained for each module MFk and each third party Tj. This embodiment advantageously makes it possible to predict the load LOADk; and the latency LAT^j of each of these modules MFk to process future requests from third party i over a time period PTj.
[0072] Each of these modules can be re-trained when a drift is observed between the actual load or latency and the predicted load or latency of this module.
[0073] In the embodiment described here, the M0D models of the platform or M0Dk models of the trained MFk functional modules of the platform are exposed in the form of APIs.
[0074] [Fig.5] represents a load determination LPS entity according to the invention. In the embodiment described here, this LPS entity offers three APIs APL, API2 and API3. In this example, it is assumed that models M0D;, M0Dk; have been trained for I thirds (i=1 to I).
[0075] The first APL API makes it possible to obtain for a third party i, the load LOADlj and the latency LAT;j of the platform for a future period of time PTj. This load and latency information is obtained directly by using the MOD; model.
[0076] The second API API2 makes it possible to obtain the load LOADj and the latency LATj of the platform for a future time period PTj. This load and latency information is obtained by taking into account the load and latency information obtained by use for each of the third parties i.
[0077] The first API API3 makes it possible to obtain for a third party i, the load LOADlj and the latency LATLj of each of the functional modules MFk of the platform for a future period of time PTj. This load and latency information per functional module is obtained directly by using the models M0Dk;.
[0078] [Fig. 6] represents a PS network platform according to the invention. This platform comprises a COM communication module configured to interrogate the LPS load determination entity in order to obtain an estimate of the load (and in this example of the latency) of this platform or of functional modules of this platform over a period of time PTj to come.
[0079] If this network platform serves a plurality of third parties in, n=1 to N, it can query the LPS entity for each of these third parties in so as to know the resources it will need to process the requests from these third parties during the time period PTj.
[0080] This query can be carried out regularly, for example every hour.
[0081] In this example, the PS network platform includes an adjustment MAR module resources to adjust in advance, that is to say before the time period PTj, the functional resources of the platform, or of some of its modules, to its needs.
[0082] This adjustment of resources may for example consist of allocating or freeing new execution environments, new containers, RAM, disk space, starting or terminating software processes.
[0083] Preferably, the platform only instantiates the predicted resources for a time period PTj.
[0084] In a particular embodiment of the invention, if the predicted latency for a period of time is zero or very low (for example less than a predetermined threshold), no resource is allocated.
[0085] The PS network platform comprises an MTR module configured to process requests issued by a third party to access a network service provided by the platform using the adjusted resources.
[0086] [Fig.7] represents in the form of a flowchart the main steps of a PCL method for determining load and the main steps of a PFS method for providing service in accordance with the invention.
[0087] It is assumed here that the PCL load determination and PFS service provision methods are implemented by the LPS load prediction entity and by the PS network platform.
[0088] In the embodiment described here, and according to a frequency configurable by an administrator of the network platform PS, the PFS method invokes (step S10) one of the APIs of the load determination entity LPS to obtain an estimate of the load (and in this example of the latency) of this platform or of functional modules of this platform to process the requests of third parties to which the platform PS provides services during a future time range PTj.
[0089] The call to this API triggers a use of the M0Din or M0Dkin models trained for each of these third parties by the PCL load estimation method (step P10).
[0090] During a step P20 of the PCL load estimation method, the LPS entity sends the load and latency estimates to the PS service platform. These estimates are received by the PS platform during a step S20.
[0091] These estimates can be sent in aggregate and / or by third parties and / or by MFk functional module.
[0092] During a step S30 of the PFS service provision method, the platform adjusts its resources in accordance with these estimates.
[0093] In one embodiment, no new resources are allocated if the expected latency is zero or very low (less than a threshold).
[0094] In one embodiment, at least one functional component of the platform is associated with a sensitivity coefficient, and the amount of resource to be allocated to this component is a function of the predicted load for this component and of this sensitivity coefficient.
[0095] During a general step S40, the network platform PS uses said adjusted resources to process the requests issued by third parties during the time period PTj to access said service using said adjusted resources.
[0096] In one embodiment of the invention, the device used to train the model, the load determination entity and the network platform each have the hardware architecture of a computer as shown in [Fig.8]. It comprises in particular a processor 10, a random access memory 11, a read only memory 12 and communication means 13.
[0097] The read-only memory 12 constitutes a recording medium within the meaning of the invention in which a computer program can be recorded. In the example of [Fig. 8], three programs are represented but the invention covers embodiments in which the read-only memory comprises at least one of these programs.
[0098] In particular, this computer program may be a PGE program comprising instructions for executing the steps of a PTR training method as described with reference to [Fig.2].
[0099] In particular, this computer program may be a PGD program comprising instructions for executing the steps of a PCL method for determining the load of at least one module of a network platform as described with reference to [Fig.7],
[0100] In particular, this computer program may be a PGS program comprising instructions for executing the steps of a PFS method for providing a service as described with reference to [Fig.7].
Claims
Claims
1. Method (PTR) for training a model (MOD;) for determining the load of a network platform (PS), this method being implemented by a computer and comprising the following steps: - obtaining (E10) a history of requests (RQir) sent by at least one third party (T;) and received by the network platform (PS) during at least one past period of time (PTr); - obtaining (E20) at least one load (LOADir, LOADkir) of at least one module of said network platform (PS) during said at least one period of time (PTr); - training (E80) the model (MOD;) with training data comprising at least one temporal data (HD, CAL, EVT) representative of said at least one period of time and said at least one load (LOADi r, LOADki r), a use of the trained model making it possible to determine the load (LOADi r, LOADki r) of said at least one module of said network platform (PS) during a period of time (PTr) provided as input to the model.;
2. Training method according to claim 1 further comprising: - obtaining (E20) at least one latency (LATir, LATki>r) of said at least one module of said network platform (PS) during said at least one time period (PTr); - said training data comprising said at least one latency (LATi>r, LATki>r), a use of the trained model further making it possible to determine the latency (LATi>r, LATki>r) of said at least one module of said network platform (PS) during said time period (PTr) provided as input to the model.
3. Training method according to claim 1 or 2 wherein the training data further comprises an identifier of the third party (Tj) issuing said at least one request or an identifier of an application within the framework of which said at least one request was issued.
4. Training method according to any one of claims 1 to 3 further comprising: - a step (E70) of comparing a real load or latency (LOADRir, LOADRkirLATRir, LATRki>r) of at least one module of said network platform with the determined load or latency (LOADR; r, LOAD^rLAT^r, LAT^J of said at least one module over a period of time and; - a step (E80) of retraining said model based on a result of said comparison.
5. Training method according to any one of claims 1 to 4 wherein said model (MODi) is a recurrent neural network or wherein said model (MODi) uses a forest comprising at least one decision tree.
6. Method (PCL) for determining the load of at least one module of a network platform (PS), this method comprising a step (P20) of using a load determination model driven by a method according to any one of claims 1 to 5 to determine at least the load of the network platform (PS) during a period of time provided as input to the network.
7. Method (PFS) for providing a network service implemented by a network platform and comprising the following steps: - obtaining (P 10) an estimate of the load of at least one module of said network platform (PS) during a period of time by implementing a determination method according to claim 6; - adjusting (P30) resources of said at least one module for said period of time as a function of a result of said determination; - processing (P40) at least one request issued by a third party to access said service using said adjusted resources.
8. Entity (LPS) for determining the load of at least one module of a network platform, this entity exposing at least one application program interface (APIb API2, API3) making it possible to trigger the execution of at least one use of at least one load determination model trained by a training method according to any one of claims 1 to 5 to determine the load of at least one module (MFk) of a network platform (PS) during a period of time provided as input to the network.
9. Network platform (PS) comprising: - a module (COM) configured to interrogate a load determination entity according to claim 6 to obtain an estimate of the load of at least one module of said network platform (PS) over a period of time; - a module (MAR) for adjusting resources of said at least one module for said period of time depending on a result of said determination; - a module (MTR) for processing at least one request issued by a third party to access a network service provided by said platform using said adjusted resources.
10. Computer program (PGE, PGD, PGS) comprising instructions for executing the steps of a training method according to one of claims 1 to 5, or the steps of a method for determining the load of at least one module of a network platform according to claim 6 or the steps of a method for providing a service according to claim 7 when said program is executed by a computer.
11. A computer-readable recording medium on which a computer program (PGE, PGD, PGS) is recorded according to claim 10.
Citation Information
Patent Citations
Proactively accomodating predicted future serverless workloads using a machine learning prediction model and a feedback control system
US20210184941A1
Dynamic autoscaling of server resources using intelligent demand analytic systems
US20220383324A1