Resource management in a machine learning environment

The system addresses the challenges of resource management and security in machine learning environments by implementing a resource management system with a workflow module, resource controllers, and a controller manager, ensuring efficient sharing and secure access to generative machine learning models.

WO2025131556A1PCT designated stage expired Publication Date: 2025-06-26BRITISH TELECOM PLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/083425
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-11-25
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The increasing demand for generative machine learning models in multiple applications poses challenges in resource management, including limited access and inefficient use of computing resources, as well as security concerns related to data access and malicious use.

Method used

A system for resource management in a machine learning environment that includes a machine learning application API, a workflow module, first and second resource controllers, and a controller manager. This system manages traffic and applies application-specific flow control configurations to ensure efficient sharing and secure access to generative machine learning models.

Benefits of technology

The system enables efficient sharing of machine learning resources, prioritizes traffic, and ensures secure access, thereby addressing the challenges of limited resource access and security in machine learning environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024083425_26062025_PF_FP_ABST
    Figure EP2024083425_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A system for managing resources in a machine learning environment comprising one or more machine learning applications, the system comprising a machine learning application API; a workflow module of the machine learning environment; a generative machine learning model, a first resource controller configured to intercept traffic going to and from the workflow module and configured to perform flow control operations on the intercepted traffic in accordance with a first flow configuration; a second resource controller configured to intercept traffic going to and from the generative machine learning model and configured to perform flow control operations on the intercepted traffic in accordance with a second flow configuration; and a controller manager in communication with the first resource controller and the second resource controller, the controller manager configured to determine the first flow configuration for the first resource controller and the second flow configuration for the second resource controller.
Need to check novelty before this filing date? Find Prior Art

Description

Resource management in a machine learning environment

[0001] The present disclosure relates to managing resources in a machine learning environment to ensure the security of the machine learning environment and to enable generative machine learning models in the machine learning environment to be efficiently shared between multiple machine learning applications.BACKGROUND

[0002] Generative machine learning applications are becoming increasingly popular as a way to produce content and to help with daily business activities such as writing, coding and creating content. However, generative machine learning models require a large amount of computing resources to run and access to such generative machine learning models can be limited compared to the number of applications which wish to access the generative machine learning models.

[0003] In addition, as organizations and individuals increase the number of machine learning applications they use, it becomes increasingly important to have a manageable way to control the content generative machine learning models receive from and return to certain machine learning applications to ensure security of company and personal data and to prevent malicious use for machine learning models.

[0004] As the number of machine learning applications increases it is becoming increasingly impracticable in terms of computing resources to provide a separate instance of a generative machine learning model for each machine learning application. Thus, ways to manage and control access of machine learning applications to machine learning resources such as generative machine learning models is becoming increasingly important.

[0005] The examples described herein are not limited to examples which solve problems mentioned in this background section.SUMMARY

[0006] Examples of preferred aspects and embodiments of the invention are as set out in the accompanying independent and dependent claims.

[0007] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0008] In one aspect, the present invention accordingly provides a system for resource management in a machine learning environment with one or more machine learning applications, comprising: a machine learning application API, application programming interface, to receive and respond to requests for using a generative machine learning model; a workflow module to process requests from the API, determine and execute a necessary workflow involving the generative model, and generate and return a response, wherein the generative machine learning model is arranged to generate content based on the operation of the workflow; first and secondresource controllers managing traffic to the workflow module and the generative model, respectively, wherein the first and second resource controllers manage traffic using applicationspecific flow control configurations; and a controller manager to devise and supply the flow configurations to both resource controllers.

[0009] In a first embodiment a system for managing resources in a machine learning environment comprising one or more machine learning applications is provide. The system comprises a machine learning application API, application programming interface, of a machine learning application of the one or more machine learning application, wherein the machine learning application API is configured to receive a request to use a generative machine learning model in the machine learning environment from a user interface and to provide, to the user, a response to the request via the user interface. The system also comprises a workflow module of the machine learning environment, wherein the workflow module of the machine learning application is configured to: receive the request from the machine learning API; determine a workflow associated with the request, wherein the workflow comprises an operation to be performed by the generative machine learning model; send an indication of the operation to be performed by the generative machine learning model to the generative machine learning model; receive content from the generative machine learning model in response to the indication; generate the response to the request based on the content from the generative machine learning model; and send the response to the request to the machine learning application API. The system further comprises a generative machine learning model configured to generate content in response to the indication of the operation to be performed provided by the workflow; a first resource controller configured to intercept traffic going to and from the workflow module and configured to perform flow control operations on the intercepted traffic in accordance with a first flow configuration received from a controller manager, wherein the first flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; a second resource controller configured to intercept traffic going to and from the generative machine learning model and configured to perform flow control operations on the intercepted traffic in accordance with a second flow configuration received from the controller manager, wherein the second flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; and a controller manager in communication with the first resource controller and the second resource controller, the controller manager configured to determine the first flow configuration and second flow configuration and provide the first flow configuration to the first resource controller and the second flow configuration to the second resource controller. The use of a system with a first and second resource controller implementing flow control in accordance with flow configurations from a controller manager enables easy control of machine learning resource use. In some examples the request can comprise a user request,which in some examples, may be received by a user interface. In some examples alternatively or in addition the request may comprise an API call.

[0010] In some examples, the system further comprises a data store configured to store information as data vectors and provide returned data to the workflow module in response to requests for data; and a third resource controller configured to intercept traffic going to and from the data store and configured to perform flow control operations on the intercepted traffic in accordance with a third flow configuration received from the controller manager, wherein the third flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications. In these examples the workflow determined by the workflow module further comprises a request to obtain data from a data store; the workflow module is further configured send the request to the data store to obtain the data from the data store and receive the returned data from the data store; the indication of the operation to be performed by the generative machine learning model further comprises the returned data; and the controller manager is further configured to determine the third flow configuration and provide the third flow configuration to the third resource controller. Hence, access control and resource management can also be applied in a machine learning environment that supplements a generative machine learning model with content from a data store.

[0011] In some examples, at least one of the first, second or third flow configuration for the respective resource controller comprises a respective traffic level trigger condition and respective traffic level flow control rules, wherein: the respective traffic level condition specifies a condition under which the respective traffic level flow control rules apply, and the respective traffic level flow control rules comprise: a respective batch size, wherein the respective batch size defines a number of messages that should be in a queue of the respective resource controller before the message are processed; and respective machine learning application specific control rules that define the machine learning application specific flow control configurations, wherein the respective machine learning application specific control rules define a priority for the corresponding machine learning application and either a rate limit of messages or a quota limit of messages for the corresponding machine learning application and cause the respective resource controller to reorder the messages, delay at least some messages, drop at least some messages, or reroute at least some messages to meet the machine learning application specific control rules. Hence, the resource controllers can be used to prioritise traffic or limit traffic from certain machine learning applications.

[0012] In some examples, the controller manager is further configured to: receive a request to add a new machine learning application to the one or more machine learning applications; run the new machine learning in a test environment to obtain a maximum allowable response time for the workflow module, the machine learning model, and the data store in order to run the new machine learning application; and use a heuristic search algorithm to determine the first flowconfiguration, the second flow configuration, and the third flow configuration. Hence, the use of a controller manager allows new machine learning applications to be added to the machine learning environment without degrading the performance of a machine learning environment beyond an acceptable level. In some examples, the heuristic search algorithm is a genetic learning algorithm.

[0013] In some examples, the controller manger being configured to use a heuristic search algorithm to determine the first flow configuration, the second flow configuration, and the third flow configuration comprises the controller manager being configured to: assign initial machine leaning application specific control rules for the first flow configuration, the second flow configuration, and the third flow configuration, to each machine learning application of the one or more machine learning applications and the new machine learning application; and iteratively: check whether the new machine learning application and the one or more machine learning applications are meeting their maximum response times for the workflow module, the machine learning model and the data store given the assigned machine learning application specific control rules; and in response to all maximum response times being met, return the application specific control rules for the first flow configuration, the second flow configuration, and the third flow configuration; else in response to the maximum response times not all being met, perturb the machine learning application specific control rules to form new assigned machine learning application specific control rules. Thus, the heuristic search algorithm can find a suitable solution for adding a new machine learning application into a machine learning environment. In some examples, the heuristic search algorithm is a genetic learning algorithm.

[0014] In a further aspect, the present invention accordinglyh provides computer-implemented method for managing resources in a machine learning environment running one or more machine learning applications, comprising: receiving and responding to a request for using a generative machine learning model via an application API, application programming interface, wherein the request is processed by a workflow module of the machine learning environment to determine and execute a necessary workflow involving the generative model, and generate and return a response, wherein the generative machine learning model is arranged to generate content based on the operation of the workflow; applying first and second resource controllers to manage traffic to the workflow module and the generative model, respectively, wherein the first and second resource controllers manage traffic using application-specific flow control configurations; and devising and supplying the flow configurations to both resource controllers.

[0015] A second embodiment of the application defines a computer-implemented method for managing resources in a machine learning environment that is running one or more machine learning applications. The method comprises receiving at a machine learning application API, application programming interface, a request to use a generative machine learning model in the machine learning environment to generate content, wherein the request is received from a user interface module for a first machine learning application of the one or more machine learning application in the machine learning environment; sending, by the user interface module, a firstmessage comprising the request to a workflow module of the machine learning environment; intercepting, by a first resource controller associated with the workflow module, the first message; applying, by the first resource controller, flow control to the first message, wherein the flow control is in accordance with a first flow configuration received from a controller manager, wherein the first flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; forwarding, by the first resource controller, the first message to the workflow module in accordance with the flow control; determining, by the workflow module, a workflow associated with the request from the first message, wherein the workflow comprises an operation to be performed by the generative machine learning model; sending, by the workflow module, to the generative machine learning model, a second message comprising an indication of the operation to be performed by the generative machine learning model; intercepting, by the first resource controller, the second message; applying, by the first resource controller, flow control to the second message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; sending, by the first resource controller, the second message to the generative machine learning model in accordance with the flow control; intercepting, by a second resource controller associated with the generative machine learning model, the second message; applying, by the second resource controller, flow control to the second message, wherein the flow control is in accordance with a second flow configuration received from the controller manager, wherein the second flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; forwarding, by the second resource controller, the second message to the generative machine learning model in accordance with the flow control; generating, by the generative machine learning model, content in accordance with the indication of the operation in the second message; sending, by the generative machine learning model, to the workflow module, the content to the workflow module as a third message; intercepting, by the second resource controller, the third message sent to the workflow module; applying, by the second resource controller, flow control to the third message, wherein the flow control is in accordance with the second flow configuration received from the controller manager; forwarding, by the second resource controller, the third message to the workflow module in accordance with the flow control; intercepting, by the first resource controller, the third message; applying, by the first resource controller, flow control to the third message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; forwarding, by the first resource controller, the third message to the workflow module; generating, by the workflow module, a response to the request based on the content in the third message received from the generative machine learning model; sending, by the workflow module, the response to the request to the machine learning application API as a fourth message; intercepting, by the first resource controller, the fourth message; applying, by the first resource controller, flow control tothe fourth message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; forwarding, by the first resource controller, the fourth message to the machine learning application API in accordance with the flow control; and providing, by the machine learning application API, the response to the request from the fourth message to the user. Thus, a method of using resource controllers to manage message flow to and from resources in a machine learning environment is defined. In some examples the request can comprise a user request, which in some examples, may be received by a user interface. In some examples alternatively or in addition the request may comprise an API call.

[0016] In some examples, the first flow configuration comprises a first traffic level trigger condition, and first traffic level flow control rules, wherein the first traffic level control rules comprise a first batch size and first machine learning application specific flow control rules that define the machine learning application specific flow control configurations and applying flow control in accordance with the first flow configuration comprises: checking to see if the first traffic level trigger condition is met; and in response to the first traffic level trigger condition not being met, forwarding any messages received at the first resource controller to the workflow module; else in response to the first traffic level trigger condition being met: storing messages received at the first resource controller in queue until a number of messages in the queue is equal to the first batch size; and in response to the number of messages in the queue being equal to the first batch size, processing the messages in the queue in accordance with the first machine learning application specific control rules by reordering the messages, delaying at least some messages, dropping at least some messages, or rerouting at least some messages to ensure the first application specific control rules are met. Hence, the first resource controller can prioritize or restrict traffic from certain machine learning applications.

[0017] In some examples, the second flow configuration comprises a second traffic level trigger condition, and second traffic level flow control rules, wherein the second traffic level control rules comprise a second batch size and second machine learning application specific flow control rules that define the machine learning application specific flow control configurations, and applying flow control in accordance with the second flow configuration comprises: checking to see if the second traffic level trigger condition is met; and in response to the second traffic level trigger condition not being met, forwarding any messages received at the first resource controller to the machine learning model; else in response to the second traffic level trigger condition being met: storing messages received at the second resource controller in queue until a number of messages in the queue is equal to the second batch size; and in response to the number of messages in the queue being equal to the second batch size, processing the messages in the queue in accordance with the second machine learning application specific control rules by reordering the messages, delaying at least some messages, dropping at least some messages, or rerouting at least some messages to ensure the second application specific control rules are met. Hence, the second resource controller can prioritize or restrict traffic from certain machine learning applications.

[0018] In some examples, the second routing configuration further comprises access control rules for the machine learning model, wherein the access control rules indicate whether particular users and / or machine learning applications are allowed access to the machine learning model; and applying flow control in accordance with the second flow configuration further comprises: determining if any messages received at the second resource controller meets the access control rules; and in response to the messages received at the second resource controller meeting the access control rules implementing the flow control described above; else in response to a message received at the second resource controller not meeting the access control rules, rerouting the message to another machine learning model or workflow module. Thus, the second resource controller can ensure only machine learning applications or users of machine learning applications that should access a generative machine learning model have access to that generative machine learning model.

[0019] In some examples, the first or second machine learning application specific control rules define a priority for the corresponding machine learning application and either a rate limit of messages or a quota limit of messages for the corresponding machine learning application. This allows certain machine learning applications to be prioritized or deprioritized.

[0020] In some examples, the workflow determined by the workflow module further comprises a request to obtain data from a data store of the machine learning environment, and the method further comprises: sending, by the workflow module, the request to obtain data to the data store as a fifth message; intercepting, by the first resource controller, the fifth message; applying, by the first resource controller, flow control to the fifth message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; forwarding, by the first resource controller, the fifth message in accordance with the flow control; intercepting, by a third resource controller associated with the data store, the fifth message; applying, by the third resource controller, flow control to the fifth message, wherein the flow control is in accordance with a third flow configuration received from the controller manager, wherein the third flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; forwarding, by the third resource controller, the fifth message to the data store in accordance with the flow control; obtaining, from the data store, returned data in response to the request to obtain data from the fifth message; sending, by the data store, to the workflow module, the returned data as a sixth message; intercepting, by the third resource controller, the sixth message; applying, by the third resource controller, flow control to the sixth message, wherein the flow control is in accordance with the third flow configuration received from the controller manager; forwarding, by the third resource controller, the sixth message to the workflow module in accordance with the flow control; intercepting, by the first resource controller, the sixth message; applying, by the first resource controller, flow control to the sixth message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; and forwarding, by the first resourcecontroller, the sixth message to the workflow module in accordance with the flow control; wherein: the second message sent by the workflow module to the generative machine learning model further comprises the returned data from the data store. Thus, a data store can be added to the machine learning environment and access to the data store can be controlled by a third resource controller.

[0021] In some examples, in response to a request to add a new machine learning application to the machine learning environment: running the new machine learning application in a test environment to determine a maximum response time for the workflow module and a maximum response time for the machine learning model required to implement the new machine learning application; providing, by the test environment, the determined maximum response time for the workflow module and the maximum response time for the machine learning model to a optimizer of the controller manager; using, by the controller manager, a heuristic search algorithm to determine the first flow configuration and the second flow configuration. Thus, the controller manager can be used to determine if a new machine learning application can be added to the machine learning environment without overly negatively impacting other machine learning applications. In some examples, the heuristic search algorithm is a genetic learning algorithm.

[0022] In some examples, the method further comprising: providing, by the test environment, the determined maximum response time for the data store to the optimizer of the controller manager; and wherein using, by the controller manager, the heuristic search algorithm to determine the first flow configuration and the second flow configuration further comprises using, by the controller manager, the heuristic search algorithm to determine the third flow configuration. Thus, the controller manager can also be used to determine if a new machine learning application can be added to a machine learning environment when the machine learning environment also comprises a data store. In some examples, the heuristic search algorithm is a genetic learning algorithm.

[0023] In some examples, using the heuristic search algorithm to determine the first flow configuration and the second flow configuration comprises: assigning machine leaning application specific control rules for the first flow configuration, the second flow configuration, and where a data store is used, the third flow configuration, to each machine learning application of the one or more machine learning applications and the new machine learning application, wherein the application specific control rules specify for the respective flow configuration and corresponding machine learning application a priority for the machine learning application and either a rate limit of messages or a quota limit of messages for the machine learning; and iteratively: checking whether the new machine learning application and the one or more machine learning applications are meeting their maximum response times for the workflow module, the machine learning model and, where applicable, the data store; and in response to all maximum response times being met, returning the application specific control rules for the first flow configuration, the second flow configuration, and where applicable, the third flow configuration; else in response to the maximumresponse times not all being met, perturbating the machine learning application specific control rules. Thus, a heuristic search algorithm to determine suitable flow configurations is defined. In some examples, the heuristic search algorithm is a genetic learning algorithm.

[0024] In some examples, using the heuristic search algorithm to determine the first flow configuration, the second flow configuration and, where applicable, the third flow configuration further comprises, in response to the maximum response times not all being met: confirming whether a stop condition is met, where the stop condition specifies a maximum number of iterations to be performed; and in response to the stop condition being met returning an indication that it is not possible to add the new machine learning application; else in response to the stop condition not being met perturbating the machine learning application specific control rules. Hence, a stop condition can be added to the heuristic search algorithm to ensure that if a new machine learning application cannot be added to the machine learning environment without negatively impacting performance, a user is alerted and the new machine learning application is not added. In some examples, the heuristic search algorithm is a genetic learning algorithm.

[0025] It will also be apparent to anyone of ordinary skill in the art, that some of the preferred features indicated above as preferable in the context of one of the aspects of the disclosed technology indicated may replace one or more preferred features of other ones of the preferred aspects of the disclosed technology. Such apparent combinations are not explicitly listed above under each such possible additional aspect for the sake of conciseness.

[0026] Other examples will become apparent from the following detailed description, which, when taken in conjunction with the drawings, illustrate by way of example the principles of the disclosed technology.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 shows a machine learning environment in accordance with the present disclosure;

[0028] Figure 2 shows a method of managing machine learning resources in a machine learning environment;

[0029] Figure 3 shows a machine learning environment including a data store in accordance with the present disclosure;

[0030] Figure 4 shows a method of managing a data store in a machine learning environment;

[0031] Figure 5 shows an example resource controller in accordance with the present disclosure;

[0032] Figure 6 shows a method for adding a new machine learning application into a machine learning environment; and

[0033] Figure 7 shows a computing device that can be used to implement the resource controllers, resources or controller manager of the present disclosure.

[0034] The accompanying drawings illustrate various examples. The skilled person will appreciate that the illustrated element boundaries (e.g., boxes, groups of boxes, or other shapes) in the drawings represent one example of the boundaries. It may be that in some examples, oneelement may be designed as multiple elements or that multiple elements may be designed as one element. Common reference numerals are used throughout the figures, where appropriate, to indicate similar features.DETAILED DESCRIPTION

[0035] The following description is made for the purpose of illustrating the general principles of the present technology and is not meant to limit the inventive concepts claimed herein. As will be apparent to anyone of ordinary skill in the art, one or more or all of the particular features described herein in the context of one embodiment are also present in some other embodiment(s) and / or can be used in combination with other described features in various possible combinations and permutations in some other embodiment(s).

[0036] Generative machine learning models are becoming increasingly common for use by both individuals and organizations. While generative machine learning models can be useful, they often consume a large number of resources and access to the generative machine learning models may be limited compared to the number of machine learning applications. An organization may wish to run several machine learning applications which share generative machine learning models and potentially data stores that contain information used to supplement the generative machine learning models. As the number of machine learning applications sharing machine learning resources such as generative machine learning models increases, management of the machine learning resources needs to improve.

[0037] Figure 1 shows a system 100 for managing resources in a machine learning environment that comprises one or more machine learning applications. Each machine learning application is formed from a machine learning application API (application programming interface) 110, a workflow module orworkflow component 120, and a generative machine learning model 130. The skilled person would be aware of generative machine learning models and any suitable generative machine learning model could be used as the generative machine learning model in the machine learning environment. In some examples, the generative machine learning model may be a bespoke generative machine learning model. However, in other examples, known or otherwise available generative machine learning models may be used. Examples of known generative machine learning models include ChatGPT, GPT-4, Cohere, Claude, Palm, Flan-T5, Falcon, llama, code Llama, star coder, Amazon Titian, Gemini, Bloom, Copilot, Stable diffusion and Bard.

[0038] In Figure 1 , the machine learning environment comprises only a single machine learning application AP1 110, workflow module 120 and generative machine learning model 130. However, in other examples further machine learning APIs, workflow modules 12 and generative machine learning models 130 can be used. In particular, each machine learning application of the one or more machine learning applications comprises its own machine learning application API through which it engages with a user interface and the workflow modules 120 and hence generative machine learning models 130 in the machine learning environment. The workflow module 120 or workflow modules 120 and the generative machine learning model 130 or generative machinelearning models 130 can comprise shared resources that are used for several machine learning applications. Thus, multiple machine learning applications can each comprise a separate machine learning application API 110 which engages with the same workflow module 120 and generative machine learning model 130. In other examples, a machine learning application may comprise its own workflow module 120. In additional examples, some machine learning applications in the environment may share a workflow module 120 while other machine learning applications may have their own workflow module 120. When the machine learning environment comprises more than one machine learning application, then the generative machine learning model 130 is generally a shared resource between multiple machine learning applications within the environment. In some examples, the environment may comprise more than one generative machine learning model 130 and some machine learning applications in the environment may use a first generative machine learning model 130 while other machine learning applications in the environment may use a second generative machine learning model 130. In addition, one or more applications in the machine learning environment may use more than one generative machine learning model 130.

[0039] When using a machine learning application in a machine learning environment, a user or another computing system provides a request to the machine learning application which is received at the machine learning application API either directly or through being forwarded by a user interface. In some examples the request can comprise a user request, which in some examples, may be received by a user interface. In some examples alternatively or in addition the request may comprise an API call. The request or user request can be authenticated by the machine learning API to confirm that the user or user interface or user computing device providing the requested is authorised to use the machine learning application. The authentication can also confirm which machine learning application is making the request. The request is then sent to a workflow module 120 or workflow component 120 and the workflow module 120 or workflow component 120 determines a workflow associated with the request. To this end the workflow module 120 or workflow component 120 determines operations to be performed by a generative machine learning model 130 in the environment and / or prompts to be sent to the generative machine learning model 130 in the environment to cause the generative machine learning model 130 to generate output. In some examples, where the machine learning application needs to use more than one request to one or more generative machine learning models 130, the workflow module 120 can determine a workflow comprising the necessary requests and / or prompts to send to the one or more generative machine learning models 130. The workflow module 120 sends the requests or prompts to the relevant machine learning models 130 in the required order and receives responses from the genitive machine learning models 130. Once all necessary responses have been received, the workflow module 120 uses the responses from the generative machine learning models 130 to produce a response to the request and returns the response to the request to the machine learning application API 110.

[0040] In the system 100 shown in Figure 1 , the system 100 further comprises a first resource controller 125, a second resource controller 135 and a controller manager 140. The first resource controller 125 is configured to intercept messages or traffic going to and from the workflow module 120 and the second resource controller 135 is configured to intercept messages or traffic going to and from the machine learning model 130. The controller manager 140 is in communication with the first resource controller 125 and the second resource controller 135 and provides the first resource controller 125 with a first flow configuration 127 and the second resource controller 135 with a second flow configuration 137. The first resource controller 125 enacts routing control on messages or traffic going to and from the workflow module 120 in accordance with the first flow configuration. The second resource controller 135 enacts routing control on messages or traffic going to and from the generative machine learning model 130 in accordance with the second flow configuration.

[0041] The first flow configuration 127 specifies machine learning application specific flow configurations to be implemented by the first resource controller 125 for each machine learning application in the machine learning environment. Similarly, the second flow configuration 137 specifies machine learning application specific flow configurations to be implemented by the second resource controller 135 for each machine learning application in the machine learning environment.

[0042] The use of a first resource controller 125 that intercept messages or traffic going to and from the workflow module 120 and a second resource controller 135 that intercepts messages or traffic going to and from the generative machine learning application 130 enables improved management and sharing of machine learning resources between multiple machine learning applications in a machine learning environment. In addition, having a first resource controller 125 intercept messages or traffic to a workflow module 120 and a second resource controller 135 intercept messages or traffic going to a generative machine learning application 130 enables messages or traffic to be reviewed for permissions to access the workflow module 120 or machine learning application 130 thus improving the operational security of a machine learning environment.

[0043] It is noted that while the above system 100 for managing a machine learning environment has been described with both a first resource controller 125 for a workflow module 120 and a second resource controller 135 for a machine learning model 130, in some examples where, for example, a machine learning application has its own workflow module 120 the first resource controller 125 may not be present as the workflow module 120 is not a shared resource so management of traffic / message flow can be performed by the machine learning application API 110 or may not be necessary. In addition, in examples where the user input is sufficient to interact directly with the generative machine learning model 130, the workflow module 120 and the corresponding first resource controller 125 may not be present and the machine learning application API 110 may communicate directly with the generative machine learning model 130with traffic or messages in either direction being intercepted by the second resource controller 135.

[0044] Figure 2 is a flowchart illustrating a method 200 for managing resources in a machine learning environment that can be performed in a system for managing resources in a machine learning environment such as system 100 described above.

[0045] In a first step 212, a request is received at an API 110 of a machine learning application wherein the request is a request to use a generative machine learning model 130 in the machine learning environment. In some examples the request can comprise a user request, which in some examples, may be received by a user interface. In some examples alternatively or in addition the request may comprise an API call. The request can comprise a direct request to use the generative machine learning model 130 or can comprise a request for content that, for example, the workflow module 120 can determine requires the use of the generative machine leaning model 130. Hence, the user may not know what generative machine learning model 130 they are interacting with as part of the request.

[0046] At step 214, the machine learning application API 110 forwards the request to a workflow module 120 for further processing. The forwarded request can be considered a first message. In some examples, before forwarding the request, the machine learning API 110 can do any appropriate authentication, for example checking the user or user computing device from which the request is received is authorised to use the machine learning application associated with the machine learning application API 110 and / or confirming a machine learning application to which the request should be assigned.

[0047] At step 216, the first resource controller 125 associated with the workflow module 120 intercepts the first message. In some examples, the first resource controller 125 comprises a message bus and the machine learning application API 110 forwarding the first message to the workflow module 120 can comprise the first message being placed in a message bus queue. In another example, the first resource controller 125 can comprise a shim that can intercept and / or otherwise receive the first message. Any other suitable way of implementing the first response controller 125 can be used.

[0048] In step 218, the first resource controller 125 applies flow control to the first message in accordance with the first flow configuration 127 received from the controller manager 140 and forwards the first message to the workflow module 120 in accordance with the flow control determined by the first resource controller. The first flow configuration specifies machine learning application specific flow control configurations for each machine learning application in the machine learning environment. The first flow configuration 127 enables the first resource controller 125 to determine a priority for messages to / from the workflow module 120 and also enables the first resource controller 125 to determine if any messages to / from the workflow module 120 should be dropped, delayed or rerouted. As explained in more detail later, the flow control performed is machine learning application specific by virtue of the machine learningapplication specific flow control configurations and may comprise reordering messages being sent to / from the workflow module 120, dropping messages, delaying messages and / or rerouting messages.

[0049] At step 220, the workflow module 120 processes the received first message. This comprises the workflow module 120 obtaining or determining a workflow associated with the request, wherein the workflow comprises at least an operation to be performed by a generative machine learning model 130. The operation can comprise a prompt to the generative machine learning model 130 to cause the generative machine learning model 130 to generate content. The workflow module 120 processing the request may comprise the workflow module 120 inputting the request and then determining a workflow from the request either based on a workflow provided in the request or by processing the request to split it into individual steps. Such workflow modules 120 exist in machine learning applications and the skilled person would be familiar with how to implement them. For example, the workflow module 120 could be implemented using the LangChain language as would be known to the person skilled in the art.

[0050] At step 222, the workflow module 120 sends an indication of the operation to be performed by the generative machine learning model 130 or a prompt for the generative machine learning model 130 to the generative machine learning model 130. This indication or prompt is sent as a second message. In some examples, the workflow module 120 may be designed or configured without awareness of the first resource controller 125. Thus, the workflow module 120 may send the second message without awareness of the first resource controller 125, which as defined in the next step, will intercept the second message. However, in other examples, the workflow module 120 may be aware of the first resource controller 125 so may send the second message via the first resource controller 125.

[0051] In step 224, the first resource controller 125 intercepts the second message as the second message is sent from the workflow module 120. As above, the first resource controller 125 can comprise a message bus or a shim and the first resource controller 125 can intercept the second message as described above. In addition, in some examples, the workflow module 120 may send the second message directly to the first resource controller 125.

[0052] At step 226, the first resource controller 125 applies flow control to the second message in accordance with the first flow configuration 127 and then sends the second message to the generative machine learning model 130 in accordance with the flow control. As mentioned above, the first flow configuration 127 specifies machine learning application specific flow control configurations for each machine learning application in the machine learning environment. The first flow configuration 127 enables the first resource controller 125 to determine a priority for messages to / from the workflow module 120 and also enables the first resource controller 125 to determine if any messages to / from the workflow module 120 should be dropped, delayed or rerouted. As explained in more detail later, the flow control performed is machine learning application specific by virtue of the machine learning application specific flow controlconfigurations and may comprise reordering messages being sent to / from the workflow module 120, dropping messages, delaying messages and / or rerouting messages.

[0053] In step 228, a second resource controller 135 associated with the generative machine learning model 130 intercepts the second message. In some examples, the second resource controller 135 comprises a message bus queue and the second resource controller 135 intercepting the second message comprises the second message being added to a message bus queue of the message bus. In other examples, the second resource controller 135 comprises a shim that receives the second message and then performs further processing on the second message before forwarding the second message.

[0054] At step 230, the second resource controller 135 applies flow control to the second message in accordance with a second flow configuration 137 received from the controller manager 140 and forwards the second message in accordance with the flow control. The second flow configuration 137 specifies machine learning application specific flow control configurations for each machine learning application in the machine learning environment. The second flow configuration 137 enables the second resource controller 135 to determine a priority for messages to / from the generative machine learning model 130 and also enables the second resource controller 135 to determine if any messages to / from the generative machine learning model 130 should be dropped, delayed or rerouted. As explained in more detail later, the flow control performed is machine learning application specific by virtue of the machine learning application specific flow control configurations and may comprise reordering messages being sent to / from the generative machine learning application 130, dropping messages, delaying messages and / or rerouting messages.

[0055] At step 232, the second message is received by the generative machine learning model 130 and the indication of the operation and / or the prompt from the second message is provided to the generative machine learning model 130 to cause the generative machine learning model 130 to generate content. The generative machine learning model 130 can be any known generative machine learning model suitable for use with the machine learning application and the skilled person would be familiar with how to implement such generative machine learning models 130.

[0056] At step 234, the generative machine learning model 130 sends the generated content to the workflow module 120 as a third messages. In some examples, the generative machine learning model 130 may be unaware of the second resource controller 135 so may direct the third message to the workflow module 120. In other examples, the generative machine learning model 130 may be aware of the second resource controller 135 and may direct the third message to the second resource controller.

[0057] In either case, at step 236, the second resource controller intercepts the third message and at step 238, the second resource controller applies flow control to the third message in accordance with the second flow configuration then forwards the third message in accordancewith the flow control. The second flow configuration 137 specifies machine learning application specific flow control configurations for each machine learning application in the machine learning environment. The second flow configuration 137 enables the second resource controller 135 to determine a priority for messages to / from the generative machine learning model 130 and also enables the second resource controller 135 to determine if any messages to / from the generative machine learning model 130 should be dropped, delayed or rerouted. As explained in more detail later, the flow control performed is machine learning application specific by virtue of the machine learning application specific flow control configurations and may comprise reordering messages being sent to / from the generative machine learning application 130, dropping messages, delaying messages and / or rerouting messages.

[0058] At step 240, the first resource controller 125 which is associated with the workflow module 120 intercepts the third message and, in step 242, performs flow control on the third message in accordance with the first flow configuration as defined above. The first resource controller 125 then forwards the third message to the workflow module 120 in accordance with the flow control.

[0059] At step 244, the workflow module 120 uses the content from the generative machine learning model 130 in the third message to generate or form a response to the request, which may be a user request. In some examples alternatively or in addition the request may comprise an API call. In some examples, the response to the request may be formed from a single query to or interaction with the generative machine learning model 130. However, in other examples multiple queries or interactions with one or more generative machine learning models 130 may be used to form the response to the user query. In addition, as explained in more detail below, queries to or interactions with a data store may also be used in generating the response to the user query.

[0060] Once a response to the user query has been generated, the workflow module 120 sends the response to the machine learning application API 110 as a fourth message at step 246. The machine learning application API can then communicate the response to the user, for example via a user interface. In examples where the workflow module 120 is aware of the first resource controller 125, the workflow module 120 can forward the fourth message to the first resource controller 125. However, in other examples, the workflow module 120 may send the fourth message to the machine learning application API 110 and the first resource controller 125 can intercept the fourth message without the workflow module 120 being aware of the first resource controller 125.

[0061] In step 248, the first resource controller 125 intercepts the fourth message as described above with respect to the second message. At step 250, the first resource controller then applies flow control to the fourth message in accordance with the first flow configuration 127 as also described above, then forwards the fourth message to machine learning application API in accordance with the flow control.

[0062] At step 252, the machine learning API can receive the fourth message and provide the response to the request from the fourth message to the user, for example via a user interface of the machine learning application.

[0063] In the above method, the first resource controller 125 intercepts all messages being sent to / from the workflow module 120 and the second resource controller 135 intercepts all messages being sent to / from the generative machine learning model 130. This enables messages to / from the shared resources of the machine learning application to be controlled to enable messages to be prioritised and to enable filtering of messages for security purposes. While the above method has been described in an embodiment that includes a shared workflow module 120, the skilled person would understand that in some examples a machine learning application may have its own workflow module 120 in which case there may be no first resource controller 125 and messages to / from the workflow module 120 may be sent / received directly without going via the first resource controller 125 for flow control. The skilled person would also understand that in some cases the machine learning application may have the machine learning application API communicate directly with the generative machine learning model 130 without a workflow module 120. In this case, while the second resource controller 135 is still present and implementing the method steps described above, the workflow module 120 and first resource controller 125 may not be present and the method steps performed by these modules may be skipped.

[0064] In the above method, the first flow configuration and the second flow configuration are received from a controller manager in the machine learning environment. As explained in more detail later the controller manager can adapt the first and second flow configurations when new machine learning applications are added to the machine learning environment. The first and second flow configurations comprise machine learning application specific flow configurations that specify for example priorities and maximum response times for the machine learning applications.

[0065] While the description has concentrated on an example where a machine learning application uses a workflow module 120 and a generative machine learning model 130, in some examples the machine learning application may also make use of a data store 150.

[0066] Figure 3 shows an example system 100 in a machine learning environment that also comprises a data store 150 and a third resource controller 155. The data store 150 can for example be an enterprise specific data store that stores data for an organization or other enterprise. The data stored in the data store 150 is stored as vectors and can be used with a generative machine learning model 130 to enable the generative machine learning model 130 to provide content based on data in the data store 150 without retraining the generative machine learning model 130. For example, when the data in the data store is enterprise specific data, the data in the data store 150 can be used to adapt the generative machine learning model 130 to the enterprise without requiring the enterprise to perform additional training on the generative machine learning model 130. The use of data from data stores such as data store 150 to guidegenerative machine learning model 130 output is known and the skilled person would know how to implement such a setup.

[0067] As mentioned above, in the system shown in Figure 3, as well as a data store 150, the system 100 further comprises a third resource controller 155. The third resource controller 155 intercepts messages or traffic going to and from the data store 150 and performs flow control operations on the intercepted messages or traffic in accordance with a third flow configuration. The third flow configuration 157 is provided by the controller manager 140 and comprises application specific flow control configurations for each machine learning application in the machine learning environment. The flow control operations can comprise reordering, delaying, dropping or rerouting messages to / from the data store 150 in accordance with the third flow configuration.

[0068] When a data store 150 is used, the workflow determined by the workflow module 120 can further comprise a request to obtain data from the data store 150. This request can be sent from the workflow module 120 to the data store 150 being intercepted by the first resource controller 125 and the third resource controller 155 as described later. In response the request to obtain data, the data store 150 can provide the requested data to the workflow module 120 via the first resource controller 125 and the third resource controller 155. The workflow module 120 can provide the obtained data to the generative machine learning model 130 as part of the request or prompt sent to the generative machine learning model 130.

[0069] Using the third resource controller 155 in accordance with a third flow configuration 157, enables a data store 150 to be incorporated into the system while also allowing the data store 150 to be efficiently and fairly shared between machine learning applications and ensuring secure access control to the data store 150.

[0070] Figure 4 shows a method 400 performed as part of method 200 when the system 100 further comprises a data store 150. Method 400 is performed between step 220 and step 222 of the method from Figure 2.

[0071] In relation to method 400, the workflow obtained by the workflow module 130 in step 220 from Figure 2 comprises, as well as an operation to be performed by a generative machine learning model 130, a request to obtain data from the data store 150 of the machine learning environment. In step 410 of method 400, this request to obtain data is sent by the workflow module 130 to the data store 150 as a fifth message. In some examples, the workflow module 120 may be aware of the first resource controller 125 and send the fifth message via the first resource controller 125. However, in other examples, the workflow module 130 may not be aware of the first resource controller 125 and may direct the fifth message to the data store 150.

[0072] In step 412, the first resource controller 125 intercepts the fifth message. The first resource controller 125 can intercept the fifth message in any suitable way including those described above where the first resource controller 125 comprises a message bus or a shim.

[0073] After intercepting the fifth message, the first resource controller 125 applies flow control to the fifth message in accordance with the first flow configuration 127 discussed above and then forwards the fifth message in accordance with the flow control. This is shown in step 414.

[0074] At step 416, the third resource controller 155 intercepts the fifth message. As with the first resource controller 125 and the second resource controller 135, the third resource controller can comprise a message bus or a shim. The third resource controller 155 intercepting the fifth message can comprise the fifth message being added to a message bus queue or the shim receiving the fifth message.

[0075] At step 418, the third resource controller 155 applies flow control to the fifth message in accordance with a third flow configuration 157 received from the controller manager 110 and forwards the fifth message to the data store 150 in accordance with the third flow configuration 157. The third flow configuration 157 specifies machine learning application specific flow control configurations for each machine learning application in the machine learning environment. The third flow configuration 157 enables the third resource controller 155 to determine a priority for messages to / from the data store 150 and also enables the third resource controller 155 to determine if any messages to / from the data store 150 should be dropped, delayed or rerouted. As explained in more detail later, the flow control performed is machine learning application specific by virtue of the machine learning application specific flow control configurations and may comprise reordering messages being sent to / from the data store 150, dropping messages, delaying messages and / or rerouting messages.

[0076] At step 420 the data store 150 receives the fifth message and data is obtained from the data store 150 in accordance with the request to obtain data from the fifth message for example by a data store controller or a data store search engine. The skilled person would be familiar with how to run data stores 150 in order to obtain data from requests. In some examples, the data in the data store 150 can comprise data chunks in the form of vectors that can easily be used by generative machine learning model 130. However, any other form of data can also be used as appropriate.

[0077] At step 422, the obtained / returned data from the data store 150 is sent by the data store 150 or a controller at the data store 150 to the workflow module 130 as a sixth message. In some examples the data store 150 or data store controller may be aware of the third resource controller 155 and send the sixth message to the third resource controller 155. However, in other examples, the data store 150 or data store controller sends the sixth message to the workflow module 120 as it is unaware of the third resource controller 155.

[0078] In step 424, the third resource controller 155 intercepts the sixth message and in step 426, the third resource controller 155 applies flow control to the sixth message in accordance with the third flow configuration discussed above and then forwards the sixth message to the workflow module 130 in accordance with the applied flow control.

[0079] At step 428, the first resource controller 125 which is associated with the workflow module 130 intercepts the sixth message. In step 430, the first resource controller 125 applies flow control to the sixth message in accordance with the first flow configuration and then forwards the sixth message to the workflow module 120 in accordance with the applied flow control.

[0080] In step 222 of method 200, the workflow module 130 sends an indication of the operation to be performed by the generative machine learning model 130 or a prompt for the generative machine learning model 130 to the generative machine learning model 130. When method 400 is used, this indication or prompt further comprises the returned data from the data store 150 enabling a data store 150 to be incorporated into the workflow and enabling a data store 150 to be used to customize a generative machine learning model 130 to an enterprise.

[0081] As discussed above, the first, second and third flow configurations can be sent from a controller manager 140 to the relevant resource controller. One or more of the first, second and third flow configurations may comprise a respective traffic level trigger condition, respective traffic level control rules and, in some examples, respective access control rules.

[0082] The respective access level control rules comprise rules as to which machine learning applications and or users are able to access the resource of the machine learning application served by the respective resource controller. In other words, where they exist, first access level control rules in the first flow configuration used by the first resource controller 125 specify which machine learning applications and / or which users are able to access a workflow module 120 associated with the first resource controller 125. Where they exist, second access level control rules in the second flow configuration used by the second resource controller 135 specify which machine learning applications and / or which users are able to access a generative machine learning model 130 associated with the second resource controller 135. Where they exist, third access level control rules in the third flow configuration used by the third resource controller 155 specify which machine learning applications and / or which users are able to access the data store 150 associated with the third resource controller 155. In machine learning environments where there are multiple workflow modules, multiple generative machine learning applications and / or multiple data stores then each workflow module, generative machine learning application and / or data store has its own resource controller and own flow configuration which can have different access control rules. The access control rules can also include security requirements for the resource associated with the respective resource controller. The use of access control rules along with resource controllers enables access to be restricted to certain machine learning resources. For example, access control rules can be used to limit access to a generative machine learning model trained on sensitive or secret data of an organization to machine learning applications or users within that organizations. Similarly, access control rules can be used to limit access to a data store to machine learning applications or individuals who are allowed access to the secret or sensitive data in that data store.

[0083] When a resource controller such as first, second or third resource controller applies flow control in accordance with the resource controllers respective flow configuration this can comprise determining if any traffic or messages received at the resource controller meets the respective access control rules. For example, this can involve determining the user and / or machine learning application associated with a request is allowed to access the resource associated with the resource controller e.g. the workflow module 120 associated with the first resource controller 125, the generative machine learning model 130 associated with the second resource controller 135 or the data store 150 associated with the third resource controller 155. In addition, checking the access control rules are met can involve determining that the message or traffic meets security requirements e.g. the message or traffic can be checked for injection attacks, requests for sensitive or inappropriate content and / or the presence of sensitive or inappropriate content. In some examples, the access control rules can be implemented by checking the message against fingerprints to identify potentially malicious or otherwise insecure messages.

[0084] The checks of the access control rules can apply both to requests from the machine learning application API, which may be user requests and / or an API call, and requests form the workflow module 120 to the data store 150 and generative machine learning model 130 and to messages / traffic / information / content returned from the workflow module 120, data store 150 and generative machine learning model 130. For example, the access control rules can limit the information certain users and / or machine learning applications are allowed to request and the resources in the machine learning environment that certain users and / or machine learning applications are allowed to access and this can be confirmed by access control rules at the first, second and third resource controllers. In this regard, incoming messages received at a resource controller can be checked to confirm the user or machine learning application is allowed to access the resource. The incoming messages can also be checked to confirm they meet any security requirements and / or privacy requirements. If the incoming messages satisfy the access control rules they can be further processed by the relevant resource controller. In contrast, if the incoming messages do not satisfy the access control rules, they can be rerouted to either a resource in the machine learning environment the user and / or machine learning application is able to access or in the case of suspected injection attacks or other malicious messages to a security system or module for further analysis.

[0085] Outgoing messages can also be governed by access control rules. For example, the access control rules can specify limits on the content that can be returned from a generative machine learning model 130 to a particular user to machine learning application and / or limits on content that can be returned from a data store 150 to a particular machine learning application or user. In addition, the access control rules can specify general limits on content and or information that can be returned from a generative machine learning model 130 or data store 150 to all users. A resource controller applying flow control can comprise the resource controller confirming all outgoing messages from a resource associated with the resource controller meet the accesscontrol rules. If an outcoming message meets the access control rules, the resource controller can further processing the message. If the outgoing message does not meet the access control rules the resource controller can drop the message, remove sensitive or private content from the message or forward the message to a security module for further processing or analysis.

[0086] The use of resource controllers with flow configurations containing access control rules enables increased control of information in a machine learning environment and thus enhances the security of sensitive information and data as it prevents such data being accidentally returned to an incorrect user or machine learning application. Having the access control rules implemented by resource controllers also enables new resources such as generative new machine learning models 130 or new data stores 150 to be incorporated into the machine learning environment without the resources having to be specifically adapted to take into account which users and or machine learning applications can use them. In addition, having the flow configurations contain the access control rules be provided to the resource controllers from a controller manager 140 enables the access control rules to be easily updated in a quick and easy manner throughout the machine learning environment if the security status of a particular machine learning application or user changes.

[0087] One or more of the first, second and third flow control configurations 127, 137, and 157 can also comprise a respective traffic level trigger condition and respective traffic level control rules. The traffic level trigger condition for a respective flow control configuration defines a criteria at which the respective resource controller applying the flow control in accordance with the respective flow control configuration should apply the respective traffic level control rules. For example, the traffic level trigger condition can comprise a maximum queue size at which messages are queued at the respective resource controller, a maximum queue delay that messages queued at the respective resource controller experience, a maximum jitter experienced by messages queued at the respective resource controller (e.g. a maximum variation in resource response a message queued at the respective resource controller should experience) or a minimum level of service expected for a high priority message queued at the respective resource controller (e.g. a maximum delay, jitter or processing time for a high priority message). Applying by resource controller flow control in accordance with its respective flow control configuration may comprise checking to see if the respective traffic level trigger condition in the flow control configuration is met. If the respective traffic level trigger condition is not met in that traffic / messages are being forwarded to the resource controlled by the resource controlled at a suitable rate, then the flow control can comprise simply forwarding the messages received to the respective resource served by the resource controlled (e.g. the workflow module 120 served by the first resource controller 125, the generative machine learning model 130 served by the second resource controller 135 or the data store 150 served by the third resource controller 155). In contrast, if the respective traffic level trigger condition is met in that the number of messages in the queue is above the maximum queue size, the message delay is above the maximum queuedelay, the amount of jitter experienced by successive messages is above the maximum jitter or a high priority message is not experiencing a minimum level of service etc. then the method can comprise implementing further flow control as described below.

[0088] As mentioned above a flow control configuration can comprise a respective traffic level control rules. The respective traffic level control rules for a specific flow control configuration may comprise a respective batch size and respective machine learning application specific control rules. The machine learning application specific control rules can define the machine learning application specific flow control configurations. The implementing further flow control comprises storing messages in a queue of the respective resource controller until the number of messages in the queue reaches the respective batch size. The respective batch size can also be referred to as a microbatch size since the number of messages processed together can be smaller than in general batch processing. For example, the number of messages processed together can be around 10, or in other examples around 5, 15 or 20. Once the number of messages in the queue at the respective resource controller reaches the batch size, the resource controller processes the messages in the queue in accordance with the machine learning application specific control rules for the resource controller from the flow configuration.

[0089] In some examples, the machine learning applications in the machine learning environment are given service level classifications based on, for example, their priority level or on restrictions in speed or number of requests to a generative machine learning model the machine learning application should be able to perform. For example, a high priority or critical machine learning application can be placed in a service level classification indicating that messages related to that application should always be given highest priority and no artificial limits should be placed on access to any machine learning resources. In contract, a low priority application that may be used primarily for leisure or non-safety critical purposes may be placed in a service level classification that indicates messages related to that application are relatively low priority and access to machine learning resources for that application should be slowed or restricted in busy periods to ensure the low priority machine learning application does not consume valuable computing resources in a time when these resources are restricted. Messages related to a machine learning application can comprise a header or other metadata indicating either which machine learning application they relate to or the service level classification of the machine learning application that they relate to. Alternatively, the machine learning application a message relates to can be determined from a URL from which the message is sent or received.

[0090] The resource controller can determine a machine learning application that provided a message or is otherwise associated with the message based on a URL to which the message is directed or is from. Alternative the machine learning application associated with the message may be indicated in the message as metadata. The resource controller processing messages in the queue in accordance with the machine learning application specific control rules comprises the resource controller re-ordering messages in the queue, delaying at least some messages in thequeue, dropping at least some messages from the queue, or rerouting at least some messages in the queue in accordance with or in order to meet the machine learning application specific control rules. For example, this can involve reordering messages in the queue based on a priority for the messages established based on the service level classification of the machine learning application. In addition, this can involve delay some messages in the queue until a next cycle (e.g. a next batch) in response to a number of messages in the queue for an specific application with a low service level classification being greater than a number of allowed request indicated by the application specific control rules for that service level classification or in response to a rate of messages from that application being greater than a rate of messages allowed for low service level classification messages. In other examples, instead of delaying the messages, the messages can be dropped or routed to another machine learning resource that provides a lower service level.

[0091] In some examples, the flow control configuration for a specific resource controller can also comprise quotas for specific machine learning applications. These can be part of the machine learning application specific control rules and processed at that time or can be a separate form of flow control that is processed irrespective of a batch size being met. When quotas are used, applying flow control in accordance with a flow configuration comprises checking to see if a machine learning application has met its quota for use of the machine learning resource and if the machine learning application has met its quota dropping subsequent messages for the machine learning application and potentially sending an error to the machine learning application. In some examples the quotas can be for specific time periods e.g. for a day, week, month or year. The use of quotas enforced by resource controllers can allow an organization to easily limit the use of machine learning resources by specific users or machine learning applications. In addition, since the quotas are part of a flow configuration provided by a controller manager 140 which can easily be updated, quotas of this form allow easy updates to limitations on use of machine learning resources by specific machine learning applications.

[0092] It is noted that flow control can be applied to both incoming and outgoing messages of a resource to ensure that any priority and service level requirements are met in both directions if processing time at a resource impacts the order of messages returned from the resource.

[0093] While the above has focused on machine learning application specific control rules, the skilled person would understand that user specific control rules could also be used in the same way. In particular, a user associated with a message could be identified either from a URL of the message or metadata or a header of the message. Users could be given different service level classifications or access restrictions and applying flow control in accordance with a flow configuration could involve applying user specific flow control rules when reordering, delaying, dropping or rerouting messages in the fashion described above with respect to application specific control rules.

[0094] Figure 5 shows an example system 500 comprising a resource controller 515 and a resource 510. The resource controller 515 can comprise the first resource controller 125 and the resource 510 the workflow module 120, or the resource controller 515 can comprise second resource controller 135 and the resource 510 the generative machine learning model 130, or the resource controller 515 can comprise the third resource controller 155 while the resource can comprise the data store 550. The resource controller 515 is in communication with the controller manager 140. The controller manager 140 provides a flow configuration 517 to the resource controller 515 and in return the resource controller 5151 can provide logs or monitoring information to the controller manager 140.

[0095] The resource controller 515 manages incoming queue 520 which comprises messages to resource 510 and outcoming queue 530 which comprises messages from resource 510. Incoming messages can be received by a first message format adapter or protocol adapter 540before being placed in the incoming queue 520 and then processed by a second message format adapter or protocol adapter 550 before send to the resource 510. Similarly, outgoing messages can be received by the second HTTP request adaptor 550 before being placed in the outgoing queue 530 and then processed by the first HTTP request adaptor 540. In some examples, the first message format adapter or protocol adapter 540 can be a first HTTP (Hypertext Transfer Protocol) request adapter, a first SOAP (Simple Object Access Protocol) request adapter or first message middleware. Similarly, in some examples, the second message format adapter or a protocol adapter 550 can be a second HTTP (Hypertext Transfer Protocol) request adapter, a second SOAP (Simple Object Access Protocol) request adapter or second message middleware.

[0096] As mentioned above, the flow control configurations are provided to the resource controllers 125, 135, 155, 515 from a controller manager 140. This enables the controller manager 140 to centrally manage different machine learning applications and their consumption of machine learning resources. The resource controller 515 uses application specific control rules to reorder, drop, delay or reroute messages in the incoming queue 520 and outgoing queue 530 based on the priority of the machine learning application providing the messages to ensure all machine learning applications receive an acceptable minimal level of service.

[0097] As described in more detail below, when a new machine learning application is added to the machine learning environment, the controller manager 140 can update the flow control configurations to ensure adequate performance of all machine learning applications in the machine learning environment.

[0098] Figure 6 shows a method 600 for adding a new machine learning application to a machine learning environment.

[0099] In step 610, the new machine learning application is run in a test environment for example with a traffic generator to determine a minimum number of resources required by the new machine learning application. The performance of the new machine learning application could be split into three categories best possible performance, minimum acceptable performance and unacceptableperformance. The minimum number of resources required by the new machine learning application can be determined based on a minimum number of resources to achieve best possible performance and a minimum number of resources to achieve a minimum acceptable performance. In some examples, a maximum number of resources needed to achieve a best possible performance or a minimum acceptable performance is also determined. The performance can be an end-to-end response time for all resources in the machine learning environment used by the new machine learning application. These resources can include the workflow module 120, the generative machine learning model 130 and the data store 150. As part of determining the performance, maximum response times for each individual machine learning resource to meet the best possible performance and minimal acceptable performance limits can also be determined.

[0100] At step 620 details of the new machine learning application are added to a store containing details for the machine learning applications in the environment. The store contains maximum response times for best possible performance and minimal acceptable performance for each machine learning application in the environment. The maximum response times can be on both an end-to-end and a per resource basis. The store also contains data on limitations of the resources (e.g. the workflow module 120, the generative machine learning model 130 and the data store 150) to enable the performance of the resource to be determined e.g. the store contains a number of messages or amount of data the resource is abled to process. The performance of the resource varies depending on how it is managed by the flow configurations thus this information is also stored. In addition, the performance of the resource varies depending upon CPU size, number of CPU cores , memory type and size, disk type and size, for generative machine learning models the model, its version, model size and token, so this information can also be stored to enable to performance of the generative machine learning model to be estimated. Thus, the store stores all information to enable a performance of each resource in the machine learning environment to be estimated and also to estimate response times for the resource given different priority machine learning applications.

[0101] At step 630, a heuristic search algorithm is used based on the data in the store to determine flow configurations for each of the resources in the machine learning environment to ensure that the performance limit for each machine learning application in the machine learning environment including the new machine learning application are met. In some examples, the heuristic search algorithm is a genetic learning algorithm. The heuristic search algorithm is used to determine application specific flow control rules for each flow configuration that ensures the resource associated with the resource controller processes messages from each machine learning application in accordance with the maximum response time indicated for that machine learning application. As mentioned above, the application specific control rules can include a priority for a machine learning application, a limit on a number of messages or requests a machine learning application is able to send to a resource of the machine learning environment and / or arate limit (for example a rate per second limit) on the number of messages a particular machine learning application is able to send to a machine learning resource. These application specific control rules can be based on a service level classification for each machine learning application. The application specific control rules can be different for each resource as each resource controller has its own flow configuration and application specific control rules.

[0102] The heuristic search algorithm in step 630 is run for a limited number of iterations until a stop condition has been met. The stop condition can be a maximum number of iterations for the heuristic search algorithm or a maximum time the heuristic search algorithm should be run for. If once the stop condition has been met, the heuristic search algorithm has not found flow configurations that satisfy the minimum performance requirements for all machine learning applications, then, at step 640, a message is sent indicating the new machine learning application cannot be added. If a suitable solution is found by the heuristic search algorithm, then in step 650 the controller manager 140 updates the flow configurations in each of the resource controllers 125, 135, 155, 515 in accordance with the newly found flow configurations containing the newly found machine learning application specific control rules found by the heuristic algorithm.

[0103] Further details of the heuristic search algorithm, which may be a genetic learning algorithm, are now provided, still with reference to Figure 6. The heuristic search algorithm can be implemented by an optimizer 660.

[0104] At step 622, the heuristic search algorithm, genetic learning algorithm or optimizer sets initial values or solutions for the machine learning application specific control rules. In some examples these initial values can be randomly assigned. However, in other examples, the initial values may be based on current values for the application specific control rules for machine learning applications already in the machine learning environment and randomly assigned for the new machine learning application. Each resource controller has its own flow control and hence its own application specific control rules. Thus, the initial application specific control rules can be set differently for each flow control and hence each resource controller and its associated resource. As mentioned above, the application specific control rules can include a priority of each machine learning application, a limit of a maximum number of requests a machine learning application can provide in a particular time period e.g. a second and / or a limit on a total number of requests a particular machine learning application can provide or a total number of requests a resource can service in a particular time period e.g. one month. The machine learning application specific control rules can be determined based on a service level classification of the machine learning application. In some examples, as well as the machine learning application specific control rules being input and determined by the heuristic search algorithm or genetic learning algorithm, the batch size for each resource controller can also be input and determined by the heuristic search algorithm or genetic learning algorithm.

[0105] After step 622, an iterative process is performed for each of steps 624 to 626. At step 624, the optimiser applies a fitness test to the machine learning specific control rules fromeach flow configuration and based on a determined performance of the resources e.g. how many messages a second the resources can processes works out a time the resources will take to process the messages from each machine learning application. The fitness test thus determines if the maximum response time per resource and maximum end-to-end response time requirements for each machine learning application are met given the machine learning application specific control rules. In some examples, the maximum response time for best possible performance should be met. In other examples, the maximum response time for minimum acceptable performance should be met. In yet further examples, some machine learning applications may be required to meet the maximum response time for best possible performance while other machine learning applications may be required to meet the maximum response time for minimum performance acceptable performance.

[0106] If the fitness test is passed and all machine learning applications meet their performance requirements then the machine learning application specific control rules and corresponding flow configurations can be returned in step 630 and used to update the resource controllers in step 650.

[0107] On the other hand, if the fitness test is not passed, then at step 626 the optimizer, the heuristic search algorithm or the genetic learning algorithm can confirm if the stop condition has been met. If the stop condition has been met then the heuristic search or genetic learning algorithm stops and an error message indicating the new machine learning application cannot be added is provided in step 640. If the stop condition has not been met the heuristic search or genetic learning algorithm proceedings to step 628.

[0108] At step 628, the flow controls are considered separately. If the flow control for a resource controller of a resource meets the maximum response times for all machine learning applications using that resource, then the machine learning application specific control rules are kept for that flow control. The machine learning application specific control rules for the other flow configurations are mutated or perturbated. After this mutation or perturbation random mutations or perturbations are applied to any of the machine learning application specific control rules to increase the chances of a near optimal solution being found. These revised machine learning application specific control rules and flow configurations are then used for the next iteration where a fitness test is applied as detailed in step 624.

[0109] The above method thus enables the controller manager 140 to determine if a new machine learning application can be added to a machine learning environment without degrading performance of the machine learning applications in the environment below an acceptable level.

[0110] The examples described in this application provide ways to manage resources in a machine learning environment enabling efficient sharing of expensive machine learning resources such as data stores and generative machine learning model between multiple machine learning applications. The examples also enable an organization to better control how machinelearning resources are used to control and limit access to sensitive information and to ensure users are only provided with appropriate information.

[0111] Figure 7 illustrates various components of an example computing device 700 in which the resource controllers, the controller manager or even one of the resources such as the workflow module, the generative machine learning model or the data store can be implemented. The computing device is of any suitable form such as a smart phone, a desktop computer, a tablet computer, a laptop computer, or a server.

[0112] The computing device 700 comprises one or more processors 702 which are microprocessors, controllers or any other suitable type of processors for processing computer executable instructions to control the operation of the device in order to perform the methods implemented by the resource controllers or any other component of the machine learning environment. In some examples, for example where a system on a chip architecture is used, the processors 702 include one or more fixed function blocks (also referred to as accelerators) which implement a part of the method implemented by the resource controllers in hardware (rather than software or firmware). That is, the methods described herein are implemented in any one or more of software, firmware, hardware. The computing device has a data store holding instructions to implement the methods performed by the resource controllers or any other component of the machine learning environment. In some examples, the data store can be a data store 150 or the storage for the controller manager 140. In some examples, the data store has platform software comprising an operating system 704 or any other suitable platform software is provided at the computing-based device to enable application software 706 to be executed on the device. Although the computer storage media (memory 708) is shown within the computing-based device 700 it will be appreciated that the storage is, in some examples, distributed or located remotely and accessed via a network or other communication link (e.g. using communication interface 710). In addition, the communications interface 710 can be used to cause the resource controller, controller manager, or machine learning resource to communicate with other resources in the machine learning environment.

[0113] The computing-based device 700 may also comprises an input / output interface 712 arranged to output display information to a display device 714 which may be separate from or integral to the computing-based device 700. The display information may provide a graphical user interface. The input / output interface 712 is also arranged to receive and process input from one or more devices, such as a user input device 716 (e.g. a mouse, keyboard, camera, microphone or other sensor). In some examples the user input device 716 detects voice input, user gestures or other user actions. In an embodiment the display device 714 also acts as the user input device 716 if it is a touch sensitive display device. The input / output interrace 712 outputs data to devices other than the display device in some examples.

[0114] Any reference to 'an' item refers to one or more of those items. The term 'comprising' is used herein to mean including the method blocks or elements identified, but thatsuch blocks or elements do not comprise an exclusive list and an apparatus may contain additional blocks or elements and a method may contain additional operations or elements. Furthermore, the blocks, elements and operations are themselves not impliedly closed.

[0115] The steps of the methods described herein may be carried out in any suitable order, or simultaneously where appropriate. The arrows between boxes in the figures show one example sequence of method steps but are not intended to exclude other sequences or the performance of multiple steps in parallel. Additionally, individual blocks may be deleted from any of the methods without departing from the spirit and scope of the subject matter described herein. Aspects of any of the examples described above may be combined with aspects of any of the other examples described to form further examples without losing the effect sought. Where elements of the figures are shown connected by arrows, it will be appreciated that these arrows show just one example flow of communications (including data and control messages) between elements. The flow between elements may be in either direction or in both directions.

[0116] Where the description has explicitly disclosed in isolation some individual features, any apparent combination of two or more such features is considered also to be disclosed, to the extent that such features or combinations are apparent and capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein. In view of the foregoing description it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.

Claims

CLAIMS1. A system for resource management in a machine learning environment with one or more machine learning applications, comprising: a machine learning application API, application programming interface, to receive and respond to requests for using a generative machine learning model; a workflow module to process requests from the API, determine and execute a necessary workflow involving the generative model, and generate and return a response, wherein the generative machine learning model is arranged to generate content based on the operation of the workflow; first and second resource controllers managing traffic to the workflow module and the generative model, respectively, wherein the first and second resource controllers manage traffic using application-specific flow control configurations; and a controller manager to devise and supply the flow configurations to both resource controllers.

2. The system of claim 1 wherein: the machine learning application API is configured to receive a request to use the generative machine learning model in the machine learning environment from a user interface and to provide, to the user, a response to the request via the user interface; the workflow module is configured to: receive the request from the machine learning API; determine a workflow associated with the request, wherein the workflow comprises an operation to be performed by the generative machine learning model; send an indication of the operation to be performed by the generative machine learning model to the generative machine learning model; receive content from the generative machine learning model in response to the indication; generate the response to the request based on the content from the generative machine learning model; and send the response to the request to the machine learning application API; the generative machine learning model is configured to generate content in response to the indication of the operation to be performed provided by the workflow; the first resource controller is configured to intercept traffic going to and from the workflow module and configured to perform flow control operations on the intercepted traffic in accordance with a first flow configuration received from a controller manager, wherein the first flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications;the second resource controller is configured to intercept traffic going to and from the generative machine learning model and configured to perform flow control operations on the intercepted traffic in accordance with a second flow configuration received from the controller manager, wherein the second flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; and the controller manager is operable in communication with the first resource controller and the second resource controller and is configured to determine the first flow configuration and second flow configuration and provide the first flow configuration to the first resource controller and the second flow configuration to the second resource controller.

3. The system of claim 2, further comprising: a data store configured to store information as data vectors and provide returned data to the workflow module in response to requests for data; a third resource controller configured to intercept traffic going to and from the data store and configured to perform flow control operations on the intercepted traffic in accordance with a third flow configuration received from the controller manager, wherein the third flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; and wherein: the workflow determined by the workflow module further comprises a request to obtain data from a data store; the workflow module is further configured send the request to the data store to obtain the data from the data store and receive the returned data from the data store; the indication of the operation to be performed by the generative machine learning model further comprises the returned data; and the controller manager is further configured to determine the third flow configuration and provide the third flow configuration to the third resource controller.

4. The system of claim 2 or claim 3, wherein at least one of the first, second or third flow configuration for the respective resource controller comprises a respective traffic level trigger condition and respective traffic level flow control rules, wherein: the respective traffic level condition specifies a condition under which the respective traffic level flow control rules apply, and the respective traffic level flow control rules comprise: a respective batch size, wherein the respective batch size defines a number of messages that should be in a queue of the respective resource controller before the message are processed; andrespective machine learning application specific control rules that define the machine learning application specific flow control configurations, wherein the respective machine learning application specific control rules define a priority for the corresponding machine learning application and either a rate limit of messages or a quota limit of messages for the corresponding machine learning application and cause the respective resource controller to reorder the messages, delay at least some messages, drop at least some messages, or reroute at least some messages to meet the machine learning application specific control rules.

5. The system of any of claims 2 to 4 wherein the controller manager is further configured to: receive a request to add a new machine learning application to the one or more machine learning applications; run the new machine learning in a test environment to obtain a maximum allowable response time for the workflow module, the machine learning model, and the data store in order to run the new machine learning application; and use a heuristic search algorithm to determine the first flow configuration, the second flow configuration, and the third flow configuration.

6. The system of claim 5, wherein the controller manger being configured to use a heuristic search algorithm to determine the first flow configuration, the second flow configuration, and the third flow configuration comprises the controller manager being configured to: assign initial machine leaning application specific control rules for the first flow configuration, the second flow configuration, and the third flow configuration, to each machine learning application of the one or more machine learning applications and the new machine learning application; and iteratively: check whether the new machine learning application and the one or more machine learning applications are meeting their maximum response times for the workflow module, the machine learning model and the data store given the assigned machine learning application specific control rules; and in response to all maximum response times being met, return the application specific control rules for the first flow configuration, the second flow configuration, and the third flow configuration; else in response to the maximum response times not all being met, perturb the machine learning application specific control rules to form new assigned machine learning application specific control rules.

7. A computer-implemented method for managing resources in a machine learning environment running one or more machine learning applications, comprising: receiving and responding to a request for using a generative machine learning model via an application API, application programming interface, wherein the request is processed by a workflow module of the machine learning environment to determine and execute a necessary workflow involving the generative model, and generate and return a response, wherein the generative machine learning model is arranged to generate content based on the operation of the workflow; applying first and second resource controllers to manage traffic to the workflow module and the generative model, respectively, wherein the first and second resource controllers manage traffic using application-specific flow control configurations; and devising and supplying the flow configurations to both resource controllers.

8. The method of claim 7 wherein: the request is a request to use the generative machine learning model in the machine learning environment to generate content, and the request is received from a user interface module for a first machine learning application of the one or more machine learning applications in the machine learning environment; the method further comprising: sending, by the user interface module, a first message comprising the request to the workflow module of the machine learning environment; intercepting, by the first resource controller, the first message; applying, by the first resource controller, flow control to the first message, wherein the flow control is in accordance with a first flow configuration received from the controller manager, wherein the first flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; forwarding, by the first resource controller, the first message to the workflow module in accordance with the flow control; determining, by the workflow module, the workflow associated with the request from the first message, wherein the workflow comprises an operation to be performed by the generative machine learning model; sending, by the workflow module, to the generative machine learning model, a second message comprising an indication of the operation to be performed by the generative machine learning model; intercepting, by the first resource controller, the second message; applying, by the first resource controller, flow control to the second message, wherein the flow control is in accordance with the first flow configuration received from the controller manager;sending, by the first resource controller, the second message to the generative machine learning model in accordance with the flow control; intercepting, by the second resource controller associated with the generative machine learning model, the second message; applying, by the second resource controller, flow control to the second message, wherein the flow control is in accordance with a second flow configuration received from the controller manager, wherein the second flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; forwarding, by the second resource controller, the second message to the generative machine learning model in accordance with the flow control; generating, by the generative machine learning model, content in accordance with the indication of the operation in the second message; sending, by the generative machine learning model, to the workflow module, the content to the workflow module as a third message; intercepting, by the second resource controller, the third message sent to the workflow module; applying, by the second resource controller, flow control to the third message, wherein the flow control is in accordance with the second flow configuration received from the controller manager; forwarding, by the second resource controller, the third message to the workflow module in accordance with the flow control; intercepting, by the first resource controller, the third message; applying, by the first resource controller, flow control to the third message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; forwarding, by the first resource controller, the third message to the workflow module; generating, by the workflow module, a response to the request based on the content in the third message received from the generative machine learning model; sending, by the workflow module, the response to the request to the machine learning application API as a fourth message; intercepting, by the first resource controller, the fourth message; applying, by the first resource controller, flow control to the fourth message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; forwarding, by the first resource controller, the fourth message to the machine learning application API in accordance with the flow control; and providing, by the machine learning application API, the response to the request from the fourth message to the user.

9. The method of claim 8, wherein the first flow configuration comprises a first traffic level trigger condition, and first traffic level flow control rules, wherein the first traffic level control rules comprise a first batch size and first machine learning application specific flow control rules that define the machine learning application specific flow control configurations and applying flow control in accordance with the first flow configuration comprises: checking to see if the first traffic level trigger condition is met; and in response to the first traffic level trigger condition not being met, forwarding any messages received at the first resource controller to the workflow module; else in response to the first traffic level trigger condition being met: storing messages received at the first resource controller in queue until a number of messages in the queue is equal to the first batch size; and in response to the number of messages in the queue being equal to the first batch size, processing the messages in the queue in accordance with the first machine learning application specific control rules by reordering the messages, delaying at least some messages, dropping at least some messages, or rerouting at least some messages to ensure the first application specific control rules are met.

10. The method of claim 8 or claim 9, wherein the second flow configuration comprises a second traffic level trigger condition, and second traffic level flow control rules, wherein the second traffic level control rules comprise a second batch size and second machine learning application specific flow control rules that define the machine learning application specific flow control configurations, and applying flow control in accordance with the second flow configuration comprises: checking to see if the second traffic level trigger condition is met; and in response to the second traffic level trigger condition not being met, forwarding any messages received at the first resource controller to the machine learning model; else in response to the second traffic level trigger condition being met: storing messages received at the second resource controller in queue until a number of messages in the queue is equal to the second batch size; and in response to the number of messages in the queue being equal to the second batch size, processing the messages in the queue in accordance with the second machine learning application specific control rules by reordering the messages, delaying at least some messages, dropping at least some messages, or rerouting at least some messages to ensure the second application specific control rules are met.11 . The method of claim 10, wherein:the second routing configuration further comprises access control rules for the machine learning model, wherein the access control rules indicate whether particular users and / or machine learning applications are allowed access to the machine learning model; and applying flow control in accordance with the second flow configuration further comprises: determining if any messages received at the second resource controller meets the access control rules; and in response to the messages received at the second resource controller meeting the access control rules implementing the method of claim 10; else in response to a message received at the second resource controller not meeting the access control rules, rerouting the message to another machine learning model or workflow module.

12. The method of any of claims 9, 10 or 11 , wherein the first or second machine learning application specific control rules define a priority for the corresponding machine learning application and either a rate limit of messages or a quota limit of messages for the corresponding machine learning application.

13. The method of any of claims 8 to 12, wherein the workflow determined by the workflow module further comprises a request to obtain data from a data store of the machine learning environment, and the method further comprises: sending, by the workflow module, the request to obtain data to the data store as a fifth message; intercepting, by the first resource controller, the fifth message; applying, by the first resource controller, flow control to the fifth message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; forwarding, by the first resource controller, the fifth message in accordance with the flow control; intercepting, by a third resource controller associated with the data store, the fifth message; applying, by the third resource controller, flow control to the fifth message, wherein the flow control is in accordance with a third flow configuration received from the controller manager, wherein the third flow configuration specifies machine learning application specific flow control configurations for each machine learning application of the one or more machine learning applications; forwarding, by the third resource controller, the fifth message to the data store in accordance with the flow control; obtaining, from the data store, returned data in response to the request to obtain data from the fifth message;sending, by the data store, to the workflow module, the returned data as a sixth message; intercepting, by the third resource controller, the sixth message; applying, by the third resource controller, flow control to the sixth message, wherein the flow control is in accordance with the third flow configuration received from the controller manager; forwarding, by the third resource controller, the sixth message to the workflow module in accordance with the flow control; intercepting, by the first resource controller, the sixth message; applying, by the first resource controller, flow control to the sixth message, wherein the flow control is in accordance with the first flow configuration received from the controller manager; and forwarding, by the first resource controller, the sixth message to the workflow module in accordance with the flow control; wherein: the second message sent by the workflow module to the generative machine learning model further comprises the returned data from the data store.

14. The method of any of claims 8 to 13, further comprising, in response to a request to add a new machine learning application to the machine learning environment: running the new machine learning application in a test environment to determine a maximum response time for the workflow module and a maximum response time for the machine learning model required to implement the new machine learning application; providing, by the test environment, the determined maximum response time for the workflow module and the maximum response time for the machine learning model to a optimizer of the controller manager; using, by the controller manager, a heuristic search algorithm to determine the first flow configuration and the second flow configuration.

15. The method of claim 14 when dependent upon claim 13, the method further comprising: providing, by the test environment, the determined maximum response time for the data store to the optimizer of the controller manager; and wherein using, by the controller manager, the heuristic search algorithm to determine the first flow configuration and the second flow configuration further comprises using, by the controller manager, the heuristic search algorithm to determine the third flow configuration.

16. The method of claim 14 or claim 15, wherein using the heuristic search algorithm to determine the first flow configuration and the second flow configuration comprises: assigning machine leaning application specific control rules for the first flow configuration, the second flow configuration, and where a data store is used, the third flow configuration, to each machine learning application of the one or more machine learning applications and the newmachine learning application, wherein the application specific control rules specify for the respective flow configuration and corresponding machine learning application a priority for the machine learning application and either a rate limit of messages or a quota limit of messages for the machine learning; and iteratively: checking whether the new machine learning application and the one or more machine learning applications are meeting their maximum response times for the workflow module, the machine learning model and, where applicable, the data store; and in response to all maximum response times being met, returning the application specific control rules for the first flow configuration, the second flow configuration, and where applicable, the third flow configuration; else in response to the maximum response times not all being met, perturbating the machine learning application specific control rules.

17. The method of claim 16, wherein using the heuristic search algorithm to determine the first flow configuration, the second flow configuration and, where applicable, the third flow configuration further comprises, in response to the maximum response times not all being met: confirming whether a stop condition is met, where the stop condition specifies a maximum number of iterations to be performed; and in response to the stop condition being met returning an indication that it is not possible to add the new machine learning application; else in response to the stop condition not being met perturbating the machine learning application specific control rules.

18. A computer system including a processor and memory storing computer program code for performing the steps of the method of any of claims 7 to 17.

19. A computer program element comprising computer program code to, when loaded into a computer system and executed thereon, cause the computer to perform the steps of a method as claimed in any of claims 7 to 17.

Citation Information

Patent Citations

  • Intelligently managing resource utilization in desktop virtualization environments

    US20210385281A1

  • Multi-Tenant Control Plane Management on Computing Platform

    US20220188170A1