Method and system for executing generative artificial intelligence and fine tuning data models
By allocating dedicated AI clusters to generative AI models and monitoring GPU resources, the high demand for computing resources and reliability issues of generative AI models are addressed, achieving efficient resource management and reducing operating costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2026-04-14
AI Technical Summary
Generative AI models have high demands on computing resources and face challenges in the efficient management of GPU resources, especially during training and deployment, where resource waste and reliability issues exist.
By identifying the client ID associated with the client, allocating a dedicated AI cluster, using asymmetric key authentication requests, and monitoring the performance parameters of GPU resources, resource optimization and replacement are performed, and dummy operations are executed to improve resource utilization.
It achieves efficient allocation of computing resources for generative AI models, reduces resource waste, improves the reliability and response speed of GPU resources, and reduces operating costs.
Smart Images

Figure CN121866570A_ABST
Abstract
Description
[0001] Cross-referencing
[0002] This application claims priority to the following applications: U.S. Non-Provisional Application No. 18 / 676,248, filed May 28, 2024, entitled “Method and System for Performing Generative Artificial Intelligence and Fine Tuning The Data Model”; U.S. Non-Provisional Application No. 18 / 741,656, filed June 12, 2024, entitled “Resource Allocation for Accessing Cloud Based Services”; U.S. Non-Provisional Application No. 18 / 676,239, filed May 28, 2024, entitled “Method and System for Resource Optimization to Perform an Operation”; U.S. Provisional Application No. 63 / 583,167, filed September 15, 2023, entitled “Secure Gen-AI Platform Integration on a Cloud Service”; and U.S. Provisional Application No. 63 / 583,167, filed September 15, 2023, entitled “Method and System for Performing Generative Artificial Intelligence and Fine Tuning the Data Model”. This application relates to U.S. Provisional Application No. 63 / 583,169, entitled “Secure Generative-Artificial Intelligence Platform Integration on a Cloud Service,” filed May 28, 2024. Each of these applications is hereby incorporated herein by reference in its entirety for all purposes. Background Technology
[0003] Generative artificial intelligence (AI) is based on one or more models and / or algorithms configured to generate new content, such as new text, images, music, or videos. Generative AI models frequently receive complex cues (e.g., in natural language formats, audio / video files, images, etc.) and generate complex outputs. Each input cue and / or each output can be represented in a high-dimensional space, which may include one or more dimensions representing time, individual pixels, frequencies, higher-dimensional features, etc. Cue processing is often complex because evaluating a given part of the input query can be important, given another part of the input query. Furthermore, while many older machine learning models can generate outputs as simple as scores or classifications, the outputs from generative AI models are typically more complex and have much larger data sizes. To handle all the complexity of the cues and outputs, many generative AI models have millions or even billions of parameters. Therefore, efficient configuration of powerful computing resources is required to train and deploy generative AI models. Summary of the Invention
[0004] In one embodiment, a computer-implemented method includes receiving a request to allocate graphics processing unit (GPU) resources for performing an operation. The request includes metadata identifying a client identifier (ID) associated with a client, the latency of the operation, and throughput. Resource constraints for performing the operation are determined based on the metadata. Additionally, attributes associated with each of a plurality of GPU resources available for allocation are obtained. These attributes indicate the capacity of the corresponding GPU resource. The attributes are analyzed relative to resource constraints. A group of GPUs is identified from the plurality of GPU resources based on the analysis. A dedicated AI cluster is generated by grouping the group of GPUs into a single cluster. The dedicated AI cluster reserves a portion of the computing capacity of the computing system for a period of time, and the dedicated AI cluster is assigned to a client associated with the client ID.
[0005] Before allocating a dedicated AI cluster, requests are authenticated based on the client ID associated with the client. The private key extracted from the asymmetric key pair associated with the client ID is used for authentication. Additionally, a pre-approved quota associated with the request can be obtained. If the pre-approved quota exceeds a predefined request limit corresponding to the client ID, the request can be blocked. If the pre-approved quota is within the predefined request limit, the request can be forwarded for further processing. Furthermore, the type of operation can be determined based on the request. A set of GPU resources is selected from multiple nodes or one of a single node to generate a dedicated AI cluster. For example, if the request relates to fine-tuning a data model, the set of GPU resources is selected from a single node to form a dedicated AI cluster. In this case, the data model to be fine-tuned is obtained, and the fine-tuning logic is executed on the data model using the dedicated AI cluster.
[0006] In another embodiment, a computer-implemented method includes monitoring a set of performance parameters corresponding to each GPU resource in a first set of graphics processing unit (GPU) resources included in a dedicated AI cluster. The set of performance parameters includes at least one of the physical or logical conditions of the corresponding GPU resource. The set of performance parameters is compared with a predefined set of performance parameters. Based on the comparison, an anomaly is identified in a first GPU resource within the first set of GPU resources. The anomaly indicates that the set of performance parameters deviates from the predefined set of performance parameters. Additionally, in response to the anomaly identified in the first GPU resource, a second GPU resource is identified from a second set of GPU resources. The second GPU resource is identified by matching the computational capacity of the second GPU resource with the computational capacity of each GPU resource in the first set of GPU resources. The second set of GPU resources is reserved for replacement. When the second GPU resource is detected, the first GPU resource is released from the dedicated AI cluster, and the second GPU resource is patched into the dedicated AI cluster.
[0007] The physical condition of each GPU includes the temperature of the corresponding GPU resource, the clock cycles per core associated with the corresponding GPU resource, the internal memory of the corresponding GPU resource, and the power supply of the corresponding GPU resource. The logical condition of each GPU includes the failure of the plug-in associated with the corresponding GPU resource, the startup problem associated with the corresponding GPU resource, runtime failures, and security vulnerabilities. Anomalies in the first GPU are identified using a rotating hash value. A rotating hash value is calculated based on the compute image, compute shape, and resources of each GPU resource in the second set of GPU resources. The rotating hash value indicates the security and compliance status of the corresponding GPU resource. Patching the second GPU resource to the dedicated AI cluster is terminated based on predefined conditions. The predefined conditions are failure of the second GPU resource during startup, failure of the second GPU resource to join the dedicated AI cluster, workload failure of the second GPU resource, and software errors detected in the second GPU resource. In response to determining the predefined conditions, a tag is associated with the second set of GPU resources. The tag indicates the inapplicability of patching the second set of GPU resources.
[0008] In another embodiment, a computer-implemented method includes monitoring the computing capacity of each GPU resource in a set of graphics processing unit (GPU) resources reserved for a client. Based on the computing capacity, a first GPU resource from the set of GPU resources is identified that will be utilized to perform an operation associated with the client. Additionally, attributes of the operation are determined based on analysis of the inputs and outputs of the operation. A dummy operation is generated using the attributes of the operation performed on the first GPU resource, and the dummy operation is performed on a second GPU resource from the set of GPU resources that has not been utilized to perform the operation.
[0009] During the dummy operation on the second GPU resource, a request to access the second GPU resource to perform an actual operation can be received from the client. When a request to access one or more second GPU resources is received, the dummy operation from one or more second GPUs is terminated, and the actual operation is performed using the second GPU resource.
[0010] In some embodiments, a computer-implemented method is provided, comprising: receiving a request for allocating graphics processing unit (GPU) resources for performing an operation, wherein the request includes metadata identifying a client identifier (ID) associated with a client, a target latency of the operation, and a target throughput; determining resource constraints for performing the operation based on the metadata; obtaining at least one attribute associated with each of a plurality of GPU resources available for allocation in a computing system, wherein the at least one attribute indicates the capacity of the corresponding GPU resource; analyzing the at least one attribute associated with each GPU resource relative to resource constraints; identifying a set of GPU resources from the plurality of GPU resources based on the analysis; generating a dedicated AI cluster by patching the set of GPU resources into a single cluster, wherein the dedicated AI cluster retains a portion of the computing capacity of the computing system for a period of time; and assigning the dedicated AI cluster to a client associated with a client ID.
[0011] The disclosed method may include: authenticating a request based on a client ID associated with the client before allocating a dedicated AI cluster, wherein the request is authenticated using a private key extracted from an asymmetric key pair associated with the client ID.
[0012] The disclosed method may include: comparing a set of performance parameters corresponding to each GPU resource in the set of GPU resources with a predefined set of performance parameters; determining an anomaly in a first GPU resource in the set of GPU resources based on the comparison, wherein the anomaly indicates that the set of performance parameters deviates from the predefined set of performance parameters; and replacing the first GPU resource with a second GPU resource within a dedicated AI cluster, wherein the hash value of the second GPU resource is the same as the hash value of the first GPU resource.
[0013] The disclosed method may include: determining a pre-approved quota associated with the request; determining whether the pre-approved quota exceeds a predefined request limit corresponding to a client ID; and blocking the request based on the determination that the pre-approved quota exceeds the predefined request limit.
[0014] The disclosed method may include: determining the type of operation based on a request; and selecting the set of GPU resources from one of multiple nodes or a single node based on the type of operation to generate a dedicated AI cluster. Based on determining the request, a fine-tuning operation can be indicated: a data model to be fine-tuned can be obtained; and fine-tuning logic can be performed on the data model using the dedicated AI cluster, wherein the dedicated AI cluster is generated using the set of GPU resources selected from the single node.
[0015] The disclosed method may include: identifying at least one underutilized GPU resource from the set of GPU resources of a dedicated AI cluster; and, in response to the identification of the at least one GPU resource, performing a dummy operation on the at least one GPU resource, wherein the dummy operation is identical to an operation performed on the at least one GPU resource.
[0016] In some embodiments, a computer-implemented method is provided, comprising: monitoring a set of performance parameters corresponding to each GPU resource in a first set of graphics processing unit (GPU) resources included in a dedicated AI cluster, wherein the set of performance parameters includes at least one of a physical condition or a logical condition of the corresponding GPU resource; comparing the set of performance parameters corresponding to each GPU resource with a predefined set of performance parameters; determining an anomaly in a first GPU resource in the first set of GPU resources based on the comparison, wherein the anomaly indicates that the set of performance parameters deviates from the predefined set of performance parameters; in response to the anomaly determined in the first GPU resource, identifying a second GPU resource from the second set of GPU resources by matching the computing capacity of the second GPU resource with the computing capacity of each GPU resource in the first set of GPU resources, wherein the second set of GPU resources is reserved for replacement; releasing the first GPU resource from the dedicated AI cluster; and patching the second GPU resource to the dedicated AI cluster.
[0017] The physical condition of each GPU resource may include at least one of the following: the temperature of the corresponding GPU resource, the clock cycle of each core associated with the corresponding GPU resource, the internal memory of the corresponding GPU resource, and the power supply of the corresponding GPU resource; and the logical condition of each GPU resource may include at least one of the following: a failure of a plugin associated with the corresponding GPU resource, a startup problem associated with the corresponding GPU resource, a runtime failure, and a security vulnerability.
[0018] The disclosed method may include: calculating a rotation hash value for each GPU resource in the second group of GPU resources based on a computational image, a computational shape, and at least one of the resources of each GPU resource in the second group of GPU resources; comparing the rotation hash value of each GPU resource in the second group of GPU resources with the rotation hash value of a dedicated AI cluster; and identifying the second GPU resource based on the comparison. The rotation hash value may indicate the security and compliance status of the corresponding GPU resource.
[0019] The disclosed method may include terminating the patching of a second GPU resource to a dedicated AI cluster based on predefined conditions, wherein the predefined conditions are one of the following: a failure of the second GPU resource during startup; the second GPU resource failing to join the dedicated AI cluster; a workload failure of the second GPU resource; and a software error detected in the second GPU resource.
[0020] The disclosed method may include: in response to determining the predefined conditions, associating a tag with a second set of GPU resources, wherein the tag indicates inapplicability in patching of the second set of GPU resources.
[0021] Each GPU resource in the second group can be configured to continuously perform dummy operations, which are exactly the same as the actual operations performed by each GPU resource in the first group.
[0022] In some embodiments, a computer-implemented method is provided, comprising: monitoring the computing capacity of each GPU resource in a set of graphics processing unit (GPU) resources reserved for a client; identifying, based on the computing capacity, one or more first GPU resources from the set of GPU resources that are utilized to perform an operation associated with the client; determining attributes of the operation based on analysis of the inputs and outputs of the operation; generating a dummy operation using the attributes of the operation performed on the one or more first GPU resources; and performing the dummy operation on one or more second GPU resources in the set of GPU resources that are not utilized to perform the operation.
[0023] The disclosed method may include: receiving a request to access the one or more second GPU resources to perform an operation; terminating a dummy operation from the one or more second GPU resources in response to the request to access the one or more second GPU resources; and performing the operation using the one or more second GPU resources.
[0024] The disclosed method may include: determining an anomaly in one or more first GPU resources based on a deviation of a set of performance parameters of a first GPU resource from a predefined set of performance parameters; identifying a second GPU resource from the one or more second GPU resources in response to the anomaly determined in the first GPU resource by matching the computing capacity of the second GPU resource with the computing capacity of each of the one or more first GPU resources; and performing an operation on the second GPU resource.
[0025] The set of performance parameters may include at least one of the physical or logical conditions of the corresponding GPU resources.
[0026] The physical condition of each GPU resource may include at least one of the following: the temperature of the corresponding GPU resource, the clock cycle of each core associated with the corresponding GPU resource, the internal memory of the corresponding GPU resource, and the power supply of the corresponding GPU resource; and the logical condition of each GPU resource may include at least one of the following: a failure of a plugin associated with the corresponding GPU resource, a startup problem associated with the corresponding GPU resource, a runtime failure, and a security vulnerability.
[0027] The disclosed method may include: determining a pre-approved quota associated with a request received from a client; determining whether the pre-approved quota exceeds a predefined request limit corresponding to the client; and blocking the request based on the determination that the pre-approved quota exceeds the predefined request limit.
[0028] The disclosed methods may include authenticating requests using a private key extracted from an asymmetric key pair associated with the client.
[0029] In various aspects, a system is provided, including one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the methods disclosed herein.
[0030] In various respects, a computer program product is provided, tangibly implemented in a non-transitory machine-readable storage medium and including instructions configured to cause one or more data processors to perform some or all of the methods disclosed herein.
[0031] The techniques described above and below can be implemented in a variety of ways and in a variety of contexts. Several example implementations and contexts are provided with reference to the following figures, which are described in more detail below. However, the following implementations and contexts are only a few of the many implementations and contexts. Attached Figure Description
[0032] This disclosure is described in conjunction with the accompanying drawings:
[0033] Figure 1 The diagram illustrates a block diagram of a system for providing and supporting a generative artificial intelligence (AI) platform, according to an exemplary embodiment.
[0034] Figure 2 The illustration shows an exemplary architecture of the data plane of a generative AI platform according to an exemplary embodiment.
[0035] Figure 3 The diagram illustrates a block diagram of an API server according to an exemplary embodiment.
[0036] Figure 4 The figure illustrates a data flow diagram of a rate limiter according to an exemplary embodiment.
[0037] Figure 5 The illustration shows a block diagram of a metering worker according to an exemplary embodiment.
[0038] Figure 6 The illustration shows a high-level design diagram for creating a dedicated AI cluster (DAC) according to an exemplary embodiment.
[0039] Figure 7 The diagram illustrates a block diagram of a system configured to create a dedicated AI according to an exemplary embodiment.
[0040] Figure 8 The illustration shows a flow diagram of a control plane request being propagated to the management plane and subsequently to the data plane of a generative AI platform, according to an exemplary embodiment.
[0041] Figure 9A and Figure 9B The illustration shows a sequence diagram of the process of creating a DAC according to an exemplary embodiment.
[0042] Figure 10 The illustration shows a flow diagram indicating the work of a DAC operator according to an exemplary embodiment.
[0043] Figure 11 The illustration shows a flow diagram of the processing of a model operator according to an exemplary embodiment.
[0044] Figure 12 The illustration shows a sequence diagram of the processes for creating a base model and fine-tuned model resources according to an exemplary embodiment.
[0045] Figure 13 The illustration shows a sequence diagram of a method for fine-tuning a base AI model according to an exemplary embodiment.
[0046] Figure 14The diagram illustrates a data flow diagram for fine-tuning a data model according to an exemplary embodiment.
[0047] Figure 15A and Figure 15B The illustration shows a sequence diagram of the process of creating a DAC according to an exemplary embodiment.
[0048] Figure 16 The diagram illustrates a data flow diagram for managing a fine-tuned inference server using an endpoint operator, according to an exemplary embodiment.
[0049] Figure 17 The illustration shows a sequence diagram of logical processing within an inference server according to an exemplary embodiment.
[0050] Figure 18 The diagram illustrates a data flow diagram of the work of an operator at an instruction model endpoint according to an exemplary embodiment.
[0051] Figure 19 The illustration shows a sequence diagram for creating a basic inference service, creating a fine-tuned inference service, and deleting a fine-tuned inference service, according to an exemplary embodiment.
[0052] Figure 20 The illustration shows a flowchart of a process for allocating a dedicated AI cluster according to an exemplary embodiment.
[0053] Figure 21 The illustration shows a flowchart of a fault management process according to an exemplary embodiment.
[0054] Figure 22 The illustration shows a flowchart of a process for managing operations performed using GPU resources, according to an exemplary embodiment. Detailed Implementation
[0055] In the following description, specific details are set forth for purposes of explanation to provide a thorough understanding of certain embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The accompanying drawings and description are not intended to be limiting. The word “exemplary” is used herein to mean “serves as an example, instance, or illustration.” Any embodiment or design described herein as “exemplary” is not necessarily to be construed as being more preferred or advantageous than other embodiments or designs.
[0056] As mentioned above, generative AI models are typically very large and require enormous computational resources for training and deployment. This is especially true given that many deployments of such models occur in environments where users expect immediate output in response to prompts.
[0057] While general-purpose and capable of handling a wide range of tasks, CPUs lack the parallel processing power to efficiently train and deploy such large and complex generative AI models. In contrast, GPUs (Graphics Processing Units) excel at parallel computing. These hardware accelerators significantly speed up training and inference processes, enabling faster experimentation, deployment, and real-time application of generative AI across diverse domains.
[0058] However, the use of GPUs presents several challenges. For example, the purchase and maintenance of GPUs, especially high-end models optimized for deep learning tasks, can be expensive. Furthermore, GPUs are power-intensive devices, consuming significant amounts of electricity during training and inference. This can lead to high operating costs, particularly for large-scale deployments using multiple GPUs simultaneously. Therefore, efficient use of GPU resources is a priority. However, some entities prioritize ensuring reliable and rapid availability of GPUs and GPU processing to support real-time or near real-time responses to prompts. Consequently, cloud providers serving requests from multiple entities have competing priorities: using GPUs most efficiently versus ensuring GPU processing is performed in a manner that guarantees GPU availability.
[0059] Certain aspects and features of this disclosure relate to a technique for allocating dedicated GPU resources to various clients. Task requests are then routed in a manner that ensures any dedicated GPU resource is reserved only for the client to which it was allocated. If a client's tasks are numerous or complex enough to fully consume GPU resources, this may result in the GPU resources being idle for a period of time. However, any such GPU resource assigned to a client can then be assigned as a standby GPU and perform a dummy operation for a given task, which replicates the operation being performed by another GPU resource assigned to the client. Performance variables for each GPU resource can be monitored, and predefined conditions can be evaluated based on one or more performance variables. For example, ping response and / or operation completion latency can be monitored and compared to corresponding thresholds. When it is determined that predefined conditions are not met, a standby GPU can be assigned to handle the given task. Instead of subsequently needing to initiate the full execution of the task from that point, results can be generated and / or returned relatively promptly using the previously initiated dummy operation.
[0060] In the example, a client might request to train or fine-tune a data model implemented on a computing system. Such operations can involve complex computations and may require significant throughput and latency. This request is received from the client through the client system. Each client and / or each client system is associated with a unique client identity (ID). The request includes metadata identifying the client ID.
[0061] The computing system identifies a set of GPU resources available for assignment from its total GPU resources. This set of GPU resources is then grouped together to form a single cluster. This cluster is also assigned to clients and / or client systems associated with those clients. Assignment is performed using the client ID associated with the client and / or client system. This type of cluster is also known as a dedicated AI cluster. A dedicated AI cluster reserves a portion of the computing system's computing capacity for a period of time requested by the client.
[0062] Once a dedicated AI cluster is assigned to a client and / or a client system associated with that client, the operations requested by the client begin to be executed using the GPU resources patched into the dedicated AI cluster. Assigning a dedicated AI cluster to a client ensures that the workload associated with the operations requested by one client is not mixed with or matched with the workload associated with the operations requested by another client. Therefore, the computing system is able to provide computational capacity for operations that require significant computing power, such as training or fine-tuning data models.
[0063] In some embodiments, GPU resources may suffer interruptions and / or failures during operation. Such failures may correspond to plug-in failures, provisioning failures, and / or response failures. In a plug-in failure, the GPU resource fails to start properly. In a provisioning failure, the GPU resource fails to provision while performing an operation. In a response failure, the GPU resource fails to respond at runtime. Certain aspects and features of this disclosure provide techniques for overcoming the scenarios mentioned above. A computing system monitors a set of performance parameters corresponding to each GPU resource in a set of GPU resources included in a dedicated AI cluster. This set of performance parameters includes the physical and / or logical condition of the corresponding GPU resource. In embodiments, the physical condition of each GPU resource includes the temperature of the corresponding GPU resource, the clock cycles of each core associated with the corresponding GPU resource, the internal memory of the corresponding GPU resource, and / or the power supply of the corresponding GPU resource. The logical condition of each GPU resource includes plug-in failures associated with the corresponding GPU resource, startup problems associated with the corresponding GPU resource, runtime failures, and / or security vulnerabilities.
[0064] A set of performance parameters corresponding to each GPU resource is compared with a predefined set of performance parameters. This predefined set of performance parameters indicates the performance parameters of an ideal GPU resource. Based on the comparison, anomalies in the GPU resources within this set can be identified. These anomalies indicate that the set of performance parameters deviates from the predefined set of performance parameters. In some embodiments, the predefined performance parameters can be a range of values within which GPU resources are considered fault-free. When an anomaly is detected in a GPU resource, another GPU resource is identified from the remaining available GPU resources in the computing system. Replaceable GPU resources are identified by matching the computational capacity of the replaceable GPU resource with the computational capacity of each GPU resource in the dedicated AI cluster. After identification, the faulty GPU resource is released from the dedicated AI cluster, and the replaceable GPU resource is patched into the dedicated AI cluster. In this way, a set of replaceable GPU resources can be reserved to replace faulty GPU resources in the dedicated AI cluster.
[0065] When predefined conditions are identified, the process of patching replaceable GPU resources into a dedicated AI cluster can be terminated. These predefined conditions include: failure of the replaceable GPU resource during startup, failure of the replaceable GPU resource to join the dedicated AI cluster, workload failure of the replaceable GPU resource, and / or software errors detected in the replaceable GPU resource. Once the predefined conditions are identified, a tag can be associated with the replaceable GPU resource. This tag indicates the inapplicability of patching that set of replaceable GPU resources.
[0066] In some embodiments, the computing system may monitor the computational capacity of each GPU resource in a group of GPU resources. This group of GPU resources is reserved for the client for a period of time. During this period, the operation may not consume all the computational capacity reserved for the client. In this case, some GPU resources in the group are not utilized to execute the operation requested by the client. The computing system identifies the unutilized GPU resources based on computational capacity. Additionally, the attributes of the operation are determined based on analysis of the operation's inputs and outputs. A dummy operation is generated using the operation's attributes. In some embodiments, the dummy operation may be the same operation as the actual operation performed by the client (e.g., identical). For example, the dummy operation may process the same data as the actual operation, and the way the data is processed across the dummy operation and the actual operation may be the same. The dummy operation is executed on the unutilized GPU resources.
[0067] When the computing system receives a request to access GPU resources to perform an operation, it terminates the virtual operation and proceeds to execute the actual operation requested by the client. This reduces the time required to load GPU resources from idle states, thus improving the overall processing latency of the operation.
[0068] Figure 1 The illustration shows a block diagram of a computing system 100 for providing and supporting a generative artificial intelligence (AI) platform 102 according to an exemplary embodiment. The computing system 100 supports providing generative AI in response to requests received from clients. The generative AI platform 102 is communicatively coupled to a client system 104 via a network 106. It should be noted that a single client system 104 (e.g., shown as client system 104) is for illustrative purposes, and the scope of this disclosure is not limited thereto. The computing system 100 may host several client systems 104 sequentially, alternately, or in parallel. Each client system 104 is associated with a corresponding client 108.
[0069] Network 106 may include suitable logic, circuitry, and interfaces configured to provide multiple network ports and communication channels for transmitting and receiving data and / or instructions related to the operation of generative AI platform 102 and client system 104. In various embodiments, network 106 may include a wireless local area network (WAN), a local area network (LAN), and / or the Internet. Some computing systems 100 may have multiple hardware stations distributed across warehouses or factory production lines, connected via network 106. Distributed factories located in different locations and bound together by network 106 may exist.
[0070] Client system 104 may include various types of computing systems, such as personal assistant (PA) devices, portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computing devices can run various types and versions of software applications and operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX® or UNIX-like operating systems, Linux or Linux-like operating systems, such as Google Chrome™ OS), including various mobile operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android™, BlackBerry®, Palm OS®). Portable handheld devices may include cellular phones, smartphones (e.g., iPhone®), tablets (e.g., iPad®), personal digital assistants (PDAs), etc. Wearable devices may include Google Glass® head-mounted displays and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices (e.g., Microsoft Xbox® consoles with or without Kinect® gesture input, Sony PlayStation® systems, various gaming systems offered by Nintendo®, etc.). Client systems 104 and 108 can run various applications, such as various Internet-related apps and communication applications (e.g., email applications, short message service (SMS) applications), and can use various communication protocols.
[0071] Generative AI platform 102 includes an operator 110 that controls the modules of generative AI platform 102. Operator 110 manages the entire lifecycle of resources such as dedicated AI clusters, fine-tuning jobs, and / or service models. Operator 110 utilizes Kubernetes to perform various operations. In some embodiments, operator 110 includes a dedicated AI cluster (DAC) operator, a model operator, a model endpoint operator, and a machine learning (ML) job operator.
[0072] DAC operators leverage Kubernetes to reserve capacity within the generative AI data plane. Generative AI is responsible for reserving a specific number of GPU resources for clients. DAC operators manage the DAC lifecycle, including capacity allocation, monitoring, and patching. Model operators handle the model lifecycle, including both base and fine-tuned models. Model endpoint operators manage the provisioning and hosting of the lifecycle (storage, access, encryption) of pre-trained or fine-tuned models. ML job operators provide services for orchestrating the execution of long-running processes or workflows using ML models.
[0073] The generative AI platform 102 also includes a storage device 112 for temporarily or permanently storing data used to perform various operations. For example, the storage device 112 stores instructions that can be executed by an operator to perform operations. In addition, the storage device 112 stores training and testing data for ML models.
[0074] Generative AI platform 102 includes several modules for performing individual tasks. In some embodiments, generative AI platform 102 may include allocation module 114 for allocating / reserving GPU resources for a specific client in response to a request received from a client system 104 via network 106. Allocation module 114 is responsible for managing uptime and patching nodes to form a DAC (Dedicated Computing Capacity). DAC is a certain amount of computing capacity reserved for an extended period of time (e.g., at least one month). Allocation module 114 can then provide and maintain computing capacity for clients that require predictable performance and throughput to perform their operations. The functionality of allocation module 114 is described in detail in subsequent paragraphs.
[0075] The generative AI platform 102 also includes a repair module 116 for repairing the DAC when an anomaly is detected in the DAC's GPU resources. The DAC is repaired by identifying new GPU resources compatible with it. Alternatively, the defective GPU resources can be replaced with new GPU resources by releasing them and patching them into the DAC. The functionality of the repair module 116 is described in detail in subsequent paragraphs.
[0076] The generative AI platform 102 also includes a dummy operation module 118 for performing dummy operations on GPU resources not utilized by the client. To this end, GPU resources reserved for the client but not used are identified. Furthermore, the dummy operation module 118 performs dummy operations on the identified GPU resources. When the generative AI platform 102 receives a request to access the identified GPU resources to perform an actual task, the dummy operation module 118 terminates the dummy operations on the identified GPU resources and makes the GPU resources available for performing the actual task. In this way, the cold start of GPU resources for performing the actual task is mitigated. Therefore, the latency of performing the actual task is reduced. The functionality of the dummy operation module 118 is described in detail in subsequent paragraphs.
[0077] Figure 2 An exemplary architecture 200 of the data plane of a generative AI platform 102 according to an exemplary embodiment is illustrated. Figure 2 As shown, the generative AI data plane 202 receives service requests from a client system 104 associated with client 108. Service requests can be requests to perform operations and / or may include requests to perform operations. For example, client 108 may host a website configured to (e.g., by selecting one or more options or providing text input) receive input requesting specific information. Service requests can be defined as including input or a transformed version thereof (e.g., introducing some structure). Requests can be transmitted over a network (e.g., the Internet) and / or a load balancer as a service (LBaaS). The generative AI data plane 202 includes a CPU node pool 204 with multiple API servers 206. Each API server can receive a corresponding request and can forward the request to a GPU node pool 208. Each component of the API server 206 is described in further detail in subsequent paragraphs.
[0078] GPU node pool 208 includes fine-tuned inference servers 210 for transforming service requests to generate one or more prompts that can be executed by one or more data models, such as partner models or open-source models. Generative AI data plane 202 can be connected to streaming module 212, object storage module 214, file storage module 216, and / or identity and access management (IAM) module 218. Each module (including streaming module 212, object storage module 214, file storage module 216, and IAM module 218) can be configured to perform configured processing. For example, streaming module 212 can be configured to support metering and / or billing, object storage module 214 can be configured to store instances of executable programs, file storage module 216 can be configured to store ML models and other supporting applications, and IAM module 218 can be configured to manage client authentication and authorization.
[0079] In one implementation, the generative AI platform 102 provides various services to the client. Some of these services are provided via Table 1.
[0080] Table 1
[0081] Figure 3 A block diagram 300 of an API server 206 according to an exemplary embodiment is illustrated. The API server 206 is capable of receiving and managing service requests from clients. Service requests may be requests to generate text or fine-tune AI models. The API server 206 utilizes various components to process service requests, embedding service requests into inference components and returning responses to service requests to clients.
[0082] API server 206 is integrated with identity service module 302 to perform authentication of service requests. Identity service module 302 extracts a client identifier (ID) from the service request and authenticates the service request based on the client ID. In some embodiments, a private key extracted from an asymmetric key pair associated with the client ID is used to authenticate the request.
[0083] Following authentication, service requests are provided to rate limiter 304. Rate limiter 304 tracks each service request received from client 108 from client system 104. Rate limiter 304 integrates with limiting service 306 to obtain predefined limits, such as the number of requests per minute (RPM) allowed for each client system 104 and a pre-approved quota. The pre-approved quota (input / output) varies from request to request. For example, a service request may consume a predefined pre-approved quota, such as between 10 and 2048 tokens for small and medium-sized LLMs, while larger and more powerful LLMs are allowed up to 4096 tokens.
[0084] Rate limiter 304 extracts the number of input tokens associated with a service request and a pre-approved quota for the output. Rate limiter 304 uses predefined limits obtained from the limiting service 306 to limit the rate of incoming service requests from client system 104. Rate limiting can be implemented as RPM-based rate limiting or token-based rate limiting. Rate limiter 304 can allow a predefined number of tokens per minute per model for each client system 104. In yet another implementation, the limit can be imposed based on the lease ID or client ID associated with client system 104 as a key. The functionality of rate limiter 304 is related to... Figure 4 Detailed explanation.
[0085] API server 206 also includes a content moderator 308 for filtering content included in service requests by removing sensitive or toxic information from the service requests. In this implementation, the content moderator 308 filters training data prior to the training and fine-tuning of the LLM. In this way, the LLM response does not provide responses about sensitive or toxic information, such as how to commit crimes or engage in illegal activities, and the model response is monitored by filtering or stopping response generation if the results contain unwanted content.
[0086] API server 206 also includes a model meta-repository 310. Model meta-repository 310 can be a storage unit configured to store LLM-related metadata, such as LLM capabilities, display names, and creation times. Model meta-repository 310 can be implemented in two ways. The first implementation is an in-memory model meta-repository, where metadata is stored in an in-memory cache. Model meta-repository 310 is populated using one or more resource files that have undergone code review and are checked into the GenAI API repository. During the deployment of API server 206, model-related metadata is loaded into storage. The second implementation uses a persistent model meta-repository. Persistent model meta-repository enables custom model training or fine-tuning. In persistent model meta-repository, data related to (pre-trained and custom-trained) LLM models is stored in a database. API server 206 can query model meta-repository 310 for LLM-related metadata. Model meta-repository 310 can provide LLM-related metadata in response to the query.
[0087] API server 206 also includes a metering worker 312 for scheduling service requests received from client system 104. Metering worker 312 is communicatively coupled to billing server 314. Metering worker 312 processes service requests and communicates with billing server 314 to generate invoices for each process performed for the service request. API server 206 is communicatively connected to streaming service 316 for providing a lightweight streaming solution to stream tokens back to client system 104 whenever a token is generated.
[0088] API server 206 is communicatively connected to Prometheus T2 unit 318. Prometheus T2 unit 318 is a monitoring and alerting system developed within the infrastructure. Prometheus T2 unit 318 enables API server 206 to scrape metrics from various sources at regular intervals. By leveraging Prometheus T2, API server 206 gains the ability to collect a wide range of metrics, such as GPU utilization, latency, CPU utilization, LLM performance, etc. These metrics can then be used to create alerts and visualizations in Grafana, providing valuable insights into the performance of generative AI services associated with LLM models, the health of LLM models, and the resource utilization of LLM models.
[0089] The generative AI service utilizes a logging architecture that combines FluentBit and Fluentd. FluentBit is deployed as a set of daemons running on each GPU resource and acts as a log forwarder. It efficiently collects logs from various GPU resources within API server 206 and sends them to Fluentd. Fluentd, acting as an aggregator, receives logs from FluentBit and performs further processing or filtering as needed. Logs are periodically flushed to the Lumberjack 320 log transport protocol at predetermined intervals. This architecture enables efficient log collection, aggregation, and transmission, ensuring comprehensive and reliable logging for the generative AI service.
[0090] Figure 4 The diagram 400 illustrates a data flow diagram of a rate limiter 304 according to an exemplary embodiment. The rate limiter 304 can utilize platform throttling, a proven and widely adopted distributed rate limiting solution. This solution is based on a peer-to-peer protocol. The API server 206 periodically broadcasts its local traffic history, stored in its local cache, to its peers using the Google Remote Procedure Call (gRPC) protocol. When a service request is received, the rate limiter 304 in the API server 206 makes a local decision based on its local cache (a representation of the global traffic history).
[0091] At box 402, the service request is received by API server 206 from client system 104. This request may be received in the local cache. At box 404, rate limiter 304 associated with API server 206 determines whether the service request exceeds any throttling policy, such as the number of input tokens received from the client before the current time. If the service request does not exceed a predefined limit, the local cache can be updated at box 406 using the number of input tokens. If the service request exceeds the predefined limit, a response such as "429" "Too many requests returned to the client" can be returned.
[0092] Additionally, rate limiter 304 can access the LLM model and identify the number of output tokens. In streaming scenarios, rate limiter 304 can accumulate output tokens as they are streamed. Rate limiter 304 can use the number of output tokens to update its local cache. If the number of response tokens exceeds the token limit at that time, the service can refrain from throttling service requests to maintain a good client experience, given that the client has already waited some time and all work has been completed.
[0093] Figure 5 A block diagram 500 of a metering worker 312 according to an exemplary embodiment is illustrated. API server 206 may include a metering buffer 502 for temporarily storing service requests received from client system 104. Service requests may be sequentially transmitted to metering publisher 504. Metering publisher 504 may send service requests to metering worker 312, where each request is scheduled for processing. Billing server 314 may generate a bill for each processing performed for the service requests.
[0094] Figure 6 The illustration shows a high-level design diagram 600 for creating a dedicated AI cluster (DAC) according to an exemplary embodiment. The generative AI platform receives requests to allocate GPU resources for performing operations or requests to change the current allocation of GPU resources.
[0095] At box 602, a request is received to allocate GPU resources or change the current allocation of GPU resources. The change in current allocation may correspond to a request to increase or decrease the GPU resources currently utilized by the client. In some embodiments, the generative AI platform may generate a request to increase or decrease the GPU resources allocated to the client in response to an increase / decrease in traffic managed by client 108.
[0096] Requests to change the current allocation of GPU resources are received by administrator 604, who can allow or deny the request using the restriction service module 606. This ensures manual monitoring of GPU resource allocation before allowing clients to consume scarce and expensive GPU resources. It also provides a communication channel with the client in cases where GPU resources are unavailable in the client's preferred region(s). In this situation, administrator 604 can contact client 108 to suggest alternative regions(s) or provide a timeline for changing the current allocation via a restriction request support ticket.
[0097] The request can correspond to actions such as: creating a dedicated AI cluster (DAC) for fine-tuning at box 608, creating a fine-tuning model on the DAC at box 610, creating a DAC for hosting a model that provides generative AI services to clients at box 612, creating a model endpoint for interacting with the model hosting the generative AI services at box 614, and creating an interface with the model endpoint at box 616.
[0098] Blocks 608-616 are supported via communication with the Generative AI Control Plane (GenAI CP) 618 and the Generative AI Data Plane (GenAI DP) 620.
[0099] The GenAI DP 620 includes multiple Kubernetes operators, such as a Dedicated AI Cluster (DAC) operator 622, a model operator 624, an ML job operator 626, and a model endpoint operator 628. The DAC operator 622 performs functions such as creating multiple DACs by reserving a set of GPU resources corresponding to each DAC, and monitoring for physical and logical failures of each GPU resource in each of the multiple DACs. Physical failures of GPU resources correspond to overheating of GPU resources and other physical parameters associated with GPU resources, such as the clock cycles of the core associated with the GPU resource, the internal memory of the GPU resource, and the power supply of the corresponding GPU resource, which are derived by comparing various physical parameters of the GPU resource with a predefined set of physical parameters. Similarly, the DAC operator 622 monitors for logical failures of GPU resources. Logical failures involve vulnerabilities in security protocols associated with GPU resources, failures in plugins associated with GPU resources, and / or failures in GPU startup and runtime failures of GPU resources.
[0100] The model operator 624, communicatively coupled to the ML job operator 626, performs functions such as managing the lifecycle of various underlying models associated with the generative AI platform. The model operator 624 also performs functions such as managing the lifecycle of fine-tuned models and managing training and artifacts used for fine-tuning the underlying models.
[0101] ML job operator 626 is invoked by model operator 624 to perform training for creating a fine-tuned model and manage the training workload. ML job operator 626 also implements logic for managing the lifecycle of workflow resources, including task and host failure recovery. ML job operator 626 also provides a simplified and auditable surface area for access permissions. ML job operator 626 has a limited, well-defined scope and access control. Resources managed by ML jobs are visible throughout the AI Platform (AIP) ecosystem (console, logs, ML Ops tools).
[0102] Model endpoint operator 628 creates a model endpoint through communication between GenAI CP 618 and GenAI DP 620. At box 630, the generative AI platform provides a DAC for hosting operations associated with requests received from client 108.
[0103] Figure 7 The diagram illustrates a block diagram of a system 700 configured to create specialized AI according to an exemplary embodiment. System 700 includes a GenAI CP 618. The GenAI CP 618 may have a control plane API (CP API) server 704, a control plane (CP) workflow worker 706, and a management plane API (MP API) server 708. The CP API server 704 serves as the entry point for the GenAI CP 618. The CP API server 704 receives requests from clients via the public internet. The CP API server 704 is a Dropwizard application based on templates generated by Pegasus. The CP API server 704 provides a Workflow-as-a-Service (WFaaS) in response to requests. These requests correspond to operations such as: creation, reading, updating, or detection (CRUD) of GPU resources associated with the client; disposal of GPU resources; retrying and idempotency of GPU resources; logging into GPU resources associated with the client; and issuing metrics corresponding to GPU resources associated with the client.
[0104] The CP workflow worker 706 is communicatively coupled to the CP API server 704 and is also a Dropwizard application derived from templates generated by Pegasus. Many CRUD operations are inherently long-running and multi-step. Therefore, WfaaS is used by the CP workflow worker 706 to orchestrate CRUD operations or long-running jobs and manage the lifecycle of workflows. Workflows include verification steps, GPU resource CRUD steps, polling steps, update Kiev steps, and cleanup steps. In the verification step, the CP workflow worker 706 performs synchronization checks and / or re-verifications on conditions that may have changed between the time the request is accepted by the CP API server and the time the request is picked up by the CP workflow worker 706. In the GPU resource CRUD step, the CP workflow worker 706 passes the operation to the MP API server 708 by calling the corresponding management plane API. In the polling step, if the job is still in progress, the CP workflow worker 706 periodically polls the job's status via the MP API server 708 until the job reaches a termination state. In the Kiev update step, after the job corresponding to the request is completed, the CP workflow worker 706 updates the job metadata in Kiev as a Service (Kaas), including lifecycle status, lifecycle details, and job completion percentage to reflect the updated progress. In the cleanup step, in the event of a workflow failure, the CP workflow worker 706 cleans up items generated in previous steps of the failed workflow to ensure an overall consistent state. WFaas and KaaS are referenced together in box 702, as... Figure 7 As shown in the image.
[0105] The MP API server 708 infers the intent or configuration of a request from the client. The MP API server 708 is communicatively coupled to the GenAI DP 620 via a local peer-to-peer gateway (LPG). The GenAI DP 620 hosts resources such as dedicated AI clusters (DACs), LLM / AI models, and AI model inference endpoints, and assists the MP API server 708 in executing the intent inferred from the request.
[0106] The GenAI DP 620 constructs the DAC, which is a logical construct of the reserved capacity within the GenAI DP 620. Generative AI is responsible for maintaining and running the underlying infrastructure (GPUs) for specific customers. The DAC operator manages the lifecycle of the dedicated AI cluster, including capacity allocation, monitoring, and patching.
[0107] Examples of such intentions include creating dedicated AI clusters, fine-tuning jobs on AI models, and creating dedicated inference endpoints on specific base AI models or fine-tuned AI models. The state of MP API server 708 is stored back to the control plane Kiev.
[0108] Additionally, the intent could include using DAC to host various AI services for clients, such as streaming services, object storage, file storage systems (FSS), identity and access management (IAM), vaults, etc. References to such AI services can be placed in box 710, such as... Figure 7 As shown in the image.
[0109] MP API server 708 is responsible for supplying DACs, scheduling fine-tuning or dedicated inference endpoint processing on DACs, monitoring DACs and processing, and performing automatic maintenance and upgrades on DACs and processing. MP API server 708 invokes GenAI resource operator 712 and ML job operator 626 to manage the entire lifecycle of GenAI DP 620 resources. GenAI DP 620 includes a data plane API server 716 for controlling model endpoint 718 and inference server 720. Model endpoint 718 includes fine-tuned weights corresponding to the fine-tuned model. The base model can be fine-tuned based on the weights assigned to the corresponding model. Inference server 720 is used to host various AI services for clients, such as streaming services, object storage, file storage system (FSS), identity and access management (IAM), vaults, etc. Model endpoint 718 and inference server 720 can be controlled by data plane API server 716 to fine-tune model 722.
[0110] The data plane also includes a model repository 724. Model repository 724 stores data related to fine-tuning model 722, model endpoint 718, and inference server 720. Model endpoint 718 and inference server 720 are used by the data plane API server 716 to provide responses to clients based on intents. Responses to requests can be in the form of a single response, i.e., a token-based response.
[0111] Figure 8 The illustration shows a flow diagram 800 illustrating how, according to an exemplary embodiment, a control plane request is propagated to the management plane and subsequently to the data plane of the generative AI platform 800. (See diagram 800 for details.) Figure 8 As shown, the generative AI platform 800 receives a request from client 108. In some embodiments, the request may correspond to the allocation of GPU resources for a fixed time period, or to changing the amount of GPU resources currently utilized by the client by increasing or decreasing the GPU resources currently allocated to the client. In some other embodiments, the request corresponds to operations such as: creation, reading, updating, or detection (CRUD) of GPU resources associated with the client, disposal of GPU resources, retries and idempotency of GPU resources, logging into GPU resources associated with the client, and issuing metrics corresponding to GPU resources associated with the client.
[0112] Requests can be received at GenAI CP 618 via SPLAT 804. SPLAT 804 provides services including authorization of requests based on the client's client identifier (ID). SPLAT 804 verifies the identity of the client system 104 associated with the client. For example, SPLAT 804 verifies the scope of services / requests allowed to the client based on the client ID. The scope of services includes the number of requests allowed to the client. In some embodiments, a private key extracted from the asymmetric key pair associated with the client ID is used to authenticate the request. After authentication, SPLAT 804 forwards the request to the generative AI platform 800.
[0113] The generative AI platform 800 includes GenAI CP 618, Generative AI Management Plane (GenAI MP) 708, and GenAIDP 620. Requests can be received by GenAI CP 618. GenAI CP 618 includes a Control Plane API (CP API) server 704, a Control Plane (CP) Keiv 806, a Control Plane (CP) worker 808, and a Control Plane Workflow as a Service (CPWFaaS) 810. CP API server 704 serves as the entry point for GenAI CP 618. CP API server 704 is configured to provide workflow as a service in response to requests.
[0114] The CP worker 808 is communicatively coupled to the CP API server 704, orchestrating CRUD operations as long-running jobs and managing the lifecycle of the workflows used to execute these operations. CRUD operations can be managed by various steps, such as verification steps, GPU resource CRUD steps, polling steps, Kiev update steps, and cleanup steps. The steps mentioned above have been referenced... Figure 7 It has been described.
[0115] The CP Kiev 806 stores metadata associated with requests. This metadata includes the identity of the client system 104 associated with client 108, data about various types of operations to be performed, and the requirements associated with those operations. For example, these operations could include fine-tuning of a basic AI model, requests for predefined throughput and latency when using a generative AI platform, requests to obtain AI models for streaming or other services, etc. The CP Kiev 806 also stores... Figure 7 The status of the operations / tasks included in the requests discussed herein.
[0116] CP worker 808 interacts with GenAI MP 802. GenAI MP 802 includes a management plane API (MP API) server 708 and an operator 814. MP API server 708 receives requests from GenAI CP 618 and infers the intent / configuration of the request based on the metadata associated with the request.
[0117] Examples of such intents include creating dedicated AI clusters, fine-tuning jobs on AI models, and creating dedicated inference endpoints on specific base or fine-tuned AI models to ensure predefined throughput and latency. Additionally, intents could include using a DAC to host various AI services for clients, such as streaming services.
[0118] Operator 814 fully leverages Kubernetes custom resources to perform various operations based on requests. These operations can be executed within GenAI DP 620. In one implementation, GenAI MP 802 can invoke DAC operator 622 to create DAC 816 using GPU resources. In another implementation, GenAI MP 802 can invoke ML job operator 626 to manage the entire lifecycle of resources for model 818, such as CRUD operations on a dedicated AI cluster based on request metadata. In yet another implementation, GenAI MP 802 can invoke model endpoint operator 628 to create and manage AI model endpoint 820, used to establish an interface with the AI model requested by client 108.
[0119] In some embodiments, when the request metadata suggests that a certain amount of GPU resources are needed in the GenAI DP 620 for a fixed period of time, the DAC operator retrieves the attributes of the available GPU resources in the GenAI DP 620. The attributes are retrieved from KubernetesCP 822. The attributes of the available GPU resources are compared with the GPU resource requirements suggested by the request metadata. The DAC operator then creates a dedicated AI cluster (DAC) 816 that includes a set of GPU resources that meet the requirements suggested by the request metadata.
[0120] The DAC operator creates the DAC 816 based on the type of operation the client wants to perform. For example, if the request metadata indicates that the operation type is a fine-tuning operation, the DAC operator creates the DAC 816 by acquiring all GPU resources from a single location or a single GPU resource pool. For other operations, the DAC operator can generate the DAC 816 using a set of GPU resources, where the GPU resources are acquired from different locations or multiple GPU resource pools.
[0121] Once a DAC is created, it is associated with the client system ID 104 and reserved for use by client 108.
[0122] Within each DAC, the DAC operator runs a dummy procedure until the DAC is requested by its corresponding client, and terminates the dummy procedure when the client requests the DAC after its creation. This helps avoid latency caused by cold starts of GPU resources.
[0123] Additionally, the DAC operator monitors for physical and logical faults in the GPU resources within the DAC. Physical faults in GPU resources can correspond to overheating and other physical parameters associated with the GPU resource, such as the clock cycles of the core associated with the GPU resource (derived by comparing various physical parameters of the GPU resource to a predefined set of physical parameters), the internal memory of the GPU resource, and the power supply to the corresponding GPU resource. Similarly, the DAC operator identifies logical faults in GPU resources. Logical faults can involve vulnerabilities in security protocols associated with the GPU resource. Furthermore, logical faults can involve faults in plugins associated with the GPU resource, GPU resource startup failures, and runtime failures of the GPU resource.
[0124] Once a GPU resource in the DAC is determined to be erroneously activated due to a physical or logical fault, the DAC operator identifies a replacement for the erroneous GPU resource from the cluster of GPU resources reserved for replacement purposes. The replacement GPU resource is selected from multiple GPU resources in the cluster reserved for replacement purposes, and its rotated hash value has the DAC's rotated hash value.
[0125] Once the replacement GPU resource is identified, the DAC operator releases the erroneous GPU resource from the DAC and patches the DAC with the replacement GPU resource.
[0126] Figure 9A and Figure 9B The illustration shows a sequence diagram 900 of the DAC creation process according to an exemplary embodiment. This process can be sequentially followed by various modules, such as a CP API server, CP Kiev, CP worker, MP API server, K8s API server, and DAC operator. At step 902, the CP API server receives a DAC request for creating a DAC. The DAC request includes the client's client ID. In some embodiments, the DAC request may correspond to allocating or increasing computing capacity. Computing capacity corresponds to the amount of GPU resources available to the client.
[0127] In response to this request, the CP API server creates a work request for creating the DAC at step 904 and associates the work request with the client ID. Metadata for creating the DAC (DAC metadata) is provided to the CP API server via this request. The DAC metadata may indicate the lease, compartment, and capacity reserved for the DAC.
[0128] The CP API server stores the job request and DAC metadata in the CP Kiev at step 906. The job request ID and DAC metadata can be transferred to the client at step 908.
[0129] At step 910, the CP worker receives a DAC creation job request from the CP Kiev. At step 912, the CP worker forwards the request to the MP API server.
[0130] In response to this request, the MP API server creates a DAC control request at step 914. The DAC control request can indicate the number of GPU resources available for the operation requested by the client. At step 914, the MP API server uses the DAC control request to store the request and DAC metadata in a Kubernetes custom resource in the data plane.
[0131] At step 916, the DAC operator uses a Kubernetes custom resource in the data plane to trigger the creation of the DAC and monitors the DAC being created by the Kubernetes operator. In this type of step, the Kubernetes custom resource triggers the DAC operator to collect data associated with each GPU resource available for the cluster. The DAC operator compares the collected data with DAC metadata to identify the GPU resources used for the cluster. The identified GPU resources are patched together to form the DAC.
[0132] At step 918, the CP worker requests the polling status of the work requests from the MP API server. At step 920, the MP API server forwards the request for the polling status of the work requests to the Kubernetes operator. At step 922, the Kubernetes operator returns the polling status of the work requests to the MP API server. The polling status can be further forwarded to the CP worker at step 924. The polling status can be updated in the CP Kiev at step 926.
[0133] At step 928, the CP API server receives requests to list, retrieve, and change compartments. These requests can correspond to operations such as: list (list): to list the AI models available within the generative AI platform; get (get): to return an AI model; and change compartment (change compartment): to change the compartment to which an AI model belongs. At step 930, the CP API server reads or writes DAC metadata stored in CP Kiev.
[0134] At step 932, the CP API server transmits a response to the client. This response includes one of the following: listing the AI models available within the generative AI platform, returning the AI model corresponding to the model ID derived from the request metadata, or changing the compartment to which the AI model whose model ID the client provided in the request metadata belongs.
[0135] At step 934, the CP API server receives a request to update / delete the DAC. At step 936, the CP API server creates a work request for creating or deleting the DAC and associates the work request with the client's system ID (104 ID). Metadata for creating the DAC (DAC metadata) is provided to the CP API server via this request. The DAC metadata may indicate the type of operation the client wants to perform on the DAC, the amount of time required by the DAC, and other information related to the DAC's lifecycle and management.
[0136] The CP API server stores the job request and DAC metadata in the control plane Kiev at step 938. The job request ID and DAC metadata can then be transferred to the client at step 940.
[0137] At step 942, the CP worker picks up the DAC delete / create job request from the CP API server. At step 944, the CP worker forwards the request to the management plane API.
[0138] In response to this request, the MP API creates a DAC control request at step 946. The MP API server uses the DAC control request to store the request metadata and DAC metadata in a custom Kubernetes resource in the data plane.
[0139] At step 948, the DAC operator triggers the update / deletion of the DAC using a Kubernetes custom resource in the data plane and monitors the DACs that the Kubernetes operator is creating. In this type of step, the Kubernetes custom resource triggers the DAC operator to collect data related to each GPU resource patched into the DAC. The DAC operator updates / deletes the DAC based on requests.
[0140] At step 950, the CP worker requests the polling status of the work requests from the MP API server. At step 952, the MP API server forwards the request for the polling status of the work requests to the Kubernetes operator. At step 954, the Kubernetes operator returns the polling status of the work requests to the MP API server. The polling status can be further forwarded to the CP worker at step 956. The polling status can be updated in the CP Kiev at step 958.
[0141] Figure 10 A flow diagram 1000 is illustrated, indicating the operation of a DAC operator 622 according to an exemplary embodiment. (See diagram 1000 for details.) Figure 10 As shown, MP API server 708 receives a request to perform CRUD operations. MP API server 708 uses the CRUD operations to determine DAC metadata. DAC metadata includes custom resource definitions (CRDs) stored in Kubernetes CP 822. Kubernetes CP provides the DAC metadata to DAC operator 622. DAC operator 622 performs CRUD operations on the DAC and uses the CRDs to manage the DAC's lifecycle.
[0142] DAC operator 622 triggers functions to manage the compute capacity 1008 for the DAC. For example, DAC operator 622 can deploy GPU resources to a dedicated AI cluster or remove GPU resources from a dedicated AI cluster. Additionally, DAC operator 622 triggers functions to ensure that the capacity limit 1010 configured in the DAC metadata is met. For example, DAC operator 622 determines the compute capacity available for each GPU resource in the cluster and selects a number of GPU resources such that the total compute capacity of that number of GPU resources is within the capacity limit 1010. Furthermore, the DAC operator triggers Kubernetes controller 1012 to perform various operations, such as monitoring and managing the health and uptime of GPU resources, detecting node problems, and patching new GPU resources if node problems are detected.
[0143] The DAC operator also updates the DAC based on DAC metadata, such as compartment changes. The DAC operator performs CRUD operations and updates the DAC. These operations provide the client with the DAC's status. For example, the status can be compared to the actual status in box 1014. Based on this comparison, an action is performed at box 1016. The operation remains suspended until any discrepancies or changes in events are detected at box 1018.
[0144] Figure 11The diagram illustrates a flow chart 1100 of the processing of a model operator 624 according to an exemplary embodiment. The model operator 624 manages all LLM / AI models and fine-tuned models within the generative AI platform. The model operator 624 can access and store all types of AI models. The AI models in the generative AI platform are either open-source AI models or AI models obtained through collaboration with one or more AI model providers. Therefore, each AI model can have different specifications (e.g., one AI model may be specifically designed for summarization, while another may be specifically designed for text completion). Additionally, each AI model can potentially have different dimensions or parameter counts. Furthermore, AI models from some providers may have specialized storage and access methods to protect the intellectual property rights of the models. Metadata describing each AI model is persistently stored in CPKiev. CPKiev acts as a source of canonical facts for all AI model-related information.
[0145] The DAC operator also updates the DAC based on DAC metadata, such as compartment changes. The DAC operator performs CRUD operations and updates the DAC. These operations provide the client with the DAC's status. For example, the status can be compared to the actual status in box 1014. Based on the comparison, an action is performed at box 1016. The operation remains suspended until any difference or change in event is detected at box 1018.
[0146] Figure 12 The diagram illustrates a sequence 1200 of the process of creating a base model and fine-tuning custom resources for the model using model operator 624 according to an exemplary embodiment. At step 1202, the control plane Kubernetes receives a request to create custom resources for the base AI model weights. At step 1204, the control plane Kubernetes waits for the model operator to create a Kubernetes job to download the base AI model weights.
[0147] At step 1206, the model operator creates a Kubernetes job to download the base AI model weights. At step 1208, the model operator invokes the ML job controller to begin downloading the AI model weights. At step 1210, the model download agent downloads the base AI model weights from the object storage bucket. The model download agent also encrypts the base AI model weights at step 1212. Additionally, the model download agent stores the base AI model weights in the FSS at step 1214 and updates the base AI model status to the CP API server at step 1216.
[0148] At step 1218, the Kubernetes control plane receives a request to create custom resources for the fine-tuned model. At step 1220, the model operator examines the fine-tuned model's custom resources. At step 1222, the model operator creates an ML job for the ML job operator to fine-tune the underlying AI model. At step 1224, the ML job operator creates a Kubernetes job for the Kubernetes job controller and updates the job status to in progress.
[0149] At box 1226, the Kubernetes job controller creates a job to fine-tune the base AI model. At box 1230, training processing is performed by the Kubernetes job controller to fine-tune the base AI model. At step 1232, the fine-tuned model weights are pushed to the object bucket. At step 1234, the Kubernetes job controller updates the status of the job to fine-tuning the base AI model to the ML job operator, indicating success. At step 1236, the ML job operator updates the status of the job to the model operator, indicating success, which subsequently leads to the removal of the ML job custom resource from the ML job operator at step 1238 and an update to the control plane API server at step 1240 that the fine-tuned model is ready for use.
[0150] A challenge inherent in training or fine-tuning LLM / AI models is that the size and complexity of AI models require the use of multiple powerful GPU resources. Generative AI platforms currently require DACs for model training to provide predictable training performance, including training job queue time, training throughput, and cost. While on-demand training (e.g., without requiring reserved capacity) may be available in the future, the current inventory of GPU resources is too low to meet all demands. Therefore, fine-tuning is limited to clients who have already ordered DACs to ensure that key customers (including internal teams such as Fusion App) have access to this capability and to maintain a good client experience for these clients.
[0151] To fine-tune the model, the client provides a training corpus, typically containing multiple examples. The training corpus includes prompt request / response pairs. The fine-tuned model can be used for various tasks, such as text classification, summarization, and entity extraction.
[0152] Text classification involves reusing the ability of a fine-tuned model to make predictions based on input text / speech. For example, a fine-tuned model might assign a rating to a given product based on reviews provided by a client in text or speech. For instance, a rating of 0 might be given if a review indicates the product arrived damaged, while a rating of 1 might be given if the review indicates super-fast delivery or friendly customer service.
[0153] Summarizing involves reusing the fine-tuned model's ability to summarize the input text / speech. For example, the input "Jack and Jill went up the mountain..." is summarized as "The two children, despite the difficulty of climbing the mountain to fetch water, still persevered in fetching a bucket of water for their mother."
[0154] Entity extraction involves reusing the ability of a fine-tuned model to extract entities of interest from input text or speech. For example, if the input is “Doxycycline is a class of antibiotics called tetracyclines. It treats infections by stopping the growth and spread of bacteria…”, then the entity of interest could be the drug name [doxycycline, tetracycline…] extracted by the fine-tuned model.
[0155] Fine-tuning the base AI model involves various steps. For example, the training data is validated for use in fine-tuning the base AI model. After validation, the base AI model is downloaded and decrypted in preparation for fine-tuning. Once the base AI model is ready, model-specific fine-tuning logic is executed, and the fine-tuning operation is monitored until completion. If fine-tuning is successful, the fine-tuned model weights and any relevant metrics / logs corresponding to the fine-tuned model are stored.
[0156] Figure 13 The illustration shows a sequence diagram 1300 of a method for fine-tuning a base AI model according to an exemplary embodiment. At step 1302, a request to create a fine-tuning job is received by the Kubernetes job module. At step 1304, the model operator invokes the Kubernetes job in response to the fine-tuning job. The Kubernetes job executes a fine-tuning initiator (init) container to prepare training data and decrypt the base model. At step 1306, the fine-tuning initi container requests a file storage server (FSS) to provide the base AI model for fine-tuning and decrypts the base AI model received from the FSS. At step 1308, the fine-tuning initi container downloads the training dataset for the base AI model from the client's object storage device.
[0157] Once the base AI model and training dataset are available, the Kubernetes job invokes a fine-tuning container on the DAC associated with the client at step 1310 to tune the base AI model. At step 1312, a fine-tuning sidecar service deployed along with the fine-tuning container monitors the fine-tuning process.
[0158] When the fine-tuning sidecar observes the completion of the fine-tuning process at step 1314, it stores the fine-tuned weights and fine-tuned metrics into the FSS at steps 1316 and 1318.
[0159] Figure 14The illustration shows a data flow diagram 1400 for fine-tuning a data model according to an exemplary embodiment. Initially, training data for a job is received by GenAI CP 618 to train the data model. The training data is further transmitted from GenAI CP 618 to model operator 624 via MP API server 708. In response to receiving the training data, model operator 624 generates a trigger signal for ML job operator 626, which is responsible for performing training for creating the fine-tuned model and managing the training workload. The ML job operator generates a management signal that enables fine-tuning unit 1402 to fine-tune the data model based on the training data.
[0160] The fine-tuning unit 1402 includes a fine-tuning initiator (FT init) 1404, a fine-tuning server 1406, and a fine-tuning sidecar 1408. In response to receiving a management signal, the FT init 1404 provides a fine-tuning environment for the data model. In some aspects of this disclosure, to provide the fine-tuning environment, the FT init 1404 sets the weights of the LLM model stored in the internal memory of the GPU resources. The FT init 1404 also downloads and verifies the training data for the job. In some aspects of this disclosure, the FT init 1404 can receive the training data for the job from the client object storage device 1412 or as inline data. Preferably, a delegate (OBO) token can be used to access client data from the client object storage device 1412. The FT init 1404 also receives the base model for training from the GenAI File Storage System (FSS) 1414. The fine-tuning server 1406 provides a quasi-black-box implementation for fine-tuning the data model. Specifically, the fine-tuning server 1406 runs a container that exposes multiple APIs to initiate training of the data model, obtain training metrics, obtain training status, and close the container after training the data model. The fine-tuning sidecar 1408 tracks the status of the fine-tuning job, exports training metrics of interest (such as accuracy and loss) to the GenAI object storage device 1410, and exports the fine-tuned model weights of the data model to the GenAI object storage device 1410.
[0161] Figure 15A and Figure 15BThe illustration shows a sequence diagram 1500 of the DAC creation process according to an exemplary embodiment. This process can be sequentially followed by various modules, such as a CP API server, CP Kiev, CP worker, MP API server, K8s API server, and endpoint operator. At step 1502, the CP API server receives an endpoint request for creating an endpoint. The endpoint request includes the client's client ID. In some embodiments, the DAC request may correspond to allocating or increasing computing capacity. Computing capacity corresponds to the amount of GPU resources available to the client.
[0162] In response to this request, the CP API server creates a work request for creating the endpoint at step 1504 and associates the work request with the client ID. Metadata for creating the endpoint (endpoint metadata) is provided to the CP API server via this request. Endpoint metadata may indicate the lease, compartment, and capacity reserved for the endpoint.
[0163] The CP API server stores the job request and endpoint metadata in CP Kiev at step 1506. The job request ID and endpoint metadata can be transmitted to the client at step 1508.
[0164] At step 1510, the CP worker receives an endpoint create job request from the CP Kiev. At step 1512, the CP worker forwards the request to the MP API server.
[0165] In response to this request, the MP API server creates an endpoint control request at step 1514. The endpoint control request can indicate the amount of GPU resources required for the operation requested by the client. The MP API server uses the endpoint control request to store the request and endpoint metadata in a custom Kubernetes resource in the data plane.
[0166] At step 1516, the DAC operator triggers endpoint creation using a Kubernetes custom resource in the data plane and monitors the endpoints being created by the Kubernetes operator. In this type of step, the Kubernetes custom resource triggers the endpoint operator to collect data associated with each GPU resource available for the cluster. The endpoint operator compares the collected data with endpoint metadata to identify the GPU resources used for the cluster. The identified GPU resources are patched together to form the endpoint.
[0167] At step 1518, the CP worker requests the polling status of the work requests from the MP API server. At step 1520, the MP API server forwards the request for the polling status of the work requests to the Kubernetes operator. At step 1522, the Kubernetes operator returns the polling status of the work requests to the MP API server. The polling status can be further forwarded to the CP worker at step 1524. The polling status can be updated in the CP Kiev at step 1526.
[0168] At step 1528, the CP API server receives a request to list / get endpoints. This request can correspond to operations such as: list: to list the AI models available within the generative AI platform; and get: to return the AI models. At step 1530, the CP API server reads or writes the endpoint metadata stored in CP kiev.
[0169] At step 1532, the CP API server transmits a response to the client. This response includes one of the following: listing the AI models available within the generative AI platform, returning the AI model corresponding to the model ID derived from the request metadata, or changing the compartment to which the AI model whose model ID the client provided in the request metadata belongs.
[0170] At step 1534, the CP API server receives a request to update / delete an endpoint. At step 1536, the CP API server creates a work request for creating or deleting the endpoint and associates the work request with the client's system ID (104 ID). Metadata for creating the endpoint (endpoint metadata) is provided to the CP API server via this request. Endpoint metadata may indicate the type of operation the client wants to perform on the endpoint, the amount of time the endpoint requires, and other information related to the endpoint's lifecycle and management.
[0171] The CP API server stores the job request and DAC metadata in the control plane Kiev at step 1538. The job request ID and endpoint metadata can then be transmitted to the client at step 1540.
[0172] At step 1542, the CP worker retrieves the endpoint update / change request from the CP API server. At step 1544, the CP worker forwards the request to the management plane API.
[0173] In response to this request, the MP API creates an endpoint control request at step 1546. The MP API server uses the endpoint control request to store request metadata and endpoint metadata in a custom Kubernetes resource in the data plane.
[0174] At step 1548, the endpoint operator triggers updates / changes to the endpoint using Kubernetes custom resources in the data plane and monitors the endpoints the Kubernetes operator is creating. In this type of step, the Kubernetes custom resource triggers the endpoint operator to collect data related to each GPU resource patched into the endpoint. The endpoint operator updates / changes the DAC based on requests.
[0175] At step 1550, the CP worker requests the polling status of the work requests from the MP API server. At step 1552, the MP API server forwards the request for the polling status of the work requests to the Kubernetes operator. At step 1554, the Kubernetes operator returns the polling status of the work requests to the MP API server. The polling status can be forwarded to the CP worker at step 1556. The polling status can be updated in the CP Kiev at step 1558.
[0176] Figure 16 The illustration shows a data flow diagram 1600 for managing a fine-tuned inference server 210 throughout its entire lifecycle using a model endpoint operator 628, according to an exemplary embodiment. The dedicated model inference endpoint provides the ability to perform inference on its pre-trained or fine-tuned models with predictable latency and throughput. Each deployed dedicated model (fine-tuned or pre-trained) is hosted in a dedicated AI cluster (DAC) and replicated across all GPU resources in the DAC. Preferably, the DAC can serve a base model with up to 50 instances of fine-tuned weights. Model endpoint metadata is synchronized from the CP Kiev806 to the internal cache of the generative AI data plane 202 for authorization, routing, and rate limiting of requests.
[0177] In operation, GenAI CP 618 receives one or more inputs (such as requests) from clients for creating one or more model endpoints. In parallel, Data Plane API Server 716 receives inference inputs from clients. Based on the received inputs, GenAI CP 618 enables CP Kiev 806 to provide Kiev stream model metadata to Data Plane API Server 716. Based on the inputs received from clients and the Kiev stream model metadata received from CP Kiev 806, Data Plane API Server 716 generates service module identities (i.e., authorized OCI identities for requests) and service module limits (i.e., OCI computation limits for each model). Model endpoint operator 628 is coupled to GenAI CP 618 via MPAPI Server 708. Model endpoint operator 628 manages one or more inference services 1601 for activating / deactivating one or more fine-tuning weights. Based on the inputs received for creating model endpoints(1,2,3)2, model endpoint operator 628 modifies the model weights of inference server 1602 using one or more inference services(1,2,3)2. Inference server 1602 is also coupled to and receives inference inputs from data plane API server 716. In some aspects of this disclosure, inference server 1602 includes inference server 1604 supporting service sidecar 1606 and service initiator 1608. Service sidecar 1606 receives one or more instructions from model endpoint operator 628 for activating / deactivating one or more fine-tuning weights(1,2,3)2. Service initiator 1608 receives the base model from GenAI FSS 1414. Service sidecar 1606 retrieves the fine-tuned weights from GenAI object storage device 1410 and modifies one or more fine-tuned weights of the base model(s) for multiple instances based on the received instructions(s,2,3)2.
[0178] Figure 17 The illustration shows a sequence diagram of logical processing 1700 within an inference server 1602 according to an exemplary embodiment. Processing 1700 includes sequential steps between the data plane API server 716 and the inference server 1604. The processing steps corresponding to the inference server 1604 include operations of the Dori module, the Nemo module, and the Triton module.
[0179] Processing 1700 begins at step 1702, at which point the data plane API server 716 receives an inference request from the client with one or more dedicated endpoints.
[0180] At step 1704, the data plane API server 716 generates a request from the model metadata in the internal memory (hereinafter referred to interchangeably as "internal memory") to obtain the Domain Name System (DNS) corresponding to the request with one or more dedicated endpoints.
[0181] At step 1706, in response to the request to obtain DNS, the internal memory transmits the corresponding DNS along with a cloud identifier (e.g., an Oracle cloud identifier) associated with one or more dedicated endpoints to the data plane API server 716.
[0182] At step 1708, the data plane API server 716 constructs an inference request to the inference server 1604 based on the DNS and the identifier of the fine-tuned weights (fine-tuning ID).
[0183] In steps 1710-1718, the inference server 1604 (via Dory) receives the inference request and determines its validity. After verifying the inference request, Dory routes the traffic corresponding to the specific fine-tuned weights. Preferably, in step 1710, Dory transmits the inference request to Nemo for further processing. In response, in step 1712, Nemo batches the inference request and forwards it to Triton. Based on the inference request, in step 1714, Triton returns an inference return token to Nemo. Based on the received token, Nemo returns the inference result to Dory in step 1716, and Dory forwards the result to the DP API server 716 in step 1718.
[0184] Figure 18 The diagram illustrates a flow chart 1800 instructing the work of a model endpoint operator according to an exemplary embodiment. A model endpoint operator 628 manages the lifecycle of a service endpoint and continues to reconcile the service endpoint. Typically, the inference service endpoint is specific to either the base model or a fine-tuned model. For the base model endpoint, the model endpoint operator 628 reconciles low-native Kubernetes resources such as ingress 1802, deployment 1804, HPA (Horizontal Cluster Autoscaler) 1806, and the base model inference service. For the fine-tuned model endpoint, the model endpoint operator 628 reconciles the resources associated with the base model endpoint. Furthermore, the model endpoint operator 628 reconciles the state of the fine-tuned model weights. A service sidecar 1606 receives traffic from the model endpoint operator associated with the state of the fine-tuned weights. When a model weight is absent, the model endpoint operator 628 triggers the service sidecar 1606 to retrieve one or more fine-tuned weights across the inference server 1604.
[0185] The DAC operator also updates the DAC based on DAC metadata, such as compartment changes. The DAC operator performs CRUD operations and updates the DAC. These operations provide the client with the DAC's status. For example, the status can be compared to the actual status in box 1014. Based on the comparison, an action is performed at box 1016. The operation remains suspended until any difference or change in event is detected at box 1018.
[0186] pass Figure 19 Sequence diagram 1900 in the document presents an embodiment for creating a base inference service, a fine-tuned inference service, and deleting a fine-tuned inference service.
[0187] Figure 20 The illustration shows a flowchart of a process 2000 for allocating a dedicated AI cluster according to an exemplary embodiment. At block 2002, a request for GPU resource allocation is received from a client system associated with the client. This request may be received by a computing system. In some embodiments, the request may be related to the execution of an operation. The request includes metadata identifying a client ID associated with the client, the latency of the operation, and its throughput. The request is authenticated using the client ID associated with the client and / or the client system. The client ID may be associated with a Hypertext Transfer Protocol (HTTP) signature used for authentication of the request. The authentication process involves signing the HTTP signature header using a private key extracted from an asymmetric key pair.
[0188] In some embodiments, the computing system can obtain a pre-approved quota associated with the request. If the pre-approved quota exceeds a predefined request limit corresponding to the client ID, the request can be blocked. If the pre-approved quota is within the predefined request limit, the request can be forwarded for further processing.
[0189] In box 2004, when authenticating a request using metadata, resource constraints for performing an operation are determined based on the metadata. To determine resource constraints, the metadata can be analyzed to determine the scope of the operation. This scope indicates the computation to be performed to execute the operation. The scope of the operation is used to determine the resource constraints for performing the operation. Resource constraints indicate the computational capacity used to perform the operation.
[0190] Additionally, the computing system identifies available GPU resources for allocation. Each GPU resource has an attribute indicating its capacity. At box 2006, the attributes of the GPU resources are obtained for analysis. Furthermore, the attributes of the GPU resources are analyzed relative to resource limits to determine a set of GPU resources for performing the operation. For example, the attributes of the GPU resources can be compared with resource limits. At box 2008, if the attributes of a GPU resource do not match the resource limits, the corresponding GPU resource can be rejected at box 2010. If the attributes of a GPU resource match the resource limits, the corresponding GPU resource can be selected at box 2012.
[0191] At box 2014, the GPU resources selected for allocation can be combined and patched together to form a dedicated AI cluster. The dedicated AI cluster reserves a portion of the computing capacity of the computing system for a period of time. In some embodiments, the type of operation can be determined based on the request. Alternatively, the group of GPU resources can be selected from multiple nodes or one of a single node to generate the dedicated AI cluster. For example, if the request relates to fine-tuning a data model, then the group of GPU resources is selected from a single node to form a dedicated AI cluster. In this case, the data model to be fine-tuned is obtained, and the fine-tuning logic is executed on the data model using the dedicated AI cluster.
[0192] At box 2016, a dedicated AI cluster is assigned to the client. Once the dedicated AI cluster is assigned to the client and / or the client system associated with the client, the operation requested by the client begins execution using the GPU resources patched into the dedicated AI cluster. Assigning the dedicated AI cluster to the client ensures that the workload associated with the operation requested by one client is not mixed with and matched with the workload associated with an operation requested by another client. Therefore, the computing system is able to provide computational capacity for operations that require a large amount of computing power, such as training or fine-tuning a data model.
[0193] Figure 21The illustration shows a flowchart of a fault management process 2100 according to an exemplary embodiment. At block 2102, performance parameters of GPU resources are monitored. Performance parameters include the physical and / or logical condition of the corresponding GPU resource. In an embodiment, the physical condition of each GPU includes the temperature of the corresponding GPU resource, the clock cycles per core associated with the corresponding GPU resource, the internal memory of the corresponding GPU resource, and / or the power supply of the corresponding GPU resource. The logical condition of each GPU resource includes faults in plugins associated with the corresponding GPU resource, startup problems associated with the corresponding GPU resource, runtime failures, and / or security vulnerabilities. The performance parameters corresponding to each GPU resource can be compared with predefined performance parameters. The predefined performance parameters indicate the performance parameters of an ideal GPU resource. In some embodiments, the predefined performance parameters can be a range of values within which the GPU resource is considered a fault-free GPU resource.
[0194] At box 2104, if the performance parameters do not deviate from the predefined performance parameters, the process flow of 2100 moves to box 2102 and the performance parameters of the GPU resources can be continuously monitored. If the performance parameters deviate from the predefined performance parameters, an anomaly in the GPU resources within this group can be identified at box 2106. When an anomaly is detected in the GPU resources, a replacement GPU resource is identified from the remaining available GPU resources in the computing system. To identify a replacement GPU resource, the computational capacity of the faulty GPU resource can be matched with the computational capacity of the replacement GPU resource at box 2108.
[0195] If the computing capacity of the replaceable GPU resource does not match that of the faulty GPU resource, the corresponding GPU resource can be rejected for replacement at step 2110. If the computing capacity of the replaceable GPU resource matches that of the faulty GPU resource, the replaceable GPU resource can be selected for replacement at step 2112. Replaceable GPU resources are identified by matching their computing capacity with that of each GPU resource in the dedicated AI cluster. For example, a rotation hash value can be calculated for the replaceable GPU resource. The rotation value is calculated based on the computational image and / or computational shape. The rotation hash value indicates the security and compliance status of the corresponding GPU resource. The rotation value of the replaceable GPU resource is further compared with the rotation hash value of the dedicated AI cluster. When the rotation value of the replaceable GPU resource matches that of the dedicated AI cluster, the replaceable GPU resource can be selected to replace the faulty GPU resource. After identification, at step 2114, the faulty GPU resource is released from the dedicated AI cluster, and the replaceable GPU resource is patched into the dedicated AI cluster. In this way, a set of replaceable GPU resources can be reserved for replacing faulty GPU resources in the dedicated AI cluster.
[0196] Patching replaceable GPU resources to a dedicated AI cluster can be terminated when predefined conditions are identified. These predefined conditions include: failure of the replaceable GPU resource during startup, failure of the replaceable GPU resource to join the dedicated AI cluster, workload failure of the replaceable GPU resource, and / or software errors detected in the replaceable GPU resource. Once the predefined conditions are identified, a label can be associated with the replaceable GPU resource. This label indicates the inapplicability of patching a set of replaceable GPU resources.
[0197] Figure 22 The illustration shows a flowchart of a process 2200 for managing operations performed using GPU resources, according to an exemplary embodiment. At block 2202, the computational capacity of the GPU resources is monitored. The GPU resources are reserved for the client for a period of time. During this period, the operation may not consume all the computational capacity reserved for the client. In this case, some GPU resources in this group are not utilized to perform the operation requested by the client.
[0198] At box 2204, determine whether the GPU resources in this group are being used by the client. If the GPU resources are being used by the client, then process 2100 moves to box 22022 and continues to monitor the computing capacity of this group of GPU resources. If the GPU resources are not being used by the client, then select the unused GPU resources at box 2206.
[0199] At box 2208, the attributes of the operation being performed by the client can be determined. These attributes are determined based on analysis of the operation's input and output. At box 2210, a dummy operation is generated based on these attributes. For example, at box 2212, the attributes of the dummy operation are matched against the attributes of the dummy operation. In some embodiments, the dummy operation may be the same operation as the actual operation performed by the client (e.g., the exact same operation). The dummy operation is executed on GPU resources not utilized by the client.
[0200] When the computing system receives a request to access GPU resources to perform an operation, it can terminate the dummy operation from its execution and patch it for the actual operation requested by the client. This reduces the time required to load GPU resources from idle states. As a result, it improves the overall processing latency of the operation.
[0201] EEE 1. A computer-implemented method comprising: receiving a request for allocating graphics processing unit (GPU) resources for performing an operation, wherein the request includes metadata identifying a client identifier (ID) associated with a client, a target latency of the operation, and a target throughput; determining resource constraints for performing the operation based on the metadata; obtaining at least one attribute associated with each of a plurality of GPU resources available for allocation in a computing system, wherein the at least one attribute indicates the capacity of the corresponding GPU resource; analyzing the at least one attribute associated with each GPU resource relative to resource constraints; identifying a set of GPU resources from the plurality of GPU resources based on the analysis; generating a dedicated AI cluster by patching the set of GPU resources into a single cluster, wherein the dedicated AI cluster retains a portion of the computing capacity of the computing system for a period of time; and assigning the dedicated AI cluster to a client associated with a client ID.
[0202] EEE 2. The method as described in EEE 1 further includes authenticating the request based on a client ID associated with the client before allocating the dedicated AI cluster, wherein the request is authenticated using a private key extracted from an asymmetric key pair associated with the client ID.
[0203] EEE 3. The method of claim EEE 1 or EEE 2, further comprising: comparing a set of performance parameters corresponding to each GPU resource in the set of GPU resources with a predefined set of performance parameters; determining an anomaly in a first GPU resource in the set of GPU resources based on the comparison, wherein the anomaly indicates that the set of performance parameters deviates from the predefined set of performance parameters; and replacing the first GPU resource with a second GPU resource within a dedicated AI cluster, wherein the hash value of the second GPU resource is the same as the hash value of the first GPU resource.
[0204] EEE 4. The method of any one of EEE 1-3, further comprising: determining a pre-approved quota associated with the request; determining whether the pre-approved quota exceeds a predefined request limit corresponding to a client ID; and blocking the request based on the determination that the pre-approved quota exceeds the predefined request limit.
[0205] EEE 5. The method of any of EEE 1-4 further includes: determining the type of operation based on the request; and selecting the set of GPU resources from one of a plurality of nodes or a single node based on the type of operation to generate a dedicated AI cluster.
[0206] EEE 6. The method as described in EEE 5 further includes, based on determining the request, instructing a fine-tuning operation: obtaining a data model to be fine-tuned; and performing fine-tuning logic on the data model using a dedicated AI cluster, wherein the dedicated AI cluster is generated using the set of GPU resources selected from the single node.
[0207] EEE 7. The method of any one of EEE 1-6 further comprises: identifying at least one underutilized GPU resource from the set of GPU resources of the dedicated AI cluster; and performing a dummy operation on the at least one GPU resource in response to the identification of the at least one GPU resource, wherein the dummy operation is identical to an operation performed on the at least one GPU resource.
[0208] EEE 8. A computer-implemented method comprising: accessing a request for allocating graphics processing unit (GPU) resources for performing an operation, wherein the request includes metadata identifying a client identifier (ID) associated with a client, latency of the operation, and throughput; determining, based on the metadata, predicted resource constraints for performing the operation; obtaining at least one parameter of a plurality of GPU resources existing in a plurality of nodes, wherein the at least one parameter includes a status indicating whether a corresponding GPU resource is occupied for performing another operation; determining a GPU resource utilization value for each of the plurality of nodes based on the status of each GPU resource, wherein the GPU resource utilization value indicates the amount of GPU resource utilization of the corresponding node; comparing the GPU resource utilization value of each node with a predefined resource utilization threshold; rescheduling the plurality of GPU resources based on the predicted resource constraints in response to determining that the GPU resource utilization value is less than the predefined resource utilization threshold; and allocating a set of GPU resources from the plurality of rescheduled GPU resources for performing the operation.
[0209] EEE 9. The method as described in EEE 8, further comprising: simulating a request on each of the plurality of nodes; determining a percentage of resource utilization for each node based on the simulation of the request; identifying the node with the highest percentage of resource utilization from the plurality of nodes; and allocating the set of GPU resources from the identified node to a client.
[0210] EEE 10. The method as described in EEE 8 or EEE 9 further includes: determining the type of operation based on the request; and allocating the set of GPU resources from the plurality of GPU resources based on the type of operation.
[0211] EEE 11. The method of any of EEE 8-10 further comprises: generating a dedicated AI cluster by patching the set of GPU resources into a single cluster, wherein the dedicated AI cluster retains a portion of the computing capacity of the computing system for a period of time; and assigning the dedicated AI cluster to a client associated with a client ID.
[0212] EEE 12. The method of any of EEE 8-11 further comprises:
[0213] Before allocating the set of GPU resources, the request is authenticated based on the client ID associated with the client, wherein the private key extracted from the asymmetric key pair associated with the client ID is used to authenticate the request.
[0214] EEE 13. The method of any of EEE 8-12 further comprises: determining the number of tokens associated with the request; determining whether the number of tokens exceeds a predefined request limit corresponding to a client ID; and blocking the request based on determining that the number of tokens exceeds the predefined request limit.
[0215] EEE 14. The method of any of EEE 8-13 further comprises: determining the patching of the set of GPU resources based on predefined conditions, wherein the predefined conditions are one of the following: a failure of the set of GPU resources during startup; a workload failure of the set of GPU resources; and a software error detected in the set of GPU resources.
[0216] EEE15. A computer-implemented method comprising: allocating a first target resource amount and a second target resource amount for using a service to a first client and a second client respectively; receiving a request from a third client for allocating resources for using the service; estimating (i) that the first client is using a first subset of the first target resource amount and not using a second subset of the first target resource amount, and (ii) that the second client is using a third subset of the second target resource amount and not using a fourth subset of the second target resource amount; determining that the second subset of the first target resource amount is greater than the fourth subset of the second target resource amount; and at least partially in response to determining that the second subset of the first resource amount is greater than the fourth subset of the second resource amount, allocating at least a portion of the second subset of the first target resource amount as a third target resource amount to the third client.
[0217] EEE16. A computer-implemented method as described in EEE15, wherein allocating at least a portion of a second subset of a first target resource quantity to a third client comprises: estimating that the amount of system-level unallocated resources not yet allocated to any client is less than a threshold; and at least partially in response to estimating that the amount of system-level unallocated resources is less than the threshold, allocating at least a portion of the second subset of the first resource quantity to the third client.
[0218] EEE17. A computer-implemented method as described in EEE15 or EEE16, wherein the system-level unallocated resource amount does not include: (i) a first buffer resource amount reserved for allocation to the one or more new clients when no other resources are available for allocation to the one or more new clients, and wherein the first buffer resource amount is not used for allocation to any active client using the service, and (ii) a second buffer resource amount not explicitly allocated to any new client or active client.
[0219] EEE18. A computer-implemented method as described in any of EEE 15-17, wherein the second buffer resource amount serves at least in part as a safety margin to account for inaccuracies in resource allocation and / or resource estimation.
[0220] EEE19. A computer-implemented method as described in any of EEE 15-18, wherein the threshold is zero.
[0221] EEE20. A computer-implemented method as described in any of EEE 15-19, wherein the service is the use of an artificial intelligence (AI) model.
[0222] EEE21. A computer-implemented method as described in EEE20, wherein various resource quantities are measured based on the number of requests per time period (RPT) or the number of tokens per time period (TPT) for the AI model.
[0223] EEE22. A computer-implemented method as described in EEE20, wherein various resource quantities are measured based on requests per minute (RPM) or tokens per minute (TPM) for the AI model.
[0224] EEE23. A computer-implemented method as described in EEE20, wherein: a first client is a first lease in a cloud environment, the first lease hosting a first cloud application using services; a second client is a second lease in a cloud environment, the second lease hosting a second cloud application using services; a third client is a third lease in a cloud environment, the third lease hosting a third cloud application using services; and an AI model is leased and hosted by an AI provider in the cloud environment.
[0225] EEE24. A computer-implemented method as described in any of EEE 15-23, wherein a request is received from a third client and a third target resource amount is allocated during a first time period, and wherein the method further comprises: allocating a first modified target resource amount and a second modified target resource amount to a first client and a second client respectively during a second time period different from the first time period and for the purpose of using the service; receiving a request from a fourth client for allocating resources for the purpose of using the service; estimating that the amount of system-level unallocated resources not yet allocated to any client is greater than a threshold; and allocating at least a portion of the system-level unallocated resources as a fourth target resource amount to the fourth client.
[0226] EEE25. A computer-implemented method as described in any of EEE 15-24, wherein a request is received from a third client and a third target resource amount is allocated during a first time period, and wherein the method further comprises: receiving requests from a fourth client and a fifth client, and during a second time period different from the first time period, for allocating resources for using the service; determining that each of a plurality of active clients using the service has been allocated a minimum target resource for using the service, wherein the plurality of active clients includes the fourth client but excludes the fifth client; determining that a buffer resource amount reserved for allocation to one or more new clients has a non-zero value; allocating at least a portion of the buffer resource amount to the fifth client; and at least partially rejecting a request received from the fourth client during the second time period in response to determining that each of the plurality of active clients using the service has been allocated a minimum target resource.
[0227] EEE26. A computer-implemented method as described in any of EEE 15-25, wherein a request is received from a third client and a third target resource amount is allocated during a first time period, and wherein the method further comprises: allocating a fourth target resource amount and a fifth target resource amount to a fourth client and a fifth client respectively during a second time period different from the first time period and for the purpose of using the service; estimating that a portion of the fourth target resource amount not used by the fourth client is less than a low threshold; estimating that a portion of the fifth target resource amount not used by the fifth client is greater than a high threshold; selecting a fifth client for reallocating resources to the fourth client; and at least in part in response to selecting the fifth client, reallocating a portion of the fifth target resource amount of the fifth client to the fourth client.
[0228] EEE27. A computer-implemented method as described in EEE26, wherein selecting a fifth client for reallocating resources to a fourth client comprises: determining that, among a plurality of active clients using the service, the fifth client has the highest amount of unused resources; and selecting the fifth client for reallocating resources to the fourth client, at least in part, in response to determining that the fifth client has the highest amount of unused resources.
[0229] EEE28. A computer-implemented method as described in any of EEE 15-27, further comprising: periodically checking resource allocations of a plurality of active clients using the service; determining that a first active client among the plurality of active clients has an amount of unused resources less than a low threshold; allocating additional resources to at least one active client, wherein the additional resources allocated to the at least one active client are from one or more of: (i) system-level unallocated resources not yet allocated to any of the plurality of active clients, (ii) a second active client among the plurality of active clients having the highest amount of unused resources, and / or (iii) a third active client among the plurality of active clients having the highest amount of allocated resources.
[0230] EEE29. A system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more of the methods described in EEE 1-28.
[0231] EEE30. A computer program product tangibly implemented on a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to execute some or all of one or more of the instructions in EEE 1-28.
[0232] EEE31. One or more components for performing one or more of the EEEs as described in 1-28.
[0233] Some embodiments of this disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the methods disclosed herein and / or some or all of the processes disclosed herein. Some embodiments of this disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform some or all of the methods disclosed herein and / or some or all of the processes disclosed herein.
[0234] The terms and expressions used are intended to be descriptive rather than limiting, and there is no intention to use them or any part thereof. However, it should be recognized that various modifications can be made within the scope of the claimed invention. Therefore, while the claimed invention has been specifically disclosed by way of embodiments and optional features, modifications and variations of the concepts disclosed herein can be adopted by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.
[0235] This description provides preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. This description of the preferred exemplary embodiments will provide those skilled in the art with an implementation description for carrying out various embodiments. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the spirit and scope set forth in the appended claims.
[0236] Specific details are set forth in this description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments.
Claims
1. A computer-implemented method, comprising: Receive a request for allocating graphics processing unit (GPU) resources for performing an operation, wherein the request includes metadata identifying a client identifier (ID) associated with the client, the target latency of the operation, and the target throughput; Determine the resource limits for performing operations based on metadata; Obtain at least one attribute associated with each of a plurality of GPU resources available for assignment in a computing system, wherein the at least one attribute indicates the capacity of the corresponding GPU resource; The at least one attribute associated with each GPU resource in relation to resource constraint analysis; Based on the analysis, a set of GPU resources is identified from the plurality of GPU resources; A dedicated AI cluster is generated by patching the set of GPU resources into a single cluster, wherein the dedicated AI cluster retains a portion of the computing capacity of the computing system for a period of time. as well as Assign a dedicated AI cluster to a client associated with a client ID.
2. The method of claim 1, further comprising authenticating the request based on a client ID associated with the client before allocating the dedicated AI cluster, wherein the request is authenticated using a private key extracted from an asymmetric key pair associated with the client ID.
3. The method of claim 1, further comprising: The set of performance parameters corresponding to each GPU resource in the set of GPU resources is compared with a predefined set of performance parameters. Based on the comparison, an anomaly is determined in a first GPU resource within the set of GPU resources, wherein the anomaly indicates that the set of performance parameters deviates from a predefined set of performance parameters; and Within a dedicated AI cluster, the first GPU resource is replaced with a second GPU resource, wherein the hash value of the second GPU resource is the same as the hash value of the first GPU resource.
4. The method of claim 1, further comprising: Determine the pre-approved quota associated with the request; Determine whether the pre-approved quota exceeds the predefined request limit corresponding to the client ID; as well as The request is blocked based on the determination that the pre-approved quota exceeds the predefined request limit.
5. The method of claim 1, further comprising: The type of operation is determined based on the request; as well as The set of GPU resources is selected from one of multiple nodes or a single node based on the type of operation to generate a dedicated AI cluster.
6. The method of claim 5, further comprising determining a fine-tuning operation based on the request: Obtain the data model to be fine-tuned; and Fine-tuning logic is performed on the data model using a dedicated AI cluster, which is generated using the set of GPU resources selected from the single node.
7. The method of claim 1, further comprising: Identify at least one underutilized GPU resource from the set of GPU resources in the dedicated AI cluster; as well as In response to the identification of the at least one GPU resource, a dummy operation is performed on the at least one GPU resource, wherein the dummy operation is identical to an operation performed on the at least one GPU resource.
8. A system comprising: One or more processors; as well as A memory coupled to the one or more processors, the memory storing a plurality of instructions executable by the one or more processors, the instructions, when executed by the one or more processors, causing the one or more processors to perform a set of operations, including: Receive a request for allocating graphics processing unit (GPU) resources for performing an operation, wherein the request includes metadata identifying a client identifier (ID) associated with the client, the latency of the operation, and the throughput. Determine the resource limits for performing operations based on metadata; Obtain at least one attribute associated with each of a plurality of GPU resources available for assignment in the system, wherein the at least one attribute indicates the capacity of the corresponding GPU resource; The at least one attribute associated with each GPU resource in relation to resource constraint analysis; Based on the analysis, a set of GPU resources is identified from the plurality of GPU resources; A dedicated AI cluster is generated by patching the aforementioned set of GPU resources into a single cluster, wherein the dedicated AI cluster retains a portion of the computing capacity of the computing system for a period of time; and Assign a dedicated AI cluster to a client associated with a client ID.
9. The system of claim 8, wherein the set of operations further comprises: Before allocating a dedicated AI cluster, requests are authenticated based on the client ID associated with the client, using a private key extracted from the asymmetric key pair associated with the client ID.
10. The system of claim 8, wherein the set of operations further comprises: The set of performance parameters corresponding to each GPU resource in the set of GPU resources is compared with a predefined set of performance parameters. Based on the comparison, an anomaly is determined in a first GPU resource within the set of GPU resources, wherein the anomaly indicates that the set of performance parameters deviates from a predefined set of performance parameters; and Within a dedicated AI cluster, the first GPU resource is replaced with a second GPU resource, wherein the hash value of the second GPU resource is the same as the hash value of the first GPU resource.
11. The system of claim 8, wherein the set of operations further comprises: Determine the pre-approved quota associated with the request; Determine whether the pre-approved quota exceeds the predefined request limit corresponding to the client ID; as well as The request is blocked based on the determination that the pre-approved quota exceeds the predefined request limit.
12. The system of claim 8, wherein the set of operations further comprises: The type of operation is determined based on the request; as well as The set of GPU resources is selected from one of multiple nodes or a single node based on the type of operation to generate a dedicated AI cluster.
13. The system of claim 12, wherein the set of operations further includes fine-tuning operations based on determining the request: When the request indicates a fine-tuning operation, the data model to be fine-tuned is obtained; and Fine-tuning logic is performed on the data model using a dedicated AI cluster, which is generated using the set of GPU resources selected from the single node.
14. The system of claim 8, wherein the set of operations further comprises: Identify at least one underutilized GPU resource from the set of GPU resources in the dedicated AI cluster; as well as In response to the identification of the at least one GPU resource, a dummy operation is performed on the at least one GPU resource, wherein the dummy operation is identical to an operation performed on the at least one GPU resource.
15. A non-transitory computer-readable medium storing a plurality of instructions executable by one or more processors to cause the one or more processors to perform a set of operations, comprising: Receive a request for allocating graphics processing unit (GPU) resources for performing an operation, wherein the request includes metadata identifying a client identifier (ID) associated with the client, the latency of the operation, and the throughput. Determine the resource limits for performing operations based on metadata; Obtain at least one attribute associated with each of a plurality of GPU resources available for assignment in a computing system, wherein the at least one attribute indicates the capacity of the corresponding GPU resource; The at least one attribute associated with each GPU resource in relation to resource constraint analysis; Based on the analysis, a set of GPU resources is identified from the plurality of GPU resources; A dedicated AI cluster is generated by patching the set of GPU resources into a single cluster, wherein the dedicated AI cluster retains a portion of the computing capacity of the computing system for a period of time. as well as Assign a dedicated AI cluster to a client associated with a client ID.
16. The non-transitory computer-readable medium of claim 15, wherein the set of operations further comprises: Before allocating a dedicated AI cluster, requests are authenticated based on the client ID associated with the client, using a private key extracted from the asymmetric key pair associated with the client ID.
17. The non-transitory computer-readable medium of claim 15, wherein the set of operations further comprises: The set of performance parameters corresponding to each GPU resource in the set of GPU resources is compared with a predefined set of performance parameters. Based on the comparison, an anomaly is determined in a first GPU resource within the set of GPU resources, wherein the anomaly indicates that the set of performance parameters deviates from a predefined set of performance parameters; and Within a dedicated AI cluster, the first GPU resource is replaced with a second GPU resource, wherein the hash value of the second GPU resource is the same as the hash value of the first GPU resource.
18. The non-transitory computer-readable medium of claim 15, wherein the set of operations further comprises: Determine the pre-approved quota associated with the request; Determine whether the pre-approved quota exceeds the predefined request limit corresponding to the client ID; as well as The request is blocked based on the determination that the pre-approved quota exceeds the predefined request limit.
19. The non-transitory computer-readable medium of claim 15, wherein said set of operations further comprises: The type of operation is determined based on the request; as well as The set of GPU resources is selected from one of multiple nodes or a single node based on the type of operation to generate a dedicated AI cluster.
20. The non-transitory computer-readable medium of claim 15, wherein said set of operations further comprises: Identify at least one underutilized GPU resource from the set of GPU resources in the dedicated AI cluster; as well as In response to the identification of the at least one GPU resource, a dummy operation is performed on the at least one GPU resource, wherein the dummy operation is identical to an operation performed on the at least one GPU resource.
Citation Information
Patent Citations
Secure generative-artificial intelligence platform integration on a cloud service
US20250097013A1