AI TOOLS IN A PRIVATE CLOUD ENVIRONMENT
Pre-configured server racks with AI tools simplify the deployment and management of AI tools in private clouds by providing secure, scalable, and efficient integration and updates through remote cloud management.
Patent Information
- Application Number
- DE102025110413
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-05
- Filing Date
- 2025-03-18
- Publication Date
- 2026-03-05
AI Technical Summary
Managing and updating AI tools in a private on-premises cloud environment is complex due to the integration of hardware and software from multiple disparate sources, requiring streamlined setup and update processes.
A pre-configured server rack with AI tools is delivered to the customer's site, simplifying the setup by establishing network and user access, and enabling remote management through a cloud management system for seamless integration and updates.
Simplifies the deployment and management of AI tools in a private cloud by reducing setup complexity and enabling efficient, secure, and scalable integration and updates.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Artificial intelligence (AI) is a method in which a non-human system learns from experience through machine learning and imitates human intelligent behavior. AI thus offers powerful tools that can be used for the efficient processing and / or analysis of large datasets. AI tools can be deployed on suitable computing equipment / hardware, such as in a cloud computing system.
[0002] Cloud computing systems can be implemented in numerous different ways, including public and private clouds. Public clouds can be used where users from the (often subscribing) public have access to cloud services, while private clouds are restricted to one or more organizations. The simplest private cloud can be managed by a single organization for internal use, without providing services to others. One type of private cloud is the on-premises cloud, where the managing entity controls or manages all the hardware and software implemented in the private cloud at its own location. DRAWINGS
[0003] Features, aspects and advantages of the present disclosure will be better understood if the following detailed description is read with reference to the accompanying drawings, in which identical symbols represent identical parts in the drawings, wherein: Fig. Figure 1 is a diagram showing a private cloud system (on-premises) with pre-configured AI tools, managed via a remote cloud management system in accordance with the aspects of this disclosure; Fig. Figure 2 is a flowchart that illustrates a process for accessing, updating, and / or using AI tools in the private cloud system of Fig. 1 illustrated in accordance with aspects of the present revelation Fig. 3 is an example screen of an interface that was created using the method of Fig. 2 can be used in accordance with aspects of the present revelation; Fig. Figure 4 is a diagram illustrating the process of adding new racks to the private cloud system of Fig. 1 in accordance with aspects of the present revelation; Fig. 5 is a sample screen that appears in a user interface as part of the process of Fig. 4 can be presented in accordance with aspects of the present revelation Fig. 6 is an example of a network setup screen that appears in the user interface as part of the process of Fig. 4 can be indicated in accordance with aspects of the present disclosure; Fig. 7 is an example of a virtualization setup screen that appears in the user interface as part of the process of Fig. 4 can be indicated in accordance with aspects of the present disclosure; Fig. 8 is an example of a control plane setup screen that appears on the user interface as part of the procedure of Fig. 4 can be indicated according to the aspects of the present revelation; Fig. 9 is an example of a screen for setting up a work node, which is part of the process of the user interface. Fig. 4 can be indicated in accordance with aspects of the present disclosure; Fig. 10 is an example of a summary setup screen that appears in the user interface as part of the process of Fig. 4 can be indicated in accordance with aspects of the present disclosure; Fig. 11 is a KL stack that, in accordance with aspects of this disclosure, is deployed to the private cloud system of Fig. 1 can be recorded; Fig. 12 is an example of a status screen that can be displayed in the user interface and represents one or more components of the private cloud system of Fig. 1 shows, in accordance with aspects of the present revelation; Fig. 13 is an example of a software details screen that appears in the user interface in response to the selection of a component of the status screen. Fig. 12 may be indicated in accordance with aspects of the present disclosure; Fig. 14 is a pre-check screen that can be displayed in the user interface to perform a software pre-check for one or more components of the private cloud system of Fig. 1 to carry out in accordance with the aspects of the present revelation; Fig. 15 is a download screen that can be displayed in the user interface to download software updates for one or more components of the private cloud system of Fig. 1 to download, in accordance with aspects of the present revelation; Fig. 16 is a flowchart that depicts a preliminary review process for the private cloud system of Fig. 1 illustrated according to the aspects of the present revelation; Fig. Figure 17 is a flowchart that illustrates a download update process for the private cloud system of Fig. 1 illustrated in accordance with aspects of the present revelation Fig. Figure 18 is a flowchart illustrating a unified process for the private cloud system in accordance with the aspects of the present disclosure; and Fig. Figure 19 is a flowchart that illustrates an update process for the private cloud system in accordance with the aspects of this disclosure. DETAILED DESCRIPTION
[0004] One or more specific aspects of this disclosure are described below. In an effort to describe these aspects concisely, not all features of an actual implementation may be described in the specification. It should be noted that, as with any engineering or design project, developing such a specific implementation involves numerous implementation-specific decisions to achieve the developers' specific goals, such as adhering to system-related and business-related constraints, which may vary from one implementation to another. Furthermore, it should be understood that while such development effort may be complex and time-consuming, it is nevertheless a routine design, fabrication, and manufacturing undertaking for professionals who benefit from this disclosure.
[0005] When introducing elements of various aspects of the present revelation, the articles "a," "an," "the," and "said" are to indicate that one or more of the elements are present. The terms "comprehensive," "including," and "with" are to be understood as comprehensive and mean that there may be other elements besides those listed.
[0006] The implementations presented here relate to techniques for setting up and updating server racks to implement artificial intelligence (AI) capabilities in private or on-premises clouds managed via remote cloud management. With the increasing availability of AI software and AI-oriented hardware from numerous sources, the possibilities for implementing AI solutions are becoming virtually limitless. Since the management entity is responsible for all hardware and software, including AI tools, deciding which hardware and software to deploy in the on-premises cloud can be complex. Managing the hardware and software can be even more challenging, as resources originate from multiple disparate sources.Furthermore, keeping software and / or hardware up to date can be challenging when managing an entire private on-premises cloud, especially when updates originate from multiple sources. Current techniques simplify the process for customers to integrate new AI-enabled hardware into their private on-premises clouds by streamlining the setup process through the delivery of a server rack with AI pre-installed. This allows a customer to select a configuration and / or use case that includes hardware and software suitable for their desired AI applications.The integrated hardware and software can be configured to a standard that utilizes Infrastructure as a Service (IaaS) to employ a subscription model for acquiring infrastructure through regular payments and / or other suitable computing model recommendations and / or sizing to meet the requirements of the model / use case / token(s). Once the standardized option has been selected and purchased, the cloud manager / vendor can arrange for the rack(s) to be delivered to the customer's site within a specified timeframe (e.g., 8 hours).
[0007] When the rack(s) is / are delivered to the customer's site, the customer and / or the manager / manufacturer can physically install and power on the rack(s). As mentioned previously, the racks may be pre-configured with AI tools, such as manufacturer-specific and / or open-source tools, models, and / or other AI support. Since the AI tools are already pre-configured in the rack(s), only the connection between the rack(s) and the remote cloud management needs to be established. This simplifies the setup by 1) configuring network access for the rack(s), 2) configuring user access for users accessing the rack(s) / on-premises private cloud, and 3) establishing the connection to the remote cloud for managing the on-premises private cloud. This setup may use default values (e.g.,This includes the location, which can be changed during setup and / or after arrival. In some implementations, if the new rack(s) is / are added to one or more racks already managed by the remote cloud, the new rack(s) can be added via the remote cloud and connected accordingly. The remote cloud can enable workload migration and / or capacity expansion to meet evolving consumer AI demands.
[0008] After installation and / or as part of the installation of the rack(s), software updates may be available for the rack(s) that must be applied. Remote cloud management can also simplify the update process. For example, the AI tools can be updated remotely via the cloud with a simple update (e.g., a single mouse click). The update can be initiated through a user interface that interacts with and / or is part of the remote cloud management system. The update can be performed in a single operation that includes downloading and applying the update, or it can be performed as separate operations, with the download and application occurring at a different time.Before downloading and / or applying updates, the remote cloud management system can check which updates are available in a cloud repository for the AI tools and are suitable for the components of the on-premises private cloud. The remote cloud management system can also verify that the system is in a non-degraded state and that no components have failed. The updates can be performed in a specific sequence, beginning with a data service connector update, which is applied to update a data service connector between the on-premises private cloud and the remote cloud management system. The updated data service connector is then used for subsequent parts of the update, such as updating a control plane, BIOS / firmware / OS, virtualization components (e.g., virtual machines (VMs) and their components), and / or other parts of the on-premises private cloud.
[0009] These installation and update techniques can represent a general improvement for consumers using private on-premises clouds, simplifying the deployment of one pre-configured rack for a first type of AI workload and another pre-configured rack for a second type of AI workload. Remote cloud management can also install updates for the entire AI stack, including the control plane, firmware, storage, operating system, AI software / nodes, etc., using a simplified update process with a specific order for applying the updates.
[0010] Against this background Fig. Figure 1 shows a diagram depicting a System 100 implementing cloud-based AI tools. System 100 includes a private cloud System 102, which is at least partially managed by a customer organization. The private cloud System 102 can be an on-premises cloud system, where the hardware connected to the cloud services implemented through the private cloud System 102 is at least partially located on-premises at at least one of the customer organization's physical locations. Customers may prefer such on-premises cloud solutions in situations where they want to keep at least some of the hardware and / or their data on-premises. This freedom to manage their own hardware can also allow the customer to create a customized configuration with specifically selected hardware and / or software solutions.However, the use of such customized configurations can also complicate the provisioning of new hardware, new software, the management of current software or hardware, and / or the updating of current software to newer versions, as these are customized solutions from multiple manufacturers and / or vendors with different sources for obtaining configurations and / or settings.
[0011] The Private Cloud System 102 can utilize a container organization system that acts as its operating system. This system can cluster one or more computers, including virtual machines (VMs) and / or bare-metal servers (BMs), to run workloads in containers. The container orchestration system can include, for example, Kubernetes®, an open-source container orchestration system, and / or any other suitable container orchestration system. The system can work with one or more container runtimes, such as HPE Ezmeral®, Docker®, Podman®, Kubernetes Container Runtime Interface (CRI-O), Containerd®, rkt, and / or other suitable runtimes, to execute the workloads within the containers.
[0012] Since the private cloud system 102 is implemented using computers with one or more servers (e.g., BMs) together with one or more VMs, the private cloud system 102 comprises a combination of hardware elements (e.g., processors) and software elements, including tangible, non-transient, and machine-readable media, such as memory 103. Memory 103 can include any suitable manufacturing object for storing data and / or executable instructions, such as random access memory, read-only memory, rewritable flash memory, hard disks, and optical media. Furthermore, programs (e.g.,(using container runtimes and / or the container orchestration system), which are coded in such a computer program product, also contain instructions that can be executed by one or more processors of the Private Cloud System 102 to enable the containers of the Private Cloud System 102 to run workloads in the containers.
[0013] As shown, the private cloud system 102 comprises worker nodes 104 and control nodes 106. Applications, such as the KL software 107, run in a cluster on the worker nodes 104. The worker nodes 104 process data and handle networking for the private cloud system 102. The worker nodes 104 host application containers in groups, which in turn run one or more containers. The worker nodes 104 are subordinate to the control nodes 106. The control nodes 106 manage the operation of the private cloud system 102 by controlling when and which of the worker nodes 104 run the containers. In other words, the control nodes 106 contain a scheduler that communicates with the worker nodes 104 to schedule container workloads. The scheduler can consider the availability of computers, such as CPU and memory, as well as the application requirements (e.g.,AI software 107). The work nodes 104 can use node-level agents that track resource consumption and facilitate the fulfillment of schedule assignments for the work nodes 104, ensuring that the work nodes 104 perform the assigned tasks.
[0014] The control nodes 106 manage the communication and control of the work nodes 104 and can include an API (Application Programming Interface) server as well as the storage of configuration and status data. In some embodiments, a control plane (e.g., the AI software control plane 110) can run across multiple control nodes 106 to ensure redundancy.
[0015] The control nodes 106 can also communicate with external services by implementing a data service connector 108. The control nodes 106 can also include the AI software control plane 110, which is used to control (e.g., schedule) workloads in application containers of the AI software 107. As explained in more detail below, the AI software 107 can include software provided by one or more different vendors, such as Hewlett Packard Enterprise Company (HPE), NVIDIA, open-source partners, and / or other tools, which can be deployed to the private cloud system 102 via a remote cloud management system 112.
[0016] System 100 comprises the remote cloud management system 112, which is used to manage the private cloud system 102 over one or more networks 114 (e.g., the internet), and the data service connector 108, which is implemented in the control nodes 106 of the private cloud system 102. The remote cloud management system 112 is physically separate from the private cloud system 102, but it can be used to perform remote management of the private cloud system 102 via a tunnel connection 116. The remote cloud management system 112 is physically separate from the private cloud system 102 in that it can be implemented on different computers / servers at different locations than those used to implement the private cloud system 102.The tunnel connector 116 is coupled with the data service connector 108 to establish a secure (“encrypted”) remote connection over one or more networks 114 so that information between the remote cloud management system 112 and the private cloud system 102 remains secure and confidential.
[0017] The remote cloud management system 112 uses a combination of hardware and software to implement a private cloud infrastructure orchestrator 118, a private cloud resource orchestrator 120, a private cloud AI platform orchestrator 122, and a private cloud AI API / user interface (UI) 124.
[0018] The Private Cloud Infrastructure Orchestrator 118 orchestrates operations related to setting up the infrastructure of the Private Cloud System 102 (e.g., setting up the control nodes 106, connecting the Private Cloud System 102 to the remote cloud management system 112, etc.). The Private Cloud Infrastructure Orchestrator 118 also manages software updates, inventories for the Private Cloud System 102, network management for the Private Cloud System 102, monitors / controls the measurement of resources (e.g., processing power, memory, and / or electricity) used by the infrastructure during operation using the Private Cloud System 102, and / or generates dashboards to display information about the infrastructure of the Private Cloud System 102.
[0019] The Private Cloud Resource Orchestrator 120 orchestrates operations using components such as VMs, BMs, and / or the container orchestration system of the Private Cloud System 102. For example, the Private Cloud Resource Orchestrator 120 can be used to provision and manage the VMs, BMs, and / or the container orchestration system (e.g., Kubernetes) of the Private Cloud System 102.
[0020] The Private Cloud AI Platform Orchestrator 122 can be used to manage the AI platform using the AI software 107. For example, the Private Cloud AI Platform Orchestrator 122 can deploy and / or extend AI applications that are installed and available in the AI software 107 of the Private Cloud system 102. The Private Cloud AI Platform Orchestrator 122 can perform such deployment and / or extension of the AI software 107 inventory via the tunnel between the tunnel connector 116 and the data service connector 108, and / or via a sideband connection 126 between the Private Cloud AI API / UI 124 and the AI software 107. For example, the sideband connection 126 can be a secure tunnel that is separate from the tunnel between the data service connector 108 and the tunnel connector 116.
[0021] The private cloud API / UI 124 can provide a user interface to enable remote management of the AI software 107 in the private cloud system 102 using APIs. For example, a user can log in to the remote cloud management system 112 and use APIs, such as REST (Representational State Transfer) APIs, to control changes in the private cloud system 102 via the private cloud infrastructure orchestrator 118 and / or the private cloud resource orchestrator 120. The private cloud AI API / UI 124 includes an infrastructure manager 128 that manages infrastructure changes using API calls via the private cloud infrastructure orchestrator 118 to make changes to the management of the infrastructure and / or changes to the infrastructure itself.The Private Cloud AI API / UI 124 also includes an AI software platform manager 130, which manages changes to the AI software 107 and / or the AI software control plane 110 via the Private Cloud resource orchestrator 120 and / or the Private Cloud AI platform orchestrator 122. Additionally or alternatively, the Private Cloud AI API / UI 124 can make changes to the AI software 107 using the sideband connection 126 via an AI interface 132, which can be used to authenticate and / or encrypt the sideband connection 126 between the Private Cloud AI API / UI 124 and the Private Cloud system 102.
[0022] The remote cloud management system 112 can include additional components to support the remote management of the private cloud system 102. For example, the remote cloud management system 112 can include a software catalog 134 that stores and / or references available software for use in the private cloud system 102. The software catalog 134 can thus determine which software is suitable and available for a specific hardware and / or software configuration of the private cloud system 102. If the organization / user connected to the private cloud system 102 has subscribed to AI services (e.g., from a remote cloud management system 112 provider and / or third-party providers), the software catalog 134 can provide appropriate AI tools for installation / use in the private cloud system 102.In addition to or as an alternative to subscription-based filtering, the software catalog 134 can be filtered based on whether the organization / user connected to the private cloud system 102 has met certain requirements before providing at least some AI services. For example, the software catalog 134 can refrain from displaying AI tools from at least some vendors (e.g., third-party providers) until the organization / user has indicated that they agree to an agreement with the respective vendors. This agreement could be, for example, an End User License Agreement (EULA) and / or other license agreements.
[0023] The remote cloud management system 112 can include an auditor 138, which can be implemented using hardware and / or software, to allow a user / organization to view metrics related to current and / or historical workloads of the private cloud system 102. The remote cloud management system 112 can also include an authorizer 140, which completes the authorization process for each user attempting to access the remote cloud management system 112 and / or the private cloud system 102 before granting such access.
[0024] In the company, a user can select which AI tools can be used in the private cloud system 102. Fig. Figure 2 shows a process 150 for deploying AI tools in the private cloud system 102 using the remote cloud management system 112. The remote cloud management system 112 receives login credentials for a user via the private cloud A / Ul 124 (block 152). The remote cloud management system 112 uses the authorizer 140 to check whether the user is authorized to deploy, modify, and / or use the AI software 107 in the private cloud system 102 (block 154). If the login credentials are invalid or the user is not authorized to access, use, and / or modify the AI software 107, the credentials are not authorized, and the remote cloud management system 112 can request the credentials again. In some embodiments, the remote cloud management system 112 may only receive a limited number of login attempts (e.g.,1, 2, 3, 4 or more times), before locking the account, logging the failed authorization check and / or notifying an administrator of the failed authorization check for the login credentials.
[0025] If authorization is successful, the Remote Cloud Management System presents 112 available solutions in the Private Cloud AI API / UI 124 (Block 155). Fig. Figure 3, for example, shows a screen 160 that can be displayed in the private cloud Al / Ul 124 and shows ready-to-use Kl-Tools 162 that are pre-configured in racks of the private cloud system 102, as indicated by an initialized tag for the respective status 164. These ready-to-use Kl-Tools 162 can be deployed via the remote cloud management system 112 using a deployment button 166.
[0026] For AI tools 168 that are already provisioned and indicated by a provisioning indicator for their status 164, no provisioning button 166 is displayed, and the provisioned AI tools 168 can be opened, edited, or run by clicking on them. In some embodiments, AI tools (e.g., AI solution accelerators) that are not initialized or provisioned can be added from the software catalog 134 using an Add button 169 based on their suitability for the private cloud system 102 and / or based on subscriptions 136 available for the credentials used to log in to the remote cloud management system 112.
[0027] In some embodiments, the deployable KL tools 162 and the deployed KL tools 168 may contain a description and / or tags that specify the objectives, scope, platforms, programming languages and / or other details about the respective KL tools and / or their use.
[0028] Back to Fig. 2: One of the presented selections is received via the private Cloud AI API / UI 124 (Block 156). For example, one of the deployable AI tools 162, one of the deployed AI tools 168, and / or the "Add" button 169 could be the received selection made via screen 160 of the private Cloud AI API / UI 124. If the selection corresponds to a new deployment (or the addition of an AI tool via the "Add" button 169) (Block 157), the Remote Cloud Management System 112 deploys a new AI tool (Block 158). In some implementations, deploying the new AI tool may involve modifying the AI workloads in the work nodes 104 to accommodate the newly deployed AI tool. Deployment may also include displaying a status of the deployment before, during, and / or after its completion. For example, screen 160 may be updated to display status information, such as... B."Deployed", "Deploying" with an indicator for the percentage of completeness and / or other suitable indicators of the deployment status.
[0029] If the selected solution is already deployed, the remote cloud management system 112 can perform an operation (block 159). This operation could include, for example, adjusting the workload of the selected AI tool, executing a process with the selected AI tool, pausing the execution of the selected AI tool, displaying data results from the execution of the AI tool, executing the AI tool with different input data, and / or other suitable operations that can utilize the selected AI tool.
[0030] In some situations, adding or deploying new AI tools can consume a significant portion of the Private Cloud System 102. In this case, or during the initial setup of the Private Cloud System 102, hardware must be added to the Private Cloud System 102 to implement the AI tools. However, the user or AI administrator who completes this process can be a different user category (e.g., Cloud Administrator) with the necessary authority / capability to add new hardware to the Private Cloud System 102. Fig. Figure 4 shows a process 170 for acquiring and deploying hardware in the private cloud system 102 with pre-configured AI tools. The remote cloud management system 112 can present options for hardware to be deployed via the private cloud AI / UL 124 (block 172). The presentation of the options can occur in response to an authentication check, as described above. Fig. 2 described. Fig. Figure 5 shows a screen 200 that can be displayed in the private cloud AI API / UI 124. Screen 200 can be displayed when an authorized user requests options for creating and / or extending the private cloud system 102 with new and / or replacement hardware. Screen 200 contains a number of different configurations 202, 204, 206, and 208 that can be added to the private cloud system 102. In some implementations, configurations 202, 204, 206, and 208 may be general options suitable for implementing AI operations. Additionally or alternatively, configurations 202, 204, 206, and / or 208 may be recommendations based on the hardware already present in the private cloud system and / or information provided by the cloud administrator or the cloud administrator's organization.For example, the recommended configurations may favor the use of configurations similar to those already deployed in the Private Cloud System 102. Additionally or alternatively, the recommended configurations may be based on information about the types of AI operations expected, the compute requirements for the expected AI operations, the storage requirements for the expected AI operations, the network requirements for the expected AI operations, and / or an energy budget for the new additions.
[0031] Each of the configurations 202, 204, 206, and 208 can include corresponding AI functions, such as inferencing, retrieval-augmented generation (RAG), model fine-tuning, other AI functions, or a combination thereof. RAG is an AI framework that uses conventional information retrieval systems, such as databases in memory. RAG optimizes large language models (LLMs) by allowing them to access current information from the curated databases and incorporate it into their responses and analyses. This additional knowledge can enable more accurate, relevant, and up-to-date analysis and / or preference-based suggestions. Model fine-tuning uses more training examples than few-shot learning by employing a model (e.g., a few-shot learning-based model) and performing iterative supervised or unsupervised training stages to fine-tune the model.Configurations 202, 204, 206, and / or 208 can be assigned a suitability flag 210, indicating the operations for which the respective configurations are best suited. Additionally or alternatively, the cloud administrator can select the expected AI functions in the private cloud AI / UL 124 at the time of option review, or they can be preconfigured and saved in the remote cloud management system settings 112. Alternatively, the cloud administrator can specify which configuration (e.g., small, medium, large, or extra-large) they require.
[0032] Screen 200 can also provide information about the various configurations 202, 204, 206, and / or 208. For example, screen 200 can display information about the computing components 212 for configurations 202, 204, 206, and / or 208. These computing components can include, for example, graphics processing units (GPUs), central processing units (CPUs), application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), other processors suitable for use in AI calculations, or a combination thereof.
[0033] Screen 200 can also display information about the memory components 214 for configurations 202, 204, 206, and / or 208. The memory component 214 display can include the amount of memory and / or storage space. In some configurations, the type of memory (e.g., latency and / or frequency) can be selected via the memory component 214 display or through another interface. The memory component 214 information can specify a base amount of memory and / or one or more upgraded memory amounts that can be selected for use in the rack(s) when installed in the private cloud system 102. Additionally or alternatively, the memory component 214 information can specify a maximum amount of memory that can be added to the configuration at a later time.
[0034] Screen 200 can also display network indicators 216 for configurations 202, 204, 206, and / or 208. The network indicators 216 can display the data transfer rates for the respective configurations. Screen 200 can also include a power consumption indicator 218, which displays an estimated power consumption for the respective configurations.
[0035] Back to Fig. 4: The private cloud AI API / UI 124 can receive a selection from one of the presented options and order the corresponding configuration (Block 174). In some embodiments, the selection and ordering can be done offline (e.g., by telephone and / or using paper or softcopy order forms), with a representative of the remote cloud management system provider 112 performing the selection and ordering. The provider then supplies the ordered rack(s) with pre-installed AI tools and / or access to AI tools. The provider and / or the ordering organization then sets up the hardware (Block 175). The following are described below. Fig. 6-9 refer to this hardware setup.
[0036] Before, during, and / or after ordering the hardware, the cloud administrator adds users and / or user roles to the rack(s) (Block 176). For example, the cloud administrator can add the following: Fig. Add the login credentials used to set up, modify, and / or use the AI software 107 in the private cloud system 102 via the newly configured hardware. This allows some users to be permitted to modify the AI tools (e.g., to start fine-tuning), while other users can only be granted permission to access or view the results of AI operations. The cloud administrator can then use the private cloud API / UI 124 to manage the private cloud system 102, including monitoring workloads and / or other AI dashboards available in the private cloud API / UI 124. From this point, the cloud administrator can extend the job / private cloud system 102 to include additional available hardware and / or software components (Block 180). For example, the private cloud API / UI 124 can return to displaying options in Block 172.After the management / monitoring step, the cloud administrator can use the Private Cloud AI API / UI 124 to update the AI stack in the rack(s) in Private Cloud system 102 (block 182). The following are explained below. Fig. 11-14 refer to this update mechanism. Hardware setup
[0037] As already mentioned, this refers to Fig. 6 on the hardware setup of racks in the Private Cloud System 102 using the Private Cloud Al / Ul 124 and / or another part of the Remote Cloud Management System 112 and / or the Private Cloud System 102. As such, it shows Fig. 6. A screen 240, which the Private Cloud AI API / UI 124 and / or other parts of the Remote Cloud Management System 112 may present to the cloud administrator as part of the setup. The screen 240 contains a progress indicator that includes an infrastructure part 242 and a private cloud AI part 244. As shown in the progress indicator, the infrastructure is set up before the private cloud AI (e.g., the AI software 107 of the private cloud system 102) is set up; the infrastructure part 242 expands while the private cloud AI part 244 collapses. Within the infrastructure part 242, the progress indicator includes a network part 246, which is used to configure the network as part of the setup. The network part 246 is printed in bold and / or otherwise highlighted to indicate that the network is currently being set up.During network configuration, screen 240 displays text 248 that can provide instructions for completing the network configuration section of the setup. Screen 240 also includes a server management network input 250, which allows the input of an IP address for the remote cloud management system 112 to establish a connection between the remote cloud management system 112 and the private cloud system 102.
[0038] Screen 240 also includes an input 252 for an integrated lights-out (iLO) management network, which allows the input of an IP address for an iLO management system used to remotely and securely manage, simplify, and / or automate server operations. The iLO management network input 252 can have the same address as the remote cloud management system 112, which is specified in the server management network input 250. Accordingly, screen 240 includes a selector 254, which the cloud administrator can use to specify that both the remote cloud management system 112 and the iLO management system have the same IP address. If selector 254 indicates that the IP address is the same for both networks, the iLO management network input 252 and / or the server management network input 250 can be hidden or otherwise disabled.When using the same or different IP addresses for the iLO management network and the remote cloud management system 112, screen 240 includes an iLO subnet input 256, which allows the cloud administrator to specify a particular subnet mask and gateway for the iLO management network. Since the private cloud system 102 also contains data to be used in the AI operations in storage 103 and / or elsewhere, screen 240 allows the input of an IP address 258, a data subnet mask 260, and / or a gateway 262 for accessing the data to be used in the AI operations in the private cloud system 102.
[0039] Once the network has been configured on screen 240, a button 263 can be selected to proceed to section 264 of the control plane server progress indicator. As with the management networks, locations (e.g., IP addresses, subnet masks, and / or gateways), serial numbers, models, CPU families, and / or other information about the control plane servers can be entered into the private cloud AI API / UI 124. Once the control plane server location has been determined and / or a corresponding button has been selected, virtualization is also set up using a virtualization section 266 of the progress indicator. Once the infrastructure setup is complete, a summary section 268 of the progress indicator can be selected (e.g., using a "Next" button in the virtualization section of the setup).
[0040] Fig. Figure 7 is a diagram of a screen 280, which may correspond to the virtualization section 266 of the progress tracker. As such, screen 280 may be displayed using the Private Cloud AI API / UI 124 and / or another part of the Remote Cloud Management System 112 and / or the Private Cloud System 102. Screen 280 contains a hostname input 282, which is configured to receive a hostname for server management software to control virtual machine environments.
[0041] For example, the server management software can include any suitable server management software, such as NVIDIA Virtual GPU (vGPU) software, VMware vCenter, and / or other suitable server management software packages. Screen 280 also includes a credential input 284, which is used to enter hypervisor (HV) credentials for the server management software, the remote cloud management system 112, and / or the private cloud system 102. The credentials can include, for example, single sign-on (SSO) credentials, which are digital credentials that can be used for multiple applications or websites, such as the server management software and / or other locations within the remote cloud management system 112 and / or the private cloud system 102. The credential input 284 may include a drop-down menu and / or a pop-up menu that allows the selection of credentials.
[0042] Screen 280 also includes an input 286 for the hypervisor (HV) root credentials, which allows the entry of the HV root credentials via manual input, a drop-down menu, a pop-up menu, or another mechanism for entering the credentials to be used for authentication with the hypervisor. The hypervisor is installed on the rack(s) and used to partition the rack(s) into VMs. Similarly, an input 288 for iLO admin credentials allows the entry of iLO admin credentials for authentication with the iLO management software described earlier. Screen 280 may also display an acceptance button 290 along with a statement, which accepts one or more agreements with vendors of the various application software (e.g., AI software 107) and / or management tools.The statement may contain a link to the various agreements, including those between different providers. Once the virtualization data has been entered, a button (292) can be selected to proceed.
[0043] Once the infrastructure setup is complete, the progress indicator can move to Private Cloud AI Part 244, as shown in screen 300 of Fig. Figure 8 is shown. On screen 300, the Private Cloud AI part 244 is expanded, while the Infrastructure part 242 is collapsed, indicating that the setup has moved to the Private Cloud AI part 244 of the setup. As shown, the Private Cloud AI setup can be broken down into a control plane setup, a worker node setup, and a summary, as shown by screens 302, 304, and 306. Because the control plane is currently being configured using screen 300, screen 300 contains information 302, which is bold, italic, underlined, and / or otherwise highlighted, while information 304 and 306 are less highlighted. Screen 300 contains a control plane VM name prefix input 308, which is used to enter a name prefix for one or more (for example, 3 for redundancy) control plane VMs.Screen 300 also includes a key input 310 for entering a key that is used to access VMs and can be selected from a list of stored secrets and / or uploaded to the private cloud Al API / UI 124.
[0044] Screen 300 also contains network details for management, storage, or worker nodes by selecting a network from a network menu 312. The control plane VMs then receive a start IP address via a start IP input 316, which is used to specify the first IP address for the first VM. Subsequent VMs receive the next available addresses. The network details also allow specifying a cluster IP address via a cluster IP input 314 and an ingress IP address via an ingress IP input 316. Once the information has been entered, the next step of the AI setup regarding the worker nodes can be accessed by activating a button 318 after completing the data entry.
[0045] As soon as the next key 318 is pressed, the private cloud Al API / UI 124 causes the display of screen 320, as shown in Fig. Figure 9 is shown. Screen 320 contains the highlighted entry 304 and the extensions to entry 304, including entries 326 and 328. Entry 326 corresponds to the servers for the work nodes 104 and can be used to enter an IP address, name, serial number, device type, processor type, amount of memory, and / or other useful information about the servers implementing the work nodes 104. Entry 328 can be used to configure the work nodes 104. As indicated by the highlighting of entry 328, the setup for configuring the work nodes 104 is in progress. To facilitate the configuration of the work nodes 104, screen 320 includes an input 330 for the work node name, allowing each work node 104 to be assigned a human-readable prefix.Screen 320 also includes key inputs 332 and 334, which allow the entry of keys for the respective access using iLO and a work node operating system, such as Red Hat Enterprise Linux (RHEL). The configuration may include an participation input to allow the specification of a partition of a multi-instance GPU (MIG) 336, if applicable. Screen 320 may also allow the entry of a start IP 338 for work nodes 104 for a specific number (e.g., 4) of work nodes 104, assigning each work node 104 the next available number. Once the work node 104 configuration has been entered, the setup can be continued by selecting the next button 340.
[0046] Once the configuration is complete, as described in Fig. As shown in Figure 10, a screen 350 is displayed using the Private Cloud AI / UL 124, which contains a summary of the setup of the Private Cloud AI part 244, as indicated by the emphasis of the display 306. The summary contains control plane details 352 about the control plane / control node 106. For example, the summary may include the VM access key, network name, subnet mask, gateway for the control nodes, names for each control node 106 and their individual IP addresses, or any combination thereof. The summary also includes information 354 about work nodes 104, such as an iLO access key, an operating system access key, names for the work nodes 104, IP addresses for the work nodes 104 in the management network, IP addresses for the work nodes 104 in the iLO network, and IP addresses for the work nodes 104 in a storage network 103.In some implementations, the summary may include a status indicator showing the configuration status, such as 0, 5, 10, 25, or more percent complete. Once the summary details have been confirmed, a submit button (356) can be selected. Alternatively, a back button (358) on this screen (350) or on previous screens can be used to navigate back and modify / update information about the control plane and / or work nodes as part of the setup.
[0047] Fig. Figure 11 shows an example AI stack 400, which includes a Data Service Connector (DSC) 402, such as the DSC 108 from Fig. 1, includes. The AI stack 400 also includes a hypervisor (HV) 404 (e.g., ESXi), a virtualization platform 406 (e.g., vSphere), server firmware 408 for control nodes 106, storage 410, network connectivity 412, an operating system 414 for the worker nodes 104 used to implement the AI software 107, server firmware 416 for the worker nodes 104, Kubernetes 418 and / or other container orchestration systems, and other AI tools 420. The AI tools 420 may include, for example, tools provided by the remote cloud management system provider 112, another provider (e.g., a third-party vendor such as NVIDIA), open-source tools, and / or other AI tools made available to the private cloud system 102 through the software catalog 134 of the Remote cloud management system 112 will be provided.The HV 404, the virtualization platform 406, the server firmware 408, and the storage 410 can be part of the control window for the private cloud system 102, as indicated in 422. The operating system 414, the server firmware 416, Kubernetes 418, and the KL tools 420 can be part of the worker nodes 104 and / or implemented with them, as indicated in 424. As one can imagine, the various sources of updates for the different objects in the KL stack 400 can make such updates more difficult and / or complicated. To simplify this process, the Private Cloud Infrastructure Orchestrator 118, the Private Cloud Resource Orchestrator 120, the Private Cloud AI Platform Orchestrator 122, and the Private Cloud AI API / UI 124 can be used to implement a simplified update involving multiple (e.g.,All objects in the Kl stack 400 can be updated in a single operation and / or with a single action (e.g., a click). In some embodiments, at least some components of the Al stack 400 can be updated through a separate operation. For example, the DSC 402 can be updated using a VM image stored in the remote cloud management system 112, which the customer can update directly via the remote cloud management system 112. Additionally or alternatively, the network connectivity 412 can be updated by the customer through a direct update. Software update
[0048] Fig. Figure 12 shows a screen 440, which can be displayed via the private cloud AI API / UI 124 and / or another part of the remote cloud management system 112. Screen 440 displays a list of one or more software records for which an update is available. In some embodiments, records can be displayed for all deployed AI tools, but in other embodiments, only records that have a potential update in a part of the AI tool's AI stack can be displayed. All other AI tools can be hidden.The data records each contain a name 442 of the object corresponding to the data record, a state status 444 indicating the known state of the object, a hypervisor cluster indicator 446 indicating which hypervisor cluster the object belongs to, a last update date indicator 448 indicating the date of the last update if the object has been previously updated, and an update status 450 indicating whether an update is available for the object. Interacting with the update status 450 by clicking, mouseover, or similar action displays a details window 452 showing the current versions of objects in the AI stack 400, such as...A current AI tool version, a current operating system version, a current hypervisor version, a current memory version, and / or a current firmware version, all of which are parts of and / or used by the current version of the object. Window 452 may contain a "Show Details" button 454, which displays more detailed information about the object. After selecting the "Show Details" button 454 and / or clicking Window 452, a software details screen, such as screen 470, may appear. Fig. 13, which displays a current version 472 and one or more update versions 474 and 476. The current version 472 may be marked to indicate that it is the current version. In some embodiments, one of the update versions 474 and 476 may contain a mark to indicate that the corresponding update version is the latest version (e.g., version 6.9.9).
[0049] After selecting a record on screen 440 of Fig. 12. A pre-check button 456 can be used to pre-check the compatibility of an update to the AI Stack 400 for the rack(s) with a proposed update. The pre-check button 456 can be used to check the compatibility and download of the update and then apply the update at a later time. However, if the update is to be applied during or after the update download without waiting for a later update start, the update can be applied via an update button 458, which causes the update and the pre-check to be confirmed sequentially in response to the selection of the update button 458. Additionally or alternatively, the pre-check button 456 can be used to download the update, and the update button 458 can be disabled until the update has been downloaded.
[0050] After selecting the pre-check button 456 in Fig. 12. The API / UI 124 of the private cloud can cause a pre-verification screen, such as screen 490 in Fig. 14. Screen 490 contains a title 492, which indicates that a selected hypervisor cluster has been chosen for a pre-check. A menu 494 allows you to select which version should be checked for an update. To start the pre-check, a button 496 is displayed. Selecting this button instructs the private Cloud Al API / UI 124 to pre-check and / or download the selected update. If you do not want to start the pre-check, a cancel button 498 can be selected to return to a previous screen without initiating the pre-check of the selected update.
[0051] After selecting the refresh button 458 in Fig. 12. The Private Cloud AI API / UI 124 can display an update screen, such as screen 510 in Fig. 15. Screen 510 contains a title 512 indicating that a selected hypervisor cluster has been chosen for update. A menu 514 allows you to select which version to update. To begin the update, a button 516 appears. Selecting this button instructs the private Cloud AI API / UI 124 to download and / or distribute the selected update. If you do not want to start the update, you can select a cancel button 518 to return to a previous screen without deploying the selected update.
[0052] Fig. Figure 16 is a flowchart of a software pre-test process 550, showing the exchange of operations and / or data between components of the remote cloud management system 112 and / or the private cloud system 102 as part of a software pre-test initiated with button 496 of Fig. 14 can be initiated. A client (552) can be an application that is implemented in and / or presented through the private cloud AI API / UI (124) and that the customer / user / organization can access. A gateway (GW) (554) can be part of a container orchestration system used for load balancing workloads. For example, if the container orchestration system includes Kubernetes, the GW (554) can be an Istio gateway that defines the load balancing. An API aggregator (API) (556) can be part of the remote cloud management system (112) (e.g., the private cloud AI API / UI (124)) used to view resources using the API inventory. An update mechanism (Update) (558) can be implemented using orchestration in the remote cloud management system (112) to perform updates and / or retrieve information from the private cloud system (102).A Communication Mechanism (CM) 560 can be a communication mechanism used for the container orchestration system. For example, if the container orchestration system includes Kubernetes, the CM 560 can include Kafka. An Authorizer (auth) 562 can be used to perform authorizations and can be part of the remote cloud management system 112 or a related platform, such as the Authorizer 140. A Task Manager (task) 564 can be part of the infrastructure of the remote cloud management system 112 and / or the private cloud system 102, providing a tracking framework for ongoing and / or scheduled tasks. An analyzer 566 can be part of the remote cloud management system 112 and / or the infrastructure of the private cloud system 102, which collects information from the on-prem infrastructure services to check the health of the on-prem infrastructure components.
[0053] Client 552 initiates the software pre-check process by requesting system details (568) for the private cloud system 102 from gateway 554. Gateway 554 then forwards the request (570) to API 556. API 556 returns the system details (572) to gateway 554, which then forwards the system details (574) to client 552. These details ensure that the client's system details are up to date. Client 552 then sends a request (576) to gateway 554 to initiate a pre-check to verify that the private cloud system 102 is in a non-degraded state with no failed components and / or to check if an update is compatible. Gateway 554 then forwards the request (578) to the update gateway 558. Update 558 requests authorization (580) from the authorizing authority 562 using credentials entered into the client and / or stored during setup.If authorization is successful, update 558 creates a task (582) in task 564 and returns a task ID (584) to gateway 554, which forwards the task ID (586) to client 552 to enable tracking of the software pre-check task. Using this task ID, client 552 and / or another component can query the status of the software pre-check. For example, client 552 can display a graphical interface showing the percentage progress of the software pre-check, which can be updated by querying update 558, command module 560, task 564, and / or other appropriate components.
[0054] After successful authorization and task creation, Update 558 initiates the software pre-check (588) with Analyzer 566 to gather information from on-premises infrastructure services for the private cloud system 102. This information can indicate, for example, whether components are in a compromised state and / or suitable / compatible for a planned update. Update 558 can monitor the progress (590) of the software pre-check and transmit all software information events (592) to CM 560. When the task is complete and the software pre-check is finished, Task 564 can notify GW 554 by sending task details (594) back to GW 554, which then forwards the task details (596) to Client 552.
[0055] As mentioned previously, in some embodiments the software pre-tests may be included in a download of an update package and / or be separate from the download of the update package. Fig. 17 is a download process 600 that shows the exchange of operations and / or data between components of the remote cloud management system 112 and / or the private cloud system 102 as part of a software pre-check initiated with the "Submit" button 516 of Fig. 15 can be initiated. Some of the components, such as the client 552, the GW 554, the API 556, the update 558, the CM 560, the authorization 562 and the task 564, can be transferred between process 600 and process 550. Fig. 16 are common. Process 600 also uses an AI service 602 and a data service connector (DSC) 604. The AI service 602 can include the AI software 107 of the private cloud system 102 and / or the platform on which the AI software is implemented. The DSC 604 can include the DSC 108 from Fig. 1 in the private cloud system 102, which is used as an interface to the remote cloud management system 112.
[0056] Client 552 initiates the software pre-check process by requesting system details (606) for private cloud system 102 from gateway 554. Gateway 554 then forwards the request (608) to API 556. API 556 returns the system details (610) to gateway 554, which then forwards the system details (612) to the client. As mentioned earlier, these details ensure that the client's system details are up to date. Client 552 then sends a request (614) to gateway 554 to initiate a pre-check to verify that private cloud system 102 is not degraded and has no failed components, and / or to check if an update is compatible. Gateway 554 then forwards the request (616) to update gateway 558.Update 558 requests authorization (618) from the authorization authority 562 using credentials entered into the client and / or stored during setup. If authorization is successful, Update 558 creates a task (620) in Task 564 and returns a task ID (622) to GW 554, which forwards the task ID (624) to Client 552 to enable tracking of the download and / or software pre-check task. Using this task ID, Client 552 and / or another component can query the status of the download and / or software pre-check. For example, Client 552 can present a graphical interface displaying the percentage of the download and / or software pre-check completed, which can be updated by querying Update 558, CM 560, Task 564, and / or other appropriate components.
[0057] Update 558 also initiates the download of updates for AI tools from the software catalog using an orchestrator (626) for the AI service 602, such as the Private Cloud AI Platform Orchestrator 122 or the Remote Cloud Management System 112. Fig. 1. Update 558 then monitors the download progress (628). During and / or after the download for AI service 602, Update 558 can download a hypervisor package (HV package) (630) using DSC 604 and monitor the download (632). After downloading the HV package, Update 558 copies the HV package (634) to a data store for the HV. During and / or after downloading / copying the HV package, Update 558 downloads a firmware package (636) using DSC 604 and monitors the download progress (638). When the task is complete and the software pre-checks / downloads are finished, Task 564 can notify GW 554 by sending task details (640) back to GW 554, which forwards the task details (642) to Client 552.In some implementations, such errors can be indicated in the returned task details if an operation fails (e.g., a download).
[0058] Once the software has been downloaded and / or the software pre-tests have been successfully completed, updates can be applied to the Private Cloud System 102. Fig. Figure 18 shows an example of an update process 650 that can be used by the remote cloud management system 112 and / or the private cloud system 102 to apply updates to the private cloud system 102. Process 650 uses update 558, the AI service 602, and the DSC 604. Process 650 also uses a data operations manager (Data Ops Man) 654 to interface with storage 103, for example, an interface with its operating system. Process 650 also uses a DSC VM manager (DSC Man) 656 to manage a VM of the DSC 604. Additionally, process 650 includes HV hosts 662.
[0059] Since DSC 604 is on-premises software that connects the rack(s) to the remote cloud management system 112, a DSC VM used to implement the DSC can be the first targeted update. Accordingly, update 558 retrieves the DSC VM version (664) from DSC Manager 656 and obtains a list of available DSC VM versions (666) from DSC Manager 656. Update 558 then initiates an update of the DSC VM (668) for DSC Manager 656 using one of the available DSC versions, such as the latest full version. Each of the updates discussed in process 650 can involve the creation, tracking, and / or communication of tasks with CM 560, as previously described in relation to the Fig. 16 and Fig. 17 discussed software pre-checking and download tasks.
[0060] After the DSC-VM update is complete, update 558 begins updating the AI tools by retrieving an AI service version (670) from AI service 602. Update 558 can also perform software pre-checks on the AI tools to be downloaded using the orchestrator (e.g., the Private Cloud AI Platform Orchestrator 122) and the workload clusters of work nodes 104. If the pre-checks are successful, the update initiates the download of AI tool updates (672) via AI service 602. Update 558 then initiates the downloaded updates on the orchestrator (674) and applies the downloaded updates to each of the workload clusters (676).
[0061] After the AI tools update is complete, Update 558 forwards an operating system update for storage 103 (678) to Data Ops Man 654 to update the operating system. Once the storage operating system update is complete, Update 558 downloads an HV update package (680) via DSC 604. Before, after, or during the download of the HV update bundles, Update 558 can download server firmware bundles and extract the firmware (682). Update 558 then performs an iLO firmware update (684) for each HV host 662 in the HV cluster. Update 558 can also perform a dry run of firmware updates (686) on all HV hosts 662 in the HV cluster. Update 558 then updates every HV (688) for every HV host 662.Updating each HV may involve first placing each HV host (662) into maintenance mode before applying the update. After updating each HV host (662), update (558) causes each HV host to restart (690) and, after the restart, checks for version matching (692) with the target firmware version to confirm that the firmware update was successful.
[0062] Update 558 then updates the firmware control nodes 106 (694), such as the AI software control plane 110, and causes a restart of the HV hosts 662 (696). Update 558 can verify success (698) by checking an iLO installation queue to determine if the firmware update was completed after the restart.
[0063] Update 558 then updates the work nodes 104 by first placing the work nodes into maintenance mode (700), updating the firmware in the work nodes 104 (702), and removing the work nodes from maintenance mode (704).
[0064] In some embodiments, at least some of the processes described above may include more or fewer operations. For example, process 650 may include fewer or more steps in the software update without deviating from the teachings presented here. For example, Fig.Block 19 describes a process 720 that includes receiving a notification via a user interface (e.g., the private cloud Al / Ul 124) of the remote cloud management system 112 to update a batch (e.g., the Al batch 400) of artificial intelligence (AI) tools in the private cloud system 102 (Block 722). In response to receiving the notification, the remote cloud management system 112 updates a virtual machine of the DSC 604 of the private cloud system 102 using the remote cloud management system 112 (Block 724). Following the receipt of the notification and the virtual machine update, the remote cloud management system 112 uses the updated virtual machine to update an AI service platform used to deploy the AI tools (Block 726).Updating the Kl service platform may involve retrieving a current version of the Kl service platform, pre-checking the compatibility of an update for the Kl service platform and the health of the Kl service platform components, downloading the update via an AI application programming interface (API) of the remote cloud management system 112, and installing the update. Such updates may involve updating any part of the Kl stack, such as the storage operating system, the HV, the control node firmware, and / or the work node firmware, using one of the techniques discussed in relation to Process 650.
[0065] Although certain features of the present disclosure have been illustrated and described here, the person skilled in the art will think of many modifications and changes. It is therefore to be understood that the attached claims are intended to cover all modifications and changes that correspond to the true spirit of the present disclosure.
Claims
[1] A system, encompassing: a remote cloud management system that includes a private cloud AI platform orchestrator located away from a private cloud system and configured to: forms an interface with the private cloud system; Artificial intelligence (“AI”) operations are orchestrated on the private cloud system using the private cloud AI platform orchestrator; and Manages AI software installed in the private cloud system. [2] System according to claim 1, wherein the remote cloud management system comprises an AI application programming interface (API) management system configured to manage the AI software in the private cloud system. [3] System according to claim 2, wherein the AI API management system uses a plurality of APIs to form an interface with the AI software in the private cloud system. [4] System according to claim 2, wherein managing the AI software includes updating the AI software in the private cloud system remotely using the remote cloud management system. [5] System according to claim 4, wherein the remote cloud management system is configured to interface with the private cloud system via a tunnel implemented using a data service connector of the private cloud system. [6] System according to claim 5, wherein the remote cloud management system is configured to update the data service connector before updating the AI software. [7] System according to claim 4, wherein the remote cloud management system is configured to receive a connection from a rack in the private cloud system and complete an initial configuration of the rack. [8] System according to claim 7, wherein the initial configuration includes updating the AI software. [9] System according to claim 8, wherein updating the AI software comprises sequentially updating multiple components of the private cloud system in a hierarchical order based on a selection of an update option. [10] System according to claim 9, wherein the selection of the update option comprises a one-click update selection. [11] A computer-implemented method comprising the following: Receiving a notification via a user interface of a remote cloud management system to update a set of artificial intelligence (AI) tools in a private cloud system, where the remote cloud management system is configured to manage the private cloud system remotely. In response to receiving the notification, update a virtual machine of a data service connector (DSC) of the private cloud system via the remote cloud management system; In response to receiving the notification and updating the virtual machine, using the virtual machine to update a AI service platform used to deploy the AI tools, including updating the AI service platform: Obtaining an up-to-date version of the AI service platform; Preliminary checks of the compatibility of an update for the AI service platform and the state of the AI service platform's components; Downloading the update via a KL application programming interface (API) of the remote cloud management system; and installing the update. [12] Computer-implemented method according to claim 11, wherein the hint comprises a single input to update the entire stack including the AI service platform and the virtual machine. [13] Computer-implemented method according to claim 11, wherein the private cloud system comprises a local cloud that is implemented at least partially on-site at a customer's premises, and the remote cloud management system is implemented at one or more locations managed by a provider of the remote cloud management system. [14] Computer-implemented method according to claim 11, comprising: in response to the update of the AI service platform, the update of an operating system of the storage of the private cloud system. [15] Computer-implemented method according to claim 14, comprising: in response to updating the operating system of the storage of the private cloud system, updating the hypervisors of the private cloud system using the remote cloud management system. [16] Computer-implemented method according to claim 15, comprising: in response to the updating of the hypervisors, the updating of the firmware of control nodes and worker nodes in the private cloud system. [17] Computer-implemented method according to claim 11, wherein updating the virtual machine of the DSC comprises: Determining a version of the DSC's virtual machine; Determining one or more available versions for the DSC virtual machine; and Updating the DSC virtual machine to one or more of the available versions of the DSC virtual machine. [18] Computer-implemented method according to claim 17, wherein one of the one or more available versions comprises a latest stable version of the virtual machine. [19] Computer-implemented method according to claim 11, wherein installing the update comprises: Installation of updates for a private cloud AI platform orchestrator of the remote cloud management system; and Installation of updates on the work nodes of the private cloud system used to implement the KL tools. [20] A tangible, non-transitory, and computer-readable medium on which instructions are stored which, when executed by one or more processors of one or more computers, are configured to cause the one or more computers to: to present a user interface via a remote cloud management system configured to remotely manage a fixed cloud system using a tunnel implemented using a data service connector (DSC) of the fixed cloud system, wherein the fixed cloud system includes a variety of artificial intelligence (AI) tools; to receive a notification about updating to KL tools; in response to the suggestion to update a virtual machine used to implement the DSC; After completing the virtual machine update, update an AI service platform that is used to implement the AI tools using the DSC implemented with the updated virtual machine; to update a hypervisor on one or more hypervisor hosts of the on-premises cloud system after the AI service platform update is complete; and After completing the hypervisor update, update the firmware of the control nodes and worker nodes of the local cloud system.