AI tools in private cloud environments
By pre-configuring server racks and utilizing a remote cloud management system for online setup and updates, the complexity of managing AI tools in a local private cloud is resolved, achieving efficient and simplified management of hardware and software and ensuring system stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2026-03-10
AI Technical Summary
Managing and updating AI tools in a local private cloud presents challenges such as the complexity of hardware and software selection and the difficulty of keeping them up-to-date, especially the challenges of managing updates from multiple sources.
By pre-configuring server racks and utilizing a remote cloud management system, the hardware and software setup process is simplified. Online setup and updates are performed using the remote cloud management system, including network access, user access, and remote connections. Simplified software updates are performed using the remote cloud management system's user interface, following a specific update sequence to ensure system stability.
It improves the efficiency of installing and updating AI tools in local private clouds, simplifies the management process of hardware and software, ensures the stability and consistency of the system, and meets the ever-changing AI needs.
Smart Images

Figure CN121644358A_ABST
Abstract
Description
Background Technology
[0001] Artificial intelligence (“AI”) is a method of using non-human systems to learn from experience and mimic intelligent human behavior through machine learning. Therefore, AI provides powerful tools that can be used to efficiently process and / or analyze large amounts of data. AI tools can be deployed on suitable computing engines / hardware, such as cloud computing systems.
[0002] Cloud computing systems can be implemented in a variety of different ways, including public clouds and private clouds. Public clouds can be deployed where public users (typically subscribers) can access cloud services, while private clouds are limited to one or more organizations. In fact, the simplest private cloud can be managed by a single organization for internal use without providing services to others. One type of private cloud includes on-prem cloud, where the management entity controls or manages all the hardware and software implemented in the private cloud at its own site. Attached Figure Description
[0003] The features, aspects, and advantages of this disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings, in which similar characters denote similar parts throughout the drawings, wherein:
[0004] Figure 1 This is a diagram illustrating a private cloud (local) system with pre-configured AI tools and managed using a remote cloud management system, according to various aspects of this disclosure;
[0005] Figure 2 The illustration depicts various aspects of access, updates, and / or deployment according to this disclosure. Figure 1 A flowchart of the process of using AI tools in a private cloud system;
[0006] Figure 3 It is applicable to aspects of the present invention. Figure 2 Example screens of the interface during the process;
[0007] Figure 4 It is an illustration of various aspects according to this disclosure. Figure 1 A diagram illustrating the process of adding (multiple) new racks to a private cloud system;
[0008] Figure 5 Based on the various aspects of this disclosure, it can be used as Figure 4 A sample screen presented in the user interface as part of the process;
[0009] Figure 6 Based on the various aspects of this disclosure, it can be used as Figure 4 An example network settings screen presented in the user interface as part of the process;
[0010] Figure 7 Based on the various aspects of this disclosure, it can be used as Figure 4 An example virtualization settings screen is presented in the user interface as part of the process;
[0011] Figure 8 Based on the various aspects of this disclosure, it can be used as Figure 4 An example control plane settings screen presented in the user interface as part of the process;
[0012] Figure 9 Based on the various aspects of this disclosure, it can be used as Figure 4 An example worker node setup screen is presented in the user interface as part of the process;
[0013] Figure 10 Based on the various aspects of this disclosure, it can be used as Figure 4 An example setup summary screen is presented in the user interface as part of the process;
[0014] Figure 11 It can be included in various aspects of this disclosure. Figure 1 AI stack in a private cloud system;
[0015] Figure 12 Based on the various aspects shown in this disclosure Figure 1 An example status screen of one or more components of a private cloud system that can be presented in the user interface;
[0016] Figure 13 It is in accordance with the various aspects of this disclosure that one can respond to the choice Figure 12 The status screen components are presented in the example software details screen of the user interface;
[0017] Figure 14 It is a pre-check screen that can be presented in a user interface according to various aspects of this disclosure, for the purpose of checking Figure 1 One or more components of a private cloud system perform software pre-checks;
[0018] Figure 15 A download screen that can be presented in a user interface according to various aspects of this disclosure, for downloading Figure 1 Software updates for one or more components of a private cloud system;
[0019] Figure 16 It is based on the illustrations of various aspects of this disclosure. Figure 1 A flowchart of the pre-inspection process for a private cloud system;
[0020] Figure 17 It is based on the illustrations of various aspects of this disclosure. Figure 1 A flowchart of the download and update process for a private cloud system;
[0021] Figure 18 The flowchart illustrates the unified process of a private cloud system based on various aspects of this disclosure; and
[0022] Figure 19 The flowchart illustrates the update process of a private cloud system based on various aspects of this disclosure. Detailed Implementation
[0023] One or more specific aspects of this disclosure will now be described. To provide a concise description of these aspects, not all features of an actual implementation may be described in the specification. It should be understood that, as in any engineering or design project, the development of any such actual implementation requires numerous implementation-specific decisions to achieve the developer's specific goals, such as complying with system-related and business-related constraints, which may vary from implementation to implementation. Furthermore, it should be understood that such development work may be complex and time-consuming, but it remains routine work in design, manufacture, and production for those skilled in the art who benefit from this disclosure.
[0024] In describing the elements of various aspects of this disclosure, the articles “a,” “an,” “the,” and “described” are intended to mean that there are one or more elements. The terms “comprising,” “including,” and “having” are intended to be inclusive and mean that other elements may exist in addition to those listed.
[0025] The embodiments provided herein relate to a technique for setting up and updating server racks to enable artificial intelligence (“AI”) capabilities in a private cloud or on-premises private cloud managed using remote cloud management. With AI software and AI-oriented hardware becoming increasingly widely available from numerous sources, the options for implementing AI solutions have become virtually limitless. Even selecting which hardware and software to implement in the on-premises cloud can become complex, as the management entity is responsible for all hardware and software (including AI tools). Managing hardware and software can be even more difficult due to resources coming from multiple different sources. Furthermore, being responsible for all management of the on-premises private cloud can make it difficult to keep such software and / or hardware up-to-date, especially when updates come from multiple different sources. Current technology simplifies the process for customers to incorporate new AI-enabled hardware into their on-premises private clouds by shipping server racks(s) already equipped with AI, making the setup process more efficient. For example, customers can select configurations and / or use cases that include hardware and software suitable for their desired AI applications. Infrastructure as a Service (IaaS) can be used to standardize integrated hardware and software to acquire infrastructure using subscription modeling by fulfilling model / use case / or (multiple) token requirements through recurring payments and / or the use of other suitable compute model recommendations and / or sizes. Once the standardization option is selected and purchased, the cloud administrator / manufacturer can deliver (multiple) racks to the customer's site within a specified time period (e.g., 8 hours).
[0026] When racks are delivered to the customer site, the customer and / or administrator / manufacturer can physically install and power on the racks. As previously noted, the racks may be pre-configured with AI tools, such as vendor-specific and / or open-source tools, models, and / or other AI support. Since the AI tools are pre-configured in the racks, setup only depends on establishing connectivity between the racks and remote cloud management. Therefore, setup can be simplified to online setup, including 1) configuring network access to the racks, 2) configuring user access for users using the racks / on-premises private cloud, and 3) establishing a link to the remote cloud to manage the on-premises private cloud. This setup may include default values (e.g., location), which may be changed during setup and / or upon arrival. In some embodiments, if new racks are added to racks already managed by the remote cloud, the new racks can be added via the remote cloud and linked accordingly. Remote cloud enables the migration of workloads and / or the expansion of capacity to meet evolving consumer AI needs.
[0027] After rack installation and / or as part of rack (multiple) installation, rack (multiple) may have software updates to be applied to rack (multiple) of them. Updates can also be simplified using remote cloud management. For example, AI tools can be updated using a simple (e.g., click) update via remote cloud. Updates can begin by interacting with and / or using a user interface that is part of the remote cloud management system. Updates can occur as a single operation including downloading and applying the update, or they can be separate operations where the download occurs at a different time than the application. Before downloading and / or applying the update, the remote cloud management system can check which updates are available in the cloud repository for AI tools and applicable to components of the on-premises private cloud. The remote cloud management system can also check to ensure the system is in a non-degraded state and free of faulty components. Updates can follow a specific sequence, where a data service connector update is applied to update the data service connector between the on-premises private cloud and the remote cloud management system. The updated data service connector is then used for subsequent update components, such as updating the control plane, updating the BIOS / firmware / OS, virtualization components (e.g., virtual machines (VMs) and their components), and / or other parts of the on-premises private cloud.
[0028] These installation and update technologies can provide overall improvements for consumers using local private clouds, simplifying the installation of pre-configured racks for Type 1 AI workloads and different pre-configured racks for Type 2 AI workloads. Using a simplified update process with a specified application update order, remote cloud management can also install updates to the entire AI stack, including the control plane, firmware, storage devices, OS, AI software / nodes, etc.
[0029] Considering the above, Figure 1 This diagram illustrates a system 100 for implementing cloud-based AI tools. System 100 includes a private cloud system 102, at least partially managed by the customer organization. The private cloud system 102 may be a local cloud system, wherein the hardware associated with cloud services implemented using the private cloud system 102 is at least partially located on-premises and at at least one physical site within the customer organization. This local cloud approach may be desirable where the customer expects to retain at least some of the hardware and / or its data on-premises. This freedom to manage its own hardware also provides the customer with the ability to create custom configurations using specially selected hardware and / or software solutions. However, using such custom configurations can also complicate the deployment of new hardware, new software, the management of current software or hardware, and / or the updating of current software to newer versions, because solutions from multiple manufacturers and / or providers are custom in nature, and the sources of configuration and / or settings differ.
[0030] Private cloud system 102 can use a container orchestration system, which acts as the operating system for private cloud system 102. The container orchestration system can assemble one or more computers, including virtual machines (VMs) and / or bare metal servers (BMs), into a cluster that executes workloads within containers. For example, the container orchestration system may include... An open-source container orchestration system and / or any other suitable container orchestration system. The container orchestration system should work correctly with one or more container runtimes (such as HPE). Kubernetes Container Runtime Interface (CRI-O) (rkt, and / or any other suitable container runtime) to execute workloads within a container.
[0031] Since the private cloud system 102 is implemented using computers, one or more servers (e.g., BMs), and one or more VMs, the private cloud system 102 includes a combination of hardware elements (e.g., processors) and software elements (including tangible, non-transient, and computer-readable media), such as storage device 103. Storage device 103 may include any suitable article of art for storing data and / or executable instructions, such as random access memory, read-only memory, rewritable flash memory, hard disk drive, and optical disk. Additionally, programs encoded on such computer program products (e.g., using container runtimes and / or container orchestration systems) may also include instructions executable by the processors(s) of the private cloud system 102, enabling containers of the private cloud system 102 to execute workloads within containers.
[0032] As illustrated, the private cloud system 102 includes worker nodes 104 and a control node 106. Worker nodes 104 run applications such as AI software 107 within the cluster. Worker nodes 104 process data and control networking of the private cloud system 102. Worker nodes 104 host application containers in a group format, and these application containers, in turn, run one or more containers. Worker nodes 104 report to the control node 106. The control node 106 manages the operation of the private cloud system 102 by controlling which node in the worker nodes 104 runs containers and when. In other words, the control node 106 includes a scheduler that communicates with worker nodes 104 to schedule container workloads. The scheduler may consider computational availability (such as CPU and memory availability) and the needs of applications (e.g., AI software 107) to determine which worker node 104 performs which tasks at what time. Worker nodes 104 may utilize node-level agents to track resource consumption and facilitate the completion of worker node 104 allocations, thereby ensuring that worker nodes 104 execute the assigned tasks.
[0033] Control node 106 manages the communication and control of worker node 104 and may include an application programming interface (API) server and store configuration and status data. In some embodiments, a control plane (e.g., AI software control plane 110) may operate across multiple control nodes 106 to provide redundancy.
[0034] Control node 106 can also communicate with external services by implementing data service connector 108. Control node 106 may also include an AI software control plane 110, which is used to control (e.g., schedule) workloads within the application container of AI software 107. As discussed in more detail below, AI software 107 may include software provided by one or more different providers (such as HPE, NVIDIA, open-source partners), and / or other tools that may be provided to private cloud system 102 via remote cloud management system 112.
[0035] System 100 includes a remote cloud management system 112 and a data service connector 108. The remote cloud management system 112 is used to manage a private cloud system 102 via one or more networks 114 (e.g., the Internet), while the data service connector 108 is implemented in the control node 106 of the private cloud system 102. The remote cloud management system 112 is located remotely from the private cloud system 102, but it can be used to perform remote management of the private cloud system 102 via a tunnel connector 116. The remote cloud management system 112 is located remotely from the private cloud system 102 because it can be implemented using a different site than the computer / server used to implement the private cloud system 102. The tunnel connector 116 pairs with the data service connector 108 to create a secure (“encrypted”) remote connection over one or more networks 114 to maintain information security and confidentiality between the remote cloud management system 112 and the private cloud system 102.
[0036] The remote cloud management system 112 uses a combination of hardware and software to implement a private cloud infrastructure orchestrator 118, a private cloud resource orchestrator 120, a private cloud AI platform orchestrator 122, and a private cloud AI API / user interface (UI) 124.
[0037] The private cloud infrastructure orchestrator 118 orchestrates operations related to setting up the infrastructure of the private cloud system 102 (e.g., setting up control node 106, connecting the private cloud system 102 to the remote cloud management system 112, etc.). The private cloud infrastructure orchestrator 118 also controls software updates and resource inventory of the private cloud system 102, controls network management of the private cloud system 102, monitors / controls the metering of resources (e.g., processing, RAM, and / or power) used by the infrastructure during operation of the private cloud system 102, and / or generates dashboards for displaying information about the infrastructure of the private cloud system 102.
[0038] Private cloud resource orchestrator 120 uses components such as VMs, BMs, and / or container orchestration systems of private cloud system 102 to orchestrate operations. For example, private cloud resource orchestrator 120 can be used to provide and manage VMs, BMs, and / or container orchestration systems (e.g., Kubernetes) of private cloud system 102.
[0039] The private cloud AI platform orchestrator 122 can be used to manage the AI platform using AI software 107. For example, the private cloud AI platform orchestrator 122 can deploy and / or extend AI applications installed and available in the AI software 107 of the private cloud system 102. The private cloud AI platform orchestrator 122 can perform such deployment and / or extension of the resource inventory of the AI software 107 using the tunnel between tunnel connector 116 and data service connector 108 and / or via sideband connection 126 between the private cloud AI API / UI 124 and the AI software 107. For example, sideband connection 126 can be a secure tunnel separate from the tunnel between data service connector 108 and tunnel connector 116.
[0040] The private cloud AI API / UI 124 provides a UI that enables remote management of AI software in the private cloud system 102 using APIs. For example, a user can log in to the remote cloud management system 112 and use APIs (such as Representational State Transfer (REST) APIs) via the private cloud infrastructure orchestrator 118 and / or the private cloud resource orchestrator 120 to control changes in the private cloud system 102. The private cloud AI API / UI 124 includes an infrastructure manager 128, which manages infrastructure changes via API calls from the private cloud infrastructure orchestrator 118, for the management of the infrastructure and / or changes to the infrastructure itself. The private cloud AI API / UI 124 also includes an AI software platform manager 130, which manages changes to the AI software 107 and / or the AI software control plane 110 via the private cloud resource orchestrator 120 and / or the private cloud AI platform orchestrator 122. Alternatively or additionally, the private cloud AIAPI / UI 124 can modify the AI software 107 via the sideband connection 126 through the AI interface 132. The AI interface 132 can be used to authenticate and / or encrypt the sideband connection 126 between the private cloud AIAPI / UI 124 and the private cloud system 102.
[0041] The remote cloud management system 112 may include add-ons to aid in the remote management of the private cloud system 102. For example, the remote cloud management system 112 may include a software catalog 134 that stores and / or links to available software that can be used in the private cloud system 102. For example, the software catalog 134 may determine which software is suitable and usable for the specific hardware and / or software configuration of the private cloud system 102. For example, if an organization / user associated with the private cloud system 102 has subscribed to (e.g., from the provider of the remote cloud management system 112 and / or from a third-party provider) AI services, the software catalog 134 may provide corresponding AI tools for installation / use in the private cloud system 102. In addition to subscription-based filtering, or as an alternative, the software catalog 134 may filter based on whether the organization / user associated with the private cloud system 102 meets requirements before providing at least some AI services. For example, the software catalog 134 may avoid displaying AI tools from at least some providers (e.g., third-party providers) until the organization / user has indicated that they have agreed to an agreement with the respective provider. For example, the agreement could be an End User License Agreement (EULA) and / or other licensing agreements.
[0042] The remote cloud management system 112 may include an auditor 138, which may be implemented using hardware and / or software to enable users / organizations to view metrics related to ongoing and / or historical workloads on the private cloud system 102. The remote cloud management system 112 may also include an authorizer 140, which authorizes any user attempting to access the remote cloud management system 112 and / or the private cloud system 102, and then grants such access.
[0043] During operation, users can choose which AI tools are available in the private cloud system 102. Figure 2 The process 150 of deploying AI tools in a private cloud system 102 using a remote cloud management system 112 is illustrated. The remote cloud management system 112 receives a user's login credentials via a private cloud AI API / UI 124 (box 152). The remote cloud management system 112 uses an authorizer 140 to check whether the user is authorized to deploy, modify, and / or use AI software 107 in the private cloud system 102 (box 154). If the credentials are invalid or the user is not authorized to access, use, and / or modify AI software 107, the credentials are not authorized, and the remote cloud management system 112 can re-request login credentials. In some embodiments, the remote cloud management system 112 may only receive a limited number of credential attempts (e.g., 1, 2, 3, 4, or more) before locking the account, logging the failed authorization check, and / or notifying the administrator of the credential authorization check failure.
[0044] If authorization is successful, the remote cloud management system 112 presents the available solutions (box 155) in the private cloud AI API / UI 124. For example, Figure 3 A screen 160, which can be presented in the private cloud AIAPI / UI 124, shows deployable AI tools 162, which are pre-configured in multiple racks of the private cloud system 102, as indicated by the initialized label of the corresponding status 164. These deployable AI tools 162 can be deployed via a remote cloud management system 112 using a deployment button 166.
[0045] For deployed AI tools 168, as indicated by the deployed label in their status 164, no deployment button 166 is displayed, and the deployed AI tool 168 can be opened, edited, or run by clicking on it. In some embodiments, using the add button 169, uninitialized or undeployed AI tools (e.g., AI solution accelerators) can be added from the software catalog 134 based on their suitability for the private cloud system 102 and / or based on a subscription 136 that can be used to log in to the remote cloud management system 112.
[0046] In some embodiments, the deployable AI tool 162 and the deployed AI tool 168 may include descriptions and / or labels that indicate the objective, application area, platform, programming language, and / or other details about the respective AI tool and / or how to use it.
[0047] Back Figure 2 The system receives one of the presented options via the private cloud AI API / UI 124 (box 156). For example, an AI tool among a plurality of deployable AI tools 162, an AI tool among a plurality of deployed AI tools 168, and / or an add button 169 can be the received option selected via screen 160 of the private cloud AI API / UI 124. If the selection corresponds to a new deployment (or an AI tool is added via add button 169) (box 157), the remote cloud management system 112 deploys the new AI tool (box 158). In some embodiments, deploying a new AI tool may include modifying the AI workload in worker node 104 to accommodate the newly deployed AI tool. Deployment may also include displaying the deployment status before, during, and / or after deployment is complete. For example, screen 160 may be updated to display status information such as deployed, deployment with a percentage completion indicator, and / or other suitable indicators of deployment status.
[0048] If the selected solution has been deployed, the remote cloud management system 112 can perform operations (box 159). For example, these operations may include adjusting the workload of the selected AI tool, running a process using the selected AI tool, stopping the operation of the selected AI tool, viewing the data results of the AI tool's execution, running the AI tool for different input data, and / or any other suitable operation that can be performed using the selected AI tool.
[0049] In some scenarios, adding or deploying new AI tools may consume a significant portion of the private cloud system 102. In such cases, or during the initial setup of the private cloud system 102, hardware will be added to enable the AI tools. However, the user or AI administrator performing this operation may be a different category of user (e.g., a cloud administrator) with the permissions / capabilities to add new hardware to the private cloud system 102. Figure 4 Process 170 is illustrated, which is used to acquire and provision hardware in a private cloud system 102 with pre-configured AI tools. A remote cloud management system 112 can present options for the hardware to be implemented (box 172) via a private cloud AI API / UI 124. Options can be presented in response to authentication verification, as described above regarding... Figure 2 The subject of discussion. Figure 5Screen 200, which can be displayed in the private cloud AI API / UI 124, is shown. Screen 200 can be presented when an authorized user requests to view options for creating and / or expanding the private cloud system 102 with new hardware and / or replacing hardware. Screen 200 includes a set of different configurations 202, 204, 206, and 208 that can be added to the private cloud system 102. In some embodiments, configurations 202, 204, 206, and 208 may be general options suitable for implementing AI operations. Additionally or alternatively, configurations 202, 204, 206, and / or 208 may be recommendations based on existing hardware in the private cloud system and / or based on information provided by the cloud administrator or the organization of the cloud administrator. For example, a recommended configuration may prioritize configurations similar to those already deployed in the private cloud system 102. Additionally or alternatively, the recommended configuration may be based on several indicators, which indicate the type of AI operation expected, the expected compute requirements of the AI operation, the expected storage requirements of the AI operation, the expected network requirements of the AI operation, and / or the newly added power budget.
[0050] Each of configurations 202, 204, 206, 208 can have a corresponding AI function, such as inference, retrieval augmentation generation (RAG), model fine-tuning, other AI functions, or a combination thereof. RAG is an AI framework that utilizes traditional information retrieval systems (such as databases in storage device 103). RAG optimizes large language models (LLMs) by enabling them to access up-to-date information from a curated database and incorporate it into their responses and analyses. This additional knowledge enables more accurate, relevant, and up-to-date preference-based analyses and / or recommendations. Model fine-tuning involves taking a model (e.g., based on a small amount of learning) and performing iterative supervised or unsupervised training on it, using more training examples than a small amount of learning. Configurations 202, 204, 206, and / or 208 can be labeled with applicability tags 210, indicating the appropriate operation for the corresponding configuration. Additionally or alternatively, cloud administrators can select the desired AI functionality in the private cloud AIAPI / UI 124 during the review of options, or it can be pre-configured and stored in the preferences of the remote cloud management system 112. Alternatively, cloud administrators can indicate which configuration is required (e.g., small, medium, large, or extra-large).
[0051] Screen 200 may also provide information about different configurations 202, 204, 206, and / or 208. For example, screen 200 may display computing component indications 212 for configuring 202, 204, 206, and / or 208. For example, computing components may include a graphics processing unit (GPU), a central processing unit (CPU), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), other processors suitable for AI computing, or combinations thereof.
[0052] Screen 200 may also display storage component indication 214 for configuring 202, 204, 206, and / or 208. Storage component indication 214 may include the number of memory and / or storage devices. In some configurations, the type of memory (e.g., latency and / or frequency) may be selected via storage component indication 214 or through another interface. Storage component indication 214 may indicate the number of basic storage devices and / or one or more upgraded storage devices that can be deployed in rack(s) when installed in the private cloud system 102. Additionally or alternatively, storage component indication 214 may indicate the maximum number of storage devices that can later be added to the configuration.
[0053] The screen 200 may also display network indication 216 for configurations 202, 204, 206, and / or 208. Network indication 216 may indicate the data transmission speed for the corresponding configuration. The screen 200 may also include a power indicator 218, which indicates a power consumption estimate for the corresponding configuration.
[0054] Back Figure 4 The private cloud AI API / UI 124 can receive a selection of one of the presented options and order the corresponding configuration (box 174). In some embodiments, the selection and ordering can be performed offline (e.g., by phone and / or using an order form with a hard or soft copy), by an agent of the provider of the remote cloud management system 112. The provider then provides the ordered rack(s) with pre-loaded AI tools and / or access to the AI tools. The provider and / or the ordering organization then configures the hardware (box 175). The following discusses... Figures 6-9 This is related to the hardware settings.
[0055] Before, during, and / or after ordering hardware, the cloud administrator adds users and / or user roles to (multiple) racks (box 176). For example, the cloud administrator can add users to (multiple) racks. Figure 2The login credentials used are employed to set up, modify, and / or use the AI software 107 in the private cloud system 102 via newly configured hardware. Therefore, some users may be allowed to modify AI tools (e.g., begin fine-tuning), while others may only be granted access to or permission to view the results of AI operations. The cloud administrator can then use the private cloud AI API / UI 124 to manage the private cloud system 102, including observing workloads and / or other AI dashboards available in the private cloud AI API / UI 124. From here, the cloud administrator can extend the order / private cloud system 102 to include additional available hardware and / or software components (box 180). For example, the private cloud AI API / UI 124 can return to the presentation options in box 172. From the management / observation steps, the cloud administrator can use the private cloud AI API / UI 124 to update the AI stack in (multiple) racks within the private cloud system 102 (box 182). The following discusses... Figures 11-14 This is related to the update mechanism.
[0056] Hardware settings
[0057] As previously noted, Figure 6 This involves hardware setup of racks(s) within the private cloud system 102 using the private cloud AI API / UI 124 and / or any other part of the remote cloud management system 112 and / or the private cloud system 102. Therefore, Figure 6 Screen 240 is shown, in which the private cloud AI API / UI 124 and / or other parts of the remote cloud management system 112 can be presented to the cloud administrator as part of the setup. Screen 240 includes a progress tracker comprising an infrastructure section 242 and a private cloud AI section 244. As illustrated in the progress tracker, the infrastructure is set up before the private cloud AI (e.g., AI software 107 of private cloud system 102) is set up; the infrastructure section 242 is expanded, while the private cloud AI section 244 is collapsed. Within the infrastructure section 242, the progress tracker includes a network section 246, which is used to configure the network as part of the setup. The network section 246 is bolded and / or otherwise emphasized to indicate that the network is currently being set up. In the network setup, screen 240 displays text 248, which can be used to provide an explanation of the completed network configuration section. Screen 240 also includes a server management network input 250, which enables the input of the IP address of the remote cloud management system 112, thereby enabling the establishment of a connection between the remote cloud management system 112 and the private cloud system 102.
[0058] Screen 240 also includes an integrated lights-out (iLO) management network input 252, which allows input of the IP address of an iLO management system used to remotely and securely manage, simplify, and / or automate server operations. The iLO management network input 252 may be located at the same address as the remote cloud management system 112 indicated in the server management network input 250. Therefore, screen 240 includes a selector 254 that allows a cloud administrator to instruct the remote cloud management system 112 and the iLO management system to share the same IP address. If the selector 254 indicates that the IP addresses of the two networks are the same, the iLO management network input 252 and / or the server management network input 250 can be hidden or otherwise disabled. When the same or different IP addresses are used for the iLO management network and the remote cloud management system 112, screen 240 includes an iLO subnet input 256, which allows a cloud administrator to specify a specific subnet mask and gateway for the iLO management network. Since the private cloud system 102 also has data for AI operations in storage device 103 and / or other locations, screen 240 enables the input of IP address 258, data subnet mask 260, and / or gateway 262 to access the data that will be used for AI operations in the private cloud system 102.
[0059] Once the network is configured on screen 240, you can select the Next button 263 to proceed to the control plane server section 264 of the progress tracker. Similar to managing the network, location (e.g., IP address, subnet mask, and / or gateway), serial number, model, CPU series, and / or other information about the control plane server can be entered into the private cloud AIAPI / UI124. Once the location of the control plane server is specified and / or the appropriate Next button is selected, virtualization is configured using the virtualization section 266 of the progress tracker. Once the infrastructure setup is complete, you can select the summary section 268 of the progress tracker (e.g., using the Next button in the virtualization settings section).
[0060] Figure 7This is a diagram that corresponds to the virtualization portion 266 of the progress tracker, screen 280. Thus, screen 280 can be presented using the private cloud AIAPI / UI 124 and / or any other portion of the remote cloud management system 112 and / or the private cloud system 102. Screen 280 includes a hostname input 282 configured to receive the hostname of server management software used to control the virtual machine environment. For example, the server management software may include any suitable server management software, such as NVIDIA Virtual GPU (vGPU) software, VMware vCenter, and / or other suitable server management packages. Screen 280 also includes a credential input 284 used to input hypervisor (HV) credentials for the server management software, the remote cloud management system 112, and / or the private cloud system 102. For example, this credential may include a single sign-on (SSO) credential, which is a digital credential that can be used for multiple applications or websites, such as other locations within the server management software and / or the remote cloud management system 112 and / or the private cloud system 102. The credential input 284 may include drop-down menus and / or pop-up menus that enable the selection of credentials.
[0061] Screen 280 also includes a Hypervisor (HV) root credential entry 286, which enables the entry of HV root credentials using manual input, drop-down menus, pop-up menus, or other mechanisms. These credentials will be used to authenticate with the hypervisor. The hypervisor is installed on rack(s) and used to partition the racks into VMs. Similarly, an iLO administrator credential entry 288 enables the entry of iLO administrator credentials to authenticate with the iLO management software discussed earlier. Screen 280 may display an Accept button 290 and a statement accepting one or more agreements with providers of various application software (e.g., AI software 107) and / or management tools. This statement may include links to various different agreements, even between different providers. Once the virtualization credentials are provided, a Next button 292 can be selected to proceed.
[0062] Once the infrastructure setup is complete, the progress tracker can proceed to the private cloud AI section 244, such as... Figure 8As illustrated in screen 300. In screen 300, the Private Cloud AI section 244 is expanded, while the Infrastructure section 242 is collapsed, indicating that the setup has been moved to the Private Cloud AI section 244 of the setup. As illustrated, the Private Cloud AI setup can be divided into Control Plane Setup, Worker Node Setup, and Summary, as indicated by indicators 302, 304, and 306. Since screen 300 is currently being used to configure the Control Plane, screen 300 includes bold, italic, underlined, and / or otherwise emphasized indicator 302, while indicators 304 and 306 are not emphasized. Screen 300 includes a Control Plane VM Name Prefix Input 308, which is used to enter the name prefixes for one or more (e.g., three for redundancy) Control Plane VMs. Screen 300 also includes a Key Input 310 for entering a key used to access the VMs, which can be selected from a stored list of secrets and / or uploaded to the Private Cloud AI API / UI 124.
[0063] Screen 300 also includes networking details for managing, storing, or serving nodes by selecting a network from the network menu 312. Then, a starting IP address is provided for the control plane VM using the starting IP input 316, which indicates the first VM's IP address. The next available address is given for the next VM. The networking details also enable the use of the cluster IP input 314 to indicate the cluster IP address and the ingress IP input 316 to indicate the ingress IP address. Once this information is entered, the next step in the AI settings associated with the serving node can be accessed when the next button 318 is activated due to completion of the input data.
[0064] Once the next button 318 is pressed, the private cloud AIAPI / UI 124 will make the screen 320 as... Figure 9The diagram is shown in the figure. Screen 320 includes an emphasis on and expansion of indicator 304 to include indicators 326 and 328. Indicator 326 corresponds to the server of worker node 104 and can be used to enter IP address, name, serial number, device type, type of processor used, number of storage devices, and / or any other useful information about the server implementing worker node 104. Indicator 328 can be used to configure worker node 104. As indicated by the emphasis on indicator 328, setup for configuring worker node 104 is in progress. To aid in completing the configuration of worker node 104, screen 320 includes a worker node name prefix input 330, which allows each worker node 104 to be appended with a human-readable prefix. Screen 320 also includes key inputs 332 and 334, which allow the input of corresponding keys to use iLO and perform appropriate access using a worker node operating system such as Red Hat Enterprise Linux (RHEL). This configuration may include, if applicable, the partition specification for enabling the Multi-Instance GPU (MIG) 336. Screen 320 may also allow input of the starting IPs 338 for a number (e.g., four) of the worker nodes 104, with each worker node 104 assigned the next available number. Once the configuration for the worker nodes 104 has been entered, the setup can continue by receiving a selection on the Next button 340.
[0065] After configuration, as follows Figure 10 As illustrated, screen 350 can be presented using Private Cloud AI API / UI 124, which displays a settings summary of Private Cloud AI section 244, as indicated by the emphasis in indicator 306. This summary includes control plane details 352 regarding the control plane / control node 106. For example, the summary may include VM access keys, network names, subnet masks, control node gateways, the name of each control node 106 and its respective IP address, or any combination thereof. The summary also includes information 354 about worker nodes 104, such as iLO access keys, OS access keys, the name of worker node 104, the IP address of worker node 104 on the management network, the IP address of worker node 104 on the iLO network, and the IP address of worker node 104 in the storage device 103 network. In some embodiments, the summary may include a status indicator indicating the configuration status, such as a completion percentage of 0, 5, 10, 25, or more. Once the summary details are confirmed, the submit button 356 can be selected. Alternatively, you can use this screen 350 or the back button 358 on the previous screen to navigate back and change / update information about the control plane and / or working nodes as part of the settings.
[0066] Figure 11 An example AI stack 400 is shown, which includes a data service connector (DSC) 402, such as Figure 1 The AI stack 400 also includes a hypervisor (HV) 404 (e.g., ESXi), a virtualization platform 406 (e.g., vSphere), server firmware 408 for control node 106, storage device 410, network connectivity 412, an operating system 414 for worker nodes 104 implementing AI software 107, server firmware 416 for worker nodes 104, Kubernetes 418 and / or other container orchestration systems, and other AI tools 420. For example, AI tools 420 may be tools provided by the provider of the remote cloud management system 112, another provider (e.g., a third-party provider such as NVIDIA), open-source tools, and / or other AI tools that can be provided to the private cloud system 102 via the software catalog 134 of the remote cloud management system 112. HV 404, virtualization platform 406, server firmware 408, and storage device 410 may be part of the control panel of the private cloud system 102, as indicated by instruction 422. Operating system 414, server firmware 416, Kubernetes 418, and AI tools 420 may be part of and / or implemented using worker node 104, as indicated by instruction 424. It should be understood that different update sources for different objects in the AI stack 400 make such updates more difficult and / or complex. To simplify this process, private cloud infrastructure orchestrator 118, private cloud resource orchestrator 120, private cloud AI platform orchestrator 122, and private cloud AI API / UI 124 can be used to implement simplified updates, where multiple (e.g., all) objects in the AI stack 400 are updated in a composite operation and / or using a single action (e.g., a single click). In some embodiments, at least some components of the AI stack 400 can be updated using individual operations. For example, DSC 402 can be updated using a VM image stored in a remote cloud management system 112, which the customer can directly use to update the VM image. Additionally or alternatively, the customer can use direct updates to update network connection 412.
[0067] Software update
[0068] Figure 12Screen 440 is shown, which can be presented via the private cloud AI API / UI 124 and / or any other part of the remote cloud management system 112. Screen 440 displays a list of one or more software records that may have available updates. In some embodiments, all deployed AI tools may display records, but in some embodiments, only records that have potential updates in any part of the AI stack of an AI tool may be displayed. All other AI tools may be hidden. Each record includes a name 442 of the object corresponding to the record, a health status 444 indicating the known health status of the object, a hypervisor cluster indicator 446 indicating which hypervisor cluster the object belongs to, a last update date indicator 448 indicating the last update date (if the object has been updated previously), and an update status 450 indicating whether an update is available for the object. Interacting with the update status 450 via clicking, mouse hovering, or other means will trigger a detailed information window 452. This window indicates the current version of objects in the AI stack 400, such as the current AI tool version, current operating system version, current hypervisor version, current storage version, and / or current firmware version. These are all part of and / or used by the current object version. Window 452 may include a "View Details" button 454 for viewing more detailed information about the object. Selecting the "View Details" button 454 and / or clicking window 452 can display a software details screen (such as...). Figure 13 The screen 470 displays the current version 472 and one or more updated versions 474 and 476. The current version 472 may be labeled to indicate that it is the current version. In some embodiments, the updated versions 474 and 476 may include labels to indicate that the corresponding updated version is the latest version (e.g., version 6.9.9).
[0069] In selection Figure 12 After recording on screen 440, the pre-check button 456 can be used to pre-check the compatibility of the AI stack 400 of (multiple) racks with the proposed update. The pre-check button 456 can be used to check the compatibility of the update and download it, then apply the update at a later time. However, if the update needs to be deployed during or after the update download without waiting for a later update to be initiated, the update button 458 can be used to apply the update, which will cause a confirmation of the update and pre-check in response to selecting the update button 458. Alternatively or additionally, the update can be downloaded using the pre-check button 456, and the update button 458 can be disabled until the update is downloaded.
[0070] In selection Figure 12 After the pre-inspection button 456, the private cloud AIAPI / UI 124 can make the pre-inspection screen appear, such as... Figure 14Screen 490 includes a header 492 that clearly indicates the selected hypervisor cluster for the pre-check. Menu 494 allows selection of which version of the update to pre-check. To begin the pre-check, a submit button 496 is presented; upon selection, this button causes the private cloud AIAPI / UI 124 to initiate the pre-check and / or download of the selected update. If you do not wish to begin the pre-check, you can select the cancel button 498 to return to the previous screen without initiating the pre-check for the selected update.
[0071] In selection Figure 12 After clicking the update button 458, the private cloud AIAPI / UI 124 can make the update screen appear, such as... Figure 15 Screen 510 includes a header 512 that clearly indicates the selected hypervisor cluster to be updated. Menu 514 allows selection of which version to update. To begin the update, a submit button 516 is presented; upon selection, this button causes the private cloud AIAPI / UI 124 to download and / or deploy the selected update. If you do not wish to begin the update, you can select the cancel button 518 to return to the previous screen without deploying the selected update.
[0072] Figure 16 This is a flowchart of the software pre-check process 550, which illustrates the operation and / or data exchange between components of the remote cloud management system 112 and / or the private cloud system 102, which can be used... Figure 14The submit button 496 is part of a software pre-check. Client 552 can be an application implemented in and / or presented via the private cloud AIAPI / UI 124, which customers / users / organizations can access. Gateway (GW) 554 can be part of a container orchestration system used for load balancing workloads. For example, if the container orchestration system includes Kubernetes, then GW 554 can be an Istio gateway defining the load balancer. API aggregator (API) 556 can be part of a remote cloud management system 112 (e.g., private cloud AI API / UI 124) used to display resources using API manifests. An update mechanism 558 can be implemented using orchestration in the remote cloud management system 112 to perform updates and / or obtain information from the private cloud system 102. Communication mechanism (CM) 560 can be a communication mechanism used for the container orchestration system. For example, when the container orchestration system includes Kubernetes, CM 560 can include Kafka. Authorization (auth) 562 can be used to perform authorization and can be part of the remote cloud management system 112 or related platforms (such as authorizer 140). Task administrator (task) 564 can be part of the infrastructure of the remote cloud management system 112 and / or private cloud system 102, providing a tracking framework for ongoing and / or planned tasks. Analyzer 566 can be part of the infrastructure of the remote cloud management system 112 and / or private cloud system 102, collecting information from local infrastructure services to check the health of local infrastructure components.
[0073] Client 552 initiates a software pre-check process by requesting system details (568) of private cloud system 102 from GW 554. GW 554 then forwards the request to API 556 (570). API 556 returns the system details to GW 554 (572), which then forwards the system details to client 552 (574). This information ensures that the system details in the client are up-to-date. Client 552 then sends a request (576) to GW 554 to initiate a pre-check to verify that private cloud system 102 is in a non-degraded state and has no faulty components and / or to verify that updates are compatible. GW 554 then forwards the request to update 558 (578). Update 558 requests authorization (580) from authorization 562 using credentials entered in the client and / or stored during setup. If authorization is successful, update 558 creates a task (582) in task 564 and returns a task identifier (task ID) (584) to GW 554, which forwards the task ID to client 552 (586) to enable tracking of the software pre-check task. Using this task ID, client 552 and / or any other component can poll and request the status of the software pre-check. For example, client 552 can present a graphical interface displaying the percentage of completion of the software pre-check that can be updated by polling update 558, CM 560, task 564, and / or any other suitable component.
[0074] After successful authorization and task creation, update 558 initiates a software pre-check (588) using analyzer 566, causing it to collect information from the local infrastructure services of private cloud system 102. This information may indicate whether any components are in a degraded state and / or suitable / compatible with the planned update. Update 558 may monitor the progress of the software pre-check (590) and transmit any software information events to CM 560 (592). When the task is complete and the software pre-check is finished, task 564 can notify GW 554 by returning task details to GW 554 (594), which then forwards the task details to client 552 (596).
[0075] As previously noted, in some embodiments, software pre-checks may be included in and / or may be separate from the update package download. Figure 17 This is download process 600, which illustrates the operation and / or data exchange between components of the remote cloud management system 112 and / or private cloud system 102, which can be used... Figure 15 The submit button 516 is part of the software pre-check initiated by the application. Some components (such as client 552, GW 554, API 556, update 558, CM 560, license 562, and task 564) can... Figure 16 Process 550 and process 600 share resources. Process 600 also utilizes AI service 602 and data service connector (DSC) 604. AI service 602 may include AI software 107 of private cloud system 102 and / or a platform on which AI software is implemented. DSC 604 may be... Figure 1 The DSC 108 in the private cloud system 102 is used to interface with the remote cloud management system 112.
[0076] Client 552 initiates a software pre-check process by requesting system details of private cloud system 102 from GW 554 (606). GW 554 then forwards the request to API 556 (608). API 556 returns the system details to GW 554 (610), which in turn forwards the system details to the client (612). As previously noted, this information ensures that the system details in the client are up-to-date. Client 552 then sends a request to GW 554 to initiate a pre-check (614) to verify that private cloud system 102 is in a non-degradable state and has no faulty components and / or to verify that updates are compatible. GW 554 then forwards the request to update 558 (616). Update 558 requests authorization from authorizer 562 using credentials entered into the client and / or stored during setup (618). If authorization is successful, update 558 creates a task in task 564 (620) and returns a task identifier (task ID) to GW 554 (622). GW 554 forwards the task ID to client 552 (624) to enable tracking of the download and / or software pre-check task. Using this task ID, client 552 and / or any other component can poll and request the status of the download and / or software pre-check. For example, client 552 can present a graphical interface displaying the completion percentage of the download and / or software pre-check, which can be updated by polling update 558, CM 560, task 564, and / or any other suitable component.
[0077] Update 558 also uses the orchestrator of AI service 602 (such as...) Figure 1The remote cloud management system 112's private cloud AI platform orchestrator 122 initiates an update to download AI tools from the software catalog (626). The update 558 then monitors the download progress (628). During and / or after the AI service 602 download, the update 558 can use DSC 604 to download the management program (HV) package (630) and monitor the download (632). After downloading the HV package, the update 558 copies the HV package to the HV's data repository (634). During and / or after downloading / copying the HV package, the update 558 uses DSC 604 to download the firmware package (636) and monitor the download progress (638). When the task is complete and the software pre-check / download is finished, the task 564 can notify GW 554 by returning task details to GW 554 (640), which forwards the task details to client 552 (642). In some embodiments, if any operation such as downloading fails, this failure can be indicated in the returned task details.
[0078] Once the software download and / or software pre-check have been successfully completed, the update can be applied to the private cloud system 102. Figure 18 An example update process 650 is illustrated, which can be used by a remote cloud management system 112 and / or a private cloud system 102 to apply updates to the private cloud system 102. Process 650 uses update 558, AI service 602, and DSC 604. Process 650 also uses data operations manager 654 to interface with storage device 103, such as with its operating system. Process 650 also uses DSC VM manager 656 to manage the VMs of DSC 604. Furthermore, process 650 involves HV host 662.
[0079] Since DSC 604 is local software that connects (multiple) racks to the remote cloud management system 112, the DSC VM used to implement DSC can be the first target update. Therefore, update 558 obtains the DSC VM version from DSC man 656 (664) and a list of available DSC VM versions from DSC man 656 (666). Update 558 then initiates an update to the DSC VM in DSC man 656 using one of the available DSC versions (such as the latest full release) (668). Each update discussed in process 650 may include task creation, tracking, and / or communication using CM 560, as previously discussed. Figure 16 and Figure 17 The discussion covers software pre-checking and download tasks.
[0080] After completing the DSC VM update, update 558 begins updating the AI tools by obtaining the AI service version (670) from AI service 602. Update 558 can also perform a software pre-check on the AI tools to be downloaded using an orchestrator (e.g., private cloud AI platform orchestrator 122) and a workload cluster of worker nodes 104. If the pre-check is successfully completed, the update initiates the download of the AI tool update using AI service 602 (672). Update 558 then initiates the download of the update on the orchestrator (674) and applies the downloaded update to each workload cluster (676).
[0081] After completing the AI tool update, update 558 initiates an update (678) to the OS of storage device 103 to update the OS to dataops man 654. After completing the storage OS update, update 558 uses DSC 604 to download the HV update package (680). Before, after, or during the download of the HV update package, update 558 may download the server firmware bundle and extract the firmware (682). Then, update 558 performs an iLO firmware update (684) on each HV host 662 in the HV cluster. Update 558 may also perform a firmware update trial run (686) on all HV hosts 662 in the HV cluster. Then, update 558 updates each HV for each HV host 662 (688). Updating each HV may include first placing each HV host 662 into maintenance mode before applying the update. After updating each HV host 662, update 558 causes each HV host to reboot (690), and after rebooting, checks for a version match with the target firmware version (692) to confirm that the firmware update has been successfully completed.
[0082] Then, update 558 updates the firmware control node 106 (694) such as the AI software control plane 110, and reboots the HV host 662 (696). Update 558 verifies success by checking the iLO installation queue to confirm whether the firmware update has been completed after the reboot (698).
[0083] Then, update 558 updates worker node 104 by first putting the worker node into maintenance mode (700), updating the firmware in worker node 104 (702), and removing the worker node from maintenance mode (704).
[0084] In some embodiments, at least some of the previously discussed processes may include more or fewer operations. For example, process 650 may include fewer or more steps in a software update without departing from the teachings of this document. Figure 19Process 720 includes receiving an instruction to update the artificial intelligence (AI) tool stack (e.g., AI stack 400) in private cloud system 102 via a user interface (e.g., private cloud AI API / UI 124) of remote cloud management system 112 (box 722). In response to receiving the instruction, remote cloud management system 112 updates the virtual machine of DSC 604 of private cloud system 102 (box 724). In response to receiving the instruction and updating the virtual machine, remote cloud management system 112 uses the updated virtual machine to update the AI service platform used to deploy the AI tools (box 726). Updating the AI service platform may include obtaining the current version of the AI service platform, pre-checking compatibility with updates to the AI service platform and the health of the AI service platform components, downloading updates using the AI application programming interface (API) of remote cloud management system 112, and installing the updates. Such updates may include updating any part of the AI stack, such as the storage OS, HV, control node firmware, and / or worker node firmware, using any of the techniques discussed with respect to process 650.
[0085] While certain features of this disclosure have been described and illustrated herein, many modifications and variations will occur to those skilled in the art. Therefore, it should be understood that the appended claims are intended to cover all such modifications and variations, provided they fall within the true spirit of this disclosure.
Claims
1. A system comprising: a remote cloud management system comprising a private cloud artificial intelligence (AI) platform orchestrator remote from a private cloud system, the remote cloud management system configured to: interface with the private cloud system; orchestrate AI operations on the private cloud system using the private cloud AI platform orchestrator; and manage AI software installed in the private cloud system.
2. The system of claim 1, wherein the remote cloud management system comprises an AI application programming interface (API) management system configured to manage the AI software in the private cloud system.
3. The system of claim 2, wherein the AI API management system interfaces with the AI software in the private cloud system using a plurality of APIs.
4. The system of claim 2, wherein managing the AI software comprises using the remote cloud management system to remotely update the AI software in the private cloud system.
5. The system of claim 4, wherein the remote cloud management system is configured to interface with the private cloud system using a tunnel implemented using a data service connector of the private cloud system.
6. The system of claim 5, wherein the remote cloud management system is configured to update the data service connector prior to updating the AI software.
7. The system of claim 4, wherein the remote cloud management system is configured to receive a connection from a rack in the private cloud system and complete an initial configuration of the rack.
8. The system of claim 7, wherein the initial configuration comprises updating the AI software.
9. The system of claim 8, wherein updating the AI software comprises sequentially updating a plurality of components of the private cloud system in a hierarchical order based on a selection of an update option.
10. The system of claim 9, wherein the selection of the update option comprises a one-click update selection.
11. A computer-implemented method comprising: receiving, via a user interface of a remote cloud management system, an indication to update a stack of artificial intelligence (AI) tools in a private cloud system, wherein the remote cloud management system is configured to remotely manage the private cloud system; in response to receiving the indication, using, via the remote cloud management system, the remote cloud management system to update a virtual machine of a data service connector (DSC) of the private cloud system; in response to receiving the indication and updating the virtual machine, using the virtual machine to update an AI service platform used to deploy the AI tools, wherein updating the AI service platform comprises: obtaining a current version for the AI service platform; pre-checking compatibility of an update to the AI service platform and health of components of the AI service platform; using an AI application programming interface (API) of the remote cloud management system to download the update; and installing the update.
12. The computer-implemented method of claim 11, wherein the indication comprises a single input to update an entirety of the stack comprising the AI service platform and the virtual machine.
13. The computer-implemented method of claim 11, wherein the private cloud system comprises a local cloud implemented at least partially on a customer’s site, and the remote cloud management system is implemented at one or more sites maintained by a provider of the remote cloud management system.
14. The computer-implemented method of claim 11, comprising: updating an operating system of a storage device of the private cloud system in response to updating the AI service platform.
15. The computer-implemented method of claim 14, comprising: updating a hypervisor of the private cloud system using the remote cloud management system in response to updating the operating system of the storage device of the private cloud system.
16. The computer-implemented method of claim 15, comprising: updating firmware of control nodes and worker nodes in the private cloud system in response to updating the hypervisor.
17. The computer-implemented method of claim 11, wherein updating the virtual machine of the DSC comprises: determining a version of the virtual machine of the DSC; determining one or more available versions for the virtual machine of the DSC; and updating the virtual machine of the DSC to one of the one or more available versions of the virtual machine of the DSC.
18. The computer-implemented method of claim 17, wherein the one of the one or more available versions comprises a latest stable version of the virtual machine.
19. The computer-implemented method of claim 11, wherein installing the update comprises: installing an update to a private cloud AI platform orchestrator of the remote cloud management system; and installing an update to a worker node of the private cloud system used to implement the AI tool.
20. A tangible, non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors of one or more computers, are configured to cause the one or more computers to: present a user interface via a remote cloud management system configured to remotely manage a local cloud system using a tunnel implemented using a data services connector (DSC) of the local cloud system, wherein the local cloud system comprises a plurality of artificial intelligence (AI) tools; receive an indication to update an AI tool; update a virtual machine used to implement the DSC in response to the indication; update an AI service platform used to implement the AI tool using the DSC implemented using the updated virtual machine after completing the update to the virtual machine; update a hypervisor on one or more hypervisor hosts of the local cloud system after completing the update to the AI service platform; and update firmware of control nodes and worker nodes of the local cloud system after completing the update to the hypervisor.