AUTOMATICALLY DEPLOYED INFORMATION TECHNOLOGY (IT) SYSTEM AND METHOD - Patent application

An automated IT system with a core controller and self-assembly rules addresses setup, configuration, and management challenges, enhancing security and scalability while reducing human error and improving documentation.

JP7797459B2Active Publication Date: 2026-01-13NET THUNDER LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023198584
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-07-06
Filing Date
2023-11-22
Publication Date
2026-01-13
Estimated Expiration
2038-12-07

AI Technical Summary

Technical Problem

Current IT systems face challenges in infrastructure setup, configuration, management, and updates, leading to inefficiencies, security vulnerabilities, and difficulties in troubleshooting and scalability, with issues exacerbated by human error and lack of comprehensive documentation.

Method used

An automated IT system setup and management system utilizing a core controller with self-assembly rules, templates, and system state to facilitate flexible, secure, and scalable deployment and configuration of physical and virtual resources, including automated documentation and change management.

Benefits of technology

The system reduces human error, enhances security, improves scalability, and ensures efficient resource utilization and management, providing comprehensive documentation and enabling seamless system updates and reversibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007797459000001
    Figure 0007797459000001
  • Figure 0007797459000002
    Figure 0007797459000002
  • Figure 0007797459000003
    Figure 0007797459000003
Patent Text Reader

Abstract

To provide a scalable computer system that can serve as a turnkey scalable private cloud.SOLUTION: A system 100 has controller logic 205, a global system rule database 210, and a controller 200 which has an IT system state 220 and a template 230. A global system rule includes minimum requirements declaring rules for setting up, configuring, booting, allocating, and managing resources including computing, storage and networking so that the system is in a correct or a desired state, and those requirements include IT tasks which are expected to be completed and an updatable list of expected hardware required to construct a desired system as expected. With the list, the controller can confirm that necessary resources are available.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims priority to U.S. Provisional Patent Application No. 62 / 596,355, filed December 8, 2017, and entitled "Automatically Deployed Information Technology (IT) System and Method," the entire disclosure of which is incorporated herein by reference.

[0002] This patent application also claims priority to U.S. Provisional Patent Application No. 62 / 694,846, filed July 6, 2018, and entitled "Automatically Deployed Information Technology (IT) System and Method," the entire disclosure of which is incorporated herein by reference. [Background technology]

[0003] The demand, usage, and needs for computing have grown exponentially over the past few decades. This combined demand for greater storage, speed, computational power, applications, and accessibility has rapidly transformed the field of computing, providing tools for entities of diverse types and sizes. As a result, the use of public virtual and cloud computing systems has increased, providing more computing resources to a greater number of users and user types. This exponential growth is expected to continue. At the same time, infrastructure setup, management, change management, and updates have become more complex and costly, with increased risks of failure and security. Scalability, or the ability to grow systems over time, has also become a major challenge in the field of information technology.

[0004] Most IT system problems, many of which are performance- and security-related, can be difficult to diagnose and address. Time and resource constraints allowed for system setup, configuration, and deployment can lead to errors and result in future IT problems. Over time, many different administrators may be involved in modifying, patching, or updating IT systems, including users, applications, services, security, software, and hardware. Often, configuration and change documentation and history are inadequate or lost, making it difficult to later understand how a particular system is configured and functions. This can make future changes or troubleshooting difficult. When problems or failures occur, IT configurations and settings can be difficult to recover and reproduce. System administrators can also easily make mistakes, such as issuing incorrect commands or other errors, that can bring down computers and web databases and services. Furthermore, an increased risk of security breaches is common, while changes, updates, and patches to prevent security breaches can cause unwanted downtime.

[0005] Once critical infrastructure is in place, functioning, and operational, the costs or risks may often appear to outweigh the benefits of changing the system. Issues associated with making changes to live IT systems or environments can cause significant, sometimes catastrophic, problems for users or entities that rely on these systems. At the very least, the time it takes to troubleshoot and resolve failures or issues that arise during change management can require significant time, personnel, and financial resources. Technical issues that may arise when changes are made to a live environment can have cascading effects and may not be resolved simply by reversing the changes that were made. Many of these issues contribute to the inability to quickly rebuild systems when failures exist during change management.

[0006] Additionally, bare metal cloud nodes or resources within an IT system may be vulnerable to security issues or may be compromised or accessed by unauthorized users. Hackers, attackers, or unauthorized users may pivot from that node or resource to access or hack other parts of the IT system or the network connected to the node. Bare metal cloud nodes or controllers of an IT system may also be vulnerable through resources connected to application networks that may expose the system to security threats or otherwise compromise the system. According to various exemplary embodiments disclosed herein, an IT system may be configured to interface with the Internet or improve the security of bare metal cloud nodes or resources from application networks with or without connections to external networks.

[0007] According to an exemplary embodiment, an IT system includes bare-metal cloud nodes or physical resources. When the bare-metal cloud nodes or physical resources are powered on, set up, managed, or used, they may be connected to a network with nodes that other people or customers may be using. In-band management may be omitted from the controller, switchable, disconnectable, or filtered. Also, an application network or networks of multiple applications in the system may be disconnected, disconnectable, switchable, or filtered from the controller via the resource(s) through which the application network is coupled to the controller.

[0008] Physical resources, including virtual machines or hypervisors, may also be vulnerable to security issues and may be compromised or accessed by unauthorized users when a hypervisor may be used to pivot to other hypervisors that are shared resources. An attacker may be able to break out of a virtual machine and gain network access to a management system and / or supervisory system via a controller. According to various exemplary embodiments disclosed herein, an IT system may be configured to improve security, where one or more physical resources, including virtual resources on a cloud platform, may be disconnected, disconnectable, filtered, filterable, or disconnected from a controller via an in-band management connection.

[0009] According to an example embodiment, the physical resources of the IT system may include one or more virtual machines or hypervisors, and the in-band management connection between the controller and the physical resource may be omitted from the resource, may be disconnected, may be disconnectable, or may be filtered / filterable. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a schematic diagram of a system in accordance with an exemplary embodiment. [Figure 2A] FIG. 2 is a schematic diagram of an exemplary controller for the system of FIG. 1. [Figure 2B] 1 illustrates an example flow of operation for an example set of storage expansion rules. [Figure 2C] 2B shows an alternative for performing steps 210.1 and 210.2 of FIG. 2B. [Figure 2D] 2B shows an alternative for performing steps 210.1 and 210.2 of FIG. 2B. [Figure 2E] 1 shows an exemplary template. [Figure 2F] 10 illustrates an exemplary process flow of the controller logic for processing a template. [Figure 2G] 2F shows an exemplary process flow for steps 205.11, 205.12, and 205.13. [Figure 2H] 2F shows an exemplary process flow for steps 205.11, 205.12, and 205.13. [Figure 2I] 10 shows another exemplary template. [Figure 2J] 10 illustrates another exemplary process flow of the controller logic for processing a template. [Figure 2K] 1 illustrates an exemplary process flow for managing service dependencies. [Figure 2L] 4 is a schematic diagram of an exemplary image derived from a template in accordance with an exemplary embodiment; [Figure 2M] 1 illustrates an exemplary set of system rules. [Figure 2N] 2C illustrates an exemplary process flow for the controller logic to process the system rules of FIG. 2M. [Figure 2O] 1 illustrates an exemplary process flow for composing storage resources from a group of file system blobs or other files. [Figure 3A] FIG. 2B is a schematic diagram of the controller of FIG. 2A with additional computing resources. [Figure 3B] 4 is a schematic diagram of an exemplary image derived from a template in accordance with an exemplary embodiment; [Figure 3C] 1 illustrates an exemplary process flow for adding resources, such as computing resources, storage resources, and / or networking resources, to a system. [Figure 4A] FIG. 2B is a schematic diagram of the controller of FIG. 2A with added storage resources. [Figure 4B] 4 is a schematic diagram of an exemplary image derived from a template in accordance with an exemplary embodiment; [Figure 5A] FIG. 2B is a schematic diagram of the controller of FIG. 2A with added JBOD and storage resources. [Figure 5B] 1 illustrates an exemplary process flow for adding a storage resource and direct attached storage of a storage resource to a system. [Figure 6A] FIG. 2B is a schematic diagram of the controller of FIG. 2A with added networking resources. [Figure 6B] 4 is a schematic diagram of an exemplary image derived from a template in accordance with an exemplary embodiment; [Figure 7A] 1 is a schematic diagram of a system according to an exemplary embodiment in an exemplary physical deployment. [Figure 7B] 1 illustrates an exemplary process for adding a resource to an IT system. [Figure 7C] 1 illustrates an exemplary process flow for deploying an application across multiple computing resources, multiple servers, multiple virtual machines, and / or multiple sites. [Figure 7D] 1 illustrates an exemplary process flow for deploying an application across multiple computing resources, multiple servers, multiple virtual machines, and / or multiple sites. [Figure 8A] FIG. 1 is a schematic diagram of a system according to an exemplary embodiment in an exemplary deployment. [Figure 8B] 1 illustrates an exemplary process flow for expanding from a single node system to a multi-node system. [Figure 8C] 1 illustrates an exemplary process flow for migrating storage resources to new physical storage resources. [Figure 8D] 1 illustrates an exemplary process flow for migrating virtual machines, containers, and / or processes on a single node of a multi-tenant system to a multi-node system that may have separate hardware for computing and storage. [Figure 8E] 10 illustrates another exemplary process flow for expanding from a single node to multiple nodes in a system. [Figure 9A] 1 is a schematic diagram of a system according to an exemplary embodiment in an exemplary physical deployment. [Figure 9B]4 is a schematic diagram of an exemplary image derived from a template in accordance with an exemplary embodiment; [Figure 9C] Here is an example of installing an application from an NT package: [Figure 9D] FIG. 1 is a schematic diagram of a system according to an exemplary embodiment in an exemplary deployment. [Figure 9E] 1 illustrates an exemplary process flow for adding a virtual computing resource host to an IT system. [Figure 10] FIG. 1 is a schematic diagram of a system according to an exemplary embodiment in an exemplary deployment. [Figure 11A] 1 illustrates a system and method of an exemplary embodiment. [Figure 11B] 1 illustrates a system and method of an exemplary embodiment. [Figure 12] 1 illustrates a system and method of an exemplary embodiment. [Figure 13A] FIG. 1 is a schematic diagram of a system in accordance with an exemplary embodiment. [Figure 13B] FIG. 2 is another schematic diagram of a system in accordance with an exemplary embodiment. [Figure 13C] 3 illustrates an exemplary process flow for a system according to an exemplary embodiment. [Figure 13D] 3 illustrates an exemplary process flow for a system according to an exemplary embodiment. [Figure 13E] 3 illustrates an exemplary process flow for a system according to an exemplary embodiment. [Figure 14A] 1 illustrates an exemplary system in which a main controller deploys controllers in different systems. [Figure 14B] 1 shows an exemplary flow illustrating possible steps for provisioning a controller with a main controller. [Figure 14C] 1 shows an exemplary flow illustrating possible steps for provisioning a controller with a main controller. [Figure 15A] 1 illustrates an exemplary system in which a main controller creates an environment. [Figure 15B]10 illustrates an exemplary process flow for a controller to set up an environment. [Figure 15C] 10 illustrates an exemplary process flow for a controller to set up multiple environments. [Figure 16A] 1 illustrates an exemplary embodiment in which a controller acts as a main controller for setting up one or more controllers. [Figure 16B] 1 illustrates an exemplary system in which environments can be configured to write to other environments. [Figure 16C] 1 illustrates an exemplary system in which environments can be configured to write to other environments. [Figure 16D] 1 illustrates an exemplary system in which environments can be configured to write to other environments. [Figure 16E] 1 illustrates an exemplary system that allows users to purchase new environments that are created by a controller. [Figure 16F] 1 illustrates an exemplary system provided with a user interface for interfacing with an environment created by a controller. [Figure 17A] Here are some examples of change management tasks for new environments: [Figure 17B] Here are some examples of change management tasks for new environments: [Figure 18A] Here are some examples of change management tasks for new environments: [Figure 18B] Here are some examples of change management tasks for new environments: DETAILED DESCRIPTION OF THE INVENTION

[0011] In an effort to provide technical solutions to the needs in the art as described above, the inventors disclose various inventive embodiments relating to systems and methods for information technology that provide automated IT system setup, configuration, maintenance, testing, change management, and / or upgrades. For example, the inventors disclose a controller configured to automatically manage a computer system based on a plurality of system rules, a system state of the computer system, and a plurality of templates. As another example, the inventors disclose a controller configured to automatically manage the physical infrastructure of a computer system based on a plurality of system rules, a system state of the computer system, and a plurality of templates. Examples of automated management that can be performed by a controller may include remotely or locally accessing and changing settings or other information on computers that may run applications or services, configuring an IT system, modifying an IT system, configuring individual stacks within an IT system, creating a service or application, loading a service or application, configuring a service or application, migrating a service or application, modifying a service or application, removing a service or application, cloning a stack to another stack on a different network, creating, adding, removing, setting up, configuring, reconfiguring, and / or modifying resources or system components, automatically adding, removing, and / or restoring resources, services, applications, IT systems, and / or IT stacks, configuring interactions between applications, services, stacks, and / or other IT systems, and / or monitoring the health of IT system components. In an exemplary embodiment, the controller may be embodied as a physical or virtual computing resource, which may be remote or local.Additional examples of controllers that may be employed include, but are not limited to, any or any combination of processes, virtual machines, containers, remote computing resources, applications deployed by other controllers, and / or services. Controllers may be distributed across multiple nodes and / or resources and may be located in other locations or networks.

[0012] IT infrastructures are almost always built from discrete hardware and software components. Hardware components typically include servers, racks, power supplies, interconnects, display monitors, and other communications equipment. The methods and techniques for selecting and interconnecting these discrete components are highly complex due to the numerous possible configurations that operate with varying degrees of efficiency, cost-effectiveness, performance, and security. Individual technicians / engineers skilled in connecting these infrastructure components are expensive to hire and train. Furthermore, the numerous possible iterations of hardware and software complicate maintenance and updates. This creates additional challenges when updates cannot be performed by the individuals and / or engineering firms that originally installed the IT infrastructure. Software components, such as operating systems, are either generically designed to work with a wide range of hardware or highly specialized for a particular component. Complex plans or blueprints are often written and implemented. Change, growth, scaling, and other challenges require updating the complex plans.

[0013] Some IT users purchase cloud computing services from a growing supplier industry, but this does not solve infrastructure setup problems and challenges, but rather shifts them from the IT user to the cloud service provider. Furthermore, large cloud service providers address infrastructure setup challenges and challenges in ways that can reduce flexibility, customization, scalability, and rapid adoption of new hardware and software technologies. Cloud computing services also do not offer ready-to-use bare-metal setup, configuration deployment, and updates, nor do they enable migration to, from, or between bare-metal and virtualized IT infrastructure components. These and other limitations of cloud computing services can result in numerous computing, storage, and networking inefficiencies. For example, speed or latency inefficiencies in computing and networking can be incurred by cloud services or in applications or services that use cloud services.

[0014] The systems and methods of the exemplary embodiments provide for the deployment, utilization, and management of a novel and unique IT infrastructure. According to the exemplary embodiments, the complexity of resource selection, installation, interconnection, management, and updates is rooted in a core controller system and its parameter files, templates, rules, and IT system state. The system includes a set of self-assembly and operational rules that configure components to self-assemble rather than requiring a technician to assemble, connect, and manage them. Furthermore, the systems and methods of the exemplary embodiments use self-assembly rules to enable greater customization, scalability, and flexibility without requiring external blueprints, typical of today. They also enable efficient resource use and reuse.

[0015] Systems and methods are provided that ameliorate many of the problems and issues of current IT systems, whether wholly or partially physical or virtual. The systems and methods of the exemplary embodiments provide a structure that allows flexibility, reduces variability and human error, and can improve system security.

[0016] While some individual solutions may exist to one or more of the problems of current IT systems, such solutions do not comprehensively address the multitude of problems that are addressed by the exemplary embodiments described herein. Moreover, while such existing solutions may address certain problems, they may exacerbate others.

[0017] Current challenges addressed include, but are not limited to, issues related to setup, configuration, infrastructure deployment, asset tracking, security, application deployment, service deployment, maintenance and compliance documentation, maintenance, scaling, resource allocation, resource management, load balancing, software failures, software and security updates / patching, testing, IT system recovery, change management, and hardware updates.

[0018] As used herein, IT systems may include, but are not limited to, servers, virtual and physical hosts, databases and database applications, such as, but not limited to, IT services, business computing services, computer applications, customer-facing applications, web applications, mobile applications, back-end, case number management, customer tracking, ticketing, business tools, desktop management tools, accounting, email, documentation, compliance, data storage, backup, and / or network management.

[0019] One of the problems users may face before setting up an IT system is predicting their infrastructure needs. Users may not know how much storage, computing power, or other requirements they will need, either initially or over time as they grow or change. According to exemplary embodiments, IT systems and infrastructures allow for flexibility in that, as the system needs change, the self-deploying infrastructure (both physical and / or virtual) of the exemplary embodiments can be used to automatically add, remove, or reallocate resources within the infrastructure later. Thus, the challenge of predicting future needs presented at the time of system setup is addressed by providing the ability to add to the system using global rules, templates, and system states, and by tracking changes to such rules, templates, and system states.

[0020] Other challenges may relate to correct configuration, uniformity of configuration, interoperability, and / or interdependencies, which may include, for example, future incompatibilities due to changes in configured system elements or their configuration over time. For example, when an IT system is initially set up, there may be missing elements or some elements may have been misconfigured. Also, for example, when iterations of elements or infrastructure components are set up, there may be a lack of uniformity between the iterations. Changes to the system may require a reconfiguration. A difficult choice is presented between optimal configuration and flexibility for future infrastructure changes. According to an exemplary embodiment, when a system is initially deployed, global system rules are used to self-deploy configurations from templates to infrastructure components, resulting in uniform, repeatable, or predictable configurations, enabling optimal configurations. While such initial system deployments may be performed on physical components, subsequent components may be added or modified, which may or may not be physical. Furthermore, while such initial system deployments may be performed on physical components, subsequent environments may be cloned from the physical structure, which may or may not be physical. This allows for optimal system configuration while minimizing future problematic changes.

[0021] During the deployment phase, there are typically challenges with interoperability of bare-metal and / or software-defined infrastructure. There may also be challenges with interoperability of software with other applications, tools, or infrastructure. These may include, but are not limited to, challenges with deployed products from different vendors. The inventors disclose an IT system that can provide infrastructure interoperability, regardless of whether the infrastructure is bare-metal, virtual, or any combination thereof. Thus, interoperability, i.e., the ability of parts to work together, can be built into the disclosed infrastructure deployment, where the infrastructure is automatically configured and deployed. For example, different applications may depend on each other and reside on different hosts. To enable such applications to interact with each other, the controller logic, templates, system states, and system rules described herein contain information and configuration instructions used to configure and track application interdependencies. Thus, the infrastructure features described herein provide a way to manage how each application or service interacts with each other. For example, ensuring that an email service communicates properly with an authentication service and / or ensuring that a groupware service communicates properly with an email service. Furthermore, such management can go down to the infrastructure level, making it possible to track, for example, how computing resources are communicating with storage resources. Otherwise, the complexity of IT systems would increase by O(n n ) may rise.

[0022] As disclosed, automatic deployment of resources does not require pre-configuration of operating system software due to the controller's ability to deploy based on global system rules, templates, and IT system state / system self-awareness. According to exemplary embodiments, a user or IT professional may not need to know whether resource additions, allocations, or reallocations are coordinated to ensure interoperability. Additional resources according to exemplary embodiments may be automatically added to the network.

[0023] Using an application typically requires many different resources, including computing, storage, and networking. It also requires interoperability between resources and system components, including knowledge of what's in place and running, and interoperability with other applications. Applications may need to connect to other services to retrieve configuration files and ensure all components work properly together. Configuring an application can be time- and resource-intensive. If there are interoperability issues with other applications, configuring an application can have cascading effects on the rest of the infrastructure, potentially leading to outages or compromises. The inventors disclose automated application deployment to address these concerns. Thus, as the inventors disclose, applications can self-deploy by intelligently configuring themselves using knowledge of the system's current state, reading from IT system state, global system rules, and templates. Furthermore, according to an exemplary embodiment, pre-deployment testing of the configuration can be performed using the change management functionality described herein.

[0024] Another issue addressed by the exemplary embodiments relates to problems that may arise with intermediate configurations when it is desirable to switch to a different vendor or other tool. According to one aspect of the exemplary embodiments, template translation is provided between the rules and templates of the controller and the application templates of a particular vendor, allowing the system to automatically change vendors of software or other tools.

[0025] Many security issues stem from misconfigurations, patching failures, and the inability to test patching before deployment. Security issues can often arise during the configuration phase of setup. For example, misconfigurations can leave sensitive applications exposed to the Internet or allow email forgery from a mail server. The inventors disclose a system setup that automatically configures itself to protect against attackers, avoid unnecessary exposure to attackers, and provide security engineers and application security architects with more knowledge about the system. Automation reduces security flaws due to human error or misconfiguration. The disclosed infrastructure also provides introspection between services, enables rule-based access, and can limit communication between services to only what is actually necessary. The inventors disclose a system and method capable of securely testing patches before deployment, as described, for example, with respect to change management.

[0026] Documentation is often a problematic area of ​​IT management. During setup and configuration, the primary goal can typically be getting components to work together. This usually involves troubleshooting and trial-and-error processes, and it can be difficult to know what actually made the system work. While the exact commands executed are usually documented, the troubleshooting or trial-and-error process that may have resulted in a functioning system is often poorly documented or not documented at all. Problems or deficiencies in documentation can create problems with audit trails and audits. Documentation issues that arise can create problems when demonstrating compliance. Often, compliance issues may not be well known when building a system or its components. Applicable compliance determinations may only be known after the IT system is set up and configured. Thus, documentation is essential for audits and compliance. The inventors disclose a system that provides automatically documented setup and configuration, including a global system rules database, templates, and an IT system state database. Every configuration that occurs is recorded in the database. According to an exemplary embodiment, the automatically documented configuration provides an audit trail and can be used to demonstrate compliance. Inventory management may involve automatically documenting and tracking information.

[0027] Another challenge arising from the setup, configuration, and operation of IT systems concerns hardware and software inventory management. For example, it is typically important to know how many servers there are, whether they are up and running, what their functions are, which rack each server is in, which power supply is connected to which server, which network card and port each server uses, which IT system components are running on, and many other important matters. In addition to inventory information, passwords and other sensitive information used for inventory management must be effectively managed. Collecting and maintaining this information is a time-consuming task, especially in large IT systems, data centers, or data centers where equipment changes frequently, and it is often managed manually or using various software tools. Compliant protection of secure passwords is a significant risk factor that can become a key issue in ensuring a secure computing environment. The inventors disclose an IT system in which the collection and maintenance of the inventory and operational status of all servers and other components is automatically updated, stored, and protected as part of the IT system state, global system rules, templates, and controller logic of the controller.

[0028] In addition to issues related to IT system setup and configuration, the inventors disclose an IT system that can address problems and issues arising in IT system maintenance. For example, many problems arise with the continued functioning of a data center with hardware failures, such as power failures, memory failures, network failures, network card failures, and / or CPU failures, among others. Migrating a host during a hardware failure introduces additional failures. Therefore, the inventors disclose dynamic resource migration, e.g., migrating resources from one resource provider to another when a host goes down. In such situations, according to exemplary embodiments, the IT system can be migrated to another server, node, or resource, or to another IT system. The controller can report the status of the system. Data replicas are located on other hosts with known, automatically set-up configurations. When a hardware failure is detected, any resources that the hardware may have provided can be automatically migrated after automatically detecting the failure.

[0029] A key issue with many IT systems is scalability. Growing businesses or other organizations typically add or reconfigure IT systems as they grow and their needs change. Problems arise when more resources are needed for existing IT systems, such as adding hard drive space, storage capacity, CPU processing, more network infrastructure, more endpoints, more clients, and / or more security. Configuration, setup, and deployment also present challenges when different services and applications or infrastructure changes are required. According to exemplary embodiments, data centers can be automatically scaled. Nodes or resources can be dynamically and automatically added or removed from a pool of resources. Resources added and removed from a resource pool can be automatically allocated or reallocated. Services can be provisioned and quickly moved to new hosts. A controller can discover and dynamically add more resources to a resource pool and know where to allocate / reallocate resources. A system according to exemplary embodiments can scale from a single-node IT system to a scaled system requiring numerous physical and / or virtual nodes or resources across multiple data centers or IT systems.

[0030] The inventors disclose a system that enables flexible resource allocation and management. The system includes computing, storage, and networking resources that can be in a resource pool and dynamically allocated. A controller can recognize new nodes or hosts on the network and then configure them to become part of the resource pool. For example, when a new server is plugged in, the controller can configure it as part of the resource pool and add it to the resources, which can then be dynamically started. Nodes or resources can be discovered by the controller and added to different pools. Resource requests can be made, for example, via API requests to the controller. The controller can then deploy or allocate the needed resources from the pool according to rules. This allows the controller and / or applications via the controller to load balance and dynamically distribute resources based on the needs of the requests.

[0031] Examples of load balancing include, but are not limited to, deploying new resources in the event of a hardware or software failure, deploying one or more instances of the same application in response to an increase in user load, and deploying one or more instances of the same application in response to an imbalance in storage, computing, or networking demand.

[0032] Problems involving making changes to a live IT system or environment can create significant, sometimes catastrophic, issues for users or entities that rely on these systems to operate consistently. These outages not only represent a potential loss in system use, but also an economic loss due to data loss and the considerable resources of time, personnel, and money required to resolve the issue. Problems can be exacerbated by the difficulty of rebuilding the system if the configuration documentation is incorrect or there is a lack of understanding of the system. Because of this issue, many IT system users are reluctant to patch their IT resources to eliminate known security risks, thus leaving them more vulnerable to security breaches.

[0033] Many problems that arise in maintaining IT systems are related to software failures due to change management or control that may require configuration. Situations in which such failures may occur include, but are not limited to, upgrading to a new software version, migrating to different software, changing passwords or authentication management, and switching between services or between different providers of a service.

[0034] Manually configured and maintained infrastructure is typically difficult to recreate. Recreating the infrastructure can be important for several reasons, including, but not limited to, rolling back problematic changes, power outage, or other disaster recovery. It is difficult to diagnose problems in manually configured systems. It is difficult to recreate manually configured and maintained infrastructure. Additionally, system administrators can easily make mistakes, such as an incorrect command that is known to bring down a computer system.

[0035] Making changes to a live IT system or environment can cause significant, sometimes catastrophic, problems for users or entities that rely on these systems to operate consistently. These outages not only represent a potential loss in system use, but such outages can also cause economic loss due to data loss as well as the considerable resources of time, personnel, and money required to resolve the problem. Problems can be exacerbated by difficulties in rebuilding the system if the configuration documentation is incorrect or there is a lack of understanding of the system. Also, it is often very difficult to restore a system to its previous state after a significant or major change.

[0036] Furthermore, technical issues that may arise when changes are made to a live environment can have cascading effects. These cascading effects can make it difficult, and in some cases impossible, to revert to a state prior to the change. Thus, even if a problem with an implemented change requires that the change be reverted, the state of the system has already been altered. In recent years, reverting infrastructure and system administration errors, as well as incomplete changes to a production environment, has been described as an unsolvable problem. Furthermore, testing changes to a system before deployment to a live environment is known to be problematic.

[0037] Accordingly, the inventors disclose several exemplary embodiments of systems and methods configured to revert changes to a running system to a state prior to the changes. Furthermore, the inventors disclose systems and methods configured to enable significant reversion of the state of a system or environment that has undergone a running change, which may prevent or ameliorate one or more of the problems discussed above.

[0038] According to a variant of an exemplary embodiment, an IT system has a complete system knowledge through global system rules, templates, and IT system state. The infrastructure can be cloned using the complete system knowledge. A system or system environment can be cloned as a software-defined infrastructure or environment. The system environment, including the volatile database in use, called the production environment, can be written to a non-volatile read-only database and used as the development environment in the development and testing process. Desired changes can be made to the development environment and tested there. A user or controller logic can make changes to the global rules and create a new version. Rule versions can be tracked. Then, according to another aspect of the exemplary embodiment, the newly developed environment can be automatically implemented. The previous production environment can also be retained or fully functional, allowing modifications to a previous state of the production environment without data loss. The development environment can then be booted with the new specifications, rules, and templates, and the database or system can be synchronized with the production database and switched to a writable database. The original production database can then be switched to a read-only database, and the system can be reverted to it if recovery is required.

[0039] With respect to software upgrades or patching, new hosts may be deployed if services are detected that require an upgrade or patch. In the event of a failure due to the upgrade or patch, new services may be deployed if it is possible to revert the changes as described above.

[0040] Hardware upgrades are important in many situations, especially when the latest hardware is essential. One example of this type of situation occurs in the high-frequency trading industry, where IT systems that take advantage of millisecond speeds can enable users to achieve superior trading results and profits. In particular, issues arise in ensuring interoperability with the current infrastructure, such that the new hardware understands how to communicate using protocols and work with the existing infrastructure. In addition to ensuring component interoperability, the components also need to be integrated with the existing setup.

[0041] 1, an exemplary embodiment of an IT system 100 is shown. System 100 may be one or more types of IT systems, including but not limited to those described herein.

[0042] The user interface (UI) 110 is shown coupled to the controller 200 via an application program interface (API) application 120, which may or may not reside on a standalone physical or virtual server. The controller 200 may be deployed on one or more processors and one or more memories to perform any control operations described herein. Instructions executed by the processor(s) to perform such control operations may reside on a non-transitory computer-readable storage medium, such as processor memory. The API 120 may include one or more API applications, which may be redundant and / or operate in parallel. The API application 120 receives requests to configure system resources, parses the requests, and passes them to the controller 200. The API application 120 receives one or more responses from the controller, parses the response(s), and passes them to the UI (or application) 110. Alternatively, or in addition, applications or services may communicate with the API application 120. The controller 200 is coupled to computing resource(s) 300, storage resource(s) 400, and networking resource(s) 500. Resources 300, 400, 500 may or may not reside on a single node. One or more of resources 300, 400, 500 may be virtual. Resources 300, 400, 500 may or may not reside on multiple nodes, or in various combinations on multiple nodes. Physical devices may include one or more or each of resource types, including, but not limited to, computing resources 300, storage resources 400, and networking resources 500. Resources 300, 400, 500 may also include pools of resources, whether or not located in different physical locations and whether or not virtual. Bare metal computing resources may also be used to enable the use of virtual or containerized computing resources.

[0043] In addition to known definitions of a node, a node as used herein may be any system, device, or resource connected to a network(s) or other functional unit that performs a function on a standalone or networked device. A node may also include, for example, but is not limited to, a server, a service / application / multiple services on a physical or virtual host, a virtual server, and / or multiple or single services running on a multi-tenant server or within a container.

[0044] The controller 200 may include one or more physical or virtual controller servers, which may also be redundant and / or operate in parallel. The controller may operate on a physical or virtual host that functions as a computing host. As an example, the controller may be configured as a controller operating on a host that serves other purposes, e.g., because it has access to sensitive resources. The controller may receive requests from the API application 120, analyze the requests, optimize task distribution to other resources, instruct the other resources, monitor and receive information from the resources, maintain a history of system states and changes, and communicate with other controllers in the IT system. The controller may also include the API application 120.

[0045] A computing resource, as defined herein, may include a single computing node, real or virtual, or a resource pool containing one or more computing nodes. A computing resource or computing node may include one or more physical or virtual machine or container hosts that may host one or more services or run one or more applications. A computing resource may be on hardware designed for multiple purposes, including, but not limited to, computing, storage, caching, networking, and specialized computing, including, but not limited to, GPUs, ASICs, coprocessors, CPUs, FPGAs, and other specialized computing methods. Such devices may be added using PCI Express switches or similar devices and may be dynamically added in such manner. A computing resource or computing node may include or run one or more hypervisors or container hosts that include multiple different virtual machines that run services or applications or may be virtual computing resources. A computing resource may be focused on providing computing functionality but may also include data storage and / or networking functions.

[0046] As defined herein, a storage resource may include a storage node or a pool of storage resources. A storage resource may include any data storage medium, e.g., high-speed, low-speed, hybrid, cache, and / or RAM. A storage resource may include one or more types of networks, machines, devices, nodes, or any combination thereof, which may or may not be directly connected to other storage resources. According to aspects of an exemplary embodiment, a storage resource may be bare metal or virtual, or a combination thereof. A storage resource may be focused on providing storage functionality, but may also include computing and / or networking functionality.

[0047] The networking resource(s) 500 may include a single networking resource, multiple networking resources, or a pool of networking resources. The networking resource(s) may include physical or virtual device(s), tool(s), switch, router, or other interconnections between system resources, or applications for managing networking. Such system resources may be physical or virtual and may include computing, storage, or other networking resources. The networking resources may provide connectivity between external networks and application networks and may host core network services, including, but not limited to, DNS, DHCP, subnet management, Layer 3 routing, NAT, and other services. Some of these services may be deployed on computing, storage, or networking resources on physical or virtual machines. The networking resources may utilize one or more fabrics or protocols, including, but not limited to, Infiniband, Ethernet, RoCE, Fibre Channel, and / or Omnipath, and may include interconnections between multiple fabrics. The networking resources may or may not be SDN-enabled. The controller 200 may be able to configure the topology of the IT system by directly modifying the networking resources 300 using SDN, VLANs, etc. The networking resources may be focused on providing networking functionality, but may also comprise computing and / or storage functionality.

[0048] As used herein, an application network refers to networking resources, or any combination thereof, for connecting or coupling applications, resources, services, and / or other networks, or for connecting users and / or clients to applications, resources, and / or services. An application network may include a network used by servers to communicate with other application servers (physical or virtual) and to communicate with clients. An application network may communicate with machines or networks external to system 100. For example, an application network may connect a web front end to a database. Users may connect to web applications via the Internet or other networks that may or may not be managed by a controller.

[0049] According to an example embodiment, computing, storage, and networking resources 300, 400, 500 may each be automatically added, removed, set up, allocated, reallocated, configured, reconfigured, and / or deployed by controller 200. According to an example embodiment, additional resources may be added to the resource pool.

[0050] While illustrated is a user interface 110, such as a Web UI or other user interface through which a user 105 may access and interact with the system, alternatively or additionally, applications may communicate or interact with the controller 200 via API application(s) 120 or in another manner. For example, a user 105 or application may send requests including, but not limited to, building an IT system, building individual stacks within an IT system, creating a service or application, migrating a service or application, modifying a service or application, deleting a service or application, cloning a stack to another stack on a different network, creating, adding, deleting, setting up or configuring, reconfiguring a resource or system component.

[0051] 1 may include servers having connections or other communication interfaces with various elements, components, or resources, which may be physical or virtual, or any combination thereof. According to a variant, the system 100 shown in FIG. 1 may include bare metal servers having connections.

[0052] As described in more detail herein, the controller 200 may be configured to add resources, allocate resources, manage resources, and update available resources by powering on resources or components and automatically setting up, configuring, and / or controlling the boot-up of resources. The power-on process may begin with powering on the controller so that the order of devices booted is consistent and independent of the user powering on the devices. This process may also include detecting powered-on resources.

[0053] 2A-10, a controller 200, controller logic 205, a global system rules database 210, an IT system state 220, and a template 230 are shown.

[0054] System 100 includes global system rules 210. Global system rules 210 may declare rules for setting up, configuring, booting, allocating, and managing resources, which may include computing, storage, and networking, among other things. Global system rules 210 include minimum requirements for system 100 to be in a correct or desired state. These requirements may include IT tasks expected to be completed and an updatable list of expected hardware needed to predictably build the desired system. The updatable list of expected hardware allows the controller to verify that the required resources are available (e.g., before initiating a rule or using a template). Global rules may include lists of operations required for various tasks and corresponding instructions related to sequencing of operations and tasks. For example, rules may specify the order in which components are powered on, the order in which resources, applications, and services are booted, dependencies, when various tasks should be initiated, such as when to initiate an application load, configuration, start, reload, or hardware update. The rules 210 may also include, for example, one or more of: a list of resource allocations required for applications and services; a list of templates that may be used; a list of applications to load and how to configure them; a list of services to load and how to configure a list of application networks and which applications fit into which networks; a list of various application-specific configuration variables and user-specific application variables; an expected state that allows the controller to check the system state to ensure that the state is as expected and that the results of each instruction are as expected; and / or a version list that includes a list of rule changes (e.g., snapshots) that may allow tracking of rule changes and the ability to test or revert to different rules in different situations. The controller 200 may be configured to apply the global system rules 210 to the IT system 100 on the physical resources.The controller 200 may be configured to apply the global system rules 210 to the IT system 100 on virtual resources. The controller 200 may be configured to apply the global system rules 210 to the IT system 100 on a combination of physical and virtual resources.

[0055] FIG. 2M illustrates an exemplary set of system rules 210, which may take the form of global system rules. The exemplary set of system rules 210 illustrated in FIG. 2M may be loaded into controller 200 or derived by querying system state (see 210.1). In the example of FIG. 2M, system rules 210 include a set of instructions, which may take the form of configuration routines 210.2, and also include data 210.3 for creating and / or recreating an IT system or environment. Configuration rules within system rules 210 may know how to find templates 230 via required template list 210.7 (templates 230 may reside in a file system, disk, storage resource, or may be located within the system rules). Controller logic 205 may also look for templates 230 before processing and enable system rules 210 after verifying their presence. System rules 210 may include subsets 210.15 of system rules, which may be executed as part of configuration routines 210.2.

[0056] Subsystem rules 210.15 can also be used, for example, as a tool for building a system of integrated IT applications (where they are processed by system rule execution routine 210.16 to update the system state and current configuration rules to reflect the addition of 210.15). Subsystem rules 210.15 can also be located elsewhere and loaded into system state 220 by user interaction. For example, one could have subsystem rules 210.15 as a playbook, making them available and operational (the global system rules 210 are then updated so the playbook can be replayed if one wants to clone a system).

[0057] Configuration routine 210.2 may be a set of instructions used to build a system. Configuration routine 210.2 may also include subsystem rules 210.15 or system state pointers 210.8 if desired by the implementer. When running configuration routine 210.2, controller logic 205 processes a series of templates in a specific order (210.9), optionally allowing parallel deployment while maintaining proper dependency processing (210.12). Configuration routine 210.2 may optionally invoke API calls 210.10, which may set configuration parameters 210.5 for applications that may be configured by processing templates according to 210.9. Additionally, required services 210.11 are services that must be running when the system makes API call(s) 210.10.

[0058] Routines 210.2 may include procedures, programs, or methods for data loading (210.13) for volatile data 210.6, including but not limited to copying data, transferring databases to computing resources, pairing computing resources with storage resources, and / or updating system state 220 with the location of volatile data 210.6. Pointers (see 210.4) to the volatile data may be maintained with data 210.3 to locate volatile data that may be stored elsewhere. Data loading routines 210.13 may also be used to load configuration parameters 210.5 if they are located in a non-standard data store (e.g., contained in a database).

[0059] System rules 210 may also include resource lists 210.18, which can dictate which components are assigned to which resources and allow controller logic 205 to determine whether appropriate resources and / or hardware are available. System rules 210 may also include alternative hardware and / or resource lists 210.19 for alternative deployments (e.g., for a development environment where software engineers may want to perform demonstration testing but do not want to allocate an entire data center). System rules may also include data backup / standby routines 210.17, which provide instructions on how to back up the system and use standby for redundancy.

[0060] After any action is taken, the system state 220 may be updated and the query (which may include writing) may be saved as a system state query 210.14.

[0061] Figure 2N illustrates an example process flow for controller logic 205 processing system rules 210 (or subsystem rules 210.15) of Figure 2M. In step 210.20, controller logic 205 checks to ensure that appropriate resources are available (see 210.18 in Figure 2M). If not, alternative configurations may be checked in step 210.21. A third option may include the user being prompted to select an alternative configuration that can be supported by template 230 referenced in list 210.7 of Figure 2M.

[0062] In step 210.22, the controller logic may then check the computing resource (or any appropriate resource) to gain access to the volatile data. This may involve connecting to a storage resource or adding a storage resource to the system state 220. In step 210.23, the configuration routines are then processed, and as each routine is processed, the system state 220 is updated (step 210.24). The system state 220 may also be queried to check whether a particular step is finished before proceeding (step 210.25).

[0063] A configuration routine processing steps as shown in FIG. 210.23 may include any of the procedures in 210.26 (or a combination thereof). It may also include other procedures. For example, processing in 210.26 may include processing templates (210.27), loading configuration data (210.28), loading static data (210.29), loading dynamic volatile data (210.30), and / or binding services, applications, subsystems, and / or environments (210.31). Such procedures in 210.26 may be repeated in a loop or may be executed in parallel so that some system components can be independent and other components can be independent. Controller logic, service dependencies, and / or system rules may dictate which services can depend on each other and may bind services to further build out the IT system from system rules.

[0064] Global system rules 210 may also include storage expansion rules. For example, storage expansion rules may provide a set of rules for automatically adding storage resources to existing storage resources in the system. In addition, trigger points may be provided that know when an application running on a computing resource(s) requests storage expansion (or allow controller 200 to know when to expand the storage of a computing resource or application). Controller 200 may allocate and manage new storage resources and may merge or consolidate storage resources with existing storage resources for a particular running resource.

[0065] Such a particular execution resource may be, but is not limited to, a computing resource within the system, an application running a computing resource within the system, a virtual machine, a container, or a physical or virtual compute host, or a combination thereof. The execution resource may signal to the controller 200 that it is running out of storage space, for example, through a storage space query. In-band management connection 270, SAN connection 280, or any networking or coupling to the controller 200 may be used in such a query. Out-of-band management connection 260 may also be used.

[0066] For resources that are not running, those storage expansion rules (or a subset of those storage expansion rules) may also be used.

[0067] Storage expansion rules dictate how to identify, connect, and set up new storage resources within the system. The controller registers new storage resources in the system state 220, informing the execution resources where the storage resources reside and how to connect to them. The execution resources use such registration information to connect to the storage resources. The controller 200 may merge the new storage resources with existing storage resources, or it may add the new storage resources to a volume group.

[0068] FIG. 2B shows an exemplary flow of operation of an exemplary set of storage expansion rules. In step 210.41, the executing resource determines that storage is low based on a trigger point or otherwise. In step 210.42, the executing resource connects to the controller 200 via an in-band management connection 270, a SAN connection 280, or another type of connection visible to the operating system. Through this connection, the executing resource can notify the controller 200 that storage is low. In step 210.43, the controller configures the storage resource to increase storage capacity for the executing resource. In step 210.44, the controller provides the executing resource with information about where the newly configured storage resource is located. In step 210.45, the executing resource connects to the newly configured storage resource. In step 210.46, the controller adds a mapping of the location of the new storage resource to the system state 220. The controller may then add the new storage resource to the volume group assigned to the running resource (step 210.47), or the controller may add the assignment of the new storage resource to the running resource to the system state 220 (step 210.48).

[0069] FIG. 2C illustrates an alternative embodiment for performing steps 210.41 and 210.42 in FIG. 2B. In steps 210.50, the controller sends a key command over the out-of-band management connection 260 to view a monitor or console for storage status updates for the executing resource. For example, the monitor may be an IPMI console through which the screen can be reviewed via the out-of-band connection 260. By way of example, the out-of-band connection 260 may be connected to a USB port as a keyboard / mouse or to a VGA monitor port. In step 210.51, the executing resource displays information on the screen. In step 210.52, the controller then reads the information presented on the monitor or console via the out-of-band management connection 260 and a screen scrape or similar operation, which may indicate a low storage condition based on a trigger point.

[0070] The process flow may then continue with step 210.43 of FIG. 2B.

[0071]

[0047] Figure 2D illustrates another alternative embodiment for performing steps 210.41 and 210.42 in Figure 2B. In step 210.55, the execution resource automatically displays information on a monitor or console for retrieval by the controller. In step 210.56, the controller automatically, periodically, or continuously retrieves the monitor or console to check the execution resource. In response to this retrieval, the controller determines that the execution resource is low on storage (step 210.57). Process flow may then continue with step 210.43 of Figure 2B.

[0072] The controller 200 also includes a library of templates 230, which may include bare metal and / or service templates. These templates may include, but are not limited to, email, file storage, voice over IP, software accounting, software XMPP, wiki, version control, account authentication management, and third-party applications that may be configurable through a user interface. A template 230 may have an association with a resource, application, or service, or it may serve as a recipe that defines how such a resource, application, or service is integrated into the system.

[0073] As such, a template may include an established set of information used to create, configure, and / or deploy a resource or an application or service loaded onto the resource. Such information may include, but is not limited to, a kernel, an initrd file, a file system or file system image, a file, a configuration file, a configuration file template, information used to determine the appropriate setup for different hardware and / or computing backends, and / or other available options for configuring resources to run applications and / or operating system images that enable and / or facilitate the creation, boot, or execution of the application.

[0074] A template may contain information that can be used to deploy an application on multiple supported hardware types and / or compute backends, including, but not limited to, multiple physical server types or components, multiple hypervisors running on multiple hardware types, and container hosts that can be hosted on multiple hardware types.

[0075] A template may derive a boot image for an application or service to run on a computing resource. Templates and images derived from the templates may be used to create applications, deploy applications or services, and / or adjust resources for various system functions that enable and / or facilitate application creation. A template may have variable parameters in files, file systems, and / or operating system images that can be overridden by configuration options from either default settings or settings provided by a controller. A template may have configuration scripts used to configure applications or other resources, which may utilize configuration variables, configuration rules, and / or default rules or variables, and these scripts, variables, and / or rules may include specific rules, scripts, or variables for specific hardware or other resource-specific parameters, e.g., hypervisor (if virtual), available memory. A template may have files in the form of binary resources, binary resources, or compilable source code that yield hardware or other resource-specific parameters, a specific set of binary resources, or source code with compilation instructions for specific hardware or other resource-specific parameters, e.g., hypervisor (if virtual), available memory. A template may contain a set of information that is independent of what is running on the resource.

[0076] A template may include a base image. The base image may include a base operating system file system. The base operating system may be read-only. The base image may also include basic tools of an operating system independent of what is running on it. The base image may include a base directory and operating system tools. A template may include a kernel. The kernel or kernels may include an initrd kernel or multiple kernels configured for different hardware and resource types. An image may be derived from a template ad loaded or deployed on one or more resources. The loaded image may also include boot files, such as the corresponding template kernel or initrd kernel.

[0077] The image may include template file system information that can be loaded into resources based on the template. The template file system may configure an application or service. The template file system may include a shared file system that is common to all resources or similar resources, for example, to conserve storage space where the file system is stored or to facilitate the use of read-only files.

[0078] A template file system or image may contain a set of files that are common to the services being deployed. The template file system may be preloaded on the controller or downloaded. The template file system may be updated. The template file system may allow for relatively quick deployment because it may not need to be rebuilt. Sharing a file system with other resources or applications may allow for reduced storage because files are not unnecessarily replicated. This may also allow for easier recovery from failures because only files that differ from the template file system need to be restored.

[0079] The template boot file may contain a kernel and / or an initrd or similar file system used to assist the boot process. The boot file can boot an operating system and set up a template file system. The initrd may contain a small temporary file system with instructions on how to set up the template so that it can be booted.

[0080] The template may further include BIOS settings. The template BIOS settings may be used to set optional settings for running an application on a physical host. If used, out-of-band management 260, as described herein with respect to FIGS. 1-12, may then be used to boot the resource or application. The physical host may boot the resource or application using the out-of-band management network 260 or a CD-ROM. The controller 200 may set application-specific BIOS settings defined in such a template. The controller 200 may use the out-of-band management system to make direct BIOS changes through APIs specific to a particular resource. Settings may be verified through a console and image recognition. Thus, the controller 200 may use console functionality and make BIOS changes through a virtual keyboard and mouse. The controller may also use a UEFI shell or type directly into the console, verifying successful results, and using image recognition to type commands correctly and ensure successful setting changes. If there is a bootable operating system available for BIOS changes or updates to a particular BIOS version, the controller 200 may remotely load a disk image or ISO boot where the operating system runs an application that updates the BIOS and allows configuration changes in a trusted manner.

[0081] A template may also include a template-specific list of supported resources or a list of resources required to run a particular application or service.

[0082] The template image or a portion of the image or template may be stored in the controller 200 , or the controller 200 may move or copy it to the storage resource 410 .

[0083] FIG. 2E shows an example template 230. The template contains all the information needed to create an application or service. The template 230 may also contain information about different hardware types, alternative data, files, and binaries that provide similar or identical functionality. For example, there may be file system blobs 232 for / usr / bin and / bin, with binaries 234 compiled for different architectures. The template 230 may also contain daemons 233 or scripts 231. The daemons 233 are binaries or scripts that can run at boot time when the host is powered on or ready. In some cases, the daemons 233 may expose APIs that are accessible by a controller and allow the controller to change the host's configuration (and the controller can subsequently update active system rules). The daemons may be powered down and restarted through out-of-band management 260 or in-band management 270, discussed above and below. These daemons may also expose generic APIs to provide dependent services for the new service (e.g., a generic web server API that communicates with an API controlling nginx or Apache). Script 231 may be an installation script that can be executed during or after booting an image, or after starting a daemon or enabling a service.

[0084] Template 230 may also contain a kernel 235 and a pre-boot file system 236. Template 230 may also include multiple kernels 235 and one or more pre-boot file systems (such as an initrd or initramfs for Linux, or a read-only RAM disk for bsd) for different hardware and different configurations. The initrd may also be used to mount a file system blob 232 presented as an overlay and mount a root file system on remote storage by booting into an initramfs 236, which may optionally be connected to storage resources through a SAN connection 280, as discussed below.

[0085] Filesystem blob 232 is a filesystem image that can be split into separate blobs. Blobs may be interchangeable based on configuration options, hardware type, and other differences in setup. A host booted from template 230 may boot from a union filesystem (such as overlayfs) that contains multiple blobs or images created from one or more filesystem blobs.

[0086] Template 230 may also include or be linked to additional information 237, such as volatile data 238 and / or configuration parameters 239. For example, volatile data 238 may be included in template 230, or it may be included externally, in the form of a file system blob 232, or other data store, including, but not limited to, a database, a flat file, files stored in a directory, a tarball of files, a git, or other version control repository. Additionally, configuration parameters 239 may be included externally or internally to template 230, and optionally be included in system rules and applied to template 230.

[0087] System 100 further includes IT system state 220, which tracks, maintains, changes, and updates the state of system 100, including, but not limited to, resources. System state 220 may track available resources, which informs controller logic whether and which resources are available for rule implementation and template use. System state may track used resources, which allows controller logic 205 to check and take advantage of efficiencies, whether there is a need to upgrade or switch for other reasons, such as to improve efficiency or for prioritization. System state may track which applications are running. Controller logic 205 may compare expected applications running to actual applications running according to the system state and whether they need to be revised. System state 220 may also track where applications are running. Controller logic 205 may use this information for purposes of evaluating efficiency, change management, updates, troubleshooting, or audit trails. System state may track networking information, such as which networks are on or currently running, or configuration values ​​and history. System state 220 may track the history of changes. System state 220 may also track which templates are used in which deployments based on global system rules that dictate which templates are used. The history may be used for auditing, alerts, change management, build reports, tracking versions correlated with hardware and applications and configurations, or configuration variables. System state 220 may maintain a history of configurations for auditing, compliance testing, or troubleshooting purposes.

[0088] The controller has logic 205 for managing all information contained in the system state, templates, and global system rules. The controller logic 205, global system rule database 210, IT system state 220, and templates 230 are managed by the controller 200 and may or may not reside on the controller 200. The controller logic or applications 205, global system rule database 210, IT system state 220, and templates 230 may be physical or virtual, and may or may not be distributed services, distributed databases, and / or files. The API applications 120 may be included along with the controller logic / controller applications 205.

[0089] The controller 200 may run on a standalone machine and / or may include one or more controllers. The controller 200 may include a controller service or application and may run inside another machine. The controller machine may launch the controller service first to ensure ordered and / or consistent booting of the entire stack or group of stacks.

[0090] The controller 200 may control one or more stacks with computing, storage, and networking resources, each of which may or may not be controlled by a different subset of rules in the global system rules 210. For example, there may be pre-built, build, deploy, test stacks, parallel, backup, and / or other stacks with different functions within the system.

[0091] Controller logic 205 may be configured to read and interpret global system rules to achieve a desired IT system state. Controller logic 205 may be configured to use templates according to the global rules to build system components, such as applications or services, and allocate, add, or remove resources to achieve a desired IT system state. Controller logic 205 may read global system rules, correct conditions, and develop a list of tasks to undertake to satisfy the rules based on available operations and issue instructions. Controller logic 205 may contain logic to perform operations, such as booting the system, adding, removing, and reconfiguring resources, and identifying what is available to do so. Controller logic may check the system state at startup time and at periodic intervals to see if hardware is available and, if so, execute the task. If the required hardware is not available, controller logic 205 uses available hardware from global system rules 210, templates 220, and system state 230 to present alternative options and modify the global rules and / or system state 220 accordingly.

[0092] The controller logic 205 can figure out what variables are needed, what the user needs to input to continue, or what the user needs in the system to function. The controller logic may use a list of templates from the global system rules and compare them to the templates required in the system state to ensure that the required templates are available. The controller logic 205 may identify from the system state database whether resources on a template-specific list of supported resources are available. The controller logic may allocate resources, update the state, and undertake the next set of tasks to implement the global rules. The controller logic 205 may start / run the application on the allocated resources as specified in the global rules. The rules can specify how to build the application from the template. The controller logic 205 may obtain the template(s) and configure the application from the variables. The templates may inform the controller logic 205 what kernel, boot files, file system, and supported hardware resources are required. The controller logic 205 may then add information about the application deployment to the system state database. After each command, the controller logic 205 may check the system state database against the expected state of the global rules to verify that the expected action was completed correctly.

[0093] The controller logic 205 may use versions according to the version rules. The system state 220 may have a database that correlates which rule versions are used in different deployments.

[0094] The controller logic 205 may include efficiency logic and an efficient order for rule optimization. The controller logic 205 may be configured to optimize resources. Information in the system state, rules, and templates associated with running or predicted running applications may be used by the controller logic to implement efficiency or priorities for resources. The controller logic 205 may use information in the "used resources" in the system state 220 to determine efficiency or the need to upgrade, repurpose, or switch resources for other reasons.

[0095] The controller may check the running application according to the system state 220 and compare it with the expected running application of the global rules. If the application is not running, it may start it.

[0096] If an application should not be running, it may be stopped and resources may be reallocated where appropriate. Controller logic 205 may contain a database of resource (compute, storage, networking) specifications. The controller logic may also contain logic to recognize the resource types available to the system that can be used. This may be performed using out-of-band management network 260. Controller logic 205 may also be configured to recognize new hardware using out-of-band management 260. Controller logic 205 may also retrieve information from system state 220 regarding change history, rules used, and versions for purposes of auditing, building reports, and change management.

[0097] FIG. 2F illustrates an exemplary process flow for controller logic 205 in processing template 230 and deriving an image to boot, power on, and / or enable resources, which for this exemplary purpose may be referred to as hosts. This process may also include configuring storage resources and binding storage and compute hosts and / or resources. Controller logic 205 keeps track of the hardware resources available in system 100, and system rules 210 may indicate which hardware resources may be utilized. Controller logic 205 parses template 230 in step 205.1, which may include an instruction file that can be executed to cause the controller logic to collect files external to template 230 shown in FIG. 2E. The instruction file may be in json format. In step 205.2, the controller logic collects a list of required file buckets. Also, in step 205.3, the controller logic 205 collects necessary hardware-specific files into buckets that are referenced by the hardware and, optionally, by the hypervisor (or container host system, or multi-tenancy type), which may be necessary if the hardware is to run on a virtual machine.

[0098] If hardware-specific files are present, the controller logic collects the hardware-specific files in step 205.4. In some cases, the file system image may include a kernel and initramfs along with a directory containing kernel modules (or kernel modules ultimately placed in a directory). Controller logic 205 then selects an appropriate compatible base image in step 205.5. The base image includes operating system files that may not be specific to the application or image derived from template 230. Compatible in this context means that the base image includes the files necessary to turn the template into a working application. The base image may be managed outside of the template as a space-saving mechanism (and, in many cases, the base image may be identical for several applications or services). Additionally, in step 205.6, controller logic 205 selects bucket(s) that have executable files, source code, and hardware-specific configuration files. Template 230 may reference other files, including, but not limited to, configuration files, configuration file templates (which are configuration files containing placeholders or variables that are filled with variables in system rules 210 that may be made known in template 230 so that controller 200 can turn the configuration template into a configuration file and, optionally, modify the configuration file through an API endpoint), binaries, and source code (which may be compiled when the image is booted). In step 205.7, hardware-specific instructions corresponding to the elements selected in steps 205.4, 205.5, and 205.6 may be loaded as part of the booted image. Controller logic 205 derives the image from the selected components. For example, there may be different pre-installation scripts for physical hosts versus virtual machines, or differences for PowerPC versus x86.

[0099] In step 205.8, controller logic 205 mounts the overlayfs and repackages the subject files into a single file system blob. When multiple file system blobs are used, an image may be created with the multiple blobs, unpacking the tarball, and / or fetching the git. If step 205.8 is not performed, the file system blobs may remain separate, and an image is created as a set of file system blobs and mounted with a file system that can mount multiple smaller file systems (e.g., overlayfs) together. Controller logic 205 may then identify a compatible kernel (or a kernel specified in system rules 210) in step 205.9 and an applicable initrd in step 205.10. A compatible kernel may be a kernel that satisfies the dependencies of the template and the resources used to implement the template. A compatible initrd may be an initrd that loads the template onto the desired computing resource. In many cases, an initrd may be used for physical resources so that storage resources can be mounted before a full boot (such as when the root file system is remote). The kernel and initrd may be packaged into a filesystem blob, which may be used for direct kernel boot using kexec or on the physical host to change the kernel on the live system after booting a spare operating system.

[0100] The controller then configures the storage resource(s) to enable the computing resource(s) to run the application(s) and / or image(s) using any of the techniques indicated by 205.11, 205.12, and / or 205.13. Per 205.11, an overlayfs file may be provided as a storage resource. Per 205.12, a file system is presented. For example, a storage resource may present multiple file system blobs that a computing resource can mount simultaneously using a combined file system or a file system similar to overlayfs. Per 205.13, the blobs are sent to the storage resource before presenting the file system.

[0101] Figures 2G and 2H show exemplary process flows for steps 205.11 and 205.12 of Figure 2F. Furthermore, the system may employ processes and rules for connecting computer resources to storage resources, which may be referred to as a storage connection process. Examples of such storage connection processes in addition to the process illustrated by Figures 2G and 2H are provided in Exhibit A attached hereto. Figure 2G shows an exemplary process flow for connecting storage resources. Some storage resources may be read-only, while others may be writable. Storage resources may manage their own write locks to ensure that there are no simultaneous writes that create race conditions, or system state 220 may track which connections can write to a storage resource and / or prevent multiple read-write connections to a resource (step 205.21) (see, e.g., step 205.20). The controller logic or the resource itself may query the controller's system state 220 for the location and transport type (e.g., iSCSI, Iser, NVMeoF, Fibre Channel, FCoE, NFS, NFS over RDMA, AFS, CIFs, Windows Share) of the storage resource (step 205.22). If the computing resource is virtual, the hypervisor (e.g., via a hypervisor daemon) may handle the connection to the storage resource (step 205.23). This may have desirable security benefits, as the virtual machine may not be aware of the SAN 280.

[0102] See step 205.24. The process of connecting computing resources and storage resources may be instructed in system rules 210. The controller logic then queries system state 220 to verify that the resources are available and, if necessary, writable (step 205.22). System state 220 may be queried via any number of techniques, such as an SQL query (or other type of database query), JSON parsing, etc. The query returns the necessary information for the computing resource to connect to the storage resource. The controller 200, system state 220, or system rules 210 may provide authentication credentials for the computing resource to connect to the system state (step 205.25). The computing resource then updates system state 220 (step 205.26), either directly or via the controller.

[0103] FIG. 2H illustrates an exemplary boot process for a physical, virtual, or other type of computing resource, application, service, or host to power on and connect to a storage resource. The storage resource may optionally utilize a merged file system and / or expandable volumes. In situations where a controller or other system enables a physical host, the physical host may be preloaded with an operating system to configure the system. Thus, in step 205.31, the controller may preload a boot disk with an initramfs. The controller 200 may also network boot a spare operating system (step 205.30) and then, optionally, use the out-of-band management connection 260 to preload the host with the spare operating system (step 205.31). The initramfs then loads in step 205.32, and the storage resource is connected in step 205.33 using the method shown in FIG. 2G.

[0104] If an expandable volume exists, the subvolumes or devices that are combined together are then optionally assembled into a volume group in step 205.34 if Logical Volume Management (LVM) is in use, or they may be combined using other methods of combining disks in step 205.34.

[0105] If a merged file system is in use, then in step 205.36 the files may be combined and then the boot process may continue (step 205.46). If overlayfs is in use in Linux to fix some known issues, then the following subprocesses may be performed: A / data directory may be created in each mounted file system blob, which may be volatile (step 205.37). Then, in step 205.38 a new_root directory may be created and in step 205.39 the overlayfs is mounted to the directory. Then initramfs runs exec_root on / new_root (step 205.40).

[0106] If the host is a virtual machine, additional tools such as direct kernel boot may be available. In this situation, the hypervisor may connect to storage resources before booting the VM (step 205.41), or it may do so while booting. The VM may then be direct kernel booted along with loading an initramfs (step 205.42). The initramfs then loads in step 205.43, and the hypervisor may connect to storage resources, which may be remote, at this point (step 205.44). For this to be accomplished, the hypervisor host may need to traverse an interface (e.g., if inifiniband is required to connect to an iSER target, it may traverse an SR-IOV-based virtual function using pci-passtru, or in some circumstances, may use a paravirtualized network interface). Those connections are available through the initramfs. The virtual machine may then connect to storage resources in step 205.45, if not already. It may also receive its storage resources through the hypervisor (optionally through paravirtualized storage). The process may optionally be similar for virtual machines mounting blended file systems and LVM-style disks.

[0107] FIG. 2O illustrates an exemplary process flow for configuring storage resources from file system blobs or other groups of files, as at 205.13. The blobs are collected in step 205.75, and they may be copied directly to the storage resource host in 205.73 (if the storage resource host is different from the device holding the file system blob 232). Once the storage resources are in place, the system state is then updated with the storage resource's location and available transport (e.g., iSER, nvmeof, iSCSI, FcoE, Fibre Channel, nfs, nfs over rdma) in 205.74. Some of those blobs may be read-only, in which case the system state remains the same and new computing resources or hosts may connect to that read-only storage resource (e.g., when connecting to a base image). In some cases, it may be desirable to place the files into a single file system image to avoid the overhead of any merging file systems, as shown in 205.70. This may be accomplished by mounting the blobs as a merged file system (step 205.71), then copying them to a new file system or repackaging them as a single file system (step 205.72), and then optionally copying the new file system image to an appropriate location for it to be presented as a storage resource. Some merged file systems can be merged without first mounting it in step 205.71, allowing them to be merged in a single step.

[0108] FIG. 2I illustrates another example template 230 such as that shown in FIG. 2E. In this example, a controller may be configured to use template 230 such as that shown in FIG. 2I in conjunction with an intermediate configuration tool. According to an example embodiment, the intermediate configuration tool may include a common API used to combine a new application or service with dependent applications or services. Thus, template 230 may additionally include a list 244 of dependencies that may be needed to set up the template's services. Template 230 may also include connection rules 245, which may include calls to the dependencies' common APIs. Template 230 may also include one or more common APIs 243 and a list 242 of common APIs and versions. Common API 243 may have methods, functions, scripts, or instructions, which may be callable (or not) from an application or controller, that enable the controller to configure dependent applications or services so that they can be combined into a new application built by template 230. The controller may communicate with common API 243 and / or make API calls to configure the binding of the new service or application and the dependent service or application. Alternatively, the instructions may enable the application or service to communicate with common API 243 and / or send calls to common API 243 directly on the dependent application or service. Connection rules 245 of template 230 are a set of rules and / or instructions that may encompass the API calls related to connecting the new service or application with the dependent service or application.

[0109] The system state 220 may further include a list 246 of running services, which may be queried by the controller logic 205 to satisfy dependencies 244 from the templates 230. The controller may also include a list 247 of different common APIs available for a particular service / application or type of service / application, and may also include templates that encompass the common APIs. The list may reside in the controller logic 205, the system rules 210, the system state 220, or template storage accessible to the controller. The controller also maintains an index 248 of common APIs compiled from all existing or loaded templates.

[0110] FIG. 2J illustrates an example process flow for controller logic 205 processing template 230, as shown by FIG. 2F, where the controller manages service dependencies in step 255. FIG. 2K shows an example process flow for step 255 of FIG. 2J. In step 255.1, the controller collects a list of dependencies 244 from the template. The controller also collects a list of common APIs 243 from the template. (A) In step 255.2, the controller reduces the list of possible dependent applications or services by comparing the list of common APIs 243 from the template with the index of common APIs 248 and based on the type of application or service that is considered to satisfy the dependency. In step 255.3, the controller determines whether system rules 210 specify a manner in which the dependency is satisfied.

[0111] If step 255.3 is yes, the controller determines whether the dependent service or application is running by querying a list of running templates (step 255.4); if step 255.4 is no, a service application is executed (and / or configured and then executed) (step 255.5), which may include controller logic for processing the dependent service / application's template. If the dependent service or application is found to be running in step 255.4, then process flow proceeds to step 255.6. In step 255.6, the controller uses the template to bind the new service or application being constructed to the dependent service or application. In binding the new service or application and the dependent application / service, the controller considers the template it is processing and executes connection rules 245. Based on the connection rules 245, the controller sends commands to the common API 243 on how to satisfy the dependencies 244 and / or how to bind the applications / services. The common API 243 translates instructions from the controller, which may include, but are not limited to, calling API functions of services, changing configurations, executing scripts, calling other programs, connecting new services or applications and dependent applications or services. Following step 255.6, process flow proceeds to step 205.2 of Figure 2J.

[0112] If step 255.3 determines that the system rules 210 do not specify a manner in which the dependency is satisfied, the controller then queries the system state 220 to see if the appropriate dependent application or service is running in step 255.7. In step 255.8, the controller makes that determination based on the query as to whether the appropriate dependent application or service is running. If no in step 255.8, the controller may then notify an administrator or user of the action (step 255.9). If yes in step 255.8, process flow then proceeds to step 255.6, which may operate as discussed above. The user may optionally be queried as to whether the new application should connect to the running dependent application, in which case the controller may bind the new application or service to the dependent application or service in step 255.6 as follows: the controller considers the template 230 it is processing and executes the connection rules 245. The controller then sends commands to the common API 243 based on connection rules 245 regarding how to satisfy the dependencies 244. The common API 243 translates the instructions from the controller to connect the new service or application and the dependent applications or services.

[0113] A user communicates with the controller 200 through an external user interface or web UI, or an application, through an API application 120 which may also be incorporated into the controller application or logic 205 .

[0114] The controller 200 communicates with the stack or resources by one or more of a number of networks, interconnects, or other connections through which the controller can direct the computing, storage, and networking resources to operate. Such connections may include an out-of-band management connection 260, an in-band management connection 270, a SAN connection 280, and an optional on-network in-band management connection 290.

[0115] Out-of-band management may be used by the controller 200 to discover, configure, and manage components of the system 100 through the controller 200. The out-of-band management connection 260 may enable the controller 200 to discover resources that are plugged in and available but not turned on. When plugged in, the resources may be added to the IT system state 220. The out-of-band management may be configured to load boot images and configure and monitor resources belonging to the system 100. The out-of-band management may also boot a temporary image for operating system diagnostics. The out-of-band management may be used to change BIOS settings and may use console tools to execute commands on a running operating system. Settings may also be changed by the controller using visual recognition of video signals from the console, keyboard, and physical or virtual monitor ports on hardware resources such as VGA, DVI, or HDMI ports, and / or APIs provided by the out-of-band management, e.g., Redfish.

[0116] As used herein, out-of-band management may include, but is not limited to, a management system capable of connecting to a resource or node independent of the operating system and main motherboard. Out-of-band management connection 260 may include a network or multiple types of direct or indirect connections or interconnections. Examples of types of out-of-band management connections include, but are not limited to, IPMI, Redfish, SSH, telnet, other management tools, keyboard, video, and mouse (KVM) or KVM over IP, serial console, or USB. Out-of-band management is a tool that can be used over a network to power a node or resource on and off, monitor temperatures and other system data, make BIOS and other low-level changes that may be outside the control of the operating system, connect to a console, send commands, and control inputs including, but not limited to, a keyboard, mouse, and monitor. Out-of-band management may also be coupled to out-of-band management circuitry within a physical resource. Out-of-band management may also connect a disk image as a disk that can be used to boot installation media.

[0117] A management network or in-band management connection 270 can enable a controller to gather information about computing, storage, networking, or other resources and communicate directly with the operating systems on which the resources are running. Storage, computing, or networking resources may include a management interface that interfaces with connections 260 and / or 270, allowing them to communicate with controller 200, inform the controller of what is running and available to the resources, and receive commands from the controller. As used herein, an in-band management network includes a management network that can communicate directly with resources and their operating systems. Examples of in-band management connections may include, but are not limited to, SSH, telnet, other management tools, a serial console, or USB.

[0118] Although out-of-band management is described herein as a network that is physically or virtually separate from the in-band management network, they may be combined with or operate in conjunction with each other for efficiency, as described in more detail herein. Accordingly, out-of-band and in-band management, or aspects thereof, may communicate through the same port of a controller or may be combined with a combined interconnection. Optionally, one or more of connections 260, 270, 280, 290 may be separate or combined from others of such networks, and may or may not include the same fabric.

[0119] Additionally, the computing resources, storage resources, and controller may or may not be coupled to a storage network (SAN) 280 in a manner that allows the controller 200 to use the storage network to boot each resource. The controller 200 may send a boot image or other template to a separate storage or other resource so that the other resource can boot off the storage or other resource. The controller may instruct the resource where to boot in such a situation. The controller may power on the resource and instruct the resource where to boot from and how to configure itself. The controller 200 may instruct the resource how to boot, which image to use, and where the image is located if it is on another resource. BIOS resources may be pre-configured. The controller may additionally or alternatively configure the BIOS through out-of-band management so that they boot off a storage area network. The controller 200 may also be configured to boot an operating system from an ISO and allow the resource to copy the data to a local disk. The local disk may then be used to subsequently boot. A controller may configure other resources, including other controllers, in such a manner that the resources can boot. Some resources may include applications that provide computing, storage, or networking functionality. In addition, a controller may boot up a storage resource and then have the storage resource serve to provide boot images for subsequent resources or services. Storage may also be managed through a different network used for another purpose.

[0120] Optionally, one or more of the resources may be coupled to an on-network in-band management connection 290. The connection 290 may include one or more types of in-band management such as those described with respect to the in-band management connection 270. The connection 290 may connect the controller to an application network to utilize or manage the network through an in-band management network.

[0121] FIG. 2L illustrates an image 250 that can be loaded directly or indirectly (through another resource or database) from template 230 onto a resource to boot the resource or an application or service loaded on the resource. Image 250 may include boot files 240 for the resource type and hardware. Boot files 240 may include kernels 241 corresponding to the resource, application, or service to be deployed. Boot files 240 may also include an initrd or similar file system used to assist the boot process. Boot system 240 may include multiple kernels or initrds configured for different hardware and resource types. In addition, image 250 may include file systems 251. File systems 251 may include base images 252 and corresponding file systems, as well as service images 253 and corresponding file systems and volatile images 254 and corresponding file systems. The file systems and loaded data may vary depending on the resource type and the application or service to be executed. Base image 252 may include a base operating system file system. The base operating system may be read-only. The base image 252 may also contain the basic tools of an operating system independent of what is being run. The base image 252 may include base directories and operating system tools. The service file system 253 may contain configuration files and specifications for resources, applications, or services. The volatile file system 254 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including, but not limited to, passwords, session keys, and private keys.File systems may be mounted as one single file system using techniques such as overlayFS, allowing some read-only and some read-write file systems to reduce the amount of duplicated data used for applications.

[0122] As described above, the controller 200 may be used to add resources, such as computing, storage, and / or networking resources, to the system. FIG. 11A illustrates an exemplary method for adding a physical resource, such as a bare metal node, to the system 100. A resource, i.e., a computing, storage, or networking resource, is plugged into the controller via a network connection 1110. The network connection may include an out-of-band management connection. The controller recognizes that the resource is plugged in through the out-of-band management connection 1111. The controller recognizes information associated with the resource, which may include, but is not limited to, the resource's type, capabilities, and / or attributes 1112. The controller adds the resource and / or information associated with the resource to its system state 1113. An image derived from a template is loaded onto a physical component of the system, which may include, but is not limited to, the resource, another resource, such as a storage resource, or a controller 1114. The image includes one or more file systems, which may include configuration files. Such configuration may include BIOS and boot parameters. The controller instructs the physical resource to boot using the file system of the image 1115. Additional resources, or multiple bare metal or physical resources of different types, may be added in this manner using the template image, or at least a portion thereof.

[0123] 11B illustrates an exemplary method for automatically allocating resources using global system rules and templates in an exemplary embodiment. A request is made to the system (1120) that requires resource allocation to fulfill the request. The controller discovers its resource pool based on its system state database (1121). The controller uses the template to determine the required resources (1122). The controller allocates the resources and stores the information in the system state (1123). The controller deploys the resources using the template (1124).

[0124] Referring to Figure 12, an exemplary method for automatically deploying an application or service using the system 100 described herein is illustrated. A user or application makes a request for a service (1210). The request is translated to an API application (1220). The API application routes the request to a controller (1230). The controller interprets the request (1240). The controller considers the state of the system and its resources (1250). The controller uses its rules and templates for service deployment (1260). The controller sends the request to the resources (1270), deploys an image derived from the template (1280), and updates the IT system state.

[0125] Additional and more detailed examples of operations such as adding resources, allocating resources, and deploying applications or services are discussed in more detail below.

[0126] Adding Computing Resources to the System 3A, the addition of a computing resource 310 to the system 100 is illustrated. When the computing resource 310 is added, it may be coupled to the controller 200 and powered off. If the computing resource 310 is preloaded with an image, alternative steps may follow, in which any of the network connections may be used to communicate with the resource, boot the resource, and add information to the system state. If the computing resource and the controller are on the same node, the service running the computing resource is off.

[0127] As shown in FIG. 3A , computing resource 310 is coupled to the controller 200 by networks: out-of-band management connection 260, in-band management connection 270, and optionally, SAN 280. Computing resource 310 is also coupled to one or more application networks 390, through which services, application users, and / or clients can communicate with each other. Out-of-band management connection 260 may be coupled to a separate out-of-band management device 315 or circuitry within computing resource 310 that is turned on when computing resource 310 is plugged in. Device 315 may enable functions including, but not limited to, powering the device on / off, attaching to a console, typing commands, monitoring temperature and other computer health-related factors, setting BIOS settings, and other functions outside of the operating system. Controller 200 can see computing resource 310 through out-of-band management network 260. It can also use in-band or out-of-band management to identify the type of computing resource and its configuration. Controller logic 205 is configured to consult out-of-band management 260 or in-band management 270 to add hardware. If computing resource 310 is detected, controller logic 205 may then use global system rules 220 to determine whether the resource is configured automatically or by interacting with a user. If it is added automatically, the setup follows global system rules 210 in controller 200. If it is added by a user, global system rules 210 in controller 200 may query the user to confirm the addition of the resource and what the user wants to do with the computing resource.The controller 200 may query an API application or otherwise request a user or any program controlling the stack to confirm that the new resource is authorized. The authorization process may be completed automatically and securely using cryptography to verify the legitimacy of the new resource. The controller logic 205 adds the computing resource 310 to the IT system state 220, including the switch or network to which the computing resource 310 is plugged.

[0128] If the computing resource is physical, the controller 200 may power on the computing resource through the out-of-band management network 260, and the computing resource 310 may boot off an image 350 loaded from the template 230 using the global system rules 210 and the controller logic 205, for example, by the SAN 280. The image may also be loaded indirectly through other network connections or by another resource. Once booted, information related to the computing resource 310 received through the in-band management connection 270 may also be collected and added to the IT system state 220. The computing resource 310 may then be added to a storage resource pool, and it becomes a resource managed by the controller 200 and tracked in the IT system state 220.

[0129] If the computing resource is virtual, the controller 200 may power on the computing resource either through the in-band management network 270 or the out-of-band management 260. The computing resource 310 may boot off an image 350 loaded from a template 230 using global system rules 210 and controller logic 205, for example, by the SAN 280. The image may also be loaded indirectly through other network connections or by another resource. Once booted, information related to the computing resource 310 received through the in-band management connection 270 may also be collected and added to the IT system state 220. The computing resource 310 may then be added to a storage resource pool, and it becomes a resource managed by the controller 200 and tracked in the IT system state 220.

[0130] The controller 200 may be capable of automatically turning resources on and off and updating the IT system state according to global system rules for reasons determined by the IT system user, such as turning resources off to save power, or turning resources on to improve application performance, or for other reasons the IT system user may have.

[0131] FIG. 3B illustrates an image 350 that is loaded directly or indirectly (through another resource or database) from template 230 onto computing resource 310 to boot the computing resource and / or load an application. Image 350 may include boot files 340 for the resource type and hardware. Boot files 340 may include kernels 341 corresponding to the resources, applications, or services to be deployed. Boot files 340 may also include an initrd or similar file system used to support the boot process. Boot system 340 may include multiple kernels or initrds configured for different hardware and resource types. In addition, image 350 may include file systems 351. File systems 351 may include base images 352 and corresponding file systems, as well as service images 353 and corresponding file systems and volatile images 354 and corresponding file systems. The file systems and loaded data may vary depending on the resource type and the running applications or services. Base image 352 may include the file system of a base operating system. The base operating system may be read-only. The base image 352 may also contain the basic tools of an operating system independent of what is being run. The base image 352 may include base directories and operating system tools. The service file system 353 may contain configuration files and specifications for resources, applications, or services. The volatile file system 354 may contain information or data specific to that deployment, such as binary applications, specific addresses, and other information that may or may not be configured as variables, including, but not limited to, passwords, session keys, and private keys.File systems may be mounted as one single file system using techniques such as overlayFS, allowing some read-only and some read-write file systems to reduce the amount of duplicated data used for applications.

[0132] FIG. 3C illustrates an exemplary process flow for adding a resource, such as computing resource 310, to system 100. In this example, the target resource is described as computing resource 310; however, it should be understood that the target resource for the process flow of FIG. 3C may be storage resource 410 and / or networking resource 510. In the example of FIG. 3C, resource 310 to be added is not on the same node as controller 200. In step 300.1, resource 310 is coupled to controller 200 in a powered-off state. In the example of FIG. 3C, out-of-band management connection 260 is used to connect resource 310. However, it should be understood that other network connections may be used if desired by the implementer. In steps 300.2 and 300.3, controller logic 205 examines the system's out-of-band management connection and uses out-of-band management connection 260 to recognize and identify the type and configuration of resource 310 to be added. For example, the controller logic may check the BIOS or other information about the resource (such as serial number information) as a reference to obtain type and configuration information.

[0133] In step 300.4, the controller uses global system rules to determine whether a particular resource 310 should be automatically added. If not, the controller waits until its use is approved (step 300.5). For example, a user may respond to a query that they do not want to use a particular resource 310, or that it may be automatically withheld until it is to be used in step 300.4. If step 300.4 determines that the resource 310 should be automatically added, the controller then uses the rules for automatic setup (step 300.6) and proceeds to step 300.7.

[0134] In step 300.7, the controller selects and uses a template 230 associated with the resource to add the resource to the system state 220. In some cases, the template 230 may be specific to a particular resource. However, some templates 230 may encompass multiple resource types. For example, some templates 230 may be hardware-independent. In step 300.8, the controller powers on the resource 310 through its out-of-band management connection 260 according to the global system rules 210. In step 300.9, using the global system rules 210, the controller finds and loads a boot image for the resource from the selected template(s). The resource 310 is then booted from the image derived from the target template 230 (step 300.10). Additional information about the resource 310 may then be received from the resource 310 through the in-band management connection 270 after the resource 310 has booted (step 300.11). Such information may include, for example, firmware version, network card, or any other devices to which the resource may be connected. In step 300.12, the new information may be added to the system state 220. The resource 310 may then be considered to be added to the resource pool and is prepared for allocation (step 300.13).

[0135] With respect to FIG. 3C , it should be understood that if the resource and controller are on the same node, the service running the resource may be outside of that node. In such a case, the controller may use an inter-process communication technique with the resource, such as, for example, a unix socket, loopback adapter, or other inter-process communication technique to communicate with the resource. From the system rules, the controller may install a virtual host, or a hypervisor or container host, to run the application using a known template from the controller. The resource application information can then be added to the system state 220, and the resource will be ready for allocation.

[0136] Adding storage resources to the system 4A illustrates adding a storage resource 410 to system 100. In an exemplary embodiment, the exemplary process flow of FIG. 3C may be followed to add a storage resource 410 to system 100, where the added storage resource 410 is not on the same node as controller 200. It should also be noted that if the storage resource 410 is preloaded with an image, alternative steps may be followed in which any network connection may be used to communicate with storage resource 410, boot storage resource 410, and add information to system state 220.

[0137] When a storage resource 410 is added, it may be coupled to the controller 200 and powered off. The storage resource 410 is coupled to the controller via networks: out-of-band management network 260, in-band management connection 270, SAN 280, and, optionally, connection 290. The storage resource 410 may also be coupled to one or more application networks 390 through which services, application users, and / or clients can communicate with each other. An application or client may directly or indirectly access the resource's storage through the application, whereby the application or client is not accessed through a SAN. The application network may have storage built into it, or it may be accessed and identified as a storage resource in the IT system state. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 415 or circuitry for the storage resource 410 that turns on when the storage resource 410 is connected. Device 415 may enable features including, but not limited to, powering the device on / off, attaching to a console and typing commands, monitoring temperature and other computer health-related factors, and configuring BIOS settings and other features outside the scope of the operating system. Controller 200 may reference storage resources 410 through out-of-band management network 260. It may identify the type of storage resource and its configuration using in-band or out-of-band management. Controller logic 205 is configured to consult out-of-band management 260 or in-band management 270 for added hardware. If storage resource 410 is detected, controller logic 205 may use global system rules 220 to determine whether resource 410 is configured automatically or through user interaction.If the resource 410 is added automatically, the setup follows global system rules 210 in the controller 200. If the resource 410 is added by a user, the global system rules 210 in the controller 200 may ask the user to confirm the addition of the resource and what the user wants to do with the storage resource. The controller 200 may query the API application(s) or otherwise request the user or any program controlling the stack to confirm that the new resource is authorized. The authorization process may also be completed automatically and securely using cryptography to verify the validity of the new resource. The controller logic 205 adds the storage resource 410 to the IT system state 220, including the switch or network to which the storage resource 410 is connected.

[0138] Controller 200 may power on storage resource 410 through out-of-band management network 260, and storage resource 410 boots from image 450 loaded from template 230, for example, via SAN 280, using global system rules 210 and controller logic 205. The image may also be loaded through other network connections or indirectly through another resource. Once booted, information received through in-band management connection 270 about storage resource 410 may be collected and added to IT system state 220. Here, storage resource 410 is added to a storage resource pool, and it becomes a resource managed by controller 200 and tracked in IT system state 220.

[0139] A storage resource may comprise a storage resource pool or multiple storage resource pools that an IT system may use or access, singly or simultaneously. When a storage resource is added, the storage resource may provide a storage pool, multiple storage pools, a portion of a storage pool, and / or multiple portions of a storage pool to the IT system state. A controller and / or storage resource may manage various storage resources in a pool, or a group of such resources within a pool. A storage pool may include multiple storage pools running on multiple storage resources. For example, flash storage disks or arrays caching platter disks or arrays, or a storage pool on dedicated compute nodes coupled to a pool on dedicated storage nodes, simultaneously optimizing bandwidth and latency.

[0140] FIG. 4B illustrates an image 450 that is loaded directly or indirectly from template 230 (from another resource or database) onto storage resource 410 to boot the storage resource and / or load an application. Image 450 may include boot files 440 for the resource type and hardware. Boot files 440 may include a kernel 441 corresponding to the resource, application, or service being deployed. Boot files 440 may include an initrd or similar file system used to assist the boot process. Boot system 440 may include multiple kernels or initrds configured for different hardware and resource types. Additionally, image 450 may include file systems 451. File systems 451 may include a base image 452 and corresponding file systems, a service image 453 and corresponding file systems, and a volatile image 454 and corresponding file systems. The file systems and data loaded may vary depending on the resource type and application or service being run. Base image 452 may include a base operating system file system. The base operating system may be read-only. Base image 452 may also include the basic tools of an operating system that are independent of what is being run. Base image 452 may include base directories and operating system tools. Service file system 453 may include configuration files and specifications for resources, applications, or services. Volatile file system 454 may include information or data specific to that deployment, such as binary applications, unique addresses, and other information, which may or may not be configured as variables, including, but not limited to, passwords, session keys, and private keys.File systems may be mounted as one single file system using techniques such as overlayFS, allowing some file systems to be read-only and some to be read-write, reducing the amount of duplicate data used for applications.

[0141] 5A illustrates an example in which another storage resource, direct-attached storage 510, which may take the form of a JBOD or other type of direct-attached storage node, is coupled to storage resource 410 as an additional storage resource for the system. A JBOD is typically an external disk array connected to a node that provides the storage resource, and while FIG. 5A uses a JBOD as an exemplary form of direct-attached storage 510, it should be understood that other types of direct-attached storage may be employed as 510.

[0142] Controller 200 may add storage resources 410 and JBODs 510 to its system, for example, as described with respect to FIG. 5A . JBODs 510 are coupled to controller 200 via out-of-band management connections 260. Storage resources 410 are coupled to a network: out-of-band management connections 260, in-band management connections 270, SAN 280, and, optionally, connections 290. Storage nodes 410 communicate with the JBODs 510's storage through a SAS or other disk drive fabric 520. JBODs 510 may also include an out-of-band management device 515 that communicates with the controller through out-of-band management connections 260. Through out-of-band management 260, controller 200 may discover JBODs 510 and storage resources 410. Controller 200 may also discover other parameters not controlled by the operating system, for example, as described herein with respect to various out-of-band management circuits. Controller 200 global system rules 210 provide configuration startup rules for booting or starting up JBODs and storage nodes that have not yet been added. The order in which storage resources are turned on may be controlled by controller logic 205 using global rules 220. According to one set of global system rules 220, the controller may first power on JBOD 510, and then controller 200 may power on storage resources using the loaded image 450 in a manner similar to that described with respect to FIG. 4. In another set of global system rules, controller 200 may first power on storage resources 410 and then JBOD 510. Other global system rules may specify power-on timing or delays between various devices. Detecting the readiness or operational state of various resources may be determined by controller logic 205, global system rules 210, and / or templates 230 and / or may be used in device allocation management by controller 200.The IT system state 220 may be updated through communication with the storage resources 410. The storage nodes 410 are aware of the storage parameters and configuration of the JBODs 510 by accessing the JBODs through the disk fabric 520. The storage resources 410 then provide information to the controller 200, which updates the IT system state 220 with information regarding the amount of available storage and other attributes. The controller updates the IT system state 220 when the storage resources 410 are booted and recognized as part of the pool of storage resources 400 in the system 100. The storage nodes handle the logic for controlling the JBOD storage resources using the configuration set by the controller 200. For example, the controller may instruct the storage nodes to configure the JBODs to create a pool from a RAID-10 or other configuration.

[0143] 5B shows an exemplary process flow for adding storage resource 410 and direct-attached storage 510 for storage resource 410 to system 100. In step 500.1, direct-attached storage 510 is coupled to powered-off controller 200 via out-of-band management connection 260. In step 500.2, storage resource 410 is coupled to powered-off controller 200 via out-of-band management connection 260 and in-band management connection 270, while storage resource 410 is simultaneously coupled to direct-attached storage 510 via SAS 520, e.g., a disk drive fabric.

[0144] Controller logic 205 may then examine out-of-band management connection 260 (step 500.3) to discover storage resources 410 and direct-attached storage 510. While any network connection may be used, in this example out-of-band management may be used for the controller logic to recognize and identify the types of resources being added (in this case storage resources 410 and direct-attached storage 510) and their configuration (step 500.4).

[0145] In step 500.5, the controller 200 selects and uses a template 230 for a particular type of storage for each type of storage device to add the resources 410 and 510 to the system state 220. In step 500.6, the controller follows global system rules 210, which can specify a boot order for powering on the direct storage and storage nodes in that order through the out-of-band management connection 260 (500.6). Using the global system rules 210, the controller finds and loads a boot image for the storage resource 410 from the template 230 selected for that storage resource 410, and the storage resource is then booted from the image (step 500.7). The storage resource 410 knows the storage parameters and configuration of the direct-attached storage 510 by accessing the direct-attached storage 510 through the disk fabric 520. Additional information about the storage resource 410 and / or the direct-attached storage 510 may then be provided to the controller through the in-band management connection 270 to the storage resource (step 500.8). In step 500.9, the controller updates the system state 220 with the information obtained in step 500.8. In step 500.10, the controller sets the configuration for the storage resources 410 that handle the direct-attached storage 510 and how to configure the direct-attached storage. Then, in step 500.11, the new resource comprising the storage resources 410 together with the direct-attached storage 510 may be added to the resource pool and is ready to be allocated within the system.

[0146] According to another aspect of the exemplary embodiment, the controller may use out-of-band management to be aware of other devices in the stack that may not be involved in computing or services. For example, such devices may include, but are not limited to, cooling towers / air conditioners, lighting, temperature, sound, alarms, power systems, or any other devices associated with the system.

[0147] Adding networking resources to a system 6A illustrates the addition of a networking resource 610 to system 100. In an exemplary embodiment, the exemplary process flow of FIG. 3C may be followed to add a networking resource 610 to system 100, where the added networking resource 610 is not on the same node as controller 200. It should also be noted that if the networking resource 610 has a preloaded image, alternative steps may be followed in which any network connection may be used to communicate with network resource 610, boot network resource 610, and add information to system state 220.

[0148] When a networking resource 610 is added, it may be coupled to the controller 200 and powered off. The networking resource 610 may be coupled to the controller 200 via connections, i.e., out-of-band management connection 260 and / or in-band management connection 270. The networking resource 610 is optionally connected to the SAN 280 and / or connection 290. The networking resource 610 may also be coupled to one or more application networks 390 through which services, application users, and / or clients can communicate with each other, or may not be coupled. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 615 or circuitry for the networking resource 610 that turns on when the networking resource 610 is connected. The device 615 may enable features including, but not limited to, powering the device on / off, attaching to a console and typing commands, monitoring temperature and other computer health-related factors, and setting BIOS settings and other features outside the scope of the operating system. The controller 200 may reference the networking resource 610 through the out-of-band management connection 260. It may identify the type of networking resource and / or network fabric and may identify the configuration using in-band or out-of-band management. The controller logic 205 is configured to consult the out-of-band management 260 or in-band management 270 for the hardware being added. If the networking resource 610 is detected, the controller logic 205 may use the global system rules 220 to determine whether the networking resource 610 is configured automatically or through user interaction. If the resource 610 is added automatically, the setup follows the global system rules 210 in the controller 200. If added by a user, the global system rules 210 in the controller 200 may ask the user to confirm the addition of the resource and what the user wants to do with the resource.The controller 200 may query the API application(s) or otherwise request a user or any program controlling the stack to confirm that the new resources have been authorized. The authorization process may also be completed automatically and securely using cryptography to verify the validity of the new resources. The controller logic 205 may then add the networking resources 610 to the IT system state 220. For switches that cannot identify themselves to the controller, the user may manually add them to the system state.

[0149] If the networking resource is physical, the controller 200 may power on the networking resource 610 through the out-of-band management connection 260, and the networking resource 610 may boot from an image 605 loaded from the template 230, for example, via the SAN 280, using the global system rules 210 and the controller logic 205. The image may also be loaded through other network connections or indirectly via other resources. Once booted, information received through the in-band management connection 270 about the networking resource 610 may be collected and added to the IT system state 220. The networking resource 610 may then be added to a storage resource pool, and it becomes a resource managed by the controller 200 and tracked in the IT system state 220. Optionally, some networking resource switches may be controlled through a console port connected to the out-of-band management 260, configured at power-on, or have a switch operating system installed through a boot loader, for example, through ONIE.

[0150] If the networking resource is virtual, the controller 200 may power on the networking resource either through the in-band management network 270 or the out-of-band management 260. The networking resource 610 may boot from an image 650 loaded from a template 230 via the SAN 280 using the global system rules 210 and the controller logic 205. Once booted, information received through the in-band management connection 270 about the networking resource 610 may be collected and added to the IT system state 220. The networking resource 610 may then be added to the storage resource pool, where it becomes a resource managed by the controller 200 and tracked in the IT system state 220.

[0151] The controller 200 may instruct networking resources, whether physical or virtual, to assign, reassign, or move ports to connect to different physical or virtual resources, i.e., connectivity, storage, or compute as defined herein. This may be done using techniques including, but not limited to, SDN, InfiniBand partitioning, VLAN, and vXLAN. The controller 200 may instruct virtual switches to move or assign virtual interfaces for network or interconnect communication with the virtual switch or the resource hosting the virtual switch. Some physical or virtual switches may be controlled by an API coupled to the controller.

[0152] When such a change is possible, the controller 200 may also instruct the compute, storage, or networking resource to change fabric type. A port may be configured to switch to a different fabric, for example, to switch between fabrics of a hybrid InfiniBand / Ethernet interface.

[0153] The controller 200 may provide instructions to networking resources, which may include switches or other networking resources that switch multiple application networks. The switches or network devices may include different fabrics, or, for example, they may be connected to InfiniBand switches, ROCE switches, and / or other switches, preferably with SDN functionality and multiple fabrics.

[0154] FIG. 6B illustrates an image 650 that is loaded directly or indirectly from template 230 (e.g., via another resource or database) onto networking resource 610 to boot the networking resource and / or load an application. Image 650 may include boot files 640 for the resource type and hardware. Boot files 640 may include a kernel 641 corresponding to the resource, application, or service being deployed. Boot files 640 may include an initrd or similar file system used to assist the boot process. Boot system 640 may include multiple kernels or initrds configured for different hardware and resource types. Additionally, image 650 may include file systems 651. File systems 651 may include a base image 652 and corresponding file systems, a service image 653 and corresponding file systems, and a volatile image 654 and corresponding file systems. The file systems and data loaded may vary depending on the resource type and application or service being run. Base image 652 may include a base operating system file system. The base operating system may be read-only. The base image 652 may also include the basic tools of an operating system that are independent of what is being run. The base image 652 may include base directories and operating system tools. The service file system 653 may include configuration files and specifications for resources, applications, or services. The volatile file system 654 may include information or data specific to that deployment, such as binary applications, unique addresses, and other information, which may or may not be configured as variables, including, but not limited to, passwords, session keys, and private keys.File systems may be mounted as one single file system using techniques such as overlayFS, allowing some file systems to be read-only and some to be read-write, reducing the amount of duplicate data used for applications.

[0155] Deploying an application or service on a resource 7A illustrates system 100 comprising controller 200, physical and virtual computing resources comprising first computing node 311, second computing node 312, and third computing node 313, storage resources 410, and network resources 610. The resources are shown set up and added to IT system state 220 in the manner described herein with respect to FIGS.

[0156] Although multiple compute nodes are shown in this figure, a single compute node may be used according to an example embodiment. The compute node may host physical or virtual computing resources and may run applications on a physical or virtual compute node. Similarly, while a single network provider node and storage node are shown, it is contemplated that multiple resource nodes of these types may or may not be used in the system of the example embodiment.

[0157] Services or applications may be deployed to any of the systems according to an example embodiment. An example of deploying services on compute nodes may be described with reference to FIG. 7A , but may similarly be used in different configurations of system 100. For example, controller 200 of FIG. 7A may automatically configure computing resources 310 in the form of compute nodes 311, 312, and 313 according to global system rules 210. These may then be added to IT system state 220. Thus, controller 200 may recognize computing resources 311, 312, and 313 (which may or may not be powered off) and, in some cases, any physical or virtual applications running on the computing resources or nodes. Controller 200 may automatically configure storage resource(s) 410 and networking resource(s) 610 according to global system rules 210 and templates 230 and add them to IT system state 220. Controller 200 may recognize storage resources 410 and networking resources 610, which may or may not start in a powered-off state.

[0158] FIG. 7B shows an exemplary process for adding a resource to IT system 100. In step 700.1, a new physical resource is coupled to the system. In step 700.2, the controller becomes aware of the new resource. The resource may be connected to remote storage (step 700.4). In step 700.3, the controller configures how to boot the new resource. All connections made to the resource may be recorded in system state 220 (step 700.5). FIG. 3C, discussed above, provides further details regarding an exemplary embodiment of a process flow such as that shown in FIG. 7B.

[0159] 7C and 7D show an example process flow for deployment of an application to multiple computing resources, multiple servers, multiple virtual machines, and / or multiple sites. The process for this example differs from standard template deployment in that IT system 100 requires components that link redundant and interrelated applications and / or services. Controller logic may process a meta-template at step 700.11, where the meta-template may include multiple templates 230, file system blobs 232, and other components (which may be in the form of other templates 230) needed to configure the multi-homed service.

[0160] In step 700.12, controller logic 205 checks system state 220 for available resources. However, if there are not enough resources, controller logic may reduce the number of redundant services that may be deployed (see 700.16, where the number of redundant services is identified). In step 700.13, controller logic 205 configures the networking resources and interconnections needed to connect the services. If the service or application is deployed across multiple sites, the meta-template may include (or controller logic 205 may configure) services that are optionally configured from the template to enable data synchronization and interoperability across sites (see 700.15).

[0161] In step 700.16, controller logic 205 may determine the number of redundant services (if the redundant services are on multiple hosts) from system rules, meta template data, and resource availability. In step 700.17, it concatenates with other redundant services and concatenates with the master. If there are multiple redundant hosts, controller logic 205 or logic in the template (binaries 234, daemons 232, or file system blobs that may contain configuration files directing settings in the operating system) can prevent network address and hostname conflicts. Optionally, controller logic provides a network address (see 700.18) and registers each redundant service in DNS (700.19) and system state 220 (700.18). System state 220 tracks redundant services, and if controller logic 205 notices a redundant service with conflicting parameters such as a hostname, DNS name, etc., it disallows the duplicate registration and the network address already in system state 220.

[0162] The configuration routine illustrated by FIG. 7D processes the template(s) of the meta-template. The configuration routine processes all redundant services, deploys multi-host or clustered services to multiple hosts, and deploys services that link hosts. Any process capable of deploying IT systems from system rules can execute the configuration routine. For multi-host services, an exemplary routine may process the service template as in 700.32, provision storage resources as in 700.33, power on the hosts as in 700.35, and link the hosts / computing resources with the storage resources (and register with system state 220) as in 700.36 (then repeat for the number of redundant services (700.38)). Each time, it registers with system state 220 (see 700.20) and uses controller logic to record information to track individual services and prevent conflicts (see 700.31).

[0163] Some service templates may include services and tools that allow multi-host services to be chained together. Some of these services may be treated as dependencies (700.39), and then the chaining routines of 700.40 may be used to chain the services and register the chaining with system state 220. Furthermore, one of the service templates may be a master template, in which case the dependent service templates of 700.39 are slave or secondary services, and the chaining routines of 700.40 connect them. Routines may be defined in meta templates. For example, for redundant DNS configuration, the chaining routines of 700.40 may include connecting slave DNSs to master DNSs and configuring zone transfers with DNSSEC. Some services may use physical storage (see 700.34) to improve performance, which may be loaded with a backup OS as disclosed in FIG. 5B. Tools for chaining services may be included in the template itself, and inter-service configuration may be performed by an API accessible on the controller and / or other hosts in a multi-node application / service.

[0164] The controller 200 may enable a user or a controller to determine the appropriate compute backend to use for an application. The controller 200 may enable a user or a controller to optimally place an application on an appropriate physical or virtual computing resource by determining resource usage. When hypervisors or other compute backends are deployed on compute nodes, they may report resource usage statistics back to the controller through the in-band management connection 270. When the controller decides to create an application on a virtual computing resource, either from its own logic and global system rules or from user input, it may automatically select the hypervisor on the most suitable host and power on the virtual computing resource on that host.

[0165] For example, the controller 200 deploys an application or service to one or more computing resources using the template(s) 230. Such an application or service may be, for example, a virtual machine that executes the application or service. In one example, FIG. 7A illustrates the deployment of multiple virtual machines (VMs) on multiple computing nodes, where the controller 200 can recognize that multiple computing resources 310 are in the computing resource pool in the form of computing nodes 311, 312, and 313. The computing nodes may, for example, be deployed with a hypervisor, or alternatively, may be on bare metal where the use of virtual machines may be undesirable due to speed. In this example, the computing resource 310 has VM(1) 321 and VM(2) 322 configured and deployed on the computing node 311 with a hypervisor application loaded. For example, if compute node 311 does not have the resources for an additional VM, or if other resources are preferred for a particular service, controller 200 may recognize based on stack state 220 that there are no available resources on compute node 311, or that it is preferable to provision a new VM on different resources. It may also be recognized that a hypervisor is loaded on computing resource 312 and not on resource 313, which may be, for example, a bare-metal compute node used for other purposes. Thus, according to the requirements of installed service or application templates and the status of system state 220, the controller in this example may select compute node 313 for deployment of the next required resource VM(3) 323.

[0166] The computing resources of the system may be configured to share storage on the storage resources for the storage nodes.

[0167] A user may request through the user interface 110 or an application that services be set up for the system 100. Services may include, but are not limited to, email services, web services, user management services, network providers, LDAP, Dev tools, VOIP, authentication tools, and accounting.

[0168] The API application 120 translates the user or application request and sends a message to the controller 200. The service template or image 230 of the controller 200 is used to identify which resources are needed for the service. The resources to be used are then identified based on availability according to the IT system state 220. The controller 200 makes a request to one or more of the compute nodes 311, 312, or 313 for the needed computing services, a request to the storage resources 410 for the needed storage resources, and a request to the network resources 610 for the needed networking resources. The IT system state 220 is then updated to identify the resources to be allocated. The service is then installed on the allocated resources using the global system rules 210 according to the template 230 for the service or application.

[0169] According to an exemplary embodiment, multiple computing nodes may be used, whether for the same service or different services, while, for example, storage services and / or network provider pools may be shared among computing nodes.

[0170] Referring to FIG. 8A , system 100 is shown in which controller 200, computing resources 300, storage resources 400, and networking resources 600 are on the same or shared physical hardware, such as a single node. Various features shown and described in FIGS. 1-10 may be incorporated into a single node. When a node is powered on, a controller image is loaded onto the node. Computing resources 300, storage resources 400, and networking resources 600 are configured by templates 230 using global system rules 210. Controller 200 may be configured to load compute backends 318, 319 as computing resources, which may or may not be added on the node or on different node(s). Such backends 318, 319 may include, but are not limited to, virtualization, containers, and multi-tenant processes to create virtual computing, networking, and storage resources.

[0171] Applications or services 725, for example, web, email, core network services (DHCP, DNS, etc.), collaboration tools, may be installed on virtual resources on nodes / devices shared with the controller 200. These applications or services may be moved to physical or virtual resources independent of the controller 200. Applications may run on virtual machines on a single node.

[0172] FIG. 8B illustrates an exemplary process flow for expanding a system from a single node to a multiple node system (e.g., with nodes 318 and / or 319 as shown in FIG. 8A). Now, with reference to FIGS. 8A and 8B, one can consider an IT system with controller 200 running on a single server, and it is desirable to expand the IT system to a multi-node IT system. Thus, prior to expansion, the IT system is in a single-node state. As shown in FIG. 8A, controller 200 runs on the multi-tenant single-node system to run various IT system management applications and / or resources, which may include, but are not limited to, storage resources, computing resources, a hypervisor, and / or a container host.

[0173] In step 800.2, a new physical resource is coupled to a single node system by connecting the new physical resource through out-of-band management connection 260, in-band management connection 270, SAN 280, and / or network 290. For this example, this new physical resource may also be referred to as hardware or a host. Controller 200 may discover the new resource on the management network and then query the device. Alternatively, the new device may broadcast a message announcing itself to controller 200. For example, the new device may be identified by MAC address, out-of-band management, and / or booting a standby OS and using in-band management to identify its hardware type. In either case, in step 800.3, the new device provides information about its node type and its currently available hardware and software resources to the controller. Controller 200 then recognizes the new device and its capabilities.

[0174] In step 800.4, tasks allocated to the system running controller 200 may be assigned to the new host. For example, if the host is preloaded with an operating system (such as a storage host operating system or a hypervisor), controller 200 allocates new hardware resources and / or functionality. The controller may then provide an image to provision the new hardware, or the new hardware may request an image from the controller and configure itself using the methods disclosed above and below. If the new host is capable of hosting storage resources or virtual computing resources, the new resources may be made available to controller 200. Controller 200 may then move and / or allocate existing applications to the new resources, or may use the new resources for newly created or subsequently created applications.

[0175] In step 800.5, the IT system may keep its current applications running on the controller or may migrate them to new hardware. When migrating virtual computing resources, VM migration techniques (e.g., qemu+kvm migration tools) may be used to update the system state with new system rules. Change management techniques discussed below can be used to make these changes reliably and safely. As more applications may be added to the system, the controller may use any of a variety of techniques to determine how to allocate the system's resources, including, but not limited to, round robin, weighted round robin, least utilization, weighted least utilization, utilization-based forecasting with assisted training, planning, desired capacity, and maximum size techniques.

[0176] 8C shows an exemplary process flow for migration of storage resources to new physical storage resources. The storage resources may then be mirrored, migrated, or a combination thereof (e.g., the storage may be mirrored and then the original storage resource disconnected). In step 820, the storage resource is coupled to the system either by contacting the controller or by having the controller discover the new storage resource. This can be done with out-of-band management connection 260, in-band management connection 270, SAN network 280, or in a flat network, application network, a combination thereof may be used. With in-band management, the operating system may be pre-booted and the new resource may connect to the controller.

[0177] In step 822, a new storage target is created on the new storage resource, which can be recorded in the database in step 824. In one example, the storage target may be created by copying a file. In another example, the storage target may be created by creating a block device and copying data (which may be in the form of file system blob(s)). In another example, the storage target may be created by mirroring two or more storage resources between block devices (e.g., creating a RAID), optionally connecting through remote storage transport(s), including but not limited to iSCSI, ISER, NVMeOF, NFS, NFS over RDMA, FC, FCOE, SRP, etc. The entry in the database in step 824 may include information for the computing resource (or other type of resource and / or host) to connect to the new storage resource remotely or locally if the storage resource is on the same device as other resources or hosts.

[0178] In step 826, storage resources are synchronized. For example, the storage can be mirrored. As another example, the storage can be synchronized offline. Technologies such as RAID1 (or other types of RAID, typically RAID1 or RAID0, but optionally RAID110 (mirrored RAID10)) (mdadm, zfs, btrfs, hardware raid) may be employed in step 826.

[0179] The data from the old storage resource is then optionally connected after database login in step 828 (the database may include information related to the status of copying the data if such data must be recorded when that occurs later). If the storage target is being migrated away from the previous host (e.g., moving from a single node system to a multi-node and / or distributed IT system as previously shown in FIGS. 8A and 8B), the new storage resource may be referred to as the primary storage resource in step 830 by the controller, system state, computing resource, or a combination thereof. This may be done as a step to remove the old storage resource. In some cases, the physical or virtual hosts connected to the resource may then need to be updated, and in some cases may be powered off (and then powered back on) during the migration in step 832 (the techniques disclosed herein for powering on physical or virtual hosts can be used).

[0180] FIG. 8D shows an exemplary process flow for migrating virtual machines, containers, and / or processes on a single node of a multi-tenant system to a multi-node system, which may have separate hardware for compute and storage. In step 850, controller 200 creates new storage resources, which may be on the new node (see, for example, nodes 318 and 319 in FIG. 8A ). Next, in step 852, the old application host may be powered off. Then, in step 854, data is copied or synchronized. By powering down in step 852 before copying / synchronizing in step 854, the migration may be safer if the migration involves migrating VMs from a single node. Powering down may also be beneficial for moving VMs to physical machines. Step 854 may be accomplished before powering down via data pre-synchronization step 862, which can minimize associated downtime. Furthermore, the host may not need to be powered down as in step 852; in that case, the old host remains online until the new host is ready (or new storage resources are ready). Techniques for avoiding the power off step 852 are discussed in more detail below. In step 854, data may optionally be synchronized unless the storage resources are mirrored or synchronized using hot standby.

[0181] The new storage resource may now be up and running and logged into the database in step 856 so that the controller 200 can connect the new host to the new storage resource in step 858. When multiple virtual hosts are migrated from a single node, this process may need to be repeated for multiple hosts (step 860). Boot order may be determined by the controller logic using application dependencies if they are tracked.

[0182] 8E shows another exemplary process flow for scaling a system from a single node to multiple nodes. In step 870, a new resource is coupled to a single node system. The controller may have a set of system rules and / or scaling rules for the system (or it may derive scaling rules based on service executions, their templates, and the dependencies of the services on each other). In step 872, the controller checks such rules to be used to facilitate the scaling.

[0183] If the new physical resources include storage resources, the storage resources may be moved from the single node or other form of the simpler IT system in step 874 (or the storage resources may be mirrored). If storage resources are moved, after the storage resources are moved, the computing resources or running resources may be reloaded or rebooted in step 876. In another example, the computing resources may be connected to the mirrored storage resources and left running in step 876, while the old storage resources or hardware resources of the previous system on the single-node system may be disconnected or disabled. For example, a running service may be attached to two mirrored block devices, one on the single-node server (e.g., using mdadm raid1) and the other on the storage resources, and then, once the data is synchronized, the drives on the single-node server may be disconnected. The previous hardware may remain part of the IT system and may run on the same node as a controller in mixed mode (step 878). The system may continue to repeat this migration process until the original node is the only one running the controller, at which point the system is distributed (step 880). Furthermore, at each step of the process flow of Figure 8E, the controller may update the system state 220 and record the changes to the system in a database (step 882).

[0184] Referring to FIG. 9A , application 910 is installed on resource 900. Resource 900 may be computational resource 310, storage resource 410, or networking resource 610 with respect to FIGS. 1-10 as described herein. Resource 900 may be a physical resource. A physical resource may comprise a physical machine or a physical IT system component. Resource 900 may be, for example, a physical computational resource, a physical storage resource, or a physical networking resource. Resource 900 may be coupled to controller 200 of system 100 by other resources, such as computational resources, networking resources, or storage resources, as described with respect to FIGS. 2A-10 herein.

[0185] The resource 900 may be initially powered down. The resource 900 may be coupled to a controller via a network, i.e., out-of-band management connection 260, in-band management connection 270, SAN 280, and / or network 290. The resource 900 may also be coupled to one or more application networks 390 through which services, application users, and / or clients can communicate with each other. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 915 or circuitry for the resource 900 that turns on when the resource 900 is connected. The device may enable features including, but not limited to, powering the device on / off, attaching to a console and typing commands, monitoring temperature and other computer health-related factors, and setting BIOS settings 195 and other features outside of the scope of the operating system.

[0186] Controller 200 may detect resource 900 through out-of-band management network 260. It may identify the type of resource and its configuration using in-band or out-of-band management. Controller logic 205 may be configured to consult out-of-band management 260 or in-band management 270 for additional hardware. If resource 900 is detected, controller logic 205 may use global system rules 220 to determine whether resource 900 is configured automatically or through user interaction. If resource 900 is added automatically, setup follows global system rules 210 in controller 200. If resource 900 is added by a user, global system rules 210 in controller 200 may ask the user to confirm the addition of the resource and what the user wants to do with the computing resource. Controller 200 may query an API application or otherwise request the user or any program controlling the stack for confirmation that the new resource has been authorized. The authorization process may also be completed automatically and securely using cryptography to verify the validity of the new resource. The resource 900 is then added to the IT system state 220, including the switch or network to which the resource 900 is connected.

[0187] The controller 200 may power on resources through the out-of-band management network 260. The controller 200 may use the out-of-band management connection 260 to power on physical resources and configure the BIOS 195. The controller 200 may automatically use the console 190 to select desired BIOS options, which may be achieved by the controller 200 reading a console image with image recognition and controlling the console 190 through out-of-band management. The boot-up state may be determined by image recognition through the console of the resource 900, out-of-band management with a virtual keyboard, querying a service utilizing the resource, or querying a service of the application 910. Some applications may have processes that allow the controller 200 to monitor, or in some cases, change the settings of the application 910 using in-band management 270.

[0188] An application 910 on a physical resource 900 (or of resources 300, 310, 311, 312, 313, 400, 410, 411, 412, 600, 610 as described with respect to FIGS. 1-10 herein) may boot over a SAN 280 or another network using a BIOS boot option or other method of configuring a remote boot, such as enabling a PXE boot or a Flex boot. Additionally or alternatively, the controller 200 may use out-of-band management 260 and / or in-band management connection 270 to instruct the physical resource 900 to boot an application image of image 950. The controller may configure a boot option on the resource or may use an existing available remote boot method, such as a PXE boot or a Flex boot. The controller 200 may optionally or alternatively use out-of-band management 260 to boot from an ISO image, configure a local disk, and then instruct the resource to boot from local disk(s) 920. The local disk(s) may be loaded with boot files. This may be accomplished using out-of-band management 260, image recognition, and a virtual keyboard. The resource may also have a boot file and / or boot loader installed. The resource 900 and application may boot from an image 950 loaded from a template 230, for example, via SAN 280, using global system rules 210 and controller logic 205. The global system rules 220 may specify the boot order. For example, the global system rules 220 may require that the resource 900 be booted first, followed by the application 910. Once the resource 900 is booted using the image 950, information received through the in-band management connection 270 about the resource 900 may be collected and added to the IT system state 220.Resource 900 may be added to a storage resource pool, which becomes a resource managed by controller 200 and tracked in IT system state 220. Applications 910 may be booted in the order specified in global system rules 220 using image 950 or application image 956 loaded on resource 900.

[0189] The controller 200 may configure networking resources 610 that connect the application 910 to the application network 390 via an out-of-band management connection 260 or another connection. The physical resource 900 may be connected to remote storage, such as block storage resources including, but not limited to, ISER (ISCSI over RDMA), NVMeOF FCOE, FC, or ISCSI, or another storage backend, such as SWIFT, GFUSTER, or CEPHFS. When a service or application is operational, the IT system state 220 may be updated using the out-of-band management connection 260 and / or the in-band management connection 270. The controller 200 may use the out-of-band management connection 260 or the in-band management connection 270 to determine the power state of the physical resource 900, i.e., whether it is on or off. The controller 200 may use the out-of-band management connection 260 or the in-band management connection 270 to determine whether a service or application is running or in a boot-up state. The controller may perform other functions based on the information it receives and the global system rules 210 .

[0190] 9B shows an image 950 that is loaded directly or indirectly (e.g., via another resource or database) from template 230 onto a compute node to boot application 910. Image 950 may comprise a custom kernel 941 for application 910.

[0191] Image 950 may include boot files 940 for the resource type and hardware. Boot files 940 may include kernels 941 corresponding to the resources, applications, or services being deployed. Boot files 940 may include an initrd or similar file system used to support the boot process. Boot system 940 may include multiple kernels or initrds configured for different hardware and resource types. Additionally, image 950 may include file systems 951. File systems 951 may include base images 952 and corresponding file systems, service images 953 and corresponding file systems, and volatile images 954 and corresponding file systems. The file systems and data loaded may vary depending on the resource type and application or service being run. Base image 952 may include a base operating system file system. The base operating system may be read-only. Base image 952 may also include basic tools of an operating system that is independent of what is being run. Base image 952 may include base directories and operating system tools. Service file systems 953 may include configuration files and specifications for resources, applications, or services. The volatile file system 594 may contain information or data specific to that deployment, such as binary applications, unique addresses, and other information, which may or may not be configured as variables, including but not limited to passwords, session keys, and private keys. The file systems may be mounted as one single file system using techniques such as overlayFS, allowing some read-only and some read-write file systems to reduce the amount of duplicate data used for applications.

[0192] FIG. 9C shows an example of installing an application from an NT package, which may be one type of template 230. In step 900.1, the controller determines that a package blob needs to be installed. In step 900.2, the controller creates a storage resource on the default data store for the blob type (block, file, file system). In step 900.3, the controller connects to the storage resource via an available storage transport for the storage resource type. In step 900.4, the controller copies the package blob to the connected storage resource. The controller then disconnects from the storage resource (step 900.5) and sets the storage resource to read-only (step 900.6). The package blob is then successfully installed (step 900.7).

[0193] In another example, Appendix B provides exemplary details regarding how a system may connect a computing resource to an overlayfs. Such techniques may be used to facilitate installing an application on a resource in accordance with Figure 9A or booting a computing resource from a storage resource in accordance with step 205.11 of Figure 2F.

[0194] 9D illustrates an application 910 deployed on a resource 900. The resource 900 may comprise a virtual computing resource, e.g., a compute node that may comprise a hypervisor 920, one or more virtual machines 921, 922, and / or containers. The resource 900 may be configured in a manner similar to that described herein with respect to FIGS. 1-10 using an image 950 loaded onto the resource 900. In this example, the resource 920 is shown as a hypervisor that manages the virtual machines 921, 922. The controller 200 may use in-band management 270 to communicate with the resource 900 hosting the hypervisor 920 that creates the resource and configures the resource to allocate appropriate hardware resources, including, but not limited to, CPU, RAM, GPU, a remote GPU (which may use RDMA to remotely connect to another host), a network connection, a network fabric connection, and / or virtual and physical connections to a partitioned and / or segmented network. The controller 200 may use a virtual console 190 (e.g., including but not limited to SPICE or VNC) and image recognition to control the resources 900 and the hypervisor 920. Additionally or alternatively, the controller 200 may use an out-of-band management 260 or an in-band management connection 270 to instruct the hypervisor 920 to boot an application image 950 from a template 230 using global system rules 210. The images 950 may be stored on the controller 200, or the controller 200 may move or copy them to the storage resource 410.The boot image for VMs 921, 922 may be stored locally, for example, as image 950, or as a file on a block device or remote host, and shared via file sharing, for example, NFS over RDMA / NFS using image types such as qcow2 or raw, or it may use a remote block device using ISCSI, ISER, NVMEOF, FC, FCOE. Portions of image 950 may be stored on storage resources 410 or compute nodes 310. Controller 200 may use global rules and / or templates to configure networking resources 610 to appropriately support applications over out-of-band management connection 260 or another connection. An application 910 on a resource 900 may boot using an image 950 loaded over a SAN 280 or another network using a BIOS boot option, or by enabling a hypervisor 920 on the resource 900 to connect to block storage resources, including, but not limited to, ISER (ISCSI over RDMA), NVMEOF FCOE, FC, or ISCSI, or another storage backend such as SWIFT, GFUSTER, or CEPHFS. Storage resources may be copied from a template target on the storage resource. IT system state 220 may be updated by querying the hypervisor 920 for information. An in-band management connection 270 may communicate with the hypervisor 920 and may be used to determine the power state of the resource, i.e., whether it is on or off, or to determine the boot-up state. The hypervisor 920 may use a virtual in-band connection 923 to the virtualized application 910 or may use the hypervisor 920 for functions similar to out-of-band management. This information may indicate whether a service or application is available to run because it has been powered on or booted.

[0195] The boot-up state may be determined by visual recognition through the console 190 of the resource 900, out-of-band management 260 with a virtual keyboard, querying a service utilizing the resource, or querying the application 910's own services. Some applications may have processes that allow the controller 200 to monitor and, in some cases, change the configuration of the application 910 using in-band management 270. Some applications may reside on virtual resources, which the controller 200 may monitor by communicating with the hypervisor 920 using in-band management 270 (or out-of-band management 260). The application 910 may not have such processes for monitoring and / or adding input (or such processes may be switched to save the resource). In such cases, the controller 200 may use the out-of-band management connection 260 and / or visual processing and / or a virtual keyboard to log on to the system to make changes and / or switch on management processes. Similar to virtual computing resources, a virtual machine console 190 may be used.

[0196] FIG. 9E illustrates an exemplary process flow for adding a virtual computing resource host to IT system 100. In step 900.11, a host available as a virtual computing resource is added to the system. The controller may configure a bare-metal server according to the process flow of FIG. 15B (step 900.12). Alternatively, an operating system may be preloaded and / or the host may be preconfigured (step 900.13). The resource is then added to system state 220 as a virtual computing resource pool (step 900.14), and the resource becomes accessible via an API from controller 200 (step 900.15). The API is typically accessed through in-band management connection 270. However, in-band management connection 270 may be selectively enabled and / or disabled with a virtual keyboard. The controller may then use out-of-band management connection 260 and the virtual keyboard and monitor to communicate over out-of-band connection 260 (step 900.16). Now, in step 900.17, the controller can utilize the new resource as a virtual computing resource.

[0197] Exemplary Multi-Controller System Referring to Figure 10, a system 100 is shown having computing resources 300, 310 as described herein with reference to Figures 1-10, comprising multiple physical computing nodes 311, 312, 313, storage resources 400, 410 as described herein in the form of multiple storage nodes 411, 412 and JBODs 413, multiple controllers 200a, 200b including components 205, 210, 220, 230 (Figures 1-9C) and configured as controller 200 as described herein, networking resources 600, 610 as described herein, comprising multiple fabrics 611, 612, 613, and an application network 390.

[0198] FIG. 10 illustrates one possible configuration of components of system 100 in an exemplary embodiment, but does not limit the possible configurations of components of system 100.

[0199] User interface or application 110 communicates with API application 120, which communicates with either or both of controllers 200a and 200b. Controllers 200a and 200b may be coupled to out-of-band management connection 260, in-band management connection 270, SAN 280, or network in-band management connection 290. As described with reference to Figures 1-9C herein, controllers 200a and 200b are coupled to compute nodes 311, 312, and 313, storage 411 and 412, including JBOD 413, and networking resources 610 via connections 260, 270, 280, and optionally 290. Application network 390 is coupled to compute nodes 311, 312, and 313, storage resources 411, 412, and 413, and networking resources 610.

[0200] The controllers 200a, 200b may operate in parallel. Either controller 200a or 200b may initially operate as the master controller 200, as described with respect to FIGS. 1 through 9C herein. The controller(s) 200a, 200b may be configured to configure the entire system 100 from a powered-off state. One of the controllers 200a, 200b may further generate the system state 220 from an existing configuration by probing the other controllers through one of the out-of-band and in-band connections 260, 270. Either of the controllers 200a, 200b may access or receive resource status and related information from resources or other controllers through one or more connections 260, 270. The controller or other resources may update the other controller. Thus, when an additional controller is added to the system, the controller may be configured to return the system 100 to the system state 220. If one of the controllers or the master controller fails, the other controller may be designated as the master controller. The IT system state 220 may further be reconstructable from status information available or stored on the resource. For example, an application may be deployed on a computing resource where the application is configured to create a virtual computing resource on which the system state is stored or replicated. The global system rules 210, system state 220, and templates 230 may further be saved or copied to a resource or combination of resources. Thus, if all controllers go offline and a new controller is added, the system may be configured so that the new controller can restore the system state 220.

[0201] Networking resources 610 may include multiple network fabrics. For example, as shown in FIG. 10, the multiple network fabrics may include one or more of an SDN Ethernet switch 611, a ROCE switch 612, an Infiniband switch 613, or other switches or fabrics 614. A hypervisor provides virtual machines on compute nodes that may connect to physical or virtual switches utilizing one or more of the necessary fabrics. Network configurations may allow restriction of the physical network, for example, through segmented networks, for example, for security or other resource optimization.

[0202] The system 100, through the controller 200 described in Figures 1-10 herein, may automatically set up services or applications. A user may request, through the user interface 110 or an application, that a service be set up for the system 100. The service may include, but is not limited to, email services, web services, user management services, network providers, LDAP, developer tools, VOIP, authentication tools, and accounting software. The API application 120 translates the user or application request and sends a message to the controller 200. The service template or image 230 of the controller 200 is used to identify the resources required for the service. The required resources are identified based on their availability according to the system state 220. The controller 200 makes a request to the computing resources 310 or compute nodes 311, 312, or 313 for the required computational services, to the storage resources 410 for the required storage resources, and to the networking resources 610 for the required networking resources. The system state 220 is then updated to identify the resources to be allocated. The service is then installed on the allocated resources using global system rules 210 according to the service template.

[0203] Improving system security Referring to FIG. 13A , an IT system 100 is shown, where the system 100 includes a resource 1310, which may be a bare metal or physical resource. While FIG. 13A shows only a single resource 1310 connected to the system 100, it should be understood that the system 100 may include multiple resources 1310. The resource 1310(s) may be or comprise a bare metal cloud node. A bare metal cloud node may include, but is not limited to, a resource connected to an external network 1380 that allows remote access to physical hosts or virtual machines, allows the creation of virtual machines, and allows external users to execute code on the resource(s). The resource(s) 1310(s) may be directly or indirectly connected to the external network 1380 or the application network 390. The external network 1380 may be the Internet or other resource(s) not managed by the controller 200 or a controller of the IT system 100. The external network 1380 may include, but is not limited to, the Internet, an internet connection(s), a resource(s) not managed by the controller, other wide area networks (e.g., Stratcom, a peer-to-peer mesh network, or other external networks whether or not publicly accessible), or other networks.

[0204] When a physical resource 1310 is added to IT system 100a, it may be coupled to controller 200 and powered off. Resource 1310 is coupled to controller 200a via one or more networks, such as out-of-band management (OOBM) connection 260, optionally in-band management (IBM) connection 270, and optionally SAN connection 280. As used herein, SAN 280 may or may not comprise a constituent SAN. A constituent SAN may comprise a SAN used to power on or configure a physical resource. A constituent SAN may be part of SAN 280 or may be separate from SAN 280. In-band management may further comprise a constituent SAN, which may or may not be SAN 280, as illustrated herein. Furthermore, when a resource is in use, a constituent SAN may be disabled, disconnected, or unavailable. OOBM connection 260 is not visible to the OS of system 100, but IBM connection 270 and / or the constituent SAN may be visible to the OS of system 100. The controller 200 of FIG. 13A may be configured similarly to the controller 200 described with reference to FIGS. 1-12B herein. The resource 1310 may comprise internal storage. In some configurations, the controller 200 may generate storage and temporarily configure the resource to connect to a SAN to fetch data and / or information. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 315 or to circuitry in the resource 1310 that is powered on when the resource 1310 is plugged in. The device 315 may enable functions including, but not limited to, powering the device on / off, connecting to a console and entering commands, monitoring temperature and other computer health-related factors, and setting BIOS settings and other functions outside of the operating system. The controller 200 can reference the resource 1310 through the out-of-band management network 260. The controller can also identify the type of resource and its configuration using in-band or out-of-band management.13C-13E, described below, illustrate various process flows for adding physical resources 1310 to IT system 100a and / or for starting up or managing system 100 in a manner that enhances system security.

[0205] The term "disable" as used herein with reference to a network, networking resource, network device, and / or network interface refers to the action of such network, networking resource, network device, and / or network interface being powered off (manually or automatically), physically disconnected, and / or virtually disconnected, or disconnected in some other manner (e.g., by filtering) from a network, virtual network (e.g., including, but not limited to, a VLAN, VXLAN, or InfiniBand partition). The term "disable" also includes a unidirectional or unidirectional restriction of operability, such as preventing a resource from sending or writing data to a destination (while still having the ability to receive or read data from a source), or preventing a resource from receiving or reading data from a source (while still having the ability to send or write data to a destination). Such a network, networking resource, network device, and / or network interface may be disconnected from an additional network, virtual network, or resource binding, but remain connected to the previously connected network, virtual network, or resource binding. Additionally, such a networking resource or device may be switched from one network, virtual network, or resource binding to another.

[0206] The term "enabled" as used herein with reference to a network, networking resource, network device, and / or network interface refers to the action of such network, networking resource, network device, and / or network interface being powered on (manually or automatically), physically connected, and / or virtually connected, or in some other manner connected to a network, virtual network (e.g., including, but not limited to, a VLAN, VXLAN, or InfiniBand partition). Such a network, networking resource, network device, and / or network interface, if already connected to another system component, may be connected to additional networks, virtual networks, or resource combinations. Additionally, such a networking resource or device may be switched from one network, virtual network, or resource combination to another. The term "enabled" also includes unidirectional or unidirectional restrictions on operability, such as allowing a resource to send, write, or receive data to a destination (while having the ability to restrict data from the source) or allowing a resource to send, receive, or read data from a source (while having the ability to restrict data from the destination).

[0207] The controller logic 205 is configured to examine the out-of-band management connection 260 or in-band management connection 270 and / or configuration SAN 280 of the added hardware. If resource 1310 is detected, the controller logic 205 may use global system rules 220 to determine whether to configure the resource automatically or by interacting with a user. If added automatically, the setup will follow global system rules 210 in the controller 200. If added by a user, the global system rules 210 in the controller 200 may query the user about adding the resource and what the user wants to do with the resource 1310. The controller 200 may query an API application or otherwise issue a request to the user or any program controlling the stack to verify that the new resource is authenticated. The authentication process may also be completed automatically and securely using cryptography to verify the legitimacy of the new resource. The controller logic 205 then adds the resource 1310 to the IT system state 220, including the switch or network to which the resource 1310 is plugged.

[0208] If the resource is physical, the controller 200 may power on the resource through the out-of-band management network 260, and the resource 1310 may use global system rules 210 and controller logic 205 to boot off the image 350 loaded from the template 230, for example, via the SAN 280. The image may be loaded through other network connections or indirectly via another resource. Once booted, information about the resource 1310 may also be collected and added to the IT system state 220. This may be done through in-band management and / or a configuration SAN or out-of-band management connection. The resource 1310 may use global system rules 210 and controller logic 205 to boot off the image 350 loaded from the template 230, for example, via the SAN 280. The image may be loaded through other network connections or indirectly via another resource. Once booted, information received through the in-band management connection 270 about the computing resource 310 may also be collected and added to the IT system state 220. The resource 1310 may then be added to the storage resource pool, which becomes a resource managed by the controller 200 and tracked in the IT system state 220 .

[0209] The in-band management and / or configuration SAN may be used by the controller 200 to set up, manage, use, or communicate with the resource 1310 and execute any command or task. However, optionally, the in-band management connection 270 may be configured by the controller 200 to be turned off or disabled at any time or during the setup, management, use, or operation of the system 100 or the controller 200. The in-band management may further be configured to be turned on or enabled at any time or during the setup, management, use, or operation of the system 100 or the controller 200. Optionally, the controller 200 may controllably or switchably disconnect the resource 1310 from the in-band management connection 270 to the controller(s) 200. Such disconnection or disconnectability may be physical, for example, using an automatic physical switch or a switch that powers off the resource's in-band management connection to the network and / or the configuration SAN. For example, the disconnection may be performed by a network switch that shuts off power to the port connected to the in-band management 270 and / or the configuration SAN 280 of the resource 1310. Such disconnection or partial disconnection may occur using the software-defined network, or may be physically filtered to the controller using the software-defined network. Such disconnection may occur via the controller, either through in-band management or out-of-band management. According to an exemplary embodiment, the resource 1310 may be disconnected from the in-band management connection 270 in response to a selective control command from the controller 200 at any time before, during, or after the resource 1310 is added to the IT system.

[0210] Using a software-defined network, the in-band management connection 270 and / or the constituent SAN 280 may or may not retain some functionality. The in-band management connection 270 and / or the constituent SAN 280 may be used as a restricted connection for communication between the controller 200 or other resources. The connection 270 may be restricted to prevent an attacker from pivoting to the controller 200, other networks, or other resources. The system may be configured to prevent devices such as the controller 200 and the resources 1310 from communicating openly and compromising the resources 1310. For example, the in-band management 270 and / or the constituent SAN 280 may only allow data transmission and prohibit any reception of the in-band management and / or constituent SAN, either through a software-defined network or hardware modification methods (such as electronic restrictions). The in-band management and / or constituent SAN may be configured, either physically or using a software-defined network that only allows writes from the controller to the resources, to be a one-way write component or as a one-way write connection from the controller 200 to the resources 1310. The one-way write nature of the connection can be further controlled or turned on or off according to desired security conditions and various stages or times of system operation. The system can be further configured to limit writes or communications from resources to the controller, for example, to communicate logs or alerts. Interfaces can also be moved to other networks or added or removed from networks through techniques including, but not limited to, software-defined networking, VLAN, VXLAN, and / or InfiniBand partitioning. For example, an interface can be connected to a configuration network, removed from that network, and moved to a network used at runtime. Controller-to-resource communications can be disconnected or limited, resulting in the controller being physically unable to respond to any data sent from the resource 1310.According to one example, once a resource 1310 is added and boots, the in-band management 270 can be switched off or filtered, either physically or using a software-defined network. The in-band management can be configured to send data to another resource dedicated to log management.

[0211] In-band management can be turned on and off using out-of-band management or software-defined networking. When in-band management is disconnected, there is no need to run a daemon and in-band management can be re-enabled using keyboard functionality.

[0212] Additionally, optionally, resources 1310 may not have in-band management connectivity and may be managed through out-of-band management.

[0213] Out-of-band management may alternatively or additionally be used to manipulate various aspects of the system, including, but not limited to, for example, connecting a keyboard, a virtual keyboard, a disk-mounted console, a virtual disk, changing BIOS settings, modifying boot parameters and other aspects of the system, executing pre-existing scripts that may reside on a bootable image or installation CD, or other functions of out-of-band management that allow the controller 200 and the resources 1310 to communicate with each other, whether or not exposed to an operating system running on the resources 1310. For example, the controller 200 may send commands using such tools via out-of-band management 260. The controller 200 may further use image recognition to assist in controlling the resources 1310. Thus, using an out-of-band management connection, the system may prevent or avoid undesired manipulation of resources connected to the system via the out-of-band management connection. The out-of-band management connection may also be configured as a one-way communication system during system operation or at selected times during system operation.

[0214] Additionally, the out-of-band management connection 260 may be selectively controlled by the controller 200 in the same manner as the in-band management connection, if the implementer so desires.

[0215] The controller 200 can automatically turn resources on and off according to global system rules and update the state of the IT system for reasons determined by the IT system user, such as turning resources off to conserve power, turning resources on to improve application performance, or any other reason the IT system user may consider. The controller may also be able to turn constituent SAN, in-band, and out-of-band management connections on and off, or designate such connections as one-way write connections at all times during system operation or for various security purposes (e.g., disabling the in-band management connection 270 or constituent SAN 280 while the resource 1310 is connected to the external network 1380 or the internal network 390). One-way in-band management may also be used, for example, to monitor the health of the system and would monitor logs and information that may be displayed to the operating system.

[0216] The resources 1310 may further be coupled to one or more internal networks 390, such as an application network, through which services, application users, and / or clients can communicate with one another. Such application networks 390 may further be connected or connectable to an external network 1380. According to example embodiments herein, including but not limited to FIGS. 2A-12B , in-band management may be disconnected or disconnectable from the resource or application network 390, or may provide one-way writes from the controller to provide additional security when the resource or application network is connected to an external network, or when the resource is connected to an application network that is not connected to an external network.

[0217] The IT system 100 of FIG. 13A may be configured similarly to the IT system 100 shown in FIG. 3B. An image 350 may be loaded directly or indirectly (through another resource or database) from the template 230 onto the resource 1310 for booting the computing resource and / or loading applications. The image 350 may include a boot file 340 for the resource type and hardware. The boot file 340 may include a kernel 341 corresponding to the resource, application, or service being deployed. The boot file 340 may further include an initrd or similar file system used to support the boot process. The boot system 340 may include multiple kernels or initrds configured for different hardware and resource types. Additionally, the image 350 may include a file system 351. The file system 351 may include a base image 352 and corresponding file systems, a service image 353 and corresponding file systems, and a volatile image 354 and corresponding file systems. The file systems and data loaded may vary depending on the resource type and the application or service being executed. The base image 352 may include a base operating system file system. The base operating system may be read-only. The base image 352 may further include the basic tools of an operating system independent of what is running. The base image 352 may include base directories and operating system tools. The service file system 353 may include configuration files and specifications for resources, applications, or services. The volatile file system 354 may include information or data specific to the deployment, such as binary applications, specific addresses, and other information, which may or may not be configured as variables, including, but not limited to, passwords, session keys, and private keys.File systems can be combined into a single file system using techniques such as overlayFS, some read-only and some read / write file systems, to reduce the amount of duplicate data used by applications.

[0218] FIG. 13B illustrates multiple resources 1310, each comprising one or more hypervisors 1311 that host or comprise one or more virtual machines. Controller 200a is coupled to resources 1310, each comprising bare-metal resources. As shown and described with reference to FIG. 13B, resources 1310 are each coupled to controller 200a. According to exemplary embodiments herein, in-band management connection 270, configuration SAN 280, and / or out-of-band management connection 260 may be configured as described with reference to FIG. 13A. One or more virtual machines or hypervisors may be compromised or become compromised. In conventional systems, other virtual machines on other hypervisors may then become compromised. For example, this may result from exploitation of a hypervisor running within a virtual machine. For example, a pivot may be moved from the compromised hypervisor to controller 200a, and then from the compromised controller 200a to another hypervisor coupled to controller 200a. For example, a pivot may occur between a compromised hypervisor and a target hypervisor using a network connected to both. In the configuration of in-band management 270, constituent SAN 280, or out-of-band management 260 of controller 200a and resource 1310 shown in Figure 13B, some or all of the in-band (or constituent SAN) and / or out-of-band connections may be selectively controlled to disable certain links between controller 200a and resource 1310, which may prevent a compromised virtual machine from originating from one hypervisor and being used to pivot to another resource.

[0219] As described with respect to Figures 1-12 above, the in-band management connection 270 and the out-of-band management connection 260 may be further configured in a manner similar to that described with respect to Figures 13A and 13B.

[0220] 13C illustrates an exemplary process flow for adding or managing physical resources, such as bare metal nodes, to system 100. Resources 1310 illustrated in FIGS. 13A and 13B herein or with respect to FIGS. 1-12 may be connected via out-of-band management connections 260 and in-band management connections 270 to controllers of system 100 and / or via a SAN.

[0221] After the instance of resource connection, the external network and / or application network are disabled in step 1370. As discussed above, this disabling can use any of a variety of techniques. For example, before setting up the system, adding resources, testing the system, updating the system, or performing other tasks or commands, components of system 100 (or only those vulnerable to attack) can be disabled, disconnected, or filtered from the external network or application network using an in-band management connection or configuration SAN, as described with respect to Figures 13A and 13B.

[0222] After step 1370, next in step 1371, the in-band management connection and / or configuration SAN is enabled. Thus, the combination of steps 1370 and 1371 isolates the resource from the external network and / or application network while the in-band management and / or SAN connection is active. Commands can then be executed on the resource under the control of controller 200 via the in-band management connection (see step 1372). For example, setup and configuration steps, including but not limited to those described herein with respect to FIGS. 1-13B, can be performed in step 1372 using the in-band management and / or configuration SAN. Alternatively or additionally, the in-band management and / or configuration SAN may be used in step 1372 to perform other tasks including, but not limited to, system operation, updating or management (which may include, but is not limited to, any change management or system updates), testing, updates, data transfers, gathering performance and health information (including, but not limited to, errors, CPU utilization, network utilization, file system information, and storage utilization) and log collection, and other commands that may be used to manage system 100 as described in Figures 1 through 13B herein.

[0223] After adding resources, setting up the system, and executing such tasks or commands as described herein with respect to FIGURES 13A and 13B, the in-band management connection 270 and / or the constituent SAN 280 between the resources and the controller or other components of the system may be disabled in one or more directions in step 1373. Such disabling may use disconnection, filtering, etc., as described above. After step 1373, connectivity to the external network and / or application network may be restored in step 1374. For example, the controller may notify the networking resource so that the resource 1310 can connect to the application network or the Internet. The same steps may be performed when testing or updating the system, i.e., the in-band management connection to the external network and / or application network may be disconnected or filtered, and then the in-band management connection to the resource (in one or both directions) may be enabled or connected. Thus, steps 1373 and 1374 operate simultaneously to isolate the resource from connectivity to the controller through the in-band management connection and / or constituent SAN while the resource is connected to the external network and / or application network.

[0224] Out-of-band management can be used to manage systems or resources, set up, configure, boot, or add systems or resources. When used in any embodiment herein, out-of-band management can use a virtual keyboard to send commands to a machine to change settings before booting, and can also send commands to the operating system by typing on the virtual keyboard. If the machine is not logged in, out-of-band management can use the virtual keyboard to enter a username and password and use image recognition to confirm the logon, verify the commands entered, and whether they were executed. If the physical resource only has a graphical console, a virtual mouse can also be used, and image recognition will allow out-of-band management to make changes.

[0225] FIG. 13D is another exemplary process flow for adding or managing a physical resource, such as a bare metal node, to system 100. At step 1380, a resource shown in FIGS. 13A and 13B or 1-12 herein may be connected to a system or resource via out-of-band management 260. A disk may be virtually connected by providing access to a disk image (e.g., an ISO image) through out-of-band management facilitated by a controller (see step 1381). The resource or system may then be booted from the disk image (step 1382), and files are then copied from the disk image to a bootable disk (see step 1383). This can also be used to boot a system in which resources have been set up in this manner using out-of-band management. This can also be used to configure and / or boot multiple resources that can be combined together (including, but not limited to, networking resources), regardless of whether the multiple resources further comprise a controller or comprise a system. Thus, a virtual disk may be used to enable a controller to connect a disk image to a resource as if a virtual disk were connected to the resource. Out-of-band management can also be used to send files to a resource. Data may be copied from the virtual disk to a local disk in step 1383. The disk image may include files that the resource can copy and use during operation. Files may be copied or used either through a scheduled program or instructions from out-of-band management. Through out-of-band management, a controller may log on to the resource using a virtual keyboard and enter commands to copy files from the virtual disk to the controller's own disk or other storage accessible to the resource. In step 1384, the system or resource is configured to boot by setting BIOS, EFI, or boot order settings so that it will boot from a bootable disk.Boot configuration may use the operating system's EFI manager, such as efibootmgr, which can be performed directly from out-of-band management or by including it in an installer script (e.g., a script using efibootmgr is automatically executed when the resource boots). Additionally, boot options or other BIOS changes may be set through an out-of-band management tool, such as Supermicro Boot Manager, using either boot order commands or by uploading a BIOS configuration (e.g., an XML BIOS configuration supported by Supermicro Update Manager). The BIOS can also be configured to set the appropriate BIOS settings, including boot order, using image recognition from the keyboard and console. The installer can be run with the configured image loaded. The configuration can be tested by viewing the screen and using image recognition. After configuration, the resource can be enabled (e.g., powered on, booted, connected to the application network, or a combination thereof) (step 1385).

[0226] 13E is another exemplary process flow for adding or managing physical resources, such as bare metal nodes, to system 100, in this case using PXE, Flexboot, or similar network booting. In step 1390, resources 1310 shown in FIGS. 13A and 13B herein or shown with respect to FIGS. 1-12 may be connected to a controller of system 100 via (1) in-band management connection 270 and / or SAN and (2) out-of-band management connection 260. Next, external network and / or application network connections may be disabled (e.g., filtered or disconnected, in whole or in part, physically or virtually using SDN) in step 1391 (similar to that described above with respect to step 1370). For example, before setting up the system, adding resources, testing the system, updating the system, or performing other tasks or commands, components of system 100 (or only those vulnerable to attack) are disabled, disconnected, or filtered from the external or application network using an in-band management connection or SAN, as described with respect to Figures 13A and 13B.

[0227] In step 1392, the type of resource is determined. For example, information about the resource can be gathered from the MAC address using an out-of-band management tool, or by connecting a disk image (e.g., an ISO image) to the resource as if a disk were connected to the resource and temporarily booting the operating system with a tool that can be used to identify the resource information. Next, in step 1393, the resource is identified as being configured or pre-configured for PXE or flexboot, etc. Next, in step 1394, the resource is powered on and a PXE, Flexboot, or similar boot is performed (or if the resource temporarily boots and is powered on again). Next, in step 1395, the resource boots from the in-band management connection or SAN. In step 1396, data is copied to a disk accessible by the resource in a manner similar to that described with reference to step 1383 of FIG. 13D. In step 1397, the resource is configured to boot from the disk(s) in a manner similar to that described above with reference to step 1384 of FIG. 13D. If the resource is identified as being pre-configured for PXE, flexboot, etc., the files may be copied in any of steps 1393 through 1396. If in-band management was enabled, it may be disabled in step 1398 and the application network or external network may be reconnected or enabled in step 1399.

[0228] Additionally, it should be understood that technologies other than OOBM can be used to remotely enable resources (e.g., power them on) and confirm that they have booted. For example, the system can prompt the user to press the power button, manually tell the controller that the system has booted (or use a console connection to the keyboard / controller). Additionally, the system can ping the controller through IBM (e.g., through ssh, telnet, or another method on the network) once the system has booted and the controller has logged on and instructed it to reboot. For example, the controller could use ssh to send a reboot command. If PXE is used and there is no OOBM, then in either case the system would have a way to either remotely instruct the resource to power on or to instruct the user to manually power on.

[0229] Deploying Controllers and / or Environments In an exemplary embodiment, controllers may be deployed in a system from an originating controller 200 (such an originating controller 200 may be referred to as a “main controller”). The main controller may therefore set up a system or environment that may be a separate or separable IT system or environment.

[0230] As described herein, an environment refers to a collection of resources within a computer system that can interoperate with each other. A computer system may, but is not required to, include multiple environments within it. The resource(s) of an environment may comprise one or more instances, applications, or sub-applications executing in the environment. Additionally, an environment may comprise one or more environments or sub-environments. An environment may or may not include a controller, and an environment may operate one or more applications. Such resources of an environment may include, for example, networking resources, computing resources, storage resources, and / or application networks used to execute a particular environment, including the applications within the environment. Thus, it should be understood that an environment may provide functionality for one or more applications. In some examples, an environment described herein may be physically or virtually isolated or separable from other environments. Additionally, in other examples, an environment may have network connections to other environments, and such connections may be disabled or enabled as needed.

[0231] Additionally, the main controller may set up, deploy, and / or manage one or more additional controllers in various environments or as separate systems. Such additional controllers may be independent or stand alone from the main controller. Even if independent or pseudo-independent from the main controller, such additional controllers may receive instructions from or send information to the main controller (or a separate monitor or environment via a monitoring application) at various times during operation. Environments may be configured for security (e.g., by making environments separable from each other and / or from the main controller) and / or for various management purposes. Environments may be connected to an external network, while other related environments may or may not be connected to an external network.

[0232] The main controller can manage environments or applications, regardless of whether they are separate systems and whether they include controllers or sub-controllers. The main controller can also manage shared storage of global configuration files or other data. The main controller can also parse global system rules (e.g., system rules 210) or a subset of the main controller for different controllers depending on their capabilities. Each new controller (sometimes referred to as a "sub-controller") can receive new configuration rules, which may be a subset of the main controller's configuration rules. The subset of global configuration rules deployed to a controller may depend on or correspond to the type of IT system being set up. The main controller can set up or deploy new controllers or separate IT systems, which are then permanently separated from the main controller, for example, for shipping or delivery or other reasons. The global configuration rules (or a subset thereof) can define a framework for setting up applications or sub-applications in various environments and how they can interact with each other. Such applications or environments can run on sub-controllers that include a subset of the global configuration rules deployed by the main controller. In some examples, such applications or environments can be managed by the main controller. However, in other instances, such applications or environments are not managed by the main controller. When a new controller is spawned from the main controller to manage an application or environment, application dependency checking can be performed across multiple applications to facilitate control by the new controller.

[0233] Thus, in an exemplary embodiment, a system may include a main controller configured to deploy other controllers, or an IT system including such other controllers. Such an implemented system may be configured to be completely disconnected from the main controller. Once independent, such a system may be configured to operate as a standalone system, or may be controlled or monitored by another controller, such as the main controller (or an environment with an application), at various discrete or continuous times during operation.

[0234] 14A illustrates an exemplary system in which a main controller 1401 has controllers 1401a and 1401b deployed on different systems 1400a and 1400b, respectively (where 1400a and 1400b may be referred to as subsystems, although it should be understood that subsystems 1400a and 1400b may also function as environments). The main controller 1401 may be configured in a manner similar to the controller 200 described above. As such, it may include controller logic 205, global system rules 210, system states 220, and templates 230.

[0235] Systems 1400a and 1400b each include a controller 1401a, 1401b coupled to resources 1420a, 1420b, respectively. The main controller 1401 may be coupled to one or more other controllers, such as controller 1401a of subsystem 1400a and controller 1401b of subsystem 1400b. Global rules 210 of the main controller 1400 may include rules that can manage and control the other controllers. Using such global rules 210 in conjunction with controller logic 205, system state 220, and templates 230, the main controller 1401 can set up, provision, and deploy subsystems 1400a, 1400b through controllers 1401a, 1401b in a manner similar to that described with reference to Figures 1-13E herein.

[0236] For example, the main controller 1401 can load the global rules 210 (or a subset thereof) into the subsystems 1400a, 1400b as rules 1410a, 1410b, respectively, in such a way that the global rules 210 (or a subset thereof) direct the operation of the controllers 1401a, 1401b and their subsystems 1400a, 1400b. Each controller 1401a, 1401b can have rules 1410a, 1410b, which can be the same or different subsets of the global rules 210. For example, which subset of the global rules 210 is provisioned to a given subsystem can depend on the type of subsystem being deployed. Additionally, the controller 1401 can load or send data to be loaded into the system resources 1420a, 1420b or the controllers 1401a, 1401b.

[0237] The main controller 1401 may be connected to the other controllers 1401 a, 1401 b through in-band management connection(s) 270(s) and / or out-of-band management connection(s) 260(s) or SAN connection 280, which may be enabled or disabled at various stages of deployment or management in a manner as described herein, for example, with reference to the resource deployment and management illustrated in Figures 13A-13E. Using selective enabling and disabling of the in-band management connections 270 or out-of-band management connections 260, the subsystems 1400 a, 1400 b may be deployed in a way that the subsystems 1400 a, 1400 b have no knowledge (or knowledge is limited, controlled, or prohibited) of the main system 100 or the controllers 1401 or each other at various times.

[0238] In an exemplary embodiment, the main controller 1401 can operate a centralized IT system with local controllers 1401a, 1401b deployed and configured by the main controller 1401, allowing the main controller 1401 to deploy and / or run multiple IT systems. Such IT systems may or may not be independent of each other. The main controller 1401 can set up monitoring as a separate application, isolated or air-gapped from the IT systems it created. A separate console for monitoring may include connections between the main controller and local controller(s) and / or connections between environments that can be selectively enabled or disabled. The controller 1401 can deploy isolated systems for various applications, each with a different controller in case of shutdown or compromise, including, but not limited to, business, manufacturing systems with data storage, data centers, and various other functional nodes. Such isolation may be complete or permanent, or may be pseudo-isolated, for example, temporarily, depending on time or task, depending on communication direction, or depending on other parameters. For example, the main controller 1401 may be configured to provide instructions to the system that may or may not be limited to some defined circumstances, while the subsystems may have limited or no ability to communicate with the main controller. Therefore, such subsystems may not be able to compromise the main controller 1401. The main controller 1401 and the sub-controllers 1401a, 1401b may be isolated from each other as described herein (in specific examples described below), for example, by disabling in-band management 270, by one-way writes, and / or by restricting communications to out-of-band management 260. For example, in the event of a breach, one or more controllers may disable in-band management connection 270 to one or more other controllers to prevent the breach or access from spreading.System sections can be turned off or isolated.

[0239] The subsystems 1400 a , 1400 b may further share resources with or connect to other environments or systems through in-band management 270 or out-of-band management 260 .

[0240] 14B and 14C are an exemplary flow illustrating possible steps for provisioning a controller with a main controller.

[0241] In FIG. 14B , in step 1460, the main controller provisions or sets up a resource, such as resource 1420a or 1420b. In step 1461, the main controller provisions or sets up a sub-controller. The main controller can use the techniques described above to set up resources in the system and perform steps 1460 and 1461. Additionally, while FIG. 14B shows that step 1460 is performed before step 1461, it should be understood that this is not required. The main controller 1401 can use its system rules 210 to determine what resources are needed and to deploy the resources on the system or network. The main controller can set up or deploy a sub-controller in step 1461 by loading system rules 210 into the system to set up the sub-controller (or by providing instructions to the sub-controller on how to set up and obtain its own system rules). These instructions may include, but are not limited to, configuring resources, configuring applications, global system rules to create the IT systems run by the sub-controllers, instructions to reconnect to the main controller to gather new or changed rules, and instructions to disconnect from the application network to make space for the new production environment. After deploying the resources, in step 1463 the main controller can allocate resources to the sub-controllers via updates to system rules 210 and / or system state 220.

[0242] Figure 14C shows an alternative process flow for deployment. In the example of Figure 14C, the main controller deploys the sub-controllers in step 1470 (which may proceed as described with respect to step 1461). Then, in step 1475, the sub-controllers deploy resources using techniques such as those shown in Figures 3C and 7B.

[0243] 15A shows an example system in which a main controller 1501 of system 100 generates environments 1502, 1503, and 1504. Environment 1502 includes resource 1522, environment 1503 includes resource 1523, and environment 1504 includes resource 1524. Additionally, environments 1502, 1503, and 1504 may share access to a pool of shared resources 1525. Such shared resources may include, for example, but are not limited to, shared data sets, APIs, or running applications that need to communicate with each other.

[0244] In the example of FIG. 15A , each environment 1502, 1503, 1504 shares a main controller 1501. The global system rules 210 of the main controller 1501 may include rules for deploying and managing the environments. Resources 1522, 1523, and / or 1524 may be required by each environment 1501, 1502, 1503 to manage one or more applications. Configuration rules for such applications may be implemented by the main controller (or local controllers within the environments, if present) to define how each such environment operates and how it interacts with other applications and environments. The main controller 1401 can use the global rules 210, along with the controller logic 205, system state 220, and templates 230, to set up, provision, and deploy environments in a manner similar to the deployment of resources and systems described with reference to FIGS. 1-14C herein. If the environment includes a local controller, the main controller 1501 can load the global rules 210 (or a subset thereof) into the local controller or associated storage, such that the global rules (or a subset thereof) define the behavior of the environment.

[0245] The controller 1501 can use configuration rules comprising the system rules 210 to deploy and configure resources 1522, 1523, 1524 and / or shared resources 1525 for each of the environments 1502, 1503, 1504. The controller 1501 can also monitor the environments or configure resources 1522, 1523, 1524 (or shared resources 1525) to enable monitoring of each of the environments 1502, 1503, 1504. Such monitoring can be through a connection to a separate monitoring console that can be enabled or disabled, or through the main controller. The main controller 1501 may be connected to one or more of the environments 1502, 1503, 1504 through in-band management connection(s) 270 and / or out-of-band management connection(s) 260 or SAN connection 280, which may be enabled or disabled at various stages of deployment or management, in a manner as described herein with reference to the deployment and management of resources in Figures 13A-13E and 14A. Using the enabling and disabling of the in-band management connections 270 or out-of-band management connections 260 or SAN connections 280, the environments 1502, 1503, 1504 may be deployed in such a way that they have, at various times, no, limited, or controlled knowledge of the main system 100 or controller 1501, or no, limited, or controlled knowledge of each other's connections.

[0246] An environment can include one or more resources that are coupled to an external network 1580 that connects to an external environment or that interact with other resources. An environment can be physical or non-physical. "Non-physical" in this context means that the environments share the same physical host(s) but are virtually isolated from each other. The environments and systems can be deployed on identical hardware, similar but different hardware, or non-identical hardware. In some examples, environments 1502, 1503, and 1504 can be valid copies of each other, while in other examples, environments 1502, 1503, and 1504 can provide different functionality from each other. As an example, a resource in an environment can be a server.

[0247] In accordance with the techniques described herein, placing systems and resources into separate environments or subsystems may enable application isolation for security and / or performance. Isolating environments may also mitigate the impact of compromised resources. For example, one environment may contain sensitive data and be configured with reduced exposure to the Internet, while another environment may host Internet-facing applications.

[0248] FIG. 15B illustrates an exemplary process flow for the controller shown in FIG. 15A to set up an environment. In such an example, the system may be tasked with creating and setting up a new environment. This may be triggered by a user request or by a system rule that executes when processing a particular task or series of tasks. FIGS. 17A-18B, described below, illustrate examples of specific change management tasks or series of tasks in which the system creates a new environment. However, there may be many situations in which the controller may create and set up a new environment.

[0249] Thus, referring to FIG. 15B, when setting up a new environment, the controller selects environment rules (step 1500.1). According to the environment rules, and using the global system rules 210 and templates 230, the controller finds resources for the environment (step 1500.2). The rules may have a hierarchy of suitable resource selections that the controller goes through until it finds the resources needed for the environment. In step 1500.3, the controller assigns the resources found in step 1500.2 to the environment, for example, using the techniques described in FIG. 3C or 7B. The controller then configures the system's networking resources for the new environment to ensure compatible and efficient connectivity between the new environment and other system components (step 1500.4). The system state is updated in step 1500.5 as each resource is enabled and each template is processed. The controller then sets up and enables the integration and interoperability of the environment's resources and powers on any applications to deploy the new environment (step 1500.6). The system state is updated again in step 1500.7 once the environment becomes available.

[0250] FIG. 15C illustrates an exemplary process flow for the controller shown in FIG. 15A to set up multiple environments. When setting up multiple environments, the environments can be set up in parallel using the technique described in FIG. 15B for each environment. However, it should be understood that the environments can be set up in sequential order or consecutively, as described in FIG. 15C. Referring to FIG. 15C, in step 1500.10, the controller sets up and deploys a first new environment (which can be performed as described with respect to step 1500.1 of FIG. 15B). Different environment rules may exist for different types of environments and how different environments interoperate. In step 1500.11, the controller selects an environment rule for the next environment. In step 1500.12, the controller finds resources according to a priority that may be defined by system rules 210. In step 1500.13, the controller allocates the resources found in step 1500.12 to the next environment. The environments may or may not share resources. In step 1500.14, the controller uses system rules 210 to configure the system's networking resources for the next environment and between environments with dependencies. The system state is updated in step 1500.15 as each resource is enabled, templates are processed, and networking resources are configured, including environment dependencies. The controller then sets up and enables the integration and interoperability of resources for the next environment and between environments, and powers on any applications to deploy the new environment (step 1500.16). The system state is updated in step 1500.17 when the next environment becomes available.

[0251] One-way communication to support monitoring 16A illustrates an exemplary embodiment in which a first controller 1601 operates as a main controller for setting up one or more controllers, such as 1601a, 1601b, and / or 1601c. Main controller 1601 may be used to create multiple cloud hosts, systems, and / or applications as environments 1602, 1603, and 1604, which may or may not be dependent on each other in operation, using the techniques described above with respect to controllers, such as controllers 200 / 1401 / 1501. As illustrated in FIG. 16A, IT systems, environments, clouds, and / or any combination(s) thereof may be created as environments 1602, 1603, and 1604. Environment 1602 includes second controller 1601a, environment 1603 includes third controller 1601b, and environment 1604 includes fourth controller 1601c. The environments 1602, 1603, 1604 may each include one or more resources 1642, 1643, 1644, respectively. The resources may include one or more applications 1642, 1643, 1644 that may run on them. These applications may connect to assigned resources, whether shared or not. These or other applications may run on one or more shared resources over the Internet or in a pool 1660, which may include a shared application or application network. The applications may provide services to one or more users or environments or clouds. The environments 1602, 1603, 1604 may share resources or databases and / or include or use resources in the pool 1660 that are specifically assigned to the particular environment. Various components of the system, including the main controller 1601 and / or one or more environments, may be connectable to an external network 1615, such as an application network or the Internet.

[0252] Between any resource, environment, or controller and another resource, environment, controller, or external connection, there may be a connection that can be configured to be selectively enabled and / or disabled in the manner described with respect to Figures 13A-13E herein. For example, any resource, controller, environment, or external connection can be disabled or disconnected from controller 1601, environment 1602, environment 1603, and / or environment 1604, resource, or application via in-band management connection 270, out-of-band management connection 270, SAN connection 280, or by physically disconnecting. As one example, in-band management connection 270 between controller 1601 and any of environments 1602, 1603, and 1604 can be disabled to protect controller 1601. As another example, such in-band management connection(s) 270 may be selectively disabled or enabled during operation of environments 1602, 1603, and 1604. 13A-13E herein, disabling or disconnecting the main controller 1601 from the environments 1602, 1603, 1604 may allow the main controller 1601 to spin the environments 1602, 1603, 1604 as clouds that can then be separated from the main controller 1601 or other clouds or environments. In this sense, the controller 1601 is configured to create multiple clouds, hosts, or systems.

[0253] Using the disabling or disconnection elements described herein, users can be granted limited access to environments through the main controller 1601 for specific uses. For example, a developer can be provided access to a development environment. As another example, an application administrator can be limited to a specific application or application network. As another example, logs can be viewed through the main controller 1601, collecting data without compromising the environment or controllers it generates.

[0254] After the main controller 1601 sets up the environment 1602, the environment 1602 is disconnected from the main controller 1601, at which point the environment 1602 can operate independently of the main controller 1601 and / or can be selectively monitored and maintained by the main controller 1601 or other applications associated with or executed by the environment 1602.

[0255] An environment, such as environment 1602, may be coupled to a user interface or console 1640 that allows a purchaser or user to access environment 1602. Environment 1602 may host the user console as an application. Environment 1602 may be accessed remotely by a user. Each environment 1602, 1603, 1604 may be accessed by a common or separate user interface or console.

[0256] 16B shows an exemplary system in which environments 1602, 1603, and 1604 may be configured to write to another environment 1641, where logs may be viewed using, for example, a console (which may be any console that can connect, either directly or indirectly, to environment 1641). In this manner, environment 1641 may function as a log server to which one or more of environments 1602, 1603, and 1604 write events. Main controller 1601 may then access log server 1641 to monitor events on environments 1602, 1603, and 1604 without maintaining a direct connection with such environments 1602, 1603, and 1604, as described below. Environment 1641 may further be configured to be selectively disconnected from main controller 1601 and read-only from the other environments 1602, 1603, and 1604.

[0257] The main controller 1601 can be configured to monitor some or all of its environments 1602, 1603, and 1604 even when the main controller 1601 is disconnected from any of those environments 1602, 1603, and 1604, as shown in FIG. 16C. FIG. 16C shows that the in-band management connection 270 between the main controller 1601 and the environments 1602, 1603, and 1604 is disconnected, which can help protect the main controller 1601 if the environments 1602, 1603, and 1604 are compromised. As shown in FIG. 16C, the out-of-band connection 260 can continue to be maintained between the main controller 1601 and an environment, such as 1602, even when the in-band connection 270 between the main controller 1601 and the environment 1602 is disconnected. Additionally, the environment 1641 can have a connection to the main controller 1601 that can be selectively enabled or disabled. The main controller 1601 can set up monitoring as a separate application in an environment 1641 that is isolated or air-gapped from the environments 1602, 1603, and 1604. The main controller 1601 can use one-way communication for monitoring. For example, logs may be provided through one-way communication from the environments 1602, 1603, and 1604 to the environment 1641. Through such one-way writes and through a connection between the environment 1641 and the main controller 1601, the main controller 1601 can collect data via the environment 1641 to monitor the environments 1602, 1603, and 1604 even when there is no in-band connection 270 between the main controller 1601 and the environments 1602, 1603, and 1604, thereby mitigating the risk that the environments 1602, 1603, and 1604 will compromise the main controller 1601. Access may be filtered or controlled and / or access may be independent of the Internet.16D , when the in-band connection 270 between the main controller 1601 and the environment 1602 is connected, the main controller 1601 can control the network switch 1650 to disconnect the environment 1602 from an external network 1615 such as the Internet. Disconnecting the environment 1602 from the external network 1615 when the environment 1602 is connected to the main controller 1601 by the in-band connection 270 can improve the security of the main controller 1601.

[0258] 16B-16D illustrate how the main controller can securely monitor environments 1602, 1603, and 1604 while minimizing exposure to those environments. Thus, the main controller 1601 can disconnect itself from environments 1602, 1603, and 1604 (or at least disconnect itself from the in-band link) while continuing to maintain a mechanism for monitoring those environments via the log server of environment 1641, to which environments 1602, 1603, and 1604 may have one-way write access. Thus, if, in the course of reviewing the logs of environment 1641, the main controller 1601 discovers that environment 1602 may be compromised by malware, the main controller 1601 can use SDN tools to isolate environment 1602 so that only out-of-band connection 260 exists (e.g., see FIG. 16C). Additionally, the controller 1601 can send a notification of the potential problem to an administrator of the environment 1602. The controller can further isolate the compromised environment 1602 by selectively disabling any connections (e.g., in-band management connections 270) between the compromised environment and any of the other environments 1603, 1604. In another example, the main controller 1601 can discover through logs that a resource in the environment 1603 is running too hot. This can cause the main controller to intervene and migrate applications or services from the environment 1603 to another environment (whether an existing environment or a newly created environment).

[0259] The controller 1601 may also set up one or more similar systems according to the requirements of a purchaser or user. As shown in FIG. 16E , a purchasing application 1650 may be provided, for example, on a console or otherwise, to enable a purchaser to purchase or request a cloud, host, system environment, or application to be set up for the purchaser. The purchasing application 1650 may instruct the controller 1601 to set up an environment 1602. The environment 1602 may include a controller 1601a that deploys or configures an IT system, for example, by allocating or assigning resources to the environment 1602.

[0260] 16F illustrates user interfaces 1632, 1633, and 1634 that may be used in environments where environments 1602, 1603, and 1604, respectively, operate as clouds and may or may not include a controller. User interfaces 1632, 1633, and 1634 (corresponding to environments 1602, 1603, and 1604, respectively) may each be connected through a main controller 1601 that manages the connection between the user interfaces and the environments. Alternatively, or additionally, interface 1640a (which may take the form of a console) may be directly coupled to environment 1602, interface 1640b (which may take the form of a console) may be directly coupled to environment 1603, and interface 1640c (which may take the form of a console) may be directly coupled to environment 1604. A user can use one or more of the interfaces to use an environment or a cloud regardless of whether the connection to main controller 1601 is isolated, disconnected, or disabled.

[0261] System cloning and backup for change management support Some of the environments 1602, 1603, 1604 can be clones of typical setup software used by developers, or they can be clones of current working environments as a way to scale, for example, to clone an environment in another data center in a different location to reduce latency due to location.

[0262] It should be appreciated that a main controller that sets up systems and resources into separate environments or subsystems may therefore enable cloning or backup of portions of an IT system, which may be used in testing and change management, as described herein. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, and / or other changes.

[0263] According to an example embodiment, an IT system or controller described herein can be configured to clone one or more environments. The new or cloned environment may or may not have the same resources as the original environment. For example, it may be desirable or necessary to use an entirely different combination of physical and / or virtual resources in the new or near-cloned environment. It may be desirable to clone an environment to a different location or time period where utilization optimization can be managed. It may be desirable to clone an environment to a virtual environment. When cloning an environment, the global system rules 210 and global templates 230 of the controller or main controller can include information on how to configure and / or run various types of hardware. Configuration rules within the system rules 210 can direct resource placement and usage to more optimally optimize resources and applications given the particular available resources.

[0264] The main controller structure provides the ability to set up systems and resources into separate environments or subsystems, provide structure for clone environments, provide structure for creating development environments, and / or provide structure for deploying a standardized set of applications and / or resources. Such applications or resources may include, but are not limited to, those that can be used to develop and / or run applications, or back up or restore portions of IT systems and other disaster recovery applications (e.g., a LAMP (Apache, MySQL, PHP) stack, a system including a web front end and servers running React / Redux, as well as resources running Node.js and a Mongo database and other standardized "stacks"). In some cases, the main controller may deploy an environment that is a clone of another environment, deriving configuration rules from a subset of the configuration rules used to create the original environment.

[0265] According to an exemplary embodiment, change management of a system or a subset of a system can be performed by cloning one or more environments and the configuration rules or a subset of configuration rules of such environments. For example, changes may be required to make code, configuration rules, security patches, changes to templates, hardware changes, adding / removing components and dependent applications, and other changes.

[0266] According to exemplary embodiments, such changes to the system may be automated to avoid errors in direct manual entry of changes. Changes may be tested by a user in a development environment before automatically implementing the changes in a live system. According to exemplary embodiments, a controller may be used to clone a live production environment by automatically powering on, provisioning, and / or configuring an environment that is configured using the same configuration rules as the production environment. The cloned environment may be run and operational (while a backup environment may be left in place, preferably for emergency use, in case changes need to be rolled back). This may be done using the controller to create, configure, and / or provision a new system or environment, such as those described with reference to Figures 1-16F above, using system rules 210, templates 230, and / or system states 220. The new environment may be used as a development environment to test changes that will later be implemented in the production environment. The controller may generate the infrastructure for such an environment from a software-defined structure in the development environment.

[0267] A production environment, as defined herein, means an environment that is used to run a system, as opposed to an environment that is dedicated to development and testing, i.e., a development environment.

[0268] When a production environment is cloned, the infrastructure or cloned development environment is configured and generated by the controller according to the global system rules 210, just like the production environment. Changes to the development environment can be made to code, templates 230 (either changes related to modifying existing templates or creating new templates), security, and / or application or infrastructure configuration. Once new changes implemented in the development environment are prepared as needed through development and / or testing, the system automatically applies the changes to the development environment, which then goes live or is deployed as the production environment. The new system rules 210 are then uploaded to either the environment's controller and / or the main controller, which applies the system rule changes to the specific environment. The system state 220 can be updated within the controller to implement added or modified templates 230. Thus, a complete system knowledge of the infrastructure, along with the ability to recreate it, can be maintained by the development environment and / or the main controller. As used herein, complete system knowledge can include, but is not limited to, system knowledge of resource state, resource availability, and system configuration. Complete system knowledge may be gathered by the controller from system rules 210, system state 220, and / or by querying resources using in-band management connection(s) 270, out-of-band management connection(s) 260, and / or SAN connection(s) 280. Resources may be queried to determine, among other things, resource, network or application utilization, configuration status or availability.

[0269] The cloned infrastructure or environment may be software defined via system rules 210, but this is not required. The cloned infrastructure or environment may or may not generally include a front end or user interface and one or more allocated resources, which may or may not include compute, network, storage, and / or application networking resources. The environment may or may not be configured as a front end, middleware, and database. A service or development environment may be bootstrapping with system rules 210 of a production environment. The infrastructure or environment allocated for use by the controller may be software defined specifically for the clone. Thus, the environment is deployable via system rules 210 and cloneable by similar means. The cloned or development environment may be automatically set up by a local or main controller using system rules 210 before or when a change is desired.

[0270] The production environment data may be written to read-only data storage until the development environment is separated from the production environment, after which it will be used by the development environment in the development and testing process.

[0271] A user or client can make and test changes in the development environment while the production environment is online. Data in data storage may be changed during development, and the changes are tested in the development environment. In volatile or writable systems, a hot synchronization with the production environment data may also be used after the development environment is set up or deployed. Desired changes to the system, application, and / or environment can be made and tested in the development environment. Desired changes are then made to the system rules 210 scripts, creating a new version for the entire environment or system and the main controller.

[0272] According to another exemplary embodiment, the newly developed environment may then be automatically implemented as the new production environment while the previous production environment remains maintained or fully functional, thus allowing for a reversion to the previous state of the production environment without significant data loss. The development environment is then booted with the new configuration rules in system rules 210, and the database is synchronized with the production database and switched to a writable database. The original production database can then be switched to a read-only database. If it is desired to revert to the previous production environment, the previous production environment is maintained as a copy of the previous production environment for as long as necessary.

[0273] An environment can be configured as a single server or instance, which can include physical and / or virtual hosts, networks, and other resources. In another exemplary embodiment, an environment can be multiple servers, including physical and / or virtual hosts, networks, and other resources. For example, there may be multiple servers forming a load-balanced Internet-facing application, which may be connected to multiple API / middleware applications (which may be hosted on one or more servers). The environment's database can include one or more databases to which APIs communicate queries within the environment. Environments can be constructed from system rules 210 in static or volatile form. Environments or instances can be virtual, physical, or a combination of each.

[0274] The configuration rules for an application or a system in system rules 210 can specify different compute backends (e.g., bare metal, AMD epyc server, Intel Haswell on qemu / kvm) and can include rules on how to run an application or service on a new compute backend. Thus, for example, an application can be virtualized if there is a situation where resources for testing are less available.

[0275] Using examples described herein, according to which a test environment may be deployed on virtual resources where the original environment uses physical resources, using the controller described herein with reference to Figures 1-18B, and as further described herein, a system or environment may be cloned from a physical environment to an environment that may or may not include virtual resources in whole or in part.

[0276] 17A shows an exemplary embodiment in which system 100 comprises controller 1701 and one or more environments, e.g., 1702, 1703, 1704. System 100 may be a static system, i.e., a system in which active user data does not constantly change the state of the system or frequently manipulate data, e.g., a system that hosts only static web pages. The system may be coupled to a user (or application) interface 110.

[0277] The controller 1701 can be configured in a manner similar to the controllers 200 / 1401 / 1501 / 1601 described herein and can similarly include global system rules 210, controller logic 205, templates 230, and system state elements 220. The controller 1701 can be coupled to one or more other controllers or environments in a manner similar to that described with reference to Figures 14A-16F herein. The global rules 210 of the controller 1701 can include rules that can manage and control other controllers and / or environments. Using such global rules 210, controller logic 205, system state 220, and templates 230, a system or environment can be set up, provisioned, and deployed through the controller 1701 in a manner similar to that described with reference to Figures 1-16F herein. Each environment can be configured with a subset of the global system rules 210 that define the behavior of the containing environment with respect to other environments.

[0278] The global system rules 210 may also comprise change management rules 1711. The change management rules 1711 comprise a set of rules and / or instructions that may be used when changes to the system 100, the global system rules 210, and / or the controller logic 205 may be desired. The change management rules 1711 may be configured to allow a user or developer to develop changes, test the changes in a test environment, and then implement the changes by automatically converting the changes into a new set of configuration rules in the system rules 210. The change management rules 1711 may be a subset of the global system rules 210 (as shown in FIG. 17A ) or may be separate from the global system rules 210. The change management rules may use a subset of the global system rules 210. For example, the global system rules 210 may comprise a subset of environment creation rules configured to create a new environment. The change management rules 1711 may be configured to set up and use a system or environment configured and set up by the controller 1701 to copy and clone some or all aspects of the system 100. Change control rules 1711 can be configured to allow testing of proposed new changes to a system before implementation by using a clone of the system for testing and implementation.

[0279] A clone 1705, as shown in FIG. 17A , can comprise the rules, logic, applications, and / or resources of a particular environment or portion of system 100. The clone 1705 can comprise similar or different hardware than system 100 and may or may not use virtual resources. The clone 1705 can be set up as an application. The clone 1705 can be set up and configured using configuration rules in system rules 210 of system 100 or controller 1701. The clone 1705 may or may not comprise a controller. The clone 1705 can have allocated network, computing resources, application network, and / or data storage resources, as described in more detail above. Such resources can be allocated using change management rules 1711 controlled by controller 1701. The clone 1705 can be coupled to a user interface that allows a user to make changes to the clone 1705. The user interface can be the same as or different from user interface 110 of system 100. A clone 1705 may be used for the entire system 100 or for a portion of the system 100, such as one or more environments and / or controllers. The clone 1705 may or may not be a complete copy of the system 100. The clone 1705 may be coupled to the system 100 via an in-band management connection 270, an out-of-band management connection 260, and / or a SAN connection 280, which may be selectively enabled and / or disabled and / or converted to a unidirectional read and / or write connection. Thus, when the clone environment 1705 is separated from the production environment during testing, or until the clone environment 1705 is ready to be brought online as the new production environment, the connection to the data in the clone environment 1705 can be changed to make the clone data read-only. For example, if the clone 1705 has a data connection to environment 1702, this data connection can be made read-only for the separation.

[0280] Optional backup 1706 may or may not be used for the entire system or for portions of the system, such as one or more environments and / or controllers. Backup 1706 may comprise network, computing, application network, and / or data storage resources, as described in more detail above. Backup 1706 may or may not comprise a controller. Backup 1706 may be a complete copy of system 100. Backup 1706 may be set up as an application or using similar or different hardware than system 100. Backup 1706 may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which may be selectively enabled and / or disabled entirely and / or converted to a unidirectional read and / or write connection.

[0281] FIG. 17B shows an exemplary process flow for using the clone and backup system of FIG. 17A in system change management. In step 1785, a user or management application initiates a change to the system. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, hardware changes, adding / removing components and / or dependent applications, and other changes. In step 1786, controller 1701 sets up the environment in the manner described with respect to FIGS. 14A-16F to become clone environment 1705 (where the clone environment may have its own new controller or may use the same controller as the original environment).

[0282] In step 1787, controller 1701 can use global rules 210, including change management rules 1711, to clone all or part of one or more environments of the system (e.g., a "production environment") to clone environment 1705 (e.g., where clone environment 1705 can function as a "development environment"). Accordingly, controller 1701 identifies and allocates resources and uses system rules 210 to set up and allocate clone resources and copy any of the data, configuration, code, executables, and other information required to launch the application from the environment to the clone. In step 1788, controller 1701 optionally backs up the system by setting up another environment (with or without a controller) to serve as backup 1706 using configuration rules in system rules 210 and copies templates 230, controller logic 205, and global rules 210.

[0283] After the clone 1705 is created from the production environment, the clone 1705 can be used as a development environment where changes to the clone's code, configuration rules, security patches, templates, and other modifications can be made. In step 1789, changes to the development environment can be tested before implementation. During testing, the clone 1706 can be isolated from the production environment (system 100) or other components of the system. This can be done by having the controller 1701 selectively disable one or more of the connections between the system 100 and the clone 1706 (e.g., by disabling the in-band management connection 270 and / or disabling the application network connection). In step 1790, it is determined whether the modified development environment is ready. If in step 1709 it is determined that the development environment is not yet ready (a determination typically made by a developer), the process flow returns to step 1789 for further modifications to the clone environment 1705. If in step 1790 it is determined that the development environment is ready, the development environment can be switched between the production environment and the production environment in step 1791. That is, the controller can change the development environment 1705 to a new production environment and maintain the previous production environment until the transition to the development / new production environment is complete and in good condition.

[0284] Figure 18A shows another example embodiment of a system 100 that can be set up and used in system change management. In the example of Figure 18A, the system 100 includes a controller 1801 and one or more environments 1802, 1803, 1804, 1805. The system is shown with a clone environment 1807 and a backup system 1808.

[0285] The controller 1801 may be configured in a manner similar to the controllers 200 / 1401 / 1501 / 1601 / 1701 described herein and may include elements of global system rules 210, controller logic 205, templates 230, and system states 220. The controller 1801 may be coupled to one or more other controllers or environments in a manner similar to that described with reference to Figures 14A-16F herein. The global rules 210 of the controller 1801 may include rules that can manage and control other controllers and / or environments. Using such global rules 210, controller logic 205, system states 220, and templates 230, a system or environment may be set up, provisioned, and deployed through the controller 1801 in a manner similar to that described with reference to Figures 1-17B herein. Each environment may be configured with a subset of the global rules 210 that defines the behavior of the environment, including behavior with respect to other environments.

[0286] The global rules 210 may also comprise change management rules 1811. The change management rules 1811 may comprise a set of rules and / or instructions that may be used when changes to the system, global rules, and / or logic are desired. The change management rules may be configured to allow a user or developer to develop changes, test the changes in a test environment, and then implement the changes by automatically converting the changes into a new set of configuration rules in the system rules 210. The change management rules 1711 may be a subset of the global system rules 210 (as shown in FIG. 18A ) or may be separate from the global system rules 210. The change management rules 1711 may use a subset of the global system rules 210. For example, the global system rules 210 may comprise a subset of environment creation rules configured to create a new environment. The change management rules 1811 may be configured to set up and use a system or environment set up and deployed by the controller 1801 to copy and clone some or all aspects of the system 100. Change control rules 1811 can be configured to allow testing of proposed new changes to a system before implementation by using a clone of the system for testing and implementation.

[0287] 18A , clone environment 1807 may include controller 1807a with rules, controller logic, templates, system state data, and assigned resources 1820 that may be assigned to one or more environments and set up according to global system rules 210 and change management rules 1811 of controller 1801. Backup system 1808 may further include controller 1808a with rules, controller logic, templates, system state data, and assigned resources 1821 that may be assigned to one or more environments and set up according to global system rules 210 and change management rules 1811 of controller 1801. The system may be coupled to user (or application) interface 110 or another user interface.

[0288] The clone environment 1807 may include rules, logic, templates, system state, applications, and / or resources for a particular environment or part of a system. The clone 1807 may include similar or different hardware than the system 100, and the clone 1807 may or may not use virtual resources. The clone 1807 may be set up as an application. The clone 1807 may be set up and configured using configuration rules in the system rules 210 of the system 100 or the controller 1801 for the environment. The clone 1807 may or may not include a controller, and may share a controller with the production environment. The clone 1807 may include allocated network, computing resources, application network, and / or data storage resources, as described in more detail above. Such resources may be allocated using change management rules 1811 controlled by the controller 1801. The clone 1807 may be coupled to a user interface that allows a user to make changes to the clone 1807. The user interface may be the same as or different from the user interface 110 of the system 100.

[0289] Clone 1807 may be used for an entire system or for a portion of a system, such as one or more environments and / or controllers. In an exemplary embodiment, clone 1807 may include hot standby data resources 1820a coupled to data resources 1820 of environment 1802. Hot standby data resources 1820a may be used during setup of clone 1807 and during testing of changes. Hot standby data resources 1820a may be selectively disconnectable or isolated from storage resources 1820 during change management, for example, as described herein with respect to FIG. 18B. Clone 1807 may or may not be a complete copy of system 100. Clone 1807 may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which may be selectively enabled and / or disabled and / or converted to a unidirectional read and / or write connection. Thus, when the cloned environment 1807 is separated from the production environment during testing, or until the cloned environment is ready to be brought online as the new production environment, connections to volatile data in the cloned environment 1807 can be changed to make the cloned data read-only.

[0290] When switching from an old production environment to a new one, the controller 1801 can instruct front-ends, load balancers, or other applications or resources to point to the new production environment. Thus, users, application resources, and / or other connections can be redirected when the change occurs. This can be done, for example, by methods including, but not limited to, changing the list of IP / IPoib addresses, Infiniband GUIDs, DNS servers, Infiniband partitions / OpenSM configurations, or software-defined network (SDN) configuration changes, which can be performed by sending instructions to networking resources. The front-ends, load balancers, or other applications and / or resources can point to systems, environments, and / or other applications, including, but not limited to, databases, middleware, and / or other back-ends. Such load balancers can be used for change management to switch from the old production environment to the new environment.

[0291] Clone 1807 and backup 1808 can be set up and used to manage changes to the system. Such changes can include, but are not limited to, changes to code, configuration rules, security patches, templates, hardware changes, adding / removing components and / or dependent applications, and other changes. Backup 1808 can be used for the entire system or for one or more environments and / or portions of the system, such as controller 1801. Backup 1808 can include the network, computing resources, application network, and / or data storage resources, as described in more detail above. Backup 1808 may or may not include the controller. Backup 1808 can be a complete copy of system 100. Backup 1808 can include the data necessary to reconstruct the system / environment / application from the configuration rules contained in the backup and can include all application data. Backup 1808 can be set up as an application or using similar or different hardware than system 100. Backup 1808 may be coupled to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which may be selectively enabled and / or disabled and / or converted to a unidirectional read and / or write connection.

[0292] Figure 18B is an exemplary process flow illustrating the use of the system of Figure 18A in change management, particularly when the system of Figure 18A includes volatile data or when the database is writable. Such a database may be part of the storage resources used by an environment within the system. In step 1870, the system is deployed using global system rules (including a production environment).

[0293] Next, in step 1871, the production environment is cloned using global system rules 210, including change control rules 1811, and resource allocation by the main controller 1801 or a controller in the cloned environment to create a read-only environment where the cloned environment is prohibited from writing to the system. The cloned environment can then be used as a development environment.

[0294] In step 1872, hot standby 1820a is enabled and allocated to clone environment 1807 to store any volatile data that has been changed in system 100. The clone data is updated and the new version in the development environment can be tested with the updated data. Hot sync data can be turned off at any time. For example, hot sync data can be turned off when writes from the old or production environment to the development environment are being tested.

[0295] Next, in step 1873, the user can make changes using the cloned environment 1807 as a development environment. Next, in step 1874, the changes to the development environment are tested. In step 1875, it is determined whether the modified development environment is ready (typically, such a determination is made by the developer). If in step 1875 it is determined that the changes are not ready, then process flow can return to step 1873 so that the user can go back and make other changes to the development environment. If in step 1875 it is determined that the changes are ready to take effect, then process flow proceeds to step 1876, where configuration rules are updated in the system or controller for the particular environment and will be used to deploy the new, updated environment.

[0296] In step 1877, the development environment (or new environment) may then be redeployed with the changes in the desired final configuration with the desired resource and hardware allocation before going live. In a next step 1878, write capabilities of the original production environment are disabled, and the original production environment becomes read-only. While the original production environment is read-only, as part of 1878, any new data from the original production environment (or possibly also the new production environment) may be cached and identified as migrated data. As an example, the data may be cached in a database server or other suitable location (e.g., a shared environment). The development environment (or new environment) and the old production environment are then switched in step 1879, with the development environment (or new environment) becoming the production environment.

[0297] After this switchover, the new production environment is made writable in step 1880. Once the new production environment is deemed to be functioning as determined by the developer in step 1881, any data lost during the switchover process (such data was cached in step 1878) may be written to the new environment and verified in step 1884. After such verification, the change is complete (step 1885).

[0298] If step 1881 results in a determination that the new production environment is not working (e.g., an issue is identified that requires the system to be reverted to the old system), then in step 1882 the environment is reverted and the old production environment becomes the production environment again. As part of step 182, the configuration rules for the target environment in controller 1801 are reverted to the previous version that was used for the production environment that is now being reverted.

[0299] In step 1883, database changes may be determined, for example, using cached data, and the data is restored to the old production environment with the old configuration rules. To support step 1883, the database may maintain a log of changes made to the database so that step 1883 can determine changes that may need to be invalidated. A backup database that tracks and times the cached data may be used to cache the data as described above, and the clock can be turned back to determine what changes were made. Snapshots and logs may be used for this purpose.

[0300] After restoring the cache data at 1883, if it is desired to start again, the process may return to step 1871.

[0301] The example change management systems described herein may be used, for example, when updating, adding, or removing hardware or software, patching software, detecting a system failure, migrating hosts during hardware failure or detection, for dynamic resource migration, modifying configuration rules or templates, and / or making any other system-related changes. The controller 1801 or system 100 may be configured to detect failures, and upon detection, the system may automatically enforce change management rules or existing configuration rules on other hardware available to the controller. Examples of possible failure detection methods include, but are not limited to, pinging hosts, querying applications, and running various tests or test suites. The change management configuration rules described herein may be enforced upon detection of a failure. Such rules may trigger the automatic creation of a backup environment or the automatic migration of data or resources implemented by the controller upon detection of a failure. The selection of backup resources may be based on resource parameters. Such resource parameters may include, but are not limited to, usage information, speed, configuration rules, and data capacity and usage.

[0302] As described herein, whenever a change occurs, the controller creates a log of the change and what was actually performed. For security or system updates, the controller described herein may be configured to automatically turn on and off and update the IT system state according to configuration rules. The controller may turn off resources to conserve power. The controller may turn on or migrate resources for different efficiencies at different times. During migration, a backup or copy of the environment or system may be made according to configuration rules. In the event of a security breach, the controller may isolate and shut down the attacked area.

[0303] While the invention has been described with reference to exemplary embodiments, various modifications may be made that are within the scope of the invention. Such modifications to the invention are discernible in light of the teachings herein.

[0304] Appendix A: Example of Storage Attachment Process Described below are example processes and rules for sharing storage resources between multiple systems. It should be understood that this is only one example of a storage connection process, and other techniques for connecting computing resources to storage resources may be used. Unless otherwise stated, these rules apply to all systems attempting to initiate a storage connection. Definitions in this Appendix A Storage Resource: A block, file, or file system that can be shared via a storage transport. Storage Transport: How storage resources are shared locally or remotely. Examples are iSCSI / iSER, NVMEoF, NFS, and Samba file sharing. System: Anything that attempts to connect to a storage resource via a specified storage transport. A system may support any number of storage transports and may be able to determine for itself which transport to use. Read-only: A read-only storage resource does not allow modification of the data it contains. This restriction is imposed by the storage daemon that handles the export of the storage resource over the storage transport. For additional assurance, some datastores may mark the storage resources that back their data as read-only (for example, marking an LVM LV as read-only). Read-write (or volatile): A read-write (volatile) storage resource is one whose contents may be modified by the systems that connect to the storage resource. Rules: When a controller determines whether a system can connect to a given storage resource, there is a set of rules that must be followed. 1. A read / write storage resource shall be exported on only one storage transport. 2. Read / write storage resources shall only be connected by one system. 3. Read-write storage resources must not be connected as read-only. 4. A read-only storage resource may be exported over multiple storage transports. 5. A read-only storage resource may be connected to from multiple systems. 6. Read-only storage resources must not be connected as read-write. process When the connection process is considered as a function, it takes two independent variables. 1.Storage resource ID 2. List of supported storage transports (prioritized by order) First, determine whether the requested storage resource is read-only or read-write. If it is read-write, read-write storage resources are limited to one connection, so we must check if the storage resource is already connected. If it already has a connection, we verify that the system requesting the storage resource is the currently connected system (this may occur in the case of a reconnection, for example). If not, we raise an error, as multiple systems cannot connect to the same read-write storage resource. If the requesting system is the system connected to this storage resource, we verify that one of the available storage transports matches the current export of this storage resource. If there is a match, we pass the connection information to the requesting system. If there is no match, we raise an error, as read-write storage resources cannot be served by multiple storage transports. For read-only and unattached read-write storage resources, it iterates through the list of storage transports provided and attempts to export the storage resource using that transport. If the export fails, it continues to attempt the export down the list until it is successful or until it runs out of storage transports. If it runs out of storage transports, it notifies the requesting system that the storage resource could not be attached. If the export is successful, it stores the attachment information and the new (resource, transport) => (system) relationship in a database. The requesting system is then informed of the storage transport attachment information. System: Storage connectivity is currently handled by the controller and compute daemon during normal operation. However, future iterations may allow services to connect directly to storage resources, bypassing the compute daemon. This may be a requirement for the physical deployment example of a service, and it makes sense to use the same process for virtual machine deployment as well.

[0305] Appendix B: Example of connecting to OverlayFS Services use OverlayFS to reuse objects from a common file system and reduce service package size. The service in this example includes three or more storage resources. 1. Platform. This includes the base Linux file system and is accessed read-only. 2. Service. This includes all software directly related to the operation of the service (NetThunder ServiceDaemon, OpenRC scripts, binaries, etc.). This storage resource is accessed read-only. 3. Volatile. These storage resources contain all changes to the system and are managed by LVM from within the service (for physical, container, and virtual machine deployments). When running in a virtual machine, the service is Direct Kernel Booted in Qemu using a custom Linux kernel with an initramfs that contains logic to: 1. Assemble an LVM Volume Group (VG) from available read-write disks *This VG contains one Logical Volume (LV) that contains all volatile storage data for the service. 2. Mount the platform, service, and LV 3. Use a union file system (here, OverlayFS) to combine the three file systems. The same process can be used for physical deployment. One option is to remotely provide the kernel to a lightweight OS booted via PXEBoot or IPMI ISO Boot, and then use kexec to the new actual kernel. Alternatively, skip the lightweight OS and PXE boot directly into the kernel. Such systems may require additional logic in the kernel initramfs to connect to storage resources. An OverlayFS configuration could look like this: / ――――――――――――\ |Volatile Layer(LV)(RW)| +――――――――――――+ |Service Layer (RO)| +――――――――――――+ |Platform Layer (RO)| \―――――――――――― / Some restrictions in OverlayFS allow a special directory, ' / data', to be marked as "out-of-tree". This directory is made available to services by creating the ' / data' directory when creating a service package. This special directory is mounted via 'mount --rbind' to allow access to a subset of the volatile layer that is not in OverlayFS. This is necessary for applications such as NFS (Network File System), which do not support shared directories that are part of OverlayFS. Kernel file system layout: / +--platform / +--bin / +--... / +--service / +--data / [optional] +--bin / +--... +--volatile +--work / +--root / +--bin / +--data / [if present in / service / ] +--... +--new_root / +--... Create a / new_root directory and use that directory as the target for configuring OverlayFS. Once OverlayFS is configured, when you do an exec_root into / new_directory, the system will start successfully with all available resources.

Claims

1. 1. An automated IT management system comprising: a controller configured to provide automated management of an IT system including a plurality of resources and a plurality of applications or services; a plurality of templates containing information used to configure and deploy resources and applications or services to be loaded onto the resources; a system state configured to track which templates are being used to deploy the resources and the applications or services; Including, one or more of the plurality of templates having an association with a target application or service and serving as a recipe defining how the associated target application or service should be integrated into the IT system; the controller is configured to have system knowledge based on the system state; the system knowledge includes (1) the resource status, (2) resource availability for the IT system, and (3) system configuration for the IT system; the controller is configured to use (1) at least one of the one or more templates associated with the target application or service and (2) the system knowledge to automatically (i) deploy the target application or service to at least one of the resources and (ii) integrate the target application or service into the IT system. Automated IT management systems.

2. further comprising a plurality of system rules that specify which template should be used to deploy which of a plurality of resources, applications, or services; the controller is further configured to have the system knowledge based on system rules and the system state.

10. The automated IT management system of claim 1.

3. 3. The automated IT management system of claim 1, wherein the automated management includes automatic configuration of interactions between applications and / or services.

4. 4. The automated IT management system of claim 1, wherein the automated management includes automated configuration and deployment of an infrastructure, the infrastructure including multiple parts, and interoperability of the parts of the infrastructure is built into the automated configuration and deployment.

5. 5. The automated IT management system of claim 1, wherein the information includes a configuration script used to configure the associated resource, application or service.

6. 6. The automated IT management system of claim 1, wherein the information includes a boot file.

7. The automated IT management system of claim 6 , wherein the boot file is configured to boot an operating system.

8. The automated IT management system of claim 1 , wherein the information comprises a configuration file or a configuration file template.

9. The automated IT management system of claim 1 , wherein the information includes hardware setup information.

10. The automated IT management system of claim 1 , wherein the information includes computing backend setup information.

11. The automated IT management system of claim 1 , wherein the information includes a base image.

12. The automated IT management system of claim 11 , wherein the base image includes a base operating system file system.

13. 13. The automated IT management system of claim 1, wherein the information includes a small temporary file system with instructions on how to set up the template so that it can be booted.

14. The automated IT management system of claim 1 , wherein one or more of the plurality of templates includes a template BIOS setting.

15. 15. The automated IT management system of claim 14, wherein the BIOS settings are configured to set optional settings for running applications on a physical host.

16. 16. The automated IT management system of claim 1, further comprising an out-of-band management connection, the out-of-band management connection configured to be used to boot resources or applications.

17. 17. The automated IT management system of claim 1, wherein at least one of the templates includes alternative data, files or binaries for different hardware types that provide similar or identical functionality.

18. 18. The automated IT management system of claim 1, wherein the information includes a daemon or script, the daemon configured to run an API accessible by the controller to enable the controller to change the configuration of the host.

19. 19. The automated IT management system of claim 1, wherein at least one of the templates includes multiple kernels and one or more pre-boot file systems for different hardware and different configurations.

20. 20. The automated IT management system of claim 1, wherein the templates and images derived from the templates are used to create applications, to deploy applications or services, and / or to configure resources for system functionality that enables and / or facilitates the creation of applications.

21. 21. The automated IT management system of claim 1, wherein the controller is configured to deploy and configure bare metal components from the information.

22. 22. The automated IT management system of claim 21, wherein the information for at least one of the templates includes an image derived from the at least one template.

23. further comprising a pool of resources; the controller is configured to dynamically allocate the resources in the pool.

23. An automated IT management system according to any one of claims 1 to 22.

24. 24. The automated IT management system of claim 23, wherein the controller is configured to recognize new nodes or hosts on a network and configure the new nodes or hosts to become part of the pool.

25. further comprising a plurality of system rules that define how resources from the pool should be deployed or allocated; The controller is configured to (1) receive requests for resources in the pool via an API, and (2) deploy or allocate the requested resources from the pool according to the system rules.

25. An automated IT management system according to claim 23 or 24.

Citation Information

Patent Citations

  • Cluster system and software deployment method

    JP2012048330A

  • System and method for automated hardware provisioning based on application characteristics

    JP2014524608A

  • Method and apparatus for provisioning a template-based platform and infrastructure

    JP2017505494A

  • Methods and systems for deploying applications to one or more cloud systems in a mobile manner.

    JP2017529633A

  • Configuring monitoring for virtualized servers

    US20160162312A1