Automatically deployed information technology (IT) systems and methods
The IT system with a core controller automates setup and management using self-assembly rules and templates, addressing setup, configuration, and security challenges, enhancing scalability and flexibility, and ensuring consistent configuration and secure patch testing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-25
AI Technical Summary
Current IT systems face challenges in setup, configuration, infrastructure deployment, asset tracking, security, application deployment, documentation for maintenance and compliance, maintenance, scaling, resource allocation, resource management, load balancing, software failures, software and security updates/patching, testing, IT system recovery, and change management, particularly due to complexity, scalability issues, and vulnerabilities in bare metal cloud nodes and virtual resources.
An IT system with a core controller that automates setup, configuration, and management using self-assembly rules, templates, and system states, enabling flexible, scalable, and secure deployment and management of physical and virtual resources, with features for interoperability, security enhancement, and automated documentation.
The system reduces human error, enhances system security, improves scalability and flexibility, ensures consistent configuration, and facilitates efficient resource use and reuse, while providing automated documentation and secure patch testing, thus addressing multiple IT system challenges.
Smart Images

Figure 2026053601000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This patent application claims the priority of U.S. Provisional Patent Application No. 62 / 596,355, filed on December 8, 2017, entitled "Automatically Deployed Information Technology (IT) System and Method", the entire disclosure of which is incorporated herein by reference.
[0002] This patent application also claims the priority of U.S. Provisional Patent Application No. 62 / 694,846, filed on July 6, 2018, entitled "Automatically Deployed Information Technology (IT) System and Method", the entire disclosure of which is incorporated herein by reference.
Background Art
[0003] The demands, utilization, and needs for computing have increased exponentially in the past few decades. Along with this, due to the demands for greater storage, speed, computing power, applications, and accessibility, the field of computing has changed rapidly, and tools have been provided to entities of various types and sizes. As a result, the use of public virtual computing systems and cloud computing systems has advanced, and more computing resources have been provided to a large number of users and user types. This exponential growth is expected to continue. At the same time, as the risks of failures and security increase, the setup, management, change management, and updates of the infrastructure become more complex and costly. Scalability, that is, growing the system over time, has also become a major issue in the field of information technology.
[0004] Most IT system problems, while many related to performance and security, can be difficult to diagnose and address. Time and resource constraints allowed for system setup, configuration, and deployment can lead to errors and future IT problems. Over time, many different administrators may be involved in changes, patches, or updates to IT systems, including users, applications, services, security, software, and hardware. Often, insufficient or lost documentation and history of configurations and changes make it difficult to later understand how a particular system is configured and functions. This can make future changes or troubleshooting difficult. When problems or failures occur, IT configurations and settings can be difficult to recover and reproduce. Furthermore, system administrators can easily make mistakes, such as incorrect commands or other errors, which can bring down computers, web databases, and services. Additionally, while an increased risk of security breaches is not uncommon, changes, updates, and patches aimed at avoiding breaches can cause unwanted downtime.
[0005] When critical infrastructure is in place, functioning, and operational, the costs or risks often seem to outweigh the benefits of changing the system. Problems associated with making changes to a live IT system or environment can cause significant, sometimes catastrophic, problems for users or entities that depend on those systems. At the very least, troubleshooting and resolving failures or problems that occur during change management can require considerable time, personnel, and financial resources. Technical problems that may arise when changes are made to a live environment can have cascading effects and may not be resolved simply by reverting the changes. Many of these issues contribute to the inability to quickly rebuild a system if failures occur during change management.
[0006] Furthermore, bare metal cloud nodes or resources within an IT system may be vulnerable to security issues or compromised or accessed by unauthorized users. Hackers, attackers, or unauthorized users may pivot from that node or resource to access or hack other parts of the IT system or networks connected to nodes. Bare metal cloud nodes or controllers within an IT system may also be vulnerable through resources connected to application networks that could expose the system to security threats or otherwise compromise the system. According to various exemplary embodiments disclosed herein, an IT system may be configured to improve the security of bare metal cloud nodes or resources from application networks, whether or not they interface with the Internet or have connectivity to external networks.
[0007] According to an exemplary embodiment, the IT system includes bare-metal cloud nodes or physical resources. If the bare-metal cloud nodes or physical resources may be connected to a network having nodes that may be used by other people or customers when they are powered on, set up, managed, or used, in-band management may be omitted, switchable, disconnectable, or filtered from the controller. Also, application networks or multi-application networks within the system may be disconnected, switchable, switchable, or filtered from the controller via the resource(s) to which the application network is coupled.
[0008] Physical resources, including virtual machines or hypervisors, can also be vulnerable to security issues and may be compromised or accessed by malicious users if the hypervisor can be used to pivot to other hypervisors that are shared resources. An attacker may be able to escape from a virtual machine and gain network access to management and / or supervisory systems via the controller. According to various exemplary embodiments disclosed herein, IT systems may be configured to improve security, and one or more physical resources, including virtual resources on a cloud platform, may be disconnected from the controller via an in-band management connection, may be detachable, filtered, may be filtered, or otherwise disconnected from the controller.
[0009] According to an exemplary embodiment, the physical resources of the IT system may include one or more virtual machines or hypervisors, and the in-band management connection between the controller and the physical resources may be omitted, disconnected, disconnectible, or filtered from those resources. [Brief explanation of the drawing]
[0010] [Figure 1] This is a schematic diagram of a system according to an exemplary embodiment. [Figure 2A] Figure 1 is a schematic diagram of an exemplary controller for the system shown. [Figure 2B] This illustrates the exemplary flow of operation for an exemplary set of storage expansion rules. [Figure 2C] Figure 2B shows an alternative example for performing steps 210.1 and 210.2. [Figure 2D] Figure 2B shows an alternative example for performing steps 210.1 and 210.2. [Figure 2E] An example template is shown. [Figure 2F] This shows an example processing flow of controller logic related to template processing. [Figure 2G] Figure 2F shows an illustrative processing flow for steps 205.11, 205.12, and 205.13. [Figure 2H] Figure 2F shows an illustrative processing flow for steps 205.11, 205.12, and 205.13. [Figure 2I] Other example templates are shown. [Figure 2J] This shows another exemplary processing flow of controller logic related to template processing. [Figure 2K] This shows an exemplary processing flow for managing service dependencies. [Figure 2L] This is a schematic diagram of an exemplary image derived from a template according to an exemplary embodiment. [Figure 2M] A set of exemplary system rules is shown. [Figure 2N] This shows an exemplary processing flow for how the controller logic processes the system rule in Figure 2M. [Figure 2O] This illustrates an exemplary processing flow for configuring storage resources from a filesystem blob or other group of files. [Figure 3A] This is a schematic diagram of the controller in Figure 2A with added computing resources. [Figure 3B] This is a schematic diagram of an exemplary image derived from a template according to an exemplary embodiment. [Figure 3C] This illustrates an exemplary process flow for adding resources such as computing resources, storage resources, and / or networking resources to the system. [Figure 4A] This is a schematic diagram of the controller in Figure 2A with added storage resources. [Figure 4B] This is a schematic diagram of an exemplary image derived from a template according to an exemplary embodiment. [Figure 5A] Figure 2A is a schematic diagram of the controller with JBOD and storage resources added. [Figure 5B] An exemplary process flow for adding storage resources and direct attached storage of storage resources to a system is shown. [Figure 6A] A schematic diagram of the controller of FIG. 2A with networking resources added. [Figure 6B] A schematic diagram of an exemplary image derived from a template according to an exemplary embodiment. [Figure 7A] A schematic diagram of a system according to an exemplary embodiment in an exemplary physical deployment. [Figure 7B] An exemplary process for adding resources to an IT system is shown. [Figure 7C] An exemplary process flow for deploying an application to multiple computing resources, multiple servers, multiple virtual machines, and / or multiple sites is shown. [Figure 7D] An exemplary process flow for deploying an application to multiple computing resources, multiple servers, multiple virtual machines, and / or multiple sites is shown. [Figure 8A] A schematic diagram of a system according to an exemplary embodiment in an exemplary deployment. [Figure 8B] An exemplary process flow for expanding from a single-node system to a multi-node system is shown. [Figure 8C] An exemplary process flow for migrating storage resources to a new physical storage resource is shown. [Figure 8D] An exemplary process flow for migrating virtual machines, containers, and / or processes on a single node of a multi-tenant system to a multi-node system that may have separate hardware for computing and storage is shown. [Figure 8E] Another exemplary process flow for expanding from a single node to multiple nodes within a system is shown. [Figure 9A] A schematic diagram of a system according to an exemplary embodiment in an exemplary physical deployment. [Figure 9B]This is a schematic diagram of an exemplary image derived from a template according to an exemplary embodiment. [Figure 9C] This shows an example of installing an application from an NT package. [Figure 9D] This is a schematic diagram of the system according to an exemplary embodiment in an exemplary deployment. [Figure 9E] This illustrates an exemplary process flow for adding a virtual computing resource host to an IT system. [Figure 10] This is a schematic diagram of the system according to an exemplary embodiment in an exemplary deployment. [Figure 11A] The system and method of an exemplary embodiment are shown. [Figure 11B] The system and method of an exemplary embodiment are shown. [Figure 12] The system and method of an exemplary embodiment are shown. [Figure 13A] This is a schematic diagram of a system according to an exemplary embodiment. [Figure 13B] This is another schematic diagram of the system according to an exemplary embodiment. [Figure 13C] This shows an exemplary processing flow of the system according to an exemplary embodiment. [Figure 13D] This shows an exemplary processing flow of the system according to an exemplary embodiment. [Figure 13E] This shows an exemplary processing flow of the system according to an exemplary embodiment. [Figure 14A] This illustrates an exemplary system where the main controller deploys controllers to different systems. [Figure 14B] This diagram shows an exemplary flow illustrating the possible steps for provisioning a controller that has a main controller. [Figure 14C] This diagram shows an exemplary flow illustrating the possible steps for provisioning a controller that has a main controller. [Figure 15A] This illustrates an exemplary system where the main controller generates the environment. [Figure 15B]This shows an exemplary processing flow in which the controller sets up the environment. [Figure 15C] This illustrates an exemplary processing flow in which a controller sets up multiple environments. [Figure 16A] This illustrates an exemplary embodiment in which a controller acts as a main controller for setting up one or more controllers. [Figure 16B] This illustrates an exemplary system in which one environment can be configured to write to another. [Figure 16C] This illustrates an exemplary system in which one environment can be configured to write to another. [Figure 16D] This illustrates an exemplary system in which one environment can be configured to write to another. [Figure 16E] This illustrates an exemplary system where users can purchase the new environment created by the controller. [Figure 16F] This example demonstrates a system that provides a user interface for interfacing with the environment created by the controller. [Figure 17A] Here is an example of a change management task related to a new environment. [Figure 17B] Here is an example of a change management task related to a new environment. [Figure 18A] Here is an example of a change management task related to a new environment. [Figure 18B] Here is an example of a change management task related to a new environment. [Modes for carrying out the invention]
[0011] In an effort to provide technical solutions to the aforementioned needs of the art, the inventors disclose various embodiments of the invention of systems and methods for information technology that provide automated setup, configuration, maintenance, testing, change management, and / or upgrade of IT systems. For example, the inventors disclose a controller configured to automatically manage a computer system based on a plurality of system rules, the system state of a computer system, and a plurality of templates. As another example, the inventors disclose a controller configured to automatically manage the physical infrastructure of a computer system based on a plurality of system rules, the system state of a computer system, and a plurality of templates. Examples of automated management that can be performed by the controller may include remote or local access to and modification of settings or other information on computers that can run an application or service, building an IT system, modifying an IT system, building individual stacks within an IT system, creating a service or application, loading a service or application, configuring a service or application, migrating a service or application, modifying a service or application, deleting a service or application, cloning a stack to another stack on a different network, creating, adding, deleting, setting up, configuring, reconfiguring, and / or modifying resources or system components, automatically adding, deleting, and / or restoring resources, services, applications, IT systems, and / or IT stacks, configuring interactions between applications, services, stacks, and / or other IT systems, and / or monitoring the health of IT system components. In exemplary embodiments, the controller may be embodied as a physical or virtual computing resource that can be remote or local.Examples of additional controllers that can be adopted include, but are not limited to, any or any combination of, processes, virtual machines, containers, remote computing resources, applications deployed by other controllers, and / or services. Controllers may be distributed across multiple nodes and / or resources, and may be located in other locations or networks.
[0012] IT infrastructure is, in most cases, built from individual hardware and software components. The hardware components used typically include servers, racks, power supplies, interconnects, display monitors, and other communication equipment. The methods and techniques for selecting and interconnecting these individual components are highly complex due to the vast number of optional configurations that operate at varying degrees of efficiency, cost-effectiveness, performance, and security. Hiring and training individual technicians / engineers skilled in connecting these infrastructure components is costly. Furthermore, the vast number of possible hardware and software iterations complicates maintenance and updates. This creates additional challenges if the individuals and / or engineering companies that originally installed the IT infrastructure are unable to perform the updates. Software components, such as operating systems, are either designed generically to work with a wide range of hardware or are highly specialized for specific components. In most cases, complex plans or blueprints are drawn up and executed. Changes, growth, scaling, and other challenges necessitate updating these complex plans.
[0013] Some IT users purchase cloud computing services from a growing supplier industry, but this does not solve the infrastructure setup problems and challenges; rather, it shifts them from the IT user to the cloud service provider. Furthermore, large cloud service providers address infrastructure setup challenges and problems in a way that can reduce flexibility, customizability, scalability, and the rapid adoption of new hardware and software technologies. In addition, cloud computing services do not provide ready-to-use bare-metal setups, configuration deployments, and updates, nor do they enable migration to, from, or between bare-metal and virtual IT infrastructure components. These and other limitations of cloud computing services can lead to many computing, storage, and networking inefficiencies. For example, speed or latency inefficiencies in computing and networking can occur by the cloud service or in the applications or services that utilize the cloud service.
[0014] The exemplary embodiment of the system and method provides for the deployment, use, and management of new and unique IT infrastructure. According to the exemplary embodiment, the complexity of resource selection, installation, interconnection, management, and updating is rooted in a core controller system and its parameter files, templates, rules, and IT system state. This system includes a set of self-assembly rules and behavioral rules configured so that components self-assemble, rather than requiring technicians to assemble, connect, and manage them. Furthermore, the exemplary embodiment of the system and method enables greater customizability, scalability, and flexibility using self-assembly rules without requiring typical external planning documents. They also enable efficient resource use and reuse.
[0015] Systems and methods are provided to improve many of the problems and issues of current IT systems, whether entirely or partially physical or virtual. The exemplary embodiments of systems and methods provide structures that enable flexibility, reduce variability and human error, and improve system security.
[0016] While several individual solutions may exist for one or more of the problems in current IT systems, such solutions do not comprehensively address the numerous problems that are resolved by the exemplary embodiments described herein. Furthermore, such existing solutions may address specific problems but exacerbate others.
[0017] Current challenges to be addressed include, but are not limited to, issues related to setup, configuration, infrastructure deployment, asset tracking, security, application deployment, service deployment, documentation for maintenance and compliance, maintenance, scaling, resource allocation, resource management, load balancing, software failures, software and security updates / patching, testing, IT system recovery, change management, and hardware updates.
[0018] The IT systems used herein may include, but are not limited to, servers, virtual and physical hosts, databases and database applications, IT services, business computing services, computer applications, customer service applications, web applications, mobile applications, backends, case number management, customer tracking, ticketing, business tools, desktop management tools, accounting, email, documentation, compliance, data storage, backup, and / or network management.
[0019] One of the challenges users may face before setting up an IT system is predicting infrastructure needs. Initially, or over time as it grows or changes, users may not know how much storage, computing power, or other requirements will be needed. According to the exemplary embodiment, the IT system and infrastructure allow for flexibility in that, if the system requires change, components can be automatically added, removed, or reallocated later within the infrastructure using the self-deployed infrastructure (both physical and / or virtual) of the exemplary embodiment. Thus, the challenge of predicting future needs presented at system setup is addressed by providing the ability to add to the system using global rules, templates, and system states, and by tracking changes to such rules, templates, and system states.
[0020] Other challenges may relate to correct configuration, configuration consistency, interoperability, and / or interdependencies, including, for example, future incompatibilities due to changes in configured system elements or their configurations over time. For example, when an IT system is initially set up, there may be missing elements or some elements may be poorly configured. Also, for example, when iterations of elements or infrastructure components are set up, there may be a lack of consistency between the iterations. When changes are made to the system, the configuration may need to be reviewed. A difficult choice is presented between the optimal configuration and flexibility to accommodate future infrastructure changes. According to an exemplary embodiment, when the system is initially deployed, the configuration is self-deployed from a template to infrastructure components using global system rules, so the configuration is consistent, reproducible or predictable, enabling the optimal configuration. Such an initial system deployment may be performed on physical components, but subsequent components may be added or modified, whether physical or not. Furthermore, such an initial system deployment may be performed on physical components, but subsequent environments may be cloned from the physical structure, whether physical or not. This makes it possible to optimize the system configuration while minimizing changes that may cause problems in the future.
[0021] In the deployment phase, there are typically challenges in the interoperability of bare-metal and / or software-defined infrastructure. There may also be challenges in the interoperability of software with other applications, tools, or infrastructure. These may include, but are not limited to, challenges arising from deployed products from different vendors. The inventors disclose an IT system that can provide infrastructure interoperability, regardless of whether it is bare-metal, virtual, or any combination thereof. Thus, interoperability, i.e., the ability of parts to work together, can be incorporated into the disclosed infrastructure deployment, where the infrastructure is automatically configured and deployed. For example, different applications may be interdependent, and they may reside on separate hosts. To enable such applications to interact with each other, the controller logic, templates, system states, and system rules described herein include information and configuration instructions used to configure and track application interdependencies. Therefore, the infrastructure features described herein provide a way to manage how each application or service interacts with each other. For example, to ensure that an email service communicates properly with an authentication service, and / or a groupware service communicates properly with an email service. Furthermore, such management can extend down to the infrastructure level, allowing for tracking, for example, how computing resources communicate with storage resources. Otherwise, the complexity of the IT system would be O(n). n ) may rise.
[0022] As disclosed, automated resource deployment does not require prior configuration of operating system software, thanks to the controller's ability to deploy based on global system rules, templates, and IT system state / system self-awareness. According to exemplary embodiments, users or IT professionals may not need to know whether resource additions, allocations, or reallocations are coordinated to ensure interoperability. Additional resources may be automatically added to the network according to exemplary embodiments.
[0023] Using an application typically requires many different resources, including computing, storage, and networking. Interoperability between resources and system components is also necessary, which includes knowledge of what is located and operating in place, and interoperability with other applications. An application may need to connect to other services to retrieve configuration files and ensure all components work together properly. Therefore, configuring an application can be time- and resource-intensive. If there are interoperability issues with other applications, the application's configuration can have cascading effects on the rest of the infrastructure, potentially leading to downtime or breaches. To address these issues, the inventors disclose automated application deployment. Thus, as disclosed by the inventors, an application can be self-deployed by intelligently configuring itself, using knowledge of the current state of the system, reading from IT system states, global system rules, and templates. Furthermore, according to exemplary embodiments, pre-deployment testing of the configuration can be performed using the change management functions described herein.
[0024] Other issues addressed by the exemplary embodiments relate to problems that may arise with intermediate configurations when it is desirable to switch to a different vendor or other tool. According to one aspect of the exemplary embodiments, template conversion is provided between controller rules and templates and application templates of a specific vendor. This allows the system to automatically change the vendor of software or other tools.
[0025] Many security problems arise from misconfigurations, patching failures, and the inability to test patching before deployment. Often, security issues can arise during the setup configuration phase. For example, misconfigurations can leave sensitive applications exposed to the internet or allow email forgery from mail servers. The inventors disclose a system setup that, through automatic configuration, protects against attackers, avoids unnecessary exposure to attackers, and provides security engineers and application security architects with more knowledge about the system. Automation reduces security vulnerabilities caused by human error or misconfigurations. The disclosed infrastructure can also provide introspection between services, enable rule-based access, and restrict communication between services to only what is actually needed. The inventors disclose a system and method that has the ability to securely test patches before deployment, as described, for example, with respect to change management.
[0026] Documentation is often a problematic area of IT management. During setup and configuration, the primary goal may typically be to get components working together. This usually involves troubleshooting and trial-and-error processes, and it can be difficult to know what actually made the system work. While the exact commands executed are usually documented, the troubleshooting or trial-and-error processes that may have resulted in a working system are often not adequately documented, or not documented at all. Problems or deficiencies in documentation can lead to problems with audit trails and audits. The documentation problems that arise can cause problems in demonstrating compliance. When building a system or its components, compliance issues are often not well known. Determining applicable compliance may only become clear after the IT system has been set up and configured. Thus, documentation is essential for auditing and compliance. The inventors disclose a system including a global system rules database, templates, and an IT system state database that provides automatically documented setup and configuration. Any configuration that occurs is recorded in the database. According to an exemplary embodiment, the automatically documented configuration can be used to provide an audit trail and demonstrate compliance. Inventory management can utilize information that is automatically documented and tracked.
[0027] Another challenge arising from the setup, configuration, and operation of IT systems concerns hardware and software inventory management. For example, it is typically important to know how many servers there are, whether they are up and running and still functioning, what their functions are, which rack each server is in, which power supply is connected to which server, which network card and which network port each server is using, which IT systems the components are running on, and many other important details. In addition to inventory information, passwords and other sensitive information used for inventory management must be effectively managed. Collecting and maintaining this information is a time-consuming task, especially in large IT systems, data centers, or data centers where equipment changes frequently, and this is often managed manually or using various software tools. Ensuring secure password compliance is a significant risk factor that can be a critical issue in guaranteeing a secure computing environment. The inventors disclose an IT system in which the collection and maintenance of inventory and operating status of all servers and other components is automatically updated, stored, and protected as part of the IT system state, global system rules, templates, and controller logic of the controllers.
[0028] In addition to issues relating to the setup and configuration of IT systems, the inventors disclose an IT system that can also address issues and problems arising in the maintenance of IT systems. For example, many problems arise as a data center continues to function with hardware failures, such as power failures, memory failures, network failures, network card failures, and / or CPU failures. Further failures occur when migrating hosts during hardware failures. Therefore, the inventors disclose dynamic resource migration, for example, migrating resources from one resource provider to another when a host goes down. In such a situation, according to an exemplary embodiment, the IT system can migrate to another server, node, or resource, or to another IT system. A controller can report the status of the system. Data replication is on another host with a known, automatically set up configuration. If a hardware failure is detected, any resources that the hardware may have provided can be automatically migrated after the failure is automatically detected.
[0029] A key issue with many IT systems is scalability. Growing companies or other organizations typically add or reconfigure their IT systems as they grow and their needs change. Problems arise when existing IT systems require more resources, such as additional hard drive capacity, storage capacity, CPU processing power, more network infrastructure, more endpoints, more clients, and / or more security. Configuration, setup, and deployment also become problematic when different services and applications, or changes to the infrastructure are required. According to an exemplary embodiment, a data center can scale automatically. Nodes or resources can be added to or removed from a pool of resources dynamically and automatically. Resources added to and removed from the resource pool can be automatically allocated or reallocated. Services can be provisioned and quickly moved to new hosts. A controller can detect more resources and dynamically add them to the resource pool, and know where resources should be allocated / reallocated. The system in the exemplary embodiment can scale from a single-node IT system to a scaled system requiring a large number of physical and / or virtual nodes or resources across multiple data centers or IT systems.
[0030] The inventors disclose a system that enables flexible resource allocation and management. This system includes compute, storage, and networking resources that may be in a resource pool and can be dynamically allocated. A controller can recognize new nodes or hosts on the network and then configure them so that they can become part of a resource pool. For example, when a new server is plugged in, the controller can configure it as part of a resource pool, add it to the resources, and make it available for use dynamically. Nodes or resources can be discovered by the controller and added to different pools. Resource requests can be made, for example, via API requests to the controller. The controller can then deploy or allocate the necessary resources from the pool according to rules. This allows the controller and / or applications via the controller to load balance and dynamically distribute resources based on the needs of the requests.
[0031] Examples of load balancing include, but are not limited to, deploying new resources in the event of a hardware or software failure, deploying one or more instances of the same application in response to increased user load, and deploying one or more instances of the same application in response to imbalances in storage, compute, or networking demands.
[0032] Issues involving changes to live IT systems or environments can cause significant, sometimes catastrophic, problems for users or entities that depend on the consistent operation of these systems. These outages represent not only potential losses in system use but also economic losses due to data loss and substantial resources—time, personnel, and money—required to resolve the issues. Problems can be exacerbated by the difficulty of rebuilding systems due to erroneous configuration documentation or a lack of understanding of the system. Because of this issue, many IT system users are reluctant to patch their IT resources to eliminate known security risks, thus leaving them more vulnerable to security breaches.
[0033] Many problems arising in IT system maintenance are related to software failures due to change management or control that may require configuration changes. Situations in which such failures may occur include, but are not limited to, upgrading to a new software version, migrating to different software, changing password or authentication management, and switching between services or between different providers of a service.
[0034] Infrastructure configured and maintained manually is typically difficult to recreate. Recreating infrastructure can be important for several reasons, including but not limited to rolling back problematic changes, power outages, or other disaster recovery. Diagnosing problems with manually configured systems is difficult. Recreating manually configured and maintained infrastructure is difficult. Also, system administrators can easily make mistakes, such as incorrect commands that are known to have brought down a computer system.
[0035] Making changes to a live IT system or environment can cause significant, sometimes catastrophic, problems for users or entities that depend on the consistent operation of these systems. These outages not only represent potential losses in system use, but such outages can also result in economic losses due to the considerable resources of time, personnel, and money required to resolve the problem, in addition to data loss. Problems can be exacerbated by the difficulty of rebuilding the system due to errors in configuration documentation or a lack of understanding of the system. Furthermore, restoring a system to its previous state after significant or major changes is often extremely difficult.
[0036] Furthermore, technical problems that may arise when changes are made to a live environment can have cascading effects. These cascading effects can make it difficult, and in some cases impossible, to revert to the previous state. Therefore, even if it is necessary to revert changes due to problems with the implemented changes, the system state has already been altered. In recent years, it has been stated that infrastructure and system management errors, as well as reverting incomplete changes to production environments, are unresolved problems. Moreover, it is known that testing changes to a system before deployment to a live environment is problematic.
[0037] Accordingly, the inventors disclose several exemplary embodiments of systems and methods configured to revert changes made to a system in operation back to its previous state. Furthermore, the inventors disclose systems and methods configured to enable significant reversals of the state of a modified system or environment in operation, which may prevent or mitigate one or more of the problems described above.
[0038] According to variations of the exemplary embodiment, the IT system has complete knowledge of the system through global system rules, templates, and IT system state. The infrastructure can be cloned using this complete knowledge of the system. The system or system environment can be cloned as software-defined infrastructure or environment. The system environment, including the volatile database in use, called the production environment, can be written to a non-volatile read-only database and used as a development environment in the development and testing process. Desired changes can be made to the development environment and tested there. User or controller logic can make changes to the global rules to create new versions. Versions of the rules can be tracked. Then, according to other aspects of the exemplary embodiment, the newly developed environment can be automatically implemented. The previous production environment can also be preserved or fully functional so that modifications to the production environment in a previous state are possible without data loss. The development environment can then be booted with new specifications, rules, and templates, and the database or system can be synchronized with the production database and switched to a writable database. The original production database can then be switched to a read-only database, and the system can be returned to it if recovery is needed.
[0039] Regarding software upgrades or patching, a new host may be deployed if a service requiring an upgrade or patch is detected. In the event of a failure due to an upgrade or patch, a new service may be deployed if it is possible to revert the changes as described above.
[0040] Hardware upgrades are crucial, especially in many situations where the latest hardware is essential. One example of this type of situation occurs in high-frequency trading industries, where IT systems with millisecond-level speed advantages can enable users to achieve superior trading results and profits. The challenge lies particularly in ensuring interoperability with the current infrastructure, as the new hardware must understand how to communicate using protocols and interact with the existing infrastructure. In addition to ensuring component interoperability, the components require integration with the existing setup.
[0041] Referring to Figure 1, an exemplary embodiment of IT system 100 is shown. System 100 may be one or more types of IT systems, including but not limited to those described herein.
[0042] A user interface (UI) 110 is illustrated, which is coupled to a controller 200 via an application program interface (API) application 120, which may or may not reside on a standalone physical or virtual server. The controller 200 may be deployed on one or more processors and one or more memories to perform any control operations described herein. Instructions executed by the processor(s) to perform such control operations may reside on a non-temporary computer-readable storage medium, such as processor memory. The API 120 may include one or more API applications, which may be redundant and / or operate in parallel. The API application 120 receives requests that constitute system resources, parses the requests, and passes them to the controller 200. The API application 120 receives one or more responses from the controller, parses the responses, and passes them to the UI(or application) 110. Alternatively, an application or service may communicate with the API application 120. The controller 200 is coupled to computing resources(s) 300, storage resources(s) 400, and networking resources(s) 500. Resources 300, 400, and 500 may or may not reside on a single node. One or more of resources 300, 400, and 500 may be virtual. Resources 300, 400, and 500 may or may not reside on multiple nodes, or in various combinations of multiple nodes. A physical device may contain one or more, or each of, resource types, including but not limited to computing resource 300, storage resource 400, and networking resource 500. Resources 300, 400, and 500 may also include pools of resources, regardless of whether they are in different physical locations or whether they are virtual or not. Bare metal computing resources may also be used to enable the use of virtual or container computing resources.
[0043] In addition to the known definition of a node, as used herein, a node can be any system, device, or resource connected to a network(s) or other functional unit that performs functions on a standalone or networked device. Nodes may also include, but are not limited to, servers, services / applications / multiple services on a physical or virtual host, virtual servers, and / or multiple or single services running on or within a multitenant server.
[0044] The controller 200 may include one or more physical or virtual controller servers, which may also be redundant and / or operate in parallel. The controller may operate on a physical or virtual host that functions as a computing host. For example, the controller may consist of a controller operating on a host that also serves other purposes, such as accessing highly sensitive resources. The controller may receive requests from the API application 120, parse the requests, optimize task distribution to other resources, instruct other resources, monitor and receive information from resources, maintain a history of system state and changes, and communicate with other controllers in the IT system. The controller may also include the API application 120.
[0045] Computing resources as defined herein may include a single, real or virtual computing node, or a resource pool containing one or more computing nodes. A computing resource or computing node may include one or more physical or virtual machine or container hosts capable of hosting one or more services or running one or more applications. A computing resource may reside on hardware designed for multiple purposes, including but not limited to computing, storage, caching, networking, and specialized computing, including but not limited to GPUs, ASICs, coprocessors, CPUs, FPGAs, and other specialized computing methods. Such devices may be added using PCI express switches or similar devices, and may be added dynamically in such a manner. A computing resource or computing node may include, or be capable of running, one or more hypervisors or container hosts containing multiple different virtual machines that may run services or applications or be virtual computing resources. A computing resource may be focused on providing computing capabilities, but may also include data storage capabilities and / or networking capabilities.
[0046] Storage resources as defined herein may include storage nodes or pools of storage resources. Storage resources may include any data storage medium, such as fast, slow, hybrid, cache, and / or RAM. Storage resources may include one or more types of networks, machines, devices, nodes, or any combination thereof, which may or may not be directly connected to other storage resources. According to aspects of exemplary embodiments, storage resources may be bare metal, virtual, or a combination thereof. Storage resources may be focused on providing storage functionality, but may also include computing and / or networking functionality.
[0047] Networking resource(s) 500 may include a single networking resource, multiple networking resources, or a pool of networking resources. Networking resource(s) may include physical or virtual devices(s), tools(s), switches, routers, or other interconnections between system resources, or applications for managing networking. Such system resources may be physical or virtual and may include computing, storage, or other networking resources. Networking resources may provide connectivity between external networks and application networks and may host core network services, including but not limited to DNS, DHCP, subnet management, Layer 3 routing, NAT, and other services. Some of these services may be deployed on computing resources, storage resources, or networking resources on physical or virtual machines. Networking resources may utilize one or more fabrics or protocols, including but not limited to Infiniband, Ethernet, RoCE, Fibre Channel, and / or Omnipath, and may include interconnections between multiple fabrics. Networking resources may or may not be SDN-enabled. The controller 200 may be able to configure the IT system topology by directly modifying the networking resources 300 using SDN, VLANs, etc. While the networking resources may focus on providing networking functionality, they may also include computing and / or storage capabilities.
[0048] As used herein, an application network means networking resources or any combination thereof for connecting or joining applications, resources, services, and / or other networks, or for connecting users and / or clients to applications, resources, and / or services. An application network may include networks used by a server to communicate with other application servers (physical or virtual) and to communicate with clients. An application network may communicate with machines or networks outside of system 100. For example, an application network may connect a web frontend to a database. Users may connect to a web application via the internet or other networks, which may or may not be managed by the controller.
[0049] According to an exemplary embodiment, computing, storage, and networking resources 300, 400, and 500 can be automatically added, removed, set up, allocated, reallocated, configured, reconfigured, and / or deployed by the controller 200, respectively. According to an exemplary embodiment, additional resources can be added to the resource pool.
[0050] The diagram illustrates a user interface 110, such as a Web UI or other user interface, through which a user 105 can access and interact with the system. Alternatively or additionally, an application may communicate with or interact with the controller 200 via an API application 120 or by other means. For example, a user 105 or application may submit requests including, but not limited to, building an IT system, building individual stacks within an IT system, creating a service or application, migrating a service or application, modifying a service or application, deleting a service or application, cloning a stack to another stack on a different network, and creating, adding, deleting, setting up or configuring, or reconfiguring resources or system components.
[0051] The system 100 in Figure 1 may include a server having connections or other communication interfaces with various elements, components, or resources that may be physical, virtual, or any combination thereof. In a modified example, the system 100 shown in Figure 1 may include a bare-metal server with connections.
[0052] As described in more detail herein, the controller 200 may be configured to add, allocate, manage, and update available resources by powering on resources or components and automatically setting up, configuring, and / or controlling the boot-up of resources. The power-on process may begin with powering on the controller so that the order in which devices are booted is consistent and independent of the user powering on the devices. This process may also include the detection of powered-on resources.
[0053] Referring to Figures 2A to 10, the controller 200, controller logic 205, global system rule database 210, IT system state 220, and template 230 are shown.
[0054] System 100 includes global system rules 210. Global system rules 210 may declare rules for setting up, configuring, booting, allocating, and managing resources, which may include, among other things, computing, storage, and networking. Global system rules 210 include the minimum requirements for System 100 to be in the correct or desired state. These requirements may include IT tasks that are expected to be completed and an updatable list of expected hardware that is necessary to build the desired system as expected. The updatable list of expected hardware allows the controller to verify that the necessary resources are available (for example, before the rules start or before using templates). Global rules may include a list of actions and corresponding instructions required for various tasks, relating to the ordering of actions and tasks. For example, rules may specify the order in which components are powered on, the order in which resources, applications, and services are booted, dependencies, and when various tasks should be started, such as when to start loading, configuring, starting, reloading applications, or updating hardware. Rule 210 may also include, for example, a list of resource allocations required for applications and services, a list of templates that can be used, a list of applications to be loaded and how to configure them, a list of services to be loaded, and a list of application networks and how to configure which applications fit into which networks, a list of configuration variables specific to various applications and user-specific application variables, expected states that allow the controller to check the system state to ensure that the state is as expected and that the results of each instruction are as expected, and / or a version list including a list of rule changes (e.g., snapshots) that may allow tracking of rule changes and the ability to test different rules in different situations or revert to them. The controller 200 may be configured to apply the global system rule 210 to the IT system 100 on physical resources.The controller 200 may be configured to apply global system rules 210 to the IT system 100 on virtual resources. The controller 200 may be configured to apply global system rules 210 to the IT system 100 on a combination of physical and virtual resources.
[0055] Figure 2M shows an exemplary set of system rules 210 that may take the form of global system rules. The exemplary set of system rules 210 shown in Figure 2M may be loaded into the controller 200 or derived by querying the system state (see 210.1). In the example in Figure 2M, system rules 210 includes a set of instructions that may take the form of a configuration routine 210.2, and also includes data 210.3 for creating and / or recreating an IT system or environment. The configuration rules within system rules 210 may know how to find templates 230 via a required template list 210.7 (templates 230 may reside in file systems, disks, storage resources, or be placed within system rules). The controller logic 205 may also search for templates 230 before processing them and activate system rules 210 after confirming their existence. System rules 210 may include a subset of system rules 210.15, which may be executed as part of a configuration routine 210.2.
[0056] Furthermore, subsystem rules 210.15 can be used, for example, as a tool for building a system of integrated IT applications (in which case they are processed by the system rule execution routine 210.16, which updates the system state and current configuration rules to reflect the additions of 210.15). Subsystem rules 210.15 can also be located elsewhere and loaded into the system state 220 through user interaction. For example, subsystem rules 210.15 can be held as a playbook, making them available and operational (the global system rules 210 are then updated, allowing the playbook to be replayed if a system clone is desired).
[0057] The configuration routine 210.2 can be a set of instructions used to build the system. The configuration routine 210.2 may also include subsystem rules 210.15 or system state pointers 210.8, if desired by the implementer. When the configuration routine 210.2 is running, the controller logic 205 optionally enables parallel deployment by processing a set of templates in a specific order (210.9), while maintaining proper dependency handling (210.12). The configuration routine 210.2 may optionally invoke API calls 210.10 that can set configuration parameters 210.5 for an application that can be configured by processing templates according to 210.9. Additionally, required services 210.11 are services that must be running when the system makes API calls 210.10.
[0058] Routine 210.2 may include procedures, programs, or methods for data loading (210.13) related to volatile data 210.6, including but not limited to copying data, transferring the database to computing resources, pairing computing resources with storage resources, and / or updating system state 220 by the location of volatile data 210.6. A pointer to volatile data (see 210.4) can be kept along with data 210.3 to find volatile data that may be stored elsewhere. The data loading routine 210.13 may also be used to load configuration parameters 210.5 if they are located in a non-standard data store (for example, contained in a database).
[0059] System rule 210 may also include a resource list 210.18 that can instruct which components are assigned to which resources, enabling controller logic 205 to determine whether appropriate resources and / or hardware are available. System rule 210 may also include an alternative hardware and / or resource list 210.19 for alternative deployments (e.g., for a development environment where a software engineer might want to perform demonstration testing but does not want to allocate an entire data center). System rule may also include a data backup / standby routine 210.17 that provides instructions on how to back up the system and how to use standby for redundancy.
[0060] After any action has been taken, the system state 220 may be updated, and the query (which may include a write) may be saved as system state query 210.14.
[0061] Figure 2N illustrates an exemplary process flow for controller logic 205 that processes system rule 210 (or subsystem rule 210.15) in Figure 2M. In step 210.20, controller logic 205 checks to ensure that appropriate resources are available (see 210.18 in Figure 2M). Otherwise, in step 210.21, an alternative configuration may be checked. A third option may include facilitating the user to select an alternative configuration that can be supported by template 230, referenced in listing 210.7 in Figure 2M.
[0062] In step 210.22, the controller logic may then check for computing resources (or any appropriate resources) to obtain access to volatile data. This may involve connecting to a storage resource or adding the storage resource to system state 220. In step 210.23, the configuration routines are then processed, and system state 220 is updated as each routine is processed (step 210.24). System state 220 may also be queried to check whether a particular step has finished before processing (step 210.25).
[0063] A configuration routine that processes the steps shown in Figure 210.23 may include any (or a combination thereof) of the procedures in 210.26. It may also include other procedures. For example, the processing in 210.26 may include processing templates (210.27), loading configuration data (210.28), loading static data (210.29), loading dynamic volatile data (210.30), and / or combining services, applications, subsystems, and / or environments (210.31). Such procedures in 210.26 may be repeated in a loop or executed in parallel so that some system components can be independent and others can be independent. Controller logic, service dependencies, and / or system rules may dictate which services can be interdependent and may combine services to further expand the IT system from the system rules.
[0064] The global system rule 210 may also include storage expansion rules. Storage expansion rules may include, for example, a set of rules that automatically add storage resources to existing storage resources within the system. In addition, trigger points may be included to detect when an application running on a computing resource(s) requests storage expansion (or when the controller 200 can detect when the storage of a computing resource(s) or application(s) is expanded. The controller 200 may allocate and manage new storage resources, and may merge or integrate storage resources with existing storage resources for specific execution resources.
[0065] Such specific execution resources may be, but are not limited to, computing resources within the system, applications running computing resources within the system, virtual machines, containers, or physical or virtual computing hosts, or a combination thereof. The execution resources may signal to the controller 200 that storage space is being exhausted, for example, through storage space queries. Networking or joining to the in-band management connection 270, the SAN connection 280, or the controller 200 may be used in such queries. An out-of-band management connection 260 may also be used.
[0066] For resources that are not currently running, those storage expansion rules (or a subset of those storage expansion rules) may also be used.
[0067] Storage expansion rules instruct how to identify, connect to, and set up new storage resources within the system. The controller registers the new storage resource in system state 220, notifying the running resources where the storage resource resides and how to connect to it. The running resources connect to the storage resource using such registration information. The controller 200 may merge the new storage resource with existing storage resources, or it may add the new storage resource to a volume group.
[0068] Figure 2B shows an exemplary flow of operation for an exemplary set of storage expansion rules. In step 210.41, the running resource determines that storage is low based on a trigger point or otherwise. In step 210.42, the running resource connects to the controller 200 via an in-band management connection 270, a SAN connection 280, or another type of connection visible to the operating system. Through this connection, the running resource can notify the controller 200 that storage is low. In step 210.43, the controller configures the storage resources to increase storage capacity for the running resource. In step 210.44, the controller provides the running resource with information about the location of the newly configured storage resources. In step 210.45, the running resource connects to the newly configured storage resources. In step 210.46, the controller adds a mapping of the location of the new storage resources to the system state 220. Next, the controller can add the new storage resource to the volume group assigned to the execution resource (step 210.47), or the controller can add the assignment of the new storage resource to the execution resource to system state 220 (step 210.48).
[0069] Figure 2C illustrates an alternative embodiment of performing steps 210.41 and 210.42 in Figure 2B. In steps 210 and 210.50, the controller sends a key command through the out-of-band management connection 260 to view a monitor or console for storage status updates for the running resource. For example, the monitor may be an IPMI console through which the screen can be reviewed via the out-of-band connection 260. For example, the out-of-band connection 260 may be connected via USB as a keyboard / mouse or via a VGA monitor port. In step 210.51, the running resource displays information on the screen. In step 210.52, the controller then reads the information presented on the monitor or console via the out-of-band management connection 260 and screen scraping or similar operations, and this read information may indicate a low storage state based on the trigger point.
[0070] The process flow may then continue to step 210.43 in Figure 2B.
[0071] Figure 2D illustrates another alternative embodiment that performs steps 210.41 and 210.42 in Figure 2B. In step 210.55, the execution resource automatically displays information on the monitor or console for the controller to read. In step 210.56, the controller automatically, periodically, or continuously reads the monitor or console to check the execution resource. In response to this read, the controller confirms that the execution resource has low storage (step 210.57). The process flow may then continue to step 210.43 in Figure 2B.
[0072] The controller 200 also includes a library of templates 230, which may include bare-metal and / or service templates. These templates may include, but are not limited to, third-party applications that can be configured by email, file storage, voice over IP, software counting, software XMPP, wiki, version control, account authentication management, and user interfaces. Templates 230 may have associations with resources, applications, or services, or they may serve as recipes defining how such resources, applications, or services are integrated into the system.
[0073] In this manner, a template may include an established set of information used to create, configure, and / or deploy a resource, or an application or service loaded onto that resource. Such information may include, but is not limited to, a kernel, an initrd file, a file system or file system image, files, configuration files, configuration file templates, information used to determine the appropriate setup for different hardware and / or computing backends, and other available options for configuring resources to run applications and / or operating system images that enable and / or facilitate the creation, booting, or execution of applications.
[0074] A template may include information that can be used to deploy an application on multiple supported hardware types and / or compute backends, including, but not limited to, multiple physical server types or components, multiple hypervisors running on multiple hardware types, and container hosts that can be hosted on multiple hardware types.
[0075] A template may derive a boot image for an application or service to run on computing resources. Templates and images derived from templates may be used to coordinate resources for various system functions to create applications, deploy applications or services, and / or enable and / or facilitate the creation of applications. A template may have variable parameters in files, file systems, and / or operating system images that can be overridden by configuration options from either default settings or settings provided by a controller. A template may have configuration scripts used to configure applications or other resources, which may utilize configuration variables, configuration rules, and / or default rules or variables, and these scripts, variables, and / or rules may include specific rules, scripts, or variables for specific hardware or other resource-specific parameters, e.g., hypervisor (when virtual), available memory. A template may have files, binary resources, or compilable source code that results in hardware or other resource-specific parameters, a specific set of binary resources, or source code with compilation instructions for specific hardware or other resource-specific parameters, e.g., hypervisor (when virtual), available memory. A template may contain a set of information independent of what is being executed on the resource.
[0076] A template may include a base image. The base image may include a base operating system filesystem. The base operating system may be read-only. The base image may also include basic tools for an operating system independent of what is running. The base image may include a base directory and operating system tools. A template may include a kernel. The kernel or multiple kernels may include an initrd kernel or multiple kernels configured for different hardware and resource types. An image may be derived from a template advertisement loaded or deployed to one or more resources. A loaded image may also include boot files, such as the kernel of the corresponding template or the kernel of the initrd.
[0077] The image may include template file system information that can be loaded into resources based on a template. The template file system may constitute an application or service. The template file system may include a shared file system common to all resources or similar resources, for example, to save storage space on which the file system is stored or to facilitate the use of read-only files.
[0078] A template file system or image may contain a set of files common to the services being deployed. The template file system may be pre-loaded on the controller or downloaded. The template file system may be updated. Because the template file system may not require reconstruction, relatively rapid deployment can be enabled. Sharing the file system with other resources or applications can reduce storage usage because files are not unnecessarily duplicated. This also allows for easier recovery from failures because only files different from those in the template file system need to be restored.
[0079] The template boot file may include the kernel and / or an initrd or a similar filesystem used to assist the boot process. The boot file can boot the operating system and set up the template filesystem. The initrd may include a small temporary filesystem containing instructions on how to set up the template so that it can be booted.
[0080] The template may further include BIOS settings. Template BIOS settings may be used to configure optional settings for running an application on a physical host. If used, the out-of-band management 260 may then be used to boot the resource or application, as described herein with respect to Figures 1-12. The physical host may boot the resource or application using the out-of-band management network 260 or a CD-ROM. The controller 200 may configure application-specific BIOS settings defined in such a template. The controller 200 may use the out-of-band management system to make direct BIOS changes through APIs specific to a particular resource. The settings may be verified through the console and image recognition. Thus, the controller 200 may use console functionality and make BIOS changes via a virtual keyboard and mouse. The controller may also use the UEFI shell, type directly into the console, verify successful results, type accurately in the commands and use image recognition to ensure the success of the configuration changes. If a bootable operating system is available for changing or updating the BIOS to a specific BIOS version, the controller 200 may remotely load a disk image or ISO boot in which the operating system runs an application that updates the BIOS and enables the configuration change in a reliable manner.
[0081] The template may also include a list of template-specific supported resources or a list of resources required to run a particular application or service.
[0082] The template image or part of the image or template may be stored in the controller 200, or the controller 200 may move or copy it to the storage resource 410.
[0083] Figure 2E shows an exemplary template 230. The template contains all the information necessary to create an application or service. Template 230 may also contain information, alternative data, files, and binaries for different hardware types that provide similar or identical functionality. For example, there may be filesystem blobs 232 for / usr / bin and / bin, compiled for different architectures by binaries 234. Template 230 may also contain a daemon 233 or script 231. Daemon 233 is a binary or script that can be executed at boot time when the host is powered on or ready, and in some cases, daemon 233 may be accessible by the controller and run an API that allows the controller to change the host configuration (and the controller can subsequently update active system rules). Daemons may be powered down and restarted through out-of-band management 260 or in-band management 270 as discussed above and below. Those daemons may also run common APIs to provide dependent services to new services (e.g., a common web server API that communicates with an API that controls nginx or apache). Script 231 may be an installation script that can be executed while the image is booting or afterward, or after the daemon has been started or the service has been enabled.
[0084] Template 230 may also include kernels 235 and a pre-boot filesystem 236. Template 230 may also include multiple kernels 235 and one or more pre-boot filesystems for different hardware and different configurations (such as an initrd or initramfs for Linux, or a read-only RAM disk for BSD). The initrd may also be used to mount a filesystem blob 232 presented as an overlay by booting into an initramfs 236, which can optionally connect to a storage resource via a SAN connection 280, as discussed below, and to mount the root filesystem on remote storage.
[0085] The filesystem blob232 is a filesystem image that can be divided into separate blobs. The blobs may be interchangeable based on configuration options, hardware type, and other differences in setup. A host booted from template 230 may boot from a union filesystem (such as an overlayfs) that contains multiple blobs or images created from one or more filesystem blobs.
[0086] Template 230 may also include additional information 237, such as volatile data 238 and / or configuration parameters 239, or may be linked to the additional information 237. For example, volatile data 238 may be contained within template 230, or it may be contained externally. It may be, but is not limited to, a file system blob 232, or a database, flat file, files stored in a directory, a tarball of files, a git, or other data store including other version control repositories. In addition, configuration parameters 239 may be contained externally or internally within template 230, and optionally contained in system rules and applied to template 230.
[0087] System 100 further includes IT system state 220, which tracks, maintains, changes, and updates the state of System 100, including but not limited to resources. System state 220 may track available resources, which inform the controller logic whether there are resources available for rule implementation and templates, and which resources are available. System state may track used resources, which allows controller logic 205 to inspect and utilize efficiency, whether there is a need to switch for upgrades or other reasons such as improving efficiency or prioritizing. System state may track which applications are running. Controller logic 205 may compare expected running applications to actual running applications according to the system state and whether revisions are necessary. System state 220 may also track where applications are running. Controller logic 205 may use this information for the purpose of evaluating efficiency, change management, updates, troubleshooting, or audit trails. System state may track networking information, such as which networks are on or currently operational, or configuration values and history. System state 220 may track the history of changes. System state 220 may also track which templates are used in which deployments, based on global system rules that specify which templates are used. The history may be used for auditing, alerting, change management, build reports, tracked versions correlated with hardware and applications and configuration, or configuration variables. System state 220 may maintain a history of configurations for auditing, compliance inspection, or troubleshooting purposes.
[0088] The controller has logic 205 for managing all information contained in the system state, templates, and global system rules. The controller logic 205, global system rule database 210, IT system state 220, and templates 230 are managed by the controller 200 and may or may not reside in the controller 200. The controller logic, or application 205, global system rule database 210, IT system state 220, and templates 230, may or may not be physical or virtual, and may or may not be distributed services, distributed databases, and / or files. The API application 120 may be included together with the controller logic / controller application 205.
[0089] Controller 200 may run as a standalone machine and / or may contain one or more controllers. Controller 200 may contain controller services or applications and may run inside another machine. The controller machine may start controller services first to ensure that the booting of the entire stack or a group of stacks is ordered and / or consistent.
[0090] The controller 200 may control one or more stacks by computing, storage, and networking resources. Each stack may or may not be controlled by a different subset of rules within the global system rules 210. For example, there may be pre-created, created, deployed, inspected stacks, parallel, backup, and / or other stacks with different functions within the system.
[0091] The controller logic 205 may be configured to read and interpret global system rules to achieve a desired IT system state. The controller logic 205 may be configured to use templates that conform to the global rules to build system components such as applications or services and to allocate, add, or remove resources to achieve a desired IT system state. The controller logic 205 may develop a list of tasks that read the global system rules, correct the state, and issue commands to satisfy the rules based on available operations. The controller logic 205 may include logic for performing operations such as starting the system, adding, removing, and reconfiguring resources, and identifying what is available to do so. The controller logic may check the system state at startup time and at periodic intervals to determine if hardware is available, and if so, may perform tasks. If the required hardware is not available, the controller logic 205 presents the global system rules 210, templates 220, and alternative options, and uses the available hardware from the system state 230 to correct the global rules and / or system state 220 accordingly.
[0092] The controller logic 205 can determine which variables are needed, what the user needs to input to continue, or what the user needs in the system to function. The controller logic may use a list of templates from global system rules and compare them with the required templates in the system state to ensure that the required templates are available. The controller logic 205 may identify from the system state database whether resources on a template-specific list of supported resources are available. The controller logic may allocate resources, update the state, or initiate the next set of tasks to implement global rules. The controller logic 205 may start / run the application on the allocated resources as specified in the global rules. The rules can specify how to build the application from the templates. The controller logic 205 may obtain the template(s) and construct the application from the variables. The templates may inform the controller logic 205 which kernels, boot files, file systems, and supported hardware resources are needed. The controller logic 205 may then add information about the application deployment to the system state database. After each instruction, the controller logic 205 may check the system state database against the expected state of the global rule to verify whether the expected operation was completed exactly as predicted.
[0093] The controller logic 205 may use a version according to the version rule. The system state 220 may have a database that correlates which rule version is used in different deployments.
[0094] The controller logic 205 may include efficient logic and efficient ordering for rule optimization. The controller logic 205 may be configured to optimize resources. Information in system state, rules, and templates related to applications that are running or expected to run may be used by the controller logic to implement efficiency or priority for resources. The controller logic 205 may use information in "Used Resources" in system state 220 to determine efficiency or the need to upgrade, reuse for another purpose, or switch resources for other reasons.
[0095] The controller may check which applications are running according to system state 220 and compare them to the expected applications running according to the global rules. If an application is not running, the controller may start it.
[0096] If an application should not be running, it may be stopped, and resources may be reallocated as appropriate. The controller logic 205 may include a database of resource (computation, storage, networking) specifications. The controller logic may include logic for recognizing available resource types for systems that can be used. This may be done using an out-of-band management network 260. The controller logic 205 may be configured to recognize new hardware using out-of-band management 260. The controller logic 205 may also take information about change history, rules used, and versions from the system state 220 for auditing, reporting, and change management purposes.
[0097] Figure 2F shows an exemplary process flow for controller logic 205 relating to processing template 230 and deriving an image to boot, power on, and / or enable a resource that may be referred to as a host for this exemplary purpose. This process may also include configuring storage resources and joining storage and compute hosts and / or resources. Controller logic 205 is aware of the hardware resources available in system 100, and system rules 210 may indicate which hardware resources are available. In step 205.1, controller logic 205 parses template 230, which may include an instruction file that can be executed to cause controller logic to collect files that are external to template 230 as shown in Figure 2E. The instruction file may be in JSON format. In step 205.2, controller logic collects a list of required file buckets. Furthermore, in step 205.3, the controller logic 205 collects necessary hardware-specific files into a bucket that are referenced by the hardware and, optionally, by the hypervisor (or container host system, multi-tenancy type). References to the hypervisor (or container host system, or multi-tenancy type) may be necessary if the hardware will be running on a virtual machine.
[0098] If hardware-specific files exist, the controller logic collects them in step 205.4. In some cases, the filesystem image may contain the kernel and initramfs along with a directory containing the kernel modules (or kernel modules ultimately placed in a directory). The controller logic 205 then selects a compatible and appropriate base image in step 205.5. The base image contains operating system files that may not be specific to the application or image derived from template 230. Compatibility in this context means that the base image contains the files necessary to modify the template for the application on which it runs. The base image may be managed outside the template as a mechanism to save space (and often the base image may be identical for several applications or services). In addition, in step 205.6, the controller logic 205 selects a bucket(s) containing executable files, source code, and hardware-specific configuration files. Template 230 may refer to other files, including, but is not limited to, configuration files, configuration file templates (configuration files that contain placeholders or variables that are filled by variables in system rule 210, which may be known in template 230, so that the controller 200 can convert the configuration template into a configuration file and optionally modify the configuration file through an API endpoint), binaries, and source code (which can be compiled when the image is booted). In step 205.7, hardware-specific instructions corresponding to the elements selected in steps 205.4, 205.5, and 205.6 may be loaded as part of the image to be booted. Controller logic 205 derives the image from the selected components. For example, there may be different pre-installation scripts for physical hosts versus virtual machines, or differences between Powerpc and x86.
[0099] In step 205.8, controller logic 205 mounts the overlayfs and repackages the subject files into a single filesystem blob. When multiple filesystem blobs are used, the image may be created by multiple blobs, decompressing the tarball and / or fetching the git. If step 205.8 is not performed, the filesystem blobs may remain isolated, and the image may be created as a set of filesystem blobs and mounted by a filesystem capable of mounting multiple smaller filesystems (such as overlayfs) together. Controller logic 205 may then identify a compatible kernel (or a kernel specified in system rule 210) in step 205.9, and an applicable initrd in step 205.10. A compatible kernel may be a kernel that satisfies the dependencies of the template and the resources used to implement the template. A compatible initrd may be an initrd that loads the template onto the desired computing resources. Often, the initrd may be used for physical resources so that storage resources can be mounted before a full boot (so that the root filesystem may be remote). The kernel and initrd may be packaged into a filesystem blob, which may be used for direct kernel booting using kexec, or on a physical host, to modify the kernel on the live system after booting a backup operating system.
[0100] The controller then configures the storage resources(s) to enable the computing resources(s) to run applications(s) and / or images(s) using any of the techniques outlined in 205.11, 205.12, and / or 205.13. According to 205.11, an overlayfs file may be provided as a storage resource. According to 205.12, a filesystem is presented. For example, the storage resource may present multiple filesystem blobs that the computing resources can mount simultaneously using a combined filesystem or a filesystem similar to an overlayfs. According to 205.13, the blobs are sent to the storage resource before the filesystem is presented.
[0101] Figures 2G and 2H show exemplary process flows for steps 205.11 and 205.12 of Figure 2F. Furthermore, the system may employ processes and rules for connecting computer resources to storage resources, which may also be referred to as storage connection processes. Examples of such storage connection processes, in addition to those shown in Figures 2G and 2H, are provided in Appendix A attached thereto. Figure 2G shows an exemplary process flow for connecting storage resources. Some storage resources may be read-only, while others may be writable. Storage resources may manage their own write locks to prevent simultaneous writes that would create a competitive state, or system state 220 may track which connections can write to the storage resource and / or prevent multiple read-write connections to the resource (step 205.21) (see, for example, step 205.20). The controller logic or the resource itself may query the controller's system state 220 for the location and transmission type of the storage resource (e.g., iSCSI, ISER, NVMeOF, Fibre Channel, FCOE, NFS, NFS over RDMA, AFS, CIFS, Windows Share) (step 205.22). If the computing resource is virtual, the hypervisor (e.g., via the hypervisor daemon) may handle the connection to the storage resource (step 205.23). This may have desirable security advantages because the virtual machine may not be aware of SAN280.
[0102] Referring to step 205.24, the process of connecting the computing resource and the storage resource may be instructed in system rule 210. The controller logic then queries system state 220 to verify that the resource is available and, if necessary, writable (step 205.22). System state 220 may be queried via any number of techniques, such as SQL queries (or other types of database queries), JSON parsing, etc. The query returns the information necessary for the computing resource to connect to the storage resource. The controller 200, system state 220, or system rule 210 may provide authentication credentials for the computing resource to connect to the system state (step 205.25). The computing resource then updates system state 220, either directly or via the controller (step 205.26).
[0103] Figure 2H illustrates an exemplary boot process for a physical, virtual, or other type of computing resource, application, service, or host to power on and connect to a storage resource. The storage resource may optionally utilize a fused file system and / or expandable volumes. In situations where a controller or other system enables a physical host, the physical host may be preloaded with an operating system to configure the system. Therefore, in step 205.31, the controller may preload the boot disk with an initramfs. Alternatively, the controller 200 may use an out-of-band management connection 260 to network boot a backup operating system (step 205.30) and then, optionally, preload the host with the backup operating system (step 205.31). The initramfs is then loaded in step 205.32, and the storage resource is connected in step 205.33 using the method shown in Figure 2G.
[0104] Next, if expandable volumes exist, the subvolumes or devices to be joined together are optionally assembled as a volume group in step 205.34 if Logical Volume Management (LVM) is in use. Alternatively, they may be joined in step 205.34 using other methods of combining disks.
[0105] If a fused filesystem is in use, in step 205.36 the files may be combined, and then the boot process may continue (step 205.46). If overlayfs is in use in Linux to fix some known issues, the following subprocesses may be executed: A / data directory may be created in each mounted filesystem blob, which may be volatile (step 205.37). Then, in step 205.38 the new_root directory may be created, and in step 205.39 the overlayfs is mounted to the directory. Then initramfs executes exec_root on / new_root (step 205.40).
[0106] If the host is a virtual machine, additional tools such as direct kernel booting may be available. In this situation, the hypervisor may connect to the storage resource before booting the VM (step 205.41), or it may do so while booting. The VM may then boot a direct kernel along with loading the initramfs (step 205.42). The initramfs is then loaded in step 205.43, and at this point, the hypervisor may connect to a storage resource that may be remote (step 205.44). To achieve this, the hypervisor host may need to go through an interface (for example, if inifiniband needs to connect to an iSER target, it may go through a virtual function based on SR-IOV using pci-passhtru, or in some situations, a paravirtualized network interface may be used). These connections are made available by the initramfs. The virtual machine may then connect to the storage resource in step 205.45, if it has not already done so. It may also receive its storage resources through the hypervisor (optionally, through paravirtualized storage). The process may also do the same for virtual machines that have fused file systems and LVM-style disks mounted, optionally.
[0107] Figure 2O illustrates an exemplary process flow for configuring storage resources from filesystem blobs or other groups of files, as shown in 205.13. The blobs are collected in step 205.75 and may be copied directly to the storage resource host in 205.73 (if the storage resource host is different from the device holding filesystem blob232). Once the storage resources are in place, the system state is then updated in 205.74 with the location of the storage resources and available transmissions (e.g., iSER, nvmeof, iSCSI, FCoE, Fibre Channel, NFS, NFS over RDMA). Some of those blobs may be read-only, in which case the system state remains identical and a new computing resource or host may connect to that read-only storage resource (e.g., when connecting to a base image). In some cases, it may be desirable to place the files in a single filesystem image to avoid the overhead of any of the fused filesystems, as shown in 205.70. This may be achieved by mounting the blobs as a fused file system (step 205.71), then copying them to a new file system, or repackaging them as a single file system (step 205.72), and then optionally copying the new file system image to a suitable location where the new file system image will be presented as a storage resource. Some fused file systems can be merged without first mounting them in step 205.71, allowing them to be merged into a single step.
[0108] Figure 2I illustrates another exemplary template 230, as shown in Figure 2E. In this embodiment, the controller may be configured to use template 230 as shown in Figure 2I, together with an intermediate configuration tool. According to the exemplary embodiment, the intermediate configuration tool may include a common API used to combine a new application or service with a dependent application or service. Thus, template 230 may also include a list 244 of dependencies that may be required to set up the template's services. Template 230 may also include connection rules 245 that can contain calls to the common API of dependencies. Template 230 may also include one or more common APIs 243, as well as a list 242 of common APIs and their versions. The common APIs 243 may have methods, functions, scripts, or instructions, which may be callable (or not callable) from an application or controller, enabling the controller to configure the dependent application or service so that the dependent application or service can be combined with a new application built by template 230. The controller may communicate with the common API 243 and / or make API calls to configure a coupling between a new service or application and a dependent service or application. Alternatively, instructions may allow an application or service to communicate with the common API 243 and / or send calls to the common API 243 directly on the dependent application or service. The connection rule 245 of template 230 is a set of rules and / or instructions that may contain API calls relating to coupling a new service or application with a dependent service or application.
[0109] System state 220 may further include a list 246 of running services. The list 246 of running services may be queried by controller logic 205 to satisfy dependencies 244 from template 230. The controller may also include a list 247 of different common APIs available for a particular service / application or type of service / application, and may also include templates that contain the common APIs. The list may reside in controller logic 205, system rules 210, system state 220, or template storage accessible by the controller. The controller also maintains an index 248 of common APIs compiled from all existing or loaded templates.
[0110] Figure 2J relates to the controller logic 205 processing the template 230, as shown in Figure 2F, but step 255 illustrates an exemplary process flow in which the controller manages service dependencies. Figure 2K shows an exemplary process flow for step 255 in Figure 2J. In step 255.1, the controller collects a list of dependencies 244 from the template. The controller also collects a list of common APIs 243 from the template. (A) In step 255.2, the controller narrows down the list of possible dependent applications or services by comparing the list of common APIs 243 from the template with the index 248 of the common APIs, and based on the types of applications or services that are likely to satisfy the dependency. In step 255.3, the controller determines whether system rule 210 specifies a way in which the dependency is satisfied.
[0111] If the answer in step 255.3 is yes, the controller determines whether the dependent service or application is running by querying a list of execution templates (step 255.4). If the answer in step 255.4 is no, the service application, which can contain controller logic to process the template for the dependent service / application, is executed (and / or configured and then executed) (step 255.5). If it is found that the dependent service or application is running in step 255.4, the process flow then proceeds to step 255.6. In step 255.6, the controller uses the template to join the new service or application being built with the dependent service or application. When joining the new service or application with the dependent application / service, the controller considers the template it is processing and executes connection rule 245. Based on connection rule 245, the controller sends commands to the common API 243 on how to satisfy dependency 244 and / or how to join the application / service. Common API 243 translates instructions from the controller to connect new services or applications and dependent applications or services, and may include, but is not limited to, calling API functions of a service, changing its configuration, executing scripts, and calling other programs. Following step 255.6, the process flow proceeds to step 205.2 in Figure 2J.
[0112] If step 255.3 determines that system rule 210 does not specify a way to satisfy dependency, the controller then queries system state 220 in step 255.7 to check whether a suitable dependency application or service is running. In step 255.8, the controller makes a determination based on the query about whether a suitable dependency application or service is running. If the answer in step 255.8 is no, the controller may then notify the administrator or user about the action (step 255.9). If the answer in step 255.8 is yes, the process flow then proceeds to step 255.6, which can operate as discussed above. The user may optionally query whether they should connect to a dependency application that a new application is running, in which case the controller may, in step 255.6, connect the new application or service to the dependency application or service as follows: the controller considers the template 230 it is processing and executes connection rule 245. The controller then sends a command to the common API 243 based on connection rules 245 regarding how to satisfy dependency 244. The common API 243 translates the instructions from the controller to connect the new service or application and the dependent application or service.
[0113] The user communicates with the controller 200 through an external user interface, web UI, or application via an API application 120, which may also be incorporated into the controller application or logic 205.
[0114] The controller 200 communicates with the stack or resources through one or more networks, interconnections, or other connections through which the controller can instruct computing, storage, and networking resources to operate. Such connections may include out-of-band management connections 260, in-band management connections 270, SAN connections 280, and optional on-network in-band management connections 290.
[0115] Out-of-band management may be used by the controller 200 to discover, configure, and manage components of system 100 through the controller 200. The out-of-band management connection 260 can enable the controller 200 to discover resources that are plugged in and available but not turned on. Resources may be added to the IT system state 220 when plugged in. Out-of-band management may be configured to load a boot image and configure and monitor resources belonging to system 100. Out-of-band management may also boot a temporary image for operating system diagnostics. Out-of-band management may be used to change BIOS settings and may also use console tools to execute commands on the running operating system. Settings may also be changed by the controller using image recognition of video signals from a console, keyboard, and physical or virtual monitor ports on hardware resources such as VGA, DVI, or HDMI ports, and / or APIs provided by out-of-band management, such as Redfish.
[0116] Out-of-band management as used herein may include, but is not limited to, a management system capable of connecting to a resource or node independent of the operating system and the main motherboard. The out-of-band management connection 260 may include a network, or multiple types of direct or indirect connections or interconnections. Examples of types of out-of-band management connections include, but are not limited to, IPMI, Redfish, SSH, telnet, other management tools, keyboard video and mouse (KVM), or KVM over IP, serial console, or USB. Out-of-band management is a tool that can be used over a network, capable of powering on and off nodes or resources, monitoring temperature and other system data, making BIOS and other low-level changes that may be beyond the control of the operating system, connecting to a console, sending commands, and controlling input including, but not limited to, keyboards, mice, and monitors. Out-of-band management may be coupled to out-of-band management circuitry within a physical resource. Out-of-band management may connect a disk image as a disk that can be used to boot installation media.
[0117] The management network or in-band management connection 270 can enable the controller to collect information about computing, storage, networking, or other resources and communicate directly with the operating system running the resources. Storage resources, computing resources, or networking resources may include management interfaces that interface with connections 260 and / or 270, thereby enabling them to communicate with the controller 200, inform the controller of what is running and available for the resources, and receive commands from the controller. An in-band management network as used herein includes resources and a management network that is capable of communicating directly with the operating system of the resources. Examples of in-band management connections may include, but are not limited to, SSH, telnet, other management tools, serial consoles, or USB.
[0118] Out-of-band management is described herein as a network physically or virtually separated from the in-band management network, but they may be combined with each other for efficiency, or operate together with each other, as will be described in more detail herein. Furthermore, out-of-band and in-band management, or embodiments thereof, may communicate through the same port of the controller, or be coupled with a combined interconnection. Optionally, one or more of connections 260, 270, 280, and 290 may be separated from or combined with others in such networks, and may or may not be part of the same fabric.
[0119] In addition, the computing resources, storage resources, and controllers may or may not be coupled to the storage network (SAN) 280 in such a manner that the controller 200 can use the storage network to boot each resource. The controller 200 may send a boot image or other template to separate storage or other resources or to other resources, so that the other resources can boot off the storage or other resources. The controller may instruct where to boot from in such a situation. The controller may power on the resources and instruct the resources on where to boot from and how to configure themselves. The controller 200 may instruct the resources on how to boot, which image to use, and where the image is located if it resides on another resource. The resources in the BIOS may be pre-configured. In addition, or instead, the controller may configure the BIOS through out-of-band management so that they boot off the storage area network. The controller 200 may also be configured to boot an operating system from an ISO and allow the resources to copy data to a local disk. The local disk may then be used for booting. The controller may configure other resources, including other controllers, in such a manner that resources can be booted. Some resources may include applications that provide computing, storage, or networking functions. In addition, the controller may be responsible for booting up storage resources and then supplying boot images for subsequent resources or services to the storage resources. Storage may also be managed through different networks used for other purposes.
[0120] Optionally, one or more of the resources may be coupled to an on-network in-band management connection 290. Connection 290 may include one or more types of in-band management, as described with respect to the in-band management connection 270. Connection 290 may connect controllers to the application network to utilize or manage the networks through the in-band management network.
[0121] Figure 2L illustrates an image 250 that can be loaded directly or indirectly (through another resource or database) from template 230 to a resource, or an application or service loaded on the resource, to boot the resource. Image 250 may include boot files 240 for the resource type and hardware. Boot files 240 may include kernels 241 corresponding to the resource, application, or service to be deployed. Boot files 240 may also include an initrd or similar filesystem used to assist the boot process. The boot system 240 may include multiple kernels or initrds configured for different hardware and resource types. In addition, image 250 may include a filesystem 251. Filesystem 251 may include a service image 253 and its corresponding filesystem, as well as a volatile image 254 and its corresponding filesystem, along with a base image 252 and its corresponding filesystem. The filesystems and loaded data may vary depending on the resource type and the application or service to be executed. Base image 252 may include a base operating system filesystem. The base operating system may be read-only. Base image 252 may also include basic operating system tools independent of what is being run. Base image 252 may also include base directories and operating system tools. Service file system 253 may include configuration files and specifications for resources, applications, or services. Volatile file system 254 may include information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may or may not be configured as variables, including but not limited to passwords, session keys, and private keys.The file systems may be mounted as a single file system using techniques such as overlayFS, allowing several read-only and several read-write file systems to reduce the amount of duplicate data used for applications.
[0122] As described above, the controller 200 may be used to add resources such as compute, storage, and / or networking resources to the system. Figure 11A illustrates an exemplary method of adding a physical resource, such as a bare metal node, to system 100. The resource, i.e., compute, storage, or networking resource, is plugged into the controller by a network connection 1110. The network connection may include an out-of-band management connection. The controller recognizes that the resource is plugged in through the out-of-band management connection 1111. The controller recognizes information related to the resource, including, but not limited to, the resource's type, capabilities, and / or attributes 1112. The controller adds the resource and / or information related to the resource to its system state 1113. The image derived from the template is loaded into a physical component of the system, including, but not limited to, another resource such as a resource, a storage resource, or a controller 1114. The image includes one or more file systems, which may include configuration files. Such configurations may include BIOS and boot parameters. The controller instructs the physical resource to boot using the file system of image 1115. Additional resources, or multiple bare metal or physical resources of different types, may be added in this manner using the template image or at least a part thereof.
[0123] Figure 11B illustrates an exemplary method for automatically allocating resources using global system rules and templates of an exemplary embodiment. A request is made to the system that requires resource allocation to fulfill the request (1120). The controller becomes aware of its resource pool based on its system state database (1121). The controller uses templates to determine the required resources (1122). The controller allocates the resources and stores the information in the system state (1123). The controller deploys the resources using templates (1124).
[0124] Referring to Figure 12, an exemplary method for automatically deploying an application or service using the system 100 described herein is illustrated. A user or application makes a request for the service (1210). The request is translated into an API application (1220). The API application routes the request to the controller (1230). The controller interprets the request (1240). The controller considers the state of the system and its resources (1250). The controller uses its rules and templates for service deployment (1260). The controller sends the request to the resources (1270), deploys the image derived from the template (1280), and updates the IT system state.
[0125] Additional and more detailed examples of actions such as adding resources, allocating resources, and deploying applications or services are discussed in more detail below.
[0126] Adding computing resources to the system Referring to Figure 3A, the addition of a computing resource 310 to system 100 is illustrated. When the computing resource 310 is added, it may be coupled to the controller 200 and powered off. If the computing resource 310 is preloaded by an image, alternative steps may follow, which may involve using one of the network connections to communicate with the resource, boot the resource, and add information to the system state. If the computing resource and controller are on the same node, the service running the computing resource is off.
[0127] As shown in Figure 3A, the computing resource 310 is coupled to the controller by a network, namely, an out-of-band management connection 260, an in-band management connection 270, and optionally, a SAN 280. The computing resource 310 is also coupled to one or more application networks 390, from which services, application users, and / or clients can communicate with each other. The out-of-band management connection 260 may be coupled to an independent out-of-band management device 315 or to the circuitry of the computing resource 310, which is turned on when the computing resource 310 is plugged in. The device 315 may enable functions including, but are not limited to, powering on / off the device, attaching to a console, typing commands, monitoring temperature and other computer health-related elements, setting BIOS settings, and other functions outside the scope of the operating system. The controller 200 can see the computing resource 310 through the out-of-band management network 260. It can also identify the type of computing resource and its configuration using in-band or out-of-band management. The controller logic 205 is configured to look up out-of-band management 260 or in-band management 270 to add hardware. If a computing resource 310 is discovered, the controller logic 205 may then use the global system rule 220 to determine whether the resource is configured automatically or through user interaction. If it is added automatically, the setup follows the global system rule 210 in the controller 200. If it is added by a user, the global system rule 210 in the controller 200 may query the user to confirm the addition of the resource and what the user wants to do with the computing resource.The controller 200 may query the API application to verify that the new resource is authorized, or it may request this from the user or any program controlling the stack. The authorization process may be completed automatically and securely using encryption to verify the legitimacy of the new resource. The controller logic 205 adds the computing resource 310 to the IT system state 220, which includes a switch or network into which the computing resource 310 is plugged.
[0128] If the computing resource is physical, the controller 200 may power on the computing resource via the out-of-band management network 260, and the computing resource 310 may boot off an image 350 loaded from template 230 using global system rules 210 and controller logic 205, for example, by SAN 280. The image may be loaded indirectly via other network connections or by another resource. Once booted, information related to the computing resource 310 received via the in-band management connection 270 may also be collected and added to the IT system state 220. The computing resource 310 may then be added to a storage resource pool, which becomes a resource managed by the controller 200 and tracked in the IT system state 220.
[0129] If the computing resource is virtual, the controller 200 may power on the computing resource through either the in-band management network 270 or out-of-band management 260. The computing resource 310 may boot off an image 350 loaded from template 230 using global system rules 210 and controller logic 205, for example, by SAN 280. The image may be loaded indirectly through other network connections or by another resource. Once booted, information related to the computing resource 310 received through the in-band management connection 270 may also be collected and added to the IT system state 220. The computing resource 310 may then be added to a storage resource pool, which becomes a resource managed by the controller 200 and tracked in the IT system state 220.
[0130] The controller 200 may also be capable of automatically turning resources on and off in accordance with global system rules and updating the IT system state for reasons determined by the IT system user, such as turning off resources to save power, turning on resources to improve application performance, or for other reasons that the IT system user may have.
[0131] Figure 3B shows an image 350 loaded directly or indirectly (through another resource or database) from template 230 to a computing resource 310 in order to boot the computing resource and / or load an application. Image 350 may include a boot file 340 for the resource type and hardware. The boot file 340 may include a kernel 341 corresponding to the resource, application, or service to be deployed. The boot file 340 may also include an initrd, or a similar file system used to assist the boot process. The boot system 340 may include multiple kernels or initrds configured for different hardware and resource types. In addition, image 350 may include a file system 351. File system 351 may include a service image 353 and its corresponding file system, as well as a volatile image 354 and its corresponding file system, along with a base image 352 and its corresponding file system. The file systems and loaded data may vary depending on the resource type and the application or service being run. The base image 352 may include a file system for the base operating system. The base operating system may be read-only. The base image 352 may also include basic operating system tools independent of what is being executed. The base image 352 may also include a base directory and operating system tools. The service file system 353 may include configuration files and specifications for resources, applications, or services. The volatile file system 354 may include information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may or may not be configured as variables, including but not limited to passwords, session keys, and private keys.The file systems may be mounted as a single file system using techniques such as overlayFS, allowing several read-only and several read-write file systems to reduce the amount of duplicate data used for applications.
[0132] Figure 3C shows an exemplary process flow for adding resources, such as a computing resource 310, to system 100. In this embodiment, the resource in question is described as computing resource 310, but it should be understood that the resource in question for the process flow in Figure 3C may also be storage resource 410 and / or networking resource 510. In the embodiment of Figure 3C, the resource 310 to be added is not on the same node as controller 200. In step 300.1, resource 310 is coupled to controller 200, which is powered off. In the embodiment of Figure 3C, an out-of-band management connection 260 is used to connect resource 310. However, it should be understood that other network connections may be used if required by the implementer. In steps 300.2 and 300.3, controller logic 205 examines the system's out-of-band management connection and uses the out-of-band management connection 260 to recognize and identify the type and configuration of the resource 310 to be added. For example, controller logic may look at the BIOS or other information about the resource (such as serial number information) as a reference to obtain type and configuration information.
[0133] In step 300.4, the controller uses a global system rule to determine whether a particular resource 310 should be added automatically. If it is not needed, the controller waits until its use is approved (step 300.5). For example, a user may respond to the query by indicating that they do not wish to use a particular resource 310, or that it may be automatically held in abeyance until it is used in step 300.4. If step 300.4 determines that resource 310 should be added automatically, the controller then uses that rule for automatic setup (step 300.6) and proceeds to step 300.7.
[0134] In step 300.7, the controller selects and uses a template 230 associated with a resource to add the resource to system state 220. In some cases, the template 230 may be specific to a particular resource. However, some templates 230 may cover multiple resource types. For example, some templates 230 may be hardware-independent. In step 300.8, the controller powers on the resource 310 through its out-of-band management connection 260 in accordance with the global system rule 210. In step 300.9, using the global system rule 210, the controller finds and loads a boot image for the resource from the selected template(s). The resource 310 is then booted from the image derived from the target template 230 (step 300.10). Additional information about the resource 310 may then be received from the resource 310 through the in-band management connection 270 after the resource 310 has been booted (step 300.11). Such information may include, for example, the firmware version, the network card, and any other devices to which the resource can be connected. In step 300.12, new information may be added to the system state 220. Resource 310 may then be considered to be added to the resource pool and prepared for allocation (step 300.13).
[0135] Regarding Figure 3C, it should be understood that if the resource and controller are on the same node, the service running the resource may be outside that node. In such cases, the controller may use inter-process communication techniques with the resource, such as Unix sockets, loopback adapters, or other inter-process communication techniques to communicate with the resource. According to system rules, the controller may install a virtual host, or a hypervisor or container host from the controller to run the application using a known template. Resource application information can then be added to system state 220, and the resource will be ready for allocation.
[0136] Adding storage resources to the system Figure 4A shows the addition of a storage resource 410 to system 100. In one exemplary embodiment, the exemplary process flow in Figure 3C may be followed to add a storage resource 410 to system 100, where the storage resource 410 being added is not on the same node as the controller 200. It should also be noted that if an image is preloaded onto the storage resource 410, alternative steps may be followed in which any network connection can be used to communicate with the storage resource 410, boot the storage resource 410, and add information to system state 220.
[0137] When a storage resource 410 is added, it may be connected to the controller 200 and powered off. The storage resource 410 is connected to the controller via a network, namely the out-of-band management network 260, the in-band management connection 270, the SAN 280, and optionally, the connection 290. The storage resource 410 may also be connected to one or more application networks 390, from which services, application users, and / or clients can communicate with each other, or may not be connected. Applications or clients may access the resource's storage directly or indirectly through the application, thereby not being accessed through the SAN. Application networks may have built-in storage, or may be accessed and identified as storage resources in the IT system state. The out-of-band management connection 260 may be connected to an independent out-of-band management device 415 or circuit of the storage resource 410, which is turned on when the storage resource 410 is connected. Device 415 may enable features including, but not limited to, powering on / off the device, attaching to and entering commands to the console, monitoring temperature and other computer health-related elements, and setting BIOS settings and other features out of scope from the operating system. Controller 200 may reference storage resources 410 through an out-of-band management network 260. It may identify the type of storage resource and identify its configuration using in-band or out-of-band management. Controller logic 205 is configured to look at out-of-band management 260 or in-band management 270 for additional hardware. If storage resource 410 is detected, controller logic 205 may use global system rules 220 to determine whether resource 410 is configured automatically or through user interaction.If resource 410 is added automatically, the setup follows the global system rules 210 in controller 200. If resource 410 is added by a user, the global system rules 210 in controller 200 may ask the user to confirm the addition of the resource and what the user wants to do with the storage resource. Controller 200 may query the API application(s) to confirm that the new resource is authorized, or otherwise request the user or any program controlling the stack. The authorization process may also be completed automatically and reliably using cryptography to verify the validity of the new resource. Controller logic 205 adds storage resource 410 to the IT system state 220, which includes the switch or network to which the storage resource 410 is connected.
[0138] The controller 200 may power on the storage resource 410 via the out-of-band management network 260, and the storage resource 410 boots from an image 450 loaded from template 230, for example via SAN 280, using the global system rules 210 and controller logic 205. The image may also be loaded indirectly via other network connections or other resources. Once booted, information received via the in-band management connection 270 regarding the storage resource 410 may be collected and added to the IT system state 220. Here, the storage resource 410 is added to the storage resource pool, which is managed by the controller 200 and becomes a resource tracked in the IT system state 220.
[0139] A storage resource may comprise a storage resource pool or multiple storage resource pools that an IT system may use or access, either individually or simultaneously. When a storage resource is added, the storage resource may provide a storage pool, multiple storage pools, a portion of a storage pool, and / or multiple portions of a storage pool to the IT system state. A controller and / or storage resource may manage various storage resources in a pool, or groups of such resources within a pool. A storage pool may include multiple storage pools running on multiple storage resources. For example, a flash storage disk or array caching platter disks or arrays, or a storage pool on a dedicated compute node coupled to a pool on a dedicated storage node, optimizes bandwidth and latency simultaneously.
[0140] Figure 4B shows an image 450 that is loaded directly or indirectly from template 230 (from another resource or database) onto a storage resource 410 to boot the storage resource and / or load an application. Image 450 may include a boot file 440 for the resource type and hardware. The boot file 440 may include a kernel 441 corresponding to the resource, application, or service being deployed. The boot file 440 may include an initrd or similar filesystem used to assist the boot process. The boot system 440 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 450 may include a filesystem 451. Filesystem 451 may include a base image 452 and its corresponding filesystem, a service image 453 and its corresponding filesystem, and a volatile image 454 and its corresponding filesystem. The filesystems and data loaded may vary depending on the resource type and application, or the service being executed. The base image 452 may include a base operating system filesystem. The base operating system may be read-only. Base image 452 may also contain basic operating system tools independent of what is being run. Base image 452 may include a base directory and operating system tools. Service file system 453 may contain configuration files and specifications for resources, applications, or services. Volatile file system 454 may contain information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may or may not be configured as variables including, but not limited to, passwords, session keys, and private keys.The file system may be mounted as a single file system using techniques such as overlayFS, allowing for several read-only and several read-write file systems to reduce the amount of duplicate data used by the application.
[0141] Figure 5A shows an example in which another storage resource, namely direct-attached storage 510, which may take the form of a node with a JBOD or other type of direct-attached storage, is coupled to storage resource 410 as an additional storage resource to the system. A JBOD is typically an external disk array attached to a node that provides storage resources, and in Figure 5A, a JBOD is used as an exemplary form of direct-attached storage 510, but it should be understood that other types of direct-attached storage may be employed as 510.
[0142] The controller 200 may add storage resources 410 and JBOD 510 to its system, for example, as described with respect to Figure 5A. JBOD 510 is connected to the controller 200 via an out-of-band management connection 260. Storage resources 410 are connected to a network, i.e., the out-of-band management connection 260, the in-band management connection 270, the SAN 280, and optionally, connection 290. Storage node 410 communicates with the storage of JBOD 510 through a SAS or other disk drive fabric 520. JBOD 510 may also include an out-of-band management device 515 that communicates with the controller via the out-of-band management connection 260. Through out-of-band management 260, the controller 200 may discover JBOD 510 and storage resources 410. The controller 200 may also discover other parameters not controlled by the operating system, for example, as described herein with respect to various out-of-band management circuits. The controller 200 global system rules 210 provide configuration boot rules for booting or starting up JBODs and storage nodes that have not yet been added. The order in which storage resources are turned on may be controlled by the controller logic 205 using global rules 220. According to one set of global system rules 220, the controller may first power on the JBOD 510, and then the controller 200 may power on the storage resources using the loaded image 450 in a manner similar to that described with respect to Figure 4. In another set of global system rules, the controller 200 may first power on the storage resource 410, and then the JBOD 510. Other global system rules may specify the timing or delay of power-on between different devices. The controller logic 205, global system rules 210, and / or template 230 may determine the readiness or operational state of various resources and / or use this in device allocation management by the controller 200.The IT system state 220 may be updated by communication with the storage resource 410. The storage node 410 is aware of the storage parameters and configuration of the JBOD 510 by accessing the JBOD through the disk fabric 520. The storage resource 410 then provides information to the controller 200, which updates the IT system state 220 with information about the amount of available storage and other attributes. Once the storage resource 410 is booted and recognized as part of the pool of storage resources 400 in system 100, the controller updates the IT system state 220. The storage node handles the logic for controlling the JBOD storage resource using the configuration set by the controller 200. For example, the controller may instruct the storage node to configure the JBOD to create a pool from RAID 10 or other configurations.
[0143] Figure 5B shows an exemplary process flow for adding a storage resource 410 and a directly attached storage 510 for the storage resource 410 to the system 100. In step 500.1, the directly attached storage 510 is connected to the powered-off controller 200 via an out-of-band management connection 260. In step 500.2, the storage resource 410 is connected to the powered-off controller 200 via an out-of-band management connection 260 and an in-band management connection 270, and at the same time, the storage resource 410 is connected to the directly attached storage 510 via a SAS 520, such as a disk drive fabric.
[0144] Next, the controller logic 205 may examine the out-of-band management connection 260 to discover the storage resource 410 and the directly attached storage 510 (step 500.3). Any network connection can be used, but in this example, out-of-band management may be used for the controller logic to recognize and identify the added resources (in this case, the storage resource 410 and the directly attached storage 510) and the type of their configuration (step 500.4).
[0145] In step 500.5, the controller 200 selects and uses a template 230 for a specific type of storage for each type of storage device in order to add resources 410 and 510 to the system state 220. In step 500.6, the controller uses a global system rule 210 that can specify the boot order in which to power on the direct-attached storage and storage nodes in such order via the out-of-band management connection 260 (500.6). Using the global system rule 210, the controller finds and loads a boot image for the storage resource 410 from the template 230 selected for that storage resource 410, and then the storage resource is booted from the image (step 500.7). The storage resource 410 is aware of the storage parameters and configuration of the direct-attached storage 510 by accessing the direct-attached storage 510 through the disk fabric 520. Additional information about the storage resource 410 and / or the direct-attached storage 510 may then be provided to the controller via the in-band management connection 270 to the storage resources (step 500.8). In step 500.9, the controller updates the system state 220 with the information obtained in step 500.8. In step 500.10, the controller configures the storage resource 410 that handles the directly attached storage 510, and the method for configuring the directly attached storage. Then, in step 500.11, the new resource comprising the storage resource 410 together with the directly attached storage 510 may be added to the resource pool and is ready for allocation within the system.
[0146] According to another aspect of the exemplary embodiment, the controller may use out-of-band management to recognize other devices in the stack that do not need to be involved in computing or services. For example, such devices may include, but are not limited to, cooling towers / air conditioning units, lighting, temperature, sound, alarm, power systems, or any other devices associated with the system.
[0147] Adding networking resources to the system Figure 6A shows the addition of a networking resource 610 to system 100. In one exemplary embodiment, the exemplary process flow in Figure 3C may be followed to add a networking resource 610 to system 100, where the networking resource 610 being added is not on the same node as the controller 200. It should also be noted that if an image is preloaded onto the networking resource 610, alternative steps may be followed in which any network connection can be used to communicate with the networking resource 610, boot the networking resource 610, and add information to the system state 220.
[0148] When a networking resource 610 is added, it may be connected to the controller 200 and powered off. The networking resource 610 may be connected to the controller 200 via connections, i.e., out-of-band management connection 260 and / or in-band management connection 270. The networking resource 610 may optionally be connected to SAN 280 and / or connection 290. The networking resource 610 may also be connected to one or more application networks 390 from which services, application users, and / or clients can communicate with each other, or it may not be connected. The out-of-band management connection 260 may be connected to an independent out-of-band management device 615 or circuit for the networking resource 610, which is turned on when the networking resource 610 is connected. Device 615 may enable features including, but not limited to, powering on / off the device, attaching to a console and entering commands, monitoring temperature and other computer health-related elements, and setting BIOS settings and other out-of-scope features from the operating system. The controller 200 may reference the networking resource 610 through an out-of-band management connection 260. It may identify the type of networking resource and / or network fabric, and may identify the configuration using in-band or out-of-band management. The controller logic 205 is configured to look up out-of-band management 260 or in-band management 270 for any additional hardware. If the networking resource 610 is found, the controller logic 205 may use a global system rule 220 to determine whether the networking resource 610 is configured automatically or through user interaction. If the resource 610 is added automatically, the setup follows the global system rule 210 in the controller 200. If it is added by a user, the global system rule 210 in the controller 200 may prompt the user to confirm the addition of the resource and what the user wants to do with it.The controller 200 may query the API application(s) to confirm that the new resource has been authorized, or otherwise request it from the user or any program controlling the stack. The authorization process may also be completed automatically and reliably using cryptography to verify the validity of the new resource. The controller logic 205 may then add the networking resource 610 to the IT system state 220. For switches that cannot be identified by the controller, the user may manually add them to the system state.
[0149] If the networking resource is physical, the controller 200 may power on the networking resource 610 through the out-of-band management connection 260, and the networking resource 610 may boot from an image 605 loaded from template 230, for example via SAN 280, using global system rules 210 and controller logic 205. The image may also be loaded indirectly through other network connections or other resources. Once booted, information received through the in-band management connection 270 regarding the networking resource 610 may be collected and added to the IT system state 220. The networking resource 610 may then be added to the storage resource pool, which becomes a resource managed by the controller 200 and tracked in the IT system state 220. Optionally, some networking resource switches may be controlled through a console port connected to the out-of-band management 260, configured at power-on, or have a switch operating system installed through a boot loader, for example via ONIE.
[0150] If the networking resource is virtual, the controller 200 may power on the networking resource through either the in-band management network 270 or the out-of-band management 260. The networking resource 610 may boot from an image 650 loaded from template 230 via SAN 280 using global system rules 210 and controller logic 205. Once booted, information received through the in-band management connection 270 regarding the networking resource 610 may be collected and added to the IT system state 220. The networking resource 610 may then be added to the storage resource pool, which becomes a resource managed by the controller 200 and tracked in the IT system state 220.
[0151] The controller 200 may instruct networking resources to allocate, reallocate, or move ports to connect to different physical or virtual resources, i.e., connectivity, storage, or computing as defined herein, whether physical or virtual. This may be done using technologies including, but not limited to, SDN, InfiniBand partitioning, VLANs, and vXLAN. The controller 200 may instruct virtual switches to move or assign virtual interfaces to network or interconnect communications with virtual switches or resources hosting virtual switches. Some physical or virtual switches may be controlled by APIs coupled to the controller.
[0152] When such changes are possible, the controller 200 may also instruct compute, storage, or networking resources to change the fabric type. Ports may be configured to switch to different fabrics, for example, to switch fabrics for hybrid InfiniBand / Ethernet interfaces.
[0153] The controller 200 may issue commands to networking resources, which may include switches or other networking resources that switch multiple application networks. The switches or network devices may comprise different fabrics, or, for example, they may be connected to InfiniBand switches, ROCE switches, and / or other switches, preferably by SDN functionality and multiple fabrics.
[0154] Figure 6B shows an image 650 that is loaded directly or indirectly from template 230 (for example, via another resource or database) into a networking resource 610 to boot the networking resource and / or load an application. Image 650 may include a boot file 640 for the resource type and hardware. The boot file 640 may include a kernel 641 corresponding to the resource, application, or service to be deployed. The boot file 640 may include an initrd or similar filesystem used to assist the boot process. The boot system 640 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 650 may include a filesystem 651. Filesystem 651 may include a base image 652 and its corresponding filesystem, a service image 653 and its corresponding filesystem, and a volatile image 654 and its corresponding filesystem. The filesystems and data to be loaded may vary depending on the resource type and application, or the service to be executed. The base image 652 may include a base operating system filesystem. The base operating system may be read-only. Base image 652 may also contain basic operating system tools independent of what is being run. Base image 652 may include a base directory and operating system tools. Service file system 653 may contain configuration files and specifications for resources, applications, or services. Volatile file system 654 may contain information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may or may not be configured as variables including, but not limited to, passwords, session keys, and private keys.The file system may be mounted as a single file system using techniques such as overlayFS, allowing for several read-only and several read-write file systems to reduce the amount of duplicate data used by the application.
[0155] Deployment of applications or services on resources Figure 7A shows a system 100 comprising a controller 200, physical and virtual computing resources including a first compute node 311, a second compute node 312, and a third compute node 313, storage resources 410, and network resources 610. The resources are shown to be set up and added to the IT system state 220 in the manner described herein with respect to Figures 1-6B.
[0156] Although multiple compute nodes are shown in this diagram, a single compute node may be used in one exemplary embodiment. The compute node may host physical or virtual computing resources, and applications may run on the physical or virtual compute node. Similarly, although a single network provider node and storage node are shown, multiple resource nodes of these types may or may not be used in the system of the exemplary embodiment.
[0157] A service or application may be deployed on any of the systems according to an exemplary embodiment. An example of deploying a service on a compute node may be described with respect to Figure 7A, but may be used similarly in different configurations of system 100. For example, the controller 200 in Figure 7A may automatically configure compute resources 310 in the form of compute nodes 311, 312, and 313 according to the global system rule 210. They may then be added to the IT system state 220. Thus, the controller 200 may be aware of compute resources 311, 312, and 313 (which may be powered off or not powered off), and optionally any physical or virtual applications running on the compute resources or nodes. The controller 200 may also automatically configure storage resources 410 and networking resources 610 according to the global system rule 210 and template 230, and add them to the IT system state 220. The controller 200 may be aware of storage resources 410 and networking resources 610 which may be started in a powered-off state or not started.
[0158] Figure 7B illustrates an exemplary process for adding a resource to IT system 100. In step 700.1, the new physical resource is linked to the system. In step 700.2, the controller becomes aware of the new resource. The resource may be connected to remote storage (step 700.4). In step 700.3, the controller configures how to boot the new resource. All connections made to the resource can be recorded in system state 220 (step 700.5). Figure 3C discussed above provides further details on exemplary embodiments of the process flow, such as that shown in Figure 7B.
[0159] Figures 7C and 7D show an exemplary process flow for deploying an application to multiple computing resources, multiple servers, multiple virtual machines, and / or multiple sites. The process for this example differs from a standard template deployment in that the IT system 100 requires components that link redundant and interrelated applications and / or services. The controller logic may process a meta-template in step 700.11, where the meta-template may include multiple templates 230, a file system blob 232, and other components (which may be in the form of other templates 230) required to configure a multi-homed service.
[0160] In step 700.12, the controller logic 205 checks the system state 220 of available resources. However, if there are not enough resources, the controller logic may reduce the number of redundant services that may be deployed (see 700.16, where the number of redundant services is identified). In step 700.13, the controller logic 205 configures the networking resources and interconnections required to connect the services. If a service or application is deployed across multiple sites, the meta-template may include services optionally configured from templates that enable data synchronization and interoperability across sites (or the controller logic 205 may configure them) (see 700.15).
[0161] In step 700.16, the controller logic 205 may determine the number of redundant services from system rules, meta-template data, and resource availability (if the redundant services reside on multiple hosts). In 700.17, it connects with other redundant services and connects with the master. If there are multiple redundant hosts, the controller logic 205 or logic (which may include binary 234, daemon 232, or a filesystem blob that instructs the operating system to configure) in the template can prevent network address and hostname conflicts. Optionally, the controller logic provides network addresses (see 700.18) and registers each redundant service with DNS (700.19) and system state 220 (700.18). System state 220 tracks redundant services, and if the controller logic 205 detects a redundant service with conflicting parameters such as hostname or DNS name, it will not allow duplicate registration, and the network address is already in system state 220.
[0162] The configuration routine shown in Figure 7D processes the template(s) of the metatemplate. The configuration routine handles all redundant services, deploys multi-host or clustered services across multiple hosts, and deploys services that link hosts. Any process capable of deploying an IT system from system rules can execute the configuration routine. For multi-host services, the exemplary routine may process the service template as in 700.32, provision storage resources as in 700.33, power on the hosts as in 700.35, and link the host / computing resources with the storage resources (and register them in system state 220) as in 700.36 (and then repeat for the number of redundant services (700.38)). Each time, it registers in system state 220 (see 700.20) and uses controller logic to track individual services and record information to prevent conflicts (see 700.31).
[0163] Some service templates may include services and tools that can link multi-host services. Some of these services may be treated as dependencies (700.39), and then the linking routines in 700.40 may be used to link the services and register the linking to system state 220. Furthermore, one of the service templates may be a master template, in which case the dependent service templates in 700.39 are slave or secondary services, and the linking routines in 700.40 connect them. The routines can be defined in a meta-template. For example, for a redundant DNS configuration, the linking routines in 700.40 may include connecting the slave DNS to the master DNS and configuring zone transfers with DNSsec. Some services may use physical storage (see 700.34) to improve performance, and it may be loaded with the backup OS disclosed in Figure 5B. Tools for linking services may be included in the template itself, and configuration between services may be done by APIs accessible by the controller and / or other hosts in the multi-node application / service.
[0164] The controller 200 may enable the user or controller to determine the appropriate compute backend to use for the application. The controller 200 may enable the user or controller to optimally place the application on the appropriate physical or virtual compute resources by determining resource usage. When hypervisors or other compute backends are deployed to compute nodes, they may report controller resource usage statistics back through the inband management connection 270. When the controller decides to create an application on virtual compute resources, either from its own logic and global system rules or from user input, it may automatically select the most suitable hypervisor on the host and power on the virtual compute resources on that host.
[0165] For example, the controller 200 uses a template(s) 230 to deploy an application or service to one or more computing resources. Such an application or service may be, for example, a virtual machine running the application or service. In one example, Figure 7A shows the deployment of multiple virtual machines (VMs) on multiple compute nodes, and the controller 200 shown can recognize that multiple computing resources 310 are in its computing resource pool in the form of compute nodes 311, 312, 313. The compute nodes may, for example, have a hypervisor deployed on them, or they may instead be on bare metal where the use of virtual machines may be undesirable for speed reasons. In this example, the computing resource 310 has VM(1) 321 and VM(2) 322 configured and deployed on compute node 311 with a hypervisor application loaded on them. For example, if compute node 311 does not have resources for an additional VM, or if other resources are preferred for a particular service, the controller 200 may recognize, based on the stack state 220, that there are no available resources on compute node 311, or that it is preferable to prepare a new VM on different resources. It may also recognize that the hypervisor is loaded onto compute resource 312 and not onto resource 313, which may be, for example, a bare-metal compute node used for other purposes. Thus, according to the requirements of the installed service or application template and the status of the system state 220, the controller in this example may select compute node 313 for the deployment of the next required resource VM(3)323.
[0166] The system's computing resources may be configured to share storage resources on top of the storage resources for the storage nodes.
[0167] A user may request, through the user interface 110 or an application, that services be set up for system 100. These services may include, but are not limited to, email services, web services, user management services, network providers, LDAP, Dev tools, VoIP, authentication tools, and accounting.
[0168] The API application 120 translates user or application requests and sends messages to the controller 200. The service template or image 230 of the controller 200 is used to identify which resources are required for the service. The resources to be used are then identified based on their availability according to the IT system state 220. The controller 200 makes requests for one or more of the compute nodes 311, 312, or 313 for the required compute service, for the storage resource 410 for the required storage resource, and for the network resource 610 for the required networking resource. The IT system state 220 is then updated to identify the resources to be allocated. The service is then installed on the allocated resources using the global system rule 210, according to the template 230 for the service or application.
[0169] According to an exemplary embodiment, multiple compute nodes may be used for the same service or for different services, while, for example, a storage service and / or a network provider pool may be shared among the compute nodes.
[0170] Referring to Figure 8A, System 100 is shown, where the controller 200, as well as the computing resources 300, storage resources 400, and networking resources 600, reside on the same or shared physical hardware, such as a single node. The various features shown and described in Figures 1-10 may be incorporated into a single node. When the node is powered on, the controller image is loaded onto the node. The computing resources 300, storage resources 400, and networking resources 600 are configured by template 230 using global system rules 210. The controller 200 may be configured to load compute backends 318, 319 as compute resources, which may or may not be added on a node or different nodes. Such backends 318, 319 may include, but are not limited to, virtualization, containers, and multitenant processes to create virtual computing resources, networking resources, and storage resources.
[0171] Applications or services 725, such as web, email, core network services (DHCP, DNS, etc.), and collaborative tools, may be installed on virtual resources on a node / device shared with the controller 200. These applications or services may be moved to physical or virtual resources independent of the controller 200. Applications may run on virtual machines on a single node.
[0172] Figure 8B shows an exemplary process flow for extending a system from a single node to a multi-node system (for example, by nodes 318 and / or 319 as shown in Figure 8A). Referring to Figures 8A and 8B, we can consider an IT system with a controller 200 running on a single server, and it is desirable to extend the IT system to a multi-node IT system. Therefore, before extension, the IT system is in a single-node state. As shown in Figure 8A, the controller 200 runs on a multi-tenant single-node system to run various IT system management applications and / or resources, which may include, but are not limited to, storage resources, computing resources, a hypervisor, and / or container hosts.
[0173] In step 800.2, the new physical resource is linked to a single node system by connecting it through the out-of-band management connection 260, the in-band management connection 270, the SAN 280, and / or the network 290. For the purposes of this example, this new physical resource can also be referred to as hardware or a host. The controller 200 may discover the new resource on the management network and then query the device. Alternatively, the new device may broadcast a message to inform the controller 200 of its presence. For example, the new device can be identified by its MAC address, out-of-band management, and / or the use of booting a backup OS and in-band management, and thereby the identification of its hardware type. In any case, in step 800.3, the new device provides the controller with information about its node type and its currently available hardware and software resources. The controller 200 then recognizes the new device and its capabilities.
[0174] In step 800.4, tasks assigned to the system running controller 200 may be assigned to the new host. For example, if the host is preloaded with an operating system (such as a storage host operating system or hypervisor), controller 200 assigns new hardware resources and / or capabilities. The controller may then provision the new hardware by providing an image, or the new hardware may request an image from the controller and configure itself using the methods disclosed above and below. If the new host can host storage resources or virtual computing resources, the new resources may be made available to controller 200. The controller 200 may then move and / or assign existing applications to the new resources, or use the new resources for newly created applications or applications created thereafter.
[0175] Step 800.5 allows the IT system to retain its current applications running on the controller, or to migrate them to new hardware. When migrating virtual computing resources, VM migration techniques (e.g., the QEMU+KVM migration tool) may be used to update the system state along with the new system rules. The change management techniques discussed below can be used to ensure these changes are carried out reliably and securely. As more applications may be added to the system, the controller may use any of the various techniques to determine how to allocate system resources, including but not limited to round-robin techniques, weighted round-robin techniques, least utilization techniques, weighted least utilization techniques, predictive techniques with utilization-assisted training, planning techniques, desired capacity techniques, and maximum size techniques.
[0176] Figure 8C shows an exemplary process flow for migrating a storage resource to a new physical storage resource. The storage resource may then be mirrored, migrated, or a combination thereof (for example, the storage may be mirrored, and then the original storage resource is disconnected). In step 820, the storage resource is connected to the system by either a new storage resource that is in contact with the controller or that causes the controller to discover it. This can be done via an out-of-band management connection 260, an in-band management connection 270, a SAN network 280, or, in a flat network, an application network, or a combination thereof. By in-band management, the operating system may be pre-booted and the new resource may be connected to the controller.
[0177] In step 822, a new storage target is created on a new storage resource, which can be recorded in the database in step 824. In one example, a storage target may be created by copying files. In another example, a storage target may be created by creating a block device and copying data (which may be in the form of filesystem blobs). In yet another example, a storage target may be created by mirroring two or more storage resources between block devices (for example, by creating a RAID array) and optionally connecting them through a remote storage transport (which may include, but is not limited to, iSCSI, iser, nvmeof, nfs, nfs over rdma, fc, fcoe, srp, etc.). The database entry in step 824 may include information for computing resources (or other types of resources and / or hosts) to connect to the new storage resource remotely or locally, if the storage resource resides on the same device as other resources or hosts.
[0178] In step 826, the storage resources are synchronized. For example, the storage can be mirrored. As another example, the storage can be synchronized offline. Techniques such as RAID1 (or other types of RAID, usually RAID1 or RAID0, but may be RAID110 (mirrored RAID10)) (MDADM, ZFS, BTRFS, hardware RAID) may be employed in step 826.
[0179] Next, data from the old storage resource is optionally connected after the database login in step 828 (if such data must be recorded when it is done later, the database may include information relating to the status of copying the data). If the storage target is being migrated away from the previous host (for example, moving from a single-node system to a multi-node and / or distributed IT system, as shown in Figures 8A and 8B previously), the new storage resource may be referred to in step 830 as the primary storage resource, by controller, system state, computing resources, or a combination thereof. This may be done as a step to delete the old storage resource. In some cases, it may then be necessary to update the physical or virtual host connected to the resource, and in some cases, in step 832, the power may be turned off during the migration (and then turned back on) (technologies disclosed herein for powering on the physical or virtual host may be used).
[0180] Figure 8D shows an exemplary process flow for migrating virtual machines, containers, and / or processes on a single node of a multitenant system to a multinode system which may have separate hardware for compute and storage. In step 850, the controller 200 creates new storage resources which may be on the new node (see, for example, nodes 318 and 319 in Figure 8A). Next, in step 852, the old application host may be powered off. Then, in step 854, the data is copied or synchronized. Powering off in step 852 before copying / synchronizing in step 854 would make the migration safer if the migration involves migrating VMs from a single node. Powering off would also be beneficial for moving VMs to physical machines. Step 854 may be performed before powering off via data pre-synchronization step 862, which can minimize the associated downtime. Furthermore, the host does not have to be powered off as in step 852, in which case the old host remains online until the new host (or new storage resources) is ready. Techniques to avoid power-off step 852 will be discussed in more detail later. In step 854, data can be optionally synchronized unless the storage resources are mirrored or synchronized using hot standby.
[0181] Here, the new storage resource is operational and may be logged into the database in step 856, so that the controller 200 can connect the new host to the new storage resource in step 858. When multiple virtual hosts are migrated from a single node, this process may need to be repeated for multiple hosts (step 860). If they are tracked, the boot order may be determined by the controller logic using application dependencies.
[0182] Figure 8E shows another exemplary process flow for scaling a system from a single node to multiple nodes. In step 870, the new resource is linked to the single-node system. The controller may have a set of system rules and / or extension rules for the system (or it may have extension rules based on service execution, their templates, and the interdependence of services). In step 872, the controller checks such rules to facilitate the scaling.
[0183] If the new physical resources include storage resources, the storage resources may be moved in step 874 from a single node or other form of a simpler IT system (or the storage resources may be mirrored). If the storage resources are moved, the computing resources or running resources may be reloaded or rebooted in step 876 after the storage resources have been moved. In another example, the computing resources may be connected to a mirrored storage resource in step 876 and may remain running, while the old storage resources on a single-node system or hardware resources from the previous system may be disconnected or disabled. For example, a running service may be coupled to two mirrored block devices (one on a single-node server (e.g., using mdadm raid1) and the other on a storage resource), and then, once the data is synchronized, the drive on the single-node server may be disconnected. The previous hardware may remain as part of the IT system and may run on the same node as a controller in mixed mode (step 878). The system may continue this migration process until the original node is running only the controller, at which point the system is distributed (step 880). Furthermore, at each step of the process flow in Figure 8E, the controller can update the system state 220 and record the changes to the system in the database (step 882).
[0184] Referring to Figure 9A, application 910 is installed on resource 900. Resource 900 may be a computing resource 310, a storage resource 410, or a networking resource 610 as described in Figures 1-10 herein. Resource 900 may be a physical resource. A physical resource may comprise a physical machine or a physical IT system component. Resource 900 may be, for example, a physical computing resource, a physical storage resource, or a physical networking resource. Resource 900 may be connected to the controller 200 of system 100 by other resources such as computing resources, networking resources, or storage resources as described in Figures 2A-10 herein.
[0185] Resource 900 may initially be powered off. Resource 900 may be connected to the controller via a network, i.e., an out-of-band management connection 260, an in-band management connection 270, a SAN 280, and / or network 290. Resource 900 may also be connected to one or more application networks 390, from which services, application users, and / or clients can communicate with each other. The out-of-band management connection 260 may be connected to an independent out-of-band management device 915 or circuit for Resource 900, which is turned on when Resource 900 is connected. The device may enable features including, but not limited to, powering the device on / off, attaching to and entering commands on a console, monitoring temperature and other computer health-related elements, and setting BIOS settings 195 and other features out of scope from the operating system.
[0186] The controller 200 may discover resource 900 through the out-of-band management network 260. It may identify the type of resource and identify its configuration using in-band or out-of-band management. The controller logic 205 may be configured to look at out-of-band management 260 or in-band management 270 for additional hardware. If resource 900 is discovered, the controller logic 205 may use global system rules 220 to determine whether resource 900 is configured automatically or through user interaction. If resource 900 is added automatically, the setup follows global system rules 210 within the controller 200. If resource 900 is added by a user, global system rules 210 within the controller 200 may ask the user to confirm the addition of the resource and what the user wants to do with the computing resource. The controller 200 may query the API application to confirm that the new resource has been authorized, or otherwise request it from the user or any program controlling the stack. The authorization process may also be completed automatically and reliably using cryptography to verify the validity of the new resource. Next, resource 900 is added to the IT system state 220, which includes the switch or network to which resource 900 is connected.
[0187] The controller 200 may power on the resources through the out-of-band management network 260. The controller 200 may use the out-of-band management connection 260 to power on the physical resources and configure the BIOS 195. The controller 200 may automatically use the console 190 to select the desired BIOS options, which may be achieved by the controller 200 reading the console image with image recognition and controlling the console 190 through out-of-band management. The boot-up state may be determined by image recognition through the console of resource 900, out-of-band management with a virtual keyboard, querying the services using the resources, or querying the services of application 910. Some applications may have processes that allow the controller 200 to monitor or, optionally, modify the settings of application 910 using in-band management 270.
[0188] The application 910 on the physical resource 900 (or on resources 300, 310, 311, 312, 313, 400, 410, 411, 412, 600, 610 as described with respect to Figures 1-10 of this specification) may be booted over SAN280 or another network using a BIOS boot option or other method of configuring remote booting, such as enabling PXE booting or Flex booting. Furthermore, the controller 200 may use out-of-band management 260 and / or in-band management connections 270 to instruct the physical resource 900 to boot the application image of image 950. The controller may configure boot options on the resource, or use existing available remote booting methods such as PXE booting or Flex booting. The controller 200 may optionally or alternatively use out-of-band management 260 to boot from an ISO image, configure a local disk, and then instruct the resource to boot from the local disk(s) 920. Local disks may be loaded with boot files. This may be achieved by using out-of-band management 260, image recognition, and a virtual keyboard. Resources may also have boot files and / or a boot loader installed. Resources 900 and applications may boot from image 950 loaded from template 230, for example via SAN 280, using global system rules 210 and controller logic 205. Global system rules 220 may specify the boot order. For example, global system rules 220 may require that resource 900 be booted first, followed by application 910. Once resource 900 is booted using image 950, information received through the in-band management connection 270 regarding resource 900 may be collected and added to the IT system state 220.Resource 900 may be added to the storage resource pool, which will be managed by the controller 200 and tracked in the IT system state 220. Application 910 may be booted in the order specified in the global system rule 220 using image 950 or application image 956 loaded on resource 900.
[0189] The controller 200 may configure a networking resource 610 that connects the application 910 to the application network 390 via an out-of-band management connection 260 or another connection. The physical resource 900 may be connected to remote storage such as block storage resources including but not limited to ISER (ISCSI over RDMA), NVMEOF FCOE, FC, or iSCSI, or to another storage backend such as SWIFT, GFUSTER, or CEPHFS. When a service or application is operational, the IT system state 220 may be updated using the out-of-band management connection 260 and / or the in-band management connection 270. The controller 200 may use the out-of-band management connection 260 or the in-band management connection 270 to determine the power state of the physical resource 900, i.e., whether it is on or off. The controller 200 may use the out-of-band management connection 260 or the in-band management connection 270 to determine whether a service or application is running or in a boot-up state. The controller may perform other functions based on the information it receives and the global system rule 210.
[0190] Figure 9B shows image 950, which is loaded directly or indirectly from template 230 to a compute node (for example, via another resource or database) to boot application 910. Image 950 may include a custom kernel 941 for application 910.
[0191] Image 950 may include a boot file 940 for resource types and hardware. The boot file 940 may include a kernel 941 corresponding to the resources, applications, or services to be deployed. The boot file 940 may include an initrd or similar filesystem used to assist the boot process. The boot system 940 may include multiple kernels or initrds configured for different hardware and resource types. Furthermore, image 450 may include a filesystem 951. The filesystem 951 may include a base image 952 and its corresponding filesystem, a service image 953 and its corresponding filesystem, and a volatile image 954 and its corresponding filesystem. The filesystems and data to be loaded may vary depending on the resource type and application, or the service to be executed. The base image 952 may include a base operating system filesystem. The base operating system may be read-only. The base image 952 may also include basic operating system tools independent of what is being run. The base image 952 may include a base directory and operating system tools. The service filesystem 953 may include configuration files and specifications for resources, applications, or services. A volatile filesystem 594 may contain information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may or may not be configured as variables including, but not limited to, passwords, session keys, and private keys. The filesystem may be mounted as a single filesystem using a technique such as overlayFS, allowing several read-only and several read-write filesystems to reduce the amount of duplicate data used by the application.
[0192] Figure 9C shows an example of installing an application from an NT package, which can be one type of template 230. In step 900.1, the controller determines that a package blob needs to be installed. In step 900.2, the controller creates a storage resource on the default datastore for the blob type (block, file, filesystem). In step 900.3, the controller connects to the storage resource via an available storage transport for the storage resource type. In step 900.4, the controller copies the package blob to the connected storage resource. The controller then disconnects from the storage resource (step 900.5) and sets the storage resource to read-only (step 900.6). The package blob is then successfully installed (step 900.7).
[0193] In another example, Annex B provides illustrative details on how a system connects computing resources to an overlayfs. Such techniques can be used to facilitate installing an application on a resource according to Figure 9A, or launching a computing resource from a storage resource according to step 205.11 in Figure 2F.
[0194] Figure 9D shows application 910 deployed on resource 900. Resource 900 may comprise virtual computing resources, for example, a hypervisor 920, one or more virtual machines 921, 922, and / or a compute node which may comprise containers. Resource 900 may be configured using an image 950 loaded onto resource 900 in a manner similar to that described herein with respect to Figures 1 to 10. In this example, resource 920 is shown as a hypervisor managing virtual machines 921, 922. Controller 200 may use inband management 270 to communicate with resource 900 hosting hypervisor 920 which creates the resources, configure the resources, and allocate appropriate hardware resources including, but not limited to, CPU, RAM, GPU, remote GPU (which may use RDMA to remotely connect to another host), network connectivity, network fabric connectivity, and / or virtual and physical connectivity to a partitioned and / or segmented network. The controller 200 may use a virtual console 190 (including, but not limited to, SPICE or VNC) and image recognition to control the resource 900 and the hypervisor 920. Additionally or alternatively, the controller 200 may use out-of-band management 260 or in-band management connection 270 to instruct the hypervisor 920 to boot the application image 950 from the template 230 using global system rules 210. The image 950 may be stored on the controller 200, or the controller 200 may move or copy them to the storage resource 410.The boot image for VM921 and 922 may be stored locally as, for example, image 950, or as a file on a block device or remote host, and may be shared by a file share such as NFS over RDMA / NFS using an image type such as qcow2 or raw, or it may use a remote block device using iSCSI, ISER, NVMEOF, FC, or FCOE. A portion of image 950 may be stored on storage resource 410 or compute node 310. Controller 200 may configure networking resource 610 to appropriately support the application via out-of-band management connection 260 or another connection using global rules and / or templates. The application 910 on resource 900 may boot using an image 950 loaded by SAN280 or another network using a BIOS boot option, or by enabling the hypervisor 920 on resource 900 to connect to a block storage resource including but not limited to ISER (ISCSI over RDMA), NVMEOF FCOE, FC, or iSCSI, or to another storage backend such as SWIFT, GFUSTER, or CEPHFS. The storage resource may be copied from a template target on the storage resource. The IT system state 220 may be updated by querying the hypervisor 920 for information. The in-band management connection 270 may communicate with the hypervisor 920 and may be used to determine the power state of the resource, i.e., whether it is on or off, or to determine the boot-up state. The hypervisor 920 may use a virtual in-band connection 923 to the virtualized application 910, or it may use the hypervisor 920 for functions similar to out-of-band management. This information may indicate whether a service or application is operational because it is powered on or booted.
[0195] The boot-up state may be determined by image recognition through the console 190 of resource 900, out-of-band management 260 via a virtual keyboard, querying services utilizing the resource, or querying services of application 910 itself. Some applications may have processes that allow the controller 200 to monitor or, optionally, modify the settings of application 910 using in-band management 270. Some applications may reside on virtual resources, and the controller 200 may monitor them by communicating with the hypervisor 920 using in-band management 270 (or out-of-band management 260). Application 910 may not have such processes for monitoring and / or adding input (or such processes that may be switched to to save resources). In such cases, the controller 200 may use the out-of-band management connection 260 and use image processing and / or a virtual keyboard to log on to the system to make changes and / or switch over management processes. A virtual machine console 190 may be used, as with virtual computing resources.
[0196] Figure 9E shows an exemplary process flow for adding a virtual computing resource host to IT system 100. In step 900.11, a host available for use as a virtual computing resource is added to the system. The controller may configure a bare metal server according to the process flow in Figure 15B (step 900.12). Alternatively, the operating system may be preloaded and / or the host may be preconfigured (step 900.13). The resource is then added to system state 220 as a virtual computing resource pool (step 900.14), and the resource becomes accessible by an API from controller 200 (step 900.15). The API is typically accessed through an in-band management connection 270. However, the in-band management connection 270 may be selectively enabled and / or disabled on the virtual keyboard. The controller may then use the out-of-band management connection 260, along with the virtual keyboard and monitor, to communicate through the out-of-band connection 260 (step 900.16). In step 900.17, the controller can then utilize the new resource as a virtual computing resource.
[0197] Exemplary Multi-Controller System Referring to Figure 10, a system 100 is shown having computing resources 300, 310 as described herein with respect to Figures 1-10, comprising a plurality of physical computing nodes 311, 312, 313; storage resources 400, 410 as described herein in the form of a plurality of storage nodes 411, 412 and JBOD 413; a plurality of controllers 200a, 200b, comprising components 205, 210, 220, 230 (Figures 1-9C) and configured as controllers 200 as described herein; networking resources 600, 610, comprising a plurality of fabrics 611, 612, 613 and described herein; and an application network 390.
[0198] Figure 10 shows possible configurations of the components of system 100 in an exemplary embodiment, but does not limit the possible configurations of the components of system 100.
[0199] The user interface or application 110 communicates with an API application 120 that communicates with either or both of the controllers 200a or 200b. The controllers 200a and 200b may be coupled to an out-of-band management connection 260, an in-band management connection 270, a SAN 280, or a network in-band management connection 290. As described with reference to Figures 1 to 9C of this specification, the controllers 200a and 200b are coupled to the compute nodes 311, 312, 313, storage 411, 412 including JBOD 413, and networking resources 610 via connections 260, 270, 280, and optionally 290. The application network 390 is coupled to the compute nodes 311, 312, 313, storage resources 411, 412, 413, and networking resources 610.
[0200] Controllers 200a and 200b may operate in parallel. Either controller 200a or 200b may initially operate as the master controller 200, as described with respect to Figures 1 to 9C of this specification. Controllers 200a and 200b may be configured to configure the entire system 100 from a power-off state. One of controllers 200a and 200b may further generate system state 220 from the existing configuration by exploring other controllers through either out-of-band or in-band connections 260 and 270. Either controller 200a or 200b may access or receive resource status and related information from resources or other controllers through one or more connections 260 and 270. Controllers or other resources may update other controllers. Therefore, when an additional controller is added to the system, this controller may be configured to return system 100 to system state 220. If one of the controllers or the master controller fails, another controller may be designated as the master controller. The IT system state 220 may further be reconstructible from available or stored status information on a resource. For example, an application may be deployed on a computing resource on which the application is configured to create a virtual computing resource on which the system state is stored or replicated. The global system rules 210, system state 220, and template 230 may further be stored or copied on a resource or combination of resources. Thus, if all controllers go offline and a new controller is added, the system may be configured so that the new controller can recover the system state 220.
[0201] The networking resource 610 may comprise multiple network fabrics. For example, as shown in Figure 10, the multiple network fabrics may include one or more of the SDN Ethernet switch 611, ROCE switch 612, Infiniband switch 613, or other switches or fabrics 614. The hypervisor utilizes one or more of the required fabrics to provide virtual machines on compute nodes that can connect to physical or virtual switches. The network configuration may allow restrictions on the physical network, for example, through a segmented network, for security or other resource optimization purposes.
[0202] System 100, through the controller 200 as described in Figures 1 to 10 of this specification, may automatically set up services or applications. A user may request that a service be set up for System 100 through the user interface 110 or an application. Services may include, but are not limited to, email services, web services, user management services, network providers, LDAP, developer tools, VoIP, authentication tools, and accounting software. An API application 120 translates user or application requests and sends messages to the controller 200. A service template or image 230 of the controller 200 is used to identify the resources required for a service. The required resources are identified based on their availability according to the system state 220. The controller 200 requests computing resource 310 or computing nodes 311, 312, or 313 for the required computing services, storage resource 410 for the required storage resources, and networking resource 610 for the required networking resources. The system state 220 is then updated to identify the resources that will be allocated. Next, the service is installed on the allocated resources using global system rule 210, according to the service template.
[0203] System security improvement Referring to Figure 13A, IT system 100 is shown, where system 100 includes resource 1310, which may be bare metal or physical resources. Although Figure 13A shows only a single resource 1310 connected to system 100, it should be understood that system 100 may include multiple resources 1310. Resource 1310(or more) may be or may comprise a bare metal cloud node. A bare metal cloud node may include, but is not limited to, resources connected to an external network 1380 that allows remote access to physical hosts or virtual machines, allows the creation of virtual machines, and allows external users to execute code on the resource(or more). Resource 1310(or more) may be directly or indirectly connected to the external network 1380 or application network 390. The external network 1380 may be the Internet or other resources(or more) not managed by controller 200 or the controller of IT system 100. External network 1380 may include, but is not limited to, the Internet, Internet connectivity(s), resources(s) not managed by the controller, other wide-area networks (e.g., Stratcom, peer-to-peer mesh networks, or other external networks, whether publicly accessible or not), or other networks.
[0204] When the physical resource 1310 is added to the IT system 100a, it is coupled to the controller 200 and may be powered off. The resource 1310 is coupled to the controller 200a via one or more networks, such as an out-of-band management (OOBM) connection 260, optionally an in-band management (IBM) connection 270, and optionally a SAN connection 280. The SAN 280 as used herein may or may not include a configuration SAN. The configuration SAN may include a SAN used to power on or configure the physical resource. The configuration SAN may be part of the SAN 280 or separate from the SAN 280. In-band management may further include a configuration SAN, which may or may not be the SAN 280, as shown herein. Furthermore, if the resource is in use, the configuration SAN may be disabled, disconnected, or unavailable. The OOBM connection 260 is not visible to the OS of system 100, but the IBM connection 270 and / or the configuration SAN may be visible to the OS of system 100. The controller 200 in Figure 13A may be configured similarly to the controller 200 described herein with reference to Figures 1 to 12B. Resource 1310 may have internal storage. In some configurations, the controller 200 may generate storage and temporarily configure the resource to connect to a SAN for fetching data and / or information. The out-of-band management connection 260 may be coupled to a separate out-of-band management device 315 or to the circuitry of resource 1310 which is powered on when resource 1310 is plugged in. Device 315 may enable functions including, but not limited to, powering on / off the device, connecting to a console and entering commands, monitoring temperature and other computer health-related elements, and setting BIOS settings and other functions outside the scope of the operating system. The controller 200 can reference resource 1310 through the out-of-band management network 260. The controller may also identify the type of resource and its configuration using in-band or out-of-band management.Figures 13C to 13E, described below, illustrate various process flows for adding physical resource 1310 to IT system 100a and / or for starting or managing system 100 in a way that enhances system security.
[0205] As used herein with reference to a network, networking resource, network device, and / or network interface, the term “disabled” means that such a network, networking resource, network device, and / or network interface is powered off (manually or automatically), physically disconnected, and / or virtually disconnected, or otherwise disconnected from a network, virtual network (including, but not limited to, VLANs, VXLANs, and InfiniBand partitions) by several other means (e.g., filtering). The term “disabled” also includes restrictions on the unidirectional or one-way operation, such as preventing a resource from sending or writing data to a destination (while having the ability to receive or read data from a source), or preventing a resource from receiving or reading data from a source (while having the ability to send or write data to a destination). Such a network, networking resource, network device, and / or network interface may be disconnected from an additional network, virtual network, or resource combination and remain connected to a previously connected network, virtual network, or resource combination. In addition, such a networking resource or device may be switched from one network, virtual network, or resource combination to another.
[0206] As used herein with reference to a network, networking resource, network device, and / or network interface, the term “enabled” means that such a network, networking resource, network device, and / or network interface is powered on (manually or automatically), physically connected, and / or virtually connected, or connected in several other ways to a network, virtual network (including, but not limited to, VLANs, VXLANs, and InfiniBand partitions). Such a network, networking resource, network device, and / or network interface may be connected to an additional network, virtual network, or resource combination if it is already connected to another system component. In addition, such a networking resource or device may be switched from one network, virtual network, or resource combination to another. The term “enabled” also includes one-way or unidirectional restrictions on operation, such as enabling a resource to send, write, or receive data to a destination (while having the ability to restrict data from the source), or enabling a resource to send, receive, or read data from a source (while having the ability to restrict data from the destination).
[0207] The controller logic 205 is configured to check the out-of-band management connection 260 or in-band management connection 270 and / or configuration SAN 280 of the added hardware. If resource 1310 is found, the controller logic 205 may use global system rules 220 to decide whether to configure the resource automatically or to configure it by interacting with the user. If added automatically, the setup will follow global system rules 210 in the controller 200. If added by a user, global system rules 210 in the controller 200 may query the user to add the resource and what the user wants to do with resource 1310. The controller 200 may query the API application or otherwise request the user or any program controlling the stack to verify that the new resource is authenticated. The authentication process can also be completed automatically and securely using cryptography to verify the legitimacy of the new resource. The controller logic 205 then adds resource 1310 to the IT system state 220, which includes the switch or network to which resource 1310 is plugged in.
[0208] If the resource is physical, the controller 200 may power on the resource via the out-of-band management network 260, and resource 1310 can boot off image 350 loaded from template 230, for example via SAN 280, using the global system rules 210 and controller logic 205. The image can be loaded indirectly via other network connections or through another resource. Once booted, information about resource 1310 can also be collected and added to the IT system state 220. This can be done via in-band management and / or configuration SAN or out-of-band management connections. Resource 1310 can boot off image 350 loaded from template 230, for example via SAN 280, using the global system rules 210 and controller logic 205. The image can be loaded indirectly via other network connections or through another resource. Once booted, information received via the in-band management connection 270 regarding computing resource 310 can also be collected and added to the IT system state 220. Next, resource 1310 may be added to the storage resource pool, which becomes a resource managed by controller 200 and tracked in IT system state 220.
[0209] The in-band management and / or configuration SAN may be used by the controller 200 to set up, manage, use or communicate with resource 1310 and to execute any commands or tasks. However, optionally, the in-band management connection 270 may be configured by the controller 200 to be turned off or disabled at any time or during the setup, management, use or operation of system 100 or controller 200. In-band management may be further configured to be turned on or enabled at any time or during the setup, management, use or operation of system 100 or controller 200. Optionally, the controller 200 may disconnect resource 1310 from the in-band management connection 270 to the controller 200(s). Such disconnection or disconnectibility may be physical, for example, by using an automatic physical switch or a switch that turns off power to the in-band management connection and / or configuration SAN of the resource to the network. For example, disconnection may be performed by a network switch that cuts off power to the port of resource 1310 connected to the in-band management 270 and / or configuration SAN 280. Such disconnections or partial disconnections may be performed using a software-defined network or by physically filtering the controller using a software-defined network. Such disconnections may be performed via the controller through either in-band or out-of-band management. According to an exemplary embodiment, resource 1310 can be disconnected from the in-band management connection 270 in response to a selective control command from the controller 200 at any point before, during, or after resource 1310 is added to the IT system.
[0210] Using a software-defined network, the in-band management connection 270 and / or configuration SAN 280 may or may not retain certain functionalities. The in-band management 270 and / or configuration SAN 280 may be used as a restricted connection for communication with the controller 200 or other resources. Connection 270 may be restricted to prevent an attacker from pivoting to the controller 200, other networks or other resources. The system may be configured to prevent devices such as the controller 200 and resource 1310 from communicating openly and thus prevent the resource 1310 from being compromised. For example, in the in-band management 270 and / or configuration SAN 280, data transmission only from the in-band management and / or configuration SAN may be permitted and any reception prohibited through a software-defined network or a hardware modification method (such as electronic restrictions). The in-band management and / or configuration SAN may be configured to be a one-way write component, or as a one-way write connection from the controller 200 to resource 1310, either physically or using a software-defined network that only allows writes from the controller to the resource. The one-way write nature of the connection can be further controlled or turned on or off depending on the desired security situation and various stages or times in the system's operation. The system can also be configured to restrict writes or communications from resources to the controller, for example, to communicate logs or alerts. Interfaces can further be moved to or added to or removed from other networks by technologies including, but not limited to, software-defined networks, VLANs, VXLANs, and / or InfiniBand partitioning. For example, an interface can be connected to a configuration network, removed from that network, and moved to a network used at runtime. Communications from the controller to resources may be interrupted or restricted, which may result in the controller being unable to physically respond to any data sent from resource 1310.For example, when resource 1310 is added and booted, inband management 270 can be switched off or filtered, either physically or using a software-defined network. Inband management can be configured to send data to a separate resource dedicated to log management.
[0211] In-band management can be switched on and off using out-of-band management or software-defined networking. When in-band management is disconnected, there is no need to run the daemon, and in-band management can be re-enabled using the keyboard function.
[0212] Furthermore, resource 1310 may optionally not have an in-band management connection, and the resource may be managed through out-of-band management.
[0213] Out-of-band management can be used to operate various aspects of the system in ways that include, but are not limited to, other functions of out-of-band management that enable communication between controller 200 and resource 1310, whether or not they are exposed to an operating system running on resource 1310, for example, by connecting a keyboard, virtual keyboard, disk-mounted console, or virtual disk, changing BIOS settings, changing boot parameters and other aspects of the system, executing existing scripts that may be present in a bootable image or installation CD, or by exposing them to an operating system running on resource 1310. For example, controller 200 can send commands using such tools via out-of-band management 260. Controller 200 can further use image recognition to assist in controlling resource 1310. Thus, by using an out-of-band management connection, the system can prevent or avoid undesirable operation of resources connected to the system via the out-of-band management connection. The out-of-band management connection can also be configured as a one-way communication system during system operation or at selected times during system operation.
[0214] Furthermore, the out-of-band management connection 260 can also be selectively controlled by the controller 200 in the same manner as the in-band management connection, if desired by the implementer.
[0215] The controller 200 can automatically turn resources on and off according to global system rules and update the state of the IT system for reasons determined by the IT system user, such as turning resources off to save power, turning resources on to improve application performance, or any other reason the IT system user can think of. The controller can further turn configuration SAN, in-band and out-of-band management connections on and off, or may designate such connections as one-way write connections at all times or for various security purposes during system operation (for example, disabling the in-band management connection 270 or configuration SAN 280 while resource 1310 is connected to the external network 1380 or internal network 390). One-way in-band management may further be used, for example, to monitor system health, monitoring logs and information that may be displayed to the operating system.
[0216] Resource 1310 may be further coupled to one or more internal networks 390, such as application networks, from which services, application users, and / or clients can communicate with each other. Such application networks 390 may further be connected to or capable of connecting to an external network 1380. According to exemplary embodiments of this specification, including but not limited to Figures 2A to 12B, inband management may be disconnected from or capable of disconnecting from the resource or application network 390, or may provide one-way writes from the controller to provide additional security when the resource or application network is connected to an external network, or when the resource is connected to an application network that is not connected to an external network.
[0217] The IT system 100 in Figure 13A may be configured similarly to the IT system 100 shown in Figure 3B. Image 350 may be loaded directly or indirectly (through another resource or database) from template 230 to resource 1310 for booting computing resources and / or loading applications. Image 350 may include boot files 340 for resource types and hardware. Boot files 340 may include kernels 341 corresponding to the resources, applications, or services to be deployed. Boot files 340 may further include an initrd or similar file system used to assist the boot process. The boot system 340 may include multiple kernels or initrds configured for different hardware and resource types. In addition, image 350 may include a file system 351. File system 351 may include a base image 352 and its corresponding file system, as well as a service image 353 and its corresponding file system, and a volatile image 354 and its corresponding file system. The file systems and data loaded may vary depending on the resource type and the application or service being executed. Base image 352 may include a base operating system file system. The base operating system may be read-only. The base image 352 may further contain basic tools for an operating system independent of the one currently running. The base image 352 may also contain a base directory and operating system tools. The service file system 353 may contain configuration files and specifications for resources, applications, or services. The volatile file system 354 may contain information or data specific to its deployment, such as binary applications, specific addresses, and other information, which may or may not be configured as variables, including but not limited to passwords, session keys, and private keys.File systems can be integrated as a single file system using technologies such as overlayFS, and some read-only file systems and some read / write file systems can reduce the amount of duplicate data used by applications.
[0218] Figure 13B shows multiple resources 1310, each comprising one or more hypervisors 1311 that host or have one or more virtual machines. Controller 200a is coupled to each resource 1310, each comprising bare-metal resources. As illustrated and described with reference to Figure 13B, each resource 1310 is coupled to Controller 200a. According to exemplary embodiments herein, the in-band management connection 270, configuration SAN 280 and / or out-of-band management connection 260 can be configured as described with reference to Figure 13A. One or more virtual machines or hypervisors may be compromised or become compromised. In conventional systems, other virtual machines on other hypervisors may subsequently become compromised. For example, this may result from hypervisor exploitation running within a virtual machine. For example, a pivot may move from the compromised hypervisor to Controller 200a, and there from the compromised Controller 200a to another hypervisor coupled to Controller 200a. For example, a pivot could occur between a compromised hypervisor and a target hypervisor using the network connected to both. In the configuration of controller 200a and resource 1310 with in-band management 270, configuration SAN 280, or out-of-band management 260 shown in Figure 13B, some or all of the in-band (or configuration SAN) and / or out-of-band connectivity can be selectively controlled to disable certain links between controller 200a and resource 1310, which can prevent a compromised virtual machine from originating from one hypervisor and being used to pivot to another resource.
[0219] As explained with respect to Figures 1 to 12 above, the in-band management connection 270 and the out-of-band management connection 260 can be further configured in the same manner as described with respect to Figures 13A and 13B.
[0220] Figure 13C shows an exemplary process flow for adding or managing physical resources, such as bare metal nodes, to system 100. Resources 1310, as shown in Figures 13A and 13B herein or in relation to Figures 1 to 12, may be connected to the controller of system 100 via out-of-band management connections 260 and in-band management connections 270 and / or SANs.
[0221] Following the instance of resource connection, in step 1370, the external network and / or application network are disabled. As described above, any of the following techniques can be used for this disabling. For example, before setting up the system, adding resources, testing the system, updating the system, or performing other tasks or commands, components of system 100 (or only those vulnerable to attack) are disabled, disconnected, or filtered from the external network or application network using an inband management connection or configuration SAN, as described with respect to Figures 13A and 13B.
[0222] Following step 1370, in step 1371, the in-band management connection and / or configuration SAN is enabled. Thus, the combination of steps 1370 and 1371 isolates the resources from the external network and / or application network while the in-band management and / or SAN connection is established. Commands can then be executed on the resources under the control of the controller 200 via the in-band management connection (see step 1372). Setup and configuration steps, including but not limited to those described herein with respect to Figures 1 to 13B, can be performed in step 1372 using the in-band management and / or configuration SAN. Alternatively or additionally, inband management and / or configuration SANs may be used in step 1372 to perform other tasks, including but not limited to system operation, updates, or management (including, but not limited to, any change management or system updates), testing, updates, data transfer, collection of performance and health information (including, but not limited to, errors, CPU usage, network usage, file system information and storage usage), and log collection, as well as other commands that may be used to manage System 100 as described in Figures 1 to 13B of this Spec.
[0223] As described herein with respect to Figures 13A and 13B, after adding resources, setting up the system, and executing such tasks or commands, the in-band management connection 270 and / or configuration SAN 280 between the resource and the controller or other components of the system can be disabled in one or more directions in step 1373. Such disabling can be done by disconnecting, filtering, etc., as described above. After step 1373, in step 1374, the connection to the external network and / or application network can be restored. For example, the controller can notify the networking resource that resource 1310 can connect to the application network or the internet. The same steps can be performed when testing or updating the system, that is, the in-band management connection to the external network and / or application network can be disconnected or filtered and then the in-band management connection to the resource can be enabled or connected (one or both directions). Thus, while the resource is connected to the external network and / or application network, steps 1373 and 1374 operate simultaneously to isolate the resource from the controller through the in-band management connection and / or configuration SAN.
[0224] Out-of-band management can be used to manage a system or resource, set up, configure, boot, or add a system or resource. When used in any embodiment of this specification, out-of-band management can send commands to the machine using a virtual keyboard to change settings before booting, and can also send commands to the operating system by typing into the virtual keyboard. If the machine is not logged in, out-of-band management can use the virtual keyboard to enter a username and password, use image recognition to verify the logon, verify the entered commands, and confirm whether they were executed. If the physical resource only has a graphical console, a virtual mouse can also be used, and image recognition will allow out-of-band management to make changes.
[0225] Figure 13D is another exemplary process flow for adding or managing physical resources, such as bare metal nodes, to system 100. In step 1380, the resources shown in Figures 13A and 13B or Figures 1 to 12 herein may be connected to the system or resources via out-of-band management 260. Disks can be virtually connected by providing access to a disk image (e.g., an ISO image) through out-of-band management facilitated by the controller (see step 1381). The resources or system may then be booted from the disk image (step 1382), and then files are copied from the disk image to a bootable disk (see step 1383). This can also be used to boot a system in which resources are set up in this manner using out-of-band management. This can also be used to configure and / or boot multiple resources (including but not limited to networking resources) that can be joined together, regardless of whether multiple resources further comprise a controller or constitute a system. Thus, it may be possible for the controller to connect a disk image to a resource as if the virtual disk were connected to the resource using virtual disks. Out-of-band management can also be used to send files to a resource. The data can be copied from the virtual disk to the local disk in step 1383. The disk image may contain files that the resource can copy and use during operations. Files can be copied or used through either a scheduled program or instructions from out-of-band management. Through out-of-band management, the controller can log on to the resource using a virtual keyboard and enter commands to copy files from the virtual disk to the controller's own disk or other storage accessible to the resource. In step 1384, the system or resource is configured to boot by setting the BIOS, EFI, or boot order settings, which will cause it to boot from a bootable disk.The boot configuration may use the operating system's EFI manager, such as efibootmgr, which can be run directly from out-of-band management or included in an installer script (for example, a script using efibootmgr is automatically executed when the resource boots). In addition, boot options or other BIOS changes may be configured through an out-of-band management tool such as Supermicro Boot Manager, either by using boot order commands or by uploading a BIOS configuration (such as an XML BIOS configuration supported by Supermicro Update Manager). The BIOS can also be configured to set appropriate BIOS settings, including the boot order, using image recognition from the keyboard and console. The installer can be run against a loaded configured image. The configuration can be tested by viewing the screen and using image recognition. After configuration, the resource can be enabled (e.g., power on, boot, connect to the application network, or a combination thereof) (step 1385).
[0226] Figure 13E is another exemplary process flow for adding or managing physical resources, such as bare metal nodes, to system 100, in this case using PXE, Flexboot, or a similar network boot. In step 1390, the resources 1310 shown in Figures 13A and 13B herein, or shown with respect to Figures 1-12, may be connected to the controller of system 100 via (1) in-band management connections 270 and / or SAN and (2) out-of-band management connections 260. External network and / or application network connections may then be disabled in step 1391 (similar to those described above in relation to step 1370) (e.g., filtered or disconnected, entirely or partially, physically or virtually, using SDN). For example, before setting up the system, adding resources, testing the system, updating the system, or performing any other task or command, components of system 100 (or only those vulnerable to attack) are disabled, disconnected, or filtered from the external network or application network using an inband management connection or SAN, as described with respect to Figures 13A and 13B.
[0227] In step 1392, the type of resource is determined. For example, information about the resource can be collected from its MAC address by using an out-of-band management tool, or by temporarily booting an operating system that has tools that can be used to identify resource information by attaching a disk image (e.g., an ISO image) to the resource as if a disk were attached to the resource. Next, in step 1393, the resource is identified as being configured or pre-configured for PXE or flexboot, etc. Next, in step 1394, the resource is powered on and a PXE, Flexboot, or similar boot is performed (or the resource is temporarily booted and powered on again). Next, in step 1395, the resource boots from an in-band management connection or SAN. In step 1396, data is copied to a disk accessible by the resource in a manner similar to that described with reference to step 1383 in Figure 13D. In step 1397, the resource is configured to boot from a disk(s) in a manner similar to that described above with reference to step 1384 in Figure 13D. If a resource is identified as being pre-configured for PXE, flexboot, etc., the file can be copied in any of the steps from 1393 to 1396. If in-band management is enabled, it can be disabled in step 1398, and the application network or external network can be reconnected or enabled in step 1399.
[0228] Furthermore, it should be understood that resources can be remotely enabled (e.g., powered on) and verified that the resources are booted using technologies other than OOBM. For example, the system could prompt the user to press a power button and manually inform the controller that the system has booted (or use a console connection to the keyboard / controller). Additionally, the system could ping the controller through IBM (e.g., via ssh, telnet, or another method over the network) once the system has booted and the controller has logged on and instructed to reboot. For example, the controller could send a reboot command using ssh. In either case, if PXE is used and OOBM is not available, the system should have a way to instruct the resource to be remotely powered on or to be manually powered on by the user.
[0229] Deployment of controllers and / or environments In exemplary embodiments, the controller may be deployed within the system from the outgoing controller 200 (such an outgoing controller 200 may be referred to as the “main controller”). Thus, the main controller may set up a system or environment that is an isolated or separable IT system or environment.
[0230] The environments described herein refer to a collection of resources within a computer system that can interoperate with each other. A computer system may, but is not required, contain multiple environments. The resources of an environment may comprise one or more instances, applications, or sub-applications that run within that environment. Furthermore, an environment may comprise one or more environments or sub-environments. An environment may or may not include a controller, and an environment may run one or more applications. Such resources of an environment may include, for example, networking resources, computing resources, storage resources, and / or application networks used to run a particular environment that contains applications within that environment. Thus, it should be understood that an environment may provide the functionality of one or more applications. In some examples, the environments described herein may be physically or virtually isolated from other environments, or may be separable. Furthermore, in other examples, an environment may have network connectivity to other environments, and such connectivity may be disabled or enabled as needed.
[0231] In addition, the main controller can set up, deploy, and / or manage one or more additional controllers in various environments or as separate systems. Such additional controllers may be independent of the main controller or remain in an independent state. Such additional controllers, even if independent of the main controller or pseudo-independent, may receive instructions from or transmit information to the main controller (or a separate monitor or environment via a monitoring application) at various points in their operation. Environments may be configured for security purposes (e.g., by making environments separable from each other and / or the main controller) and / or for various management purposes. Environments may be connected to an external network, while other related environments may or may not be connected to an external network.
[0232] The main controller can manage environments or applications, regardless of whether they are separate systems and whether they have controllers or subcontrollers. The main controller can also manage shared storage for global configuration files or other data. The main controller can also parse global system rules (e.g., system rule 210) or a subset of the main controller's rules for different controllers, according to their functionality. Each new controller (sometimes called a “subcontroller”) can receive new configuration rules, which may be a subset of the main controller’s configuration rules. The subset of global configuration rules deployed to a controller may depend on or correspond to the type of IT system being set up. The main controller can set up or deploy new controllers or separate IT systems, which are then permanently separated from the main controller, for example, shipping or delivery or other purposes. Global configuration rules (or a subset thereof) can define a framework for setting up applications or subapplications in various environments and how they can interact with each other. Such applications or environments can run on subcontrollers with a subset of global configuration rules deployed by the main controller. In some examples, such applications or environments can be managed by the main controller. However, in other examples, such applications or environments are not managed by the main controller. When a new controller is created from the main controller to manage an application or environment, dependency checks for applications can be performed across multiple applications to facilitate control by the new controller.
[0233] Therefore, in exemplary embodiments, the system may comprise a main controller configured to deploy another controller, or an IT system comprising such other controllers. Such an implementation system may be configured to be completely isolated from the main controller. Once isolated, such a system may be configured to operate as a standalone system, or may be controlled or monitored by another controller (or an environment with applications), such as the main controller, at various discrete or continuous time points during operation.
[0234] Figure 14A shows an exemplary system in which the main controller 1401 has controllers 1401a and 1401b deployed on different systems 1400a and 1400b, respectively (where 1400a and 1400b may be called subsystems; however, it should be understood that subsystems 1400a and 1400b may also function as environments). The main controller 1401 can be configured in a similar manner to the controller 200 described above. Therefore, it may include controller logic 205, global system rules 210, system states 220, and templates 230.
[0235] Systems 1400a and 1400b each comprise controllers 1401a and 1401b, respectively, coupled to resources 1420a and 1420b. The main controller 1401 can be coupled to one or more other controllers, such as controller 1401a of subsystem 1400a and controller 1401b of subsystem 1400b. The global rules 210 of the main controller 1400 may include rules that can manage and control other controllers. Using such global rules 210 together with controller logic 205, system states 220, and templates 230, the main controller 1401 can set up, provision, and deploy subsystems 1400a and 1400b through controllers 1401a and 1401b in a manner similar to that described herein with reference to Figures 1 to 13E.
[0236] For example, the main controller 1401 can load global rule 210 (or a subset thereof) into subsystems 1400a and 1400b as rules 1410a and 1410b, respectively, in a manner that the global rule 210 (or a subset thereof) instructs the operation of controllers 1401a and 1401b and their subsystems 1400a and 1400b. Each controller 1401a and 1401b may have rules 1410a and 1410b, which may be the same or different subsets of global rule 210. For example, which subset of global rule 210 is provisioned to a given subsystem may depend on the type of subsystem being deployed. Furthermore, controller 1401 can load or send data that will be loaded into system resources 1420a and 1420b or controllers 1401a and 1401b.
[0237] The main controller 1401 may be connected to other controllers 1401a, 1401b via in-band management connections 270(or more) and / or out-of-band management connections 260(or more) or SAN connections 280, which can be enabled or disabled at various stages of deployment or management in the manner described herein, for example, with reference to the resource deployment and management described in Figures 13A to 13E. By using selective enabling and disabling of the in-band management connections 270 or out-of-band management connections 260, subsystems 1400a, 1400b can be deployed in a manner in which subsystems 1400a, 1400b have no knowledge whatsoever (or knowledge is limited, controlled, or prohibited) of the main system 100 or controller 1401 or each other at various points in time.
[0238] In exemplary embodiments, the main controller 1401 can operate a centralized IT system having local controllers 1401a, 1401b deployed and configured by the main controller 1401, so that the main controller 1401 can deploy and / or run multiple IT systems. Such IT systems may or may not be independent of each other. The main controller 1401 can set up monitoring as a separate application isolated or air-gapped from the IT systems it creates. A separate console for monitoring may have connections between the main controller and the local controller(s), and / or connections between environments that can be selectively enabled or disabled. The controller 1401 can deploy isolated systems for various applications, including, but not limited to, business, manufacturing systems with data storage, data centers, and various other functional nodes, each having a different controller in case of failure or jeopardy. Such isolation may be complete or permanent, or it may be pseudo-isolated, for example, temporary, time-dependent or task-dependent, communication direction-dependent or other parameter-dependent. For example, the main controller 1401 may be configured to provide instructions to a system that is limited to or not limited to certain predefined situations, while subsystems may have limited or no ability to communicate with the main controller. Therefore, such subsystems may not be able to compromise the main controller 1401. The main controller 1401 and subcontrollers 1401a, 1401b may be isolated from each other as described herein (in the specific examples described below), for example, by disabling in-band management 270, by restricting communication to one-way writes and / or out-of-band management 260. For example, in the event of a breach, one or more controllers may disable the in-band management connection 270 to one or more other controllers to prevent the breach or the spread of access.The system section can be turned off or isolated.
[0239] Subsystems 1400a and 1400b can further share resources with or connect to other environments or systems through in-band management 270 or out-of-band management 260.
[0240] Figures 14B and 14C are exemplary flows illustrating possible steps for provisioning a controller with a main controller.
[0241] In Figure 14B, in step 1460, the main controller provisions or sets up resources such as resource 1420a or 1420b. In step 1461, the main controller provisions or sets up subcontrollers. The main controller can use the techniques described above to set up resources in the system and perform steps 1460 and 1461. Furthermore, although Figure 14B shows that step 1460 is performed before step 1461, it should be understood that this is not mandatory. The main controller 1401 can use its system rules 210 to determine which resources are needed and place those resources on the system or network. In step 1461, the main controller can set up or deploy subcontrollers by loading system rules 210 into the system to set up subcontrollers (or by providing instructions to subcontrollers on how to set up and obtain their own system rules). These instructions may include, but are not limited to, instructions for configuring resources, configuring applications, creating global system rules for the IT systems to be run by the subcontrollers, instructions for reconnecting to the main controller to collect new or modified rules, and instructions for disconnecting from the application network to create space for a new production environment. After the resources have been deployed, in step 1463, the main controller may allocate the resources to the subcontrollers via updates to system rules 210 and / or system states 220.
[0242] Figure 14C shows an alternative process flow for deployment. In the example in Figure 14C, the main controller deploys the subcontroller in step 1470 (which can be done as described in relation to step 1461). Next, in step 1475, the subcontroller deploys resources using techniques such as those shown in Figures 3C and 7B.
[0243] Figure 15A shows an exemplary system in which the main controller 1501 of system 100 generates environments 1502, 1503, and 1504. Environment 1502 contains resource 1522, environment 1503 contains resource 1523, and environment 1504 contains resource 1524. Furthermore, environments 1502, 1503, and 1504 can share access to a pool of shared resources 1525. Such shared resources may include, but are not limited to, shared datasets, APIs, or running applications that need to communicate with each other.
[0244] In the example in Figure 15A, each environment 1502, 1503, and 1504 shares a main controller 1501. The global system rules 210 of the main controller 1501 may include rules for deploying and managing environments. Resources 1522, 1523, and / or 1524 may be required by each environment 1501, 1502, and 1503 to manage one or more applications. Configuration rules for such applications may be implemented by the main controller (or, if present, local controllers within the environment) to define how each such environment operates and how it interacts with other applications and environments. The main controller 1401 can use the global rules 210 together with controller logic 205, system states 220, and templates 230 to set up, provision, and deploy environments in a manner similar to the resource and system deployment described with reference to Figures 1 to 14C of this specification. If the environment includes a local controller, the main controller 1501 can load global rules 210 (or a subset thereof) onto the local controller or associated storage so that global rules (or a subset thereof) define the operation of that environment.
[0245] Controller 1501 can deploy and configure resources 1522, 1523, 1524 and / or shared resource 1525 for environments 1502, 1503, and 1504, respectively, using configuration rules having system rule 210. Controller 1501 can further monitor the environments or configure resources 1522, 1523, 1524 (or shared resource 1525) to enable monitoring of environments 1502, 1503, and 1504, respectively. Such monitoring may be via a connection to a separate monitoring console that can be enabled or disabled, or through the main controller. The main controller 1501 can connect to one or more of the environments 1502, 1503, and 1504 through in-band management connections 270(or more) and / or out-of-band management connections 260(or more) or SAN connections 280, which can be enabled or disabled at various stages of deployment or management, as described herein with reference to the resource deployment and management in Figures 13A-13E and 14A. Using the enabling and disabling of the in-band management connections 270 or out-of-band management connections 260 or SAN connections 280, the environments 1502, 1503, and 1504 can be deployed in a manner in which they have no knowledge of the main system 100 or controller 1501 at various points in time, or their knowledge is limited or controlled, or their knowledge of their connections to each other is limited or controlled.
[0246] An environment may be coupled with an external network 1580 that connects to an external environment, or it may comprise one or more resources that interact with other resources. An environment may be physical or non-physical. In this context, "non-physical" means that environments share the same physical host(s) but are virtually isolated from one another. Environments and systems may be deployed on identical hardware, similar but different hardware, or non-identical hardware. In some examples, environments 1502, 1503, and 1504 may be valid copies of each other, while in other examples, environments 1502, 1503, and 1504 may provide different functionalities from each other. For example, the resources of an environment may be servers.
[0247] By placing systems and resources in separate environments or subsystems according to the techniques described herein, it may be possible to isolate applications for security and / or performance purposes. Isolating environments can also mitigate the impact of compromised resources. For example, one environment may contain sensitive data and be configured to have limited exposure to the internet, while another environment may host internet-facing applications.
[0248] Figure 15B shows an exemplary process flow in which the controller shown in Figure 15A sets up an environment. In such an example, the system may be tasked with creating and setting up a new environment. This can be triggered by a user request or by system rules that are executed when processing a particular task or set of tasks. Figures 17A to 18B, described below, show examples of specific change management tasks or sets of tasks in which the system creates a new environment. However, there can be many situations in which the controller may create and set up a new environment.
[0249] Therefore, referring to Figure 15B, when setting up a new environment, the controller selects environment rules (step 1500.1). Using global system rules 210 and templates 230 according to the environment rules, the controller finds resources for the environment (step 1500.2). The rules may have a hierarchy of preferred resource selections that the controller goes through until it finds the resources required for the environment. In step 1500.3, the controller assigns the resources found in step 1500.2 to the environment, for example, using the techniques described in Figure 3C or Figure 7B. Next, the controller configures the system's networking resources with respect to the new environment to ensure compatible and efficient connectivity between the new environment and other system components (step 1500.4). The system state is updated in step 1500.5 as each resource is enabled and each template is processed. Next, the controller sets up and enables the integration and interoperability of the environment's resources and powers on any applications to deploy the new environment (step 1500.6). The system state will be updated again in step 1500.7 when the environment becomes available.
[0250] Figure 15C shows an exemplary process flow in which the controller shown in Figure 15A sets up multiple environments. When setting up multiple environments, the environments can be set up in parallel using the technique described in Figure 15B for each environment. However, it should be understood that the environments can be set up in a sequential order or sequentially, as described in Figure 15C. Referring to Figure 15C, in step 1500.10, the controller sets up and deploys the first new environment (which can be done as described with respect to step 1500.1 in Figure 15B). Different environment rules may exist for different types of environments and different interoperability methods. In step 1500.11, the controller selects environment rules for the next environment. In step 1500.12, the controller finds resources according to the priority which can be defined by system rule 210. In step 1500.13, the controller assigns the resources found in step 1500.12 to the next environment. Environments may or may not share resources. In step 1500.14, the controller uses system rule 210 to configure the system's networking resources for the next environment and between environments with dependencies. The system state is updated in step 1500.15 when each resource is enabled, the template has been processed, and the networking resources are configured, including their environment dependencies. Next, the controller sets up and enables the integration and interoperability of resources for the next environment and between environments, and powers on any applications to deploy the new environment (step 1500.16). The system state is updated in step 1500.17 when the next environment becomes available.
[0251] One-way communication to support monitoring FIG. 16A shows an exemplary embodiment in which the first controller 1601 operates as a main controller for setting up one or more controllers such as 1601a, 1601b, and / or 1601c. The main controller 1601 may use the techniques described above with respect to controllers such as controller 200 / 1401 / 1501 to generate a plurality of cloud hosts, systems, and / or applications as environments 1602, 1603, 1604 that may or may not be operationally dependent on each other. As shown in FIG. 16A, an IT system, environment, cloud, and / or any combination thereof may be generated as environments 1602, 1603, 1604. Environment 1602 includes a second controller 1601a, environment 1603 includes a third controller 1601b, and environment 1604 includes a fourth controller 1601c. Environments 1602, 1603, 1604 may each also include one or more resources 1642, 1643, 1644, respectively. The resources may include one or more applications 1642, 1643, 1644 that may be executed thereon. These applications can connect to the assigned resources regardless of whether they are shared. These or other applications can be executed on the Internet or on one or more shared resources within a pool 1660 that may also include a shared application or application network. The applications can provide services to one or more of a user or an environment or cloud. Environments 1602, 1603, 1604 can share resources or databases and / or include or use resources within a pool 1660 specifically assigned to a particular environment. Various components of the system including the main controller 1601 and / or one or more environments may be connectable to an external network 1615 such as an application network or the Internet.
[0252] There may be connections between any resource, environment or controller and another resource, environment, controller or external connection that can be selectively enabled and / or disabled in the manner described with respect to FIGS. 13A - 13E herein. For example, any resource, controller, environment or external connection can be disabled or disconnected from the controller 1601, environment 1602, environment 1603 and / or environment 1604, resource, or application via an in - band management connection 270, an out - of - band management connection 270 or a SAN connection 280, or by physically disconnecting. As an example, the in - band management connection 270 between the controller 1601 and any of the environments 1602, 1603, 1604 can be disabled to protect the controller 1601. As another example, such in - band management connection(s) 270 can be selectively disabled or enabled during the operation of the environments 1602, 1603, 1604. In addition to the security purposes described with respect to FIGS. 13A - 13E herein, by disabling or disconnecting the main controller 1601 from the environments 1602, 1603, 1604, the main controller 1601 may be able to spin up the environments 1602, 1603, 1604 as a cloud that can then be separated from the main controller 1601 or other clouds or environments. In this sense, the controller 1601 is configured to generate multiple clouds, hosts or systems.
[0253] Using the disabling or disconnecting elements described herein, a user may be permitted limited access to an environment through the main controller 1601 for a particular purpose. For example, a developer may be provided access to a development environment. As another example, an application administrator may be limited to a particular application or application network. As another example, logs may be displayed through the main controller 1601, and data may be collected without exposing it to danger by the environment or controller generated by the main controller.
[0254] After the main controller 1601 sets up the environment 1602, the environment 1602 is disconnected from the main controller 1601, and at this time, the environment 1602 can operate independently of the main controller 1601 and / or be selectively monitored and maintained by the main controller 1601 or by other applications associated with the environment 1602 or run by the environment.
[0255] Environments such as Environment 1602 may be coupled to a user interface or console 1640 that allows the purchaser or user to access Environment 1602. Environment 1602 can host a user console as an application. Environment 1602 can be accessed remotely by a user. Each of Environments 1602, 1603, and 1604 can be accessed by a common or separate user interface or console.
[0256] Figure 16B shows an exemplary system in which environments 1602, 1603, and 1604 may be configured to write to another environment 1641, where the logs can be viewed, for example, using a console (which may be any console that can connect to environment 1641 directly or indirectly). In this way, environment 1641 can function as a log server to which one or more of environments 1602, 1603, and 1604 write events. The main controller 1601 can then access the log server 1641 to monitor events on environments 1602, 1603, and 1604 without maintaining a direct connection to such environments 1602, 1603, and 1604, as will be described later. Environment 1641 may also be selectively disconnected from the main controller 1601 and configured to be read-only from the other environments 1602, 1603, and 1604.
[0257] As shown in Figure 16C, the main controller 1601 can be configured to monitor some or all of the environments 1602, 1603, and 1604 even if the main controller 1601 is disconnected from any of those environments 1602, 1603, and 1604. Figure 16C shows that the in-band management connection 270 between the main controller 1601 and environments 1602, 1603, and 1604 is disconnected, which can help protect the main controller 1601 if environments 1602, 1603, and 1604 are compromised. As shown in Figure 16C, the out-of-band connection 260 can be maintained between the main controller 1601 and environments such as 1602 even if the in-band connection 270 between the main controller 1601 and environment 1602 is disconnected. Furthermore, environment 1641 may have a connection to the main controller 1601 that can be selectively enabled or disabled. Main controller 1601 can set up monitoring as a separate application isolated from environments 1602, 1603, and 1604, or as an air-gapped environment 1641. Main controller 1601 can use one-way communication for monitoring. For example, logs may be provided via one-way communication from environments 1602, 1603, and 1604 to environment 1641. Through such one-way writing and via the connection between environment 1641 and main controller 1601, main controller 1601 can collect data via environment 1641 and monitor environments 1602, 1603, and 1604 even if there is no in-band connection 270 between main controller 1601 and environments 1602, 1603, and 1604, thereby reducing the risk of environments 1602, 1603, and 1604 endangering main controller 1601. Access may be filtered or controlled, and / or access may be independent of the internet.For example, as shown in Figure 16D, if an in-band connection 270 is connected between the main controller 1601 and the environment 1602, the main controller 1601 can control the network switch 1650 to disconnect the environment 1602 from an external network 1615 such as the internet. Disconnecting the environment 1602 from the external network 1615 when it is connected to the main controller 1601 by the in-band connection 270 can improve the security of the main controller 1601.
[0258] Therefore, it should be understood that the exemplary embodiments in Figures 16B to 16D demonstrate how the main controller can securely monitor environments 1602, 1603, and 1604 while minimizing exposure to these environments. Thus, the main controller 1601 can disconnect itself from environments 1602, 1603, and 1604 (or at least disconnect itself from the in-band link) while continuing to maintain a mechanism for monitoring those environments via a log server in environment 1641, which allows environments 1602, 1603, and 1604 to have one-way write access. Therefore, if the main controller 1601 discovers during the process of reviewing the logs of environment 1641 that environment 1602 may be compromised by malware, the main controller 1601 can use SDN tools to isolate environment 1602 so that only out-of-band connections 260 exist (see, for example, Figure 16C). Furthermore, controller 1601 can send notifications about potential problems to the administrator of environment 1602. The controller can also isolate the compromised environment 1602 by selectively disabling any connections (e.g., in-band management connection 270) between the compromised environment and any of the other environments 1603, 1604. In another example, the main controller 1601 may discover through logs that resources in environment 1603 are running excessively hot. This allows the main controller to intervene and migrate the application or service from environment 1603 to another environment (whether an existing environment or a newly created environment).
[0259] The controller 1601 can also set up one or more similar systems according to the requirements of the purchaser or user. As shown in Figure 16E, the purchase application 1650 may be provided, for example, on a console or otherwise, thereby enabling the purchaser to purchase or request a cloud, host, system environment or application to be set up for the purchaser. The purchase application 1650 can instruct the controller 1601 to set up environment 1602. Environment 1602 may include a controller 1601a that deploys or builds an IT system, for example, by allocating or assigning resources to environment 1602.
[0260] Figure 16F shows user interfaces 1632, 1633, and 1634 that may be used in environments where environments 1602, 1603, and 1604 each operate as clouds, and may or may not have a controller. User interfaces 1632, 1633, and 1634 (corresponding to environments 1602, 1603, and 1604, respectively) can each be connected through a main controller 1601 that manages the connection between the user interface and the environment. Alternatively or additionally, interface 1640a (which can take the form of a console) may be directly coupled to environment 1602, interface 1640b (which can take the form of a console) may be directly coupled to environment 1603, and interface 1640c (which can take the form of a console) may be directly coupled to environment 1604. Regardless of whether the connection to the main controller 1601 is isolated, disconnected, or disabled, the user can use one or more of the interfaces to use the environment or cloud.
[0261] System cloning and backup for change management support Some of environments 1602, 1603, and 1604 may be clones of typical setup software used by developers. These environments can also be clones of the current working environment as a means of expansion, for example, by cloning an environment in a different data center located in a different location to reduce location-related latency.
[0262] Therefore, it should be understood that a main controller that sets up systems and resources in separate environments or subsystems may enable cloning or backup of parts of the IT system. This can be used in testing and change management, as described herein. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, and / or other modifications.
[0263] According to exemplary embodiments, the IT systems or controllers described herein may be configured to clone one or more environments. The new or cloned environments may have the same resources as the original environments, or they may not. For example, it may be desirable or necessary for the new or nearly cloned environments to use a completely different combination of physical and / or virtual resources. It may be desirable to clone environments to different locations or time zones where optimization of usage can be managed. It may be desirable to clone environments to virtual environments. When cloning environments, the global system rules 210 and global templates 230 of the controller or main controller may contain information on how to configure and / or run various types of hardware. Configuration rules within system rules 210 can direct the allocation and use of resources so that resources and applications are more optimal, taking into account specific available resources.
[0264] The main controller structure provides the functionality to set up systems and resources in separate environments or subsystems, provides a structure for cloned environments, provides a structure for creating development environments, and / or provides a structure for deploying a standardized set of applications and / or resources. Such applications or resources may include, but are not limited to, those usable for application development and / or execution, or for backing up or restoring parts of IT systems and other disaster recovery applications (e.g., a system including a LAMP (apache, mysql, php) stack, a web frontend and a server running react / redux, as well as resources running node.js and a mongo database and other standardized “stacks”). In some cases, the main controller may deploy an environment that is a clone of another environment, deriving configuration rules from a subset of the configuration rules used to create the original environment.
[0265] According to exemplary embodiments, change management for a system or a subset of a system can be performed by cloning one or more environments and the configuration rules or subsets of configuration rules for such environments. Changes may be required, for example, to make changes to code, configuration rules, security patches, templates, hardware, components and dependent applications, and other modifications.
[0266] According to exemplary embodiments, such changes to the system can be automated to avoid errors in direct manual input of the changes. The changes can be tested by a user in a development environment before being automatically implemented in a live system. According to exemplary embodiments, a live production environment can be cloned by using a controller to automatically power on, provision and / or configure an environment configured using the same configuration rules as the production environment. The cloned environment can be run and operated (while the backup environment can preferably be left as is for emergency use in case changes need to be rolled back). This can be done using a controller to create, configure and / or provision a new system or environment as described with reference to Figures 1 to 16F above, using system rules 210, template 230 and / or system state 220. The new environment can be used as a development environment to test changes that will later be implemented in the production environment. The controller can generate the infrastructure of such an environment in the development environment from a software-defined structure.
[0267] In this specification, the term "production environment" refers to an environment used to run a system, as opposed to a development environment, which is an environment used solely for development and testing.
[0268] When the production environment is cloned, the infrastructure or cloned development environment is configured and generated by the controller in accordance with the global system rule 210, just like the production environment. Changes to the development environment can be made to the code, templates 230 (either changes related to modifying existing templates or creating new templates), security, and / or application or infrastructure configuration. Once the new changes implemented in the development environment are prepared as needed through development and / or testing, the system automatically applies the changes to the development environment, which then becomes operational or deployed as the production environment. The new system rule 210 is then uploaded to either the environment controller and / or the main controller that applies the system rule changes to the specific environment. The system state 220 is updated within the controller and can implement the added or modified templates 230. Thus, complete system knowledge of the infrastructure, along with the ability to recreate it, can be maintained by the development environment and / or the main controller. As used herein, complete system knowledge may include, but is not limited to, the state of resources, resource availability, and system knowledge of system configuration. Complete system knowledge can be collected by the controller from system rules 210, system states 220, and / or by querying resources using in-band management connections 270(or more), out-of-band management connections 260(or more), and / or SAN connections 280(or more). Resources can be queried, among other things, to determine resource, network or application utilization, configuration status, or availability.
[0269] The cloned infrastructure or environment may, but is not required, be software defined via system rule 210. The cloned infrastructure or environment may, generally, comprise a front-end or user interface and one or more allocated resources, which may or may not include compute, network, storage, and / or application networking resources. This environment may, or may not, be configured as a front-end, middleware, and database. A service or development environment can be booted using system rule 210 of the production environment. Infrastructure or environments allocated for use by a controller may be defined in software specifically for cloning. Therefore, the environment is deployable by system rule 210 and cloneable by similar means. A cloned or development environment may be automatically set up by the local or main controller using system rule 210 before or when changes are desired.
[0270] Production environment data may be written to read-only data storage until the development environment is separated from the production environment, and will then be used by the development environment for development and testing processes.
[0271] Users or clients can make and test changes in the development environment while the production environment is online. Data in data storage may be modified during development, and these changes are tested in the development environment. On volatile or writable systems, hot synchronization with production data can also be used after the development environment is set up or deployed. Desired changes to the system, application, and / or environment can be made to and tested in the development environment. Then, the desired changes are made to the scripts of system rule 210, creating a new version for the environment or the entire system and the main controller.
[0272] According to another exemplary embodiment, the newly developed environment may then be automatically implemented as a new production environment while the previous production environment is maintained or fully functional, thus allowing a return to the previous state of the production environment without the loss of large amounts of data. Next, the development environment is booted with new configuration rules in system rule 210, the database is synchronized with the production database, and it is switched to a writable database. The original production database can then be switched to a read-only database. If it is desirable to return to the previous production environment, the previous production environment is maintained as a copy of the previous production environment for the required period.
[0273] The environment can be configured as a single server or instance that includes physical and / or virtual hosts, a network, and other resources. In another exemplary embodiment, the environment can consist of multiple servers that include physical and / or virtual hosts, a network, and other resources. For example, there may be multiple servers forming a load-balanced internet-facing application, and these servers may be connected to multiple API / middleware applications (which may be hosted on one or more servers). The environment's database may consist of one or more databases through which the API communicates queries within the environment. The environment can be constructed from system rule 210 in static or volatile form. The environment or instance may be virtual, physical, or a combination of both.
[0274] The application configuration rules or system configuration rules within system rules 210 can specify various compute backends (e.g., bare metal, AMD epyc server, Intel Haswell on qemu / kvm) and include rules on how to run the application or service on the new compute backend. Therefore, for example, the application can be virtualized if there are situations where the availability of resources for testing is reduced.
[0275] Using the examples described herein, according to these, the test environment can be deployed on virtual resources where the original environment uses physical resources. Using the controllers described herein with reference to FIGS. 1-18B, and further as described herein, a system or environment can be cloned from a physical environment to an environment that may or may not include virtual resources, either in whole or in part.
[0276] FIG. 17A shows an exemplary embodiment in which system 100 includes controller 1701 and one or more environments, such as 1702, 1703, 1704. System 100 may be a static system, i.e., a system in which active user data does not constantly change the state of the system or frequently manipulate data, such as a system that only hosts static web pages. The system can be coupled to user (or application) interface 110.
[0277] Controller 1701 can be configured in a manner similar to controllers 200 / 1401 / 1501 / 1601 described herein, and similarly can include global system rules 210, controller logic 205, templates 230, and system state elements 220. Controller 1701 can be coupled to one or more other controllers or environments in a manner as described with reference to FIGS. 14A-16F herein. The global rules 210 of controller 1701 can include rules that can manage and control other controllers and / or environments. Using such global rules 210, controller logic 205, system state 220, and templates 230, a system or environment can be set up, provisioned, and deployed through controller 1701 in a manner similar to that described with reference to FIGS. 1-16F herein. Each environment can be configured using a subset of the global system rules 210 that define the behavior of the environment included with respect to other environments.
[0278] The global system rules 210 may also include change management rules 1711. Change management rules 1711 comprises a set of rules and / or instructions that may be used when changes to system 100, global system rules 210, and / or controller logic 205 are desired. Change management rules 1711 can be configured to allow a user or developer to develop changes, test the changes in a test environment, and then implement the changes by automatically translating them into a new set of configuration rules within system rules 210. Change management rules 1711 may be a subset of global system rules 210 (as shown in Figure 17A) or may be separate from global system rules 210. Change management rules can utilize a subset of global system rules 210. For example, global system rules 210 may include a subset of environment creation rules configured to create a new environment. Change management rules 1711 can be configured to set up and use a system or environment configured and set up by controller 1701 to copy and clone some or all aspects of system 100. Change management rule 1711 can be configured to allow testing of proposed new changes to the system before implementation by using a clone of the system for testing and implementation.
[0279] A clone 1705, as shown in Figure 17A, may have a specific environment or some rules, logic, applications, and / or resources of system 100. A clone 1705 may have hardware similar to or different from system 100, and may or may not use virtual resources. A clone 1705 can be set up as an application. A clone 1705 can be set up and configured using configuration rules within system rules 210 of system 100 or controller 1701. A clone 1705 may or may not have a controller. A clone 1705 may have allocated network, computing resources, application networks, and / or data storage resources, as described in more detail above. Such resources can be allocated using change management rules 1711 controlled by controller 1701. A clone 1705 can be coupled to a user interface that allows users to make changes to clone 1705. The user interface may be the same as or different from the user interface 110 of system 100. Clone 1705 can be used for the entire system 100, or for one or more environments and / or parts of system 100 such as controllers. Clone 1705 may or may not be a complete copy of system 100. Clone 1705 can be connected to system 100 via in-band management connections 270, out-of-band management connections 260, and / or SAN connections 280, which can be enabled and / or disabled and / or converted to unidirectional read and / or write connections. Thus, when the clone environment 1705 is isolated from the production environment during testing, or until the clone environment 1705 is ready to come online as a new production environment, the connections to the data within the clone environment 1705 can be changed to make the clone data read-only. For example, if clone 1705 has a data connection to environment 1702, this data connection can be made read-only for isolation.
[0280] Any backup 1706 may or may not be used for the entire system, or for one or more environments and / or parts of the system such as controllers. Backup 1706 may include networks, compute, application networks and / or data storage resources, as described in more detail above. Backup 1706 may or may not include controllers. Backup 1706 may be a complete copy of system 100. Backup 1706 can be set up as an application or using hardware similar to or different from system 100. Backup 1706 can be coupled to system 100 via in-band managed connections 270, out-of-band managed connections 260 and / or SAN connections 280, which can be enabled and / or disabled and / or converted to unidirectional read and / or write connections.
[0281] Figure 17B shows an exemplary process flow for using the clone and backup system of Figure 17A in system change management. In step 1785, a user or management application initiates changes to the system. Such changes may include, but are not limited to, changes to code, configuration rules, security patches, templates, hardware changes, additions / removals of components and / or dependent applications, and other modifications. In step 1786, controller 1701 sets up the environment in the manner described with respect to Figures 14A to 16F to become the clone environment 1705 (where the clone environment may have its own new controller or use the same controller for the original environment).
[0282] In step 1787, the controller 1701 may use global rules 210, including change management rule 1711, to clone all or part of one or more environments of the system (e.g., “production environment”) to a clone environment 1705 (e.g., where clone environment 1705 can function as a “development environment”). Thus, the controller 1701 identifies and allocates resources, uses system rules 210 to set up and allocate clone resources, and copies any of the data, configuration, code, executable files, and other information necessary to start the application from the environment to the clone. In step 1788, the controller 1701 optionally backs up the system by using configuration rules within system rules 210 to set up another environment (with or without a controller) that functions as a backup 1706, copying the template 230, controller logic 205, and global rules 210.
[0283] After clone 1705 is created from the production environment, clone 1705 can be used as a development environment where changes can be made to the clone's code, configuration rules, security patches, templates, and other modifications. In step 1789, changes to the development environment can be tested before implementation. During testing, clone 1706 can be isolated from the production environment (system 100) or other components of the system. This can be done by having controller 1701 selectively disable one or more connections between system 100 and clone 1706 (for example, by disabling the in-band management connection 270 and / or the application network connection). In step 1790, it is determined whether the modified development environment is ready. In step 1709, if it is determined that the development environment is not yet ready (this is usually a determination made by the developer), the process flow returns to step 1789 for further changes to clone environment 1705. In step 1790, if it is determined that the development environment is ready, in step 1791, the development environment can be switched to the production environment. In other words, the controller can change the development environment 1705 to the new production environment and maintain the previous production environment until the migration between the development and new production environments is complete and in good condition.
[0284] Figure 18A shows another exemplary embodiment of system 100 that may be set up and used in system change management. In the example of Figure 18A, system 100 comprises a controller 1801 and one or more environments 1802, 1803, 1804, 1805. The system is shown together with a clone environment 1807 and a backup system 1808.
[0285] Controller 1801 may be configured in a manner similar to that of controllers 200 / 1401 / 1501 / 1601 / 1701 described herein, and may include elements of global system rules 210, controller logic 205, templates 230, and system states 220. Controller 1801 may be coupled to one or more other controllers or environments in a manner similar to that described with reference to Figures 14A to 16F herein. The global rules 210 of controller 1801 may include rules that can manage and control other controllers and / or environments. Using such global rules 210, controller logic 205, system states 220, and templates 230, a system or environment may be set up, provisioned, and deployed through controller 1801 in a manner similar to that described with reference to Figures 1 to 17B herein. Each environment may be configured using a subset of global rules 210 that defines the behavior of the environment, including behavior with respect to other environments.
[0286] The global rules 210 may also include change management rules 1811. Change management rules 1811 may include a set of rules and / or instructions that can be used when changes to the system, global rules, and / or logic are desired. The change management rules can be configured to allow a user or developer to develop changes, test the changes in a test environment, and then implement the changes by automatically translating them into a new set of configuration rules within the system rules 210. Change management rules 1711 may be a subset of the global system rules 210 (as shown in Figure 18A) or may be separate from the global system rules 210. Change management rules 1711 can use a subset of the global system rules 210. For example, the global system rules 210 may include a subset of environment creation rules configured to create a new environment. Change management rules 1811 may be configured to set up and use a system or environment set up and deployed by the controller 1801 to copy and clone some or all aspects of system 100. Change management rule 1811 can be configured to allow testing of proposed new changes to the system before implementation by using a clone of the system for testing and implementation.
[0287] As shown in Figure 18A, the clone environment 1807 may comprise a controller 1807a having rules, controller logic, templates, system state data, and allocated resources 1820, which can be assigned to one or more environments and set up according to the global system rules 210 and change management rules 1811 of the controller 1801. The backup system 1808 may further comprise a controller 1808a having rules, controller logic, templates, system state data, and allocated resources 1821, which can be assigned to one or more environments and set up according to the global system rules 210 and change management rules 1811 of the controller 1801. The system can be coupled to a user (or application) interface 110 or another user interface.
[0288] A clone environment 1807 may have some rules, logic, templates, system states, applications, and / or resources of a particular environment or system. Clone 1807 may have hardware similar to or different from system 100, and may or may not use virtual resources. Clone 1807 can be set up as an application. Clone 1807 can be set up and configured for the environment using configuration rules within system rules 210 of system 100 or controller 1801. Clone 1807 may or may not have a controller, and may share a controller with the production environment. Clone 1807 may have allocated network, computing resources, application networks, and / or data storage resources, as described in more detail above. Such resources can be allocated using change management rules 1811 controlled by controller 1801. Clone 1807 can be coupled to a user interface that allows users to make changes to clone 1807. The user interface may be the same as or different from the user interface 110 of system 100.
[0289] The clone 1807 can be used for the entire system or for a part of the system such as one or more environments and / or controllers. In an exemplary embodiment, the clone 1807 may include a hot standby data resource 1820a coupled to a data resource 1820 of environment 1802. The hot standby data resource 1820a can be used during the setup and change testing of the clone 1807. The hot standby data resource 1820a can be selectively disconnected or isolated from the storage resource 1820 during change management, for example, as described herein with respect to Figure 18B. The clone 1807 may or may not be a complete copy of system 100. The clone 1807 can be coupled to system 100 via in-band management connections 270, out-of-band management connections 260 and / or SAN connections 280, which can be fully selectively enabled and / or disabled and / or converted to unidirectional read and / or write connections. Therefore, when clone environment 1807 is isolated from the production environment during testing, or until the clone environment is ready to come online as a new production environment, the connection to the volatile data within clone environment 1807 can be changed to make the clone data read-only.
[0290] When switching from an old production environment to a new one, controller 1801 can instruct frontends, load balancers, or other applications or resources to point to the new production environment. Therefore, users, application resources, and / or other connections may be redirected when the change occurs. This can be done by methods including, but not limited to, changing the listing of IP / IPOIB addresses, InfiniBand GUIDs, DNS servers, InfiniBand partition / OpenSM configurations, or changing software-defined network (SDN) configurations by sending instructions to networking resources. Frontends, load balancers, or other applications and / or resources may refer to systems, environments, and / or other applications, including, but not limited to, databases, middleware, and / or other backends. Such load balancers can be used for change management when switching from an old production environment to a new one.
[0291] Clone 1807 and backup 1808 may be set up and used to manage the nature of changes to the system. Such changes may include, but are not limited to, code, configuration rules, security patches, template changes, hardware changes, additions / removals of components and / or dependent applications, and other changes. Backup 1808 may be used for the entire system or for one or more environments and / or parts of the system such as controller 1801. Backup 1808 may have networks, computing resources, application networks and / or data storage resources, as described in more detail above. Backup 1808 may or may not have a controller. Backup 1808 may be a complete copy of system 100. Backup 1808 may have the data necessary to reconstruct the system / environment / application from the configuration rules contained in the backup, and may include all application data. Backup 1808 may be set up as an application or using hardware similar to or different from system 100. Backup 1808 can be connected to system 100 via in-band management connection 270, out-of-band management connection 260, and / or SAN connection 280, which can be selectively enabled and / or disabled and / or converted to one-way read and / or write connections.
[0292] Figure 18B is an exemplary process flow illustrating the use of the system in Figure 18A in change management, particularly when the system in Figure 18A contains volatile data or when the database is writable. Such a database may be part of the storage resources used by the environment within the system. In step 1870, the system is deployed using global system rules (including the production environment).
[0293] Next, in step 1871, the production environment is cloned to create a read-only environment in which the cloned environment is prohibited from writing to the system, using global system rules 210, including change management rule 1811, and resource allocation by the main controller 1801 or a controller in the cloned environment. The cloned environment can then be used as a development environment.
[0294] In step 1872, hot standby 1820a is enabled and assigned to clone environment 1807 to store any volatile data that has been modified within system 100. The clone data is updated, and the new version in the development environment can be tested with the updated data. Hot sync data can be turned off at any time. For example, hot sync data can be turned off when writing from an older or production environment to the development environment is being tested.
[0295] Next, in step 1873, the user can use clone environment 1807 as a development environment to make changes. Then, in step 1874, the changes to the development environment are tested. In step 1875, it is determined whether the modified development environment is ready (this determination is usually made by the developer). If it is determined in step 1875 that the changes are not ready, the process flow can return to step 1873 so that the user can go back and make other changes to the development environment. If it is determined in step 1875 that the changes are ready to take effect, the process flow proceeds to step 1876, where configuration rules within the system or controller are updated for the specific environment and used to deploy the new, updated environment.
[0296] In step 1877, the development environment (or new environment) may then be redeployed with modifications to a desired final configuration with desired resource and hardware allocations before going live. In the next step, 1878, the write functionality of the original production environment is disabled, and the original production environment becomes read-only. While the original production environment is read-only, as part of 1878, any new data from the original production environment (or possibly the new production environment as well) may be cached and identified as migration data. For example, the data may be cached on a database server or other appropriate location (e.g., a shared environment). Next, the development environment (or new environment) and the old production environment are switched in step 1879, with the development environment (or new environment) becoming the production environment.
[0297] Following this switchover, the new production environment is made writable in step 1880. In step 1881, if the new production environment is deemed to be functioning as determined by the developers, any data loss during the switchover process (such data is cached in step 1878) may be repaid by writing the data to the new environment and verifying it in step 1884. After such verification, the changes are complete (step 1885).
[0298] If step 1881 determines that the new production environment is not functioning correctly (for example, an issue is identified that requires reverting the system to the old system), then in step 1882, the environment is reverted, and the old production environment becomes the production environment again. As part of step 1882, the configuration rules for the target environment in controller 1801 are reverted to the previous version that was used for the production environment currently being reverted.
[0299] In step 1883, database changes may be determined, for example, using cached data, and the data is restored to an older production environment with the old configuration rules. To support step 1883, the database may maintain a log of changes made to the database so that step 1883 can determine which changes may need to be disabled. Data may be cached as described above using a backup database that tracks and records the time of cached data, allowing the clock to be turned back to determine which changes were made. Snapshots and logs may be used for this purpose.
[0300] If you want to start again after restoring the cached data in step 1883, the process can return to step 1871.
[0301] Examples of change management systems described herein may be used, for example, when updating, adding, or removing hardware or software; when patching software; when a system failure is detected; when migrating a host during or after a hardware failure; for dynamic resource migration; for changes to configuration rules or templates; and / or when making any other system-related changes. The controller 1801 or system 100 may be configured to detect failures, and upon detection, the system may automatically enforce change management rules or existing configuration rules on other hardware available to the controller. Examples of usable failure detection methods include, but are not limited to, pinging hosts, querying applications, and running various tests or test suites. Change management configuration rules described herein may be enforced upon detection of a failure. Such rules may trigger the automatic generation of a backup environment, or the automatic migration of data or resources being performed by the controller, upon detection of a failure. The selection of backup resources may be based on resource parameters, which may include, but are not limited to, usage information, speed, configuration rules, and data capacity and usage.
[0302] Whenever a change occurs as described herein, the controller shall create a log of the change and what was actually done. For security or system updates, the controllers described herein may be configured to automatically turn on and off and update the IT system state according to configuration rules. The controller may turn off resources to conserve power. The controller may turn on or migrate resources to vary efficiency over time. During migration, backups or copies of the environment or system may be created according to configuration rules. In the event of a security breach, the controller may isolate and shut down the attacked area.
[0303] While the invention has been described in terms of exemplary embodiments, various modifications within the scope of the invention are permitted. Such modifications to the invention can be identified by reading the teachings herein.
[0304] Appendix A: Example of a Storage Connection Process This document outlines examples of processes and rules for sharing storage resources across multiple systems. It should be understood that this is only one example of a storage connectivity process, and other techniques can be used to connect computing resources to storage resources. Unless otherwise noted, these rules apply to all systems attempting to initiate a storage connectivity. Definitions in this Annex A Storage resources: Blocks, files, or file systems that can be shared via storage transport. Storage transport: A method for sharing storage resources locally or remotely. Examples include iSCSI / iSER, NVMeoF, NFS, and Samba file sharing. System: Anything that attempts to connect to a storage resource via a specified storage transport. A system may support any number of storage transports, and the system itself may decide which transport to use. Read-only: A read-only storage resource does not allow modification of the data it contains. This restriction is imposed by the storage daemon that handles the export of storage resources on the storage transport. For further assurance, some datastores may set the storage resources that back up their data to read-only (for example, setting LVM LVs to read-only). Read / Write (or Volatile): A read / write (volatile) storage resource is a storage resource whose content may be modified by the system to which it is connected. Rules: When a controller determines whether a system can connect to a given storage resource, there is a set of rules that it must follow. 1. Read / write storage resources shall be exported via only one storage transport. 2. Read / write storage resources shall be connected by only one system. 3. Read / write storage resources must not be connected in read-only mode. 4. Read-only storage resources may be exported using multiple storage transports. 5. Read-only storage resources may be connected to by multiple systems. 6. Read-only storage resources must not be connected in read-write mode. process If we consider the connection process as a function, the connection process takes two independent variables. 1. Storage Resource ID 2. List of supported storage transports (ordered by priority) First, determine whether the requested storage resource is read-only or read-write. For read / write operations, a read / write storage resource is limited to one connection, so it must be checked whether the storage resource is already connected. If it is already connected, verify that the system requesting the storage resource is the currently connected system (this may occur, for example, in the case of reconnection). Otherwise, an error is issued because multiple systems cannot connect to the same read / write storage resource. If the requesting system is a system connected to this storage resource, verify that one of the available storage transports matches the current export of this storage resource. If there is a match, provide the connection information to the requesting system. If there is no match, an error is issued because multiple storage transports cannot supply the read / write storage resource. For read-only storage resources and unconnected read-write storage resources, the system processes the list of supplied storage transports in order and attempts to export the storage resources using those transports. If the export fails, the system continues to attempt exports in the order of the list until an export is successful or until no storage transports remain. When no storage transports remain, the requesting system is notified that the storage resource could not be connected. If the export is successful, the connection information and the new (resource, transport) => (system) relationship are stored in the database. The requesting system is then informed of the storage transport connection information. System: Storage connectivity is currently handled by the controller and compute daemon during normal operation. However, in future iterations, the service may connect directly to the storage resource, bypassing the compute daemon. This may be a requirement for an example of a service's physical deployment, and it would make sense to use the same process for virtual machine deployments as well.
[0305] Appendix B: Example of connecting to OverlayFS The service uses OverlayFS to reuse objects from a common file system, reducing the service package size. The service in this example includes three or more storage resources. 1. Platform. This includes the base Linux filesystem and is accessed in read-only mode. 2. Services. This includes all software directly related to the operation of the services (NetThunder ServiceDaemon, OpenRC scripts, binaries, etc.). This storage resource is accessed in read-only mode. 3. Volatility. These storage resources contain all changes to the system and are managed by LVM from within the service (for physical, container, and virtual machine deployments). When running on a virtual machine, the service is booted directly from the kernel by Qemu using a custom Linux kernel along with an initramfs containing the logic to perform the following actions: 1. Assemble an LVM volume group (VG) from available read / write disks. *This VG contains one logical volume (LV) that holds all volatile storage data for the service. 2. Mount the platform, services, and LV. 3. Combine the three file systems using a union file system (in this case, OverlayFS). The same process can be used for physical deployment. One option is to remotely provide the kernel to a lightweight OS booted via PXE Boot or IPMI ISO Boot, and then use kexec on the new actual kernel. Alternatively, one can skip the lightweight OS and PXE boot directly into the kernel. Such systems may require additional logic in the kernel initramfs to connect to storage resources. The OverlayFS configuration can be as follows: / ―――――――――――\ |Volatile Layer (LV)(RW)| +----------------------------------------+ |Service Layer (RO)| +----------------------------------------+ Platform Layer (RO) \―――――――――――― / Due to some limitations of OverlayFS, it is possible to mark a special directory ' / data' as "out of tree". This directory becomes available to services when the ' / data' directory is created during the creation of a service package. This special directory is mounted via 'mount--rbind', allowing access to a subset of volatile layers that are not within OverlayFS. This is necessary for applications such as NFS (Network File System) that do not support shared directories that are part of OverlayFS. Kernel filesystem layout: / +--platform / +--bin / +--... / +--service / +--data / [optional] +--bin / +--... Conservative--volatile +--work / +--root / +--bin / +--data / [if present in / service / ] +--... +--new_root / +--... Create the / new_root directory and use it as the target for configuring OverlayFS. Once OverlayFS is configured, executing exec_root within / new_directory will allow the system to start successfully with all available resources.
Claims
1. A controller for use in a computer system including a physical host, configured to automatically manage the physical infrastructure of the computer system based on a plurality of system rules, the system state of the computer system, and a plurality of templates. A device equipped with the following features.
2. The apparatus according to claim 1, wherein the automated management includes an automated configuration of the physical infrastructure of the computer system based on the system rules, the system state, and the templates.
3. The apparatus according to claim 1 or 2, wherein the template includes a plurality of templates for use in a plurality of different types of physical infrastructure.
4. The apparatus according to claim 3, wherein the system rules control which template is used when managing a given type of physical infrastructure.
5. The apparatus according to any one of claims 1 to 4, wherein the controller is further configured to perform the automated management of the physical infrastructure in response to user requests.
6. The apparatus according to any one of claims 1 to 5, wherein the controller comprises a processor and memory, and the memory is configured to store the system rules, the system state, and the template.
7. The apparatus according to any one of claims 1 to 6, wherein the computer system comprises a plurality of physical hosts and a plurality of virtual hosts, and the controller is further configured to deploy applications interchangeably on the physical hosts and the virtual hosts.
8. The apparatus according to any one of claims 1 to 7, wherein the system rules include global system rules for the self-assembly of the computer system.
9. The apparatus according to claim 8, wherein the global system rule includes specifying a number of IT tasks to be completed with respect to the addition of resources to the computer system.
10. The apparatus according to claim 8 or 9, wherein the global system rules include an updatable list of hardware necessary to add resources to the computer system.
11. The apparatus according to any one of claims 8 to 10, wherein the global system rules include specifying an ordered list of operations and tasks to be completed with respect to the addition of resources to the computer system.
12. The apparatus according to any one of claims 1 to 11, wherein the system state tracks, maintains, changes, and updates the status of the computer system.
13. The apparatus according to claim 12, wherein the system state is configured to track resources available to the computer system.
14. The apparatus according to any one of claims 1 to 13, wherein the template includes a set of default information used to create, configure, and / or deploy at least one of (1) a resource, (2) an application loaded onto the resource, or (3) a service loaded onto the resource on the computer system.
15. The apparatus according to claim 14, wherein the template includes a bare metal template.
16. The apparatus according to claim 14 or 15, wherein the template includes a service template.
17. The apparatus according to any one of claims 14 to 16, wherein each of the plurality of templates includes a base image of a base operating system file system.
18. The apparatus according to any one of claims 1 to 17, wherein the controller is further configured to build the infrastructure of the computer system using the template in accordance with the system rules and update the system state accordingly.
19. The aforementioned controller, Read the system rules and create a list of tasks to be completed in order to achieve the desired state of the computer system. Based on the available resources of the computer system, the system issues instructions to satisfy the read system rules. Using the system state, find the available resources of the computer system and perform the tasks in the list. If it is determined that the resources required for the tasks in the above list are available, the tasks will be executed using the available resources. The apparatus according to claim 18, further configured as follows.
20. The apparatus according to any one of claims 1 to 19, wherein the controller is further configured to automatically add computing resources based on the system rules, system state, and template.
21. The apparatus according to claim 20, wherein the computing resources include bare-metal computing resources.
22. The apparatus according to claim 20, wherein the computing resources include virtual computing resources.
23. The apparatus according to any one of claims 20 to 22, wherein the computer system includes the pool of computing resources.
24. The aforementioned controller, Based on the aforementioned system rules, system state, and template, storage resources are automatically added to the computer system. The apparatus according to any one of claims 1 to 23, further configured as follows.
25. The apparatus according to claim 24, wherein the storage resources include bare metal storage resources.
26. The apparatus according to claim 24, wherein the storage resource includes a virtual storage resource.
27. The apparatus according to any one of claims 24 to 26, wherein the computer system includes the pool of storage resources.
28. The apparatus according to any one of claims 1 to 27, wherein the controller is further configured to automatically add networking resources to the computer system based on the system rules, system state, and template.
29. The apparatus according to claim 28, wherein the networking resource includes a bare-metal networking resource.
30. The apparatus according to claim 28, wherein the networking resource includes a virtual networking resource.
31. The apparatus according to any one of claims 28 to 30, wherein the computer system includes the pool of networking resources.
32. The apparatus according to any one of claims 1 to 31, wherein the controller is further configured to manage the physical host by (1) configuring the BIOS of the physical host, (2) configuring the boot options of the physical host, (3) directing the server to storage resources, and (4) booting the physical host via an interface.
33. The apparatus according to claim 32, wherein the interface includes at least one of an Intelligent Platform Management Interface (IPMI) and a Redfish interface.
34. The apparatus according to any one of claims 1 to 33, wherein the controller is further configured to make BIOS changes on the physical host using image recognition.
35. The device according to any one of claims 1 to 34, wherein the controller is configured to automatically add the physical infrastructure of the new resource to the computer system by: (1) recognizing that the new resource has been connected to the computer system; (2) determining information about the connected new resource; (3) adding the determined information to the state of the computer system; (4) selecting one of the templates based on the determined information; (5) loading an image derived from the selected template into the computer system, the image including a file system; and (6) instructing the new resource to boot using the file system of the loaded image.
36. The apparatus according to any one of claims 1 to 35, wherein the controller is configured to automatically manage the physical infrastructure of the computer system based on out-of-band management.
37. The apparatus according to any one of claims 1 to 36, wherein the controller is configured to automatically manage the physical infrastructure of the computer system based on in-band management.
38. The apparatus according to any one of claims 1 to 30, wherein the controller is further configured to dynamically manage the pool of resources of the computer system by adding information about additional resources of the computer system to the system state.
39. The apparatus according to any one of claims 1 to 38, wherein the controller is further configured to automatically deploy applications or services on the resources of the computer system based on the system rules, the system state, and the system template.
40. The apparatus according to claim 39, wherein the controller is further configured to instruct the resources of the computer system to boot an application image derived from one of the templates in order to deploy the application or service to be run by the resources of the computer system, using an out-of-band or in-band managed connection.
41. The apparatus according to claim 40, wherein the system rule specifies a boot order such that the resource is booted from an image derived from one of the templates, and then the application or service is booted from an image derived from another of the templates.
42. The apparatus according to claim 40 or 41, wherein the controller is further configured to connect the application or service to an application network.
43. The apparatus according to claim 42, wherein the controller is further configured to connect the application or service to the application network via an out-of-band management connection.
44. A computer system for adding physical resources to an information technology (IT) system, A management network configured for out-of-band management connectivity to an out-of-band management device for the aforementioned physical resources, A controller configured to (1) recognize the physical resource via the out-of-band management connection, (2) add information about the physical resource to the state of the computer system, (3) select a template based on the recognized physical resource, (4) load an image derived from the selected template, and (5) instruct the physical resource to boot from the loaded image, The computer system including the computer system.
45. The controller accesses multiple system rules, the system state of the computer system, and multiple templates, The controller automatically manages the physical infrastructure of the computer system, including the physical host, based on the accessed system rules, system state, and templates. Methods that include...
46. The method according to claim 45, wherein the automatically managed step includes automatically configuring the physical infrastructure of the computer system based on the system rules, the system state, and the template.
47. The method according to claim 45 or 46, wherein the template includes multiple templates for use in multiple different types of physical infrastructure.
48. Based on the selected rules, select which of the templates will be used when managing a given type of physical infrastructure. The method according to claim 47, further comprising:
49. The method according to any one of claims 45 to 48, wherein the automated management step includes the controller performing the automated management of the physical infrastructure in response to a user request.
50. The method according to any one of claims 45 to 49, wherein the controller comprises a processor and memory, the memory storing the system rules, the system state, and the template.
51. The computer system includes a plurality of physical hosts and a plurality of virtual hosts, and the method is, The controller deploys applications interchangeably on the physical host and the virtual host. The method according to any one of claims 45 to 50, further comprising:
52. The method according to any one of claims 45 to 51, wherein the system rules include global system rules for the self-assembly of the computer system.
53. The method according to claim 52, wherein the global system rule includes specifying a number of IT tasks to be completed with respect to the addition of resources to the computer system.
54. The method according to claim 52 or 53, wherein the global system rule includes an updatable list of hardware necessary to add resources to the computer system.
55. The method according to any one of claims 52 to 54, wherein the global system rule includes specifying an ordered list of operations and tasks to be completed with respect to the addition of resources to the computer system.
56. The method according to any one of claims 45 to 55, wherein the system state tracks, maintains, changes, and updates the status of the computer system.
57. The method according to claim 56, wherein the system state tracks the resources available to the computer system.
58. The method according to any one of claims 45 to 57, wherein the template includes a set of default information used to create, configure, and / or deploy at least one of (1) a resource, (2) an application to be loaded onto the resource, or (3) a service to be loaded onto the resource on the computer system.
59. The method according to claim 58, wherein the template includes a bare metal template.
60. The method according to claim 58 or 59, wherein the template includes a service template.
61. The method according to any one of claims 58 to 60, wherein each of the plurality of templates includes a base image of a base operating system file system.
62. The controller constructs the infrastructure of the computer system using the template in accordance with the system rules and updates the system state accordingly. The method according to any one of claims 45 to 61, further comprising:
63. The controller reads the system rules and creates a list of tasks to be completed in order to achieve a desired state of the computer system. The controller issues instructions to satisfy the read system rules based on the available resources of the computer system, The controller uses the system state to find available resources in the computer system and perform the tasks in the list. If it is found that the resources required for the tasks in the above list are available, the tasks will be performed using the available resources. The method according to claim 62, further comprising:
64. The method according to any one of claims 45 to 63, wherein the automatically managed step includes the controller automatically adding computing resources based on the system rules, system state, and template.
65. The method according to claim 64, wherein the computing resources include bare-metal computing resources.
66. The method according to claim 64, wherein the computing resources include virtual computing resources.
67. The method according to any one of claims 64 to 66, wherein the computer system includes the pool of computing resources.
68. The method according to any one of claims 45 to 67, wherein the automatically managed step includes the controller automatically adding storage resources to the computer system based on the system rules, system state, and template.
69. The method according to claim 68, wherein the storage resource includes a bare metal storage resource.
70. The method according to claim 68, wherein the storage resource includes a virtual storage resource.
71. The method according to any one of claims 68 to 70, wherein the computer system includes the pool of storage resources.
72. The method according to any one of claims 45 to 71, wherein the automatically managed step includes the controller automatically adding networking resources to the computer system based on the system rules, system state, and templates.
73. The method according to claim 72, wherein the networking resource includes a bare-metal networking resource.
74. The method according to claim 72, wherein the networking resource includes a virtual networking resource.
75. The method according to any one of claims 72 to 74, wherein the computer system includes the pool of networking resources.
76. The method according to any one of claims 45 to 75, wherein the automatically managed step includes the controller managing the physical host by (1) configuring the BIOS of the physical host, (2) configuring the boot options of the physical host, (3) directing the server to storage resources, and (4) booting the physical host, via an interface.
77. The method according to claim 76, wherein the interface includes at least one of an Intelligent Platform Management Interface (IPMI) and a Redfish interface.
78. The method according to any one of claims 45 to 77, wherein the automatically managed step includes the controller making BIOS changes on the physical host using image recognition.
79. The method according to any one of claims 45 to 78, wherein the automatically managing step includes, in response to the new resource being connected to the computer system, the controller: (1) recognizes that the new resource has been connected to the computer system; (2) determines information about the connected new resource; (3) adds the determined information to the state of the computer system; (4) selects one of the templates based on the determined information; (5) loads an image derived from the selected template into the computer system, the image including a file system; and (6) instructs the new resource to boot using the file system of the loaded image, thereby adding the physical infrastructure of the new resource to the computer system.
80. The method according to any one of claims 45 to 79, wherein the automatically managed step includes the controller automatically adding the physical infrastructure to the computer system based on out-of-band management.
81. The method according to any one of claims 45 to 80, wherein the automatically managed step includes the controller automatically adding physical infrastructure to the computer system based on inband management.
82. The controller adds information about the additional resources of the computer system to the system state to create a dynamically managed pool of the computer system's resources. The method according to any one of claims 45 to 81, further comprising:
83. The controller automatically deploys applications or services on the computer system's resources based on the system rules, the system state, and the system template. The method according to any one of claims 45 to 82, further comprising:
84. The controller instructs the computer system resources to boot an application image derived from one of the templates in order to deploy the application or service to be executed by the computer system resources, using an out-of-band or in-band managed connection. The method according to claim 83, further comprising:
85. The method according to claim 84, wherein the system rule specifies a boot order such that the resource is booted from an image derived from one of the templates, and then the application or service is booted from an image derived from another of the templates.
86. Connecting the aforementioned application or service to the application network The method according to claim 84 or 85, further comprising:
87. The method according to claim 86, wherein the connecting step includes connecting the application or service to the application network via an out-of-band management connection.
88. A method for adding physical resources to an information technology (IT) system, Connecting a physical resource with an out-of-band management device to the out-of-band management connection of the management network, A controller associated with the management network (1) recognizes the physical resource via the out-of-band management connection, (2) adds information about the physical resource to the state of the computer system, (3) selects a template based on the recognized physical resource, (4) loads an image derived from the selected template, and (5) instructs the physical resource to boot from the loaded image. The method, including the method described above.
89. A controller configured to automatically manage a computer system including a physical host, wherein the controller is configured to automatically manage the physical infrastructure of the computer system based on a plurality of system rules, the system state of the computer system, and a plurality of templates. A device equipped with the following features.
90. The apparatus according to claim 89, further comprising the elements described in any one of claims 2 to 44.
91. Managing a computer system using the controller described in claim 89 or 90. Methods that include...
92. A computer system comprising the controller described in any one of claims 1 to 44 and 89 to 90.
93. A plurality of instructions located on a non-temporary computer-readable storage medium, configured to cause the processor to perform the method described in any one of claims 45 to 88 and 91 when executed by the processor. Computer program products, including [this].
94. A method for enhancing the security of a computer system including a system controller and network-connectable resources, The system controller selectively controls at least one of the in-band management connection and the storage area network (SAN) connection between the resource and the system controller so that (1) if the resource is connected to the network, at least one of the in-band management connection and the SAN connection is disabled, and (2) if the resource is not connected to the network, at least one of the in-band management connection and the SAN connection is enabled. The method, including the method described above.
95. The method according to claim 94, wherein the system includes the in-band management connection between the resource and the system controller, and the selective control step includes controlling the in-band management connection between the resource and the system controller such that (1) the in-band management connection is disabled when the resource is connected to the network, and (2) the in-band management connection is enabled when the resource is not connected to the network.
96. The method according to claim 94, wherein the system includes the SAN connection between the resource and the system controller, and the selective control step includes controlling the SAN connection between the resource and the system controller such that (1) the SAN connection is disabled if the resource is connected to the network, and (2) the SAN connection is enabled if the resource is not connected to the network.
97. The method according to claim 94, wherein the system includes the in-band management connection and the SAN connection between the resource and the system controller, and the selective control step includes controlling the in-band management connection and the SAN connection between the resource and the system so that (1) the in-band management connection and the SAN connection are disabled when the resource is connected to the network, and (2) the in-band management connection and the SAN connection are enabled when the resource is not connected to the network.
98. The method according to any one of claims 94 to 97, wherein the network includes a network outside the computer system.
99. The method according to any one of claims 94 to 98, wherein the network includes an internal network of the computer system.
100. The system controller maintains an out-of-band management connection between the resource and the system controller. The method according to any one of claims 94 to 99, further comprising:
101. The system controller selectively controls the out-of-band management connection between the resource and the system controller so that (1) the out-of-band management connection is disabled when the resource is connected to the network, and (2) the out-of-band management connection is enabled when the resource is not connected to the network. The method according to claim 100, further comprising:
102. Adding the aforementioned resources to the aforementioned computer system, Performing the selective control step in response to the aforementioned addition, The method according to any one of claims 94 to 101, further comprising:
103. The method according to any one of claims 94 to 102, wherein the resource includes a bare metal cloud node.
104. The method according to any one of claims 94 to 103, wherein the resource includes a physical resource.
105. The method according to claim 104, wherein the physical resources include a virtual machine.
106. The method according to claim 104 or 105, wherein the physical resource includes a hypervisor.
107. The method according to any one of claims 94 to 106, wherein the disabled in-band management connection or SAN connection has limited functionality.
108. The method according to claim 107, wherein the restricted functionality includes the ability of the controller to write to the resource, but does not include the ability of the resource to write to the controller.
109. A method for enhancing the security of a computer system including a system controller and network-connectable resources, Connecting the resources to the system controller via an out-of-band management connection, To provide access to the virtual disk to the resource via the out-of-band management connection, Copying the disk image from the virtual disk to the resource via the out-of-band management connection, The method, including the method described above.
110. Configure the resources to boot from the copied disk image. The method according to claim 109, further comprising:
111. Boot the resources from the copied disk image. The method according to claim 110, further comprising:
112. Perform the steps of the method if the resource is not connected to the system controller via at least one of an in-band management connection and a storage area network (SAN) connection. The method according to any one of claims 109 to 111, further comprising:
113. A controller for use in a computer system including a physical host and resources, wherein the resources are network-connectable, and the controller is configured to automatically manage the physical infrastructure of the computer system by selectively controlling at least one of an in-band management connection and a storage area network (SAN) connection between the resources and the controller, including (1) if the resources are connected to the network, at least one of the in-band management connection and the SAN connection is disabled, and (2) if the resources are not connected to the network, at least one of the in-band management connection and the SAN connection is enabled. A device equipped with the following features.
114. The apparatus according to claim 113, wherein the controller is further configured to perform the method described in any one of claims 95 to 112.
115. A computer system configured to perform the method described in any one of claims 94 to 112.
116. A plurality of instructions located on a non-temporary computer-readable storage medium, wherein, when executed by a processor, the instructions are configured to cause the processor to perform the method described in any one of claims 94 to 112. Computer program products, including [this].
117. A controller for a computer system, configured to access a plurality of system rules, the system state of the computing system, and a plurality of templates, wherein a subset of the system rules includes configuration rules that allow other controllers to be deployed within the computing system, the controller Includes, The controller is further configured to automatically deploy other controllers within the computing system based on the configuration rules, and the other controllers control the environment within the computing system. system.
118. The system according to claim 117, wherein the controller is further configured to control connections within the system so as to isolate the controller from other controllers within the computing system.
119. The system according to claim 118, wherein the controller is further configured to control the connection by selectively enabling and disabling the in-band management connection between the controller and the other controllers.
120. As part of the automated deployment, the controller is further configured to load the file system image derived from the template onto the other controllers. The other controller is configured to boot itself based on the loaded file system image. The system according to any one of claims 117 to 119.
121. A controller for a computing system accesses multiple system rules, the system state of the computing system, and multiple templates, wherein a subset of the system rules includes configuration rules that allow other controllers to be deployed within the computing system. The controller automatically deploys other controllers within the computing system based on the configuration rules, and the other controllers control the environment within the computing system; Methods that include...
122. To separate the controller from the other controllers within the computing system. The method according to claim 121, further comprising:
123. The aforementioned separation step is, Selectively disable the in-band management connection between the aforementioned controller and the other controller. The method according to claim 122, including the method described in claim 122.
124. The aforementioned step of automatically deploying is: Loading the file system image derived from the template onto the other controller, Booting the other controller based on the loaded file system image, The method according to any one of claims 121 to 123, including the method described in any one of claims 121 to 123.
125. A controller for a computing system, configured to access a plurality of system rules, the system state of the computing system, and a plurality of templates, wherein a subset of the system rules includes configuration rules that allow a computing environment to be deployed within the computing system, the controller Includes, The controller is further configured to automatically deploy the computing environment within the computing system based on the configuration rules. system.
126. The system according to claim 125, wherein the controller is further configured to automatically deploy a plurality of computing environments within the computing system based on the configuration rules.
127. The system according to claim 126, wherein the controller is further configured to form a larger computing system by linking multiple computing environments together using software-defined networking.
128. The system according to claim 126 or 127, wherein the computing environment includes a production environment and a development environment.
129. The system according to claim 128, wherein the controller is further configured to clone the production environment onto the deployed computing environment to form the development environment.
130. The system according to claim 128 or 129, wherein the development environment is configured to support testing of changes made to it.
131. The system according to any one of claims 128 to 130, wherein the controller is further configured to switch the development environment to operate as a new production environment within the larger computer system.
132. The system according to claim 131, wherein the controller is further configured to remove the production environment from the larger computer system.
133. The system according to any one of claims 125 to 132, wherein the controller is further configured to control connections within the computing system so as to isolate the deployed computing environment from the controller within the computing system.
134. The system according to claim 133, wherein the controller is further configured to control the connection by selectively enabling and disabling the in-band management connection between the controller and the deployed computing environment.
135. A controller for a computing system accesses a plurality of system rules, the system state of the computing system, and a plurality of templates, wherein a subset of the system rules includes configuration rules that allow a computing environment to be deployed within the computing system. The controller automatically deploys the computing environment within the computing system based on the configuration rules, Methods that include...
136. The aforementioned step of automatically deploying is: The controller automatically deploys multiple computing environments within the computing system based on the configuration rules. The method according to claim 135, including the method described in claim 135.
137. The controller links multiple computing environments together using software-defined networking to form a larger computing system. The method according to claim 136, further comprising:
138. The method according to claim 136 or 137, wherein the computing environment includes a production environment and a development environment.
139. The method according to claim 138, wherein the step of automatically deploying includes the controller cloning the production environment onto the deployed computing environment to form the development environment.
140. Test the aforementioned development environment. The method according to claim 138 or 139, further comprising:
141. The controller switches the development environment to operate as a new production environment within the larger computer system. The method according to any one of claims 138 to 140, further comprising:
142. The controller removes the production environment from the larger computer system. The method according to claim 141, further comprising:
143. To separate the controller from the deployed computing environment within the computing system. The method according to any one of claims 135 to 142, further comprising:
144. The aforementioned separation step is, Selectively disable the in-band management connection between the controller and the deployed computing environment. The method according to claim 143, including the method described in claim 143.
145. Controller and The first computing environment and A second computing environment, A system including, The controller is configured to deploy the first and second computing environments, The first computing environment is configured to write data to the second computing environment, The first computing environment is restricted from reading data from the second computing environment. The controller is configured to monitor data written to the second computing environment by the first computing environment. The aforementioned system.
146. The aforementioned second computing environment is deployed as a log server, The first computing environment is further configured to write log data relating to the first computing environment to the second computing environment. The system according to claim 145.
147. The system according to claim 145 or 146, wherein the controller is configured to determine whether or not to take action with respect to the first computing environment in response to the monitoring operation.
148. The computing system further includes a third computing environment, the first computing environment being controllably permitted by the controller to communicate with the third computing environment. The controller is further configured to separate the third computing environment from the first computing environment in response to the determination of the action. The system according to claim 147.
149. The system according to claim 148, wherein the monitoring operation includes the determination by the controller that a security risk exists with respect to the first computing environment based on the written data.
150. The system according to any one of claims 145 to 149, wherein the controller is further configured to controllably separate the first computing environment from the controller.
151. The first computing environment is the system according to any one of claims 145 to 150, which is capable of accessing an external computer network.
152. The aforementioned controller, In response to the monitoring operation, it is determined whether or not the first computing environment should be isolated from the external computer network. In response to the determination of the aforementioned action, the first computing network is isolated from the external computer network. The system according to claim 151, further configured as follows.
153. The system according to any one of claims 145 to 152, wherein the first computing environment is a production environment and the second computing environment is a development environment.
154. The first computing environment is further configured to write live data to the second computing environment. The second computing environment is testable using the live data. The system according to claim 153.
155. The controller deploys a first computing environment and a second computing environment within the computing system, The first computing environment is permitted to write data to the second computing environment, The first computing environment restricts the reading of data from the second computing environment, The first computing environment writes data to the second computing environment, The controller monitors the data written to the second computing environment by the first computing environment, Methods that include...
156. The aforementioned second computing environment is deployed as a log server, The writing step includes the first computing environment writing log data relating to the first computing environment to the second computing environment. The method according to claim 155.
157. The controller determines whether or not to take action with respect to the first computing environment in response to the monitoring. The method according to claim 155 or 156, further comprising:
158. The computing system further includes a third computing environment, the first computing environment being controllably permitted by the controller to communicate with the third computing environment, and the method is The controller, in response to the determination, separates the third computing environment from the first computing environment. The method according to claim 157, further comprising:
159. The method according to claim 158, wherein the monitoring step includes determining, based on the written data, that a security risk exists with respect to the first computing environment.
160. To separate the first computing environment from the controller in a controllable manner. The method according to any one of claims 155 to 159, further comprising:
161. The method according to any one of claims 155 to 160, wherein the first computing environment can access an external computer network.
162. The controller, in response to the monitoring, determines whether or not the first computing environment should be isolated from the external computer network. In response to the determination made by the controller, the controller isolates the first computing network from the external computer network. The method according to claim 161, further comprising:
163. The method according to any one of claims 155 to 162, wherein the first computing environment is a production environment and the second computing environment is a development environment.
164. The writing step includes the first computing environment writing live data to the second computing environment, and the method is Test the second computing environment using the aforementioned live data. The method according to claim 163, further comprising:
165. Controller and The first computing environment and A clone of the first computing environment, A system including, The first computing environment functions as the production environment, and the clone functions as the development environment. The controller is configured to maintain and apply a plurality of change management rules deployed to the production environment and the development environment, to (1) allow the modified development environment to be switched to become the new production environment within the system, and (2) allow the system to return to the production environment after the switch. The aforementioned system.
166. The development environment is configured to test changes to the first computing environment. The system according to claim 165, wherein, in response to a determination that the modified development environment is ready, the controller is further configured to (1) switch the modified development environment to become a new production environment, and (2) switch the production environment to cease being a production environment.
167. The controller is further configured to back up the first computing environment before the switchover, In response to a determination that the new production environment is not functioning properly, the controller is further configured to (1) restore the first computing environment to the production environment based on the backed-up computer system. The system according to claim 166.
168. Controller and The first computing environment and A clone of the first computing environment, A system including, The first computing environment functions as the production environment, and the clone functions as the development environment. The controller is configured to make the development environment read-only. The controller is further configured to activate a hot standby data resource in the development environment in order to store volatile data from the production environment. The aforementioned development environment is configured to receive changes made to it and test them within the development environment. The controller is further configured to (1) deploy the modified development environment as a new computing environment, (2) disable the write function of the production environment, (3) control the caching of new data from the production environment as migration data, (4) switch the new computing environment to become a new production environment, (5) switch the production environment to cease being a production environment, and (6) allow writing to the new production environment. The aforementioned system.
169. The system according to claim 168, wherein, in response to a determination that the new production environment is functioning properly, the new production environment is configured to match data between the new production environment and the cached migration data.
170. The system according to claim 168, wherein, in response to a determination that the new production environment is not functioning properly, the controller is further configured to (1) return to the first computing environment as the production environment based on the cached migration data, and (2) switch the new production environment so that it is no longer a production environment.
171. The process involves cloning a first computing environment within a system to form a clone of the first computing environment, wherein the first computing environment functions as a production environment and the clone functions as a development environment. The controller maintains and applies multiple change management rules deployed to the production environment and the development environment to (1) allow the modified development environment to be switched to become the new production environment within the system, and (2) allow the system to return to the production environment after the switch. Methods that include...
172. Testing the changes to the first computing environment in the aforementioned development environment, Based on the aforementioned tests, it is determined whether the modified development environment is ready, In response to the determination that the modified development environment is ready, (1) the modified development environment is switched to become the new production environment, and (2) the production environment is switched to cease being the production environment. The method according to claim 171, further comprising:
173. In response to the determination that the modified cloned computing environment is ready, the first computing environment is backed up before the switching step, To determine whether the aforementioned new production environment is functioning correctly, In response to a determination that the new production environment is not functioning properly, the system switches back to the first computing environment based on the backed-up computer system, The method according to claim 172, further comprising:
174. The process involves cloning a first computing environment within a computer system to form a cloned computing environment, wherein the cloned computing environment functions as a development environment and the first computing environment functions as a production environment. Making the aforementioned development environment read-only, In order to store volatile data from the aforementioned production environment, a hot standby data resource is activated in the development environment, Changing the aforementioned development environment, Test the changes in the aforementioned development environment, The modified development environment is deployed as a new computing environment, Disabling the write function in the aforementioned production environment, The new data from the aforementioned production environment is cached as migration data, Switching the aforementioned new computing environment to become the new production environment, Switching the aforementioned production environment to a non-production environment, Allowing writing to the aforementioned new production environment, Methods that include...
175. To determine whether the aforementioned new production environment is functioning properly, In response to the determination that the new production environment is functioning correctly, the data is matched between the new production environment and the cached migration data. The method according to claim 174, further comprising:
176. To determine whether the aforementioned new production environment is functioning properly, In response to the determination that the new production environment is not functioning properly, (1) the system returns to the first computing environment as the production environment based on the cached migration data, and (2) the new production environment is switched off from being a production environment. The method according to claim 174, further comprising:
177. Apparatus, system, method, and / or computer program product comprising any combination of the features disclosed herein.